judgment_version. It also records the engine version, in engine_version. For example, a payment’s answer for settles can read:
- An answer from a version that asks differently reads
staleafter you activate a new version. A query for current answers leaves it out. - Activation judges nothing again. A backfill judges the old answers again, and it shows its cost before it starts.
- A shadow report compares the new version with the active version on the same documents, before you switch.
- When the engine changes by itself, each new answer records the new engine version. The engine change never makes an answer stale.
What a new version does
To change a judgment’s question, criteria, context recipe, engine orhorizon, post the definition again under the same name. Brussle creates version n+1. The new version stays inactive until you activate it. See versions.
If the create’s cost estimate takes more than 20 seconds, the create answers 202 with a judgment_create job. When the job is done, its result holds the response of the create, for example the new version. The SDKs wait for the job for you. See when a change takes long.
Activate it with POST /namespaces/{ns}/judgments/settles/activate:
- A version that changes only thresholds becomes active at once. It judges as the active version does. Thus, the answers of the earlier version stay
fresh. - A version that asks differently starts a shadow report. The switch happens when you confirm it.
stale, with stale_reason older_version:
- Its numbers stay. The answer keeps the numbers that version 3 made, until Brussle judges the document again.
- A query for current answers leaves it out.
answers: "fresh_only"removes the row. A filter onfreshnessseesstale. - A subscription still matches it on its numbers. A subscription does not filter on freshness. Activation sends no subscription event for the version alone. Thresholds that come with the version act as any change to thresholds does.
Judge the old answers again with a backfill
A backfill judges each document in scope that has no successful answer from the active version for its current revision. SendPOST /namespaces/{ns}/judgments/settles/backfill without confirm first. Brussle returns the estimate and does nothing else:
"confirm": true to start the backfill as a job. See backfill for the fields and the job.
On a large namespace, Brussle can take more than 20 seconds to count the documents. Then the backfill request, with or without confirm, answers 202 with a backfill_start job. When that job is done, its result holds the estimate, or the job_id of the backfill that started. The SDKs wait for the job for you. See when a change takes long.
A backfill skips the answers from a version that changed only thresholds. These answers read fresh. Thus, after such a version, the estimate is zero.
Change a question safely with a shadow report
A shadow report shows what a new version of a judgment would change, before you switch to it. When version 4’s engine or definition is different from the active version, the activate call returns202 with a shadow job in awaiting_confirm. The job judges a random sample of 1,000 documents under version 4. If there are fewer documents, it judges all of them. While the job samples, report is null, and progress.documents_done counts the sampled documents. When the job is done, the report compares the two versions on the same documents:
current and candidate.
currentis the active version’s answers.candidateis the new version’s answers. Forboolandscore, each side has ameanand a 10-binhistogramofporscore. Forchoice, each side hasdist, the mean probability of each option. Thus, you can read the change option by option.threshold_flipscounts, for each named threshold, the documents that would go from false to true and from true to false. The new side uses the thresholds that will apply when the version is active. In the example, 31 more payments would wait for a person in the unmatched queue.recomputeestimates a backfill of each document under the new version: documents, tokens, judgments counted by size class times the weight of the new version’s engine, cost and duration.
POST /jobs/{id}/confirm(db.jobs.confirm(id)) switches to the new version. Confirm once the report is ready. A confirm before the report exists returnsconflictand changes nothing. The job showsrunning, thendoneabout a second later, when Brussle has committed the switch.POST /jobs/{id}/cancelkeeps the active version as it is.{"version": 4, "force": true}on activate skips the report and switches immediately.activate: trueon a create of a version does the same. A composite judgment is the exception. Its shadow job fits it on your labels, so the shadow job always runs.
shadow: true. They never make answers. They never count toward calibration. Brussle does not bill them.
Thresholds that you gave with the new version replace the current thresholds when the version becomes active.
In the dashboard, the Overview of the judgment’s page shows the report, with Confirm, Cancel and Force.
For a judgment on a template, activate the new version on the prefix path. You get one shadow report, with a sample from all the tenants.
Find answers by version
A query can filter on the version that made each answer:answers.<judgment>.judgment_versionis an integer. Compare it withEq,NotEq,Lt,Lte,Gt,Gte,InorNotIn, or test it withExists.answers.<judgment>.engine_versionis a string. Compare it withEq,NotEq,InorNotIn, or test it withExists. For an engine version that is not pinned, such as jevcurrentor gpt-6-lunacurrent, it holds the label of the epoch. The label does not name the engine, and two engines can have the same label. A judgment changes its engine only in a new version, so filter onjudgment_versiontoo when a judgment has used more than one engine.
settles answer is from a version before version 4:
POST /namespaces/{ns}/query. Use the same filter on engine_version to compare the answers of two epochs side by side. See find the answers of one version.
When the engine changes
Epochs
An epoch is a period in which the engine’s behaviour stays the same. Theengine_version of an evaluation records its epoch.
- An exact engine version is one epoch for its full life.
- An engine version that is not pinned, such as Jev
current, starts a new epoch each time its behaviour changes. The label is likecurrent+2026-09-24.1. See the engines.
engine_version. Each new answer records the new epoch.
You can see that the engine changed in these places:
- An event. Each judgment whose active version uses the changed engine version gets one
judgment.engine_changedevent for each new epoch. If the conformance run on the new epoch does worse than the last run that passed, each of these judgments also gets onejudgment.engine_regressedevent. A judgment that a template defines gets its events on the template. See the event types. - The answers. New answers have a new
engine_version. Filter on it to find them. - The calibration report.
lift.model_changedgives the earlier epoch, the current epoch and the day that the current epoch started. See read the lift. - The Calibration tab. It says that the engine changed while Brussle fits the new epoch.
current. When you move a judgment to another engine yourself, see change the engine yourself.
After the engine changes
Some minutes after an engine version that is not pinned, such as Jevcurrent, changes, Brussle fits each judgment with outcomes again for the new engine version. It uses a replay:
- Brussle takes the stored contexts of the answers that your most recent outcomes labelled, up to a fixed number.
- The same engine judges these contexts again.
- Each new answer, with its old outcome, is a sample of the new epoch.
- The Calibration tab. For the new epoch, while the replay runs, it says “Re-fitting after the engine changed on 2026-10-02: 400 of 812 outcomes replayed so far.” Then it says “Re-fitted on 812 replayed outcomes after the engine changed on 2026-10-02.”
- The report. The new epoch has
fit_source: "replay"and itsreplayprogress. Its ownoutcomesstay 0 until you post new outcomes. - Answers. Until Brussle makes the replay’s fit, usually in less than an hour, answers under the new epoch use the previous epoch’s calibration with
"from_previous_epoch": true. For a judgment without features, the previous epoch’s calibration lasts for 30 days at most. A judgment with features keeps the previous epoch’s fit until the new epoch has its own. - Evaluation history. Replays show in a document’s evaluation history with
replay_of, which names the evaluation that they replayed. Replays never change an answer. Brussle never bills them. - Outcomes. Continue to post outcomes as usual. When the new epoch has 100 of its own, with 20 of each kind, its fit uses only them, and
fit_sourcebecomesoutcomes.
status: "stopped"). Then Brussle fits the new epoch on what it replayed, if that meets the minimums. If not, Brussle fits the new epoch from its own outcomes as they arrive. The reason tells why the replay stopped:
cost_cap: your organization reached this month’s replay limit. If the epoch still has no fit next month, the replay runs again on the outcomes that it did not replay.superseded: the engine changed again. The newer epoch has its own replay.
Change the engine yourself
The engine is part of the definition. To pin an exact engine version, or to move to a different engine, post a new version with the newengine. Then activate it through its shadow report, as for any change to the question. Answers from the earlier version read stale with older_version until Brussle judges them again.
For example, to move settles from jev current to gpt-6-luna current, post version 4 with "engine": {"name": "gpt-6-luna", "version": "current"} and the same question and recipe. These things change with the engine:
- Cost. Each judgment counts its size class times the weight of the new engine: 1 on jev, 2.5 on gpt-6-luna. Each engine also counts tokens in its own way, so some documents can change size class. The shadow report’s
recomputeprices a backfill on the new engine. - What the engine reads. gpt-6-luna reads the images that the recipe selects. jev gets only their references. Thus, the context, its size class and the answers can all change.
- The answers. The shadow report compares the raw answers of the two engines on the same documents. It does not compare calibrated values.
- Calibration starts again. The new version has no fit. Its answers have no
calibrateduntil it has 100 outcomes of its own, with 20 of each kind. Brussle does not replay your earlier outcomes through the new engine. - Epochs. Each engine numbers its own epochs, so jev and gpt-6-luna can have the same label, such as
current+2026-10-07.1. To find the answers of the new engine, filter onjudgment_version, not onengine_versionalone. - Events. The judgment gets
judgment.engine_changedevents for the new engine’s epochs, and none for the old engine’s.