Skip to main content
A judgment’s definition has versions, as a database schema does. This page serves one problem: a change to the question or to the engine must not mix old answers with new answers without notice. The engine is the AI model that makes the answers. Each answer records the version of the judgment that made it, in judgment_version. It also records the engine version, in engine_version. For example, a payment’s answer for settles can read:
The example shows only some fields. See answers. What you can rely on:
  • An answer from a version that asks differently reads stale after you activate a new version. A query for current answers leaves it out.
  • Activation judges nothing again. A backfill judges the old answers again, and it shows its cost before it starts.
  • A shadow report compares the new version with the active version on the same documents, before you switch.
  • When the engine changes by itself, each new answer records the new engine version. The engine change never makes an answer stale.

What a new version does

To change a judgment’s question, criteria, context recipe, engine or horizon, post the definition again under the same name. Brussle creates version n+1. The new version stays inactive until you activate it. See versions. If the create’s cost estimate takes more than 20 seconds, the create answers 202 with a judgment_create job. When the job is done, its result holds the response of the create, for example the new version. The SDKs wait for the job for you. See when a change takes long. Activate it with POST /namespaces/{ns}/judgments/settles/activate:
  • A version that changes only thresholds becomes active at once. It judges as the active version does. Thus, the answers of the earlier version stay fresh.
  • A version that asks differently starts a shadow report. The switch happens when you confirm it.
After the switch, Brussle uses the new version for each new evaluation. Activation itself judges nothing again. Each answer from an earlier version that asks differently reads stale, with stale_reason older_version:
  • Its numbers stay. The answer keeps the numbers that version 3 made, until Brussle judges the document again.
  • A query for current answers leaves it out. answers: "fresh_only" removes the row. A filter on freshness sees stale.
  • A subscription still matches it on its numbers. A subscription does not filter on freshness. Activation sends no subscription event for the version alone. Thresholds that come with the version act as any change to thresholds does.
What judges the answer again depends on the judgment’s freshness policy:

Judge the old answers again with a backfill

A backfill judges each document in scope that has no successful answer from the active version for its current revision. Send POST /namespaces/{ns}/judgments/settles/backfill without confirm first. Brussle returns the estimate and does nothing else:
The numbers are illustrative, for a judgment on jev. On gpt-6-luna, the same backfill counts 2.5 times as many judgments. Send the same request with "confirm": true to start the backfill as a job. See backfill for the fields and the job. On a large namespace, Brussle can take more than 20 seconds to count the documents. Then the backfill request, with or without confirm, answers 202 with a backfill_start job. When that job is done, its result holds the estimate, or the job_id of the backfill that started. The SDKs wait for the job for you. See when a change takes long. A backfill skips the answers from a version that changed only thresholds. These answers read fresh. Thus, after such a version, the estimate is zero.

Change a question safely with a shadow report

A shadow report shows what a new version of a judgment would change, before you switch to it. When version 4’s engine or definition is different from the active version, the activate call returns 202 with a shadow job in awaiting_confirm. The job judges a random sample of 1,000 documents under version 4. If there are fewer documents, it judges all of them. While the job samples, report is null, and progress.documents_done counts the sampled documents. When the job is done, the report compares the two versions on the same documents:
The values are illustrative. The example leaves out current and candidate.
  • current is the active version’s answers. candidate is the new version’s answers. For bool and score, each side has a mean and a 10-bin histogram of p or score. For choice, each side has dist, the mean probability of each option. Thus, you can read the change option by option.
  • threshold_flips counts, for each named threshold, the documents that would go from false to true and from true to false. The new side uses the thresholds that will apply when the version is active. In the example, 31 more payments would wait for a person in the unmatched queue.
  • recompute estimates a backfill of each document under the new version: documents, tokens, judgments counted by size class times the weight of the new version’s engine, cost and duration.
Then decide:
  • POST /jobs/{id}/confirm (db.jobs.confirm(id)) switches to the new version. Confirm once the report is ready. A confirm before the report exists returns conflict and changes nothing. The job shows running, then done about a second later, when Brussle has committed the switch.
  • POST /jobs/{id}/cancel keeps the active version as it is.
  • {"version": 4, "force": true} on activate skips the report and switches immediately. activate: true on a create of a version does the same. A composite judgment is the exception. Its shadow job fits it on your labels, so the shadow job always runs.
Brussle writes shadow evaluations to the evaluation log with shadow: true. They never make answers. They never count toward calibration. Brussle does not bill them. Thresholds that you gave with the new version replace the current thresholds when the version becomes active. In the dashboard, the Overview of the judgment’s page shows the report, with Confirm, Cancel and Force. For a judgment on a template, activate the new version on the prefix path. You get one shadow report, with a sample from all the tenants.

Find answers by version

A query can filter on the version that made each answer:
  • answers.<judgment>.judgment_version is an integer. Compare it with Eq, NotEq, Lt, Lte, Gt, Gte, In or NotIn, or test it with Exists.
  • answers.<judgment>.engine_version is a string. Compare it with Eq, NotEq, In or NotIn, or test it with Exists. For an engine version that is not pinned, such as jev current or gpt-6-luna current, it holds the label of the epoch. The label does not name the engine, and two engines can have the same label. A judgment changes its engine only in a new version, so filter on judgment_version too when a judgment has used more than one engine.
For example, this query finds the payments whose settles answer is from a version before version 4:
Send it to POST /namespaces/{ns}/query. Use the same filter on engine_version to compare the answers of two epochs side by side. See find the answers of one version.

When the engine changes

Epochs

An epoch is a period in which the engine’s behaviour stays the same. The engine_version of an evaluation records its epoch.
  • An exact engine version is one epoch for its full life.
  • An engine version that is not pinned, such as Jev current, starts a new epoch each time its behaviour changes. The label is like current+2026-09-24.1. See the engines.
A new epoch never makes an answer stale, and it judges nothing again. Answers from before keep their epoch in engine_version. Each new answer records the new epoch. You can see that the engine changed in these places:
  • An event. Each judgment whose active version uses the changed engine version gets one judgment.engine_changed event for each new epoch. If the conformance run on the new epoch does worse than the last run that passed, each of these judgments also gets one judgment.engine_regressed event. A judgment that a template defines gets its events on the template. See the event types.
  • The answers. New answers have a new engine_version. Filter on it to find them.
  • The calibration report. lift.model_changed gives the earlier epoch, the current epoch and the day that the current epoch started. See read the lift.
  • The Calibration tab. It says that the engine changed while Brussle fits the new epoch.
Brussle fits calibration for each judgment version and each epoch, because a different engine version needs a different fit. After the engine changes by itself, the new epoch starts with no outcomes of its own. Brussle fits it again from the outcomes that you already posted. This is for a change in the behaviour of current. When you move a judgment to another engine yourself, see change the engine yourself.

After the engine changes

Some minutes after an engine version that is not pinned, such as Jev current, changes, Brussle fits each judgment with outcomes again for the new engine version. It uses a replay:
  1. Brussle takes the stored contexts of the answers that your most recent outcomes labelled, up to a fixed number.
  2. The same engine judges these contexts again.
  3. Each new answer, with its old outcome, is a sample of the new epoch.
You do not have to do anything. The replay runs on each plan. Calibration explains what Brussle sends to the engine. What you see:
  • The Calibration tab. For the new epoch, while the replay runs, it says “Re-fitting after the engine changed on 2026-10-02: 400 of 812 outcomes replayed so far.” Then it says “Re-fitted on 812 replayed outcomes after the engine changed on 2026-10-02.”
  • The report. The new epoch has fit_source: "replay" and its replay progress. Its own outcomes stay 0 until you post new outcomes.
  • Answers. Until Brussle makes the replay’s fit, usually in less than an hour, answers under the new epoch use the previous epoch’s calibration with "from_previous_epoch": true. For a judgment without features, the previous epoch’s calibration lasts for 30 days at most. A judgment with features keeps the previous epoch’s fit until the new epoch has its own.
  • Evaluation history. Replays show in a document’s evaluation history with replay_of, which names the evaluation that they replayed. Replays never change an answer. Brussle never bills them.
  • Outcomes. Continue to post outcomes as usual. When the new epoch has 100 of its own, with 20 of each kind, its fit uses only them, and fit_source becomes outcomes.
A replay can stop early (status: "stopped"). Then Brussle fits the new epoch on what it replayed, if that meets the minimums. If not, Brussle fits the new epoch from its own outcomes as they arrive. The reason tells why the replay stopped:
  • cost_cap: your organization reached this month’s replay limit. If the epoch still has no fit next month, the replay runs again on the outcomes that it did not replay.
  • superseded: the engine changed again. The newer epoch has its own replay.
Fits after an engine change come first. Recipe tuning uses only part of the monthly limit. Thus, these fits always have the rest.

Change the engine yourself

The engine is part of the definition. To pin an exact engine version, or to move to a different engine, post a new version with the new engine. Then activate it through its shadow report, as for any change to the question. Answers from the earlier version read stale with older_version until Brussle judges them again. For example, to move settles from jev current to gpt-6-luna current, post version 4 with "engine": {"name": "gpt-6-luna", "version": "current"} and the same question and recipe. These things change with the engine:
  • Cost. Each judgment counts its size class times the weight of the new engine: 1 on jev, 2.5 on gpt-6-luna. Each engine also counts tokens in its own way, so some documents can change size class. The shadow report’s recompute prices a backfill on the new engine.
  • What the engine reads. gpt-6-luna reads the images that the recipe selects. jev gets only their references. Thus, the context, its size class and the answers can all change.
  • The answers. The shadow report compares the raw answers of the two engines on the same documents. It does not compare calibrated values.
  • Calibration starts again. The new version has no fit. Its answers have no calibrated until it has 100 outcomes of its own, with 20 of each kind. Brussle does not replay your earlier outcomes through the new engine.
  • Epochs. Each engine numbers its own epochs, so jev and gpt-6-luna can have the same label, such as current+2026-10-07.1. To find the answers of the new engine, filter on judgment_version, not on engine_version alone.
  • Events. The judgment gets judgment.engine_changed events for the new engine’s epochs, and none for the old engine’s.