> ## Documentation Index
> Fetch the complete documentation index at: https://docs.brussle.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Calibration

> How answers learn from what actually happened, so a probability means what it says.

An engine's probability is its own confidence, and confidence is not the same as being right. An engine that says 0.9 might be right only 80% of the time on your documents, or 97%. Calibration measures that on your own outcomes and corrects it, so a calibrated 0.9 comes true about 90% of the time.

It matters whenever you act on a probability: a threshold that sends a document to a person, an automation that runs above 0.95, a queue sorted by risk. With calibrated numbers, you can say how often those actions will be wrong.

## Outcomes are the input

An **outcome** is what actually happened to a judged document, posted with `POST /namespaces/{ns}/outcomes`. There are two kinds:

* **Labelled examples.** A person's answer to the question for a document, such as "this ticket did need escalation". Write the documents, post the labels, and each is joined to the document's current answer. This measures a judgment before you rely on it.
* **Real-world results.** What happened later, such as "this account churned". On a judgment with a `horizon`, such as `30d`, each answer predicts the event within that long, and each result labels the answer it proves right or wrong. See [predictions with a horizon](#predictions-with-a-horizon).

You rarely need to post them by hand. When the truth already arrives in data you write, such as an account's status becoming `cancelled` or a ticket closing with its final team, the judgment's **outcome rules** derive outcomes from those writes. See [outcomes from your own data](/guides/measure-improve-tune#outcomes-from-your-own-data).

Outcomes are append-only. Each counts for the judgment version and the engine epoch of the evaluation it joins. One that joins nothing is kept, and the [calibration report](/guides/measure-improve-tune#read-the-calibration-report) counts it in `unmatched_outcomes`, with the reason.

## Predictions with a horizon

An answer to a judgment with a 30-day `horizon` says "this happens within 30 days of now". Outcomes are matched to those predictions like this:

* **Windows.** A document's first answer opens a window that runs 30 days from when the answer was made. Answers made while it is open do not open windows of their own, and the first answer after it closes opens the next one. Each window is measured on the answer that opened it.
* **Labels.** An outcome labels the window its `observed_at` falls in. A churn on day 12 counts for the answer made on day 0. Two outcomes in one window count once: one you posted wins over one an outcome rule derived, and otherwise a `true` wins.
* **"It did not happen" after the fact.** A `false` observed after a window closed, and before the next one opens, answers the window that closed: post "still active" on day 40 and it counts for the answer made on day 0. A `true` already in that window stands.
* **No prediction to measure.** An outcome observed before the document's first answer counts nowhere, and so does an event (`true`) observed after a window closed and before the next answer. The report says so in `unmatched_outcomes` (`before_first_evaluation`, `after_horizon`).
* **After the event.** Once a `true` is observed for a document, answers made from then on have nothing left to predict: the account has already churned, and the write that recorded it often triggered the answer. Their windows count nowhere, neither as a `false` nor labelled by an outcome, and an outcome in one is counted in `unmatched_outcomes.after_positive`. For an event that can happen again, such as a second complaint, only the first counts.
* **Late outcomes count.** Outcomes are matched whenever calibration or the report reads them, so a result posted weeks after it happened labels its window at the next fit.
* **Documents written again** after a delete start a new window at their first new answer, and predict again even after an event. An answer is about one life of its document: an outcome observed after the new life's first answer never labels a window from before the delete.
* **Deleted documents.** An outcome observed while its document is deleted labels no window, and the report counts it in `unmatched_outcomes.no_evaluation`. A document written again remembers its id's last 8 deletions, so this holds however late the outcome is posted. It needs the deletion to still be recorded when the outcome is posted or read, or when the id is written again.
* **Engine epochs.** A window counts for the epoch of the answer that opened it.

Each window counts once, measured on the answer that opened it, so judging often doesn't inflate the scores.

Labelled examples, with the default `0s` horizon, are simpler: each joins the document's answer for the revision current at `observed_at`. One observed while the document is deleted joins nothing, on the same terms. An outcome a rule derived joins the revision just before the write that fired it, since that write revealed the answer.

## How the fit works

Calibration is fitted separately for each judgment version and each engine epoch, because a different question or a different model needs a different correction.

| Outcomes for that version and epoch | What happens |
| - | - |
| Fewer than 100, or too few of one kind (below) | No calibration. Answers carry only the engine's raw numbers. |
| 100 to 999, with 20 of each kind | `stage: early`. The correction is deliberately conservative and its uncertainty is wider, so expect it to move more between fits. |
| 1,000 or more | `stage: full`. The correction follows the engine's own pattern of over- and under-confidence closely. |

The count is the outcomes the fit rests on, which `calibrated.outcomes` shows on every calibrated answer. `stage` is on every `calibrated` answer and on each epoch of the [calibration report](/guides/measure-improve-tune#read-the-calibration-report), where it is `null` while nothing is fitted (`not_fitted` says why).

While `stage` is `early`, don't automate on calibrated values: a calibrated probability can move noticeably from one nightly fit to the next. Keep acting on thresholds over the raw numbers, picked with the [threshold recommender](/guides/measure-improve-tune#pick-thresholds-with-the-recommender) and read with its interval, and use calibrated values for reporting until the fit reads `full`.

The first fit runs within a few minutes of the first outcome you post, or of setting the judgment's first [outcome rules](/guides/measure-improve-tune#outcomes-from-your-own-data), and it is refitted every night after that, so new outcomes improve it over time. Outcomes that rules derive later are picked up by the nightly fit. There is no call to start a fit on demand.

### Only when it beats the raw numbers

A fit is used only if it does better than the engine's own probabilities on held-out outcomes: each outcome is scored by log loss against fits made without it. If the raw probabilities score as well or better, the fit is not applied. Answers then carry no `calibrated` object, and the report's `not_fitted` says `raw_better`, with both scores in `held_out`. An engine that is already well calibrated on your data stays raw, rather than picking up the fit's noise.

### At least 20 of each kind

100 outcomes are not enough on their own. An epoch also needs outcomes on both sides:

* **`bool`:** at least 20 `true` and at least 20 `false`.
* **`choice` and `score`:** at least two different values, with 20 outcomes each. Not every option: most options are rare, and the fit needs to see the engine's top answer both right and wrong.

A fit on one kind of outcome only learns that kind. If every outcome is `true`, the "best" correction says everything is likely, which is wrong for every document that did not happen. So until an epoch has both, it gets no calibration and no threshold recommendation, and the [calibration report](/guides/measure-improve-tune#read-the-calibration-report) says why in `not_fitted`:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"reason": "too_few_negatives", "message": "Only positive outcomes so far: post outcomes for documents where it did not happen."}
```

## Post what did not happen too

Most teams record what happened: the customer churned, the ticket was escalated, the transaction was fraud. What did not happen is rarely written down, but calibration needs it just as much.

* **For labelled examples,** label documents where the answer is "no" as well as "yes", with `"value": false`.
* **For real-world results,** post `false` once you know it did not happen: when a prediction's window has passed without the event, such as an account still active 30 days after the answer. It answers the latest window that closed by its `observed_at`.
* **For a `bool` judgment with a horizon,** you can turn on [implicit negatives](#implicit-negatives) instead, when the absence of an outcome really does mean "no".

## Outcomes from a reviewed sample

Outcomes usually come from documents someone already looked at, and people look at what crossed a threshold. Then almost every outcome sits at a high raw probability, and the fit learns nothing about the rest of the range. Its calibrated numbers for low-probability answers are a guess.

The calibration report warns about this for each epoch, in `warnings`:

* **`above_thresholds`:** more than 80% of the outcomes are on answers that meet one of the judgment's thresholds.
* **`high_band`:** for a `bool` judgment, more than 80% of the outcomes are on answers with `p` of 0.7 or more.

The report's `coverage` gives the lowest and highest raw value the outcomes cover, and the reliability curve shows how many fall in each band.

**What to do about it:** label a random sample of documents too, not only the ones that were reviewed. The [labelling queue](#the-labelling-queue) picks that sample for you. Keep labelling the reviewed ones as well; both count.

## The labelling queue

The labelling queue hands out documents to label that nobody chose: a random sample of the judgment's answers, spread across the whole probability range. Labels from it are unbiased, which reviewed outcomes are not.

**Across the whole range.** Most answers sit in one or two parts of the range: a fraud judgment says "no" with `p` near 0.02 to almost everything, so a plain random sample would say nothing about how the engine does at 0.6. The queue hands out documents across the whole probability range, not just where most answers sit, and weights their labels back to your real traffic, so the fit and its [held-out test](#only-when-it-beats-the-raw-numbers) reflect traffic as it is. Each queued document names the probability `band` it was drawn from.

**How it combines with other outcomes.** Once an epoch's queue labels alone meet the [minimums](#at-least-20-of-each-kind), 100 labels with 20 of each kind, the fit uses only them, because nobody chose them. Until then, queue labels count as ordinary outcomes beside the others and widen the range the fit covers. The [calibration report](/guides/measure-improve-tune#read-the-calibration-report) shows which it is in each epoch's `fitted_on` (`queue` or `all`), and counts queue labels in `outcomes_by_source.queue`. The threshold recommender still uses every outcome, unweighted.

**Predictions with a horizon have no queue.** An answer to "will this account churn within 30 days?" is a question about the next 30 days, which nobody can answer by looking at the account today. Those judgments get honest negatives from [implicit negatives](#implicit-negatives) and positives as they happen, or from [outcome rules](/guides/measure-improve-tune#outcomes-from-your-own-data).

**Leases.** Each draw leases its documents for 7 days, so two people labelling at once are not given the same document. A label is taken while the lease holds; after that it is refused (`conflict`), since the document is back in the queue for someone else. A document labelled from the queue leaves it until its answer changes. An answer keeps one queue label, the latest recorded, so a retry or a corrected label counts once. Posted and rule outcomes do not take a document out of the queue: those are the reviewed documents, and leaving them out would make the sample a sample of the unreviewed ones.

The queue is part of the learning loop, on the Team plan and above. Every plan can preview it: the answers in each band, and how many are labelled, leased and still available. See [label a random sample](/guides/measure-improve-tune#label-a-random-sample) to use it.

## Implicit negatives

For a `bool` judgment with a non-zero `horizon`, such as "will this account churn within 30 days?", you can post only what happened and let every other judged document count as "no":

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"outcomes": {"implicit_negatives": true}}
```

Send it with `PATCH /namespaces/{ns}/judgments/{name}`. It is a setting, like thresholds, so it creates no version. The next fit and every report use it.

With it on:

* A [prediction window](#predictions-with-a-horizon) that closes with no outcome counts as a `false` outcome, for the answer that opened it. A window still open counts nothing yet: an answer from yesterday on a 30-day judgment does not count, because its outcome may still come.
* Outcomes you post, or that [outcome rules](/guides/measure-improve-tune#outcomes-from-your-own-data) derive, always win in their window. Post a positive promptly, or let a rule derive it from the write that records it: until then, its window counts as a negative once it closes.
* A window whose document was deleted before it closed is left out, unless an outcome labels it. A deleted account may well have churned, so its silence is not a "no". Once the deletion is no longer recorded, the document counts as existing until its newest answer.
* A window opened [after the event](#predictions-with-a-horizon) is left out: an account that has already cancelled does not cancel again, so its silence is not a "no" either.
* The report shows how many outcomes are implicit, in each epoch's `implicit_negatives` and in `outcomes_by_source`, and the windows left out, in `censored_predictions`.
* A judgment with many judged documents is calibrated on a representative sample of them, with the outcomes posted for those documents, so the share of positives stays true.

**When not to use it.** Only turn it on when a missing outcome really means "it did not happen". Leave it off when:

* you only record some of the events, such as churn from one billing system but not another;
* outcomes arrive long after the horizon, such as chargebacks that take 90 days to report on a 30-day judgment;
* the judgment is not a prediction at all, such as labelled examples with the default `0s` horizon. Those are refused.

On a namespace that inherits a judgment from a [template](/guides/templates), set it on the template's prefix path: every namespace under it shares one calibration.

## What changes in an answer

Once a fit rests on enough outcomes and [beats the raw numbers](#only-when-it-beats-the-raw-numbers), each answer carries a `calibrated` object beside the raw numbers:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "needs_escalation": {
    "type": "bool",
    "p": 0.91,
    "calibrated": {"p": 0.84, "stage": "full", "outcomes": 1432, "from_previous_epoch": false, "extrapolated": false},
    "freshness": "fresh"
  }
}
```

* The raw `p`, `dist` and `score` never change. `calibrated` sits beside them.
* It is computed when the answer is read, so a nightly refit updates every answer without re-judging anything.
* Thresholds, filters and ranking use the raw numbers, so a refit never moves a document across a threshold on its own. To act on calibrated numbers, pick thresholds with the [recommender](/guides/measure-improve-tune#pick-thresholds-with-the-recommender), which works from your outcomes.
* **`extrapolated: true`** means the answer's raw value is outside the range the fit was made on: below the lowest or above the highest raw value among its outcomes. For a `choice` or `score`, the raw value is the probability of the most probable option or level. The calibrated number is then a guess from the edge of what the fit has seen. Post outcomes for documents like it to cover that range.

## What calibration can and cannot do

* **It makes probabilities honest.** After calibration, the stated probability matches how often answers come true, measured on your data. The [calibration report](/guides/measure-improve-tune#read-the-calibration-report) shows both, before and after.
* **It does not change the order of documents.** The correction only ever maps a higher raw probability to an equal or higher calibrated one. So it cannot make the engine better at telling likely documents from unlikely ones. If the right documents are not near the top of the raw ranking, a better question or [context recipe](/guides/context-recipes) is what helps.
* **It can change a yes or no at 0.5.** Accuracy before and after can differ, and the report shows both.
* **It is only as good as the outcomes.** Outcomes from one kind of document, or from one part of the range, give a correction that is right there and a guess elsewhere. The report's `not_fitted`, `coverage` and `warnings` say when that is the case.

## Seeing it improve

Each nightly fit's held-out scores are kept for 400 days, so you can see what the loop has done. The calibration report's `lift` gives one number, such as "calibration cut this judgment's error by 38% since it started, on 1,240 outcomes", with its 95% interval, and the error after each fit. It is measured the way a fit is judged: log loss on outcomes the fit did not see, raw against what answers read, on the same outcomes and within one engine epoch. Below the minimums there is no number. See [read the lift](/guides/measure-improve-tune#read-the-lift).

## When the engine changes

Jev `current` follows its provider's model, and each change in its behaviour that we detect starts a new **epoch**, recorded on every answer's `engine_version`. A new epoch starts with no outcomes of its own, but the outcomes you already posted still say what came true. So the new epoch is re-fitted from them straight away, without waiting for new ones.

**The replay.** Within minutes of the change, each judgment's recent outcomes are replayed through the new model:

* **What is replayed.** The outcomes that joined an answer of the active version in an earlier epoch, newest first, up to a bounded number. Each outcome replays the exact answer it measured, so [prediction windows](#predictions-with-a-horizon) are kept.
* **What is sent.** Each of those answers' stored compiled context, the text the engine read then, with the version's question, to the same engine. It goes to the same engine provider your answers already use, on the same terms, so no new subprocessor sees your data (see [where your data goes](/behavior#where-your-data-goes)). A document you have deleted since is not replayed.
* **What is recorded.** Each call is a replay evaluation in the evaluation log, with `replay_of` naming the evaluation it replayed. It is never an answer: nothing in your answers, queries or thresholds changes. It is never billed: no judgments, and its log is not counted in your storage.
* **The fit.** Each replayed answer and its outcome is a sample of the new epoch, fitted with the same rules as any fit: the [minimums](#at-least-20-of-each-kind), the [held-out test](#only-when-it-beats-the-raw-numbers), queue weights and, on templates, each tenant's own fit. The fit is made as soon as the replay ends, usually within the hour. Once the new epoch has 100 outcomes of its own, with 20 of each kind, it is fitted on those alone.

Replay is part of calibration, so it runs on every plan. Replays never slow your answers and are never billed; a monthly allowance per organization bounds them, shared with [recipe tuning](/guides/context-recipes#let-your-outcomes-tune-the-recipe). A replay that reaches it stops (`cost_cap`), and its epoch is fitted on what was replayed if that meets the minimums; if it still has no fit the next month, the replay runs again on the outcomes it had not replayed.

**While it runs.** Answers from the new epoch use the previous epoch's calibration, marked `"from_previous_epoch": true`, so you can tell they were corrected for the old model. That lasts until the new epoch has a fit, and at most 30 days from the day the model changed: a fit for another model is a guess, so after that the answers carry no `calibrated` until the epoch is fitted. The threshold recommender uses the previous epoch's outcomes the same way, and says so with `from_previous_epoch`.

**In the report.** The new epoch's `fit_source` is `replay` while its fit rests on replayed outcomes, and `outcomes` once its own are enough, and `replay` shows the replay's progress: `planned`, `done`, and the replayed `outcomes` its fit rests on. The epoch's own `outcomes` count only outcomes that joined the new model's answers, so the report's totals count each outcome once. `lift` says the model changed and starts the new model's headline from its first fit. See [after the model changes](/guides/measure-improve-tune#after-the-model-changes).

An exact engine version never changes, so it is one epoch for its whole life.

A new judgment version starts fresh too. It needs its own outcomes before its answers carry `calibrated`.

## Templates and composites

* A [template](/guides/templates) judgment's report on its prefix pools the outcomes of every namespace under it, and a tenant without a fit of its own reads that pooled fit, so a new tenant is calibrated from its first answer. On Scale, a tenant with enough outcomes of its own (100, with 20 of each kind, in an epoch) also gets its own fit, weighted more heavily as its outcomes grow (the report shows `weight`), and used only when it clearly beats the pool on that tenant's held-out outcomes. See [each tenant grows its own fit](/guides/templates#each-tenant-grows-its-own-fit-scale).
* A [composite judgment](/guides/composite-judgments) has no `calibrated` object. Its `p` comes from weights fitted on your labels, refitted nightly in the same run, which usually leaves it close to calibrated, but not always: check the reliability curve in its report.

To post outcomes, read the report and pick thresholds, see [measure, improve and tune](/guides/measure-improve-tune).


## Related topics

- [Get the calibration report for the active version](/api-reference/outcomes-and-calibration/get-the-calibration-report-for-the-active-version.md)
- [Measure, improve, tune](/guides/measure-improve-tune.md)
- [Templates](/guides/templates.md)
- [Append observed outcomes or labelled examples](/api-reference/outcomes-and-calibration/append-observed-outcomes-or-labelled-examples.md)
- [Run this month's recipe tuning now](/api-reference/outcomes-and-calibration/run-this-months-recipe-tuning-now.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.