Skip to main content
An engine’s probability is its own confidence, and confidence is not the same as being right. An engine that says 0.9 might be right only 80% of the time on your documents, or 97%. Calibration measures that on your own outcomes and corrects it, so a calibrated 0.9 comes true about 90% of the time. It matters whenever you act on a probability: a threshold that sends a document to a person, an automation that runs above 0.95, a queue sorted by risk. With calibrated numbers, you can say how often those actions will be wrong.

Outcomes are the input

An outcome is what actually happened to a judged document, posted with POST /namespaces/{ns}/outcomes. There are two kinds:
  • Labelled examples. A person’s answer to the question for a document, such as “this ticket did need escalation”. Write the documents, post the labels, and each is joined to the document’s current answer. This measures a judgment before you rely on it.
  • Real-world results. What happened later, such as “this account churned”. On a judgment with a horizon, such as 30d, each answer predicts the event within that long, and each result labels the answer it proves right or wrong. See predictions with a horizon.
You rarely need to post them by hand. When the truth already arrives in data you write, such as an account’s status becoming cancelled or a ticket closing with its final team, the judgment’s outcome rules derive outcomes from those writes. See outcomes from your own data. Outcomes are append-only. Each counts for the judgment version and the engine epoch of the evaluation it joins. One that joins nothing is kept, and the calibration report counts it in unmatched_outcomes, with the reason.

Predictions with a horizon

An answer to a judgment with a 30-day horizon says “this happens within 30 days of now”. Outcomes are matched to those predictions like this:
  • Windows. A document’s first answer opens a window that runs 30 days from when the answer was made. Answers made while it is open do not open windows of their own, and the first answer after it closes opens the next one. Each window is measured on the answer that opened it.
  • Labels. An outcome labels the window its observed_at falls in. A churn on day 12 counts for the answer made on day 0. Two outcomes in one window count once: one you posted wins over one an outcome rule derived, and otherwise a true wins.
  • “It did not happen” after the fact. A false observed after a window closed, and before the next one opens, answers the window that closed: post “still active” on day 40 and it counts for the answer made on day 0. A true already in that window stands.
  • No prediction to measure. An outcome observed before the document’s first answer counts nowhere, and so does an event (true) observed after a window closed and before the next answer. The report says so in unmatched_outcomes (before_first_evaluation, after_horizon).
  • After the event. Once a true is observed for a document, answers made from then on have nothing left to predict: the account has already churned, and the write that recorded it often triggered the answer. Their windows count nowhere, neither as a false nor labelled by an outcome, and an outcome in one is counted in unmatched_outcomes.after_positive. For an event that can happen again, such as a second complaint, only the first counts.
  • Late outcomes count. Outcomes are matched whenever calibration or the report reads them, so a result posted weeks after it happened labels its window at the next fit.
  • Documents written again after a delete start a new window at their first new answer, and predict again even after an event. An answer is about one life of its document: an outcome observed after the new life’s first answer never labels a window from before the delete.
  • Deleted documents. An outcome observed while its document is deleted labels no window, and the report counts it in unmatched_outcomes.no_evaluation. A document written again remembers its id’s last 8 deletions, so this holds however late the outcome is posted. It needs the deletion to still be recorded when the outcome is posted or read, or when the id is written again.
  • Engine epochs. A window counts for the epoch of the answer that opened it.
Each window counts once, measured on the answer that opened it, so judging often doesn’t inflate the scores. Labelled examples, with the default 0s horizon, are simpler: each joins the document’s answer for the revision current at observed_at. One observed while the document is deleted joins nothing, on the same terms. An outcome a rule derived joins the revision just before the write that fired it, since that write revealed the answer.

How the fit works

Calibration is fitted separately for each judgment version and each engine epoch, because a different question or a different model needs a different correction. The count is the outcomes the fit rests on, which calibrated.outcomes shows on every calibrated answer. stage is on every calibrated answer and on each epoch of the calibration report, where it is null while nothing is fitted (not_fitted says why). While stage is early, don’t automate on calibrated values: a calibrated probability can move noticeably from one nightly fit to the next. Keep acting on thresholds over the raw numbers, picked with the threshold recommender and read with its interval, and use calibrated values for reporting until the fit reads full. The first fit runs within a few minutes of the first outcome you post, or of setting the judgment’s first outcome rules, and it is refitted every night after that, so new outcomes improve it over time. Outcomes that rules derive later are picked up by the nightly fit. There is no call to start a fit on demand.

Only when it beats the raw numbers

A fit is used only if it does better than the engine’s own probabilities on held-out outcomes: each outcome is scored by log loss against fits made without it. If the raw probabilities score as well or better, the fit is not applied. Answers then carry no calibrated object, and the report’s not_fitted says raw_better, with both scores in held_out. An engine that is already well calibrated on your data stays raw, rather than picking up the fit’s noise.

At least 20 of each kind

100 outcomes are not enough on their own. An epoch also needs outcomes on both sides:
  • bool: at least 20 true and at least 20 false.
  • choice and score: at least two different values, with 20 outcomes each. Not every option: most options are rare, and the fit needs to see the engine’s top answer both right and wrong.
A fit on one kind of outcome only learns that kind. If every outcome is true, the “best” correction says everything is likely, which is wrong for every document that did not happen. So until an epoch has both, it gets no calibration and no threshold recommendation, and the calibration report says why in not_fitted:

Post what did not happen too

Most teams record what happened: the customer churned, the ticket was escalated, the transaction was fraud. What did not happen is rarely written down, but calibration needs it just as much.
  • For labelled examples, label documents where the answer is “no” as well as “yes”, with "value": false.
  • For real-world results, post false once you know it did not happen: when a prediction’s window has passed without the event, such as an account still active 30 days after the answer. It answers the latest window that closed by its observed_at.
  • For a bool judgment with a horizon, you can turn on implicit negatives instead, when the absence of an outcome really does mean “no”.

Outcomes from a reviewed sample

Outcomes usually come from documents someone already looked at, and people look at what crossed a threshold. Then almost every outcome sits at a high raw probability, and the fit learns nothing about the rest of the range. Its calibrated numbers for low-probability answers are a guess. The calibration report warns about this for each epoch, in warnings:
  • above_thresholds: more than 80% of the outcomes are on answers that meet one of the judgment’s thresholds.
  • high_band: for a bool judgment, more than 80% of the outcomes are on answers with p of 0.7 or more.
The report’s coverage gives the lowest and highest raw value the outcomes cover, and the reliability curve shows how many fall in each band. What to do about it: label a random sample of documents too, not only the ones that were reviewed. The labelling queue picks that sample for you. Keep labelling the reviewed ones as well; both count.

The labelling queue

The labelling queue hands out documents to label that nobody chose: a random sample of the judgment’s answers, spread across the whole probability range. Labels from it are unbiased, which reviewed outcomes are not. Across the whole range. Most answers sit in one or two parts of the range: a fraud judgment says “no” with p near 0.02 to almost everything, so a plain random sample would say nothing about how the engine does at 0.6. The queue hands out documents across the whole probability range, not just where most answers sit, and weights their labels back to your real traffic, so the fit and its held-out test reflect traffic as it is. Each queued document names the probability band it was drawn from. How it combines with other outcomes. Once an epoch’s queue labels alone meet the minimums, 100 labels with 20 of each kind, the fit uses only them, because nobody chose them. Until then, queue labels count as ordinary outcomes beside the others and widen the range the fit covers. The calibration report shows which it is in each epoch’s fitted_on (queue or all), and counts queue labels in outcomes_by_source.queue. The threshold recommender still uses every outcome, unweighted. Predictions with a horizon have no queue. An answer to “will this account churn within 30 days?” is a question about the next 30 days, which nobody can answer by looking at the account today. Those judgments get honest negatives from implicit negatives and positives as they happen, or from outcome rules. Leases. Each draw leases its documents for 7 days, so two people labelling at once are not given the same document. A label is taken while the lease holds; after that it is refused (conflict), since the document is back in the queue for someone else. A document labelled from the queue leaves it until its answer changes. An answer keeps one queue label, the latest recorded, so a retry or a corrected label counts once. Posted and rule outcomes do not take a document out of the queue: those are the reviewed documents, and leaving them out would make the sample a sample of the unreviewed ones. The queue is part of the learning loop, on the Team plan and above. Every plan can preview it: the answers in each band, and how many are labelled, leased and still available. See label a random sample to use it.

Implicit negatives

For a bool judgment with a non-zero horizon, such as “will this account churn within 30 days?”, you can post only what happened and let every other judged document count as “no”:
Send it with PATCH /namespaces/{ns}/judgments/{name}. It is a setting, like thresholds, so it creates no version. The next fit and every report use it. With it on:
  • A prediction window that closes with no outcome counts as a false outcome, for the answer that opened it. A window still open counts nothing yet: an answer from yesterday on a 30-day judgment does not count, because its outcome may still come.
  • Outcomes you post, or that outcome rules derive, always win in their window. Post a positive promptly, or let a rule derive it from the write that records it: until then, its window counts as a negative once it closes.
  • A window whose document was deleted before it closed is left out, unless an outcome labels it. A deleted account may well have churned, so its silence is not a “no”. Once the deletion is no longer recorded, the document counts as existing until its newest answer.
  • A window opened after the event is left out: an account that has already cancelled does not cancel again, so its silence is not a “no” either.
  • The report shows how many outcomes are implicit, in each epoch’s implicit_negatives and in outcomes_by_source, and the windows left out, in censored_predictions.
  • A judgment with many judged documents is calibrated on a representative sample of them, with the outcomes posted for those documents, so the share of positives stays true.
When not to use it. Only turn it on when a missing outcome really means “it did not happen”. Leave it off when:
  • you only record some of the events, such as churn from one billing system but not another;
  • outcomes arrive long after the horizon, such as chargebacks that take 90 days to report on a 30-day judgment;
  • the judgment is not a prediction at all, such as labelled examples with the default 0s horizon. Those are refused.
On a namespace that inherits a judgment from a template, set it on the template’s prefix path: every namespace under it shares one calibration.

What changes in an answer

Once a fit rests on enough outcomes and beats the raw numbers, each answer carries a calibrated object beside the raw numbers:
  • The raw p, dist and score never change. calibrated sits beside them.
  • It is computed when the answer is read, so a nightly refit updates every answer without re-judging anything.
  • Thresholds, filters and ranking use the raw numbers, so a refit never moves a document across a threshold on its own. To act on calibrated numbers, pick thresholds with the recommender, which works from your outcomes.
  • extrapolated: true means the answer’s raw value is outside the range the fit was made on: below the lowest or above the highest raw value among its outcomes. For a choice or score, the raw value is the probability of the most probable option or level. The calibrated number is then a guess from the edge of what the fit has seen. Post outcomes for documents like it to cover that range.

What calibration can and cannot do

  • It makes probabilities honest. After calibration, the stated probability matches how often answers come true, measured on your data. The calibration report shows both, before and after.
  • It does not change the order of documents. The correction only ever maps a higher raw probability to an equal or higher calibrated one. So it cannot make the engine better at telling likely documents from unlikely ones. If the right documents are not near the top of the raw ranking, a better question or context recipe is what helps.
  • It can change a yes or no at 0.5. Accuracy before and after can differ, and the report shows both.
  • It is only as good as the outcomes. Outcomes from one kind of document, or from one part of the range, give a correction that is right there and a guess elsewhere. The report’s not_fitted, coverage and warnings say when that is the case.

Seeing it improve

Each nightly fit’s held-out scores are kept for 400 days, so you can see what the loop has done. The calibration report’s lift gives one number, such as “calibration cut this judgment’s error by 38% since it started, on 1,240 outcomes”, with its 95% interval, and the error after each fit. It is measured the way a fit is judged: log loss on outcomes the fit did not see, raw against what answers read, on the same outcomes and within one engine epoch. Below the minimums there is no number. See read the lift.

When the engine changes

Jev current follows its provider’s model, and each change in its behaviour that we detect starts a new epoch, recorded on every answer’s engine_version. A new epoch starts with no outcomes of its own, but the outcomes you already posted still say what came true. So the new epoch is re-fitted from them straight away, without waiting for new ones. The replay. Within minutes of the change, each judgment’s recent outcomes are replayed through the new model:
  • What is replayed. The outcomes that joined an answer of the active version in an earlier epoch, newest first, up to a bounded number. Each outcome replays the exact answer it measured, so prediction windows are kept.
  • What is sent. Each of those answers’ stored compiled context, the text the engine read then, with the version’s question, to the same engine. It goes to the same engine provider your answers already use, on the same terms, so no new subprocessor sees your data (see where your data goes). A document you have deleted since is not replayed.
  • What is recorded. Each call is a replay evaluation in the evaluation log, with replay_of naming the evaluation it replayed. It is never an answer: nothing in your answers, queries or thresholds changes. It is never billed: no judgments, and its log is not counted in your storage.
  • The fit. Each replayed answer and its outcome is a sample of the new epoch, fitted with the same rules as any fit: the minimums, the held-out test, queue weights and, on templates, each tenant’s own fit. The fit is made as soon as the replay ends, usually within the hour. Once the new epoch has 100 outcomes of its own, with 20 of each kind, it is fitted on those alone.
Replay is part of calibration, so it runs on every plan. Replays never slow your answers and are never billed; a monthly allowance per organization bounds them, shared with recipe tuning. A replay that reaches it stops (cost_cap), and its epoch is fitted on what was replayed if that meets the minimums; if it still has no fit the next month, the replay runs again on the outcomes it had not replayed. While it runs. Answers from the new epoch use the previous epoch’s calibration, marked "from_previous_epoch": true, so you can tell they were corrected for the old model. That lasts until the new epoch has a fit, and at most 30 days from the day the model changed: a fit for another model is a guess, so after that the answers carry no calibrated until the epoch is fitted. The threshold recommender uses the previous epoch’s outcomes the same way, and says so with from_previous_epoch. In the report. The new epoch’s fit_source is replay while its fit rests on replayed outcomes, and outcomes once its own are enough, and replay shows the replay’s progress: planned, done, and the replayed outcomes its fit rests on. The epoch’s own outcomes count only outcomes that joined the new model’s answers, so the report’s totals count each outcome once. lift says the model changed and starts the new model’s headline from its first fit. See after the model changes. An exact engine version never changes, so it is one epoch for its whole life. A new judgment version starts fresh too. It needs its own outcomes before its answers carry calibrated.

Templates and composites

  • A template judgment’s report on its prefix pools the outcomes of every namespace under it, and a tenant without a fit of its own reads that pooled fit, so a new tenant is calibrated from its first answer. On Scale, a tenant with enough outcomes of its own (100, with 20 of each kind, in an epoch) also gets its own fit, weighted more heavily as its outcomes grow (the report shows weight), and used only when it clearly beats the pool on that tenant’s held-out outcomes. See each tenant grows its own fit.
  • A composite judgment has no calibrated object. Its p comes from weights fitted on your labels, refitted nightly in the same run, which usually leaves it close to calibrated, but not always: check the reliability curve in its report.
To post outcomes, read the report and pick thresholds, see measure, improve and tune.