Skip to main content
An engine’s probabilities are its own. To know how good a judgment is on your data, give Brussle your answers to the same question. They are called outcomes. From them you get a calibration report, calibrated probabilities on every answer, and threshold recommendations. A shadow report shows what a new version of the question would change before you switch to it.

The learning loop by plan

Calibration from the outcomes you post is on every plan: the report, and calibrated values on every answer. The rest of the learning loop is on Team and Scale, and Developer sees a preview of it in the dashboard: A call your plan doesn’t include returns plan_required (HTTP 402). Wherever such a feature appears, the dashboard says what it does and which plan has it, with an upgrade for owners and a note to ask an owner for everyone else. See plans.

Post labelled examples

A labelled example is an outcome with the judgment’s default horizon of 0s. Write the documents, then post their labels. Each label is joined to the evaluation of the document revision that was current at observed_at, so a label posted before that evaluation finishes still joins it. A label observed while its document is deleted joins nothing, because no revision was current. A document written again remembers its id’s last 8 deletions, so a label posted later for a time it was deleted joins nothing too. This needs the deletion to still be recorded when the label is posted or read, or when the id is written again. Post them to POST /namespaces/{ns}/outcomes:
value is what the answer should have been:
  • Post false outcomes as well as true ones. Calibration and the recommender need at least 20 of each (for a choice or score, two values with 20 each). See calibration.
  • Label a random sample of documents too, not only the ones a person already reviewed, so the outcomes cover the whole range of answers. The labelling queue picks one for you.
  • A value that does not fit the judgment’s type is invalid_request. An unknown judgment is not_found.
  • An outcome that joins no evaluation, such as one for a document never judged, is kept but adds nothing. The calibration report counts it in unmatched_outcomes, with the reason.
  • Outcomes are append-only. Retries are safe: the same document, judgment, value and observed_at is stored as one outcome.
  • Outcomes belong to a namespace. Post them on the namespace’s own path, never on a template prefix.

Upload labels as a CSV

On the dashboard, open the judgment and go to its Labels tab. Upload a CSV with the header document_id,value and an optional observed_at column in RFC 3339:
Your browser checks every row against the judgment’s type and lists the row errors before anything is sent. Valid rows are posted in batches of 1,000. A row without observed_at uses the time of the upload.

Label a random sample

The labelling queue hands out documents to label that nobody chose, spread evenly across the probability range, so the labels are unbiased. Once an epoch has enough of them, its calibration rests on them alone, each weighted back to how common its part of the range is in real traffic. Calibration explains why. On the dashboard, open the judgment and go to its Label tab. It shows one document at a time with the judgment’s question and a button per answer: yes or no for a bool, an option for a choice, a level for a score, or skip. Each answer is posted as an outcome at once, and the tab shows how many of each band are labelled. Through the API, draw documents with POST /namespaces/{ns}/judgments/{name}/labelling-queue:
count is 1 to 200 (default 50). The response lists the items, lowest band first:
  • band is one of 10 equal bands of the raw value, and probability the raw value: p for a bool, the most probable option’s or level’s probability for a choice or score, whose value it is. The count is split evenly over the bands that have documents, so a band holding 2% of the answers gets as many items as one holding 80%.
  • document is what the labeller needs to see. evaluation_id leads to the exact context the engine read, at GET /namespaces/{ns}/evaluations/{id}.
  • The items are leased until lease_expires_at, 7 days: nobody else is given them. A document labelled from the queue leaves it until its answer changes; one not labelled in time goes back.
  • seed (default 0) fixes the random order within each band. The same documents, labels and seed give the same items.
Post each label as an ordinary outcome with its item’s queue_item_id:
  • The outcome counts as a queue label: outcomes_by_source.queue in the calibration report. Once an epoch has 100 queue labels, 20 of each kind (for a choice or score, two values with 20 each), its fitted_on becomes queue, and the next fit uses the queue labels only, weighted by band.
  • A queue_item_id whose lease has run out, 7 days after the draw, is conflict; one whose draw is more than 30 days old, or that names another document, is invalid_request. Draw new items.
  • Labelling the same item again, to retry or to correct a label, replaces the earlier label: an answer counts only its latest queue label.
  • Label what the document shows. Skipping an item is fine: it goes back to the queue when its lease ends.
GET on the same path previews the queue without handing anything out: for each band, the answers, and how many are labelled, leased and available. It is ns.judgments.labellingQueue(name) in TypeScript and labelling_queue(name) in Python. Plans. Drawing is part of the learning loop, on the Team plan and above: below it, a draw is plan_required. The preview is on every plan, so on Developer the tab shows what the queue would contain. Horizons. A judgment with a non-zero horizon has no queue (invalid_request). Its answers predict what happens next, which nobody can label by looking at the document today. Use real-world outcomes and implicit negatives instead.

Read the calibration report

GET /namespaces/{ns}/judgments/{name}/calibration returns the report for the active version. It is ns.judgments.calibration(name) in both SDKs, and the Calibration tab on the judgment’s page in the dashboard.
To keep it short, the example leaves out each epoch’s ten-bin reliability curve under raw and calibrated, and all but the first entry of lift.history. At the top, beside the total outcomes:
  • outcomes_by_source, the outcomes by where they came from: posted, derived by outcome rules (rule), labels from the labelling queue (queue), or implicit negatives (implicit). They add up to outcomes.
  • rule_outcomes_paused, rule outcomes left out because your plan didn’t include outcome rules when they were observed (after a downgrade to Developer takes effect). They are kept, and aren’t in outcomes. It stays 0 while your plan has included outcome rules throughout.
  • unmatched_outcomes, the outcomes, posted or derived, that add nothing, by reason: They are kept, and count if a later answer takes them.
  • censored_predictions, with implicit negatives on, the prediction windows with no outcome that are not counted as false: window_open (not closed yet), document_deleted (the document was deleted before the window closed) and after_positive (the window opened after a true for the document).
Epochs are listed newest first. Each one gives:
  • outcomes, the outcomes joined to this epoch’s evaluations, outcomes_by_source, the same split as at the top, rule_outcomes_paused, as at the top, and implicit_negatives, how many of them are implicit negatives (0 unless the setting is on). Outcomes replayed into the epoch after the model changed are not here but in replay.outcomes, so each outcome counts once in the report.
  • fitted_on, what the fit rests on: queue once the epoch’s queue labels alone meet the minimums (100, with 20 of each kind), each weighted back to real traffic; otherwise all, every outcome counted once.
  • fit_source, replay while the epoch’s fit rests on outcomes replayed after the model changed, and outcomes otherwise. null for a composite.
  • replay, the replay that re-fitted the epoch after the model changed, or null when none ran for it: its status (running, done or stopped, with a reason: cost_cap or superseded), the epoch its outcomes came from (from_engine_version), how many it planned and has done, the replayed outcomes the fit rests on now, started_at and finished_at.
  • stage, how far along the epoch’s calibration is, and fitted_at, when it was fitted: Calibration is refitted every night.
  • not_fitted, why the epoch’s answers have no calibration of their own, or null when they do. reason is one of: When more than one applies, reason names the first of: a kind of outcome missing altogether (all true is too_few_negatives, all false is too_few_positives, and a choice or score with outcomes of one value is too_few_values), then fewer than 100 outcomes, then fewer than 20 of a kind. So 50 outcomes, all true, read too_few_negatives, not too_few_outcomes. message says what to post, such as “Only positive outcomes so far: post outcomes for documents where it did not happen.”
  • held_out, the test a fit must pass before answers use it: the raw answers’ log loss, and the log loss of fits made without the outcome they score. The fit is applied only when calibrated_log_loss is lower. null when nothing is fitted. Fits also give raw_ece and calibrated_ece on the same held-out outcomes, and standard_error, how sure the difference in log loss is; a template tenant’s own test does not.
  • coverage, the lowest and highest raw value among the outcomes the fit rests on (the queue labels when fitted_on is queue): p for bool, the probability of the most probable option or level for choice and score.
  • warnings, when the outcomes look like a sample someone reviewed, each with a code and a message:
    • above_thresholds: more than 80% of the outcomes are on answers that meet one of the judgment’s thresholds.
    • high_band: for bool, more than 80% of the outcomes are on answers with p of 0.7 or more.
    Either means the fit is reliable only where the outcomes are. Label a random sample of documents too, with the labelling queue; see outcomes from a reviewed sample.
  • raw and calibrated, the same metrics before and after calibration, on the outcomes the fit was made on. With raw_better, calibrated shows what the fit would have given.
    • accuracy. A bool answer counts as true when p is at least 0.5. For choice and score, the most probable option or level must match the outcome exactly.
    • expected_calibration_error, how far the stated probabilities are from how often things came true. Lower is better.
    • log_loss. Lower is better.
    • reliability, 10 equal-width bins such as {"lower": 0.8, "upper": 0.9, "count": 137, "mean_predicted": 0.845, "observed": 0.883}. A well calibrated judgment has observed close to mean_predicted in every bin. For choice and score, each answer is binned by the probability of its most probable option or level. Empty bins have null means.
    • mean_level_distance, for score judgments only: how many levels the most probable level is from the observed one, on average.
  • tenant and tenants, for templates: on a tenant, whether its answers read its own fit blended with the template’s pool, the pool, or raw, and why; on the template, how many tenants read each. Both are null for a namespace’s own judgment.
The report is computed when you ask for it, so the raw metrics include outcomes you posted a minute ago. The fit itself changes when it is refitted overnight.

Read the lift

lift says how much calibration has cut the judgment’s error since the loop started. The Calibration tab leads with it: one sentence, such as “Calibration cut this judgment’s error by 20% since it started, on 1,432 outcomes”, and on Team and Scale a chart of the error after each nightly fit, raw beside what answers read, with what the newest fit rests on. On Developer the tab shows the sentence as a preview. The API returns lift on every plan.
  • history, one point per nightly fit whose numbers changed, oldest first, at most one a day per epoch, kept for 400 days: fitted_at, the outcomes it rests on and outcomes_by_source, fitted_on, its source (replay for a fit made from a replay after the model changed), stage, whether it was applied, and its held-out log loss and ECE, raw and calibrated. Every number is on outcomes the fit did not see.
  • headline, from the current epoch’s newest point:
    • error_reduction is the share of the raw answers’ held-out log loss that what answers read removes: (raw_log_loss − applied_log_loss) / raw_log_loss. 0.204 reads “cut its error by 20%”. Both are over the same outcomes, so it compares like with like.
    • interval is its 95% interval. When lower is 0 or below, the cut is not clear yet; more outcomes narrow it.
    • applied is false when the raw answers did better on held-out outcomes. Answers then read raw, and error_reduction is 0: the engine is already well calibrated on your data.
    • since is the epoch’s first fit, and outcomes what the newest fit rests on.
  • withheld, why there is no headline yet, in the same form as not_fitted, with its reason chosen the same way: fewer than 100 outcomes, or 20 of a kind, in the current epoch, or awaiting_fit when the next nightly fit adds its first point. There is never a headline below the calibration minimums.
  • model_changed, after the engine’s model changed (epochs): the earlier epoch (from), the current one (to) and the day it began (on). A headline never mixes epochs. The new epoch starts its own from its own first fit, while the chart keeps the old epoch’s points beside it.
lift is null for a composite judgment, whose p comes from its combiner, and on a template’s tenant: the template’s report has the pool’s lift. A new judgment version starts with an empty history. Precision and recall at your named thresholds are not part of the lift. Thresholds read the raw numbers, which calibration never changes, so they measure the engine. The recommender gives them.

Epochs

An epoch is the engine_version recorded on an evaluation. An exact engine version is one epoch. Jev current starts a new epoch each time its behaviour changes, labelled like current+2026-09-24.1; see the engine page. Calibration is fitted per judgment version and per epoch, because a different model needs a different fit. After a drift, the new epoch starts with no outcomes of its own, and is re-fitted from the ones you already posted.

After the model changes

Within minutes of a change to Jev current, each judgment with outcomes is re-fitted for the new model from a replay: the stored contexts of the answers your most recent outcomes labelled, up to a bounded number, are judged again by the same engine, and each new answer with its old outcome is a sample of the new epoch. You don’t need to do anything, and it runs on every plan. Calibration explains what is sent and why it is safe. What you see:
  • The Calibration tab says, for the new epoch, “Re-fitting after the model changed on 2026-10-02: 400 of 812 outcomes replayed so far.” while it runs, then “Re-fitted on 812 replayed outcomes after the model changed on 2026-10-02.”
  • In the report, the new epoch has fit_source: "replay" and its replay progress; its own outcomes stay 0 until you post new ones.
  • Until the replay’s fit is made, usually within the hour, answers under the new epoch use the previous epoch’s calibration with "from_previous_epoch": true, for 30 days at most.
  • Replays show in a document’s evaluation history with replay_of, naming the evaluation they replayed. They never change an answer and are never billed.
  • Keep posting outcomes as usual. Once the new epoch has 100 of its own, with 20 of each kind, its fit rests on them alone and fit_source becomes outcomes.
If a replay stops early (status: "stopped"), the new epoch is fitted on what was replayed if that meets the minimums, and otherwise from its own outcomes as they arrive: reason: "cost_cap" means your organization reached this month’s replay limit: if the epoch still has no fit next month, the replay runs again on the outcomes it had not replayed. superseded means the model changed again, so the newer epoch has a replay of its own. Re-fits come first: recipe tuning uses only part of the monthly limit, so re-fits always have the rest.

Calibrated answers

Once a judgment’s fit rests on enough outcomes and beats the raw numbers on held-out outcomes, every answer carries a calibrated object beside the raw numbers:
  • It has the answer type’s own fields: p for bool; value, dist and escape_p for choice; score and dist for score. It adds stage (early or full, as in the report), outcomes (what the fit rests on), from_previous_epoch and extrapolated.
  • extrapolated is true when the answer’s raw value is outside the range the fit was made on. The calibrated number is then a guess; post outcomes for documents like it.
  • It never replaces p, dist or score, which stay the engine’s raw output.
  • It is absent while there is no fit, not null.
  • It is computed when the answer is read, from the current fit for the answer’s version and epoch. A refit changes it without recomputing anything.
  • Thresholds, filters and ranking use the raw fields.

Change a question safely with a shadow report

To change a judgment’s question, criteria, context or engine, create version n+1 by posting the definition again under the same name. It stays inactive. Then activate it with POST /namespaces/{ns}/judgments/needs_escalation/activate:
When version 4’s engine or definition differs from the active one, the response is 202 with a shadow job in awaiting_confirm. The job judges a random sample of 1,000 documents under version 4, or every document if there are fewer. While it samples, report is null and progress.documents_done counts the sampled documents. When it is done, the report compares the two versions on the same documents:
  • current is the active version’s answers, and candidate the new version’s results. For bool and score each side has a mean and a 10-bin histogram of p or score, left out of the example above. For choice each side has dist, the mean probability of each option, so you can read the shift option by option.
  • threshold_flips counts, for each named threshold, the documents that would go from false to true and from true to false. The new side uses the thresholds that will apply once the version is active.
  • recompute estimates backfilling every document under the new version: documents, tokens, judgments counted by size class, cost and duration.
Then decide:
  • POST /jobs/{id}/confirm (db.jobs.confirm(id)) switches to the new version. You can confirm before the report is done. The job shows running, then done about a second later, once the switch is committed.
  • POST /jobs/{id}/cancel leaves the active version as it is.
  • {"version": 4, "force": true} on activate skips the report and switches at once. So does activate: true when you create a version. A composite judgment is the exception: its shadow job fits it on your labels, so it always runs.
Shadow evaluations are written to the evaluation log with shadow: true. They never produce answers, never count toward calibration, and are not billed. After the switch, existing answers keep their old judgment_version until their documents change. To recompute them all, run a backfill; recompute is its estimate. Thresholds you gave with the new version replace the current ones when it becomes active. In the dashboard, the judgment page’s Overview shows the report, with Confirm, Cancel and Force.

Let your outcomes tune the recipe

The context recipe decides what every answer costs, so Brussle uses your outcomes to test smaller ones. Once a judgment version’s calibration is fitted (100 outcomes, 20 of each kind), a recipe-tuning run tests cheaper variants of the recipe against your labelled outcomes. A smaller recipe saves only where it moves answers into a smaller size class: trimming tokens within a class costs the same. The variants use the contexts already stored with your labelled answers’ evaluations, sent to the same engine on the same terms, and are compared with the current recipe on the same outcomes. It runs after the nightly fit, then at most every 30 days, and again after the engine’s model changes; it is never billed. It uses only part of your organization’s monthly limit on replays, so a re-fit after a model change always has the rest. GET /namespaces/{ns}/judgments/{name}/recipe-suggestion (ns.judgments.recipe_suggestion(name) in Python, ns.judgments.recipeSuggestion(name) in TypeScript) returns the last run’s result for the active version:
  • status: ready when there is a suggestion, none when the last run found none, running, or not_run.
  • reason and message, why there is no suggestion:
  • suggestion, when ready:
    • recipe, the whole new recipe, and removed, what changed, such as “the field state.internal_notes”.
    • definition, the exact body to post as a new version: your active version with only context changed. It gives no thresholds, so your current ones carry over when you activate it.
    • cost_per_answer_before and cost_per_answer_after: the mean judgments per answer on the labelled answers, each answer counted by its size class, with the current recipe and with this one. unit_reduction is the share it saves: 0.75 when answers that were all large (4.0) all become standard (1.0), and 0.5625 when only three quarters of them do (1.75).
    • held_out: log loss and accuracy with both recipes, each outcome scored with a calibration made without it; and interval, the 95% interval of the change in log loss (above 0 is worse). A recipe qualifies only when neither log loss nor accuracy is worse beyond noise. The suggestion is the one that qualifies with the largest unit_reduction: each recipe is measured on the labelled answers it could be tested on, so its reduction compares with the others’ and its cost per answer may not.
    • labels, the labelled answers both were scored on, and engine_version, the epoch both were judged in.
  • headline, one sentence, such as “A context recipe 75% cheaper per answer, moving answers to a smaller size class, with accuracy unchanged within noise on 312 labelled answers.”
  • computed_at and next_run_after, when the last run ended and the earliest the nightly run tries again.
To apply a suggestion, create the version and activate it through its shadow report, as for any change to a question:
Nothing is applied for you: the recipe decides what you send to the engine and what makes an answer re-run, so it stays your choice. POST /namespaces/{ns}/judgments/{name}/recipe-tuning (tune_recipe, tuneRecipe) runs this month’s tuning now instead of waiting for the nightly run, and returns the result as it stands, running until it ends. There is one run per version, engine epoch and calendar month, so asking again returns the same run. It returns invalid_request when the version cannot be tuned yet, with the reason. In the dashboard, the judgment’s Recipe tab shows the current recipe and the last run: its status and reason, and when a suggestion is ready, what it changes, the cost per answer before and after, the held-out log loss and accuracy with the interval, and the labels they rest on. Create version posts the suggestion’s definition as a new version, which is not active, and starts its shadow report on the Overview, where you confirm or cancel it. Run now starts this month’s run. Recipe tuning is part of the learning loop, on Team and Scale. On Developer, recipe-suggestion shows the headline and the numbers, with recipe, removed and definition withheld and plan_required beside them, and recipe-tuning returns plan_required.

Pick thresholds with the recommender

The recommender finds the threshold that meets a precision or recall target on your outcomes. It is part of the learning loop, on the Team plan and above; on Developer it returns plan_required (see plans). The calibration report and calibrated answers are on every plan.
target is precision:<x> or recall:<x>. For a choice judgment, add the option, such as &option=fraud; the threshold is then on the raw dist[fraud], and a choice without option is refused with invalid_request. A choice that chooses among candidates (options.from) is the exception: its threshold is on raw escape_p, the unmatched queue, and option is none_of_the_above or left out (any other option is refused). score judgments get no recommendation: the call is refused with invalid_request.
The example shows the first of the curve’s 101 points. The response is always 200, and status says which of three it is:
  • A precision target gets the lowest threshold that meets it, which keeps the most recall.
  • A recall target gets the highest threshold that meets it, which keeps the most precision.
  • For a choice among candidates it is the other way round, because a document’s pick is taken when its escape_p is below the threshold. Precision is the share of taken picks that are right, and recall the share of outcomes naming a match whose pick was taken and right. So a precision target gets the highest threshold that meets it, and a recall target the lowest. Apply it as {"unmatched": {"value": "none_of_the_above", "gte": threshold}}.
  • The curve has 101 points, thresholds 0.00 to 1.00 in steps of 0.01, each with its precision and recall. The recommended threshold is one of them. Precision is null where no answer reaches the threshold.
  • A point counts what a threshold selects. It compares each answer’s number as the answer shows it, the same way a named threshold and a query filter do, so a threshold you set from a point selects exactly the answers the point counted, including answers that sit on it.
  • The outcomes are the current epoch’s. They need at least 100, with 20 positive and 20 negative: for a choice, 20 whose outcome is the option and 20 whose outcome is another. When the current epoch has too few, the previous epoch’s are used and from_previous_epoch is true.
With outcomes on one side only, the recommender does not guess:
reason is too_few_outcomes, too_few_positives or too_few_negatives, chosen as in the report’s not_fitted: a missing kind of outcome first, then the total, then 20 of a kind. Nothing changes until you apply it. Thresholds are settings, changed with PATCH /namespaces/{ns}/judgments/{name}. The thresholds you send replace the whole set, so include the ones you want to keep:
For a choice, a threshold names its option: {"fraud": {"value": "fraud", "gte": 0.62}}. {} removes every threshold. The change applies at the next read to every answer, including answers already computed. Nothing is recomputed, no version is created, and the change is recorded in your audit log. In the dashboard, the judgment’s Thresholds tab applies a recommendation in one click, after you confirm. It keeps your other thresholds. On Developer, the Thresholds tab shows a preview instead. For a yes-or-no judgment whose calibration report has enough outcomes (the same 100, with 20 of each kind, from the same epoch), it shows the threshold the recommender would pick and its precision and recall. It works these out from the report’s reliability bins, so it looks at thresholds in steps of 0.1 and gives no interval: its pick can be a little above the recommender’s own.

Real-world outcomes with a horizon

Labelled examples say what the answer should have been at the time. Some questions are predictions, and the outcome arrives later. For “will this customer churn within 30 days?”, give the judgment a horizon:
An answer to this judgment says “this customer churns within 30 days of now”. Post the outcome when it happens:
Each outcome labels the prediction whose 30-day window it falls in. A customer’s first answer opens a window that runs 30 days from when the answer was made. Answers made while it is open do not open windows of their own, and the first answer after it closes opens the next. So a churn on day 12 counts for the answer made on day 0, and a customer judged every day still counts once per 30 days. See predictions with a horizon. horizon is part of the definition, a whole number and a unit (s, m, h or d). It defaults to 0s, and changing it creates a new version. If the cancellation already reaches you as a write, such as an account’s attributes.status becoming cancelled, an outcome rule can post it for you. Calibration needs the customers who did not churn too. Post false for them once the window has passed, whenever suits you: a false observed after a window closed answers that window, unless a churn was already posted in it. Or, when every churn is recorded, turn on implicit negatives, so that each window that closes with no outcome counts as false:
It is a setting, sent with PATCH /namespaces/{ns}/judgments/{name}, so it creates no version. It is only for bool judgments with a horizon. Leave it off when a missing outcome does not mean “no”; see implicit negatives.

Outcomes from your own data

Most outcomes already reach you as ordinary writes: an account’s status becomes cancelled, a ticket closes with the team that finally handled it, a moderator reverses a removal. Outcome rules turn those writes into outcomes, so labels arrive without anyone posting them. A rule watches one path on the judged document, and when a write changes it to the value you name, that write is an outcome. Churn. “Will this account cancel within 30 days?”, with a 30-day horizon. The cancellation is a true, and with implicit negatives on, every window that closes without one is a false:
Triage. A choice judgment that routes tickets to a team. When a ticket closes, the team it closed with is the right answer. from takes the outcome from the document after the write:
Moderation. A bool judgment, “does this post break the rules?”, whose answers drive removals. A moderator reversing a removal says the answer was wrong, and one upholding it says it was right:
How a rule fires:
  • when.path is attributes.<name> or state.<key>[.<key>...] on the judged document itself. For an entity judgment, that is the entity document, such as the account.
  • becomes names one value, and in a list of up to 16. A rule fires on the write that changes the value to it, or into the list, from anything else. A write that leaves it there fires nothing, so an unrelated edit, a repeated sync or a retried request adds no second outcome. A delete never fires. A write that creates the document fires if it is created with the value.
  • value is the outcome: true or false for a bool (the only form a bool takes), an option for a choice, a level for a score. from (choice and score) takes the value at that path after the write. A value there that is not an option or level, or a missing path, is kept but adds no sample: the report counts it in unmatched_outcomes.does_not_fit, which is how to spot a rule pointing at the wrong field.
  • A document whose value goes back and forth fires each time it arrives, so a status cancelled, reactivated and cancelled again gives two outcomes. With a horizon, two in one prediction window still count once, and windows opened after the first true count for nothing (after the event).
How derived outcomes are counted:
  • They join like posted ones, and the report’s outcomes_by_source counts them as rule, beside posted, queue and implicit.
  • A posted outcome wins. When you post an outcome and a rule derives one in the same prediction window, yours labels it. Use this to correct a rule by hand.
  • With the default 0s horizon, a derived outcome labels the answer to the document as it was just before the write that fired it. The write that closes a ticket reveals the answer; it is not what the judgment was predicting.
  • The outcome’s observed_at is when the write was committed.
Limits:
  • Outcome rules are part of the learning loop, on the Team plan and above. Below it, a PATCH that sets rules returns plan_required; clearing them ("rules": []) always works. After a downgrade takes effect, rules keep running, but the outcomes they derive from then on are left out of calibration and counted in the report’s rule_outcomes_paused, which the dashboard’s Calibration tab shows. Outcomes collected before are kept and used, and upgrading again makes new ones count. See plans.
  • Up to 4 rules per judgment, each on one path. outcomes is replaced whole by each PATCH, so send implicit_negatives and every rule together.
  • Rules apply from when you set them; earlier writes are not replayed. To label the past, post outcomes for it.
  • Rules read the judged document’s own writes. A write to another document, such as a ticket that should label its account, is not an outcome for the account; post it.
  • A namespace that inherits a judgment from a template cannot set rules, and neither can the template’s prefix path. Set them on a namespace’s own judgment.
  • A rule that does not fit the judgment’s type, such as from on a bool, or more than 4 rules, is refused with invalid_request, which says which rule and why.