The learning loop by plan
Calibration from the outcomes you post is on every plan: the report, and calibrated values on every answer. The rest of the learning loop is on Team and Scale, and Developer sees a preview of it in the dashboard:
A call your plan doesn’t include returns
plan_required (HTTP 402). Wherever such a feature appears, the dashboard says what it does and which plan has it, with an upgrade for owners and a note to ask an owner for everyone else. See plans.
Post labelled examples
A labelled example is an outcome with the judgment’s defaulthorizon of 0s. Write the documents, then post their labels. Each label is joined to the evaluation of the document revision that was current at observed_at, so a label posted before that evaluation finishes still joins it. A label observed while its document is deleted joins nothing, because no revision was current. A document written again remembers its id’s last 8 deletions, so a label posted later for a time it was deleted joins nothing too. This needs the deletion to still be recorded when the label is posted or read, or when the id is written again.
Post them to POST /namespaces/{ns}/outcomes:
value is what the answer should have been:
- Post
falseoutcomes as well astrueones. Calibration and the recommender need at least 20 of each (for achoiceorscore, two values with 20 each). See calibration. - Label a random sample of documents too, not only the ones a person already reviewed, so the outcomes cover the whole range of answers. The labelling queue picks one for you.
- A value that does not fit the judgment’s type is
invalid_request. An unknown judgment isnot_found. - An outcome that joins no evaluation, such as one for a document never judged, is kept but adds nothing. The calibration report counts it in
unmatched_outcomes, with the reason. - Outcomes are append-only. Retries are safe: the same document, judgment, value and
observed_atis stored as one outcome. - Outcomes belong to a namespace. Post them on the namespace’s own path, never on a template prefix.
Upload labels as a CSV
On the dashboard, open the judgment and go to its Labels tab. Upload a CSV with the headerdocument_id,value and an optional observed_at column in RFC 3339:
observed_at uses the time of the upload.
Label a random sample
The labelling queue hands out documents to label that nobody chose, spread evenly across the probability range, so the labels are unbiased. Once an epoch has enough of them, its calibration rests on them alone, each weighted back to how common its part of the range is in real traffic. Calibration explains why. On the dashboard, open the judgment and go to its Label tab. It shows one document at a time with the judgment’s question and a button per answer: yes or no for abool, an option for a choice, a level for a score, or skip. Each answer is posted as an outcome at once, and the tab shows how many of each band are labelled.
Through the API, draw documents with POST /namespaces/{ns}/judgments/{name}/labelling-queue:
count is 1 to 200 (default 50). The response lists the items, lowest band first:
bandis one of 10 equal bands of the raw value, andprobabilitythe raw value:pfor abool, the most probable option’s or level’s probability for achoiceorscore, whosevalueit is. The count is split evenly over the bands that have documents, so a band holding 2% of the answers gets as many items as one holding 80%.documentis what the labeller needs to see.evaluation_idleads to the exact context the engine read, atGET /namespaces/{ns}/evaluations/{id}.- The items are leased until
lease_expires_at, 7 days: nobody else is given them. A document labelled from the queue leaves it until its answer changes; one not labelled in time goes back. seed(default 0) fixes the random order within each band. The same documents, labels and seed give the same items.
queue_item_id:
- The outcome counts as a queue label:
outcomes_by_source.queuein the calibration report. Once an epoch has 100 queue labels, 20 of each kind (for achoiceorscore, two values with 20 each), itsfitted_onbecomesqueue, and the next fit uses the queue labels only, weighted by band. - A
queue_item_idwhose lease has run out, 7 days after the draw, isconflict; one whose draw is more than 30 days old, or that names another document, isinvalid_request. Draw new items. - Labelling the same item again, to retry or to correct a label, replaces the earlier label: an answer counts only its latest queue label.
- Label what the document shows. Skipping an item is fine: it goes back to the queue when its lease ends.
GET on the same path previews the queue without handing anything out: for each band, the answers, and how many are labelled, leased and available. It is ns.judgments.labellingQueue(name) in TypeScript and labelling_queue(name) in Python.
Plans. Drawing is part of the learning loop, on the Team plan and above: below it, a draw is plan_required. The preview is on every plan, so on Developer the tab shows what the queue would contain.
Horizons. A judgment with a non-zero horizon has no queue (invalid_request). Its answers predict what happens next, which nobody can label by looking at the document today. Use real-world outcomes and implicit negatives instead.
Read the calibration report
GET /namespaces/{ns}/judgments/{name}/calibration returns the report for the active version. It is ns.judgments.calibration(name) in both SDKs, and the Calibration tab on the judgment’s page in the dashboard.
reliability curve under raw and calibrated, and all but the first entry of lift.history.
At the top, beside the total outcomes:
-
outcomes_by_source, the outcomes by where they came from:posted, derived by outcome rules (rule), labels from the labelling queue (queue), or implicit negatives (implicit). They add up tooutcomes. -
rule_outcomes_paused, rule outcomes left out because your plan didn’t include outcome rules when they were observed (after a downgrade to Developer takes effect). They are kept, and aren’t inoutcomes. It stays 0 while your plan has included outcome rules throughout. -
unmatched_outcomes, the outcomes, posted or derived, that add nothing, by reason:They are kept, and count if a later answer takes them. -
censored_predictions, with implicit negatives on, the prediction windows with no outcome that are not counted asfalse:window_open(not closed yet),document_deleted(the document was deleted before the window closed) andafter_positive(the window opened after atruefor the document).
-
outcomes, the outcomes joined to this epoch’s evaluations,outcomes_by_source, the same split as at the top,rule_outcomes_paused, as at the top, andimplicit_negatives, how many of them are implicit negatives (0 unless the setting is on). Outcomes replayed into the epoch after the model changed are not here but inreplay.outcomes, so each outcome counts once in the report. -
fitted_on, what the fit rests on:queueonce the epoch’s queue labels alone meet the minimums (100, with 20 of each kind), each weighted back to real traffic; otherwiseall, every outcome counted once. -
fit_source,replaywhile the epoch’s fit rests on outcomes replayed after the model changed, andoutcomesotherwise.nullfor a composite. -
replay, the replay that re-fitted the epoch after the model changed, ornullwhen none ran for it: itsstatus(running,doneorstopped, with areason:cost_caporsuperseded), the epoch its outcomes came from (from_engine_version), how many itplannedand hasdone, the replayedoutcomesthe fit rests on now,started_atandfinished_at. -
stage, how far along the epoch’s calibration is, andfitted_at, when it was fitted:Calibration is refitted every night. -
not_fitted, why the epoch’s answers have no calibration of their own, ornullwhen they do.reasonis one of:When more than one applies,reasonnames the first of: a kind of outcome missing altogether (alltrueistoo_few_negatives, allfalseistoo_few_positives, and achoiceorscorewith outcomes of one value istoo_few_values), then fewer than 100 outcomes, then fewer than 20 of a kind. So 50 outcomes, alltrue, readtoo_few_negatives, nottoo_few_outcomes.messagesays what to post, such as “Only positive outcomes so far: post outcomes for documents where it did not happen.” -
held_out, the test a fit must pass before answers use it: the raw answers’ log loss, and the log loss of fits made without the outcome they score. The fit is applied only whencalibrated_log_lossis lower.nullwhen nothing is fitted. Fits also giveraw_eceandcalibrated_eceon the same held-out outcomes, andstandard_error, how sure the difference in log loss is; a template tenant’s own test does not. -
coverage, the lowest and highest raw value among the outcomes the fit rests on (the queue labels whenfitted_onisqueue):pforbool, the probability of the most probable option or level forchoiceandscore. -
warnings, when the outcomes look like a sample someone reviewed, each with acodeand amessage:above_thresholds: more than 80% of the outcomes are on answers that meet one of the judgment’s thresholds.high_band: forbool, more than 80% of the outcomes are on answers withpof 0.7 or more.
-
rawandcalibrated, the same metrics before and after calibration, on the outcomes the fit was made on. Withraw_better,calibratedshows what the fit would have given.accuracy. Aboolanswer counts astruewhenpis at least 0.5. Forchoiceandscore, the most probable option or level must match the outcome exactly.expected_calibration_error, how far the stated probabilities are from how often things came true. Lower is better.log_loss. Lower is better.reliability, 10 equal-width bins such as{"lower": 0.8, "upper": 0.9, "count": 137, "mean_predicted": 0.845, "observed": 0.883}. A well calibrated judgment hasobservedclose tomean_predictedin every bin. Forchoiceandscore, each answer is binned by the probability of its most probable option or level. Empty bins havenullmeans.mean_level_distance, forscorejudgments only: how many levels the most probable level is from the observed one, on average.
-
tenantandtenants, for templates: on a tenant, whether its answers read its own fit blended with the template’s pool, the pool, or raw, and why; on the template, how many tenants read each. Both arenullfor a namespace’s own judgment.
Read the lift
lift says how much calibration has cut the judgment’s error since the loop started. The Calibration tab leads with it: one sentence, such as “Calibration cut this judgment’s error by 20% since it started, on 1,432 outcomes”, and on Team and Scale a chart of the error after each nightly fit, raw beside what answers read, with what the newest fit rests on. On Developer the tab shows the sentence as a preview. The API returns lift on every plan.
history, one point per nightly fit whose numbers changed, oldest first, at most one a day per epoch, kept for 400 days:fitted_at, theoutcomesit rests on andoutcomes_by_source,fitted_on, itssource(replayfor a fit made from a replay after the model changed),stage, whether it wasapplied, and its held-out log loss and ECE, raw and calibrated. Every number is on outcomes the fit did not see.headline, from the current epoch’s newest point:error_reductionis the share of the raw answers’ held-out log loss that what answers read removes:(raw_log_loss − applied_log_loss) / raw_log_loss. 0.204 reads “cut its error by 20%”. Both are over the same outcomes, so it compares like with like.intervalis its 95% interval. Whenloweris 0 or below, the cut is not clear yet; more outcomes narrow it.appliedisfalsewhen the raw answers did better on held-out outcomes. Answers then read raw, anderror_reductionis 0: the engine is already well calibrated on your data.sinceis the epoch’s first fit, andoutcomeswhat the newest fit rests on.
withheld, why there is no headline yet, in the same form asnot_fitted, with its reason chosen the same way: fewer than 100 outcomes, or 20 of a kind, in the current epoch, orawaiting_fitwhen the next nightly fit adds its first point. There is never a headline below the calibration minimums.model_changed, after the engine’s model changed (epochs): the earlier epoch (from), the current one (to) and the day it began (on). A headline never mixes epochs. The new epoch starts its own from its own first fit, while the chart keeps the old epoch’s points beside it.
lift is null for a composite judgment, whose p comes from its combiner, and on a template’s tenant: the template’s report has the pool’s lift. A new judgment version starts with an empty history.
Precision and recall at your named thresholds are not part of the lift. Thresholds read the raw numbers, which calibration never changes, so they measure the engine. The recommender gives them.
Epochs
An epoch is theengine_version recorded on an evaluation. An exact engine version is one epoch. Jev current starts a new epoch each time its behaviour changes, labelled like current+2026-09-24.1; see the engine page. Calibration is fitted per judgment version and per epoch, because a different model needs a different fit.
After a drift, the new epoch starts with no outcomes of its own, and is re-fitted from the ones you already posted.
After the model changes
Within minutes of a change to Jevcurrent, each judgment with outcomes is re-fitted for the new model from a replay: the stored contexts of the answers your most recent outcomes labelled, up to a bounded number, are judged again by the same engine, and each new answer with its old outcome is a sample of the new epoch. You don’t need to do anything, and it runs on every plan. Calibration explains what is sent and why it is safe.
What you see:
- The Calibration tab says, for the new epoch, “Re-fitting after the model changed on 2026-10-02: 400 of 812 outcomes replayed so far.” while it runs, then “Re-fitted on 812 replayed outcomes after the model changed on 2026-10-02.”
- In the report, the new epoch has
fit_source: "replay"and itsreplayprogress; its ownoutcomesstay 0 until you post new ones. - Until the replay’s fit is made, usually within the hour, answers under the new epoch use the previous epoch’s calibration with
"from_previous_epoch": true, for 30 days at most. - Replays show in a document’s evaluation history with
replay_of, naming the evaluation they replayed. They never change an answer and are never billed. - Keep posting outcomes as usual. Once the new epoch has 100 of its own, with 20 of each kind, its fit rests on them alone and
fit_sourcebecomesoutcomes.
status: "stopped"), the new epoch is fitted on what was replayed if that meets the minimums, and otherwise from its own outcomes as they arrive: reason: "cost_cap" means your organization reached this month’s replay limit: if the epoch still has no fit next month, the replay runs again on the outcomes it had not replayed. superseded means the model changed again, so the newer epoch has a replay of its own. Re-fits come first: recipe tuning uses only part of the monthly limit, so re-fits always have the rest.
Calibrated answers
Once a judgment’s fit rests on enough outcomes and beats the raw numbers on held-out outcomes, every answer carries acalibrated object beside the raw numbers:
- It has the answer type’s own fields:
pforbool;value,distandescape_pforchoice;scoreanddistforscore. It addsstage(earlyorfull, as in the report),outcomes(what the fit rests on),from_previous_epochandextrapolated. extrapolatedistruewhen the answer’s raw value is outside the range the fit was made on. The calibrated number is then a guess; post outcomes for documents like it.- It never replaces
p,distorscore, which stay the engine’s raw output. - It is absent while there is no fit, not
null. - It is computed when the answer is read, from the current fit for the answer’s version and epoch. A refit changes it without recomputing anything.
- Thresholds, filters and ranking use the raw fields.
Change a question safely with a shadow report
To change a judgment’s question, criteria, context or engine, create version n+1 by posting the definition again under the same name. It stays inactive. Then activate it withPOST /namespaces/{ns}/judgments/needs_escalation/activate:
202 with a shadow job in awaiting_confirm. The job judges a random sample of 1,000 documents under version 4, or every document if there are fewer. While it samples, report is null and progress.documents_done counts the sampled documents. When it is done, the report compares the two versions on the same documents:
currentis the active version’s answers, andcandidatethe new version’s results. Forboolandscoreeach side has ameanand a 10-binhistogramofporscore, left out of the example above. Forchoiceeach side hasdist, the mean probability of each option, so you can read the shift option by option.threshold_flipscounts, for each named threshold, the documents that would go from false to true and from true to false. The new side uses the thresholds that will apply once the version is active.recomputeestimates backfilling every document under the new version: documents, tokens, judgments counted by size class, cost and duration.
POST /jobs/{id}/confirm(db.jobs.confirm(id)) switches to the new version. You can confirm before the report is done. The job showsrunning, thendoneabout a second later, once the switch is committed.POST /jobs/{id}/cancelleaves the active version as it is.{"version": 4, "force": true}on activate skips the report and switches at once. So doesactivate: truewhen you create a version. A composite judgment is the exception: its shadow job fits it on your labels, so it always runs.
shadow: true. They never produce answers, never count toward calibration, and are not billed.
After the switch, existing answers keep their old judgment_version until their documents change. To recompute them all, run a backfill; recompute is its estimate. Thresholds you gave with the new version replace the current ones when it becomes active.
In the dashboard, the judgment page’s Overview shows the report, with Confirm, Cancel and Force.
Let your outcomes tune the recipe
The context recipe decides what every answer costs, so Brussle uses your outcomes to test smaller ones. Once a judgment version’s calibration is fitted (100 outcomes, 20 of each kind), a recipe-tuning run tests cheaper variants of the recipe against your labelled outcomes. A smaller recipe saves only where it moves answers into a smaller size class: trimming tokens within a class costs the same. The variants use the contexts already stored with your labelled answers’ evaluations, sent to the same engine on the same terms, and are compared with the current recipe on the same outcomes. It runs after the nightly fit, then at most every 30 days, and again after the engine’s model changes; it is never billed. It uses only part of your organization’s monthly limit on replays, so a re-fit after a model change always has the rest.GET /namespaces/{ns}/judgments/{name}/recipe-suggestion (ns.judgments.recipe_suggestion(name) in Python, ns.judgments.recipeSuggestion(name) in TypeScript) returns the last run’s result for the active version:
-
status:readywhen there is a suggestion,nonewhen the last run found none,running, ornot_run. -
reasonandmessage, why there is no suggestion: -
suggestion, whenready:recipe, the whole new recipe, andremoved, what changed, such as “the field state.internal_notes”.definition, the exact body to post as a new version: your active version with onlycontextchanged. It gives nothresholds, so your current ones carry over when you activate it.cost_per_answer_beforeandcost_per_answer_after: the mean judgments per answer on the labelled answers, each answer counted by its size class, with the current recipe and with this one.unit_reductionis the share it saves: 0.75 when answers that were all large (4.0) all become standard (1.0), and 0.5625 when only three quarters of them do (1.75).held_out: log loss and accuracy with both recipes, each outcome scored with a calibration made without it; andinterval, the 95% interval of the change in log loss (above 0 is worse). A recipe qualifies only when neither log loss nor accuracy is worse beyond noise. The suggestion is the one that qualifies with the largestunit_reduction: each recipe is measured on the labelled answers it could be tested on, so its reduction compares with the others’ and its cost per answer may not.labels, the labelled answers both were scored on, andengine_version, the epoch both were judged in.
-
headline, one sentence, such as “A context recipe 75% cheaper per answer, moving answers to a smaller size class, with accuracy unchanged within noise on 312 labelled answers.” -
computed_atandnext_run_after, when the last run ended and the earliest the nightly run tries again.
POST /namespaces/{ns}/judgments/{name}/recipe-tuning (tune_recipe, tuneRecipe) runs this month’s tuning now instead of waiting for the nightly run, and returns the result as it stands, running until it ends. There is one run per version, engine epoch and calendar month, so asking again returns the same run. It returns invalid_request when the version cannot be tuned yet, with the reason.
In the dashboard, the judgment’s Recipe tab shows the current recipe and the last run: its status and reason, and when a suggestion is ready, what it changes, the cost per answer before and after, the held-out log loss and accuracy with the interval, and the labels they rest on. Create version posts the suggestion’s definition as a new version, which is not active, and starts its shadow report on the Overview, where you confirm or cancel it. Run now starts this month’s run.
Recipe tuning is part of the learning loop, on Team and Scale. On Developer, recipe-suggestion shows the headline and the numbers, with recipe, removed and definition withheld and plan_required beside them, and recipe-tuning returns plan_required.
Pick thresholds with the recommender
The recommender finds the threshold that meets a precision or recall target on your outcomes. It is part of the learning loop, on the Team plan and above; on Developer it returnsplan_required (see plans). The calibration report and calibrated answers are on every plan.
target is precision:<x> or recall:<x>. For a choice judgment, add the option, such as &option=fraud; the threshold is then on the raw dist[fraud], and a choice without option is refused with invalid_request. A choice that chooses among candidates (options.from) is the exception: its threshold is on raw escape_p, the unmatched queue, and option is none_of_the_above or left out (any other option is refused). score judgments get no recommendation: the call is refused with invalid_request.
200, and status says which of three it is:
- A precision target gets the lowest threshold that meets it, which keeps the most recall.
- A recall target gets the highest threshold that meets it, which keeps the most precision.
- For a choice among candidates it is the other way round, because a document’s pick is taken when its
escape_pis below the threshold. Precision is the share of taken picks that are right, and recall the share of outcomes naming a match whose pick was taken and right. So a precision target gets the highest threshold that meets it, and a recall target the lowest. Apply it as{"unmatched": {"value": "none_of_the_above", "gte": threshold}}. - The curve has 101 points, thresholds 0.00 to 1.00 in steps of 0.01, each with its precision and recall. The recommended threshold is one of them. Precision is
nullwhere no answer reaches the threshold. - A point counts what a threshold selects. It compares each answer’s number as the answer shows it, the same way a named threshold and a query filter do, so a threshold you set from a point selects exactly the answers the point counted, including answers that sit on it.
- The outcomes are the current epoch’s. They need at least 100, with 20 positive and 20 negative: for a
choice, 20 whose outcome is the option and 20 whose outcome is another. When the current epoch has too few, the previous epoch’s are used andfrom_previous_epochistrue.
reason is too_few_outcomes, too_few_positives or too_few_negatives, chosen as in the report’s not_fitted: a missing kind of outcome first, then the total, then 20 of a kind.
Nothing changes until you apply it. Thresholds are settings, changed with PATCH /namespaces/{ns}/judgments/{name}. The thresholds you send replace the whole set, so include the ones you want to keep:
choice, a threshold names its option: {"fraud": {"value": "fraud", "gte": 0.62}}. {} removes every threshold.
The change applies at the next read to every answer, including answers already computed. Nothing is recomputed, no version is created, and the change is recorded in your audit log.
In the dashboard, the judgment’s Thresholds tab applies a recommendation in one click, after you confirm. It keeps your other thresholds.
On Developer, the Thresholds tab shows a preview instead. For a yes-or-no judgment whose calibration report has enough outcomes (the same 100, with 20 of each kind, from the same epoch), it shows the threshold the recommender would pick and its precision and recall. It works these out from the report’s reliability bins, so it looks at thresholds in steps of 0.1 and gives no interval: its pick can be a little above the recommender’s own.
Real-world outcomes with a horizon
Labelled examples say what the answer should have been at the time. Some questions are predictions, and the outcome arrives later. For “will this customer churn within 30 days?”, give the judgment ahorizon:
horizon is part of the definition, a whole number and a unit (s, m, h or d). It defaults to 0s, and changing it creates a new version.
If the cancellation already reaches you as a write, such as an account’s attributes.status becoming cancelled, an outcome rule can post it for you.
Calibration needs the customers who did not churn too. Post false for them once the window has passed, whenever suits you: a false observed after a window closed answers that window, unless a churn was already posted in it. Or, when every churn is recorded, turn on implicit negatives, so that each window that closes with no outcome counts as false:
PATCH /namespaces/{ns}/judgments/{name}, so it creates no version. It is only for bool judgments with a horizon. Leave it off when a missing outcome does not mean “no”; see implicit negatives.
Outcomes from your own data
Most outcomes already reach you as ordinary writes: an account’s status becomescancelled, a ticket closes with the team that finally handled it, a moderator reverses a removal. Outcome rules turn those writes into outcomes, so labels arrive without anyone posting them. A rule watches one path on the judged document, and when a write changes it to the value you name, that write is an outcome.
Churn. “Will this account cancel within 30 days?”, with a 30-day horizon. The cancellation is a true, and with implicit negatives on, every window that closes without one is a false:
choice judgment that routes tickets to a team. When a ticket closes, the team it closed with is the right answer. from takes the outcome from the document after the write:
bool judgment, “does this post break the rules?”, whose answers drive removals. A moderator reversing a removal says the answer was wrong, and one upholding it says it was right:
when.pathisattributes.<name>orstate.<key>[.<key>...]on the judged document itself. For an entity judgment, that is the entity document, such as the account.becomesnames one value, andina list of up to 16. A rule fires on the write that changes the value to it, or into the list, from anything else. A write that leaves it there fires nothing, so an unrelated edit, a repeated sync or a retried request adds no second outcome. A delete never fires. A write that creates the document fires if it is created with the value.valueis the outcome:trueorfalsefor abool(the only form abooltakes), an option for achoice, a level for ascore.from(choiceandscore) takes the value at that path after the write. A value there that is not an option or level, or a missing path, is kept but adds no sample: the report counts it inunmatched_outcomes.does_not_fit, which is how to spot a rule pointing at the wrong field.- A document whose value goes back and forth fires each time it arrives, so a status cancelled, reactivated and cancelled again gives two outcomes. With a horizon, two in one prediction window still count once, and windows opened after the first
truecount for nothing (after the event).
- They join like posted ones, and the report’s
outcomes_by_sourcecounts them asrule, besideposted,queueandimplicit. - A posted outcome wins. When you post an outcome and a rule derives one in the same prediction window, yours labels it. Use this to correct a rule by hand.
- With the default
0shorizon, a derived outcome labels the answer to the document as it was just before the write that fired it. The write that closes a ticket reveals the answer; it is not what the judgment was predicting. - The outcome’s
observed_atis when the write was committed.
- Outcome rules are part of the learning loop, on the Team plan and above. Below it, a
PATCHthat setsrulesreturnsplan_required; clearing them ("rules": []) always works. After a downgrade takes effect, rules keep running, but the outcomes they derive from then on are left out of calibration and counted in the report’srule_outcomes_paused, which the dashboard’s Calibration tab shows. Outcomes collected before are kept and used, and upgrading again makes new ones count. See plans. - Up to 4 rules per judgment, each on one path.
outcomesis replaced whole by eachPATCH, so sendimplicit_negativesand every rule together. - Rules apply from when you set them; earlier writes are not replayed. To label the past, post outcomes for it.
- Rules read the judged document’s own writes. A write to another document, such as a ticket that should label its account, is not an outcome for the account; post it.
- A namespace that inherits a judgment from a template cannot set rules, and neither can the template’s prefix path. Set them on a namespace’s own judgment.
- A rule that does not fit the judgment’s type, such as
fromon abool, or more than 4 rules, is refused withinvalid_request, which says which rule and why.