Skip to main content
A judgment is a definition attached to a namespace. It says what Brussle should know about every document in it. A judgment defined on a prefix such as acme/prod/* is a template, inherited by every namespace under it.

Types

A bool answer has no value. Booleans come only from named thresholds.

How far to trust each type

The types are not equally reliable. On our conformance set (see the engine page for the latest run):
  • bool is the strongest. It is the most accurate type and its p is close to calibrated out of the box. If a question can be yes or no, ask it that way.
  • score is right to within one level far more often than exactly. About a third of answers land on the exact level, but about three quarters are within one. Read score as a position on the scale, set thresholds between levels (such as 3.5) rather than on them, and don’t rely on the difference between neighbouring levels.
  • choice picks the right option about as often as bool, but its probabilities are overconfident. dist often puts 1.0 on one option and 0 on the rest, even when the pick is wrong. Use value as the answer, and post outcomes and read calibrated before you act on the probabilities themselves; see calibration.
A bool judgment can instead ask 2 to 8 narrow yes/no parts (1 to 8 beside features) and combine them with weights fitted on your labels. It is measured on those labels before it answers; see composite judgments.

Choose among candidates

A choice can take its options from a blocking relation: the documents that share a key with the judged one, such as the open invoices a payment could settle. "options": {"from": "candidates", "label": ["state.number", "state.amount"]} makes one option per candidate the relation reads, in its order, newest first: value is the candidate’s id, and its label fields describe it. none_of_the_above follows, meaning “none of these: no match”.
  • At most 10 candidates. The relation’s last_n is at most the candidate limit, and the candidates are judged together in one call.
  • One match per document. value names one candidate, or none_of_the_above. To find every duplicate of a record, use a pair document per candidate pair.
  • escape_p is the probability that nothing matches, and a threshold on it, the only kind this choice takes, is your unmatched queue. dist is a shortlist, not a ranking. A document with no candidates is answered none_of_the_above with escape_p 1, with no engine call and nothing billed.
  • Not readable, not discoverable. No relation can read its answers, its horizon is 0s, and discovery does not apply: its escape means “no candidate”, not “a missing option”.
Match one record to another covers keys, outcomes and cost.

Starters

You don’t have to write a definition from nothing. Starter judgments are ready-made ones for ticket triage, moderation, churn risk, lead quality and review queues, each asked as a bool where the question can be yes or no. Create one with from_starter and the paths of your own documents; with "dry_run": true the response shows the definition it would make, its cost per answer and its backfill estimate, and creates nothing. dry_run works on any definition, not only starters. Every create returns the version it made in definition. See start from a starter.

Thresholds

Thresholds are named numbers evaluated when an answer is read, and returned as booleans under answers.{name}.thresholds. For a bool, a threshold is true when p is at least the number. For a score, it is true when score is at least the number. For a choice, it names an option and a minimum probability: "escalate": {"value": "fraud", "gte": 0.7}. They are always evaluated against the raw fields, never the calibrated ones. Thresholds are settings, not part of a version. Change them with PATCH /namespaces/{ns}/judgments/{name}:
The thresholds you send replace the whole set, so include the ones you want to keep. {} removes them all. The change applies at the next read to every answer, including answers already computed, with no recompute and no new version, and is recorded in your audit log. Thresholds given when you create a version replace the current ones when that version becomes active. To pick a threshold from your own outcomes, see the recommender. When another judgment’s relation reads a threshold by name, changing it moves what that reader sees with no answer to re-judge it by. So the PATCH returns each reader’s cost as downstream and changes nothing until you send it again with confirm: true; see changing a threshold a reader reads.

Escape alert and discovery

A fixed-option choice’s escape option is where a new kind of complaint or abuse shows up first. Two things watch it for you, and neither changes the judgment on its own:
  • escape_alert is a setting: the share of the last 7 days’ answers that were none_of_the_above, above which you want to know. When the share rises above it, GET on the judgment lists the taxonomy_drift warning and the events feed records one judgment.taxonomy_drift. It needs 20 answers in the window, and is off (null) by default.
  • Discovery. POST /namespaces/{ns}/judgments/{name}/discover starts a job that reviews recent escaped documents and proposes the options they have in common, checked against those documents before you see them. It proposes a new version’s options and creates nothing: you review and create the version. It sends a sample of them to a third-party LLM provider, so an admin turns on suggestions first.
See discover new options.

Engine

Every judgment names one engine version, such as {"name": "jev", "version": "current"}. current runs whichever Jev model the provider serves now, and each answer records the epoch of model behaviour that produced it (engine_version, such as current+2026-09-24.1). Nothing picks an engine for you, unless you set a namespace’s default_engine. The version must be active in the registry; see engines.

Context recipe

context says which parts of the document the engine sees. It is the main lever on both cost and accuracy; see writing a context recipe. Every definition needs one: a create that leaves it out is refused with a recipe suggested from your documents, which you can accept with "use_context": "suggested", and "use_context": "whole_state" sends the whole state instead. See when you leave out context. A recipe can also read more than the judged document: the documents that point at it, such as an account’s tickets, the one it points at, such as an order line’s product, other judgments’ answers for them, candidates that share a key, or the document as the judgment last saw it. See relations.

Versions

Versions are immutable. Posting a definition with an existing name creates version n+1. The first version of a name is active when it is created, except a composite, which never is. A later version stays inactive until you activate it, or you pass activate: true. POST /namespaces/{ns}/judgments/{name}/activate with {"version": 4} makes version 4 the one new evaluations use. Existing answers keep their judgment_version until they are recomputed. If the new version’s engine or definition differs from the active one, the call returns 202 with a shadow job instead. The job judges a sample of 1,000 documents under the new version and reports the distribution shift, the threshold flips and the cost of recomputing everything. POST /jobs/{id}/confirm makes the switch. "force": true skips the report and switches at once, and so does activate: true on create. A composite is the exception: force still runs its shadow job, and activate: true on create is refused with invalid_request. See changing a question safely. Freshness policy, debounce, fan-out, thresholds, escape_alert and outcomes (implicit negatives and outcome rules) are settings, not part of the definition. Changing them with PATCH /namespaces/{ns}/judgments/{name} creates no version. freshness in a create sets the judgment’s freshness with its first version only: a later version keeps the settings the judgment has, and its response warns when the freshness you sent differs. So creating a version never switches a judgment to on_change or drops a debounce you tuned. See choosing a freshness policy, for implicit negatives on a bool judgment with a horizon, calibration, and for rules that derive outcomes from your writes, outcomes from your own data. horizon is part of the definition: a duration such as "30d", 0s by default. It says how far ahead a judgment predicts: an answer says the event happens within the horizon after it was made, and each outcome labels the prediction whose window it falls in. Changing it creates a version. DELETE /namespaces/{ns}/judgments/{name} deactivates a judgment. Its answers and evaluations are kept. While another judgment’s relation reads its answers, GET lists those readers, and deactivating it is refused with conflict naming them. A namespace can have 100 active judgments.

Backfill

A new judgment has no answers for existing documents. POST /namespaces/{ns}/judgments/{name}/backfill without confirm returns an estimate and does nothing:
cost_usd is what the backfill costs you: its judgment_units, the judgments it would bill counted by size class, at the prices of the tiers your organization will be in, counting what it has already used this billing period, the same on every engine. Here each document’s context and question come to 2,000 tokens, so each is a standard judgment and counts as 1, and all of them fall in the first tier in a month with no other judging. duration_s is the expected wall-clock time if nothing else is judged meanwhile. Other judging comes first and makes it longer, so read it as the least to expect. Here it is about 98 days on Jev, which is why it is shown next to the cost. With "confirm": true the backfill starts as a job, which you can follow, pause, resume or cancel under /jobs/{id}. Backfills respect the namespace budget, and a filters expression limits a backfill to matching documents. A backfill judges only the documents in its scope that need it: those within the judgment’s applies_to that match filters, except the ones whose answer from the active version already succeeded for their current revision. Its estimate counts the same documents, as they are when you ask, so a backfill after a finished one is quoted only for what was written since, and one with nothing left to judge is quoted nothing. The job’s progress.documents_done counts the documents judged so far, so it climbs to the estimate’s documents. The job reads done once it has passed the last document.