> ## Documentation Index
> Fetch the complete documentation index at: https://docs.brussle.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Judgments

> Typed questions kept answered for every document.

A judgment is a definition attached to a namespace. It says what Brussle should know about every document in it. A judgment defined on a prefix such as `acme/prod/*` is a [template](/guides/templates), inherited by every namespace under it.

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "name": "needs_escalation",
  "type": "bool",
  "question": "Does this ticket require a human to take over from the automated flow?",
  "criteria": "Escalate when the customer is at risk of leaving, mentions legal action, or the automation has failed twice.",
  "context": {"fields": ["state.subject", "state.body", "attributes.plan"], "last_n": {"state.messages": 3}, "max_tokens": 4000},
  "engine": {"name": "jev", "version": "current"},
  "freshness": {"policy": "on_change", "debounce_ms": 0},
  "thresholds": {"escalate": 0.85},
  "activate": true
}
```

## Types

| Type | Definition | Answer |
| - | - | - |
| `bool` | `question` only | `p`, the probability of yes |
| `choice` | `options`: 2 to 254 `{value, description}` entries, or `{"from": ...}` to [choose among candidates](#choose-among-candidates). Brussle appends `none_of_the_above` as a final option. | `value` (the most likely option), `dist`, and `escape_p` (the probability of `none_of_the_above`) |
| `score` | `levels`: 2 to 10 ordered `{value, label, description}` entries | `score` (the probability-weighted mean of the level values) and `dist` |

A `bool` answer has no `value`. Booleans come only from named thresholds.

### How far to trust each type

The types are not equally reliable. On our conformance set (see [the engine page](/engines/jev-current#conformance) for the latest run):

* **`bool` is the strongest.** It is the most accurate type and its `p` is close to calibrated out of the box. If a question can be yes or no, ask it that way.
* **`score` is right to within one level far more often than exactly.** About a third of answers land on the exact level, but about three quarters are within one. Read `score` as a position on the scale, set thresholds between levels (such as `3.5`) rather than on them, and don't rely on the difference between neighbouring levels.
* **`choice` picks the right option about as often as `bool`, but its probabilities are overconfident.** `dist` often puts 1.0 on one option and 0 on the rest, even when the pick is wrong. Use `value` as the answer, and post outcomes and read `calibrated` before you act on the probabilities themselves; see [calibration](/concepts/calibration).

A `bool` judgment can instead ask 2 to 8 narrow yes/no `parts` (1 to 8 beside [features](/guides/composite-judgments#features)) and combine them with weights fitted on your labels. It is measured on those labels before it answers; see [composite judgments](/guides/composite-judgments).

### Choose among candidates

A `choice` can take its options from a [blocking relation](/guides/entity-matching): the documents that share a key with the judged one, such as the open invoices a payment could settle. `"options": {"from": "candidates", "label": ["state.number", "state.amount"]}` makes one option per candidate the relation reads, in its order, newest first: `value` is the candidate's `id`, and its `label` fields describe it. `none_of_the_above` follows, meaning "none of these: no match".

* **At most 10 candidates.** The relation's `last_n` is at most the [candidate limit](/limits#relations), and the candidates are judged together in one call.
* **One match per document.** `value` names one candidate, or `none_of_the_above`. To find every duplicate of a record, use a pair document per candidate pair.
* **`escape_p` is the probability that nothing matches,** and a threshold on it, the only kind this choice takes, is your unmatched queue. `dist` is a shortlist, not a ranking. A document with no candidates is answered `none_of_the_above` with `escape_p` 1, with no engine call and nothing billed.
* **Not readable, not discoverable.** No relation can read its answers, its `horizon` is `0s`, and [discovery](#escape-alert-and-discovery) does not apply: its escape means "no candidate", not "a missing option".

[Match one record to another](/guides/entity-matching) covers keys, outcomes and cost.

## Starters

You don't have to write a definition from nothing. [Starter judgments](/guides/starter-judgments) are ready-made ones for ticket triage, moderation, churn risk, lead quality and review queues, each asked as a `bool` where the question can be yes or no. Create one with `from_starter` and the paths of your own documents; with `"dry_run": true` the response shows the definition it would make, its cost per answer and its backfill estimate, and creates nothing. `dry_run` works on any definition, not only starters. Every create returns the version it made in `definition`. See [start from a starter](/guides/context-recipes#start-from-a-starter).

## Thresholds

Thresholds are named numbers evaluated when an answer is read, and returned as booleans under `answers.{name}.thresholds`. For a `bool`, a threshold is true when `p` is at least the number. For a `score`, it is true when `score` is at least the number. For a `choice`, it names an option and a minimum probability: `"escalate": {"value": "fraud", "gte": 0.7}`. They are always evaluated against the raw fields, never the `calibrated` ones.

Thresholds are settings, not part of a version. Change them with `PATCH /namespaces/{ns}/judgments/{name}`:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"thresholds": {"escalate": 0.79, "review": 0.5}}
```

The `thresholds` you send replace the whole set, so include the ones you want to keep. `{}` removes them all. The change applies at the next read to every answer, including answers already computed, with no recompute and no new version, and is recorded in your audit log. Thresholds given when you create a version replace the current ones when that version becomes active. To pick a threshold from your own outcomes, see [the recommender](/guides/measure-improve-tune#pick-thresholds-with-the-recommender).

When another judgment's relation reads a threshold by name, changing it moves what that reader sees with no answer to re-judge it by. So the `PATCH` returns each reader's cost as `downstream` and changes nothing until you send it again with `confirm: true`; see [changing a threshold a reader reads](/guides/roll-ups#changing-a-threshold-a-reader-reads).

## Escape alert and discovery

A fixed-option `choice`'s escape option is where a new kind of complaint or abuse shows up first. Two things watch it for you, and neither changes the judgment on its own:

* **`escape_alert`** is a setting: the share of the last 7 days' answers that were `none_of_the_above`, above which you want to know. When the share rises above it, `GET` on the judgment lists the `taxonomy_drift` warning and the events feed records one `judgment.taxonomy_drift`. It needs 20 answers in the window, and is off (`null`) by default.
* **Discovery.** `POST /namespaces/{ns}/judgments/{name}/discover` starts a job that reviews recent escaped documents and proposes the options they have in common, checked against those documents before you see them. It proposes a new version's options and creates nothing: you review and create the version. It sends a sample of them to a third-party LLM provider, so an admin turns on suggestions first.

See [discover new options](/guides/discovery).

## Engine

Every judgment names one engine version, such as `{"name": "jev", "version": "current"}`. `current` runs whichever Jev model the provider serves now, and each answer records the epoch of model behaviour that produced it (`engine_version`, such as `current+2026-09-24.1`). Nothing picks an engine for you, unless you set a namespace's `default_engine`. The version must be `active` in the registry; see [engines](/engines/index).

## Context recipe

`context` says which parts of the document the engine sees. It is the main lever on both cost and accuracy; see [writing a context recipe](/guides/context-recipes). Every definition needs one: a create that leaves it out is refused with a recipe suggested from your documents, which you can accept with `"use_context": "suggested"`, and `"use_context": "whole_state"` sends the whole `state` instead. See [when you leave out context](/guides/context-recipes#when-you-leave-out-context).

A recipe can also read more than the judged document: the documents that point at it, such as an account's tickets, the one it points at, such as an order line's product, other judgments' answers for them, candidates that share a key, or the document as the judgment last saw it. See [relations](/concepts/relations).

## Versions

Versions are immutable. Posting a definition with an existing name creates version `n+1`. The first version of a name is active when it is created, except a [composite](/guides/composite-judgments#activate-it-measured-on-your-labels), which never is. A later version stays inactive until you activate it, or you pass `activate: true`.

`POST /namespaces/{ns}/judgments/{name}/activate` with `{"version": 4}` makes version 4 the one new evaluations use. Existing answers keep their `judgment_version` until they are recomputed. If the new version's engine or definition differs from the active one, the call returns `202` with a `shadow` job instead. The job judges a sample of 1,000 documents under the new version and reports the distribution shift, the threshold flips and the cost of recomputing everything. `POST /jobs/{id}/confirm` makes the switch. `"force": true` skips the report and switches at once, and so does `activate: true` on create. A composite is the exception: `force` still runs its shadow job, and `activate: true` on create is refused with `invalid_request`. See [changing a question safely](/guides/measure-improve-tune#change-a-question-safely-with-a-shadow-report).

Freshness policy, debounce, fan-out, thresholds, `escape_alert` and `outcomes` (implicit negatives and outcome rules) are settings, not part of the definition. Changing them with `PATCH /namespaces/{ns}/judgments/{name}` creates no version. `freshness` in a create sets the judgment's freshness with its first version only: a later version keeps the settings the judgment has, and its response warns when the `freshness` you sent differs. So creating a version never switches a judgment to `on_change` or drops a debounce you tuned. See [choosing a freshness policy](/guides/freshness-policies), for implicit negatives on a `bool` judgment with a `horizon`, [calibration](/concepts/calibration), and for rules that derive outcomes from your writes, [outcomes from your own data](/guides/measure-improve-tune#outcomes-from-your-own-data).

`horizon` is part of the definition: a duration such as `"30d"`, `0s` by default. It says how far ahead a judgment predicts: an answer says the event happens within the horizon after it was made, and each [outcome](/guides/measure-improve-tune#real-world-outcomes-with-a-horizon) labels the prediction whose window it falls in. Changing it creates a version.

`DELETE /namespaces/{ns}/judgments/{name}` deactivates a judgment. Its answers and evaluations are kept. While another judgment's relation [reads its answers](/guides/roll-ups#what-a-judgment-can-read), `GET` lists those `readers`, and deactivating it is refused with `conflict` naming them. A namespace can have 100 active judgments.

## Backfill

A new judgment has no answers for existing documents. `POST /namespaces/{ns}/judgments/{name}/backfill` without `confirm` returns an estimate and does nothing:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"estimate": {"documents": 87291033, "tokens": 174582066000, "judgment_units": 87291033, "cost_usd": 21822.76, "duration_s": 8475805}}
```

`cost_usd` is what the backfill costs you: its `judgment_units`, the judgments it would bill counted by [size class](/pricing#judgments), at the prices of the tiers your organization will be in, counting what it has already used this billing period, the same on every engine. Here each document's context and question come to 2,000 tokens, so each is a standard judgment and counts as 1, and all of them fall in the first tier in a month with no other judging. `duration_s` is the expected wall-clock time if nothing else is judged meanwhile. Other judging comes first and makes it longer, so read it as the least to expect. Here it is about 98 days on Jev, which is why it is shown next to the cost. With `"confirm": true` the backfill starts as a job, which you can follow, pause, resume or cancel under `/jobs/{id}`. Backfills respect the namespace budget, and a `filters` expression limits a backfill to matching documents.

A backfill judges only the documents in its scope that need it: those within the judgment's `applies_to` that match `filters`, except the ones whose answer from the active version already succeeded for their current revision. Its estimate counts the same documents, as they are when you ask, so a backfill after a finished one is quoted only for what was written since, and one with nothing left to judge is quoted nothing. The job's `progress.documents_done` counts the documents judged so far, so it climbs to the estimate's `documents`. The job reads `done` once it has passed the last document.


## Related topics

- [Deactivate a judgment](/api-reference/judgments/deactivate-a-judgment.md)
- [Composite judgments](/guides/composite-judgments.md)
- [Starter judgments](/guides/starter-judgments.md)
- [List active judgments](/api-reference/judgments/list-active-judgments.md)
- [judgment.taxonomy_drift](/api-reference/webhooks/judgmenttaxonomy_drift.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.