> ## Documentation Index
> Fetch the complete documentation index at: https://docs.brussle.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluations

> The immutable record of every computation, kept for as long as its namespace.

An evaluation is an immutable record of one computation: the document revision, judgment version, compiled context hash, engine and version, raw output, time and status. Answers point at evaluations. Every evaluation is kept, with its compiled context, until its namespace is deleted; history is included in storage pricing.

`GET /namespaces/{ns}/documents/{id}?include=history` returns a document's evaluations for its current incarnation, newest first. Add `all_incarnations=true` for earlier lives of the same id, and `history_limit` to bound the list.

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "id": "ev_01j8zq3k5v9w2x4y6z8a0b1c2d",
  "document_id": "t_123",
  "revision": 43,
  "incarnation": 7,
  "judgment": "needs_escalation",
  "judgment_version": 3,
  "engine": "jev",
  "engine_version": "current+2026-09-24.1",
  "context_hash": "sha256:9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08",
  "context_tokens": 1830,
  "context_truncated": false,
  "output": {"p": 0.91},
  "status": "success",
  "error": null,
  "shadow": false,
  "replay_of": null,
  "created_at": "2026-09-23T12:00:02Z"
}
```

On Scale, you can [export evaluation history](/guides/export-history): every evaluation in a time range, with its context and the outcomes joined to it, as files to download.

`GET /namespaces/{ns}/evaluations/{id}` (`ns.evaluation(id)` in both SDKs) returns one evaluation by its id, such as an answer's `evaluation_id`, from any document, incarnation or judgment, shadow evaluations included. It has `context` and `raw`, except a replay, whose context is its `replay_of`'s, and `latency_ms`, how long the engine request took, retries excluded. `latency_ms` is absent on evaluations recorded before it was measured.

* **`output`** holds the engine's raw numbers, each read as the answer it produced reads it: `p` for a bool; `value`, `dist` and `escape_p` for a choice; `score` and `dist` for a score. A [composite judgment](/guides/composite-judgments)'s evaluation has `parts`, each part's `p`; its combined `p` is computed when the answer is read.
* **`include=history,context`** adds `context`, the exact compiled context the engine saw. **`include=history,raw`** adds `raw`, the engine's answer to this judgment alone, in the shape of `output`, never with other judgments' answers or the engine's token counts: `context_tokens` is the size that matters to you. Both are left out by default.
* **`context_truncated`** is true when the context was cut to fit `max_tokens` or the engine's limit.
* **`status: "failed"`** comes with `error: {class, message}`. `retryable` failures are retried with backoff and then hourly. `terminal` failures wait for the document's next write. A document's answer becomes a `terminal` failure, kept until the document is written again, in two cases:

  * It failed for a reason of its own (the engine refused it or could not process it, such as output that is invalid or cannot be parsed, a document too large, or a content refusal). That is terminal at once.
  * A server error, timeout, lost connection or rate limit outlasted the retries with backoff, so the answer reads `failed` with class `retryable`. It is retried hourly, and becomes `terminal` if it is still failing on a retry made 24 hours or more after the document last changed.

  A full outage is different. Once a large share of calls to the engine are failing, Brussle stops calling it, and documents are held `pending`, never marked `failed`, however long the outage lasts. A held retry is not an attempt, so the 24-hour rule cannot turn an answer `terminal` during an outage: everything is judged when the engine recovers. `GET /engines` shows each engine version's latest health probe (`probe.status`). To re-judge answers left `failed` without rewriting the documents, run a [backfill](/concepts/judgments#backfill) of the judgment: it skips documents whose answer for the active version and current revision already succeeded, and judges the rest. Writing a document again also starts a new attempt. A context the engine will not take even cut to 80% of its limit fails `terminal` with a message starting `too_large:`: lower the recipe's `max_tokens` or send fewer fields.
* **The evaluation of a judgment with [relations](/concepts/relations)** has `watermark`, the log position its context was read at, and `related_documents`: every related document the context read, rendered or aggregated, as `{relation, document_id, revision}`. The list is part of the context, so it comes with `include=history,context` and from `GET /namespaces/{ns}/evaluations/{id}`. It is what lets an outcome or an audit see exactly what the engine read, after the documents have changed. Relations add:
  * **`evaluations`** on a `related_documents` entry whose answers the context read: the evaluation behind each answer, by judgment, such as `{"frustrated": "ev_…"}`. Evaluations are kept until the namespace is deleted, so a [roll-up](/guides/roll-ups#the-audit) can always be followed back to the answers it read.
  * **`answers_generation`**, beside `watermark`: how far the context had read other judgments' answers and block changes.
  * **`child_answers_stale: true`** when a document whose answers it read had been written and not judged again yet, so its last answer was shown. That document's new answer re-judges this one if it changes what a relation shows, so it is a flag to wait out, not to act on.
  * **`previous_revision`**, for a recipe with [`previous`](/guides/change-detection#where-it-shows): the revision the previous rendering came from. **`first_revision: true`** instead when nothing was shown under `previous`, such as a document's first evaluation or a re-created id. The stored context holds both renderings.
* **`shadow: true`** marks evaluations from an activation's [shadow report](/guides/measure-improve-tune#change-a-question-safely-with-a-shadow-report). They never produce answers and are not billed.
* **`replay_of`** names the evaluation whose stored context a replay judged again after the engine's model changed, to re-fit [calibration](/concepts/calibration#when-the-engine-changes). Replays never produce answers and are not billed. It is `null` for every other evaluation.

The context hash covers the compiled context, the judgment version and the engine version. When a document changes in a way that leaves a judgment's context unchanged, the previous answer is reused without calling the engine, and no evaluation is recorded: the answer's `evaluation_id` still names the earlier evaluation, for an earlier revision.


## Related topics

- [Export evaluation history](/api-reference/documents/export-evaluation-history.md)
- [Export audit and evaluation history](/guides/export-history.md)
- [Get one evaluation record](/api-reference/documents/get-one-evaluation-record.md)
- [Make a version the one new evaluations use](/api-reference/judgments/make-a-version-the-one-new-evaluations-use.md)
- [Judge a document with its related documents](/guides/related-documents.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.