Skip to main content
An evaluation is an immutable record of one computation: the document revision, judgment version, compiled context hash, engine and version, raw output, time and status. Answers point at evaluations. Every evaluation is kept, with its compiled context, until its namespace is deleted; history is included in storage pricing. GET /namespaces/{ns}/documents/{id}?include=history returns a document’s evaluations for its current incarnation, newest first. Add all_incarnations=true for earlier lives of the same id, and history_limit to bound the list.
On Scale, you can export evaluation history: every evaluation in a time range, with its context and the outcomes joined to it, as files to download. GET /namespaces/{ns}/evaluations/{id} (ns.evaluation(id) in both SDKs) returns one evaluation by its id, such as an answer’s evaluation_id, from any document, incarnation or judgment, shadow evaluations included. It has context and raw, except a replay, whose context is its replay_of’s, and latency_ms, how long the engine request took, retries excluded. latency_ms is absent on evaluations recorded before it was measured.
  • output holds the engine’s raw numbers, each read as the answer it produced reads it: p for a bool; value, dist and escape_p for a choice; score and dist for a score. A composite judgment’s evaluation has parts, each part’s p; its combined p is computed when the answer is read.
  • include=history,context adds context, the exact compiled context the engine saw. include=history,raw adds raw, the engine’s answer to this judgment alone, in the shape of output, never with other judgments’ answers or the engine’s token counts: context_tokens is the size that matters to you. Both are left out by default.
  • context_truncated is true when the context was cut to fit max_tokens or the engine’s limit.
  • status: "failed" comes with error: {class, message}. retryable failures are retried with backoff and then hourly. terminal failures wait for the document’s next write. A document’s answer becomes a terminal failure, kept until the document is written again, in two cases:
    • It failed for a reason of its own (the engine refused it or could not process it, such as output that is invalid or cannot be parsed, a document too large, or a content refusal). That is terminal at once.
    • A server error, timeout, lost connection or rate limit outlasted the retries with backoff, so the answer reads failed with class retryable. It is retried hourly, and becomes terminal if it is still failing on a retry made 24 hours or more after the document last changed.
    A full outage is different. Once a large share of calls to the engine are failing, Brussle stops calling it, and documents are held pending, never marked failed, however long the outage lasts. A held retry is not an attempt, so the 24-hour rule cannot turn an answer terminal during an outage: everything is judged when the engine recovers. GET /engines shows each engine version’s latest health probe (probe.status). To re-judge answers left failed without rewriting the documents, run a backfill of the judgment: it skips documents whose answer for the active version and current revision already succeeded, and judges the rest. Writing a document again also starts a new attempt. A context the engine will not take even cut to 80% of its limit fails terminal with a message starting too_large:: lower the recipe’s max_tokens or send fewer fields.
  • The evaluation of a judgment with relations has watermark, the log position its context was read at, and related_documents: every related document the context read, rendered or aggregated, as {relation, document_id, revision}. The list is part of the context, so it comes with include=history,context and from GET /namespaces/{ns}/evaluations/{id}. It is what lets an outcome or an audit see exactly what the engine read, after the documents have changed. Relations add:
    • evaluations on a related_documents entry whose answers the context read: the evaluation behind each answer, by judgment, such as {"frustrated": "ev_…"}. Evaluations are kept until the namespace is deleted, so a roll-up can always be followed back to the answers it read.
    • answers_generation, beside watermark: how far the context had read other judgments’ answers and block changes.
    • child_answers_stale: true when a document whose answers it read had been written and not judged again yet, so its last answer was shown. That document’s new answer re-judges this one if it changes what a relation shows, so it is a flag to wait out, not to act on.
    • previous_revision, for a recipe with previous: the revision the previous rendering came from. first_revision: true instead when nothing was shown under previous, such as a document’s first evaluation or a re-created id. The stored context holds both renderings.
  • shadow: true marks evaluations from an activation’s shadow report. They never produce answers and are not billed.
  • replay_of names the evaluation whose stored context a replay judged again after the engine’s model changed, to re-fit calibration. Replays never produce answers and are not billed. It is null for every other evaluation.
The context hash covers the compiled context, the judgment version and the engine version. When a document changes in a way that leaves a judgment’s context unchanged, the previous answer is reused without calling the engine, and no evaluation is recorded: the answer’s evaluation_id still names the earlier evaluation, for an earlier revision.