GET /namespaces/{ns}/documents/{id}?include=history returns a document’s evaluations for its current incarnation, newest first. Add all_incarnations=true for earlier lives of the same id, and history_limit to bound the list.
GET /namespaces/{ns}/evaluations/{id} (ns.evaluation(id) in both SDKs) returns one evaluation by its id, such as an answer’s evaluation_id, from any document, incarnation or judgment, shadow evaluations included. It has context and raw, except a replay, whose context is its replay_of’s, and latency_ms, how long the engine request took, retries excluded. latency_ms is absent on evaluations recorded before it was measured.
-
outputholds the engine’s raw numbers, each read as the answer it produced reads it:pfor a bool;value,distandescape_pfor a choice;scoreanddistfor a score. A composite judgment’s evaluation hasparts, each part’sp; its combinedpis computed when the answer is read. -
include=history,contextaddscontext, the exact compiled context the engine saw.include=history,rawaddsraw, the engine’s answer to this judgment alone, in the shape ofoutput, never with other judgments’ answers or the engine’s token counts:context_tokensis the size that matters to you. Both are left out by default. -
context_truncatedis true when the context was cut to fitmax_tokensor the engine’s limit. -
status: "failed"comes witherror: {class, message}.retryablefailures are retried with backoff and then hourly.terminalfailures wait for the document’s next write. A document’s answer becomes aterminalfailure, kept until the document is written again, in two cases:- It failed for a reason of its own (the engine refused it or could not process it, such as output that is invalid or cannot be parsed, a document too large, or a content refusal). That is terminal at once.
- A server error, timeout, lost connection or rate limit outlasted the retries with backoff, so the answer reads
failedwith classretryable. It is retried hourly, and becomesterminalif it is still failing on a retry made 24 hours or more after the document last changed.
pending, never markedfailed, however long the outage lasts. A held retry is not an attempt, so the 24-hour rule cannot turn an answerterminalduring an outage: everything is judged when the engine recovers.GET /enginesshows each engine version’s latest health probe (probe.status). To re-judge answers leftfailedwithout rewriting the documents, run a backfill of the judgment: it skips documents whose answer for the active version and current revision already succeeded, and judges the rest. Writing a document again also starts a new attempt. A context the engine will not take even cut to 80% of its limit failsterminalwith a message startingtoo_large:: lower the recipe’smax_tokensor send fewer fields. -
The evaluation of a judgment with relations has
watermark, the log position its context was read at, andrelated_documents: every related document the context read, rendered or aggregated, as{relation, document_id, revision}. The list is part of the context, so it comes withinclude=history,contextand fromGET /namespaces/{ns}/evaluations/{id}. It is what lets an outcome or an audit see exactly what the engine read, after the documents have changed. Relations add:evaluationson arelated_documentsentry whose answers the context read: the evaluation behind each answer, by judgment, such as{"frustrated": "ev_…"}. Evaluations are kept until the namespace is deleted, so a roll-up can always be followed back to the answers it read.answers_generation, besidewatermark: how far the context had read other judgments’ answers and block changes.child_answers_stale: truewhen a document whose answers it read had been written and not judged again yet, so its last answer was shown. That document’s new answer re-judges this one if it changes what a relation shows, so it is a flag to wait out, not to act on.previous_revision, for a recipe withprevious: the revision the previous rendering came from.first_revision: trueinstead when nothing was shown underprevious, such as a document’s first evaluation or a re-created id. The stored context holds both renderings.
-
shadow: truemarks evaluations from an activation’s shadow report. They never produce answers and are not billed. -
replay_ofnames the evaluation whose stored context a replay judged again after the engine’s model changed, to re-fit calibration. Replays never produce answers and are not billed. It isnullfor every other evaluation.
evaluation_id still names the earlier evaluation, for an earlier revision.