Answers are judgments, not facts
An answer is an engine’s judgment of your document, with a probability. It can be wrong, and the numbers on an engine’s page, measured on our conformance suite, may not match your data. Measure accuracy on your own documents with outcomes, and keep a person or a check in the loop wherever a wrong answer is costly.Your data
- Writes. A write is acknowledged only after it is durably committed and replicated.
- Reads after a write. After a write is acknowledged, a strong query sees the new revision and its attributes. Eventual queries can lag; see tradeoffs.
- Concurrent writes. Writes to the same document are applied in the order they commit, and none is lost: two clients can
patchorappendto one document at once, and both changes stand. This holds during a failover too. A write caught in one can be refused withrate_limitedand aRetry-Afterheader, and retrying it is safe; the SDKs do so for you. To write a document back only if nothing changed it since you read it, send anupsertwithif_revision: it is refused withconflictif the revision moved (conditional upserts). - Retries. Retrying a write is safe: a retry leaves what the first attempt left, even when the first attempt landed and only its response was lost. An
appendskips every value already in the array, however long ago it was added and whatever was written since, so a retried append adds nothing while its values are still there. Two limits. If a later write removes those values (apatchorupsertof the array, or adelete), a late retry of the append adds them again. And a retriedupsertorpatchsets its content again, so one that arrives after a later write to the same fields undoes that write, as any late write would. A conditionalupsert(if_revision) is the exception: a retry of one whose first attempt landed isconflict, since the revision moved on, so read the document to tell whether it landed. - Repeated values in an array. Because an
appendskips values already in the array, the same value can never be appended twice, in one request or in two. To keep repeats, make each element unique, such as with anidor a timestamp field. - Training. We do not fine-tune the engines we host on your data unless you opt in.
Provenance
Every answer names its document revision, judgment version, engine version, and evaluation id. An answer reused because a change left its context the same keeps the id of the evaluation that computed it, which was for an earlier revision. An answer from a judgment that reads related documents also names its watermark, except afailed one, which keeps its last good numbers without it, and its evaluation lists every related document it read, with its revision.
Engines
- Your judgment runs on the engine version you named until you change it. Jev
currentcannot be pinned: it runs whatever Jev serves now. We detect changes in its behaviour and record them rather than hide them, and each answer’sengine_versionnames the epoch of model behaviour that produced it. We re-check its quality against our conformance suite every day and after every change we detect, and publish the results on its engine page. Other third-party engines are pinned as far as the vendor allows. - Versions do not change silently. We intend to keep the engine versions we host available. If a version becomes unavailable, requests fail with a named error (
engine_version_unavailable) rather than moving to another version. - Which subprocessor sees your data depends on the engine you choose, and is written on that engine’s page. See where your data goes.
Where your data goes
- Judging sends each document’s compiled context, with the judgment’s question, to the engine you chose. On an engine we host, it stays on our network. On a third-party engine, it goes to that engine’s provider; the engine’s page, such as Jev’s, names who operates the model.
- Calibration after a model change sends the stored compiled contexts of the answers your recent outcomes labelled to the same engine once more, on the same terms. A document you have deleted is not sent.
- Suggested parts and discovery are off until an org admin turns on Suggestions. Then they send labelled examples, or a sample of documents, with the judgment’s question, to a third-party LLM provider.
- Subprocessors. Every subprocessor that receives your data for any of these is listed in the data processing agreement. How long each keeps what it receives is set by that provider’s terms, as the agreement records them.
Budgets
- Judging stops before the budget, not after it. Each request to the engine is priced before it is sent: each of its questions counted by size class over the compiled context, as your bill counts them, at the price of the tier your organization is in. It is sent only if the namespace’s spend this billing period, plus the requests already sent and not yet recorded, plus its price, is within the budget. A request that does not fit is not sent: its documents stay pending, nothing is billed for them, and judging pauses (
budget_paused: true). A large write is judged up to the last document that fits, and one that fits exactly is judged in full. - How far spend can pass the budget. When judging runs in parallel, such as a backfill beside live writes, spend can pass the budget only by what was already on its way to the engine when it was reached: about a second of judging for each kind of judging running at once. The one exception: if we briefly can’t tell your organization’s usage so far this billing period when answers are recorded, they are priced at the first tier, the highest price, which can put spend above what was estimated.
- Resuming. Raising or removing the budget, or the start of your next billing period, ends the pause. The documents that waited are still pending, and are judged at the namespace’s next write or within a few minutes, whichever comes first.
Webhooks and the events feed
Every event is recorded once in your organization’s event log. Webhooks push it to your endpoints, and the events feed (GET /events) lists the same log in order.
- Recorded once.
- A subscription’s
enteredorexitedis recorded with the change in membership that caused it, so a transition is never lost or recorded twice, including across failures and retries on our side. - A job event is recorded in the same transaction as the change it reports. A job that finished has exactly one
job.completedorjob.failed. - A budget pause or resume is recorded at most once, with the email about it. If the event log is briefly unavailable, pause and resume events in between can be collapsed or, rarely, lost; the namespace’s
budget_pausedalways shows where it stands.
- A subscription’s
- Delivered at least once. Every recorded event is delivered to each enabled endpoint it routes to, until the endpoint answers with a 2xx or its retries run out, after about 3 days.
- A retry, a redelivery or a replay repeats the same
webhook-id, so dedupe on it. - Duplicates are rare but possible. For example, a 2xx whose response is lost to a timeout is sent again.
- A retry, a redelivery or a replay repeats the same
- Ordering.
- Webhooks are unordered. A retried event can arrive after a later one.
- Subscription events carry
sequence, which is higher for each later evaluation of a namespace’s subscriptions. For one subscription and document, a highersequenceis newer, so drop anything older than what you have. It starts again when a namespace is deleted and created again (see subscriptions). - The events feed is ordered for your organization in the order events were recorded, oldest first, and
next_cursorcarries on from the last one you read. An event is never visible before one recorded ahead of it.
- Completeness for subscriptions. For each subscription and document, live events alternate between
enteredandexited. A change applied silently always sends asubscription.synced: a subscription’s first evaluation, and a bulk change such as a backfill under the default"bulk": "summary". Under"bulk": "deliver"a bulk change sends each transition instead. - Timestamps. An event’s
timestampis when it was recorded in the event log. Each delivery’swebhook-timestampis when that attempt was signed. - Latency. We aim to make the first delivery attempt within 10 seconds of the answer that caused it, for 95% of events. This is a target, not a service level.
- Retention. Events and their delivery log are kept for 30 days. A feed cursor older than that has expired.
- When the event log is down, judging carries on. Subscription transitions wait, and are recorded in order once it recovers, with none lost.