Skip to main content
Brussle is designed for a low cost per record and background judging, not for the request path. That design has costs you should know before you build on it.
Numbers marked measured on staging come from runs against our staging deployment on 28 and 29 September 2026. Unless they say they were measured in the service’s region, the client was elsewhere in the US, so they include about 50 ms of network at the median. Staging has less capacity than production, so they are not production guarantees. The other latency and freshness numbers on this page are targets. We don’t measure them continuously yet.
  • Writes take roughly 100 to 200 ms in the service’s region, because each batch is durably committed before it is acknowledged. Measured on staging from within the service’s region on 29 September 2026, a one-document write was acknowledged in 100 to 136 ms at p50 and 148 to 195 ms at p90, at 1 to 500 writes a second to one namespace. From the US client, network included, it took 172 ms at p50 and 209 ms at p90, and a write of 100 documents 193 ms and 296 ms. Writes to one namespace are acknowledged in order, so when a batch of writes is slow to store (about 1 in 250 takes over a quarter of a second), the namespace’s next writes wait for it, for up to about 300 ms. At 500 to 1,000 writes a second to one namespace, measured on staging from within the region, p99 was 290 to 350 ms. It is not a replacement for your transactional database. If writes of that speed are too slow for your use case, Brussle is not the right fit for it.
  • Answers lag writes by seconds while the engine is healthy and keeping up with your writes. A burst of writes waits its turn (see freshness lag below). Freshness is always visible.
  • It is not for the request path. A new answer takes seconds, so Brussle fits work that runs beside your product, such as triage queues, moderation and risk flags, not a decision a user’s request waits on. wait_for makes a write wait for its answers; it is for scripts and tests, not for serving a user. Reading answers that already exist is fast. Measured on staging, getting a document with its answer took 115 to 217 ms at p50, network included, on namespaces of 10,000 to 1M documents.
  • A document can end failed until it is written again: when it fails for a reason of its own (the engine refuses it or cannot process it), or when server errors or timeouts for it are still failing on a retry 24 hours after it last changed. A full engine outage holds documents pending instead, and they are judged when it recovers. A backfill re-judges what is left failed without rewriting it; see evaluations.
  • Cold queries take hundreds of milliseconds the first time after a namespace has been idle. On a 1M-document namespace the warm p90 target is under 10 ms for eventual reads, and the cold p90 target is under 500 ms. strong reads (the default) see the newest writes, so their warm number is higher: measured on staging, they added about 65 ms at p50 over the same eventual query. Also measured on staging, on 1M documents: a query with a selective attribute filter took 73 ms at p50 and 109 ms at p90 warm with eventual, network included, and 127 ms the first time after the namespace had been idle. From outside the region we can’t yet tell whether the warm part, within the region, meets the 10 ms target. Pinning keeps a namespace warm, at a price. A query on a judgment that reads the document each judged document points at can take tens of milliseconds longer the first time after a while, on namespaces with many changed referenced documents. Later queries are not slowed.
  • Filtering on an answer is fast when the judged documents were written together, and slower when they are spread thinly across a large namespace. A filter that only a document with an answer can pass, such as answers.urgent.p Gte 0.5 or answers.urgent.freshness Eq fresh, or answers: "fresh_only", skips the documents without that answer. Answer filters are fastest when the judged documents have nearby ids, for example written together or sharing an id prefix. Spread thinly across a large namespace, the query reads most of it. Ranking by an answer without such a filter, and filters a document without an answer can pass (NotEq, Not, freshness Eq pending), still read every document, because documents without an answer are results there.
  • Freshness lag for on_change judgments targets under 5 s at p90 with a healthy engine that is keeping up with your writes. Measured on staging, with a one-field recipe and one write at a time, an answer was fresh 1.3 s after the write was acknowledged at p50, 1.6 s at p90, and 2.4 s at most (n = 200). Writes that arrive faster than the engine judges wait in line. After 500 writes at once, the lag was 21 s at p50 and 33 s at p90, and the last answer came 44 s after its write. Bursts are judged as fast as the engine’s rate limit allows and queue beyond it.
  • Answers that read related documents lag them by the debounce, typically minutes, and by at most the ceiling (max_wait_ms). Measured on staging with a 5 s debounce, an account’s answer was fresh 6.2 s after a ticket was written at p50 and 6.7 s at p90. With a new ticket every few seconds and a 20 s ceiling, the account was judged 22 to 23 s after the first ticket. That is the price of judging an account once per burst of tickets rather than once per ticket. A question about a single event, such as a fraudulent payment, stays a judgment of that document.
  • The cost of a judgment with related documents follows how often they change. Creating one shows the monthly cost, replayed from your last 30 days of writes, before you confirm.
  • A judgment that reads the document each judged document points at re-judges, on each change to that document, every judged document in its re-judge scope. That is the fan-out, and it costs changes × documents in scope. So only a change to what the relation shows counts, bands hide moves inside a band, the scope defaults to documents created in the last 30 days, and a burst of changes is debounced to one fan-out. Changes that never settle still fan out at least once per ceiling, an hour by default, for as long as they keep coming. The replay estimate cannot count fan-out and says so. Every one of these defaults is yours to change.
  • A fan-out runs behind ordinary judging, on at most half of your namespace’s engine requests, so a large one takes hours; it runs with no job, and the judgment’s fanout_window shows how much of its rolling limit fan-outs have used. Answers outside the re-judge scope keep the version of the referenced document they read and read stale until their own document is judged again; their watermark says which version.
  • Probabilities are the engine’s. They are not calibrated to your outcomes until you give us outcomes. We label the difference. From 100 outcomes, with 20 of each kind, answers carry a calibrated object beside the raw fields, fitted per judgment version and engine epoch and refitted nightly. A fit is used only when it beats the raw numbers on outcomes it was not fitted on.
  • Composite judgments need at least 50 of your labels before they can answer, and cost one judgment per part. Suggested sub-questions come from a third-party LLM provider that sees a sample of your labelled examples, listed in the response; it is off unless you turn it on.
  • Jev’s answers vary from call to call. Asked the same question about the same text repeatedly, Jev gave probabilities with a standard deviation of up to about 0.031, and the widest gap between its lowest and highest value was about 0.12. So a document whose probability is within about 0.06 of a threshold can land on either side of it each time it is judged, and within about 0.03 it often will. That includes a re-judgment after a change that has nothing to do with the question. For documents that close, treat the threshold as a band rather than a line. An unchanged context is never re-judged, so a stored answer never changes on its own, with three exceptions: a periodic judgment re-judges every document on its interval whether or not it changed; a composite’s combined p moves when its combiner is refitted overnight; and a calibrated value moves when calibration is refitted, overnight or after the engine’s model changes, while the raw numbers beside it stay as they were.
  • The engines cannot abstain. We inject an escape option on choices and expose it. For bool judgments, a probability near 0.5 is the only signal.
  • No ad hoc questions. Define the judgment, then query it.
  • Large backfills can take days. Every backfill estimate shows the duration before you confirm. Measured on staging, backfills of 1,000 short documents were estimated at 90 s, and every answer was in after 79 to 87 s. A backfill’s job judges and counts only the documents its judgment applies to, so its progress reads as documents judged out of the estimate’s.
  • Laya, when it ships, takes at most 512 tokens. A longer context on a Laya judgment is cut to 512 tokens, with context_truncated: true on the evaluation; it is never sent to Jev instead. Long text needs a Jev judgment. Jev current has no versions for us to retire: it follows whatever model its provider serves, and each change we detect in its behaviour starts a new epoch, recorded in every answer’s engine_version. A judgment on any engine version that is retired fails with engine_version_unavailable, and we never re-pin silently.
  • Your system of record stays yours. You keep documents in sync through the write API. CDC connectors come later.

Consistency, stated separately

There are two consistency models, one for documents and one for answers.
  • Documents: strong. After a write acks, a strong query sees the new revision and its attributes. Measured on staging, a get and a strong query run straight after a write both saw it in 300 of 300 tries. Eventual queries may lag up to 60 seconds, and when a namespace is queried steadily they often lag that long: measured on staging, an eventual query polled after a write first saw it 60 s later at p50.
  • Answers: eventually judged, with explicit freshness. After a write acks, the answer for that revision arrives seconds later while the engine is healthy and keeping up with your writes. Until then the answer carries freshness: pending or stale and the revision it was computed against. Queries can ask for answers: "fresh_only" to exclude pending documents. A write can pass wait_for: ["needs_escalation"] to block until the answer for a revision at or after its own exists.