Skip to main content
A judgment’s context recipe decides which parts of a document the engine sees. Judging is billed per judgment, each counted by its size class, which the tokens of compiled context and question together decide, so the recipe is your biggest lever on cost. It is also a lever on accuracy: the engine answers from what it is shown, and a few recent related records are usually enough.
A definition needs a recipe. Leaving out context is refused with one suggested from your documents; see when you leave out context. A recipe decides what the engine sees of a document. Which documents a judgment judges at all is applies_to, beside the recipe: {"attributes.kind": "account"} judges, answers and bills only accounts. Each key is an attribute; a value is equality, and a list of values matches any of them. Keys combine with and, up to 8. Other documents have no answer for the judgment. applies_to is part of the definition, so changing it creates a version. The recipe is the part of Brussle that takes real thought. It decides three things:
  • Cost. The tokens it sends, with the question, set each judgment’s size class.
  • Accuracy. The engine answers from what it is shown, and nothing else.
  • When anything is judged again. A change to a field or related record that the recipe leaves out changes nothing the engine sees, so it costs nothing and re-runs nothing.
For each question, ask what a careful person would need to see to answer it, and send only that. Then check what the engine saw on a few documents.

Start from a starter

Starter judgments are ready-made definitions for the most common questions: ticket triage (is it urgent, which team), moderation, churn risk, lead quality and review queues. Each has its question, type, recipe and suggested thresholds already written, and its paths are placeholders you fill with your own. Starter judgments lists every one, with what it reads and when to use it; GET /v1/starters returns the same list. Fill one in with from_starter and paths, and add dry_run to see the definition it makes and what it would cost, without creating anything:
  • definition is exactly what creating it makes. To change its question, criteria or recipe, edit it and create it as an ordinary definition.
  • cost_per_answer is the mean judgments per answer, each counted by its size class, over a sample of up to 1,000 of the documents it would judge: 1.0 when every answer is standard, 4.0 when every one is large.
  • estimate is what backfilling the documents you already have would cost. Creating a judgment never backfills by itself.
  • For a starter that reads related documents (churn risk), replay is its estimated monthly cost.
The same body without dry_run creates it, and the response carries the same definition. A placeholder that is not optional must be given, and each must be the right kind of path; otherwise the request is refused with invalid_request and details names the problem. You can pass name, engine, freshness, activate and confirm beside from_starter, and nothing else. A starter sets no freshness policy, so without freshness the judgment is on_read and its answers cannot be filtered, sorted or subscribed to; pass "freshness": {"policy": "on_change"} for that, as the starter judgments bodies do. A starter can suggest how its outcomes are read. Churn risk turns on implicit negatives and, if you give the account’s status path, adds an outcome rule for “status becomes cancelled”. They are set when the judgment is first created: implicit negatives on every plan, and the rule only on a plan with outcome rules. On any other plan the response’s outcomes returns the rule with plan_required, so you can set it with PATCH after you upgrade. In the dashboard, “New judgment” has the same flow: pick a starter, choose each path from the fields in your documents, check the definition and its cost, then create.

When you leave out context

A definition needs a context recipe. If you leave context out, the create is refused with invalid_request, nothing is created, and details.suggested_context suggests one from a sample of your documents:
It keeps short text, numbers, booleans and small attributes. It leaves out ids, timestamps, links, binary-looking and very long values, and rarely present fields, and excluded says why for each. Long arrays keep their most recent items, and max_tokens is set to keep the context and its question together in the standard size class: the cap leaves room for the question. A longer question you add later counts too, so if cost_per_answer rises above 1.0, lower max_tokens by about the question’s growth, or shorten the question. The same documents always give the same suggestion. Then either:
  • send a recipe of your own as context, starting from the suggestion;
  • send "use_context": "suggested" to create with the suggestion as it is. The response’s definition shows the recipe it used;
  • send "use_context": "whole_state" to send the whole state, up to the engine’s limit. That is allowed but rarely what you want: it costs the most per answer, and any change to the document re-runs it. The response warns.
A namespace with no documents yet has nothing to suggest from, and suggested_context is null: pass a recipe or whole_state.

Context size is the cost

Each judgment answered counts by its size class, which the tokens of its compiled context and its question together decide, whichever engine answers: The question counts as the engine reads it, with its criteria and every option’s description, so a choice between many long-described options can be large over a context where a yes/no question is standard. It costs that again every time the document changes. For one judgment whose context and question come to 2,000 tokens, a standard judgment: Over a 6,000-token context each is large, and the same workloads are 40M and 400M. Pricing has the price per judgment, which falls as your organization’s volume in a billing period grows. A smaller context lowers the bill when it moves judgments into a smaller size class: a recipe that sends the last three messages and the plan, under 2,000 tokens with the question, instead of a 6,000-token history makes each judgment standard instead of large, a quarter of the price. Within a class, fewer tokens cost the same: cutting a 1,500-token context to 800 changes nothing on the bill.

Let your outcomes tune the recipe

Most recipes start generous: it is hard to know in advance which fields the question needs. Once you post outcomes, Brussle checks for you. When calibration is fitted, and at most monthly after that, it tests cheaper variants of your recipe against your labelled outcomes. It uses the contexts already stored with your evaluations, sent to the same engine on the same terms; nothing is read from your documents again. A smaller recipe is cheaper only where it moves answers into a smaller size class. It suggests one only when it is meaningfully cheaper and accuracy is unchanged within noise. Tuning is never billed, and never applied for you.
To apply it, post definition as a new version and activate it through its shadow report, which shows what the change does to your answers before you confirm. definition is your active version with only context changed, and it gives no thresholds, so your current ones carry over. Measure, improve, tune has every field and when there is no suggestion.

Share recipes where you can

Judgments that share a recipe are answered together, up to 32 questions per request, which is faster and lighter on rate limits, and that matters most for large backfills. Each is still billed on its own, by its own size class. Before you give a new judgment its own recipe, check whether one you already have would do. related lets a judgment read the documents that point at the judged one: an account’s tickets and invoices, a user’s posts. The answer is kept current as any of those documents change: see judge a document with its related documents. Start with a small last_n of raw text, or with aggregates; that guide says which to try first.
  • Up to 4 relations, each with fields, aggregate or both. Each one that reads the documents pointing at the judged one takes last_n, window or both; a blocking one needs window; one that reads the document it points at takes neither.
  • Newest created first. The documents are ordered by created_at, so last_n: 8 is the 8 most recently created. window applies first, then last_n. A relation reads at most the newest 1,000.
  • Aggregates count the relation’s selection. count is how many it selected. sum, min and max skip values that are not numbers; sum of nothing is 0, and min and max of nothing are null. latest is the value in the newest document that has one. To show a few documents and count many, use two relations with the same match.
  • Relations stay in the namespace. A document and everything related to it live in the same namespace.
Each relation becomes one related.<name> entry in the compiled context, after the fields entries. It holds records, the selected documents newest first, keyed by path; one key per aggregate, such as count and sum(state.amount); and capped: true when more than 1,000 documents matched, so the numbers cover only the newest 1,000 (only a relation without last_n can match that many):
Over max_tokens, the compiler drops the last relation’s oldest records first, then those of the relation before it, then any previous entries, and only then the fields. Relations are rendered, and cut, in the order of their names. Aggregates are computed before truncation and are never cut. An aggregate is a few tokens where raw text is hundreds, so it is the first thing to try when a question depends on volume rather than wording.

The document the judged one points at

With join: {"theirs": "id", "mine": "attributes.<name>"}, a relation reads the one document whose id the judged document’s own attribute holds: its referenced document. Illustrations across domains: an order line reading its product, a message reading its conversation. Here each task points at its project:
It renders like any relation, with at most one record, so it takes neither last_n nor window. It can also show that document’s answers. When the attribute is missing, is not a document id, or names a document that does not match match, the relation is empty. A change to what it shows of the referenced document re-judges the judged documents that point at it and are inside the judgment’s re-judge scope, which the judgment’s freshness.fanout settings decide. See judge a document with the document it points at for what that costs and the controls, and fan-out for the settings.

Bands

An entry of a relation’s fields can be an object that shows a number as the band it falls in, rather than the number:
  • bands are 1 to 9 cut points, strictly increasing, and labels has exactly one more entry, each up to 64 bytes.
  • A number renders as the label whose position is how many cut points are at or below it: 0.29 is low, 0.3 is medium, and 0.7 or more is high.
  • A value that is not a number renders unchanged, and a missing one is left out, as for any path.
  • Aggregates always use the raw values.
The engine sees the word, and a move inside a band leaves the context exactly as it was. In a relation that reads the documents pointing at the judged one, that means the answer is kept and nothing is billed. In a relation that reads the document the judged one points at, it means no fan-out at all. Use bands for a number that moves often but matters only past a few thresholds.

Check what the engine saw

Every evaluation stores its compiled context. Fetch it with include=history,context on a get, or open the document in the dashboard, to see the exact text the engine received. context_tokens and context_truncated on each evaluation tell you how close a recipe runs to its cap.

Engine limits

The recipe must fit the engine you pin. Jev takes up to 32,000 tokens for the state plus the longest question. Laya, which is coming, will take at most 512 tokens in total and cut a longer context to fit, so it suits short text such as titles, single messages or search queries. See limits and the engines.