When splitting helps
Splitting helps when a fuzzy question hides several distinct signals the engine can see on their own. It does not help when every part measures the same thing as the question: then a composite costs more per document and does no better. Nobody can tell in advance, so Brussle measures every composite on your labels, against the single question with a fitted threshold, before it answers.Try the threshold recommender first
A fitted threshold is cheaper than a composite. It costs nothing per document, needs no new version, and takes one call. If your judgment has outcomes, pick a threshold with the recommender first. Reach for a composite when a well-chosen threshold still gets too many documents wrong.Define the parts
A composite is abool judgment with parts. Post it like any other definition. On an existing name it is version n+1:
- 2 to 8 parts, or 1 to 8 with features, each a narrow yes/no question with a name unique within the judgment.
- Parts share the judgment’s context recipe and engine. They are answered together with the judgment’s other questions, and each part counts toward the 32 questions per request.
questionis not sent to the engine. It documents what the combination means. The engine is asked the parts.- Parts are part of the version. Changing one creates a new version, like any other change to the definition.
- One level only. A part is a question, never another judgment, so composites never nest.
- Only
bool. There are no compositechoiceorscorejudgments, and all parts use the same context recipe.
Features
A composite on a judgment that reads related documents can also take the relation’s aggregates asfeatures, such as ["tickets.count", "invoices.sum(state.amount)"]:
- Up to 8, each naming an aggregate the recipe’s
relateddeclares. - With features, one part is enough.
- Features add no questions, so they are never billed.
- The answer shows each feature’s value in
features, and the report gives each its weight beside the parts’.
Activate it, measured on your labels
A composite version is never active when it is created, not even as a judgment’s first version, andactivate: true is refused with invalid_request. It cannot answer until its combiner is fitted on your labels. Post labelled examples for the judgment, then activate the version with POST /namespaces/{ns}/judgments/clickbait/activate:
202 with a shadow job in awaiting_confirm, even with force: true: a composite cannot answer without a combiner, so there is nothing to skip to. The job judges your labelled documents with each part and scores the composite on held-out labels against a fair baseline: the active version, or for a first version the composite’s question asked on its own, with its own fitted threshold, never 0.5. Splitting gets no credit for what a fitted threshold alone would give.
When it is done, report compares them. For the two-part first version of clickbait above:
compositeandbaselinegive the accuracy on held-out labels with its 95% interval, and the ROC AUC.baseline.thresholdis the cut-off fitted for it.differenceis the composite’s accuracy minus the baseline’s, with its 95% interval over the same documents.verdictisbetteronly when the composite beats the baseline by more than that interval. Otherwise it isnot_better, and the threshold recommender on the baseline is the cheaper fix.partsgives each part’s weight in the combiner fitted on all the documents, on a common scale, so sizes compare. A negative weight means a yes points tofalse. A part with a weight near 0 adds cost and little else.recomputeestimates backfilling every document under the new version. Each part is billed as a judgment, so it is several times a single question’s estimate.
type, so "type" in report tells it apart from the report of an ordinary version change:
POST /jobs/{id}/confirm activates the version with the combiner fitted on all the labelled documents, whatever the verdict. POST /jobs/{id}/cancel leaves the active version as it is. The shadow job is free.
Labels. The job needs at least 50 labelled documents, and at least 10 of each answer. With fewer it fails: status is failed, and error starts with insufficient_labels. Post more labels and activate again.
Answers
A composite’s answer hasp, the combined probability, and parts, each part’s raw p from the engine:
- There is no
calibratedobject. The combiner is fitted to your outcomes, which usually leaves the combinedpclose to calibrated on documents like your labels. It is not guaranteed: with few labels, or labels from a reviewed sample, it can still be over- or under-confident. Check the reliability curve in the calibration report. combinersays whatpcame from:outcomes, the labelled documents it was fitted on, andfrom_previous_epoch.- The combiner is fitted per version and per engine epoch, and refitted nightly along with calibration. After a Jev drift, answers under the new epoch use the previous epoch’s combiner, marked
"from_previous_epoch": true, until the new epoch has 50 outcomes, at least 10 of each answer, and the nightly refit has run. - The calibration report measures the combined
pper epoch. For a composite,stage,fitted_atandcalibratedarenull: no second calibration map is applied on top of the combiner.
p; the combined p is computed when the answer is read. To see what the engine returned, and how long it took, read the evaluation by its id:
Thresholds, filters and ranking
Everything that readsp reads the combined p: thresholds, filters such as ["answers.clickbait.p", "Gte", 0.8], rank_by, and threshold recommendations. A part’s p is in the answer for you to read, but you cannot filter or rank on it.
Suggested parts
Brussle can ask a third-party LLM provider to propose parts from your labelled examples. A suggestion is a starting point, not a sign that a composite will help. Read each set for the question it is missing. It is off by default, because it sends your labelled examples to that provider, which is then one of the subprocessors listed in the data processing agreement. An org admin turns on Suggestions in the organization’s settings in the dashboard. Until then the call is refused withforbidden.
POST /namespaces/{ns}/judgments/{name}/suggest_parts with the number of parts to propose, 2 to 8:
- What the LLM sees: the judgment’s
questionandcriteriafrom its newest version, and a sample of the namespace’s labelled examples for the judgment, compiled with the judgment’s context recipe.exampleslists exactly which documents it saw. - It never creates a version. Review the parts, edit or drop any, then post a new version with them.
partshas the shape a definition’spartstakes. - It needs labels: at least 10 labelled examples of each answer, or it is refused with
insufficient_labels. - It is free, and limited to 20 calls per judgment and 50 per organization a day (
rate_limited, withdetails.limit, the limit reached, 20 or 50, anddetails.resets_at, the next midnight UTC). Long documents are cut to fit. When the LLM is unavailable it returnsengine_unavailablewithRetry-After, which the SDKs wait out and retry. A call that fails this way before the LLM ran does not count toward the day’s calls; one that timed out, or failed after the LLM replied, does.