Skip to main content
A fuzzy yes/no question is sometimes answered better as several narrow ones. A composite judgment asks the engine 2 to 8 narrow yes/no questions (1 to 8 beside features), called parts, and combines their answers with weights fitted on your labels. It helps some judgments and not others, so Brussle measures every composite on your labels before it answers anything.

When splitting helps

Splitting helps when a fuzzy question hides several distinct signals the engine can see on their own. It does not help when every part measures the same thing as the question: then a composite costs more per document and does no better. Nobody can tell in advance, so Brussle measures every composite on your labels, against the single question with a fitted threshold, before it answers.

Try the threshold recommender first

A fitted threshold is cheaper than a composite. It costs nothing per document, needs no new version, and takes one call. If your judgment has outcomes, pick a threshold with the recommender first. Reach for a composite when a well-chosen threshold still gets too many documents wrong.

Define the parts

A composite is a bool judgment with parts. Post it like any other definition. On an existing name it is version n+1:
  • 2 to 8 parts, or 1 to 8 with features, each a narrow yes/no question with a name unique within the judgment.
  • Parts share the judgment’s context recipe and engine. They are answered together with the judgment’s other questions, and each part counts toward the 32 questions per request.
  • question is not sent to the engine. It documents what the combination means. The engine is asked the parts.
  • Parts are part of the version. Changing one creates a new version, like any other change to the definition.
  • One level only. A part is a question, never another judgment, so composites never nest.
  • Only bool. There are no composite choice or score judgments, and all parts use the same context recipe.
Good parts are narrow and different from each other. Include a part that points the other way, where a yes means no.

Features

A composite on a judgment that reads related documents can also take the relation’s aggregates as features, such as ["tickets.count", "invoices.sum(state.amount)"]:
  • Up to 8, each naming an aggregate the recipe’s related declares.
  • With features, one part is enough.
  • Features add no questions, so they are never billed.
  • The answer shows each feature’s value in features, and the report gives each its weight beside the parts’.

Activate it, measured on your labels

A composite version is never active when it is created, not even as a judgment’s first version, and activate: true is refused with invalid_request. It cannot answer until its combiner is fitted on your labels. Post labelled examples for the judgment, then activate the version with POST /namespaces/{ns}/judgments/clickbait/activate:
Activation always returns 202 with a shadow job in awaiting_confirm, even with force: true: a composite cannot answer without a combiner, so there is nothing to skip to. The job judges your labelled documents with each part and scores the composite on held-out labels against a fair baseline: the active version, or for a first version the composite’s question asked on its own, with its own fitted threshold, never 0.5. Splitting gets no credit for what a fitted threshold alone would give. When it is done, report compares them. For the two-part first version of clickbait above:
  • composite and baseline give the accuracy on held-out labels with its 95% interval, and the ROC AUC. baseline.threshold is the cut-off fitted for it.
  • difference is the composite’s accuracy minus the baseline’s, with its 95% interval over the same documents.
  • verdict is better only when the composite beats the baseline by more than that interval. Otherwise it is not_better, and the threshold recommender on the baseline is the cheaper fix.
  • parts gives each part’s weight in the combiner fitted on all the documents, on a common scale, so sizes compare. A negative weight means a yes points to false. A part with a weight near 0 adds cost and little else.
  • recompute estimates backfilling every document under the new version. Each part is billed as a judgment, so it is several times a single question’s estimate.
Only a composite report has type, so "type" in report tells it apart from the report of an ordinary version change:
Then decide. POST /jobs/{id}/confirm activates the version with the combiner fitted on all the labelled documents, whatever the verdict. POST /jobs/{id}/cancel leaves the active version as it is. The shadow job is free. Labels. The job needs at least 50 labelled documents, and at least 10 of each answer. With fewer it fails: status is failed, and error starts with insufficient_labels. Post more labels and activate again.

Answers

A composite’s answer has p, the combined probability, and parts, each part’s raw p from the engine:
  • There is no calibrated object. The combiner is fitted to your outcomes, which usually leaves the combined p close to calibrated on documents like your labels. It is not guaranteed: with few labels, or labels from a reviewed sample, it can still be over- or under-confident. Check the reliability curve in the calibration report.
  • combiner says what p came from: outcomes, the labelled documents it was fitted on, and from_previous_epoch.
  • The combiner is fitted per version and per engine epoch, and refitted nightly along with calibration. After a Jev drift, answers under the new epoch use the previous epoch’s combiner, marked "from_previous_epoch": true, until the new epoch has 50 outcomes, at least 10 of each answer, and the nightly refit has run.
  • The calibration report measures the combined p per epoch. For a composite, stage, fitted_at and calibrated are null: no second calibration map is applied on top of the combiner.
The evaluation behind an answer stores each part’s raw p; the combined p is computed when the answer is read. To see what the engine returned, and how long it took, read the evaluation by its id:

Thresholds, filters and ranking

Everything that reads p reads the combined p: thresholds, filters such as ["answers.clickbait.p", "Gte", 0.8], rank_by, and threshold recommendations. A part’s p is in the answer for you to read, but you cannot filter or rank on it.

Suggested parts

Brussle can ask a third-party LLM provider to propose parts from your labelled examples. A suggestion is a starting point, not a sign that a composite will help. Read each set for the question it is missing. It is off by default, because it sends your labelled examples to that provider, which is then one of the subprocessors listed in the data processing agreement. An org admin turns on Suggestions in the organization’s settings in the dashboard. Until then the call is refused with forbidden. POST /namespaces/{ns}/judgments/{name}/suggest_parts with the number of parts to propose, 2 to 8:
  • What the LLM sees: the judgment’s question and criteria from its newest version, and a sample of the namespace’s labelled examples for the judgment, compiled with the judgment’s context recipe. examples lists exactly which documents it saw.
  • It never creates a version. Review the parts, edit or drop any, then post a new version with them. parts has the shape a definition’s parts takes.
  • It needs labels: at least 10 labelled examples of each answer, or it is refused with insufficient_labels.
  • It is free, and limited to 20 calls per judgment and 50 per organization a day (rate_limited, with details.limit, the limit reached, 20 or 50, and details.resets_at, the next midnight UTC). Long documents are cut to fit. When the LLM is unavailable it returns engine_unavailable with Retry-After, which the SDKs wait out and retry. A call that fails this way before the LLM ran does not count toward the day’s calls; one that timed out, or failed after the LLM replied, does.

Pricing

Each part is billed as a judgment, over the context and that part’s own question. A five-part composite whose context and each part’s question come to at most 2,000 tokens is five standard judgments per document, where a single question is one; a part whose context and question are larger counts by its own size class. Features cost nothing. The shadow job that measures it is free, like every shadow job, and suggestions are free. See pricing.