Get the calibration report for the active version
How well the active version’s answers match the outcomes joined to its
evaluations, per engine epoch, before and after calibration.
Calibration is fitted per judgment version and per engine epoch, first
within minutes of the judgment’s first posted outcome or first outcome
rules, and then nightly. From 100 outcomes it is a conservative
correction (stage: early); with many more, one that follows the
engine’s own pattern closely (stage: full). An epoch also needs 20
outcomes of each class: 20 true and 20 false for a bool,
and two values with 20 each for a choice or score. Until then it has no
calibration, its calibrated metrics are null, and not_fitted says
why and what to post. A fit is applied only when it beats the raw
answers on outcomes it was not fitted on (held_out);
otherwise answers stay raw and not_fitted is raw_better.
unmatched_outcomes counts the outcomes that joined nothing, by
reason. coverage and warnings say where the outcomes
sit, and warn when most come from answers a person reviewed. The raw
metrics are computed at request time and include every outcome posted
so far. See calibration.
On a prefix path the report pools the outcomes of every namespace under the prefix, which is the fit inherited answers use.
Authorizations
An organization API key. Keys carry a role (read_write or
read_only) and may be restricted to a namespace prefix such as
acme/*, or to one namespace such as acme/prod/tenant_1. A prefix
matches on a / boundary: acme/prod/tenant_1* covers
acme/prod/tenant_1 and everything under acme/prod/tenant_1/,
never acme/prod/tenant_12.
Path Parameters
A namespace name, or a template prefix ending in /*, with any
/ sent as %2F: acme%2Fprod%2Ftenant_123 or acme%2Fprod%2F*.
A namespace name, or a template prefix: a namespace path ending in
/*, such as acme/prod/*, which every namespace under acme/prod/
inherits judgments from. Up to 256 bytes. Never exactly
. or .., which a URL path can't carry.
1 - 256^(?!\.\.?$)[A-Za-z0-9._:/-]+(/\*)?$A judgment, attribute or threshold name. Names are path segments in field references, so they never contain ..
1 - 128^[A-Za-z0-9_-]+$Response
The report.
How well the active version's answers match their outcomes.
epochs holds one entry per engine epoch with outcomes, newest first.
An epoch is the engine_version its evaluations were computed under:
each Jev epoch label, or the exact version for a pinned engine.
A judgment, attribute or threshold name. Names are path segments in field references, so they never contain ..
1 - 128^[A-Za-z0-9_-]+$bool, choice, score The active version the report is for.
x >= 1Outcomes joined to evaluations of this version, across every epoch, implicit negatives included.
x >= 0Joined outcomes by where they came from, summing to outcomes.
Rule outcomes left out because the organization's plan did not include outcome rules when they were observed. They are kept, and count again only for outcomes observed while the plan includes rules. 0 on Team and Scale throughout.
x >= 0Outcomes, posted or derived by rules, that add no sample to the report's version, by reason. They are kept, and count if a later evaluation or window takes them.
Prediction windows with no outcome that are not counted as implicit
negatives. All are 0 unless outcomes.implicit_negatives is
on.
How much calibration has cut this version's error since the
learning loop started, on held-out outcomes. On every
plan; the dashboard shows its history on Team and above, and its
headline as a preview on Developer. Null for a composite, whose
p is its combiner's, and on a template's tenant: the template's
report has the pool's lift.
For an options.from judgment, how often the pick equals the outcome where a match existed. Reported, never fitted.
For an options.from judgment, outcomes naming a
document its evaluation did not read, by cause: created after the
evaluation's watermark, cut by last_n (the number that says
whether the oldest candidates matter), not in the block (a bad
key), or the evaluation failed.