> ## Documentation Index
> Fetch the complete documentation index at: https://docs.brussle.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Backtest a predictive judgment on history

> Replay the records that you already have one period at a time, label each answer with what happened next, and see how well the judgment ranks before you use it live.

A backtest tells you how good a predictive judgment is before you use it live. You replay records that you already have, one period at a time. After each period, you label each new answer with what happened in the next period. Your history already holds the result, so you do not wait for it.

In summary:

1. Use a `bool` judgment with the default `horizon` of `0s`. Put the period in the question.
2. Read related records with `last_n`, not `window`.
3. Write one period of history. Wait until each answer of that period exists.
4. For each judged document that the period touched, post one labelled example: what happened in the next period, with `observed_at` now.
5. Do the next period. When the replay is done, let the fit run, and read the calibration report.

## Labelled examples, not a horizon

A [labelled example](/guides/measure-improve-tune#post-labelled-examples) is an outcome with a `horizon` of `0s`. It joins the answer that was current at its `observed_at`. Thus, a label whose `observed_at` is after a period's answers exist, and before the next period's writes, joins the answer for that period. This is the backtest.

You cannot backtest a judgment with a `horizon`. Its [prediction window](/concepts/calibration#predictions-with-a-horizon) opens when Brussle makes the answer, and it closes one `horizon` later. Brussle sets the time of each write when it gets the write, and the time of each answer when it makes the answer. Nothing lets you date a write or an answer in the past. `created_at` dates only the creation of a document. Thus, each replayed answer opens its window now, and an outcome from your history is before that window. To measure a horizon judgment, you must wait for the horizon to pass.

Thus, put the period in the question, and keep `horizon` at `0s`. For example: "Will this machine have a failure in the next 90 days?" The label says whether it did.

## Read each record's past, not the time of the replay

A replay writes all of your history now. Send each record's original `created_at`. Then relations sort your records in the order that they happened. But the two bounds of a relation behave differently:

* **`last_n` reads by `created_at`.** With `last_n: 10`, a machine reads the 10 newest reports that exist at that point of the replay. Use `last_n`.
* **`window` counts back from the time of a write.** It counts back from the later of two times: the newest write to the judged document, at the time that Brussle got it, and the newest `created_at` among its related documents. In a replay, the newest write is now. Thus, a `window` of `365d` holds no record created more than 365 days before the replay, and early periods read nothing. Do not use `window` on a replay. See [newest created first](/guides/related-documents#newest-created-first-and-what-an-edit-costs).

To count more records than the engine reads as text, add a second relation on the same records, with a larger `last_n` (up to 1,000) and `aggregate`. Name its numbers in `features`. `"render": false` keeps them out of what the engine reads.

**A blocking relation reads the newest records of its block.** In a replay, a whole period arrives at one time. Thus, a record early in the period reads the records after it, which it could not have seen. Add `"relative": "before"`. Then each record reads the `last_n` records just before it by `created_at`, the oldest first. A relation with `relative` takes no `window` and no `order`.

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "related": {
    "earlier_reports": {
      "match": {"attributes.kind": "report"},
      "join": {"theirs": "attributes.machine_id", "mine": "attributes.machine_id"},
      "last_n": 10,
      "relative": "before",
      "fields": ["created_at", "state.summary"]
    }
  }
}
```

## A worked outline

This example backtests a judgment on machines and their service reports. Each report has its machine's `id` in `attributes.machine_id`. The history covers several quarters.

**1. Make a namespace for the backtest.** Set a [budget](/guides/import-existing-data#set-a-budget-first) on it. A backtest is billed as live judging: Brussle judges each touched machine again in each period.

**2. Create the judgment before the first period.** Thus, Brussle judges each period as you write it.

```json POST /v1/namespaces/acme%2Fbacktest/judgments theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "name": "fails_next_quarter",
  "type": "bool",
  "applies_to": {"attributes.kind": "machine"},
  "question": "Will this machine have a failure in the next 90 days?",
  "context": {
    "fields": ["state.model", "state.site"],
    "related": {
      "reports": {
        "match": {"attributes.kind": "report"},
        "join": {"theirs": "attributes.machine_id", "mine": "id"},
        "last_n": 10,
        "fields": ["created_at", {"path": "state.summary", "max_chars": 400}]
      },
      "report_counts": {
        "match": {"attributes.kind": "report"},
        "join": {"theirs": "attributes.machine_id", "mine": "id"},
        "last_n": 1000,
        "aggregate": {"count": true, "max": ["state.downtime_hours"], "render": false}
      }
    }
  },
  "features": ["report_counts.count", "report_counts.max(state.downtime_hours)"],
  "engine": {"name": "jev", "version": "current"},
  "freshness": {"policy": "on_change", "debounce_ms": 5000}
}
```

* If the create returns a [replay estimate](/guides/related-documents#the-replay-estimate), nothing is created yet. In a new namespace the estimate is almost zero. Send the create again with `confirm: true`.
* Wait for its `reference_index` jobs before the first period.
* A judgment with features has a debounce of 10 minutes by default. A short `debounce_ms` makes each period's answers arrive sooner.

**3. Replay each quarter, in order.**

```python Python theme={"theme":{"light":"css-variables","dark":"css-variables"}}
import time
from datetime import datetime, timezone

J = "fails_next_quarter"

for quarter in quarters[:-1]:  # the last quarter has no next quarter to label it
    written = ns.import_documents(documents_created_in(quarter))  # machines and reports, with their created_at

    # Wait until the judgment has judged through the quarter's last write.
    while ns.judgments.get(J)["judged_through"] < written.revision:
        time.sleep(5)

    # One label for each machine that this quarter touched: what happened in the next quarter.
    now = datetime.now(timezone.utc).isoformat()
    ns.outcomes.append([
        {"document_id": machine, "judgment": J, "value": failed_in(machine, next_quarter(quarter)), "observed_at": now}
        for machine in machines_touched_in(quarter)
    ])
```

* `quarters`, `documents_created_in`, `machines_touched_in`, `failed_in` and `next_quarter` stand for your own code.
* A post of outcomes is at most 2 MB. Split a large period into more posts.

**4. Fit, then read the report.** See [the fit](#the-fit) and [read the ranking](#read-the-ranking).

## Wait for each period's answers

Before you write the next period, wait until each answer of this period exists. Otherwise, a machine's answer can include records of the next period, and no answer shows the machine as it was at the end of this period. Your label then joins an older answer.

* **Compare revisions**, as in the outline. The import returns the highest `revision` that it wrote. When the judgment's `judged_through` is at or above it, each machine that the period touched has its answer, or its last attempt failed. A get of the judgment reads no documents. See [is a judgment caught up?](/concepts/freshness#is-a-judgment-caught-up).
* **A pause stops the wait.** `judged_through` does not pass a write that waits for judging to start again, or for the [rolling limit](/concepts/freshness#deferral-and-catch-up). If the [budget](/guides/import-existing-data#set-a-budget-first) runs out, the loop waits at that period until you raise the budget. Thus, set a budget that is high enough for every period.
* **Poll a query** for answers that are `pending` to see which machines are not judged yet. When it returns no row, the period is judged, or its answers wait and read `stale`.
* **`wait_for`** on a write waits for answers too, but it does less. It waits only for the answers of the documents in that write, not for a machine whose reports you wrote. It waits for at most 10 seconds, and it judges only the first 16 documents of the write at once. See [limits](/limits).

## When to set `observed_at`

A label joins the answer that was current at its `observed_at`. That is the answer for the newest write that touched the document at that time: a write to the document itself or to one of its related records.

* Thus, `observed_at` must be at or after the `updated_at` of the period's last write that touched the document. If it is earlier, the label joins an older answer, or no answer. The report counts a label with no answer in `unmatched_outcomes`.
* `observed_at` now, taken after the wait, meets this rule. Use a clock that is in sync.
* Do not use the date from your history as `observed_at`. That date is before the replay, so the label joins no answer.

## One label for each answer

An answer keeps one label. A later label for the same answer replaces the earlier label: the label with the latest `observed_at` counts, and of two with the same `observed_at`, the one that you posted last. The report counts each replaced label in `unmatched_outcomes.same_window`.

* Post one label for each touched document in each period. A document that the period did not touch keeps its answer. A second label for it replaces the label of the earlier period.
* To correct a label, post the new value as of now, before you write the next period. After the next period's writes, a label as of now joins the newer answer.

## The fit

The fit is what turns your labels into a calibrated `p` ([calibration](/concepts/calibration#how-the-fit-works)).

* While the current epoch has no fit, Brussle tries to fit it each hour. The first run after the outcomes meet the [minimums](/concepts/calibration#at-least-20-of-each-kind), 100 outcomes with at least 20 of each kind, makes the fit. After that, the fit runs each day.
* To fit now, send `POST /namespaces/{ns}/judgments/{name}/calibration/fit`, or select **Fit now** on the judgment's **Calibration** tab. You can ask one time each hour for each judgment. In that hour, the call returns `rate_limited` (429), with `Retry-After`. See [when fits run](/concepts/calibration#how-the-fit-works).

**Expect the fit to move `p`.** On a judgment with features, `p` is the [calibrated value](/concepts/calibration#judgments-with-features), which uses the features and the engine's answer together. When a fit is chosen:

* `p` changes for each answer, with no new judging.
* Documents can cross your thresholds on `p`. `GET` on the judgment shows a `fit_changed` warning until you set a threshold again. Pick new thresholds with the [threshold recommender](/guides/measure-improve-tune#pick-thresholds-with-the-recommender).
* Brussle judges again the judgments that read this one, and bills those re-judges. The event `judgment.fit_applied` reports them.

## Read the ranking

A predictive judgment usually has a low base rate. For example, few machines fail in a quarter. Then `accuracy` in the [calibration report](/guides/measure-improve-tune#read-the-calibration-report) says little. It counts `p` of 0.5 or more as `true`, so an answer of "no" for each machine is correct for most machines.

Read the ranking numbers in each epoch's `held_out_metrics` instead. The report gives them for the engine's `p` (`raw`) and for each fit (`scalar`, and with features `features`), on held-out outcomes:

* how well `p` puts the documents that turned out `true` above the documents that did not;
* `precision_at_5` and `precision_at_10`: the part of the top 5% and the top 10% of documents by `p` that turned out `true`.

With features, `features_only` scores a fit on your features alone, without the engine's answer. Compare it with `features` to see what the engine adds to your own numbers. See [how well answers rank](/concepts/calibration#how-well-answers-rank).

Compare the precision with `base_rate`. If 2% of machines fail, and 20% of the top 5% fail, the judgment finds failures at 10 times the base rate.


## Related topics

- [Measure, improve, tune](/guides/measure-improve-tune.md)
- [Judge a document with its related documents](/guides/related-documents.md)
- [Import existing data](/guides/import-existing-data.md)
- [Export evaluation history](/api-reference/documents/export-evaluation-history.md)
- [Export audit and evaluation history](/guides/export-history.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.