Skip to main content
A backtest tells you how good a predictive judgment is before you use it live. You replay records that you already have, one period at a time. After each period, you label each new answer with what happened in the next period. Your history already holds the result, so you do not wait for it. In summary:
  1. Use a bool judgment with the default horizon of 0s. Put the period in the question.
  2. Read related records with last_n, not window.
  3. Write one period of history. Wait until each answer of that period exists.
  4. For each judged document that the period touched, post one labelled example: what happened in the next period, with observed_at now.
  5. Do the next period. When the replay is done, let the fit run, and read the calibration report.

Labelled examples, not a horizon

A labelled example is an outcome with a horizon of 0s. It joins the answer that was current at its observed_at. Thus, a label whose observed_at is after a period’s answers exist, and before the next period’s writes, joins the answer for that period. This is the backtest. You cannot backtest a judgment with a horizon. Its prediction window opens when Brussle makes the answer, and it closes one horizon later. Brussle sets the time of each write when it gets the write, and the time of each answer when it makes the answer. Nothing lets you date a write or an answer in the past. created_at dates only the creation of a document. Thus, each replayed answer opens its window now, and an outcome from your history is before that window. To measure a horizon judgment, you must wait for the horizon to pass. Thus, put the period in the question, and keep horizon at 0s. For example: “Will this machine have a failure in the next 90 days?” The label says whether it did.

Read each record’s past, not the time of the replay

A replay writes all of your history now. Send each record’s original created_at. Then relations sort your records in the order that they happened. But the two bounds of a relation behave differently:
  • last_n reads by created_at. With last_n: 10, a machine reads the 10 newest reports that exist at that point of the replay. Use last_n.
  • window counts back from the time of a write. It counts back from the later of two times: the newest write to the judged document, at the time that Brussle got it, and the newest created_at among its related documents. In a replay, the newest write is now. Thus, a window of 365d holds no record created more than 365 days before the replay, and early periods read nothing. Do not use window on a replay. See newest created first.
To count more records than the engine reads as text, add a second relation on the same records, with a larger last_n (up to 1,000) and aggregate. Name its numbers in features. "render": false keeps them out of what the engine reads. A blocking relation reads the newest records of its block. In a replay, a whole period arrives at one time. Thus, a record early in the period reads the records after it, which it could not have seen. Add "relative": "before". Then each record reads the last_n records just before it by created_at, the oldest first. A relation with relative takes no window and no order.

A worked outline

This example backtests a judgment on machines and their service reports. Each report has its machine’s id in attributes.machine_id. The history covers several quarters. 1. Make a namespace for the backtest. Set a budget on it. A backtest is billed as live judging: Brussle judges each touched machine again in each period. 2. Create the judgment before the first period. Thus, Brussle judges each period as you write it.
POST /v1/namespaces/acme%2Fbacktest/judgments
  • If the create returns a replay estimate, nothing is created yet. In a new namespace the estimate is almost zero. Send the create again with confirm: true.
  • Wait for its reference_index jobs before the first period.
  • A judgment with features has a debounce of 10 minutes by default. A short debounce_ms makes each period’s answers arrive sooner.
3. Replay each quarter, in order.
Python
  • quarters, documents_created_in, machines_touched_in, failed_in and next_quarter stand for your own code.
  • A post of outcomes is at most 2 MB. Split a large period into more posts.
4. Fit, then read the report. See the fit and read the ranking.

Wait for each period’s answers

Before you write the next period, wait until each answer of this period exists. Otherwise, a machine’s answer can include records of the next period, and no answer shows the machine as it was at the end of this period. Your label then joins an older answer.
  • Compare revisions, as in the outline. The import returns the highest revision that it wrote. When the judgment’s judged_through is at or above it, each machine that the period touched has its answer, or its last attempt failed. A get of the judgment reads no documents. See is a judgment caught up?.
  • A pause stops the wait. judged_through does not pass a write that waits for judging to start again, or for the rolling limit. If the budget runs out, the loop waits at that period until you raise the budget. Thus, set a budget that is high enough for every period.
  • Poll a query for answers that are pending to see which machines are not judged yet. When it returns no row, the period is judged, or its answers wait and read stale.
  • wait_for on a write waits for answers too, but it does less. It waits only for the answers of the documents in that write, not for a machine whose reports you wrote. It waits for at most 10 seconds, and it judges only the first 16 documents of the write at once. See limits.

When to set observed_at

A label joins the answer that was current at its observed_at. That is the answer for the newest write that touched the document at that time: a write to the document itself or to one of its related records.
  • Thus, observed_at must be at or after the updated_at of the period’s last write that touched the document. If it is earlier, the label joins an older answer, or no answer. The report counts a label with no answer in unmatched_outcomes.
  • observed_at now, taken after the wait, meets this rule. Use a clock that is in sync.
  • Do not use the date from your history as observed_at. That date is before the replay, so the label joins no answer.

One label for each answer

An answer keeps one label. A later label for the same answer replaces the earlier label: the label with the latest observed_at counts, and of two with the same observed_at, the one that you posted last. The report counts each replaced label in unmatched_outcomes.same_window.
  • Post one label for each touched document in each period. A document that the period did not touch keeps its answer. A second label for it replaces the label of the earlier period.
  • To correct a label, post the new value as of now, before you write the next period. After the next period’s writes, a label as of now joins the newer answer.

The fit

The fit is what turns your labels into a calibrated p (calibration).
  • While the current epoch has no fit, Brussle tries to fit it each hour. The first run after the outcomes meet the minimums, 100 outcomes with at least 20 of each kind, makes the fit. After that, the fit runs each day.
  • To fit now, send POST /namespaces/{ns}/judgments/{name}/calibration/fit, or select Fit now on the judgment’s Calibration tab. You can ask one time each hour for each judgment. In that hour, the call returns rate_limited (429), with Retry-After. See when fits run.
Expect the fit to move p. On a judgment with features, p is the calibrated value, which uses the features and the engine’s answer together. When a fit is chosen:
  • p changes for each answer, with no new judging.
  • Documents can cross your thresholds on p. GET on the judgment shows a fit_changed warning until you set a threshold again. Pick new thresholds with the threshold recommender.
  • Brussle judges again the judgments that read this one, and bills those re-judges. The event judgment.fit_applied reports them.

Read the ranking

A predictive judgment usually has a low base rate. For example, few machines fail in a quarter. Then accuracy in the calibration report says little. It counts p of 0.5 or more as true, so an answer of “no” for each machine is correct for most machines. Read the ranking numbers in each epoch’s held_out_metrics instead. The report gives them for the engine’s p (raw) and for each fit (scalar, and with features features), on held-out outcomes:
  • how well p puts the documents that turned out true above the documents that did not;
  • precision_at_5 and precision_at_10: the part of the top 5% and the top 10% of documents by p that turned out true.
With features, features_only scores a fit on your features alone, without the engine’s answer. Compare it with features to see what the engine adds to your own numbers. See how well answers rank. Compare the precision with base_rate. If 2% of machines fail, and 20% of the top 5% fail, the judgment finds failures at 10 times the base rate.