- Use a
booljudgment with the defaulthorizonof0s. Put the period in the question. - Read related records with
last_n, notwindow. - Write one period of history. Wait until each answer of that period exists.
- For each judged document that the period touched, post one labelled example: what happened in the next period, with
observed_atnow. - Do the next period. When the replay is done, let the fit run, and read the calibration report.
Labelled examples, not a horizon
A labelled example is an outcome with ahorizon of 0s. It joins the answer that was current at its observed_at. Thus, a label whose observed_at is after a period’s answers exist, and before the next period’s writes, joins the answer for that period. This is the backtest.
You cannot backtest a judgment with a horizon. Its prediction window opens when Brussle makes the answer, and it closes one horizon later. Brussle sets the time of each write when it gets the write, and the time of each answer when it makes the answer. Nothing lets you date a write or an answer in the past. created_at dates only the creation of a document. Thus, each replayed answer opens its window now, and an outcome from your history is before that window. To measure a horizon judgment, you must wait for the horizon to pass.
Thus, put the period in the question, and keep horizon at 0s. For example: “Will this machine have a failure in the next 90 days?” The label says whether it did.
Read each record’s past, not the time of the replay
A replay writes all of your history now. Send each record’s originalcreated_at. Then relations sort your records in the order that they happened. But the two bounds of a relation behave differently:
last_nreads bycreated_at. Withlast_n: 10, a machine reads the 10 newest reports that exist at that point of the replay. Uselast_n.windowcounts back from the time of a write. It counts back from the later of two times: the newest write to the judged document, at the time that Brussle got it, and the newestcreated_atamong its related documents. In a replay, the newest write is now. Thus, awindowof365dholds no record created more than 365 days before the replay, and early periods read nothing. Do not usewindowon a replay. See newest created first.
last_n (up to 1,000) and aggregate. Name its numbers in features. "render": false keeps them out of what the engine reads.
A blocking relation reads the newest records of its block. In a replay, a whole period arrives at one time. Thus, a record early in the period reads the records after it, which it could not have seen. Add "relative": "before". Then each record reads the last_n records just before it by created_at, the oldest first. A relation with relative takes no window and no order.
A worked outline
This example backtests a judgment on machines and their service reports. Each report has its machine’sid in attributes.machine_id. The history covers several quarters.
1. Make a namespace for the backtest. Set a budget on it. A backtest is billed as live judging: Brussle judges each touched machine again in each period.
2. Create the judgment before the first period. Thus, Brussle judges each period as you write it.
POST /v1/namespaces/acme%2Fbacktest/judgments
- If the create returns a replay estimate, nothing is created yet. In a new namespace the estimate is almost zero. Send the create again with
confirm: true. - Wait for its
reference_indexjobs before the first period. - A judgment with features has a debounce of 10 minutes by default. A short
debounce_msmakes each period’s answers arrive sooner.
Python
quarters,documents_created_in,machines_touched_in,failed_inandnext_quarterstand for your own code.- A post of outcomes is at most 2 MB. Split a large period into more posts.
Wait for each period’s answers
Before you write the next period, wait until each answer of this period exists. Otherwise, a machine’s answer can include records of the next period, and no answer shows the machine as it was at the end of this period. Your label then joins an older answer.- Compare revisions, as in the outline. The import returns the highest
revisionthat it wrote. When the judgment’sjudged_throughis at or above it, each machine that the period touched has its answer, or its last attempt failed. A get of the judgment reads no documents. See is a judgment caught up?. - A pause stops the wait.
judged_throughdoes not pass a write that waits for judging to start again, or for the rolling limit. If the budget runs out, the loop waits at that period until you raise the budget. Thus, set a budget that is high enough for every period. - Poll a query for answers that are
pendingto see which machines are not judged yet. When it returns no row, the period is judged, or its answers wait and readstale. wait_foron a write waits for answers too, but it does less. It waits only for the answers of the documents in that write, not for a machine whose reports you wrote. It waits for at most 10 seconds, and it judges only the first 16 documents of the write at once. See limits.
When to set observed_at
A label joins the answer that was current at its observed_at. That is the answer for the newest write that touched the document at that time: a write to the document itself or to one of its related records.
- Thus,
observed_atmust be at or after theupdated_atof the period’s last write that touched the document. If it is earlier, the label joins an older answer, or no answer. The report counts a label with no answer inunmatched_outcomes. observed_atnow, taken after the wait, meets this rule. Use a clock that is in sync.- Do not use the date from your history as
observed_at. That date is before the replay, so the label joins no answer.
One label for each answer
An answer keeps one label. A later label for the same answer replaces the earlier label: the label with the latestobserved_at counts, and of two with the same observed_at, the one that you posted last. The report counts each replaced label in unmatched_outcomes.same_window.
- Post one label for each touched document in each period. A document that the period did not touch keeps its answer. A second label for it replaces the label of the earlier period.
- To correct a label, post the new value as of now, before you write the next period. After the next period’s writes, a label as of now joins the newer answer.
The fit
The fit is what turns your labels into a calibratedp (calibration).
- While the current epoch has no fit, Brussle tries to fit it each hour. The first run after the outcomes meet the minimums, 100 outcomes with at least 20 of each kind, makes the fit. After that, the fit runs each day.
- To fit now, send
POST /namespaces/{ns}/judgments/{name}/calibration/fit, or select Fit now on the judgment’s Calibration tab. You can ask one time each hour for each judgment. In that hour, the call returnsrate_limited(429), withRetry-After. See when fits run.
p. On a judgment with features, p is the calibrated value, which uses the features and the engine’s answer together. When a fit is chosen:
pchanges for each answer, with no new judging.- Documents can cross your thresholds on
p.GETon the judgment shows afit_changedwarning until you set a threshold again. Pick new thresholds with the threshold recommender. - Brussle judges again the judgments that read this one, and bills those re-judges. The event
judgment.fit_appliedreports them.
Read the ranking
A predictive judgment usually has a low base rate. For example, few machines fail in a quarter. Thenaccuracy in the calibration report says little. It counts p of 0.5 or more as true, so an answer of “no” for each machine is correct for most machines.
Read the ranking numbers in each epoch’s held_out_metrics instead. The report gives them for the engine’s p (raw) and for each fit (scalar, and with features features), on held-out outcomes:
- how well
pputs the documents that turned outtrueabove the documents that did not; precision_at_5andprecision_at_10: the part of the top 5% and the top 10% of documents bypthat turned outtrue.
features_only scores a fit on your features alone, without the engine’s answer. Compare it with features to see what the engine adds to your own numbers. See how well answers rank.
Compare the precision with base_rate. If 2% of machines fail, and 20% of the top 5% fail, the judgment finds failures at 10 times the base rate.