"engine": {"name": "gpt-6-luna", "version": "current"}. current runs whichever version of gpt-6-luna the provider serves. The provider can change that version without notice.
Each answer records the epoch in which it was made. An epoch is a period in which the behaviour of current does not change. A new epoch begins only when we detect a change. After each change, we re-check quality against the conformance suite.
Limits
Each request to this version carries the compiled context of one document and one or more questions. The compiled context is the parts of the document, and of its related documents, that the judgment’s context recipe selects.
Probabilities. Brussle keeps each probability as gpt-6-luna gives it, and scales the probabilities of a
choice or a score so that they add up to 1. OpenAI does not say how finely gpt-6-luna rounds its probabilities. As of 2026-10-07, gpt-6-luna gives every probability in steps of 0.01, as Jev does. Thus, many documents can have the same p, and a calibration gives each p one value. A calibration with features can separate them.
gpt-6-luna can decline a question. Then the evaluation of that judgment fails, with an error that says that the engine declined to answer. The other questions in the same request still get their answers.
OpenAI’s API for gpt-6-luna is a public beta.
You pay for judging on this version per judgment, by size class, times the weight of its engine. The weight of gpt-6-luna is 2.5. Thus, a standard judgment on this version counts as 2.5 judgments.
Changes and retirement
OpenAI can change the model behindcurrent at any time. We track each change as a new epoch.
- Epochs. Each answer records its epoch in
engine_version, for examplecurrent+2026-09-24.1. The date is the day (UTC) that the epoch began. The number after the date goes up by 1 at each new epoch. Before we record the first epoch,engine_versioniscurrent+pendingfor a short time. A new epoch does not make an answer stale. To find the answers of one epoch, filter a query onanswers.<judgment>.engine_version. - When
currentchanges. We monitor the behaviour ofcurrentall the time. When the behaviour changes, we start a new epoch. Each answer after that records the new epoch. Answers from before keep their epoch. A deployment or a restart on our side does not start a new epoch. - Events. At each new epoch, each judgment that uses
currentgets onejudgment.engine_changedevent. The event gives the old and the new epoch, and the time that we detected the change. - Calibration. At each new epoch, we fit your calibration again. For this fit, we send the stored contexts of answers that have recent outcomes to the changed engine. See when the engine changes.
- Retirement. We never retire
current. A change in the engine starts a new epoch instead. If OpenAI stops serving gpt-6-luna completely, judging withcurrentfails withengine_version_unavailable. Then we send an email to your organization’s owners and admins, if your organization judged withcurrentin the last 60 days. We never move a judgment to a different engine version for you.
Data handling
We send your data to OpenAI’s API in the cases below. Each case sends it to the same provider, on the same terms.
- Judging. For each document that this engine judges, we send the document’s compiled context and the question of each judgment. The compiled context includes any images that the judgment’s context recipe selects.
- Calibration at a new epoch. We send some stored compiled contexts once more, with the images that they name and the question of each judgment. These are the contexts of answers that have recent outcomes. We use the results to fit your calibration again for the new epoch.
- Recipe tuning. If your plan includes recipe tuning, we send smaller versions of those stored contexts, with the images that they name and the question of each judgment. We use the results to test cheaper context recipes.
Conformance
The conformance suite measures the quality of an engine version. It asks the version a fixed set of questions, the golden set. Each question is about one document. People labelled the correct answer to each question. Our conformance suite ran on 2026-10-07 againstgolden-v1, the version of the golden set. This page shows the results of that run.
current can change. Thus, we run the suite on it again every day and after every change we detect. Each of these runs also includes the relational suite.
This page does not show these later runs. GET /engines returns the latest run of the golden set, with its date. It returns the latest run also when that run did worse.
If a run of either suite finds that current does worse than before, we record the failure. The run never blocks your answers and never re-judges them. current stays active.
If the conformance run after a change does worse than the last run that passed, each judgment that uses current gets a judgment.engine_regressed event. The event gives each measure that did worse.
- Accuracy is the share of questions where the most likely answer was the correct answer.
- Expected calibration error shows how far the probabilities are from how often the answers were correct. 0 is perfect.
- Log loss grows when the engine gives a low probability to the correct answer. Lower is better.
- Option-order flips is the share of
choiceanswers that changed when the options were listed in a different order.
score, accuracy counts only an answer on the exact level. 71.7% of score answers were within one level of the correct level. On average, an answer was 1.02 levels from the correct level.
For choice questions where no option fits, this version answered none_of_the_above 69.3% of the time. For the other choice questions, it answered none_of_the_above 0.9% of the time.
This run asked each question 3 times. Each time, this version gave the same probabilities. The provider does not promise this.
To learn what these results mean for each type of judgment, see how far to trust each type.
The golden set is hard on purpose. 40% of its questions are easy, 40% are medium and 20% are hard. A hard question is one that people disagree on, or one that was written to mislead. This table shows the accuracy at each difficulty:
The latency table shows how long engine requests took. Batch 1 is one request sent alone. Batch 32 is 32 requests sent at the same time. Its time ends when the last of the 32 answers comes back. p50 is the median time. 90% of the times were shorter than p90.
The relational suite
The relational suite checks that the quality of the engine does not drop from one epoch to the next. Its questions are about documents that the engine reads with their related documents. A run passes when its quality is close to the quality of an earlier run on the same questions, or better. This page shows the run of 2026-10-07 againstrelational-v1, the version of the relational suite. GET /engines does not return the results of this suite. The suite has three tasks:
- Retention. The engine reads a user and the user’s posts. It gives the probability that the user is still active 6 to 12 months after sign-up.
- Civility. The engine reads the replies in a conversation so far. It gives the probability that the next reply is a personal attack or uncivil.
- Matching. The engine reads a document and the candidates that share a key with it. It picks the candidate that is the same thing as the document, or
none_of_the_above.
features_only shows what the engine adds to the counts.
The results for retention and civility use AUROC. AUROC compares two random cases: one where the outcome was yes, and one where the outcome was no. AUROC is the chance that the engine gives the higher probability to the yes case. 0.5 is chance. 1.0 is perfect.
The 95% interval shows how precise each result is. A narrower interval is more precise.
Matching accuracy counts two kinds of correct pick: the matching candidate, or
none_of_the_above when no candidate matches. When a candidate matched, the engine picked a matching candidate 97.5% of the time. When a candidate matched, the engine picked none_of_the_above 1.3% of the time. When no candidate matched, the engine picked none_of_the_above 78.3% of the time. 5.9% of picks changed when the candidates were listed in reverse order.