> ## Documentation Index
> Fetch the complete documentation index at: https://docs.brussle.com/llms.txt
> Use this file to discover all available pages before exploring further.

# gpt-6-luna current

> Limits, changes, data handling and conformance results for gpt-6-luna current.

An engine is the AI model that makes the answers for your judgments. Each answer records the engine version that made it. Thus, after the engine changes, you can tell old answers from new answers.

Use this version with `"engine": {"name": "gpt-6-luna", "version": "current"}`. `current` runs whichever version of `gpt-6-luna` the provider serves. The provider can change that version without notice.

Each answer records the epoch in which it was made. An epoch is a period in which the behaviour of `current` does not change. A new epoch begins only when we detect a change. After each change, we re-check quality against the conformance suite.

## Limits

Each request to this version carries the compiled context of one document and one or more questions. The compiled context is the parts of the document, and of its related documents, that the judgment's [context recipe](/guides/context-recipes) selects.

| | |
| - | - |
| Context | 32,000 tokens for the compiled context plus the longest question |
| Whole request | 64,000 tokens for the compiled context plus every question |
| Choice options | 255, including `none_of_the_above`, which every `choice` has |
| Score levels | 2 to 10 |
| Questions per request | 32 |
| [Images](/guides/images) | Up to 16 per request. Each image counts toward the context by its size, so 16 images of the largest size do not fit. |
| Tokenizer | Approximate. Our token counts can differ a little from the provider's counts. |

**Probabilities.** Brussle keeps each probability as gpt-6-luna gives it, and scales the probabilities of a `choice` or a `score` so that they add up to 1. OpenAI does not say how finely gpt-6-luna rounds its probabilities. As of 2026-10-07, gpt-6-luna gives every probability in steps of 0.01, as Jev does. Thus, many documents can have the same `p`, and a [calibration](/concepts/calibration#what-calibration-can-and-cannot-do) gives each `p` one value. A calibration with [features](/concepts/calibration#judgments-with-features) can separate them.

**gpt-6-luna can decline a question.** Then the evaluation of that judgment fails, with an error that says that the engine declined to answer. The other questions in the same request still get their answers.

**OpenAI's API for gpt-6-luna is a public beta.**

You pay for judging on this version per [judgment](/pricing#judgments), by size class, times the weight of its engine. The weight of gpt-6-luna is 2.5. Thus, a standard judgment on this version counts as 2.5 judgments.

## Changes and retirement

OpenAI can change the model behind `current` at any time. We track each change as a new epoch.

* **Epochs.** Each answer records its epoch in `engine_version`, for example `current+2026-09-24.1`. The date is the day (UTC) that the epoch began. The number after the date goes up by 1 at each new epoch. Before we record the first epoch, `engine_version` is `current+pending` for a short time. A new epoch does not make an answer stale. To find the answers of one epoch, filter a query on `answers.<judgment>.engine_version`.
* **When `current` changes.** We monitor the behaviour of `current` all the time. When the behaviour changes, we start a new epoch. Each answer after that records the new epoch. Answers from before keep their epoch. A deployment or a restart on our side does not start a new epoch.
* **Events.** At each new epoch, each judgment that uses `current` gets one `judgment.engine_changed` [event](/guides/events-feed#event-types). The event gives the old and the new epoch, and the time that we detected the change.
* **Calibration.** At each new epoch, we fit your calibration again. For this fit, we send the stored contexts of answers that have recent outcomes to the changed engine. See [when the engine changes](/concepts/calibration#when-the-engine-changes).
* **Retirement.** We never retire `current`. A change in the engine starts a new epoch instead. If OpenAI stops serving gpt-6-luna completely, judging with `current` fails with `engine_version_unavailable`. Then we send an email to your organization's owners and admins, if your organization judged with `current` in the last 60 days. We never move a judgment to a different engine version for you.

## Data handling

| | |
| - | - |
| Hosted by | OpenAI, in the US |
| Region | us |
| Subprocessor | OpenAI operates gpt-6-luna. The data processing agreement lists all subprocessors. |

We send your data to OpenAI's API in the cases below. Each case sends it to the same provider, on the same terms.

* **Judging.** For each document that this engine judges, we send the document's compiled context and the question of each judgment. The compiled context includes any [images](/guides/images) that the judgment's context recipe selects.
* **Calibration at a new epoch.** We send some stored compiled contexts once more, with the images that they name and the question of each judgment. These are the contexts of answers that have recent outcomes. We use the results to fit your [calibration](/concepts/calibration#when-the-engine-changes) again for the new epoch.
* **Recipe tuning.** If your plan includes recipe tuning, we send smaller versions of those stored contexts, with the images that they name and the question of each judgment. We use the results to test cheaper [context recipes](/guides/context-recipes#let-your-outcomes-tune-the-recipe).

After you delete a document, we do not send its stored contexts again.

The data processing agreement lists each subprocessor that your data passes through. Each subprocessor's terms in that agreement set how long it keeps your data. See [where your data goes](/behavior#where-your-data-goes).

## Conformance

The conformance suite measures the quality of an engine version. It asks the version a fixed set of questions, the golden set. Each question is about one document. People labelled the correct answer to each question.

Our conformance suite ran on 2026-10-07 against `golden-v1`, the version of the golden set. This page shows the results of that run.

`current` can change. Thus, we run the suite on it again every day and after every change we detect. Each of these runs also includes the [relational suite](#the-relational-suite).

This page does not show these later runs. `GET /engines` returns the latest run of the golden set, with its date. It returns the latest run also when that run did worse.

If a run of either suite finds that `current` does worse than before, we record the failure. The run never blocks your answers and never re-judges them. `current` stays `active`.

If the conformance run after a change does worse than the last run that passed, each judgment that uses `current` gets a `judgment.engine_regressed` [event](/guides/events-feed#event-types). The event gives each measure that did worse.

| Type | Accuracy | Expected calibration error | Log loss | Option-order flips |
| - | - | - | - | - |
| bool | 78.0% | 0.119 | 0.628 | |
| choice | 78.4% | 0.108 | 0.913 | 5.6% |
| score | 31.3% | 0.444 | 3.033 | |

* **Accuracy** is the share of questions where the most likely answer was the correct answer.
* **Expected calibration error** shows how far the probabilities are from how often the answers were correct. 0 is perfect.
* **Log loss** grows when the engine gives a low probability to the correct answer. Lower is better.
* **Option-order flips** is the share of `choice` answers that changed when the options were listed in a different order.

For `score`, accuracy counts only an answer on the exact level. 71.7% of `score` answers were within one level of the correct level. On average, an answer was 1.02 levels from the correct level.

For `choice` questions where no option fits, this version answered `none_of_the_above` 69.3% of the time. For the other `choice` questions, it answered `none_of_the_above` 0.9% of the time.

This run asked each question 3 times. Each time, this version gave the same probabilities. The provider does not promise this.

To learn what these results mean for each type of judgment, see [how far to trust each type](/concepts/judgments#how-far-to-trust-each-type).

The golden set is hard on purpose. 40% of its questions are easy, 40% are medium and 20% are hard. A hard question is one that people disagree on, or one that was written to mislead. This table shows the accuracy at each difficulty:

| Type | Easy | Medium | Hard |
| - | - | - | - |
| bool | 91.5% | 79.5% | 48.0% |
| choice | 91.0% | 75.5% | 59.0% |
| score | 38.7% | 30.5% | 18.0% |

The latency table shows how long engine requests took. Batch 1 is one request sent alone. Batch 32 is 32 requests sent at the same time. Its time ends when the last of the 32 answers comes back. p50 is the median time. 90% of the times were shorter than p90.

| Latency | p50 | p90 |
| - | - | - |
| Batch 1 | 128 ms | 242 ms |
| Batch 32 | 502 ms | 993 ms |

### The relational suite

The relational suite checks that the quality of the engine does not drop from one epoch to the next. Its questions are about documents that the engine reads with their related documents. A run passes when its quality is close to the quality of an earlier run on the same questions, or better.

This page shows the run of 2026-10-07 against `relational-v1`, the version of the relational suite. `GET /engines` does not return the results of this suite. The suite has three tasks:

* **Retention.** The engine reads a user and the user's posts. It gives the probability that the user is still active 6 to 12 months after sign-up.
* **Civility.** The engine reads the replies in a conversation so far. It gives the probability that the next reply is a personal attack or uncivil.
* **Matching.** The engine reads a document and the candidates that share a key with it. It picks the candidate that is the same thing as the document, or `none_of_the_above`.

These results are a reference point for later runs. They do not show how well the engine answers your own questions. Retention and civility ask what will happen next. On questions of this kind, keep counts as [features](/concepts/calibration#judgments-with-features), and let the fit on your outcomes weigh them. The calibration report's `features_only` shows what the engine adds to the counts.

The results for retention and civility use AUROC. AUROC compares two random cases: one where the outcome was yes, and one where the outcome was no. AUROC is the chance that the engine gives the higher probability to the yes case. 0.5 is chance. 1.0 is perfect.

<svg className="docs-chart" viewBox="0 0 480 124" role="img" aria-label="AUROC with its 95% interval: Retention 0.717, Civility 0.720. 0.5 is chance.">
  <g className="docs-chart-frame"><line className="docs-chart-grid" x1="88" x2="88" y1="20" y2="100" /><text className="docs-chart-tick" x="88" y="118" textAnchor="middle">0.4</text><line className="docs-chart-grid" x1="212" x2="212" y1="20" y2="100" /><text className="docs-chart-tick" x="212" y="118" textAnchor="middle">0.6</text><line className="docs-chart-grid" x1="336" x2="336" y1="20" y2="100" /><text className="docs-chart-tick" x="336" y="118" textAnchor="middle">0.8</text><line className="docs-chart-grid" x1="460" x2="460" y1="20" y2="100" /><text className="docs-chart-tick" x="460" y="118" textAnchor="middle">1.0</text><line className="docs-chart-axis" x1="88" x2="460" y1="100" y2="100" /><text className="docs-chart-tick" x="0" y="118">AUROC</text><line className="docs-chart-ref" x1="150" x2="150" y1="20" y2="100" /><text className="docs-chart-annot" x="156" y="18">chance</text></g>
  <g className="docs-chart-row"><title>Retention: AUROC 0.717, 95% interval 0.686 to 0.748</title><rect className="docs-chart-hit" x="0" y="30" width="480" height="28" /><text className="docs-chart-label" x="0" y="48">Retention</text><line className="docs-chart-range" x1="265.3" x2="303.9" y1="44" y2="44" /><line className="docs-chart-range" x1="265.3" x2="265.3" y1="40" y2="48" /><line className="docs-chart-range" x1="303.9" x2="303.9" y1="40" y2="48" /><circle className="docs-chart-point" cx="284.8" cy="44" r="4" /><text className="docs-chart-value" x="313.9" y="48">0.717</text></g>
  <g className="docs-chart-row"><title>Civility: AUROC 0.720, 95% interval 0.699 to 0.740</title><rect className="docs-chart-hit" x="0" y="66" width="480" height="28" /><text className="docs-chart-label" x="0" y="84">Civility</text><line className="docs-chart-range" x1="273.2" x2="298.7" y1="80" y2="80" /><line className="docs-chart-range" x1="273.2" x2="273.2" y1="76" y2="84" /><line className="docs-chart-range" x1="298.7" x2="298.7" y1="76" y2="84" /><circle className="docs-chart-point" cx="286.5" cy="80" r="4" /><text className="docs-chart-value" x="308.7" y="84">0.720</text></g>
</svg>

The 95% interval shows how precise each result is. A narrower interval is more precise.

| Task | Result | 95% interval |
| - | - | - |
| Retention | AUROC 0.717 | 0.686 to 0.748 |
| Civility | AUROC 0.720 | 0.699 to 0.740 |
| Matching | 96.4% accuracy | |

Matching accuracy counts two kinds of correct pick: the matching candidate, or `none_of_the_above` when no candidate matches. When a candidate matched, the engine picked a matching candidate 97.5% of the time. When a candidate matched, the engine picked `none_of_the_above` 1.3% of the time. When no candidate matched, the engine picked `none_of_the_above` 78.3% of the time. 5.9% of picks changed when the candidates were listed in reverse order.


## Related topics

- [Change a question or engine safely](/guides/change-a-judgment.md)
- [judgment.engine_changed](/api-reference/webhooks/judgmentengine_changed.md)
- [Engines](/engines/index.md)
- [Querying answers](/concepts/queries.md)
- [judgment.engine_regressed](/api-reference/webhooks/judgmentengine_regressed.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.