"engine": {"name": "jev", "version": "current"}. current runs whichever Jev version the provider serves. Every answer records the epoch of model behaviour that produced it; a new epoch begins only when we detect a change, and then we re-check quality against the conformance suite.
Limits
Judging on this version is billed per judgment, by size class, the same as on every engine.
Data handling
Conformance
Our conformance suite ran on 2026-09-27 againstgolden-v1. Because current can change, we run the suite on it again every day and after every change we detect, and GET /engines returns the latest run with its date.
Score accuracy counts only the exact level. The answer was within one level 75.4% of the time, 0.94 levels off on average. What this means for each type: how far to trust each type.
On choice items where no option fits, it answered
none_of_the_above 70.7% of the time, and on the others 0.7%.
The golden set is deliberately hard: 40% easy, 40% medium and 20% hard items, where the hard ones are those people disagree on or that were built to mislead. Accuracy by difficulty: