Aggregate score
| Rank# | Model | Score | ±95% | $/task |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol | Score 55.4, interval ±2.3, accuracy 74.8% | 55.4 | |
| 2 | GPT-6 Luna | Score 26.8, interval ±2.6, accuracy 59.4% | 26.8 |
By axis
Every model on every axis, in aggregate order. Click a column to sort.
| 1GPT-6.1 Sol | 55.4 | 27.6 | 83.1 | $0.0052 |
| 2GPT-6 Luna | 26.8 | −2.7 | 56.3 | $0.00056 |
Knowledge
Closed-book knowledge of real entities as they get more obscure, and what a model does when it doesn't know.
| Rank# | Model | Score | ±95% | $/task |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol | Score 27.6, interval ±4.3, accuracy 54.6% | 27.6 | |
| 2 | GPT-6 Luna | Score −2.7, interval [−7.0, 1.7], accuracy 38.0% | −2.7 |
Math
Exact calculation and ballpark estimation without tools: how precise, and how wrong when it misses.
| Rank# | Model | Score | ±95% | $/task |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol | Score 83.1, interval ±1.5, accuracy 95.0% | 83.1 | |
| 2 | GPT-6 Luna | Score 56.3, interval ±2.7, accuracy 80.7% | 56.3 |
Coming soon
Planned axes and benchmarks. They have no results yet, so they are not part of any score above.
- DiscernmentComing soon
Noticing what a question doesn't say: missing information, contradictions and assumptions that don't follow.
- Common SenseComing soon
World-state tracking, physical and social plausibility, and invariance to framing.
- StrategyComing soon
Incentive analysis, game-theoretic decisions, operational optimization, and anticipation.
- ReasoningComing soon
Constraint satisfaction, rule induction, spatial reasoning, and long-chain deduction.
- ForecastingComing soon
Inferring resolved outcomes from a fixed prior information snapshot, with calibrated probabilities.
- ResearchComing soon · v0.2
Finding and checking information on the open web: what search adds, and whether answers stay grounded.
- CodingComing soon · v0.2
Writing, fixing and reviewing code against executable tests.