Frontier model evaluation · 2026
GPT-6 Astra benchmark ledger
Reported scores · higher is better
— = not reported
| Evaluation | GPT-6Astra | GPT-5.6Sol[2] | ClaudeFable 5.1 | ClaudeFable 5 | ClaudeOpus 5 | Gemini3.8 Flash |
|---|---|---|---|---|---|---|
| ARC-AGI-3 | 98.6% [1] | 7.8% | — | — | 30.2%(high) | — |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | — |
| Agents’ Last Exam | 59.3% | 52.7% | — | 48.7%(xhigh) | 52.7% | — |
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | — |
| BenchCAD | 95.9% | 83.3% | 84.3% [5] | 67.5% [5] | 82.1% [5] | — |
| DeepSWE v1.1 | 74.1% | 70.8% | 67.4% | 69.9% | 68.8% | 73.7% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 24.7% | 29.0% | — |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.2% | 95.3% |
| GeneBench Pro | 39.0% | 28.7% | — | — | — | — |
| MedChemBench (internal) | 49.7% | 47.4% | — | — | — | — |
| HealthBench Professional (length-adjusted) | 63.4% | 60.5% | 56.6% [9] | 60.9% | [57.5%] | 52.1% |
| ExploitBench | 100.0% | 78.5% | — | — | 70% | — |
| SRE-Bench (four attempts) | 99.2% | 68.7% | — | — | — | — |
| Auto-review circumvention (internal; lower is better) | 0% | 0.29% | — | — | — | — |
* Highlighted column marks the reported GPT-6 Astra result. Bracketed values and superscripts preserve source annotations.