Frontier model evaluation · 2026

GPT-6 Astra benchmark ledger

Reported scores · higher is better
— = not reported

Evaluation GPT-6Astra GPT-5.6Sol[2] ClaudeFable 5.1 ClaudeFable 5 ClaudeOpus 5 Gemini3.8 Flash
ARC-AGI-398.6% [1]7.8%30.2%(high)
FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%
Agents’ Last Exam59.3%52.7%48.7%(xhigh)52.7%
AutomationBench41.4%18.1%31.4%17.4%26.9%
BenchCAD95.9%83.3%84.3% [5]67.5% [5]82.1% [5]
DeepSWE v1.174.1%70.8%67.4%69.9%68.8%73.7%
Terminal-Bench Science 0.164.6%22.4%52.6%24.7%29.0%
GPQA Diamond96.0%94.6%93.7%92.6%93.2%95.3%
GeneBench Pro39.0%28.7%
MedChemBench (internal)49.7%47.4%
HealthBench Professional
(length-adjusted)
63.4%60.5%56.6% [9]60.9%[57.5%]52.1%
ExploitBench100.0%78.5%70%
SRE-Bench (four attempts)99.2%68.7%
Auto-review circumvention
(internal; lower is better)
0%0.29%

* Highlighted column marks the reported GPT-6 Astra result. Bracketed values and superscripts preserve source annotations.