See what each model is actually good at.
Browse reported results benchmark by benchmark. Turn the view around to see each model's capability shape—and what extra reasoning effort buys.
LAB-REPORT VIEWTerminal-Bench 2.1
OpenAI reportOne published tableSetup varies by rowHigher is better ↑
Observations1302raw reported scores
Configurations104effort kept explicit
Benchmarks166capability lenses
Lab reports18first-party sources only
The leaders, one category at a time.
Each row names a representative benchmark and the model with its highest published-value median. The low-to-high range stays visible; no scores are mixed across different benchmarks.
Professional workAgents' Last ExamGPT-5.6 SolOpenAI · 1 value · 1 report52.7%52.7%CodingSWE-Bench ProClaude Mythos 5Anthropic · 1 value · 1 report80.3%80.3%Agentic systemsTerminal-Bench 2.1GPT-5.6 SolOpenAI · 3 values · 2 reports88.8%88.8%–91.9%ScienceGPQA DiamondClaude Mythos PreviewAnthropic · 1 value · 1 report94.6%94.6%VisionMMMU Pro (no tools)Gemini 3.5 FlashGoogle · 1 value · 1 report83.6%83.6%CybersecurityCapture-the-Flag ChallengesGPT-5.6 SolOpenAI · 1 value · 1 report96.7%96.7%Hard reasoningHLE-FullClaude Mythos PreviewAnthropic · 1 value · 1 report56.8%56.8%Long contextOpenAI MRCR v2 · 8-needle · 512K-1MGPT-5.5OpenAI · 1 value · 1 report74.0%74.0%
Snapshot 2026-08-07-first-party-v4 · reported through Aug 7, 2026
Compare every model →