See what each model is actually good at.

Browse reported results benchmark by benchmark. Turn the view around to see each model's capability shape—and what extra reasoning effort buys.

LAB-REPORT VIEWTerminal-Bench 2.1
OpenAI report
01

GPT-5.6 Sol · ultra91.9%

02

GPT-5.6 Sol88.8%

03

Claude Mythos 588.0%

04

GPT-5.6 Terra87.4%

05

GPT-5.585.6%

06

GPT-5.6 Luna84.7%

One published tableSetup varies by rowHigher is better ↑
Observations1302raw reported scores
Configurations104effort kept explicit
Benchmarks166capability lenses
Lab reports18first-party sources only

The leaders, one category at a time.

Each row names a representative benchmark and the model with its highest published-value median. The low-to-high range stays visible; no scores are mixed across different benchmarks.

Snapshot 2026-08-07-first-party-v4 · reported through Aug 7, 2026

Compare every model