OverviewBrowse 284 families
9 categories · 284 families · 5696 reported values
Who leads, by category
Top five on one exact representative benchmark per category. Lists are independent — scores are never combined across benchmarks.
CodingClaude Opus 5.554.64%AgentsClaude Opus 5.566.40%Expert reasoningClaude Mythos Preview56.80%ScienceGPT-6 Astra68.1%Visual perceptionGPT-5.6 Sol59.7%Knowledge workClaude Opus 5.538.2%MultimodalGemini 3.5 Flash83.6%Exploit capabilityClaude Mythos 578.0%Tool useMuse Spark 1.188.1%
Coding
Repository code-change mergeability
- 1Claude Opus 5.5medium54.64%
- 2Claude Fable 5xhigh53.48%
- 3Claude Opus 5medium53.38%
- 4GPT-6 Astramax53.26%
- 5Claude Fable 5.1medium50.91%
Best open weights · same scale
- 12Kimi K3noneOW44.17%
Agents
Terminal agent work
- 1Claude Opus 5.5xhigh66.40%
- 2Mythos 5.1reported60.90%
- 3GPT-6 Astramax58.18%
- 4Claude Fable 5.1max57.88%
- 5Claude Opus 5xhigh53.94%
Best open weights · same scale
- 7GLM-5.3maxOW41.82%
Expert reasoning
Broad expert questions
- 1Claude Mythos Previewreported56.80%
- 2GPT-6 Astrareported54.80%
- 3Claude Fable 5max53.30%
- 4Muse Spark 1.1xhigh52.20%
- 5Claude Opus 4.8max49.80%
Best open weights · same scale
- 14Kimi K3maxOW43.50%
Science
Scientific research workflows
- 1GPT-6 Astramax68.1%
- 2Claude Opus 5.5max63.3%
- 3Claude Fable 5.1max40.0%
- 4Claude Opus 5max30.0%
- 5GPT-5.6 Solmax22.4%
Best open weights · same scale
- 7DeepSeek V4.1 FlashmaxOW15.7%
Visual perception
Atomic visual perception
- 1GPT-5.6 Solmax59.7%
- 2Kimi K3maxOW58.5%
- 3Claude Fable 5max57.2%
- 4GPT-5.5xhigh55.8%
- 5Claude Opus 4.8max47.2%
Knowledge work
Professional deliverables
- 1Claude Opus 5.5max38.2%
- 2GPT-6 Astramax34.2%
- 3Claude Sonnet 5reported33.3%
- 4Claude Opus 5high32.2%
- 5GPT-6 Solxhigh32.2%
Best open weights · same scale
- 7DeepSeek V4.1 FlashreportedOW31.8%
Multimodal
Visual knowledge and reasoning
- 1Gemini 3.5 Flashreported83.6%
- 2GPT-5.6 Solmax83.0%
- 3Kimi K3maxOW81.6%
- 4Claude Fable 5max81.2%
- 5GPT-5.5xhigh81.2%
Exploit capability
Offensive exploitation
- 1Claude Mythos 5reported78.0%
- 2Claude Mythos Previewreported77.7%
- 3GPT-5.6 Solreported73.5%
- 4GPT-5.5reported72.3%
- 5GPT-5.6 Terrareported52.9%
Best open weights · same scale
- 11Kimi K2.6reportedOW18.4%
Tool use
Multi-tool workflows
- 1Muse Spark 1.1reported88.1%
- 2Claude Fable 5.1reported87.2%
- 3Claude Opus 5xhigh85.8%
- 4Claude Fable 5max84.7%
- 5Qwen3.8-2.4T-A95Bxhigh84.5%
Best open weights · same scale
- 6GLM-5.3reportedOW84.2%
Bars use each benchmark's published scale. First-party and maintainer-reported results; Benchmaxxer did not rerun these evaluations.