OverviewBrowse 284 families

9 categories · 284 families · 5696 reported values

Who leads, by category

Top five on one exact representative benchmark per category. Lists are independent — scores are never combined across benchmarks.

Coding

Repository code-change mergeability

FrontierCode 1.1 · Main
  1. 1Claude Opus 5.5medium54.64%
  2. 2Claude Fable 5xhigh53.48%
  3. 3Claude Opus 5medium53.38%
  4. 4GPT-6 Astramax53.26%
  5. 5Claude Fable 5.1medium50.91%
Best open weights · same scale
  1. 12Kimi K3noneOW44.17%
Higher is betterVerified Sep 23

Agents

Terminal agent work

Terminal-Bench 4.0
  1. 1Claude Opus 5.5xhigh66.40%
  2. 2Mythos 5.1reported60.90%
  3. 3GPT-6 Astramax58.18%
  4. 4Claude Fable 5.1max57.88%
  5. 5Claude Opus 5xhigh53.94%
Best open weights · same scale
  1. 7GLM-5.3maxOW41.82%
Higher is betterVerified Sep 23

Expert reasoning

Broad expert questions

Humanity's Last Exam · full
  1. 1Claude Mythos Previewreported56.80%
  2. 2GPT-6 Astrareported54.80%
  3. 3Claude Fable 5max53.30%
  4. 4Muse Spark 1.1xhigh52.20%
  5. 5Claude Opus 4.8max49.80%
Best open weights · same scale
  1. 14Kimi K3maxOW43.50%
Higher is betterVerified Sep 23

Science

Scientific research workflows

TB-Science 0.1 · official
  1. 1GPT-6 Astramax68.1%
  2. 2Claude Opus 5.5max63.3%
  3. 3Claude Fable 5.1max40.0%
  4. 4Claude Opus 5max30.0%
  5. 5GPT-5.6 Solmax22.4%
Best open weights · same scale
  1. 7DeepSeek V4.1 FlashmaxOW15.7%
Higher is betterVerified Sep 23

Visual perception

Atomic visual perception

PerceptionBench
  1. 1GPT-5.6 Solmax59.7%
  2. 2Kimi K3maxOW58.5%
  3. 3Claude Fable 5max57.2%
  4. 4GPT-5.5xhigh55.8%
  5. 5Claude Opus 4.8max47.2%
Higher is betterVerified Sep 23

Knowledge work

Professional deliverables

ALE · pass rate
  1. 1Claude Opus 5.5max38.2%
  2. 2GPT-6 Astramax34.2%
  3. 3Claude Sonnet 5reported33.3%
  4. 4Claude Opus 5high32.2%
  5. 5GPT-6 Solxhigh32.2%
Best open weights · same scale
  1. 7DeepSeek V4.1 FlashreportedOW31.8%
Higher is betterVerified Sep 23

Multimodal

Visual knowledge and reasoning

MMMU-Pro · selected setups
  1. 1Gemini 3.5 Flashreported83.6%
  2. 2GPT-5.6 Solmax83.0%
  3. 3Kimi K3maxOW81.6%
  4. 4Claude Fable 5max81.2%
  5. 5GPT-5.5xhigh81.2%
Higher is betterVerified Sep 23

Exploit capability

Offensive exploitation

ExploitBench · coverage
  1. 1Claude Mythos 5reported78.0%
  2. 2Claude Mythos Previewreported77.7%
  3. 3GPT-5.6 Solreported73.5%
  4. 4GPT-5.5reported72.3%
  5. 5GPT-5.6 Terrareported52.9%
Best open weights · same scale
  1. 11Kimi K2.6reportedOW18.4%
Higher is betterVerified Sep 23

Tool use

Multi-tool workflows

MCP Atlas
  1. 1Muse Spark 1.1reported88.1%
  2. 2Claude Fable 5.1reported87.2%
  3. 3Claude Opus 5xhigh85.8%
  4. 4Claude Fable 5max84.7%
  5. 5Qwen3.8-2.4T-A95Bxhigh84.5%
Best open weights · same scale
  1. 6GLM-5.3reportedOW84.2%
Higher is betterVerified Sep 23

Bars use each benchmark's published scale. First-party and maintainer-reported results; Benchmaxxer did not rerun these evaluations.