Anthropic: Claude Opus 4
anthropic/claude-opus-4
Context: 200,000In: $15.00/1MOut: $75.00/1MUpdated 2026-08-05
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...
Intelligence
—
Coding
—
Agentic
—
Metric profile
Trend over time
Intelligence
—
Coding
—
Agentic
—
Only 1 snapshot so far (data refreshes daily) — the trend line appears after the 2nd snapshot.
Classic evals (curated from model cards)
Anthropic Claude (Opus) — manually maintained, may lag releases
MMLU-Pro
77.5%
GPQA Diamond
78.6%
HumanEval
95%
MATH-500
96.8%
AIME 2025
32%
Design Arena (ELO by category)
| Category | ELO | Win rate | Rank |
|---|---|---|---|
| gamedev | 1215 | 59.9% | #42 |
| 3d | 1194 | 57.7% | #47 |
| codecategories | 1190 | 55.6% | #59 |
| uicomponent | 1189 | 59.2% | #55 |
| website | 1188 | 54.6% | #61 |
| svg | 1175 | 57.7% | #41 |
| dataviz | 1172 | 57.9% | #61 |