The arena for
frontier models
Benchmark, compare and track every major LLM in one place. Intelligence, coding and agentic indexes, Design Arena ELO, prices and context windows — refreshed every day.
Current podium
Top intelligence index, right now.
Everything you need to pick a model
Built on live data with no key required — refreshed by an automated pipeline every 24 hours.
Daily snapshots
A scheduled sync fetches OpenRouter every day at 06:00 UTC and stores dated benchmark history in MongoDB.
Side-by-side compare
Stack up to 6 models on intelligence, coding, agentic, context and price — with shareable URLs.
Trend charts
Watch how intelligence indexes move over time on every model page as snapshots accumulate.
Price × intelligence
A log-scale scatter finds the best value models — quality per dollar, plotted for all 300+ models.
Design Arena ELO
Per-category design ELO ratings straight from OpenRouter’s Design Arena leaderboard.
Curated classic evals
MMLU-Pro, GPQA, HumanEval, MATH & AIME figures for 160+ models, curated from public model cards.
How it works
Sync
A scheduled Netlify function fetches OpenRouter’s model catalog and benchmark payloads every day at 06:00 UTC.
Store
Every fetch is written to MongoDB Atlas as a dated snapshot — models, indexes, ELOs and prices, all versioned.
Explore
Leaderboard, radar charts, trend lines, comparisons and price analytics — served server-side, animated client-side.
Backed by real, attributed data
Curious how it\u2019s built? The pipeline, data model and refresh schedule are documented on the About page.