AI Leaderboard by Design for Online®, providing businesses with the leading AI recommendations
This week Claude Fable 5 tops the board. DeepSeek V4 Flash scores within 16 points at 1% of the price.
Our leaderboard is visited and used across 155+ countries worldwide, with nearly 2.4 million Google searches in 2026 so far. Every major AI model ranked on intelligence, speed and real cost, updated daily.
Score against price
The top of the board, plotted by DFO score and input price. Up is smarter, left is cheaper; the shaded band marks our best value picks.
| # | Model | DFO Score ▾ | Tok/s | In $/1M | Out $/1M | Ctx | |
|---|---|---|---|---|---|---|---|
| 1 | TOP PICK | 90.6 | 56.5 | $10.00 | $50.00 | 1M | |
| 2 | FRONTIER | 89.3 | 52.4 | $5.00 | $25.00 | 1M | |
| 3 | TOP PICKIN-HOUSE PICKNEW | 88.6 | 54.6 | $5.00 | $25.00 | 1M | |
| 4 | FRONTIER | 86.6 | 126 | $2.00 | $12.00 | 1M | |
| 5 | TOP PICKNEW | 84.6 | 58.5 | $3.00 | $15.00 | 1M | |
| 6 | NEW | 83.8 | 73.8 | $2.00 | $6.00 | 500K | |
| 7 | IN-HOUSE PICKNEW | 83.8 | 75.6 | $2.00 | $10.00 | 1M | |
| 8 | NEW | 83.4 | 107 | $1.25 | $4.25 | 1M | |
| 9 | BEST FOR CODING | 83.3 | 73.7 | $5.00 | $30.00 | 1.1M | |
| 10 | 82.2 | 167 | $0.6916 | $2.17 | 1M | ||
| 11 | 80.5 | 95.6 | $1.75 | $14.00 | 400K | ||
| 12 | BEST FOR AGENTS | 80.0 | 64.8 | $5.00 | $30.00 | – | |
| 13 | 79.1 | 193 | $1.50 | $9.00 | 1M | ||
| 14 | 78.7 | 45.3 | $5.00 | $25.00 | 1M | ||
| 15 | 78.5 | 211 | $0.5000 | $3.00 | 1M | ||
| 16 | 78.5 | 201 | $1.48 | $4.43 | 1M | ||
| 17 | NEW | 78.1 | 124 | $1.25 | $7.50 | 1.1M | |
| 18 | 77.9 | 60.1 | $0.4350 | $0.8700 | 1M | ||
| 19 | 77.5 | 64.4 | $5.00 | $30.00 | – | ||
| 20 | NEW | 77.5 | 64.9 | $5.00 | $30.00 | 1.1M |
We deploy these models for businesses every week. Get a recommendation for your workload.
Get StartedThis independent leaderboard tracks 650 large language models and ranks them by capability, price and speed. Design for Online aggregates third-party benchmarks with our own editorial testing, refreshing the data daily so the ranking reflects the models you can actually use today, not last year's headlines.
Leaderboards by use case
The overall table, re-ranked for the job you're hiring a model for.
Leaderboard changelog
A running log of what has changed on the board and when.
Nothing logged under this filter yet.
How we rank AI models
The Design for Online AI Model Leaderboard scores 650 models on a single 0–100 scale built from four weighted dimensions: intelligence (reasoning and knowledge benchmarks), technical capability (coding and tool use), content quality (writing and instruction-following) and value (capability per dollar).
Underlying data is aggregated from the OpenRouter API for pricing and availability, Artificial Analysis for intelligence, coding and agentic indices, and the Hugging Face Open LLM Leaderboard for open-model benchmarks. The fourth source is our own: we deploy these models in client agents, chatbots and automations every week, and that internal testing feeds the editorial layer, so a model that benchmarks well but is impractical to deploy will not automatically top the table.
Models are grouped into tiers (Frontier, Professional, Specialist, Efficient, Emerging and Legacy) to make like-for-like comparison easier, and newly released models are flagged so you can see what has just landed.
Leaderboard FAQ
How often is the leaderboard updated?
Pricing, availability and benchmark data are synced daily from our sources, and editorial scores are reviewed whenever a significant new model is released.
How is the overall score calculated?
Each model is graded 0–10 on intelligence, technical capability, content quality and value; those dimensions are weighted and combined into the 0–100 overall score used to rank the table.
Where does the data come from?
From four sources: the OpenRouter API, Artificial Analysis, the Hugging Face Open LLM Leaderboard, and internal testing from real deployments by the Design for Online team.