September 2026 · updated 6 minutes ago

AI Leaderboard by Design for Online®, providing businesses with the leading AI recommendations

This week GPT-6 Astra tops the board. Ling-3.0-flash scores within 24 points at 1% of the price.

Our leaderboard is visited and used across 155+ countries worldwide, with nearly 2.4 million Google searches in 2026 so far. Every major AI model ranked on intelligence, speed and real cost, updated daily.

Value map

Score against price

The top of the board, plotted by DFO score and input price. Up is smarter, left is cheaper; the shaded band marks our best value picks.

Editorial pick Best value Tracked model
75808489$0.1$0.3$1$3$10INPUT PRICE PER 1M TOKENS (LOG)DFO SCOREGPT-6 AstraMuse Spark 1.3Claude Fable 5.187.0 · $10.00 / 1M
Top 10 · Overall DFO score Full table with speed, pricing & context follows below.
# Model DFO Score Tok/s In $/1M Out $/1M Ctx
1GPT-6 AstraOpenAINEW88.056.0$10.00$50.001.1M
2Claude Fable 5.1AnthropicTOP PICKNEW87.065.8$10.00$50.001M
3Claude Opus 5AnthropicTOP PICKIN-HOUSE PICK86.854.3$5.00$25.001M
4Muse Spark 1.3MetaNEW83.7206$1.25$4.251M
5Qwen3.8 2.4T A95BQwenTOP PICK83.623.9$2.00$6.001M
6Claude Fable 5Anthropic83.066.8$10.00$50.001M
7GPT-5.6 SolOpenAI83.057.4$2.00$10.001.1M
8Claude Sonnet 5AnthropicIN-HOUSE PICK82.669.2$2.00$10.001M
9Claude Opus 4.8Anthropic82.057.3$5.00$25.001M
10Claude Opus 4.7Anthropic81.948.4$5.00$25.001M
11GPT-5.5OpenAI81.168.6$5.00$30.001.1M
12GLM 5.3 FlashZ.aiNEW81.1103$0.1500$0.50001.3M
13Claude Opus 5 (batch)AnthropicBEST FOR CODING81.050.4$2.50$12.501M
14Grok 4.6SpaceXAIIN-HOUSE PICK80.653.6$2.00$6.00500K
15GLM 5.3Z.aiNEW80.266.1$1.40$4.401.3M
16Kimi K3MoonshotAI79.538.1$2.65$13.281M
17GPT-5.4OpenAI79.4123$2.50$15.001.1M
18Qwen3.8 MaxQwen78.837.8$2.00$6.001M
19GPT-5.6 TerraOpenAI78.697.3$2.00$12.001.1M
20Gemini 3.8 FlashGoogleNEW78.4261$0.7500$3.751M
Showing 1–20 of 772 · Data from OpenRouter, Artificial Analysis, Hugging Face & our own testing. Scores editorially curated.

We deploy these models for businesses every week. Get a recommendation for your workload.

Get Started

This independent leaderboard tracks 772 large language models and ranks them by capability, price and speed. Design for Online aggregates third-party benchmarks with our own editorial testing, refreshing the data daily so the ranking reflects the models you can actually use today, not last year's headlines.

Leaderboards by use case

The overall table, re-ranked for the job you're hiring a model for.

Leaderboard changelog

A running log of what has changed on the board and when.

30 changes in the last 7 days
PRICE DROP DeepSeek V4 Flash 0731: output price drops 56% to $0.08 per 1M tokens.
PRICE DROP DeepSeek V4 Flash 0731: input price drops 38% to $0.04 per 1M tokens.
PRICE DROP DeepSeek V4 Flash Latest: output price drops 56% to $0.07 per 1M tokens.
PRICE DROP DeepSeek V4 Flash Latest: input price drops 40% to $0.03 per 1M tokens.
PRICE DROP Qwen3 235B A22B Instruct 2507: output price drops 60% to $0.35 per 1M tokens.
PRICE DROP Qwen3 235B A22B Instruct 2507: input price drops 60% to $0.09 per 1M tokens.
PRICE DROP Qwen3.8 27B: output price drops 15% to $2.55 per 1M tokens.
PRICE DROP Qwen3.8 27B: input price drops 49% to $0.21 per 1M tokens.
PRICE DROP DeepSeek V4 Flash: output price drops 23% to $0.13 per 1M tokens.
PRICE DROP DeepSeek V4 Flash: input price drops 23% to $0.07 per 1M tokens.
PRICE DROP MoonshotAI Kimi Latest: output price drops 12% to $10.53 per 1M tokens.
PRICE DROP MoonshotAI Kimi Latest: input price drops 12% to $2.10 per 1M tokens.
PRICE DROP DeepSeek V3 0324: output price drops 12% to $1.00 per 1M tokens.
PRICE DROP DeepSeek V3 0324: input price drops 14% to $0.25 per 1M tokens.
PRICE DROP Kimi K3: output price drops 30% to $10.53 per 1M tokens.
PRICE DROP Kimi K3: input price drops 30% to $2.10 per 1M tokens.
PRICE DROP Codestral 2508 (batch): output price drops 50% to $0.45 per 1M tokens.
PRICE DROP Codestral 2508 (batch): input price drops 50% to $0.15 per 1M tokens.
PRICE DROP Mistral Medium 3.1 (batch): output price drops 50% to $1.00 per 1M tokens.
PRICE DROP Mistral Medium 3.1 (batch): input price drops 50% to $0.20 per 1M tokens.
PRICE DROP Mistral Large 3 2512 (batch): output price drops 50% to $0.75 per 1M tokens.
PRICE DROP Mistral Large 3 2512 (batch): input price drops 50% to $0.25 per 1M tokens.
PRICE DROP Ministral 3 8B 2512 (batch): output price drops 50% to $0.08 per 1M tokens.
PRICE DROP Ministral 3 8B 2512 (batch): input price drops 50% to $0.08 per 1M tokens.
PRICE DROP Mistral Small 4 (batch): output price drops 50% to $0.30 per 1M tokens.
PRICE DROP Mistral Small 4 (batch): input price drops 50% to $0.08 per 1M tokens.
PRICE DROP GLM Latest: output price drops 37% to $2.20 per 1M tokens.
PRICE DROP GLM Latest: input price drops 37% to $0.70 per 1M tokens.
PRICE DROP DeepSeek V3: input price drops 20% to $0.26 per 1M tokens.
PRICE DROP Gemma 4 26B A4B: output price drops 35% to $0.22 per 1M tokens.

How we rank AI models

The Design for Online AI Model Leaderboard scores 772 models on a single 0–100 scale built from four weighted dimensions: intelligence (reasoning and knowledge benchmarks), technical capability (coding and tool use), content quality (writing and instruction-following) and value (capability per dollar).

Underlying data is aggregated from the OpenRouter API for pricing and availability, Artificial Analysis for intelligence, coding and agentic indices, and the Hugging Face Open LLM Leaderboard for open-model benchmarks. The fourth source is our own: we deploy these models in client agents, chatbots and automations every week, and that internal testing feeds the editorial layer, so a model that benchmarks well but is impractical to deploy will not automatically top the table.

Models are grouped into tiers (Frontier, Professional, Specialist, Efficient, Emerging and Legacy) to make like-for-like comparison easier, and newly released models are flagged so you can see what has just landed.

Leaderboard FAQ

How often is the leaderboard updated?

Pricing, availability and benchmark data are synced daily from our sources, and editorial scores are reviewed whenever a significant new model is released.

How is the overall score calculated?

Each model is graded 0–10 on intelligence, technical capability, content quality and value; those dimensions are weighted and combined into the 0–100 overall score used to rank the table.

Where does the data come from?

From four sources: the OpenRouter API, Artificial Analysis, the Hugging Face Open LLM Leaderboard, and internal testing from real deployments by the Design for Online team.

© 2026 Design for Online Ltd. Registered in England and Wales No. 10328553. VAT Registered. Design for Online® and Forerunner® are registered trademarks.