All comparisons · Scores checked 2026-10-04 · Prices 2026-10-01
AI model capability vs price
A more expensive model is not always a smarter one. This chart puts each model's general capability, measured by the Epoch Capabilities Index (ECI), against what its API costs. The highest score is Claude Opus 5.5 at 167.3. Claude Sonnet 5.5 comes within 2.2 points for 50% of the price.
The highest score on the Epoch Capabilities Index among the models priced here.
2.2 points below the top model for 50% of its price, and on the best-value frontier.
The cheapest model within 15 points of the top score: 5% of the top model's price.
These models are new, and Epoch AI has not published a capability score for them yet. They join the charts automatically once it does. Their prices today:
- GPT-6.1 Sol$2 / $10 · $4.00AA 51.8 · Astra 52.7
- GPT-6 Sol$2 / $10 · $4.00
- GPT-6 Luna$0.1 / $0.5 · $0.20
- Gemini 4 Argon$4 / $20 · $7.60
- Grok 4.7$2 / $6 · $3.00
- Mistral Large 3$0.5 / $1.5 · $0.79
Scores on a different scale cannot go on this chart, but one independent result is out: on the Artificial Analysis Intelligence Index (v4.3.2, max effort, 2026-09-30), GPT-6.1 Sol scores 51.8 against GPT-6 Astra's 52.7, near-Astra capability at a fifth of the price.
The best-value models
A model is on the best-value frontier when no other model is both cheaper and more capable. From cheapest to most capable: GPT-5 nano, DeepSeek V4 Flash, Gemini 3.8 Flash, Claude Sonnet 5.5, Claude Opus 5.5. Anything below the line costs more than a frontier model with the same or higher score.
All scores and prices
| Model | ECI (90% range) | Input / output per 1M | Blended per 1M |
|---|---|---|---|
| Claude Opus 5.5 ★ best value | 167.3 164–172 | $4 / $20 | $10.40 |
| GPT-6 Astra | 166.5 163–171 | $10 / $50 | $20.00 |
| Claude Sonnet 5.5 ★ best value | 165.2 162–169 | $2 / $10 | $5.20 |
| Claude Fable 5.1 | 164.8 162–169 | $10 / $50 | $26.00 |
| GPT-5.5 | 159.2 157–162 | $5 / $30 | $11.25 |
| Kimi K3 | 157.6 155–160 | $3 / $15 | $6.00 |
| Gemini 3.8 Flash ★ best value | 156.9 155–160 | $0.75 / $3.75 | $1.42 |
| DeepSeek V4 Pro | 155.4 154–158 | $1.32 / $3.96 | $1.98 |
| Qwen 3.8 Max | 155.2 153–157 | $2 / $6 | $3.00 |
| Gemini 3.1 Pro | 154.8 153–157 | $2 / $12 | $4.27 |
| DeepSeek V4 Flash ★ best value | 154.5 152–157 | $0.3 / $1.2 | $0.52 |
| Grok 4.20 | 152.0 149–154 | $1.25 / $2.5 | $1.56 |
| Kimi K2.6 | 151.1 149–153 | $0.95 / $4 | $1.71 |
| GPT-5 mini | 145.5 144–147 | $0.25 / $2 | $0.69 |
| Gemini 3.5 Flash-Lite | 145.1 142–147 | $0.3 / $2.5 | $0.81 |
| Gemini 3.1 Flash-Lite | 144.4 142–146 | $0.25 / $1.5 | $0.53 |
| Claude Haiku 4.5 | 142.4 140–144 | $1 / $5 | $2.10 |
| Mistral Medium 3.5 | 141.4 138–144 | $1.5 / $7.5 | $3.15 |
| GPT-5 nano ★ best value | 139.4 135–142 | $0.05 / $0.4 | $0.14 |
| GPT-4.1 | 136.8 134–138 | $2 / $8 | $3.50 |
| GPT-4.1 mini | 135.0 131–137 | $0.4 / $1.6 | $0.70 |
| GPT-4o | 128.8 124–131 | $2.5 / $10 | $4.38 |
| GPT-4o mini | 126.6 120–129 | $0.15 / $0.6 | $0.26 |
Blended price = (3 × input + output) ÷ 4 per million tokens, a typical chat mix, adjusted for each model's tokenizer (how many tokens it uses for the same English text). Standard list prices, no caching or batch discounts.
What the score means
ECI combines dozens of benchmarks (math, coding, science, reasoning and more) into one scale, so models tested on different benchmarks can still be compared. A gap of a few points is often within the margin of error; the 90% range in the table shows how sure the estimate is. It measures general capability, not your task: for a specific job, test a few frontier models on your own prompts.
To see what your own prompts cost on each model, paste them into the token counter, or compare two models side by side on the comparison pages.
What the chart means, model by model: Best value LLM in October 2026.
Capability scores: Epoch Capabilities Index by Epoch AI, used under CC BY 4.0. Prices: TokenSave, updated daily.