The LLM API Pricing Index
List prices per 1M tokens across the major model providers and inference hosts, next to each model’s coding-benchmark score — so you can spot equal-or-better quality for less. Synced daily from public sources.
Prices updated 5 August 2026
657
models tracked across 12 providers, updated 5 August 2026
83.5%
top coding score — claude-opus-4-7 (Anthropic)
$0.40 /1M out
best coding value — deepseek-v3.2 at 74.2%
| gpt-5.5 | OpenAI | $5.00 | $30.00 | 1.1M | 80.6% |
| gpt-5.4 | OpenAI | $2.50 | $15.00 | 1.1M | 76.9% |
| gpt-5.4-mini | OpenAI | $0.75 | $4.50 | 272K | — |
| o3 | OpenAI | $2.00 | $8.00 | 200K | 76.9% |
| claude-opus-4-8 | Anthropic | $5.00 | $25.00 | 1M | — |
| claude-sonnet-4-5 | Anthropic | $3.00 | $15.00 | 200K | 71.3% |
| claude-haiku-4-5 | Anthropic | $1.00 | $5.00 | 200K | — |
| gemini-2.5-pro | Google (Gemini) | $1.25 | $10.00 | 1.0M | 83.1% |
| gemini-2.5-flash | Google (Gemini) | $0.30 | $2.50 | 1.0M | 55.1% |
| deepseek-v3.2 | DeepSeek | $0.28 | $0.40 | 164K | 74.2% |
| deepseek-r1 | DeepSeek | $0.55 | $2.19 | 66K | 56.9% |
| grok-4 | xAI | $3.00 | $15.00 | 256K | — |
| mistral-large-latest | Mistral | $0.50 | $1.50 | 262K | — |
| claude-opus-4-7 | Anthropic | $5.00 | $25.00 | 1M | 83.5% |
| google/gemini-2.5-pro | DeepInfra | $1.25 | $10.00 | 1M | 83.1% |
| gemini-2.5-pro-preview-tts | Google (Gemini) | $1.25 | $10.00 | 1.0M | 83.1% |
| gemini-3.5-flash | Google (Gemini) | $1.50 | $9.00 | 1.0M | 79.3% |
| claude-opus-4-6 | Anthropic | $5.00 | $25.00 | 1M | 78.7% |
| accounts/fireworks/models/glm-5p2 | Fireworks AI | $1.40 | $4.40 | 1.0M | 78.7% |
| glm-5p2 | Fireworks AI | $1.40 | $4.40 | 1.0M | 78.7% |
| deepseek-v4-pro | DeepSeek | $0.43 | $0.87 | 1M | 77.6% |
| accounts/fireworks/models/deepseek-v4-pro | Fireworks AI | $1.74 | $3.48 | 1.0M | 77.6% |
| deepseek-v4-pro | Fireworks AI | $1.74 | $3.48 | 1.0M | 77.6% |
| accounts/fireworks/models/kimi-k2p6 | Fireworks AI | $0.95 | $4.00 | 262K | 76.7% |
| kimi-k2p6 | Fireworks AI | $0.95 | $4.00 | 262K | 76.7% |
| kimi-k2.6 | Moonshot AI | $0.95 | $4.00 | 262K | 76.7% |
| accounts/fireworks/routers/kimi-k2p6-fast | Fireworks AI | $2.00 | $8.00 | 262K | 76.7% |
| kimi-k2p6-fast | Fireworks AI | $2.00 | $8.00 | 262K | 76.7% |
| claude-opus-4-5 | Anthropic | $5.00 | $25.00 | 200K | 76.7% |
| gemini-3.1-pro-preview | Google (Gemini) | $2.00 | $12.00 | 1.0M | 75.6% |
| gemini-3.1-pro-preview-customtools | Google (Gemini) | $2.00 | $12.00 | 1.0M | 75.6% |
| gemini-3-flash-preview | Google (Gemini) | $0.50 | $3.00 | 1.0M | 75.4% |
| claude-sonnet-4-6 | Anthropic | $3.00 | $15.00 | 1M | 75.2% |
| gpt-5.3-codex | OpenAI | $1.75 | $14.00 | 272K | 74.8% |
| accounts/fireworks/models/deepseek-v3p2 | Fireworks AI | $0.56 | $1.68 | 164K | 74.2% |
| accounts/fireworks/models/glm-5p1 | Fireworks AI | $1.40 | $4.40 | 203K | 74.2% |
| glm-5p1 | Fireworks AI | $1.40 | $4.40 | 203K | 74.2% |
| accounts/fireworks/routers/glm-5p1-fast | Fireworks AI | $2.80 | $8.80 | 203K | 74.2% |
| glm-5p1-fast | Fireworks AI | $2.80 | $8.80 | 203K | 74.2% |
| moonshotai/Kimi-K2.5 | Together AI | $0.50 | $2.80 | 256K | 73.8% |
| accounts/fireworks/models/kimi-k2p5 | Fireworks AI | $0.60 | $3.00 | 262K | 73.8% |
| kimi-k2p5 | Fireworks AI | $0.60 | $3.00 | 262K | 73.8% |
| kimi-k2.5 | Moonshot AI | $0.60 | $3.00 | 262K | 73.8% |
| gpt-5.2 | OpenAI | $1.75 | $14.00 | 272K | 73.8% |
| gpt-5.2-chat-latest | OpenAI | $1.75 | $14.00 | 128K | 73.8% |
| gpt-5 | OpenAI | $1.25 | $10.00 | 272K | 73.6% |
| gpt-5-chat | OpenAI | $1.25 | $10.00 | 128K | 73.6% |
| gpt-5-chat-latest | OpenAI | $1.25 | $10.00 | 128K | 73.6% |
| claude-opus-4-1 | Anthropic | $15.00 | $75.00 | 200K | 73.3% |
| gpt-5.1 | OpenAI | $1.25 | $10.00 | 272K | 68.0% |
Showing 50 of 616 models · prices in USD per 1M tokens · benchmarks by base model: coding via Aider polyglot & SWE-bench Verified; reasoning & math via Epoch AI.
Get price-change alerts
One email when a model’s price changes, a new model launches, or a model you might rely on gets a deprecation date. No schedule, no newsletter — it only sends on days something actually changed.
Unsubscribe with one click, any time.
Methodology & sources
The StackSpend LLM Pricing Index is refreshed daily. Prices are provider list prices in USD per 1 million tokens, synced from the community-maintained LiteLLM price dataset. Coding scores are the best of the Aider polyglot benchmark and SWE-bench Verified (via the Epoch AI Benchmarking Hub, CC BY) and are attached to the underlying base model, so the identical open-weight model served by different hosts shares one score. StackSpend uses this same data to track and right-size your own AI spend. For what these models are and how the families relate, see the LLM model glossary.
- How often is this LLM pricing data updated?
- Daily. Prices are synced from the community-maintained LiteLLM price dataset; coding scores are the best of the Aider polyglot benchmark and SWE-bench Verified (via Epoch AI); this page was last refreshed on 5 August 2026. All prices are list prices in USD per 1M tokens.
- Which model gives the best coding performance for the price?
- Among models benchmarked for coding, deepseek-v3.2 (DeepSeek) offers the strongest coding score relative to its output price — a 74.2% coding-benchmark score (best of Aider polyglot and SWE-bench Verified) at $0.40 per 1M output tokens.
- What is the highest-scoring model for coding?
- claude-opus-4-7 (Anthropic) currently leads the coding benchmark at 83.5%, priced at $5 input / $25 output per 1M tokens.
- Why do the same open models cost different amounts?
- Open-weight models (like Llama or DeepSeek) are served by multiple inference hosts — Groq, Together AI, Fireworks, DeepInfra and others — at different prices for the identical weights. Comparing hosts for the same model is often the fastest saving.
Use this data
The Index is free to use under CC BY 4.0 — download it or pull it live, and cite the StackSpend LLM Pricing Index with a link back to this page.
Cite as: StackSpend LLM Pricing Index — https://www.stackspend.app/resources/llm-api-pricing (updated 5 August 2026).
Track what you actually spend on these models.
StackSpend connects to your AI and cloud providers read-only, normalises every model into one daily signal, and flags when a model or deploy pushes your spend off-baseline — with cheaper equal-quality alternatives from this very dataset.