Stack Spend

AI API Pricing in 2026: OpenAI, Anthropic, Grok, Gemini and Bedrock per Token

GuidesFebruary 25, 2026Updated September 30, 2026By 9 min read

Use this when you are evaluating AI API providers and need current pricing in a consistent format — or when you need to translate pricing into a rough monthly cost estimate for a specific workload.

The fast answer: most AI APIs charge per million tokens, split into input (the prompt you send) and output (the response you receive). Output tokens are priced higher than input tokens. The cheapest option for your workload depends on your prompt shape, not just the headline rate. A mid-tier model with short prompts can cost less than a cheap model with long prompts. This guide compares observed catalogue rates and a section on how to calculate what pricing actually means for your specific usage.

You're evaluating four AI APIs. You open four pricing pages. One charges by model family. One jumps to higher rates past 200K context. One bundles usage into seats and credits. None of them use the same units.

This guide does that work for you. Selected API models in a consistent format, with rates read from the effective-dated price database. Between guide updates, the LLM API Pricing Index tracks the same per-model list prices daily. If you want to compare what these choices mean after deployment, see AI cost monitoring.

How this page stays current

The per-token rates on this page are read from an effective-dated price database when the page is generated, with a six-hour revalidation interval, so they do not depend on
anyone remembering to edit this article. The commentary around them — which tier suits
which workload, which cost controls are worth using — is reviewed periodically and is
where any staleness will show up first.

How to Read AI API Pricing

Most AI APIs charge per token. A token is roughly 0.75 words in English — a 1,000-word document is approximately 1,300 tokens. Pricing is quoted per 1 million tokens, split into input (the prompt you send) and output (the response the model generates).

Why input and output are priced differently: Output tokens cost more to generate than input tokens are to process. For most chat-style tasks, output tokens represent 20–40% of total token volume but a higher share of cost. For long-context retrieval tasks (RAG), input tokens dominate. Because the mix shifts as your product changes, token cost monitoring — tracking input and output separately against current prices — is what keeps these numbers honest month to month.

Why context window size matters for cost: Some providers charge more for larger context windows — even at the same model tier. Sending a 200,000-token document to a model costs significantly more than sending a 2,000-token prompt, both in tokens consumed and sometimes in per-token rate. If you also need one view across infrastructure and model spend, see cloud + AI cost monitoring.

If you already know your workload shape, skip next to the more specific follow-on guides: cheapest AI API for chat, RAG, and coding, OpenAI vs Anthropic pricing, or long-context pricing above 200K tokens.

How to use this guide to make a decision

The pricing tables below tell you the list rate. To turn that into a cost estimate for your workload, you need three numbers from your own product:

  1. Average input tokens per request — how long is your prompt including system prompt, any retrieved context, and user input?
  2. Average output tokens per request — how long is the model's typical response?
  3. Requests per month — how many calls will this workflow make at your expected usage level?

Then apply:

Monthly cost = [(input_tokens/1,000,000 × input_rate)
              + (output_tokens/1,000,000 × output_rate)]
             × requests_per_month

A worked example, using fixed illustrative rates (the database-backed table below carries observed catalogue rates): A team is building a content moderation classifier. Each request has a 400-token input (content + system prompt) and a 50-token output (a classification label + brief reasoning). They expect 2 million requests per month.

Comparing two candidates:

Model Input ($/1M) Output ($/1M) Monthly cost
GPT-5 Mini $0.25 $2.00 (0.4 × $0.25 + 0.05 × $2.00) × 2,000 = $400
Claude Haiku 4.5 $1.00 $5.00 (0.4 × $1.00 + 0.05 × $5.00) × 2,000 = $1,300

At this workload shape, GPT-5 Mini is 3.25x cheaper than Claude Haiku for the same task. If quality is comparable, that difference compounds quickly at scale.

The important insight is that "cheapest model" and "cheapest for this workload" are not the same question. A model with slightly higher input rates but much lower output rates can be cheaper for long-response workflows. Check both sides of the rate before deciding. If you'd rather not do the arithmetic by hand, the LLM cost calculator runs the same estimate against current rates.


Current per-token rates

This selected table uses the database’s current effective rows when the page is generated,
with six-hour revalidation. Its timestamp is the newest observation across the catalogue,
not a guarantee that every row was verified at that instant. Missing data produces an
explicit fallback. StackSpend authors this comparison; it is not an official provider
price list. Verify the applicable rate and terms in the provider’s linked documentation.

Observed catalogue prices in USD per million tokens for selected API models
ModelProviderInput / 1MOutput / 1MCached input
o1-proOpenAI$150.00$600.00—
gpt-5.4OpenAI$2.50$15.00$0.25
gpt-5-nanoOpenAI$0.05$0.4$0.005
claude-fable-5Anthropic$10.00$50.00$1.00
claude-opus-4-8Anthropic$5.00$25.00$0.5
claude-haiku-4-5Anthropic$1.00$5.00$0.1
gemini-3.5-live-translate-previewGoogle (Gemini)$3.50$21.00—
gemini-3.8-flashGoogle (Gemini)$0.75$3.75$0.075
gemini-2.5-flash-liteGoogle (Gemini)$0.1$0.4$0.01
grok-betaxAI$5.00$15.00—
grok-4.20-beta-non-reasoningxAI$1.25$2.50$0.2
grok-build-0.1xAI$1.00$2.00$0.2
Prices read from the StackSpend catalogue; newest database observation 7 October 2026. Observed catalogue list rates are not a quote or your final bill. Context tiers, batch discounts, tools, and negotiated terms may differ. The full price index carries every tracked model and monthly price changes records what moved.

OpenAI API Pricing

Compare standard input, cached input, and output rates for the models in the database-backed table. A premium model is worthwhile only when its quality on your evaluated workload justifies the cost. Do not choose by the model name alone.

Check OpenAI’s official pricing for context tiers, batch processing, tools, and other charges that a basic token comparison excludes. Use OpenAI cost monitoring to compare observed usage with billed costs after deployment.

Anthropic API Pricing

Claude models have different price and capability trade-offs. Compare the input/output mix your workload actually uses, including cache reads and writes, rather than assuming the lowest input rate produces the lowest bill.

The table shows selected observed catalogue rows. Verify Anthropic’s official pricing for batch, caching, long-context, regional, and tool charges. See OpenAI vs Anthropic pricing for workload comparison considerations.

Google Gemini API Pricing

Gemini pricing can vary by model, context size, caching, and service tier. A short-prompt estimate may not describe a long-document workload. Check the official Gemini pricing against your expected context length before choosing a rate.

For a workload estimate, compare the table’s standard token columns first, then add the charges relevant to your use case. See AI cost monitoring for keeping the deployed workload visible across providers.

AWS Bedrock and Amazon Nova API Pricing

AWS Bedrock is different from single-vendor model APIs: it is a model platform with multiple providers and service tiers.

For Amazon's first-party family, Amazon Nova is the key lineup to watch:

  • Nova Micro (text-focused, low-latency/cost tier)
  • Nova Lite (low-cost multimodal tier)
  • Nova Pro (higher-capability multimodal tier)
  • Nova Premier (most capable tier for complex reasoning/distillation workflows)

How pricing works on Bedrock/Nova:

  • On-demand token pricing (input and output) via Bedrock pricing matrix
  • Service tiers: Standard, Priority, and Flex
  • Batch inference support for selected models
  • Region and access path affect final effective rates

Because Bedrock pricing is published in a dynamic matrix and can vary by region/tier, treat Nova pricing as a configuration exercise (model + region + tier), not a single global number.

If you are choosing between managed platforms rather than direct vendor APIs, Bedrock vs Vertex AI pricing: what teams actually pay is the better next read.


xAI Grok API Pricing

Compare the observed xAI token rates in the table, then check xAI’s official model pricing for the model, context band, and tools you intend to use. Search and other tools may have charges beyond a basic input/output calculation. A low output rate alone does not establish the cheapest workload.

Mistral API Pricing

Mistral models can be purchased through different deployment paths. Verify the exact model and hosting route against Mistral’s official pricing, including any context and processing tiers. This guide’s featured table covers OpenAI, Anthropic, Gemini, and xAI; it does not quote Mistral rates.

Cursor Pricing

Cursor combines subscription plans and usage rules rather than offering a directly interchangeable per-token API tariff. Check Cursor’s official pricing and your team’s billing settings for included usage and overage behavior. Avoid comparing a seat fee with an API token rate as though they were the same unit. See Cursor cost monitoring for connected usage and spending trends.

Hugging Face API and Endpoint Pricing

Hugging Face offers routed Inference Providers and dedicated Inference Endpoints. The billing unit and included credits depend on the product and plan. Check Inference Providers pricing and Inference Endpoints pricing, then include the provider or compute charges relevant to your deployment. See Hugging Face cost monitoring for connected cost trends.

Comparing models side by side

The table at the top of this page is generated from the same effective-dated price
database that backs the full price index, which carries
every tracked model rather than a hand-picked few, and
monthly price changes records what moved and when. This
section used to hold a second, manually maintained copy of those numbers; it was the
part of the guide that went out of date first.

What This Costs in Practice (Updated)

A typical product team using AI across a few different tasks:

  • GPT-5 Mini for app chat — 500,000 input tokens/day, 150,000 output tokens/day -> ~$0.13/day input + ~$0.30/day output -> ~$13/month
  • Gemini 3.5 Flash for high-volume classification — 5M input/day, 500K output/day -> ~$7.50/day input + ~$4.50/day output -> ~$360/month
  • Cursor Teams for 10 developers -> $400/month
  • Hugging Face Endpoint (small CPU) always-on baseline -> ~$24/month at ~$0.033/hour

Combined estimate: ~$800/month before overages, tool calls, or long-context premiums. The important operational point is not the exact number, but how quickly the combined bill becomes hard to track once several providers are live. That is where AI cost monitoring or a broader cloud + AI cost monitoring layer becomes useful.

If you are budgeting for a small company rather than a single product workflow, How Much AI API Spend Should a Startup Expect Per Month? gives a better planning lens.

Readers who land here usually branch into one of these next:


FAQ

What is the current xAI Grok API pricing?

The table shows selected observed xAI catalogue rates. Model selection spans the available price range rather than promising a particular model. Check the official xAI price list for your model, context tier, and tools.

What is the current OpenAI API pricing?

The table shows selected observed OpenAI input, output, and cached-input catalogue rates. Check OpenAI’s official pricing for your exact model and processing mode, and include applicable tool charges.

What is the current Anthropic API pricing?

The table shows selected observed Anthropic catalogue rates. Compare the actual input, output, cache, context, and processing mix of your workload, and verify it against Anthropic’s official pricing.

Where is the official OpenAI API pricing page?

Use the official OpenAI pricing pages directly:

Where is the official xAI Grok API pricing page?

Use the official xAI models/pricing documentation:

Tracking AI API Spend Across Providers

The challenge with multi-provider AI usage isn't any single bill — it's the aggregate. When OpenAI, Anthropic, Cursor, and Hugging Face all invoice separately, on different cycles, with different units, the combined picture doesn't exist unless you build it.

Connect all your AI providers to StackSpend for a single view of total AI API spend, daily anomaly detection (with webhooks to push alerts to your systems), and pace-to-forecast alerts. Setup guides: OpenAI, Anthropic, Cursor, Hugging Face, GCP (Gemini via Vertex).


References

Know what you actually pay — track your AI API spend daily.

Pricing pages tell you the rate card. StackSpend tells you the bill: connect OpenAI, Anthropic, and Grok and get a daily spend signal, model-level breakdown, and anomaly alerts before the invoice.

14-day free trial. No credit card required. Plans from $79/month.