AI API Cost Index 2026
Token-level pricing across 18 LLM API providers, covering 55 models. Costs are listed per 1 million tokens and updated monthly from verified vendor pricing pages.
Key findings, August 10, 2026
- The index covers 18 API providers and 55 model tiers, priced per 1M tokens.
- Amazon Nova is the cheapest at $0.035 per 1M input and $0.08 per 1M output, far below the $2.50 to $3.00 frontier input rate.
- The median input price is $1 per 1M, median output $4.4 per 1M, about 4x more.
- Output always costs more than input because generation is more compute-intensive than reading the prompt.
- 17 of 18 API providers offer a rate-limited free tier; consumer chat plans separately cluster near $20 per month.
Key Figures
Cheapest Input
$0.035/1M
Llama (Meta)
Median Input
$1/1M
across all models
Cheapest Output
$0.08/1M
Llama (Meta)
Providers Tracked
18
55 model tiers
Token Pricing by Provider
Prices in USD per 1 million tokens. Ranges shown when a provider offers multiple model tiers at different price points.
| Provider | Input / 1M | Output / 1M | Free Tier |
|---|---|---|---|
| Amazon NovaCheapest | $0.035 - $2.5 | $0.14 - $12.5 | No |
| Cohere | $0.0375 - $2.5 | $0.15 - $10 | Yes |
| Command R+ | $0.0375 - $2.5 | $0.15 - $10 | Yes |
| Groq | $0.05 - $0.6 | $0.08 - $3 | Yes |
| Replicate | $0.1 | $0.5 | Yes |
| Llama (Meta) | $0.11 - $0.5 | $0.34 - $0.77 | Yes |
| DeepSeek | $0.14 - $0.435 | $0.28 - $0.87 | Yes |
| Phi-3 | $0.14 | $0.56 | Yes |
| Kimi | $0.2 - $3 | $2 - $15 | Yes |
| ChatGPT | $0.25 - $5 | $2 - $30 | Yes |
| OpenAI API | $0.25 - $5 | $2 - $30 | Yes |
| Google AI Studio | $0.3 - $1.25 | $2.5 - $10 | Yes |
| Mistral AI | $0.5 | $1.5 | Yes |
| Claude | $1 - $10 | $5 - $50 | Yes |
| Anthropic API (Claude) | $1 - $10 | $5 - $50 | Yes |
| Google Gemini | $1.25 - $2 | $9 - $12 | Yes |
| Mistral Large | $2 | $6 | Yes |
| Grok 2 | $2 | $10 | Yes |
Consumer LLM Plans
These products are sold as monthly subscriptions to end users, not as API access. Token pricing does not apply.
| Product | Starting Price | Free Plan |
|---|---|---|
| Qwen 2.5 | Free | Yes |
| Meta AI | $7.99/mo | Yes |
| Hugging Face | $9/mo | Yes |
Key Takeaways
Open-weight models are 50x cheaper than frontier APIs
Llama 3 and Phi-3 start at $0.05-$0.14 per 1M input tokens. GPT-4o and Claude Sonnet start at $2.50-$3.00. For non-critical workloads, the cost difference is difficult to justify.
Output tokens cost 4-10x more than input tokens
Across all providers, output pricing consistently runs higher than input. Generation is computationally expensive. Prompt caching and shorter outputs have a significant impact on total inference cost.
Subscription plans have converged at $20/month
ChatGPT Plus, Claude Pro, Gemini Advanced, and Grok all price their primary consumer tier at $19.99-$20/month. Premium tiers (o1 Pro, Claude Max) run $100-$200/month for heavy power users.
Free tier availability is high across API providers
Most API providers offer a free tier with rate-limited access. This covers development, evaluation, and low-volume production workloads without any billing commitment.
Questions buyers ask
Which LLM API is cheapest per token in 2026?
Amazon Nova is the cheapest tracked provider, starting at $0.035 per 1M input tokens and $0.08 per 1M output. Open-weight models sit well below frontier APIs, where input alone runs $2.50 to $3.00 per 1M. The median input price across all 55 tracked models is $1.
Why do output tokens cost more than input tokens?
Generation is more compute-intensive than reading a prompt, so output is priced higher across every provider. The median output rate of $4.4 per 1M runs about 4x the median input rate. Prompt caching and shorter completions are the two levers that move total inference cost the most.
How many LLM API providers offer a free tier?
17 of the 18 tracked API providers offer a free tier, usually rate-limited. That covers development, evaluation and low-volume production without a billing commitment. Consumer chat plans are separate and cluster at $19.99 to $20 per month.
When is a frontier model worth the price over an open-weight one?
When accuracy on the task justifies the premium. Frontier input at $2.50 to $3.00 per 1M is roughly 86x the cheapest open-weight rate of $0.035. For routing, extraction and high-volume classification, the open-weight model usually wins; for hard reasoning, the frontier model earns its cost.
Compare LLM providers side by side
Feature benchmarks, context windows, and full pricing breakdowns for every model.
Token pricing sourced from vendor API documentation and pricing pages, verified monthly. Full methodology.