Verified
Reviewed byOleh Kem
Plans checked93 / 21 vendors
“Contact sales”26/93
Median entry$8.50/mo
Entry range$5.99-30
Free tier15/21
Leader100 Amazon Nova

Language models compared: subscriptions, capabilities and context

Subscriptions, per-token rate cards and open weights you host yourself: LLM vendors sell in three different currencies, and we scored all three from the vendors' own 2026 pages. Free tiers, usage multipliers and the surcharges that never reach a headline rate are called out where we found them.

What does an LLM cost a month, and what does a million tokens cost?

An LLM is sold by the seat or by the million tokens, and among 21 vendors the middle monthly plan starts at $8.50. The floor sits at $5.99, the highest starting plan at $30, and 15 vendors ask nothing to begin with. The rest sell no subscription. They meter by the token, or they hand you the weights.

  • Twenty-one vendors here disclose something you can price against: a monthly plan, a public per-million-token card, or weights that cost nothing to license.
  • We opened every one of the 93 plan tiers in this category, and twenty-six showed no price at all, almost all of them enterprise steps or custom quotes.
  • 15 of 21 vendors let you start without paying, but consumer free tiers are a standing message cap, not a trial that expires.
  • Where a card prices text generation, output always costs more than input, often several times more, so the length of the answer moves the bill more than the length of the prompt.
  • Two levers cut the rate almost everywhere, batch runs and cache hits, and neither switches itself on.
  • Paid consumer steps mostly sell more usage, not a better model.

Ranked by a transparency score: pricing transparency 60%, user satisfaction 40%. Capability is not scored. It is the condition grid below. Prices are read from vendor pricing pages and re-checked per product on the dates shown. ComparEdge sells no language models and takes no payment for placement. How the score is built.

01 / 08

Language models ranked: vision, web search, JSON mode and open weights

Read the grid as a price axis first, then as a capability check: vision, web search, JSON mode, fine-tuning, and open weights that run in house. The last two columns are money rather than features: whether anything is free to start on, and whether we found a price on every tier.

Sorted by transparency scorePriced tiers 67 / 93Full disclosure 6 / 21
How to read this table
01Amazon NovaUsage-priced, public rate cardNo free tierAll tiers pricedVision listedWeb search not on the recordJSON mode listedFine-tuning listedOpen weights not on the recordSales only100Alternatives to Amazon Nova
02KimiFree tierAll tiers pricedVision listedWeb search listedJSON mode not on the recordFine-tuning listedOpen weights listed$19/u100Alternatives to Kimi
03Hugging FaceFree tierAll tiers pricedVision not on the recordWeb search not on the recordJSON mode not on the recordFine-tuning listedOpen weights listed$9/u91Alternatives to Hugging Face
04CohereUsage-priced, public rate cardNo free tierAll tiers pricedVision not on the recordWeb search listedJSON mode listedFine-tuning listedOpen weights listedSales only89Alternatives to Cohere
05Google GeminiFree tierAll tiers pricedVision listedWeb search listedJSON mode not on the recordFine-tuning not on the recordOpen weights not on the record$7.99/u86Alternatives to Google Gemini
06OpenAI APIUsage-priced, public rate cardNo free tier1 of 2 tiers unpricedVision listedWeb search not on the recordJSON mode listedFine-tuning listedOpen weights not on the recordSales only86Alternatives to OpenAI API
07Meta AIFree tierAll tiers pricedVision listedWeb search not on the recordJSON mode not on the recordFine-tuning listedOpen weights not on the record$7.99/u86Alternatives to Meta AI
08Qwen 2.5Usage-priced, public rate cardFree tier1 of 2 tiers unpricedVision listedWeb search not on the recordJSON mode listedFine-tuning listedOpen weights listedSales only85Alternatives to Qwen 2.5
09GroqUsage-priced, public rate cardFree tier2 of 3 tiers unpricedVision not on the recordWeb search not on the recordJSON mode listedFine-tuning not on the recordOpen weights not on the recordSales only85Alternatives to Groq
10Anthropic API (Claude)Free tier2 of 6 tiers unpricedVision listedWeb search not on the recordJSON mode listedFine-tuning not on the recordOpen weights not on the record$20/u84Alternatives to Anthropic API (Claude)
11DeepSeekUsage-priced, public rate cardFree tier1 of 2 tiers unpricedVision listedWeb search not on the recordJSON mode listedFine-tuning not on the recordOpen weights listedSales only83Alternatives to DeepSeek
12ClaudeFree tier1 of 7 tiers unpricedVision listedWeb search listedJSON mode not on the recordFine-tuning not on the recordOpen weights not on the record$20/u82Alternatives to Claude
13ChatGPTFree tier2 of 8 tiers unpricedVision listedWeb search listedJSON mode not on the recordFine-tuning not on the recordOpen weights not on the record$7/u78Alternatives to ChatGPT
14ReplicateUsage-priced, public rate cardNo free tier1 of 2 tiers unpricedVision not on the recordWeb search not on the recordJSON mode not on the recordFine-tuning listedOpen weights listedSales only78Alternatives to Replicate
15Command R+Usage-priced, public rate cardNo free tier0 of 1 tiers pricedVision not on the recordWeb search listedJSON mode listedFine-tuning listedOpen weights not on the recordSales only78Alternatives to Command R+
16Google AI StudioUsage-priced, public rate cardFree tier2 of 3 tiers unpricedVision listedWeb search not on the recordJSON mode listedFine-tuning not on the recordOpen weights not on the recordSales only77Alternatives to Google AI Studio
17Mistral AIFree tier1 of 5 tiers unpricedVision listedWeb search listedJSON mode listedFine-tuning not on the recordOpen weights listed$5.99/u77Alternatives to Mistral AI
18Mistral LargeFree tier2 of 5 tiers unpricedVision not on the recordWeb search not on the recordJSON mode listedFine-tuning listedOpen weights listed$5.99/u76Alternatives to Mistral Large
19Grok 2Free tier1 of 5 tiers unpricedVision listedWeb search listedJSON mode listedFine-tuning not on the recordOpen weights not on the record$30/u71Alternatives to Grok 2
20Phi-3Usage-priced, public rate cardNo free tier0 of 7 tiers pricedVision listedWeb search not on the recordJSON mode listedFine-tuning listedOpen weights listedSales only71Alternatives to Phi-3
21Llama (Meta)Free to start, paid tier on requestFree tier1 of 2 tiers unpricedVision listedWeb search not on the recordJSON mode not on the recordFine-tuning listedOpen weights listedFree63Alternatives to Llama (Meta)
GrantedSome tiers sealedNot on the recordCheapest paid seatScore ranks pricing transparency and user ratings, not capability. Capability is the grid.
21 vendorsMedian entry $8.50Coverage span 41 of fiveSealed tiers 26 / 93

Scroll the console sideways to reach the remaining conditions.

02 / 08

Shortlist a model by seat count and what it really costs your team

The shortlist ranks by what the plan actually costs your team, not by list price. Flat plans are folded into a per-team number so they compare like for like. Only vendors rated 85 and up are eligible.

Seats
6
Budget / seat

For 6 seats at $9 per seat, start with these

Ranked by monthly team cost, vendors rated 85 and up · Your ceiling for this team: $54 / mo

Best value

Google Gemini

CE 86 free tier

Drafting and retrieval inside Docs and Gmail, where the model needs the file already open rather than one you paste in.

Entry plan$7.99 / seat
× 6 seats$47.94
Against your $54 ceiling$6.06
Team cost$47.94 / mo
Best value

Meta AI

CE 86 free tier

Quick answers and images inside the messaging apps already open on the phone.

Entry plan$7.99 / seat
× 6 seats$47.94
Against your $54 ceiling$6.06
Team cost$47.94 / mo
Best value

Hugging Face

CE 91 free tier

Working with open models day to day, pulling weights, fine-tuning them and publishing a demo without wiring up hosting yourself.

Entry plan$9 / seat
× 6 seats$54
Against your $54 ceiling$0
Team cost$54 / mo
Over your ceilingKimi $19 / seat
Not sure what to weigh? Five questions narrow it faster than the grid.Answer 5 questions
03 / 08

API rates per million tokens: cheapest and flagship models

Subscriptions above are the seat you sit in; these are the rates your code pays. Vendor-published API prices, US$ per million tokens. 17 of 21 vendors put a card on the record.

VendorCheapest model in / out per 1MFlagship in / out per 1MMax context
Amazon Nova6 modelsNova Micro$0.04 / $0.14Nova Premier$2.50 / $12.501M
Cohere2 modelsCommand R7B$0.04 / $0.15Command A$2.50 / $10256K
Command R+2 modelsCommand R7B$0.04 / $0.15Command A$2.50 / $10256K
Groq6 modelsLlama 3.1 8B Instant$0.05 / $0.08Qwen 3.6 27B$0.60 / $3131K
Replicate1 modelsVarious models$0.10 / $0.50
DeepSeek2 modelsDeepSeek-V4-Flash$0.14 / $0.28DeepSeek-V4-Pro$0.44 / $0.871M
Phi-31 modelsPhi-3 Medium (Azure)$0.14 / $0.56
Kimi8 modelsmoonshot-v1-8k$0.20 / $2kimi-k3$3 / $151.048M
OpenAI API5 modelsGPT-5 Mini$0.25 / $2GPT-5.5$5 / $30400K
ChatGPT4 modelsGPT-5 Mini$0.25 / $2GPT-5.5$5 / $30400K
Google AI Studio2 modelsGemini 2.5 Flash$0.30 / $2.50Gemini 2.5 Pro$1.25 / $10
Mistral AI1 modelsMistral Large 3$0.50 / $1.50128K
Anthropic API (Claude)4 modelsHaiku 4.5$1 / $5Opus 4.8$5 / $251M
Claude4 modelsHaiku 4.5$1 / $5Fable 5$10 / $501M
Google Gemini3 modelsGemini 2.5 Pro$1.25 / $10Gemini 3.1 Pro$2 / $121M
Mistral Large1 modelsMistral Large$2 / $6128K
Grok 21 modelsGrok 2$2 / $10128K
04 / 08

Every model's plans and the date we checked each published price

Open a record and the plans appear as the vendor publishes them, with a check date against every figure. A vendor with nothing monthly to sell gets an entry line naming the token card or the open-weight route instead.

Transparency scorePricing transparency 60%User satisfaction 40%

01

Amazon Nova, AWS's own family of models priced for high-volume work
Amazon Nova

100DisclosureSales onlyNo price

High-volume inference for teams whose data already sits in AWS, billed on the account that carries everything else. The cheapest models in the line are built for batch work, and the flexible mode runs at about half the standard rate.

Critical gapThe model demonstrates inconsistent competitive performance in complex instruction following benchmarks.

Plan table and expert take

Amazon Nova: expert take

Three units of account inside one product line: tokens for the language models, an image for Canvas, an hour of work for Act. Parallel agents each bill separately, so a fan-out multiplies the hourly charge instead of sharing it.

Where Amazon Nova holds up

  • Native wiring into S3, Lambda, and IAM, so Nova drops into an existing AWS stack without new plumbing
  • Nova Micro at $0.035 per 1M input is among the cheapest hosted text models anywhere
  • A full spread of modalities under one roof: text, image, video, voice via Sonic, and agentic automation via Act
  • Nova Premier and Nova 2 Lite reach a 1M-token context window for long documents in a single pass
  • Unified billing and security posture inside your existing AWS account, no separate vendor to onboard

Founded 2023Verified July 8, 2026

9 plans, as published
PlanMonthlyAnnual
Nova Micro$0.04 / 1M tokNot published
Nova Lite$0.06 / 1M tokNot published
Nova Pro$0.80 / 1M tokNot published
Nova Premier$2.50 / 1M tokNot published
Nova 2 Lite$0.30 / 1M tokNot published
Nova 2 Pro (Preview)$1.25 / 1M tokNot published
Nova Canvas$0.04 / imageNot published
Nova Sonic$3.40 / 1M tokNot published
Nova Act$4.75 / hrNot published
02

Kimi, Moonshot's K3 model that chases frontier reasoning at open-weight cost
Kimi

100Disclosure$19Seat / mo

Near-frontier reasoning where the API line item decides which vendor wins. Kimi runs five consumer steps from free to the top, with about a fifth off each paid step when you pay yearly.

Critical gapAlways-on reasoning makes it verbose, so output-token cost runs high.

Plan table and expert take

Kimi: expert take

A cache hit drops the input rate tenfold, automatically, with no storage fee and no expiry attached, which is the part worth designing for. The API sits as its own line inside the same consumer price grid, priced per million tokens.

Where Kimi holds up

  • Top-5 independent intelligence (Artificial Analysis Index 57, #4/189) at $3/$15 API rates, a fraction of the Opus tier
  • Automatic prompt caching drops input from $3.00 to $0.30 per 1M with nothing to configure, no TTL or storage fee
  • 1M-token context window as standard, no context-length pricing bands to reason about
  • Open weights under a Modified MIT license (announced for July 27, 2026) make fine-tuning and self-hosting a real option
  • Global platform runs on Singapore infrastructure under Singapore law, a cleaner data-residency story than a China-only stack

Founded 2023Verified July 17, 2026

6 plans, as published
PlanMonthlyAnnual
AdagioFreeFree
Moderato$19$15
Allegretto$39$31
Allegro$99$79
Vivace$199$159
API Pay-as-you-go$3 / 1M input tokensNot published
03

Hugging Face, the hub for hosting and fine-tuning open models
Hugging Face

91Disclosure$9Seat / mo

Working with open models day to day, pulling weights, fine-tuning them and publishing a demo without wiring up hosting yourself. Hugging Face is a hub rather than a model, so the paid steps buy room to run and store.

Critical gapThe interface demands significant technical expertise for deployment.

Plan table and expert take

Hugging Face: expert take

The paid step buys machine time, not a model: the daily GPU quota goes from five minutes to forty, with a terabyte of private storage attached. Egress and CDN carry no separate charge, which is unusual for a host storing this much.

Where Hugging Face holds up

  • Massive hub of 500K+ open-source models and datasets
  • Transformers library simplifies using state-of-the-art models
  • Integrated Spaces for building and sharing live ML demos
  • Strong community for collaboration and support
  • Inference Endpoints for easy, scalable model deployment

4.7CE scoreG2 4.6 · 5 reviewsCapterra 4.5Founded 2016Verified July 16, 2026

4 plans, as published
PlanMonthlyAnnual
FreeFreeFree
PRO$9Not published
Team$20Not published
Enterprise$50Not published
04

Cohere, the enterprise search and retrieval model maker
Cohere

89DisclosureSales onlyNo price

Enterprise search and retrieval that has to stay inside one legal jurisdiction. Cohere sells the parts around the model, embeddings, reranking and parsing, and prices each part separately instead of folding them into a plan.

Critical gapThe model produces inconsistent outputs during complex reasoning tasks.

Plan table and expert take

Cohere: expert take

Rerank arithmetic is where the bill surprises people: one search counts as a single request with up to a hundred documents, but any document over five hundred tokens is split, and each piece then counts as its own document.

Where Cohere holds up

  • Command A carries enterprise RAG and tool use, tuned for grounded answers over your own data rather than open chat
  • State-of-the-art multilingual embeddings across 100+ languages, which most rivals do not match at that spread
  • A dedicated Rerank endpoint that measurably lifts search relevance instead of leaning on the LLM to sort results
  • Built for private deployment: your own cloud, Azure private hosting, or on-prem, so data stays put
  • API-first with SDKs for Python, Go, Node, and Java, so it drops into an existing backend cleanly

4.6CE scoreG2 4.5 · 6 reviewsCapterra 4.4Founded 2019Verified July 8, 2026

5 plans, as published
PlanMonthlyAnnual
Command A$2.50 / per 1M input tokensNot published
Command R$0.15 / per 1M input tokensNot published
Command R7B$0.04 / per 1M input tokensNot published
Embed v3$0.10 / per 1M input tokensNot published
Rerank v3$2 / per 1M tokensNot published
05

Google Gemini, the assistant wired into Workspace
Google Gemini

86Disclosure$7.99Seat / mo

Drafting and retrieval inside Docs and Gmail, where the model needs the file already open rather than one you paste in. Gemini's advantage is proximity: the model sits inside the apps and files the work already lives in.

Critical gapThe platform fails to provide native meeting recording features within the free tier.

Plan table and expert take

Google Gemini: expert take

The tier gates credits rather than model access: a hundred a month on free, a thousand in the middle, twenty-five thousand at the top. Part of the same payment buys Google One storage shared with mail and photos, so some of it is gigabytes.

Where Google Gemini holds up

  • Answers land inside Gmail, Docs, and Sheets, so drafting and summarizing happen without leaving the file
  • Live Google Search grounding means current information instead of a stale training cutoff
  • Takes text, images, audio, and video as input, and the ~1M context window swallows long documents whole
  • The free tier runs a capable current model with a real monthly credit allowance, not a teaser
  • AI Plus at $7.99 is one of the cheaper paid AI tiers going, a soft on-ramp before the $19.99 Pro jump

4.5CE scoreG2 4.4 · 22 reviewsCapterra 4.2Founded 2023Verified July 16, 2026

4 plans, as published
PlanMonthlyAnnual
FreeFreeFree
AI Plus$7.99Not published
AI Pro$19.99Not published
AI Ultra$99.99Not published
06

The OpenAI API, the developer gateway to GPT models
OpenAI API

86DisclosureSales onlyNo price

Shipping a standard AI feature this week, with SDKs and a playground that behave. Nothing is sold as a subscription: the account meters from the first call, and only the enterprise line withholds a number.

Critical gapThe system produces inconsistent outputs during high-volume summarization tasks.

Plan table and expert take

OpenAI API: expert take

Pricing comes in layers rather than tiers: batch and flex run about half of standard, priority costs more, and endpoints with regional data handling add a tenth on the newer models. The same models bought through a marketplace bill on someone else's terms.

Where OpenAI API holds up

  • One API covers text, vision, audio in and out, embeddings, and native image generation, so you are not stitching four vendors together
  • GPT-5.5 for hard reasoning down to nano tiers for high-volume classification, priced across a wide range
  • The Batch API takes 50% off every model, and prompt caching cuts repeated context by another 50%
  • o3 and o4-mini handle multi-step reasoning tasks that trip up the general chat models
  • The largest developer community of any provider, so most integration problems are already solved somewhere

4.8CE scoreG2 4.7 · 11 reviewsCapterra 4.8Founded 2020Verified July 8, 2026

2 plans, as published
PlanMonthlyAnnual
Pay-as-you-go$0Not published
EnterpriseContact sales
07

Meta AI, the Llama-powered assistant inside Meta's messaging apps
Meta AI

86Disclosure$7.99Seat / mo

Quick answers and images inside the messaging apps already open on the phone. Meta only began selling a paid tier this year, and it is arriving one country at a time, so what you can buy depends on where you are.

Critical gapThe interface presents a steep learning curve for complex campaigns.

Plan table and expert take

Meta AI: expert take

The free tier is where the ads live: the plan card says so outright, and warns it may slow or skip heavy jobs when the service is busy. No public API and no token price exist on the consumer side, so developers get pointed at open weights.

Where Meta AI holds up

  • Lives inside WhatsApp, Instagram, and Messenger, so answers arrive in the thread with nothing to install
  • Real-time web search built in, pulling current results instead of answering from a training cutoff
  • The core assistant stays free, and the paid Meta One tiers are optional rather than a wall in front of basic use
  • Fast image generation right in the chat for quick visual ideas
  • Runs on Llama 4, so the underlying model is genuinely current rather than a stripped-down consumer engine

4.4CE scoreG2 4.3 · 5,666 reviewsCapterra 4.3Founded 2023Verified July 8, 2026

5 plans, as published
PlanMonthlyAnnual
FreeFreeFree
Meta One Plus$7.99Not published
Meta One Premium$19.99Not published
Meta One Essential$14.99Not published
Meta One Advanced$49.99Not published
08

Qwen 2.5, Alibaba's open-weight model line for coding and multilingual work
Qwen 2.5

85DisclosureSales onlyNo price

Multilingual and coding work on machines you already run, where the licence has to cost nothing. The coding variant is why most teams pick this family: it does more than its parameter count suggests, and the weights stay downloadable.

Critical gapThe model underperforms against proprietary alternatives during complex architectural reasoning tasks.

Plan table and expert take

Qwen 2.5: expert take

The flagship generation has quietly left the vendor's own commercial list, displaced by the next one, and the only live figure for the larger model now comes from a third-party marketplace. The open weights are still there to download.

Where Qwen 2.5 holds up

  • Permissive commercial license on every size, so you can fine-tune, quantize, and ship without asking anyone
  • Seven model sizes from 0.5B to 72B let you match the model to the hardware instead of paying for capacity you do not use
  • Qwen2.5-Coder is one of the strongest budget coding models you can run locally, and Qwen2.5-Math handles step-by-step arithmetic reliably
  • 128K context window swallows a full technical manual in one call, which simplifies long-document RAG
  • Genuine multilingual depth across 29-plus languages, with Chinese and East Asian performance most Western models cannot match

Founded 2023Verified July 8, 2026

2 plans, as published
PlanMonthlyAnnual
Open SourceFreeFree
API (Alibaba Cloud DashScope)Contact sales
09

Groq, the low-latency inference service for open models
Groq

85DisclosureSales onlyNo price

Voice agents and live chat, where a second of latency is the whole product. Groq serves open models on an OpenAI-compatible API, so adopting it is a config change and not a rebuild. Its developer tier lifts rate limits tenfold without publishing a price.

Plan table and expert take

Groq: expert take

Caching costs nothing to enable; the discount is simply a lower rate on hits, about half. Speech recognition charges a minimum of ten seconds per request, so a two-second clip pays for ten, and compound systems stack server-side tool charges on top of model rates.

Where Groq holds up

  • LPU inference runs open-weight models at hundreds of tokens per second, several times faster than GPU hosting
  • An OpenAI-compatible API means a near drop-in swap: change the base URL, key, and model name
  • Serves current open models including Llama 4 Maverick and Scout plus OpenAI's GPT-OSS
  • Time-to-first-token under 100ms makes real-time voice and chat feel immediate
  • A genuinely usable free tier plus per-token rates among the lowest for hosted inference

Founded 2016Verified July 8, 2026

3 plans, as published
PlanMonthlyAnnual
FreeFreeFree
DeveloperContact sales
EnterpriseContact sales
10

The Anthropic API, Claude for long-context reasoning and code
Anthropic API (Claude)

84Disclosure$20Seat / mo

Precision work on long inputs: contract review, large refactors, anything where drifting off the spec is dearer than the tokens. No subscription exists here at all, and what you pay is what the model reads and writes.

Plan table and expert take

Anthropic API (Claude): expert take

Caching cuts both ways. A cache hit costs about a tenth of base input, but writing to the cache costs more than the input itself, a quarter more on the short-lived tier and double on the long one. Reused prompts win, one-shot calls do not.

Where Anthropic API (Claude) holds up

  • Up to 1M token context on the frontier models, enough to load an entire codebase or contract set in one call
  • Opus 4.8 for the hardest reasoning, Sonnet 5 for balance, Haiku 4.5 for cheap volume, priced per workload
  • Holds long instructions and output formatting better than most under load, which matters for coding agents
  • Constitutional AI training keeps harmful output low, which is why safety-sensitive teams pick it
  • Tool use, JSON mode, and a Batch API that halves the rate for offline high-volume jobs

4.8CE scoreG2 4.7 · 297 reviewsCapterra 4.6Founded 2021Verified July 8, 2026

6 plans, as published
PlanMonthlyAnnual
FreeFreeFree
Pro$20$17
Max$100Not published
Max x20Contact sales
Team$20Not published
EnterpriseContact sales
11

DeepSeek, the low-cost open model for coding and reasoning
DeepSeek

83DisclosureSales onlyNo price

Volumes where the inference bill outgrows the engineer running it: high-volume coding, step-by-step reasoning, long multilingual jobs. DeepSeek sells nothing by the month. You top up a balance and it drains by the token.

Critical gapThe architecture requires constant internet connectivity, preventing secure offline local deployment.

Plan table and expert take

DeepSeek: expert take

The flagship is the throttled one: V4-Pro allows five hundred concurrent requests against two and a half thousand on Flash, so the bigger model is the narrower pipe. Cache hits are the deepest discount in this category, roughly fifty times cheaper than a miss.

Where DeepSeek holds up

  • V4-Flash at $0.14 per 1M input undercuts the frontier labs by a wide margin, and cache hits cut that ~98% more
  • Open weights you can actually fine-tune and self-host, no gatekeeping and no vendor lock-in
  • Strong multilingual work, Mandarin especially, where the Western models are weaker
  • Chain-of-thought reasoning that shows its steps, which compliance teams can audit line by line
  • The web and mobile chat is free, so evaluating the model costs nothing before you touch the API

4.7CE scoreG2 4.6 · 14 reviewsCapterra 4.6Founded 2023Verified July 8, 2026

2 plans, as published
PlanMonthlyAnnual
FreeFreeFree
API Pay-as-you-goContact sales
12

Claude, Anthropic's assistant for long-form writing and code
Claude

82Disclosure$20Seat / mo

Once a chat box stops holding the work: long drafts, research files, a codebase pasted in whole with a spec that has to survive it. Team plans have a floor of five seats and a ceiling of 150, which prices Claude for a working group rather than an entire company.

Critical gapAPI rate limits hinder high-volume, automated production system integration.

Plan table and expert take

Claude: expert take

Claude names its meter inside a consumer plan, which is rare: Pro runs on a five-hour window worth roughly forty-four thousand tokens, ten to forty prompts. On the enterprise tier the pretence drops and usage bills at standard API rates.

Where Claude holds up

  • Industry-leading 1M token context window for deep analysis
  • Excels at nuanced writing, summarization, and creative tasks
  • Strong constitutional AI framework prioritizes safety and ethics
  • Artifacts feature for iterative code generation and editing
  • Generous free tier with access to the powerful Sonnet model

4.7CE scoreG2 4.6 · 297 reviewsCapterra 4.5Founded 2023Verified July 8, 2026

7 plans, as published
PlanMonthlyAnnual
FreeFreeFree
Pro$20$17
Max 5x$100Not published
Max 20x$200Not published
Team Standard$25$20
Team Premium$125$100
EnterpriseContact sales
13

ChatGPT, OpenAI's general assistant for everyday work
ChatGPT

78Disclosure$7Seat / mo

Everyday drafting, rewriting and quick lookups for one person or a handful of colleagues who want a single assistant across all of it. Ten messages every five hours is where the free plan runs out, so the cheap paid step arrives sooner than people expect.

Critical gapThe model produces inaccurate mathematical calculations and unreliable citations during complex, data-heavy research tasks.

Plan table and expert take

ChatGPT: expert take

The $7 Go plan is the cheap step that can carry advertising, which the plan card says outright. Above it the ladder stops selling features and starts selling multiples of Plus usage, five times and twenty times, each at its own price.

Where ChatGPT holds up

  • Frontier models on tap: GPT-5.5 for hard reasoning, cheaper GPT-5 and mini tiers for volume
  • Largest third-party ecosystem, custom GPTs, and connector support of any assistant
  • Multimodal in one place: voice, images in and out, file uploads, and web search
  • Free tier is genuinely usable, not a teaser
  • Simple interface that non-technical people pick up in minutes

4.8CE scoreG2 4.7 · 2,268 reviewsCapterra 4.6Founded 2022Verified July 8, 2026

8 plans, as published
PlanMonthlyAnnual
FreeFreeFree
Go$7Not published
Plus$20Not published
Pro 5x$120Not published
Pro 20x$200Not published
CodexContact sales
Business$25$20
EnterpriseContact sales
14

Replicate, the API for deploying open-source models without servers
Replicate

78DisclosureSales onlyNo price

Putting an open-source model behind an API without owning a GPU or a serving stack. Replicate keeps a cost estimate on each model's page, so the money question gets answered where you choose the model.

Critical gapThe platform lacks production support and exposes developers to high-cost, inefficient execution.

Plan table and expert take

Replicate: expert take

One vendor, two billing currencies: some models bill hardware by the second, others bill input and output tokens, and which one applies is a property of the model rather than your account. Multi-GPU capacity opens only under a spend commitment.

Where Replicate holds up

  • Highly rated (4.5/5 on review platforms)
  • 12 key features including 50K+ models and Simple API
  • Growing user base (200K+)
  • API access for custom integrations

4.4CE scoreG2 4.3 · 110 reviewsCapterra 4.4Founded 2021Verified July 8, 2026

2 plans, as published
PlanMonthlyAnnual
Pay-as-you-go$0Not published
EnterpriseContact sales
15

Command R+, Cohere's retrieval-focused model with grounded citations
Command R+

78DisclosureSales onlyNo price

Retrieval with grounded citations and a couple of tool calls behind it, inside an enterprise stack. Command R+ carries exactly one line of pricing, consumption, which leaves nothing to compare but the rate itself.

Plan table and expert take

Command R+: expert take

The ceiling hides in the limits line: the context window runs to 128k tokens while output stops at four thousand, which quietly caps what a single call can return. Output is priced four times input on that same line.

Where Command R+ holds up

  • Grounded RAG with in-line citations, so every answer traces back to a source you can check
  • Multi-step tool use for chaining real business workflows, not just single-shot completions
  • Ten-language coverage for global enterprise deployments out of the box
  • A 128k context window at a mid-tier price that reads long documents without splitting them
  • Deployable in cloud, VPC, or on-prem, so regulated data can stay inside your walls

4.6CE scoreG2 4.5 · 83 reviewsCapterra 4.2Founded 2024Verified July 8, 2026

1 plans, as published
PlanMonthlyAnnual
Command R+ UsageContact sales
16

Google AI Studio, the prototyping console for the Gemini API
Google AI Studio

77DisclosureSales onlyNo price

The first hour of a project, when the only thing to settle is whether the prompt works at all. AI Studio hands over a free key and a build flow; paid access switches on only after you buy credits up front, with a floor on that first purchase.

Critical gapThe platform lacks offline local execution capabilities for secure internal development environments.

Plan table and expert take

Google AI Studio: expert take

Surcharges outnumber discounts on this one. Priority runs about double standard, a prompt past two hundred thousand tokens doubles the rate again on the larger models, and cached context is billed for storage by the hour rather than per hit.

Where Google AI Studio holds up

  • Free access to Gemini 3.1 Pro with a context window near 1M tokens, no card required to start prompting
  • An API key drops out of the same screen, so a prototype moves to code in minutes
  • System instructions pin model behavior without you rebuilding the prompt every call
  • Function calling and code execution are built in, so you can test tool use and run Python inline
  • Search grounding wires answers to live results, and text, image, audio, and video all go in as input

4.3CE scoreG2 4.2 · 1,028 reviewsCapterra 4.4Founded 2023Verified July 16, 2026

3 plans, as published
PlanMonthlyAnnual
FreeFreeFree
Paid TierContact sales
EnterpriseContact sales
17

Mistral AI, the EU model maker with open-weight and flagship options
Mistral AI

77Disclosure$5.99Seat / mo

Teams that must keep data in Europe or on their own machines and want to match model size to the task. Mistral publishes open weights and a hosted flagship together, so self-hosting and API work sit under one roof.

Critical gapPerformance drops during complex reasoning and long-context conversation windows.

Plan table and expert take

Mistral AI: expert take

The $24.99 Team plan is not the whole team price: a base account fee sits on top of the per-user rate, so the seat figure understates the invoice. Document recognition run in batch costs about half the standard rate.

Where Mistral AI holds up

  • Open-weight models you can download and self-host, so data never has to leave your own infrastructure
  • An EU base with GDPR-native handling, which clears sovereignty reviews that block US-hosted providers
  • A wide model menu, Large 3, Medium 3.5, Small 4, Codestral, Magistral, Pixtral, so you match the model to the job and the budget
  • La Plateforme bills pay-as-you-go per token with no subscription requirement, and Small 4 is genuinely cheap to run
  • Function calling, JSON mode, and the Agents API make it a real base for structured, tool-using applications

4.6CE scoreG2 4.5 · 13 reviewsCapterra 4.4Founded 2023Verified July 16, 2026

5 plans, as published
PlanMonthlyAnnual
FreeFreeFree
Pro$14.99Not published
Team$24.99Not published
EnterpriseContact sales
Education$5.99Not published
18

Mistral Large, Mistral's flagship kept on European infrastructure
Mistral Large

76Disclosure$5.99Seat / mo

Structured workflows that have to stay on European infrastructure: function calling, JSON output, a flagship kept in region. Consumer plans and API access are sold as separate lines under the same name.

Critical gapUsers report a lack of creativity and bland output.

Plan table and expert take

Mistral Large: expert take

The $5.99 Mistral Pro (Student) plan is the interesting line: the same consumer product at well under half the standard rate, proof of enrolment required. Twelve months is the cap on it, after which the student rate expires.

Where Mistral Large holds up

  • GDPR-native with EU hosting, which clears data-residency and sovereignty reviews that block US-hosted models
  • The API bills $2 per 1M input and $6 per 1M output, and batch jobs run at half that, which is gentle for a flagship-class model
  • Real multilingual depth across French, German, Spanish, Italian, and Portuguese, not just English with translation bolted on
  • Function calling and constrained JSON output make it dependable inside structured, tool-using workflows
  • Open-weight Mistral releases sit alongside it, so you can self-host for the workloads that cannot leave your walls

4.4CE scoreG2 4.3 · 13 reviewsCapterra 4.2Founded 2024Verified July 8, 2026

5 plans, as published
PlanMonthlyAnnual
Mistral FreeFreeFree
Mistral Pro$14.99Not published
Mistral Pro (Student)$5.99Not published
API (Mistral Large)Contact sales
EnterpriseContact sales
19

Grok 2, xAI's model wired to live data on X
Grok 2

71Disclosure$30Seat / mo

Current-events work that needs the live feed rather than a training cut-off, with image generation sitting beside the chat. Grok's paid steps buy compute priority rather than new features, so what you are choosing is how hard you plan to push it.

Plan table and expert take

Grok 2: expert take

The personal ladder makes one jump and it is tenfold: $30 SuperGrok to $300 SuperGrok Heavy, with nothing in between. A Business seat costs the same $30 as a personal one, so the organisation pays no premium here.

Where Grok 2 holds up

  • Live X data access for genuinely current answers on breaking topics
  • Native image generation from the Aurora model, built into the same interface
  • An unfiltered, direct conversational tone that suits informal exploration
  • Bundled with X Premium at no extra charge for existing subscribers
  • Open model weights released for researchers to inspect and build on

4.3CE scoreG2 4.2Capterra 4.1Founded 2024Verified July 8, 2026

5 plans, as published
PlanMonthlyAnnual
FreeFreeFree
SuperGrok$30Not published
SuperGrok Heavy$300Not published
Business$30Not published
EnterpriseContact sales
20

Phi-3, Microsoft's small open model for on-device work
Phi-3

71DisclosureSales onlyNo price

On-device and edge work where a small footprint and tight instruction-following beat raw breadth. The family ships seven variants under one listing, none of them carrying a plan price, which leaves the rate card as the whole commercial story.

Plan table and expert take

Phi-3: expert take

Fine-tuning is billed in three separate layers here: training per million tokens, hosting by the hour, and inference on top of both. The two medium variants cost the same whether you take the short or the long context window.

Where Phi-3 holds up

  • Runs efficiently on-device, putting offline AI on phones, IoT hardware, and modest laptops with no cloud call
  • MIT license allows commercial use with almost no restrictions, and self-hosting carries no per-token fee
  • Beats several larger models on reasoning benchmarks like MMLU and GSM8K for its parameter count
  • Quantized builds run on CPU, so you avoid the expensive GPU requirement of bigger models
  • A 128K context option on a model this small and cheap to serve, which is unusual at the size

4.1CE scoreG2 4Capterra 4Founded 2024Verified July 8, 2026

7 plans, as published
PlanMonthlyAnnual
Phi-3-mini-4k-instructContact sales
Phi-3-mini-128k-instructContact sales
Phi-3.5-mini-instructContact sales
Phi-3-small-8k-instructContact sales
Phi-3-small-128k-instructContact sales
Phi-3-medium-4k-instructContact sales
Phi-3-medium-128k-instructContact sales
21

Llama, Meta's open-weight models you run yourself
Llama (Meta)

63DisclosureFreePaid tier unpriced

Running a model where nothing leaves your network and no vendor sits in the loop. Llama's weights cost nothing to license until the service using them reaches hundreds of millions of monthly users, at which point the licence stops being free.

Critical gapThe deployment process demands high technical overhead and extensive documentation for effective model integration.

Plan table and expert take

Llama (Meta): expert take

The rates attached to this model belong to a third-party host, not to Meta: the first-party API is waitlisted and publishes no price at all. What you actually pay for is your own hardware, which is the only meter that applies.

Where Llama (Meta) holds up

  • A permissive license that allows commercial use and modification, not just research
  • Strong current-generation performance for open weights across the Llama 4 family
  • Full data control and privacy, since the model runs entirely on infrastructure you own
  • A huge ecosystem of fine-tuned variants and tooling on Hugging Face to build from
  • Multiple parameter sizes, so you can match the model to the hardware you actually have

4.7CE scoreG2 4.6 · 152 reviewsCapterra 4.7Founded 2023Verified July 8, 2026

2 plans, as published
PlanMonthlyAnnual
Open WeightsFreeFree
Enterprise LicenseContact sales
05 / 08

Compare any two models: context window, capabilities and our score

vs

What the records say

Amazon Nova publishes no monthly figure, so there is no team bill to line up against Kimi.

Kimi carries 4 of the 5 capability columns on the record; Amazon Nova shows 3.

Kimi runs a free tier to start on; Amazon Nova does not.

Both publish every tier they sell.

On the meter, Amazon Nova starts at $0.04 in / $0.14 out per million tokens; Kimi starts at $0.20 / $2.

Pick Amazon Nova for: High-volume inference for teams whose data already sits in AWS, billed on the account that carries everything else.

Pick Kimi for: Near-frontier reasoning where the API line item decides which vendor wins.

01

Amazon Nova

CE 100
Published plans, US$/mo
Nova Micro$0.04 / 1M tok
Nova 2 Lite$0.30 / 1M tok
Nova Act$4.75 / hr
Team of 6Not published
API meter, from$0.04 / $0.14 per 1M tokens

Verified July 8, 2026

02

Kimi

CE 100
Published plans, US$/mo
AdagioFree
Allegretto$39
API Pay-as-you-go$3 / 1M input tokens
Team of 6$114 / mo
API meter, from$0.20 / $2 per 1M tokens

Verified July 17, 2026

Both price lists on the category axis

Amazon NovaUsage-priced, public rate card
Kimi

Where they differ

Only Amazon Nova has on the record

  • JSON mode

Only Kimi has on the record

  • $0 tier
  • Web search
  • Open weights
06 / 08

LLM questions: subscription against tokens, top tiers, rate conversion

Is a subscription cheaper than paying by the token?

The answer turns on who does the work. A seat used all day is predictable and capped, which is what makes it comfortable. Metered access costs nothing while idle and nothing holds it down in a busy week. Teams shipping a product usually pay both ways: seats for the people, tokens for the software they built.

What does the top consumer tier actually buy?

More of the same model, in most cases. The premium steps on the big chat plans are sold as multiples of the tier below: five times the usage, or twenty. The model underneath is the one you already had. If your problem is answer quality rather than running out of messages, that upgrade will disappoint you.

How do you turn a per-million-token rate into a monthly bill?

Count tokens, not requests. Take one typical exchange, reckon roughly four characters to a token, multiply input and output separately by their own rates, then multiply by the calls you expect in a month. Output dominates almost every real bill, so estimate that side generously. The calculator above does the seat half of the same arithmetic.

Why does output cost more than input?

Writing costs more than reading. A prompt goes through in a single pass, while the answer is produced one token at a time, and every card here prices the second job above the first. The effect sneaks up on people: a short question with a long answer runs dearer than a long document summarised in one line.

What does a free LLM tier not include?

Capacity and patience, mostly. Free plans cap messages inside a rolling window, hold back the newest models, and put you behind paying traffic when the service is busy. One consumer free tier openly carries advertising and warns it may slow or skip heavy jobs. On the developer side, free normally means a key with tight rate limits.

Do batch and caching discounts actually lower the bill?

They do, more than most teams expect, but nothing happens on its own. Batch endpoints trade latency for roughly half the standard rate. Cache hits go deeper, in one case to about a fiftieth of a miss. Both have to be built for: batch is a separate endpoint, and caching only pays off when the prompt is structured to repeat.

Is the cheapest rate card the cheapest bill?

Not reliably. A cheap model that reasons out loud emits far more output than a pricier one that answers straight, and output is the expensive half. Retries, a system prompt resent on every call and a document pasted into each turn all land on the meter too. Price per million is an input to the estimate, not the estimate.

Why does one lab sell both a chat plan and an API key?

They are two products with two meters. A subscription buys one person a seat and a usage allowance. The API sells capacity to your code at a published rate. Several labs appear twice on this page for that reason, and the meter shows matching rates on both entries: buying a seat does not change what the code pays.

Are open-weight models really free to run?

The licence is free. The hardware is not. Serving a mid-sized open model in production means GPUs, someone to keep them healthy, and a stack you now own. Under a certain volume a hosted API costs less than the machine you would buy. Above it the math reverses, so the open route is a volume decision before it is a philosophical one.

Can a small team share one paid plan?

Up to a point, and the team tiers set that point. Business and team plans here start at a floor of two to five seats depending on the vendor, and at least one has a ceiling of 150. Sharing a single personal login is worse value than it looks, because the usage window belongs to the account and not to each person using it.

Does a longer context window change the price?

Sometimes the rate itself changes. One vendor doubles the token rate on its larger models once a prompt passes the two-hundred-thousand-token mark. Long context is also where cached input starts to matter, because the same document is resent on every turn. Even at a flat rate, a bigger window simply means more tokens per call.

Which LLMs publish no per-token rate at all?

Silence takes a few forms here. Meta AI sells a consumer product with no public API and no token price. Hugging Face bills GPU minutes and storage instead of tokens. Qwen's larger models no longer appear on Alibaba's own commercial list. Llama's first-party API sits behind a waitlist, so circulating rates belong to whoever hosts the weights.
Field note 01

What raises an LLM token rate after you have picked the model?

Discounts get advertised. Surcharges do not. Priority service runs at a premium over the standard rate. A prompt past two hundred thousand tokens can double the rate on the larger models. Endpoints that keep processing inside one region add a percentage on newer models. Writing to a prompt cache is dearer than the input it saves, and speech recognition bills a floor of ten seconds however short the clip.

None of that reaches the headline figure. Treat the rate as the floor of the meter, then go hunting for the modifiers, because they decide whether your estimate resembles the invoice.

Field note 02

Stale LLM token charts, and why the figures here disagree

Aggregated token charts age badly and rarely say when they were built. Every number here comes with the date of its last check at the source that sets it. Where nothing is printed, nothing is invented, and no marketplace listing is quietly promoted into a vendor price. One card in the catalog belongs to a third-party host rather than the model's own maker, so we keep it off the meter.

What is worth checking is not the second decimal on a rate. It is whether the price came from the maker, whether anyone read it recently, and whether the tiers underneath it were opened at all.

The verdict on language modelsSigned review · Updated
Oleh KemFounder & Lead AnalystComparEdge Editorial

Work out which of the two things you are buying before you compare anything. A monthly seat and a per-million-token card answer different questions, and a figure from one tells you nothing about the other. Most teams end up paying both. The subscription covers the people, the card covers the code.

The trap sits on either side. Consumer tiers sell multiples of usage rather than a better model, so the step up buys headroom you cannot see until you hit it. On the card side the headline is the input rate, while output is where the bill actually lands. What we score is disclosure, not model quality. A vendor that gives its weights away scores well on openness and still leaves you the entire hardware bill.

MethodWe read 93 plans and every published rate card across twenty-one vendors, checking each figure at its source on .
DisclosureCollection is tool-assisted; every verdict is written and signed by a human analyst.
08 / 08

Read next: cost guides for language models, plus related categories

How this review is made. Prices are read from vendor pricing pages and re-checked on the dates shown against each product. Condition columns reflect the feature set recorded on the vendor’s own pages on that date. ComparEdge sells no language models and takes no vendor payment for placement. Where a vendor publishes nothing, this page says so rather than estimating. Ranking is by transparency score: pricing transparency 60%, user satisfaction 40%. What a product can do is shown in the condition columns and carries no weight in the number.