Groq pricing plans
★★★★★ 4.5 CE

Groq Pricing: Plans & API Token Cost Calculator 2026

Free tier plus usage-based per-token rates among the lowest for hosted inference. You trade a narrower model catalog for the speed.

Groq interface screenshot

Groq plans and pricing

High· Verified July 8, 2026
Free

Free

A rate-limited free tier with access to the full open-weight catalog, enough for development and low-traffic apps

Free
Build and Test on Groq APIs
Community Support
Zero-data Retention Available
Get started
Custom

Developer

Custom
Build and Test on Groq APIs
Community Support
Zero-data Retention Available
Higher Token Limits
Chat Support
Flex Service Tier
Batch Processing
Spend Limits
Prompt Caching
Contact sales
Custom

Enterprise

Custom
Build and Test on Groq APIs
Community Support
Zero-data Retention Available
Higher Token Limits
Chat Support
Flex Service Tier
Batch Processing
Spend Limits
Prompt Caching
Custom Models
Regional Endpoint Selection
Performance Tier
Scalable Capacity
Dedicated Support
LoRA Fine-Tunes
Contact sales

Groq pricing: the quick answer

Quick answerHigh· Verified July 8, 2026

Groq bills pure pay-as-you-go token rates as of July 8, 2026, with no subscription tiers and a free plan for testing. Free gets limited usage, Developer is usage-based at 10x higher rate limits, Enterprise is custom with dedicated capacity. From the catalog, Llama 4 Scout runs a fraction of a cent per 1M tokens, GPT-OSS 120B runs $0.15 in and $0.60 out per 1M, and batch jobs take 50% off. No flat fee to plan around, so your bill is whatever the rate card and your volume add up to.

  • FreeFree
  • DeveloperCustom
  • EnterpriseCustom
Run your token volumes through the cost calculator to estimate the actual monthly API bill.
Free tier
Yes
Billing model
Freemium
Annual discount
Not offered

Groq is free to start, against a $8.50/mo median across 10 large language models tools we track.


API Token Pricing

Llama 3.1 8B Instant$0.05$0.08
GPT OSS 20B$0.075$0.3
GPT OSS Safeguard 20B$0.075$0.3
GPT OSS 120B$0.15$0.6
Llama 3.3 70B Versatile$0.59$0.79
Qwen 3.6 27B$0.6$3

GroqCloud on-demand rates, groq.com/pricing, read 2026-07-22

Groq cost calculator

Groq Hidden Costs & Pitfalls

What sits on top of the plan fee

There is no flat subscription to budget against here. The cost is the per-model, per-tool rate card, and it runs wider than most buyers expect.

Model rate spread
Cheapest to priciest LLM on the catalog
$0.05-$0.60 per 1M tokens (in), $0.08-$3.00 per 1M (out)
Batch API discount
Async jobs, 24h-7-day processing window
50% off standard rate
Prompt caching
Cache hit only, no fee to enable caching
~50% off cached input tokens
Web search tool
Built-in Compound tool, basic vs advanced search
$5-$8 per 1,000 requests
Whisper transcription
ASR billed with a 10-second per-request minimum
$0.04-$0.111 per hour
The renewal jump, credit expiry and every discount worth having are in the Groq true cost guide, with an email generator for negotiating team and enterprise rates.
Groq Cost Analysis

Groq pricing, read against its live plans and category

Positioning

Groq skips the subscription entirely: a free tier plus pay-as-you-go token pricing. Where the category median sits at $8.4/mo, Groq's Developer and Enterprise tiers bill by usage rather than a flat monthly fee. For agentic workloads that make a lot of fast calls, per-token pricing on cheap open models is efficient, and the platform is worth it precisely when your priority is raw inference speed and low token cost rather than model breadth.

Cost drivers

  • 1Pricing is straightforward; no documented hidden fees or overage traps found.

Watch-outs

Groq's inference is cheap and fast, but leaning on the free tier for production is risky: strict rate limits and latency that spikes under load. Users report speeds that are generally unmatched alongside consistency that can swing hard when traffic is heavy.

Strengths

The free tier already delivers the headline capability: the fastest inference speeds on the market, at no cost.

  • World's fastest inference speed (500+ tokens/sec)
  • Custom LPU hardware eliminates sequential processing bottlenecks
  • OpenAI-compatible API for near drop-in integration

What users say

I built a Study OS with Llama 4 + Groq because Otter was too expensive.

Reddit

I tested the playground inference on their website. Insane speeds.

Reddit

Editor’s take

Individual developers and startups should start on the Free tier to test the API, then move to the Developer tier for higher token limits and the Flex service tier as production scales. For dedicated capacity and support, the custom Enterprise tier is the path. If you want a predictable flat-rate plan with built-in frontier models instead, Google Gemini at $20/mo covers it.

Oleh KemOleh KemFounder & Lead Analyst
ComparEdge EditorialUpdated: July 8, 2026

Groq price history


Price & Data Intelligence SyncLast verified: July 31, 2026 · CE-LLM-2026W31-E8D96F · No changes detected
Up to date

Cheaper Large Language Models tools


Frequently asked questions


Sources & verification

Verified by ComparEdgeMethod: Vendor docs, official pages, and selected independent sources
SourceWhat was checkedLast checked
Official Pricing PageSource of verified tiersJuly 8, 2026
Official WebsiteOfficial vendor website
PeerSpotPeerSpot enterprise peer reviews

Every fact on this Groq pricing page is tied to a named source and a verification date. Freshness-sensitive figures trace to the sources above; verify against the vendor before relying on them.

Explore Groq

Every page on Groq in one place, you are on pricing.