Replicate pricing plans
★★★★ 4.4 CE

Replicate Pricing: Plans & API Token Cost Calculator 2026

Free tier access lets you run models with zero setup costs, transitioning to a pay-as-you-go model with no flat monthly fees.

Replicate interface screenshot

Replicate plans and pricing

High· Verified July 8, 2026
Custom

Enterprise

Get volume discounts, SOC 2 compliance, and dedicated support

Custom
Volume discounts on compute usage
SOC 2 Type II compliance
Dedicated support channel and custom SLAs
Private deployments and VPC peering
Consolidated billing and custom invoicing
Contact sales

Replicate pricing: the quick answer

Quick answerHigh· Verified July 8, 2026

Replicate has no subscription; it bills per second of compute as of July 8, 2026, with a custom Enterprise tier for volume discounts. There is no monthly floor, so an idle account costs nothing, and you pay only while a model runs. Hardware sets the rate: a small CPU is $0.09/hr, an Nvidia T4 is $0.81/hr, an A100 80GB is $5.04/hr, and an H100 is $5.49/hr. Some models bill by tokens instead, around $0.10 per 1M input and $0.50 per 1M output, though the exact rate varies by model. Scale-to-zero keeps prototyping cheap, but a busy production model on an A100 can run into real money fast.

  • Pay-as-you-goFree
  • EnterpriseCustom
Run your token volumes through the cost calculator to estimate the actual monthly API bill.
Free tier
Yes
Billing model
Token-Based
Annual discount
Not offered

Replicate is free to start, against a $8.50/mo median across 10 large language models tools we track.


API Token Pricing

Various models$0.1$0.5

Replicate cost calculator

Replicate Hidden Costs & Pitfalls

What sits on top of the plan fee

Per-second billing sounds friendly until a model stays warm. The cost drivers are the hardware tier you land on, the cold-start seconds you still pay for, and the multi-GPU rigs locked behind contracts.

Hardware tier decides your per-second rate
The same prediction costs wildly different amounts depending on the GPU. A T4 at $0.81/hr is fine for light image work, but an A100 80GB at $5.04/hr or an H100 at $5.49/hr adds up quickly under load. A model kept warm on an A100 for a full day costs about $121, and if traffic keeps it running around the clock that is roughly $3,600 a month. Match the model to the cheapest hardware that runs it well, because the tier is where the money goes.
$0.09 to $5.49 per hour
Cold starts bill before any work happens
Scale-to-zero means an idle model spins down, but the next request has to spin it back up, and that cold-start time is billed at the per-second rate before your actual inference begins. For a rarely-hit endpoint on expensive hardware, you can pay a meaningful slice of every request just waiting for the model to load. High-traffic services amortize this away; bursty ones feel it on every wake-up.
billed at hardware rate
Multi-GPU capacity needs a committed-spend contract
The single-GPU tiers are self-serve, but the 4x and 8x rigs are gated behind committed-spend contracts, not available on demand. An 8x H100 lists at $43.92/hr, and you cannot simply switch it on; you negotiate access through Enterprise. If your workload genuinely needs that much parallel compute, budget for a contract commitment rather than assuming pay-as-you-go covers it.
$43.92/hr, contract-gated
The renewal jump, credit expiry and every discount worth having are in the Replicate true cost guide, with an email generator for negotiating team and enterprise rates.
Replicate Cost Analysis

Replicate pricing, read against its live plans and category

Positioning

Replicate bypasses the traditional subscription model, making its free entry point look highly attractive compared to the category median of $8.4/mo. Users pay strictly for what they use via per-second billing on CPU and GPU compute, with the ability to scale to zero instances automatically. While the Pay-as-you-go plan provides access to thousands of public models and custom deployments via Cog, heavy production workloads can quickly become expensive. For high-volume organizations, the Enterprise plan offers volume discounts, custom SLAs, and VPC peering, but requires contacting sales for custom pricing.

Cost drivers

  • 1Compute scaling inefficiencies: Idle cold-start times can inflate per-second billing before actual processing begins.
  • 2Infrastructure overhead: High-volume API calls can quickly outpace the cost of self-hosting on dedicated cloud providers.

Watch-outs

Users have reported sudden shifts in billing infrastructure, such as being forced from standard monthly invoicing to a pre-paid credit system. There are also warnings regarding high infrastructure costs when running automated pipelines at scale.

Strengths

Even the free tier offers highly rated (4.5/5 on review platforms) - strong value at no cost.

  • Highly rated (4.5/5 on review platforms)
  • 12 key features including 50K+ models and Simple API
  • Growing user base (200K+)

What users say

easiest to use option for trying out new image or video models

Reddit

each image generated is typically 1-2c... this adds up to a few dollars.

Reddit

Editor’s take

Individual developers and teams prototyping new AI concepts should start with the Pay-as-you-go plan to exploit the scale-to-zero efficiency. Large enterprises running continuous pipelines should negotiate the Enterprise plan to secure volume discounts and prevent runaway compute bills. If you require predictable flat-rate pricing for standard conversational AI instead of raw model hosting, consider ChatGPT at $20/mo.

Oleh KemOleh KemFounder & Lead Analyst
ComparEdge EditorialUpdated: July 8, 2026

Replicate price history


Price & Data Intelligence SyncLast verified: July 30, 2026 · CE-LLM-2026W31-CB7FB0 · No changes detected
Up to date

Cheaper Large Language Models tools


Frequently asked questions


Sources & verification

Verified by ComparEdgeMethod: Vendor docs, official pages, and selected independent sources
SourceWhat was checkedLast checked
Official Pricing PageSource of verified tiersJuly 8, 2026
Official WebsiteOfficial vendor website
G2G2 verified user reviews · 4.3/5 · 110 reviews
CapterraCapterra verified user reviews · 4.4/5
TrustRadiusTrustRadius verified reviews
PeerSpotPeerSpot enterprise peer reviews

Every fact on this Replicate pricing page is tied to a named source and a verification date. Freshness-sensitive figures trace to the sources above; verify against the vendor before relying on them.

Explore Replicate

Every page on Replicate in one place, you are on pricing.