Comparing Rate Limits and Pricing Tiers Across Major AI APIs as You Scale

Rows of server racks in a data center

Rate limits and pricing tiers differ across Anthropic, OpenAI, and Google’s Gemini API in three concrete ways: how you unlock a higher tier, what the published ceiling actually measures, and how much a million tokens costs once you’re past any free allowance. Anthropic and OpenAI both move you up automatically based on cumulative spend and time since your first payment, while Google layers a similar spend-and-time test on top of a separate rolling-window spend cap. The ceilings themselves aren’t apples-to-apples either: Anthropic publishes exact requests-per-minute and tokens-per-minute numbers per tier, OpenAI publishes the dollar thresholds but not a public RPM/TPM table, and Google publishes some of the qualification math but directs you to a live per-account dashboard for the actual numbers.

Every figure below is pulled from each provider’s own documentation as published in October 2026, with the source URL next to the number it supports. This is a fast-moving area — all three providers have changed tier structures or prices within the last year — so treat this as a snapshot, not a permanent reference, and check the linked pages before you build a capacity plan on top of these figures.

What “Rate Limit,” “RPM/TPM,” and “Usage Tier” Actually Mean

A rate limit is a ceiling an API provider enforces on how much traffic one account can send in a given time window, independent of how much you’re willing to pay per request. RPM (requests per minute) counts API calls; TPM (tokens per minute) counts the tokens processed, usually split into input and output because output is slower and more expensive to generate; RPD (requests per day) is a secondary daily ceiling some providers add on top of the per-minute ones. A usage tier is the account-status bucket a provider assigns you — it decides which RPM/TPM numbers apply to your account and, in most cases, is not something you select directly; it changes automatically as your billing history changes.

How Each Provider Decides Which Tier You’re In

All three providers tie tier progression to money spent and, in two cases, to time elapsed since your first payment — but the specific thresholds and the information they publish about the resulting limits vary a lot.

  • Anthropic defines five tiers: Evaluation (new organizations, reduced limits as an abuse safeguard), Start (monthly spend cap of $500), Build ($1,000), Scale ($200,000), and Custom (negotiated with an account team, no published cap). Progression through Start, Build, and Scale is automatic based on usage history and account standing, according to Anthropic’s own rate-limits documentation (platform.claude.com/docs/en/api/rate-limits).
  • OpenAI uses six tiers — Free, and Tier 1 through Tier 5 — gated by cumulative dollars paid: Tier 1 requires $5 paid, Tier 2 requires $50, Tier 3 requires $100, Tier 4 requires $250, and Tier 5 requires $1,000, each tier also carrying its own monthly organizational spend limit ($100 for Free and Tier 1, $500 for Tier 2, $1,000 for Tier 3, $5,000 for Tier 4, and $200,000 for Tier 5). OpenAI’s guide states plainly that “as your spend on our API goes up, we automatically graduate you to the next usage tier,” which usually raises rate limits across most models (developers.openai.com/api/docs/guides/rate-limits).
  • Google’s Gemini API adds a time condition on top of the spend condition. Its documentation describes a Free tier (available to any project, or on a free trial, with no spend-based rate cap), a Tier 1 that requires linking an active billing account (billing cap $250), a Tier 2 that requires having paid at least $100 and waiting at least 3 days from your first successful payment (billing cap $2,000), and a Tier 3 that requires having paid at least $1,000 and waiting at least 30 days from your first successful payment (billing cap in the $20,000–$100,000+ range) (ai.google.dev/gemini-api/docs/rate-limits).

The Comparison Table: How Tiers Scale

The table below uses each provider’s own published numbers; where a provider does not publish a static figure, that is stated explicitly rather than estimated.

Provider How a tier is unlocked What scales with tier (RPM / TPM) Notable scaling mechanic
Anthropic (Claude API) Automatic, based on usage history and monthly spend cap: Start $500/mo, Build $1,000/mo, Scale $200,000/mo, Custom negotiated (source) Published per model, e.g. Start tier: 1,000 RPM / 2,000,000 input TPM / 400,000 output TPM; Build tier: 5,000 RPM / 5,000,000 input TPM / 1,000,000 output TPM; Scale tier: 10,000 RPM / 10,000,000 input TPM / 2,000,000 output TPM (source) Only uncached input tokens count against the input-TPM limit — tokens served from the prompt cache are excluded, so a high cache-hit workload can push effective throughput well above the published number (source)
OpenAI (API) Automatic, based on cumulative dollars paid: Tier 1 at $5 paid, Tier 2 at $50, Tier 3 at $100, Tier 4 at $250, Tier 5 at $1,000; each tier also has its own monthly org-wide spend limit ($100 to $200,000) (source) OpenAI’s guide does not publish a static RPM/TPM table by tier; it states limits “vary by the model being used” and directs accounts to check their own dashboard for current numbers (source) Batch API queue limits are measured in total tokens queued per model rather than RPM/TPM, and very high-volume customers can negotiate a separate Scale Tier or Reserved Tier for dedicated, more predictable capacity (source)
Google (Gemini API) Free tier needs no payment; Tier 1 needs a linked billing account; Tier 2 needs $100 paid and 3 days elapsed since first payment; Tier 3 needs $1,000 paid and 30 days elapsed (source) Google’s rate-limits page states that exact RPM/TPM/RPD figures “depend on a variety of factors” and directs users to view their account’s live numbers in AI Studio rather than publishing one universal table (source, dashboard: aistudio.google.com/rate-limit) Separate from RPM/TPM, Google enforces a rolling 10-minute spend-based rate limit: $10 per 10 minutes at Tier 1, $50 at Tier 2, $200 at Tier 3 — a ceiling on dollars burned in a short window, not just on request count (source)

What a Million Tokens Actually Costs

Price per tier is a separate question from rate limits, and the three providers’ current flagship, mid-tier, and lightweight model prices (per million tokens, standard non-batch rate) are shown below.

Provider / model Input ($/M tokens) Output ($/M tokens) Notes
Anthropic — Claude Opus 5.5 $4.00 $20.00 Prompt caching: 5-minute cache write at 1.25x input rate, 1-hour write at 2x, cache read at 0.1x; Batch API gives a 50% discount on both input and output (source)
Anthropic — Claude Sonnet 5.5 $2.00 $10.00 Same caching/batch mechanics as above (source)
Anthropic — Claude Haiku 4.5 $1.00 $5.00 Cheapest current Claude model on the standard rate card (source)
OpenAI — GPT-5.6 Sol $5.00 ($0.50 cached) $30.00 Flagship agentic model; cached-input rate applies to repeated prefix tokens (source)
OpenAI — GPT-5.6 Terra $2.00 ($0.20 cached) $12.00 Mid-tier balanced model (source)
OpenAI — GPT-5.6 Luna $0.20 ($0.02 cached) $1.20 Fastest, lowest-cost current model in the line-up (source)
Google — Gemini 2.5 Pro $1.25 (prompts ≤200K tokens) / $2.50 (>200K) $10.00 (≤200K) / $15.00 (>200K) Price steps up once a single prompt exceeds 200,000 tokens; a separate, more expensive Priority service tier is also available (source)
Google — Gemini 3.8 Flash $0.75 (promotional, through Dec 31, 2026) $3.75 (promotional, through Dec 31, 2026) Google’s own pricing page states these rates rise to $1.50 input / $7.50 output starting January 1, 2027 — a scheduled increase, not a hypothetical one (source)
Google — Gemini 3.5 Flash-Lite $0.30 $2.50 Positioned as Google’s cost-efficient high-volume option (source)

A few things are worth flagging explicitly rather than glossing over. OpenAI’s pricing page did not show a simple flat “GPT-5” line item at the time of this research; its current flagship line is branded GPT-5.6 with three variants (Sol, Terra, Luna) — if you’re budgeting from an older article or your own memory of GPT-5 pricing, re-check the live page, because the model names and numbers have moved. Likewise, Google’s Gemini 2.5 Pro remains the current Pro-tier model on the official pricing page even though Flash-tier models have iterated further (3.5 through 3.8) — the two product lines do not version in lockstep. Neither OpenAI’s rate-limits guide nor Google’s rate-limits page publishes a single static RPM/TPM number per tier the way Anthropic does; both explicitly tell you to check your account’s live dashboard, so any blog post (including this one) that prints a specific OpenAI or Gemini RPM figure without a dashboard screenshot is likely quoting an unofficial or outdated source.

What “Scaling Up” Practically Means for a Team

Moving from a free or entry tier to a higher one changes which constraint hits you first, and it’s rarely the one teams expect going in.

  • The first wall is usually RPM, not cost. A team doing bursty traffic — a product launch, a batch of agent runs kicked off at once — hits the requests-per-minute ceiling of their current tier long before they hit a monthly dollar cap. Anthropic’s Start tier ceiling of 1,000 RPM per model (source) is a real number to design retry/backoff logic around; OpenAI and Gemini accounts should check their own dashboard number before assuming headroom.
  • Output TPM is the quiet bottleneck for agentic or long-generation workloads. Output tokens are generated more slowly and capped more tightly than input tokens — Anthropic’s Start tier allows 2,000,000 input TPM but only 400,000 output TPM per model (source), a 5:1 ratio. A workload that generates long responses (code, reports, chain-of-thought reasoning) will exhaust its output allowance well before its input allowance.
  • Monthly spend caps become the binding constraint at the top of each tier. Anthropic enforces a hard monthly spend cap per tier and returns an enforced_spend_limit_reached error with no retry-after header once you hit it, pausing access until the first of the next month unless you move up a tier (source). OpenAI’s tiers carry the same shape of monthly organizational limit, from $100 at Free/Tier 1 up to $200,000 at Tier 5 (source).
  • Google adds a short-window spend brake most teams don’t plan for. Even inside your monthly budget, Gemini’s rolling 10-minute spend limit ($10 at Tier 1, $50 at Tier 2, $200 at Tier 3) can trip during a short, expensive burst — a handful of large Gemini 2.5 Pro calls in quick succession — well before the monthly cap is anywhere close (source).
  • Prompt caching changes the math differently per provider. On Anthropic, cached input tokens are excluded from the TPM limit entirely, so a high cache-hit-rate workload can process far more raw tokens per minute than the published number suggests (source). On OpenAI, cached input is simply billed at a steep discount (as low as $0.02 per million tokens on GPT-5.6 Luna) rather than exempted from a rate limit (source) — it lowers cost but not your RPM/TPM exposure the same way.

What to Monitor as Usage Grows

Three signals matter more than raw traffic volume once a team is past the free tier: the ratio of 429 (rate-limit) errors to total requests by endpoint and model, the gap between input and output TPM consumption (since output is almost always the tighter ceiling), and — for Gemini specifically — spend rate inside rolling short windows, not just monthly totals, since that is the limit Google enforces separately from RPM/TPM. Teams that only dashboard their monthly bill will miss the Gemini 10-minute spend cap and the Anthropic output-TPM ceiling until a production incident surfaces them.

These Figures Change Often — Verify Before You Plan Capacity

All three providers have changed either their tier thresholds, their rate-limit numbers, or their per-model prices within roughly the last year, and OpenAI’s and Google’s documentation explicitly point to live, per-account dashboards rather than a fixed public table for the exact RPM/TPM ceilings. Treat every number above as accurate as of early October 2026 sourced directly from each provider’s own page, and check the linked docs — Anthropic’s rate limits page, Anthropic’s pricing page, OpenAI’s rate-limits guide, OpenAI’s pricing page, Gemini’s rate-limits page, and Gemini’s pricing page — before finalizing a procurement decision or a capacity plan.

Frequently Asked Questions

Do rate limits reset instantly, or on a fixed schedule?

Anthropic’s documentation describes its enforcement as a continuously replenishing token-bucket model rather than a hard reset at the top of each minute, meaning short bursts beyond the per-minute number can succeed if the bucket has capacity, while sustained traffic above the limit will not (source). OpenAI and Google do not publicly document their internal enforcement mechanism in the same detail, so assume a standard rolling or fixed-window limiter unless your dashboard says otherwise.

Can you get custom limits above the highest published tier?

Yes, on all three. Anthropic’s top published tier (Scale) can be followed by a negotiated Custom tier through an account team (source). OpenAI points high-volume customers to a Scale Tier or Reserved Tier for dedicated capacity beyond Tier 5 (source). Google’s documentation does not detail an equivalent negotiated tier beyond Tier 3 in the pages reviewed here, so that would need to be confirmed directly with Google Cloud sales.

Does moving up a tier happen immediately after you hit the spend threshold?

For Anthropic and OpenAI, progression is described as automatic once the spend condition is met, with Anthropic noting that moving to a higher tier restores access right away if you were paused on a spend-limit error (source). Google’s Tier 2 and Tier 3 add an explicit waiting period — 3 days and 30 days respectively after your first successful payment — on top of the dollar amount, so reaching the spend number alone does not immediately unlock the next Gemini tier (source).

Capacity planning around these numbers usually touches a few related topics: prompt caching, since it changes the TPM math differently per provider as noted above; a model’s context window, which is a separate ceiling from its rate limit but interacts with it on large requests; and evaluating which tier actually fits your workload rather than defaulting to the flagship model for everything, especially for high-volume tool-calling workloads that burn through output tokens fast.

Image: “BalticServers data center” by BalticServers.com, licensed under CC BY-SA 3.0, via Wikimedia Commons.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top