← All posts
LLMReliabilityEngineering

Rate limits will find you: backpressure for LLM APIs

429s aren't an edge case at scale, they're a Tuesday. Here's how to design queueing, concurrency caps, and retries so a rate limit degrades gracefully instead of cascading.

Most teams discover their rate limit strategy the day it fails: a marketing campaign drives traffic, a batch job kicks off at the same time as peak usage, and suddenly every request is getting 429s. The naive fix — retry immediately — turns a momentary limit into an outage, because every failed request just adds another retry on top of an already-saturated queue. Rate limits aren’t a bug in the provider’s API. They’re a constraint you have to design for, the same way you design for disk space or database connections.

The failure mode: retry storms

A retry storm looks like this:

  1. Traffic exceeds your requests-per-minute (RPM) or tokens-per-minute (TPM) limit.
  2. The provider returns 429.
  3. Your code retries immediately, adding to the same second’s request volume.
  4. The retry also gets rate-limited, so it retries again.
  5. Load compounds until every client is retrying, and the limit is nowhere near clearing.

This is the same shape as a database connection pool getting exhausted, and it needs the same fix: don’t let failures generate more load than the requests that caused them.

Three layers that actually work

1. A concurrency cap in front of the provider, not just a retry policy

Before you touch retries, cap how many in-flight requests you allow to a given model/provider at once. A semaphore sized to roughly 70-80% of your known RPM/TPM limit keeps you under the ceiling most of the time, so you’re handling rate limits as an exception, not a steady-state condition.

const limit = pLimit(20) // max 20 concurrent calls to this model
const results = await Promise.all(
  requests.map(req => limit(() => callModel(req)))
)

This alone eliminates most rate-limit errors, because you stop sending more concurrent work than the provider will accept.

2. Exponential backoff with jitter, and respect Retry-After

When you do get a 429, back off — and randomize the wait so a batch of clients doesn’t resync and retry in lockstep:

async function callWithBackoff(fn, attempt = 0) {
  try {
    return await fn()
  } catch (err) {
    if (err.status !== 429 || attempt >= 5) throw err
    const retryAfter = err.headers?.['retry-after']
    const base = retryAfter ? Number(retryAfter) * 1000 : 500 * 2 ** attempt
    const delay = base + Math.random() * base * 0.25
    await sleep(delay)
    return callWithBackoff(fn, attempt + 1)
  }
}

Most providers (OpenAI, Anthropic) return a retry-after header or equivalent — use it instead of guessing. A capped attempt count matters too: infinite retries just move the outage later and make it invisible until a queue backs up.

3. A queue with priority, not a flat FIFO

Not all requests are equal. An interactive chat response blocking a user is not the same as a nightly embeddings backfill. When you’re near the limit, a flat queue lets the backfill starve the user-facing request. Two queues — interactive and batch — with the interactive queue always drained first, keeps degradation proportional to what users actually notice.

flowchart LR
    A[Interactive request] --> C{Concurrency gate}
    B[Batch/background request] --> C
    C -->|priority: interactive first| D[Provider API]
    D -->|429| E[Backoff + jitter]
    E --> C
    D -->|200| F[Response]

What to actually monitor

Rate limiting is invisible until it isn’t, so track it before you need to:

The part teams skip: capacity planning per provider, not per app

If you call multiple models or providers, your rate limits are per-provider-account, not per-feature. A new feature that adds calls to the same underlying account eats into the same TPM budget as everything else already running. Track headroom at the account level and treat “add a new LLM-backed feature” the same way you’d treat “add a new heavy database query” — check the budget before you ship, not after the first traffic spike.

Rate limits are a resource constraint like any other. Teams that treat them as an occasional error to retry around get outages. Teams that build a concurrency gate, real backoff, and priority queueing get a system that slows down gracefully under load and comes back on its own — which is the actual bar for production reliability.

Want something like this built for your team?

Get a quote →