Skip to main content

Rate Limits

Rate limits are enforced per API key, on a rolling 60-second window, at a requests-per-minute (RPM) ceiling that depends on your org's plan. Chat completions and transcription are tracked in separate 60-second windows — heavy transcription traffic doesn't eat into your chat completions RPM, and vice versa — but both use the same ceiling for your plan:

PlanRequests / minute (each endpoint)Included tokens (chat)Included minutes (transcription)
Free1020,000,000 lifetime, or 400,000/day once the lifetime pool is exhausted100 minutes, lifetime
Starter6015,000,000 / month60 minutes / month
Growth300100,000,000 / month360 minutes / month
Scale1,000Committed volume (contact sales)Committed volume (contact sales)

Exceeding the RPM ceiling returns:

HTTP/1.1 429 Too Many Requests
Retry-After: 12

{"detail": "Rate limit exceeded: max 10 requests per 60s"}

Retry-After is in seconds — wait at least that long before retrying. A reasonable client backs off with jitter rather than retrying at the exact instant Retry-After elapses, since every client that does that at once just recreates the burst.

Free tier quota​

Chat completions: a one-time 20,000,000-token pool, plus a 400,000-token daily allowance that kicks in once the lifetime pool runs out. A request is only blocked once both are exhausted:

HTTP/1.1 402 Payment Required
{"detail": "Free tier quota exhausted for today. Upgrade to continue, or try again after midnight UTC."}

Transcription: a flat 100-minute lifetime pool — no daily top-up. Once it's used, every request is blocked until you upgrade:

HTTP/1.1 402 Payment Required
{"detail": "Free tier transcription quota exhausted. Upgrade to continue transcribing."}

Paid plans don't hit either of these — overage beyond the included amount is billed (per-token for chat, per-minute for transcription) rather than blocked. These are tracked as separate line items, since they're different units. See the dashboard billing page for current overage rates per plan.

AudioHook streaming transcription draws from this same transcription-minutes pool rather than a separate one, and isn't subject to the RPM ceiling above — a session is one long-lived connection, not a burst of requests. Quota is checked when the connection opens and every 30 seconds thereafter; if it's exhausted mid-call, the connection is closed rather than continuing to accrue overage silently.

Design your client for 429s and 402s​

Both are ordinary, expected responses under load or heavy usage — not exceptional failures. Retry 429s with backoff; for 402 on the Free plan, surface an upgrade prompt rather than retrying, since retrying won't succeed until the quota resets.