RPM
rpm_limit is an atomic rolling request limit. A denied request returns HTTP 429 with rate_limit_exceeded and a Retry-After header.
TPM
tpm_limit reserves estimated input plus requested output tokens before the provider call. The estimator is intentionally conservative and is not an exact tokenizer.
For chat and Responses, a TPM-limited request must set an output limit:
max_completion_tokensormax_tokensfor chat;max_output_tokensfor Responses.
policy.tpm_default_output_tokens.
Organization limits
An organization can carry aggregate ceilings applied across all of its keys:rpm_limit— requests/minute for the whole org. Over the limit returns HTTP 429org_rate_limit_exceededwithRetry-After.tpm_limit— estimated tokens/minute for the whole org (same estimator as the per-key TPM). Over the limit returns HTTP 429org_rate_limit_exceeded.max_concurrency— simultaneous in-flight requests for the whole org (see Concurrency).
/admin/organizations/{org_id}/limits by a system_admin only;
tenant roles cannot raise their own ceilings. Like per-key limits, organization rate limits
require Redis shared state unless the deployment acknowledges single-process
memory enforcement.
Concurrency
max_concurrency bounds the number of requests an organization can have in
flight at once. The slot is taken before the provider call and held until the
response fully completes — including the drain of a streamed response — then
released on every terminal path (success, provider error, client disconnect), so
it cannot leak. A request that arrives while the organization is at its cap
returns HTTP 429 concurrency_limit_exceeded with Retry-After; the reservation
it would have made is rolled back. Unlike RPM (a rolling count over a minute),
this is a point-in-time ceiling on overlapping requests.
Distributed enforcement
Memory limits apply to one process. Redis-backed limits are required when more than one gateway process may serve the same key or organization. By default, configuration and Admin API mutations reject rate/concurrency limits in memory mode. A deliberate single-process local deployment must explicitly setshared_state.allow_process_local_throttling: true; persisted limited keys or organizations otherwise block startup. Redis outages fail readiness, limited requests fail closed with HTTP 503, and lost concurrency leases terminate affected requests rather than bypassing policy.
Other controls
Self-managed deployments can configure:- maximum request body bytes;
- audio upload bytes/duration reservation bounds, TTS character limits, upload/provider timeouts, and per-key bounded-operation concurrency held through stream cleanup;
- maximum estimated input tokens;
- blocked client parameters;
- model-specific unsupported parameters;
- model and endpoint scopes;
- tools, vision, reasoning, embeddings, audio transcription, and text-to-speech capabilities.
Retry-After, apply jitter, and cap attempts.