Skip to main content

Error codes

Errors use the OpenAI shape, plus the gateway's request id:

{
"error": {
"message": "RPM limit exceeded for key (limit 60/min)",
"type": "rate_limit_error",
"code": "rate_limit_exceeded",
"request_id": "req_6f1c0d9a2b7e4c3f8a5d1e0b"
}
}

type mapping: 401 → authentication_error, 403 → permission_error, 429 → rate_limit_error, ≥ 500 → api_error, else invalid_request_error.

CodeHTTPWhen
missing_api_key401No / empty bearer token
invalid_api_key401Unknown key hash
key_revoked401Revoked
key_expired401Expired
key_inactive401Non-active status
key_blocked403Blocked key
model_access_denied403Model not in the key allowlist
scope_blocked403User, team, org or app is blocked
customer_blocked403Customer is blocked
guardrail_blocked403Pre or post guardrail block
budget_exceeded429Key, scope, tag, customer or first-class (Budgets page) budget exhausted
budget_approval_required429A first-class budget reached a require-approval threshold or limit
rate_limit_exceeded429RPM or TPM window exhausted
invalid_request400Bad JSON or missing model
model_not_found404The model (and its fallbacks) has no deployments configured
upstream_unavailable502Every deployment attempt failed, or every deployment of the model and its fallbacks is disabled
internal_error500Key context / DB failure
unhealthy503Health / readiness probe failure

Rate-limit messages include the scope: RPM limit exceeded for key (limit N/min), TPM limit exceeded for team (limit N tokens/min). Scopes: key, team, application, customer (chat only).

Budget messages include the entity and spend vs max, for example Budget has been exceeded! Key=<alias> Current cost: X, Max budget: Y.

Request IDs​

Every gateway response — success or error, including 401/403/404/429/5xx and streaming responses — carries the request id in two headers, set before the first byte is written:

HeaderValue
x-zeallm-request-idreq_ followed by 24 hex characters
X-Request-IdThe same value. OpenAI SDKs read this header, e.g. err.request_id on an APIStatusError in Python
x-client-request-idOnly when the request sent its own X-Request-Id: that value echoed back (up to 128 visible ASCII characters)

The gateway always generates its own id; a client-supplied X-Request-Id is never used as the request id. Error bodies repeat the id as error.request_id, so it is available to clients that only log bodies.

Finding a request in the portal​

Open Logs and paste the id into the Request ID filter, or deep-link to /logs?requestId=req_… (see Request Logs). The id is the requestId stored on the request's log entry; logs are written asynchronously, so a request may take a moment to appear.

Chat completions, embeddings and responses calls are logged whether they succeed or fail, including rejections for authentication, model access, budgets, rate limits, guardrails, model_not_found and upstream_unavailable. Requests rejected as malformed (400, e.g. invalid_request) and internal errors before routing are not logged; quote the request id and time to support anyway. /v1/models, cost-event and tokenomics-observation calls carry an id but have no log entry.

Retry-After​

Some errors include a Retry-After header: a whole number of seconds (rounded up, at least 1) after which retrying can succeed. OpenAI SDKs honour it when they retry. It is omitted when the gateway does not know when the condition will clear.

ErrorRetry-After
429 rate_limit_exceededSeconds until the current RPM/TPM window rolls over (1–60). Windows are fixed one-minute windows aligned to the clock minute, shared across gateway replicas when Redis is configured. Omitted when the request's own prompt is larger than the TPM limit: no window will ever admit it, so shorten the prompt or ask for a higher limit.
429 budget_exceeded, budget_approval_requiredSeconds until the budget named in the message resets: for key, user, team, org, app, tag and customer budgets, the next reset of a budget with a duration; for first-class budgets, the next period reset or the budget's end date, whichever is first. Omitted for budgets that never reset (for example a max budget without a duration). If the reset time has passed but the reset job (every minute) has not yet applied it, 60.
502 upstream_unavailableOnly when providers answered 429: the gateway does not pass upstream 429s through, it retries other deployments and fallbacks first. If it runs out, it forwards the shortest upstream Retry-After (or retry-after-ms) it received.

An approved budget change request can lift a budget block before Retry-After elapses.