GET /v1/models and POST /v1/chat/completions uses the same JSON envelope. The error event inside a stream has no param — see Streaming. Branch on error.code, never on error.message.
A path the API does not serve returns a bare status with an empty body, before the key is checked: 404 for an unknown path under /v1 (for example /v1/embeddings or /v1/models/{id}), and 405 for a method the path does not accept. Check the status before you parse the body.
The error envelope
string
The identifier your client branches on. Every code is in the table below.
string
The broad class:
invalid_request_error, insufficient_quota, permission_error, rate_limit_error or server_error.string
A description for a person to read. It can change; do not parse it.
string | null
The request field the error refers to, or
null.string
The same value as the
x-request-id header, which every response carries. Include it in a support request.Codes
GET /v1/models returns only 401, 403 and 503 codes from this table.
After a stream has started, every failure arrives as the last event with the code upstream_error — see Streaming.
What to do about each code is in Common problems.
Retries
A code marked Yes always comes with aRetry-After header, in seconds. Wait at least that long, then send the request again. rate_limit_exceeded can ask for up to 3,600 seconds; the others ask for 60 seconds or less. A refused request is not queued: nothing runs unless you send it again.
A code marked No is not cleared by retrying. Change the request, the account or the limit first — see Common problems.
502 and 504 happen after the request started, so part of it may have run and be charged — see What is billed. A retry is a new request and is charged separately.
Every other error is returned before generation starts and is not charged.
There is no idempotency key, and a response that broke off is not resumed.