Recommended Free Tools
AI provider errors differ because the same HTTP status can represent problems with different causes and remedies. A 429 may mean traffic needs to slow down, or it may mean an account has hit a spend or usage limit. If you are asking, “How should I handle AI API errors across providers?”, keep the HTTP status and provider details, then add a stable application-level category that tells your code what action is appropriate.
Why HTTP status codes are not enough
HTTP status is a useful first signal, but it does not fully diagnose an AI API failure. Providers attach their own error types, codes, messages, and recovery guidance to responses. Even within one provider, a status such as 429 can describe conditions that call for very different responses.
As an Amazon Associate I earn from qualifying purchases.
OpenAI, for example, documents 429 rate_limit_error responses with slow_down when traffic rises too quickly. It also uses 429 for usage or spend limits that require an account-level fix. Those are not interchangeable: reduce request pressure for the first; address the account limit for the second. OpenAI separately identifies model overload as a 503 with server_is_overloaded. See the OpenAI rate limits guide and OpenAI error codes guide.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsAs OpenAI notes, “A slow_down error can occur even when your traffic is within its requests-per-minute and tokens-per-minute limits.” In other words, published quotas do not make every traffic-related 429 predictable from your own counters.
#1 Best Overall
What should a normalized AI error contain?
Use a small internal error record that preserves the provider’s original evidence while adding a stable category for application policy. The category is your own classification, not a shared industry standard.
| Field | Purpose |
|---|---|
provider |
Identifies which API returned the failure. |
operation |
Records the application operation or endpoint being attempted. |
http_status |
Keeps the transport-level status code. |
provider_error_type and provider_error_code |
Retain the provider’s own classification and code when available. |
message |
Stores the provider’s message for diagnosis, subject to your data-handling rules. |
request_id |
Preserves a provider request identifier when returned, which can help with support and tracing. |
retry_after |
Captures a provider retry-timing instruction when present. |
attempt |
Tracks the attempt number across SDK and application retries where you can observe them. |
category |
Stores the normalized class your application uses to choose a remedy. |
Useful initial categories include invalid_request, authentication_or_permission, rate_limited, quota_or_billing, overloaded, transient_provider_failure, and unknown_provider_error. Keep the raw status, provider code or type, message, and request ID alongside the category. If you flatten everything to a number or replace the original error with your own label, you lose information needed for debugging and future reclassification.
Rank #2
How do provider errors and retry defaults differ?
Compare provider semantics and recovery signals, not just the status number. These documented examples illustrate why a shared policy needs provider-aware inputs.
| Provider | Documented distinctions | Retry behavior in official SDKs |
|---|---|---|
| OpenAI | 429 can identify traffic pressure with rate_limit_error / slow_down, or usage and spend limits; 503 can identify overload with server_is_overloaded. |
Official SDKs automatically retry eligible 429 and 503 responses. The rate-limit guide says to follow Retry-After when available; otherwise increase delay and add a small random delay. Billing, spend, or quota errors will not be fixed by retrying. |
| Anthropic | The Claude API documents 500 api_error and 529 overloaded_error, among other conditions. |
The official SDK retries transient failures, including connection errors, rate limits, and 5xx server errors, with exponential backoff; it retries twice by default and honors retry-after when present. |
| Google Gemini | The API error reference describes a structured error object for standard non-streaming requests and status categories including 400, 401, 429, and 503. | Official Gemini SDKs include default exponential-backoff retries for transient timeouts, network issues, and 429/5xx responses, according to Google’s troubleshooting guide. |
Sources: OpenAI rate limits, OpenAI error codes, Anthropic API errors, Gemini troubleshooting, and Gemini API errors reference. Retry details and SDK behavior can change; check the current provider documentation for the SDK version you deploy.
Rank #3
How should an application decide whether to retry?
Classify the failure first, then choose a remedy. A retry is appropriate only when the cause may clear with time or reduced pressure; it is not a general response to every non-success status.
- Correct request or access problems before trying again. Malformed requests and authentication or permission failures need a request or configuration change. Repeating the same call does not correct them.
- Separate account limits from traffic limits. For usage, spend, or quota conditions, surface an account-level action rather than retrying. For a traffic-related rate limit, reduce pressure and respect provider timing guidance.
- Retry transient failures with a limit. For network interruptions, overload, or eligible server failures, honor
Retry-Afterwhen present. Otherwise use bounded exponential backoff with jitter, and enforce an overall attempt count or time budget. - Account for retries already performed by the SDK. OpenAI, Anthropic, and Google document automatic retries in their official SDKs, but their eligibility and defaults differ. Set application-level limits with those retries in mind so layers do not multiply attempts unexpectedly.
- Return an actionable outcome. Tell a user or operator whether to correct input, check credentials, wait, reduce request rate, or address account configuration. Keep provider details in logs or diagnostics appropriate to your privacy and security requirements.
What should you avoid assuming?
- Do not treat every 429 as a pacing issue. It may indicate an account usage or spend limit that backoff cannot resolve.
- Do not assume retry metadata is identical across providers. The cited documentation describes retry timing in provider-specific terms; it does not establish one universal location or interpretation for every API response.
- Do not assume streaming failures behave like ordinary responses. The cited Google error reference describes standard non-streaming requests, and equivalent streaming semantics are not established here.
- Do not automatically fail over and replay a request at another provider. Replay safety, billing consequences, and semantic equivalence between models depend on the request and integration. Confirm those properties before implementing cross-provider fallback.
A practical error-handling policy
In code, make normalization a translation layer rather than a replacement for the original response. Parse what the provider exposes, retain its identifiers and details, map known conditions to your categories, and leave unrecognized cases as unknown_provider_error until you have enough information to classify them safely. Keep retry decisions separate from user-facing messages: the category can drive policy, while the preserved provider evidence supports debugging.
This structure makes provider differences manageable without pretending they have disappeared. The application gets consistent decisions; engineers retain the context needed to understand and revise those decisions as APIs and SDKs evolve.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




