Error handling & retries
Production workloads hit rate limits, network hiccups, and transient errors. This recipe shows how to wrap your Radium calls with robust retry logic and automatic fallback to a backup model.The script
resilient_chat.py
Run it
What it does
- Retries
RateLimitErrorwith exponential backoff + jitter — the standard recipe for 429 responses. - Retries
APIConnectionErrorwith shorter backoff for transient network issues. - Retries 5xx errors from the API itself, but gives up on unrecoverable 4xx client errors.
- Falls back to
tycho-1.0ifclarke-1.0is overloaded or unavailable. - Fails loudly only after both models are exhausted — no silent swallowing.
Adding circuit-breaker logic
For production services, add a simple circuit breaker that stops hammering a struggling endpoint:Tips
- Start with
max_retries=3and tune based on your traffic patterns. - Monitor
Retry-Afterheaders from rate-limit responses — the example uses fixed backoff; production code should read the header. - Log every retry with
model_nameandattemptso you can spot recurring patterns in your observability tool. - Reserve your fallback model for a different tier (e.g. switch from
clarke-1.0totycho-1.0) so you don’t hit the same capacity constraint.
Next steps
Switching with fallback
Route requests intelligently across models
Tool calling agent
Build an agent that recovers from tool failures