Switching with fallback
Radium supports multiple models. This recipe shows how to route requests intelligently — picking the right model for the task, and automatically falling back if one is overloaded or unavailable.The script
smart_router.py
Run it
Sample output
Adding cost-based routing
Route based on estimated cost or latency requirements:Tips
- Use
tycho-1.0for high-volume, simple tasks — it’s fastest and cheapest. - Use
clarke-1.0for general reasoning, coding, and longer-context work. - Use
hal-1.0for the hardest reasoning and most nuanced outputs. - Log which model served each request so you can tune your routing rules over time.
- Monitor latency per model — measure in your own workload since latency varies by prompt length and time of day.
Next steps
Error handling & retries
Add exponential backoff and circuit breakers
Batch processing
Process thousands of requests with smart routing