Skip to main content

Switching with fallback

Radium supports multiple models. This recipe shows how to route requests intelligently — picking the right model for the task, and automatically falling back if one is overloaded or unavailable.

The script

smart_router.py

Run it

Sample output

Adding cost-based routing

Route based on estimated cost or latency requirements:

Tips

  • Use tycho-1.0 for high-volume, simple tasks — it’s fastest and cheapest.
  • Use clarke-1.0 for general reasoning, coding, and longer-context work.
  • Use hal-1.0 for the hardest reasoning and most nuanced outputs.
  • Log which model served each request so you can tune your routing rules over time.
  • Monitor latency per model — measure in your own workload since latency varies by prompt length and time of day.

Next steps

Error handling & retries

Add exponential backoff and circuit breakers

Batch processing

Process thousands of requests with smart routing