Migration testing
Move from your incumbent to Radium in stages. Start with a small canary, measure output quality and cost, then scale up.Overview
Phase 0: Pre-flight validation
Before routing any live traffic, run automated checks against both providers with identical prompts.comparison-results.json for:
- Output correctness on structured tasks
- Reasoning quality on multi-step tasks
- Token efficiency (lower is cheaper)
Phase 1: 5% canary with LiteLLM
LiteLLM is the simplest way to route a percentage of traffic to Radium without changing application code.LiteLLM config
Application side
Your app still calls the same proxy endpoint. LiteLLM handles the split:Monitoring
Track these metrics per provider during the canary:Phase 2: 25% expansion
Gradually increase the Radium ratio in LiteLLM or via your own weighted router:- Streaming completions
- Tool-calling workflows
- Structured-output pipelines
- Multi-turn conversations
Phase 3: Full cutover
When Radium has proven reliable, switch the default to 100%. Keep the incumbent route live as a hot rollback:Rollback procedure
- Change the default model alias back to the incumbent
- Verify the incumbent is responding within 30 seconds
- Root-cause the Radium failure (check status page)
- Re-enable Radium once the issue is resolved
Side-by-side cost worksheet
Fill in real numbers from your dashboard and the Radium dashboard. Radium is typically cheaper per token; the worksheet proves the total cost for your specific workload.