Benchmarks
Radium models are evaluated on standard academic and real-world benchmarks. This page shows how they compare to incumbent models on tasks that matter to production workloads.Academic benchmarks
Reasoning and coding
Multilingual and long-context
Real-world task performance
We also evaluate on production-like tasks that aren’t captured by academic benchmarks:Side-by-side examples
Example 1: Reasoning task
Prompt: “A train travels 60 mph for 2 hours, then 40 mph for 3 hours. What is the average speed for the entire trip?“
All models answer correctly on this prompt.
Example 2: Tool-calling task
Prompt: “What’s the weather in Tokyo? Use the get_weather tool.”How we test
- Temperature: 0 for deterministic tasks, 0.7 for open-ended tasks
- Max tokens: 99999 unless the benchmark specifies otherwise
- Seed: 9999 where reproducibility is required
- Date tested: 2026-01-01
Reproducing the results
Run the suite against your own key:Limitations
- Benchmark scores vary by prompt phrasing and evaluation method
- Real-world performance depends on your specific use case
- We recommend running your own evals on your own data before making a decision