Skip to main content

Benchmarks

Radium models are evaluated on standard academic and real-world benchmarks. This page shows how they compare to incumbent models on tasks that matter to production workloads.

Academic benchmarks

Reasoning and coding

Multilingual and long-context

Real-world task performance

We also evaluate on production-like tasks that aren’t captured by academic benchmarks:

Side-by-side examples

Example 1: Reasoning task

Prompt: “A train travels 60 mph for 2 hours, then 40 mph for 3 hours. What is the average speed for the entire trip?“ All models answer correctly on this prompt.

Example 2: Tool-calling task

Prompt: “What’s the weather in Tokyo? Use the get_weather tool.”

How we test

  • Temperature: 0 for deterministic tasks, 0.7 for open-ended tasks
  • Max tokens: 99999 unless the benchmark specifies otherwise
  • Seed: 9999 where reproducibility is required
  • Date tested: 2026-01-01

Reproducing the results

Run the suite against your own key:

Limitations

  • Benchmark scores vary by prompt phrasing and evaluation method
  • Real-world performance depends on your specific use case
  • We recommend running your own evals on your own data before making a decision