Skip to main content

Performance

Radium is optimized for low latency and high throughput. Understand what to expect and how to measure it.

Time to first token (TTFT)

TTFT is the time from sending a request to receiving the first streamed token. Radium measures TTFT internally and will publish live P50/P99 numbers on the Radium status page in a future release. Longer prompts increase TTFT linearly. Cache hits (repeated prompts) reduce TTFT by approximately 50%.

Measuring latency in your code

Temperature-0 determinism

At temperature=0 and seed fixed, Radium produces deterministic output for a given model version. Minor variance (< 1% token difference) may occur across model hot-swaps or infrastructure updates. For maximum reproducibility, pin the model version in your request headers and pass a fixed seed.

Streaming performance

Always use streaming for latency-sensitive applications. It reduces perceived TTFT to near-zero and allows partial processing of long outputs.

Batch processing

For offline workloads, use batch to amortize overhead:

Performance tuning checklist

  • Use streaming for interactive UIs
  • Set temperature=0 for deterministic tasks
  • Use hal-1.0 for fastest TTFT
  • Use tycho-1.0 for lower cost on simpler tasks
  • Enable prompt caching for repeated queries
  • Batch offline workloads
  • Keep prompts within each model’s context window (Tycho 125K, Hal 250K, Clarke 1M)