Spend controls
Limit your spend at the request level (viamax_tokens) and handle oversized prompts gracefully before you ship to production.
Budget alerts, daily hard/soft limits, and auto-suspend are on the roadmap. For now the only spend guards are per-request
max_tokens and manual monitoring through the dashboard.Per-request token cap
Setmax_tokens on every request to prevent runaway generation from unbounded prompts or accidental infinite loops.
Rejecting oversized requests
Radium returns400 with code context_length_exceeded if the prompt exceeds the model’s context window. Handle this gracefully:
Usage audit
Export a CSV of recent requests from the dashboard for custom analysis:Pre-production checklist
Before shipping to users, verify:-
max_tokensis set on every request - Your app handles
context_length_exceedederrors - Usage is reviewed regularly in the dashboard
- Budget guardrails (alerts, auto-suspend) — coming soon on our roadmap