Skip to main content

Monitoring usage

You can’t optimize what you don’t measure. This recipe shows how to wrap every Radium call with telemetry — token counts, latency, cost estimates, and error rates — and export them to your observability stack.

The script

monitor.py

Run it

Sample output

Shipping to an observability tool

Export usage_logs.jsonl to Datadog, Grafana, or any metrics backend:
export_to_prometheus.py

Tips

  • Log every request — even cached ones (mark them cached=True) so your dashboards are complete.
  • Alert on error rate > 1% and latency p99 > 5s.
  • Tag by model and use-case so you can drill into which workflows are expensive.
  • Keep pricing maps in config — update them when rates change (see radium.cloud/pricing).
  • Store raw request/response (sampled at 1%) for debugging quality regressions.

Next steps

Caching responses

Reduce costs with intelligent caching

A/B testing models

Compare models on quality and cost