Skip to main content

Caching responses

Many prompts repeat — FAQs, classification labels, standard summaries. Caching identical or similar requests can cut costs by 30-70% and eliminate latency for cache hits.

The script

cache.py

Run it

Sample output

The semantic-caching example below calls client.embeddings.create() through the Radium API. The OpenAI compatibility guide states that POST /v1/embeddings is currently not supported. Retarget to a third-party embedding provider before using this pattern in production.

Semantic caching with embeddings

Exact-match caching misses when phrasing changes. Use embeddings to find similar past queries:
semantic_cache.py

Cache invalidation strategies

Tips

  • Use exact-match caching for deterministic tasks like classification, extraction, and FAQ answering.
  • Use semantic caching for open-ended Q&A where users rephrase the same question.
  • Never cache requests with temperature > 0 unless you want identical outputs — stochasticity defeats the cache.
  • Monitor hit rate — a 50%+ hit rate usually means caching is worth the complexity.

Next steps

Monitoring usage

Track cache hit rates and cost savings

Batch processing

Process large datasets with caching