Caching responses
Many prompts repeat — FAQs, classification labels, standard summaries. Caching identical or similar requests can cut costs by 30-70% and eliminate latency for cache hits.The script
cache.py
Run it
Sample output
Semantic caching with embeddings
Exact-match caching misses when phrasing changes. Use embeddings to find similar past queries:semantic_cache.py
Cache invalidation strategies
Tips
- Use exact-match caching for deterministic tasks like classification, extraction, and FAQ answering.
- Use semantic caching for open-ended Q&A where users rephrase the same question.
- Never cache requests with
temperature > 0unless you want identical outputs — stochasticity defeats the cache. - Monitor hit rate — a 50%+ hit rate usually means caching is worth the complexity.
Next steps
Monitoring usage
Track cache hit rates and cost savings
Batch processing
Process large datasets with caching