
Semantic caching cuts latency and reduces inference costs by reusing answers for semantically similar prompts. In this talk, Chaitanya Nuthalapati will explore how to build a production-grade semantic caching system for multi-agent systems with Valkey and Strands. Beyond the basics, this talk focuses on techniques for improving cache accuracy, including handling multi-turn interactions, applying conversation-state filters, protecting personally identifiable information (PII), and navigating the trade-offs of personalization.