Register and share your invite link to earn from video plays and referrals.

Akshay 🚀
@akshay_pachaar
Simplifying LLMs, AI Agents, RAG, and Machine Learning for you! • Co-founder @dailydoseofds_• BITS Pilani • 3 Patents • ex-AI Engineer @ LightningAI
Joined July 2012
501 Following    290K Followers
Redis built a cache that cuts LLM costs by 70%! Production LLM apps often receive different versions of the same question. For instance, an internal developer assistant might receive: - "How do I rotate an API key?" - "Where can I replace my API key?" The wording is different, but both questions have the same answer. Prefix caching cannot handle this repeated generation because it only reuses computation when the cache matches bit-by-bit. And if the cache hits, an LLM response is still generated again. A semantic cache stores the complete question-response pair outside the model. When a query arrives, it searches for previously answered questions with similar meaning. A valid match returns the stored response with no LLM call. If you want to use this in practice, @Redisinc already implements it as a managed service called Redis LangCache. Under the hood, LangCache embeds the incoming question, searches stored responses, and applies the configured similarity threshold and filters. A cache hit returns the earlier response. A miss falls back to the LLM, after which the new response can be stored for future requests. Redis also handles access scopes, custom filters, embedding selection, TTL, eviction, and monitoring through a REST API, without another database to deploy or manage. While the actual savings depend on how much safe repetition exists in the workload, cache-hit responses are up to 15x faster and 70% cheaper. Try Redis LangCache: I built a small interface comparing LangCache with direct LLM inference. The video below shows this in action, and I worked with Redis on this post to put it together. To dive deeper, my co-founder published a detailed article on KV, prefix, prompt, and semantic caching. Read it below.
Show more