Tokens cost money. Today the pricing is decent and we absorb the redundancy without flinching. But as LLM usage scales — more users, more sessions, richer context — the cost curve turns real fast.
I hit this personally while building VegAi.in, a hobby app that generates Indian veg recipes from natural language prompts. Small scale, but the principle became immediately obvious: not every LLM call needs to be an LLM call. Some of them are just the same question wearing a different shirt.
That's the problem semantic caching solves. And for architects building LLM-powered products, it deserves a serious design conversation — not just as a cost trick, but as a pattern with real tradeoffs.
Why Traditional Caching Breaks Down for LLMs
Traditional caching is exact-match. You hash the input, look up the key, return the value. Works perfectly for deterministic APIs.
LLM inputs are not deterministic strings. Users express the same intent in dozens of ways:
"How do I reset my password?" "I forgot my login credentials" "I want to change my password, I can't get in"
Three different strings. Same intent. Same correct answer. A key-value cache misses all three after the first. You're paying the LLM for work it already did.
Semantic caching solves this by caching meaning, not strings.
How Semantic Caching Works — The Technical Flow
The mechanism has four steps:
1. Embed the query When a user query arrives, it is converted into a high-dimensional vector embedding (typically 768 or 1,536 dimensions) using an embedding model. This vector is a mathematical representation of the query's semantic meaning.
2. Similarity search The new vector is compared against previously cached query vectors using cosine similarity. This search happens in a vector database.
3. Cache hit or miss decision If the similarity score exceeds a configured threshold (typically 0.85–0.95), it's a cache hit — return the stored response immediately. No LLM call. If it falls below threshold, it's a cache miss — proceed to the LLM.
4. Store on miss The LLM response comes back. Both the query vector and the response are stored in the cache for future hits.
The vector similarity search costs 5–20ms. An LLM round trip costs 1–5 seconds. For cache hits, you're looking at 2–4x faster responses on average, with documented cases of 15x improvement on high-repetition query sets. Cost reduction on LLM API spend typically lands at 50%+ for applications with moderate query overlap.
The Threshold Problem — The Most Underestimated Design Decision
The similarity threshold is where most implementations get into trouble. It looks like a single config value. It's actually a business logic decision.
Too tight (e.g., 0.98): Almost no cache hits. You've added infrastructure overhead for near-zero benefit.
Too loose (e.g., 0.75): You're returning semantically approximate but contextually wrong answers. In a recipe context, "Jain recipe" and "vegetarian recipe" have high semantic overlap — but they are not the same. Jain cooking excludes root vegetables entirely. A wrong cache hit here is a product bug, not just an accuracy issue.
Practical calibration approach:
Start at 0.90–0.95
Monitor false positive rate for 2–4 weeks
If false positives exceed 3–5%, improve with domain-specific embedding models or tighten threshold
Use rejected cache response logging to identify where the model is confused
There is no universal correct threshold. It depends on your domain, the cost of a wrong answer, and the diversity of your user query patterns.
Context-Aware Semantic Caching — Where It Gets Architecturally Serious
Stateless caching works well when the query alone determines the response. But most production LLM applications are stateful — they carry user history, preferences, session context, and personalization signals.
The same query means different things for different users:
"Give me a quick breakfast suggestion"
For a diabetic user: low-GI, no refined carbs. For a fitness-focused user: high protein. For a Jain user: no root vegetables.
If you return a cached response that ignores this context, you've broken personalization entirely.
Design approaches for context-aware caching:
1. Context-augmented embeddings Embed the query together with a summarized user context vector. The combined embedding encodes both intent and context. Cache hits only occur when both are semantically similar. More accurate, but harder to tune — context representations need to be stable and consistently formatted.
2. Per-segment caching Group users into semantic segments (dietary preference, region, use-case profile) and maintain separate cache namespaces per segment. Simpler to reason about, but cache hit rates drop proportionally with segment granularity.
3. Layered cache architecture Maintain a global cache for context-independent queries (factual, structural questions) and per-user or per-session caches for personalized responses. The global cache captures the broad hit volume; the session cache handles personalized continuations.
4. Context as cache filter Embed the query globally, but apply context as a post-similarity filter — only return a cached result if the stored response was generated under a compatible context profile. More flexible, but requires metadata attached to every cached entry.
Each approach trades cache hit rate against accuracy and infrastructure complexity. There is no free lunch. The right design depends on how much personalization variance exists in your query space.
Cache Invalidation — Still the Hard Problem
Phil Karlton's observation holds: "There are only two hard things in Computer Science: cache invalidation and naming things."
For LLM response caches, invalidation strategy should be driven by how quickly the underlying source of truth changes:
Data TypeRecommended TTL StrategyRapidly changing (prices, inventory, live data)5–15 minute TTLsModerately dynamic (product descriptions, policies)1–4 hour TTLsStable content (recipes, FAQs, documentation)24-hour or longer TTLsLLM model upgrade / prompt changeFull cache flush
Beyond TTLs, implement content-triggered invalidation — when source data changes, proactively flush related cache entries rather than waiting for TTL expiry. This is especially important in RAG (Retrieval-Augmented Generation) architectures where the knowledge base is updated frequently.
One often-missed edge case: model upgrades. If you swap your LLM from GPT-4o to Claude 3.5 Sonnet, or even update your system prompt, old cached responses may no longer match the quality or style of new responses. Plan for full or partial cache invalidation on model/prompt changes.
Streaming Responses — A Practical Consideration
Many modern LLM UX patterns use streaming responses (tokens rendered as they arrive). Semantic caching needs to accommodate this:
Stream-then-cache: Stream the full response to the user in real time. Once complete, cache the full response. Subsequent identical queries return the full cached string instantly. Works well; the first user pays full latency, all subsequent users get instant responses.
Early-exit caching: Pre-generate and cache responses for high-probability queries before users ask. When a query matches, the response is already waiting. Better UX, higher operational complexity, requires query prediction logic.
For most implementations, stream-then-cache is the right starting point.
Infrastructure — Redis as the Pragmatic Choice
Semantic caching requires three components: an embedding model, a vector store, and a cache layer. Running these as separate systems (e.g., OpenAI embeddings + Pinecone + Redis) works but adds operational surface area.
Redis with vector search capability collapses all three into one: sub-millisecond vector search, built-in TTL-based cache eviction, and support for complex data structures in a single system. For most applications — especially those already using Redis for session or app caching — this is the lowest-friction path.
For very large-scale deployments with billions of cached vectors, purpose-built vector databases (Pinecone, Weaviate, Qdrant) may be warranted. But for the majority of LLM application scales, Redis handles it cleanly.
What This Pattern Is Not
Semantic caching is not a replacement for good prompt engineering. It does not reduce the cost of genuinely novel queries — and in a rich, contextual AI product, novel queries are often the majority.
It is also not a substitute for output quality controls. A cache hit that returns a fast but wrong answer is worse than a cache miss that returns a slow but correct one. Monitor hit quality, not just hit rate.
The goal is to avoid paying the LLM tax on work that has already been done, while preserving full LLM capability for everything that is genuinely new.
The Design Principle
LLM tokens are expensive fuel. Semantic caching is the fuel efficiency layer. Like any optimization, it works best when applied with clear understanding of the tradeoffs — threshold tuning, context awareness, cache invalidation, and the inherent tension between hit rate and accuracy.
For architects and engineers building production LLM systems: this is not optional thinking at scale. Design for it early. Retrofitting a caching layer into a stateful, context-rich LLM application is significantly harder than building it in from the start.
Manoj writes The Practical Technologist — a newsletter for builders who think clearly about technology and build things that work.
First published on LinkedIn.
