October 7, 2026 · 19 min read
Caching a RAG App with Redis: What to Cache, When It Goes Stale, and What Breaks
A plain-English guide to putting Redis in front of a RAG app. Learn cache-aside and TTLs, cache the three slow steps of a RAG request, watch a cached answer go stale when a help page changes, and see what happens when Redis is down or full. Every number is measured.
Most RAG apps answer the same questions again and again, and pay for the model every time. A cache fixes that in an afternoon. It also brings new ways to be wrong: answers that go stale, a "similar question" that isn't, and a dead cache that makes the app slower than having no cache at all. This guide builds the cache step by step on a small help centre, and every one of those failures is shown with real output.
You don't need to know Redis. Each idea is explained the first time it appears, and every section ends with a one-line Remember. There's a ten-line summary at the end.
The running example is Parcelo, the same invented online shop from the RAG debugging guide. Its help centre has 12 pages, and customers ask questions like "Can I get my money back for the yearly Plus membership?"
The guide has four parts. Jump to the one you need:
- Part 1 · Understand it: why RAG requests are slow, what Redis is, cache-aside and TTLs, sections 1–4.
- Part 2 · Cache a RAG app: three cache layers and what they saved, sections 5–7.
- Part 3 · When the cache lies: stale answers and "similar" questions, sections 8–10.
- Part 4 · Run it for real: Redis down, Redis full, Azure Managed Redis, sections 11–14.
Part 1
Understand it
What makes a RAG request slow, and the one pattern that fixes it.
1. A RAG request pays three times
In the RAG guide we split a RAG app into a clerk and a writer. The clerk finds the right pages. The writer answers from them. Every request makes them do three jobs:
- Embed the question: turn it into a list of numbers (a vector) so it can be compared with the pages.
- Retrieve: find the 5 chunks of text closest to that vector.
- Write: send those chunks and the question to a language model, which writes the answer.
When I ran these on my laptop, embedding took about 49 ms and retrieval about 4 ms. The model is the slow and expensive part. A hosted model usually takes a second or more and is billed per token. My lab uses a stand-in that waits a fixed 1.2 seconds, so the totals below are realistic in shape, not a benchmark of any model.
Now notice what a help centre looks like from the inside. The same few questions arrive all day: refunds, delivery times, lost parcels. Each one goes through all three jobs, and gets the same answer as last time.
2. Redis is a fast notebook
Imagine the clerk keeps a notebook on the desk. Before going to the archive, they check the notebook: "Did someone ask this already?" If yes, they read out the answer. If not, they do the work and write the answer down for next time.
Redis is that notebook. It's a database that keeps everything in memory (RAM) instead of on disk, so reading a value takes well under a millisecond. You store a value under a name, called a key, and read it back by the same name:
> SET greeting "hello"
OK
> GET greeting
"hello"
> GET something-else
(nil)
(nil) means "nothing stored under that key". In Python, with the redis package, the same thing reads r.set("greeting", "hello") and r.get("greeting"), and a missing key comes back as None.
Redis can hold more than plain text (hashes, lists, sets, streams), but for caching, a string per key covers almost everything.
3. Cache-aside: look first, then work, then write it down
The notebook routine has a name: cache-aside. Your app, not Redis, does all of it:
- Look in Redis for the key.
- Hit: it's there. Use it.
- Miss: it's not. Do the slow work, then store the result in Redis.
Watch two identical questions go through it:
7. Answered in 0.6 ms. The model never ran.
In code it's five lines around the slow part:
cached = r.get(key)
if cached is not None: # hit
return json.loads(cached)
answer = do_the_slow_work(question) # miss: embed, retrieve, write
r.set(key, json.dumps(answer), ex=3600)
return answer
Redis never fetches anything by itself. If the app forgets to store the result, or stores the wrong thing, the cache stays empty or wrong. That's why it's called cache-aside: the cache sits next to the real work, and the app keeps the two in step.
4. Every copy gets an expiry date
A copy in the notebook can go out of date. Parcelo might change its refund policy tomorrow. The simplest protection is to write an expiry date on every page of the notebook.
In Redis that's a TTL (time to live): a countdown attached to a key. When it reaches zero, Redis deletes the key by itself. You set it when you write, with ex= (seconds):
Three TTL answers to remember: a positive number is seconds left, -1 means the key exists but never expires, -2 means there is no such key.
The TTL is a promise about how wrong you're willing to be. A one-hour TTL means that, without any other help, an answer can be up to an hour out of date. Section 8 shows how to do better than waiting.
The key name matters as much as the expiry. In the lab, every key has three parts:
All three normalise to the same text, so all three hit the same answer (0.5–0.6 ms each when I ran it). “Is the yearly Plus fee refundable?” means the same thing but is a different key: a miss.
The hash is a short fingerprint of the question. A fingerprint keeps keys short, and the same question always produces the same fingerprint. Normalising the question first (lower case, single spaces, no trailing "?") lets small differences in typing share one entry.
Part 2
Cache a RAG app
Three cache layers on the Parcelo help centre, measured.
5. Three things worth caching
Each of the three jobs from section 1 can have its own cache. They are worth different amounts and go stale in different ways:
| Layer | Key | Stores | TTL in the lab | Goes stale when |
|---|---|---|---|---|
| Answer cache | ans:v1:<hash> | the final answer and its source page | 1 hour | any page it used changes |
| Retrieval cache | ret:v1:<hash>:k5 | the top 5 chunks | 1 hour | any of those pages changes |
| Embedding cache | emb:minilm:<hash> | the question's vector (1,536 bytes) | 24 hours | only if you change the embedding model |
The answer cache is checked first. A hit skips everything. On a miss, the two lower layers still save their own steps.
The embedding cache is the safest of the three. A question's vector depends only on the question and the model, so it can't go stale when pages change. That's why its key names the model (minilm) instead of the document version.
6. The lab
Here's the setup. Everything was run on 7 October 2026.
- Redis 7.4 in Docker (
redis:7.4), and PostgreSQL 17 with pgvector in Docker for the chunks. - Python 3.14 with
redis8.0.1 (the client),sentence-transformers6.1.0 and theall-MiniLM-L6-v2embedding model. - The 12-page Parcelo help centre, cut at headings with the page title added to each chunk (the fix from the RAG guide).
- No language model. The "write" step picks the retrieved sentence that best matches the question, then waits a fixed 1.2 seconds to stand in for a model call. The 1.2 seconds is my choice, not a measurement. Everything else is measured.
The heart of it is one function per layer, each following the routine from section 3. Here is the embedding layer:
def embed(cache, q):
key = f"emb:minilm:{h(normalise(q))}"
if (hit := cache.get(key)) is not None:
return np.frombuffer(hit, dtype=np.float32), "hit"
vec = model.encode([q], normalize_embeddings=True)[0].astype(np.float32)
cache.set(key, TTL_EMBEDDING, vec.tobytes()) # 384 floats = 1,536 bytes
return vec, "miss"
The retrieval and answer layers look the same, with one addition: they put the corpus version in the key, and they record which pages they were built from. Section 9 explains why.
To run it yourself:
./run.sh # Redis on 6390, Postgres on 5545, venv
./.venv/bin/python cached_rag.py layers # also: keys, stale, semantic, down
7. What the cache saved
I asked all 10 test questions three times: once with nothing cached, once with only the lower two layers warm, and once with the answer cache on.
median per request cold (all miss) warm, no answer cache
embed the question 49.1 ms 0.8 ms
retrieve top 5 chunks 4.4 ms 1.4 ms
whole request 1262.4 ms 1209.0 ms
warm, answer cache on: whole request 0.6 ms
Two things stand out.
First, the lower layers barely move the total. Saving 48 ms on embedding is real, but next to a model call it's 4% of the request. Those layers matter on misses: a question that's new as an answer can still reuse its vector and chunks.
Second, the answer cache removes the model call completely. 1,262 ms became 0.6 ms. If 40% of your questions are repeats, 40% of your model bill and latency goes away.
Here's what the notebook held after one question, from the keys demo:
ans:v1:9a66ab4fb01269b7 string 232 bytes TTL 3600 s
corpus:version string 56 bytes TTL -1 s
deps:membership set 128 bytes TTL 86400 s
deps:returns set 120 bytes TTL 86400 s
emb:minilm:9a66ab4fb01269b7 string 1864 bytes TTL 86399 s
ret:v1:9a66ab4fb01269b7:k5 string 1608 bytes TTL 3599 s
The two deps: sets are the record of which pages this answer was built from. You'll need them in a moment.
Part 3
When the cache lies
Stale answers, and why "similar" isn't "same".
8. A help page changes
Parcelo changes its policy: Plus membership fees are now refundable for 30 days instead of 14. Someone edits the returns page, and the app re-embeds it. The database is correct.
The cache isn't. Step through what happened in my run:
5. Or: bump the version
INCR corpus:version changes every key name from v1 to v2. Nothing old can be found any more. The 2 old keys wait in memory until their TTL ends.
User sees: “… refundable only within 30 days …” ✓ correct
Step 3 is the trap. Deleting the cached answer feels like enough. But the answer was rebuilt from the retrieval cache, which still held the old chunks, so the "fresh" answer said 14 days again. This is the real output:
2) no invalidation answer hit -> … refundable only within 14 days …
3) delete the answer key only answer miss retrieval hit -> … refundable only within 14 days …
4) delete every key built from 'returns' (2 keys) answer miss retrieval miss -> … within 30 days …
A TTL would have fixed it eventually, after up to an hour of wrong answers. For a refund policy, an hour of wrong answers means real support tickets.
9. Two ways to invalidate
Invalidation means removing copies you know are wrong, instead of waiting for them to expire. The lab shows two ways.
Precise: delete what depended on the page. Each time the app stores an answer or a retrieval result, it adds that key to a set named after every page it used:
for doc in {c["doc"] for c in chunks}:
r.sadd(f"deps:{doc}", key) # deps:returns -> {ans:v1:…, ret:v1:…}
When the returns page changes, read deps:returns and delete every key in it. Only answers that used that page are dropped, and everything else stays cached.
Blunt: change every key name at once. The corpus version is part of every retrieval and answer key. After any page changes, increase it:
> INCR corpus:version
(integer) 2
Now the app builds keys starting ans:v2: and ret:v2:. Nothing from v1 can be found any more, so every request rebuilds from the database. The old keys aren't deleted. They sit in memory until their TTL ends (2 keys in my run, thousands in a real one).
| Delete dependants | Bump the version | |
|---|---|---|
| What's dropped | only answers that used the changed page | everything except embeddings |
| Extra work | a set per page, written on every miss | one INCR |
| Old keys | deleted at once | left until their TTL ends |
| Easy to get wrong | yes: forget one layer and it's stale (section 8) | hard to get wrong |
| Good for | frequent small edits | rare bulk re-indexing |
Both need one more thing: something has to tell the app that a page changed. If you run three copies of the app, all three need to hear it. That's a job for Redis pub/sub, and it's the subject of the next guide.
10. "Similar" is not "same"
Exact-key caching only helps when people type the same question. "Is the yearly Plus fee refundable?" means the same as the question we cached, but it's a different key, so it's a miss (1,225 ms in my run).
A semantic cache tries to fix that. It embeds each new question and reuses a cached answer if an old question is similar enough: if the cosine similarity of the two vectors is above a threshold. It sounds like a free upgrade. Here's what the similarity scores looked like with the same embedding model:
0.800 same question, new words 'Can I get a refund on my Plus membership?' vs 'Is the Plus membership fee refundable?'
0.671 different answer 'Can I get a refund on my Plus membership?' vs 'Can I get a refund on my order?'
0.786 different answer 'How long does standard delivery take?' vs 'How long does express delivery take?'
0.909 different answer 'What does error E-4013 mean?' vs 'What does error E-4012 mean?'
Try the threshold yourself:
At 0.800: 2 reused · 1 wrong answer
The real paraphrase scored 0.800. Two questions about different error codes scored 0.909, because the sentences differ by one character. Any threshold low enough to catch the paraphrase also gives the E-4012 answer to someone asking about E-4013. That answer would be confidently and specifically wrong.
This is one small model and four pairs, so the exact numbers don't generalise. The lesson does: similarity measures how alike the words are, not whether the answers are the same. If you add a semantic cache, keep it for questions where a near-miss answer is harmless, and test it on pairs that differ by one important detail.
Part 4
Run it for real
Redis down, Redis full, and the same code on Azure Managed Redis.
11. When Redis is down
A cache is an optimisation. If Redis is unreachable, the app should still answer, just slower. My lab wraps every Redis call so an error counts as a miss. That part worked: every request still answered.
The timing didn't. I pointed the app at a port where nothing was listening and asked three questions:
redis-py 8.0.1 defaults 50313 ms 59951 ms 58370 ms errors 40, calls skipped 0
tuned: 1 retry + skip cache for 30 s 1249 ms 1223 ms 1221 ms errors 1, calls skipped 39
With the default settings, redis-py 8.0.1 retries a failed connection up to 10 times, waiting a little longer each time (I checked the defaults in its source: 10 retries, 10 ms base backoff, 1 second cap). One GET against the dead port took 3.3 seconds on its own. A request makes about 13 cache calls, so a request with no cache took about a minute.
The fix has two parts:
from redis.backoff import NoBackoff
from redis.retry import Retry
r = redis.Redis(host=..., socket_timeout=0.5, socket_connect_timeout=0.5,
retry=Retry(NoBackoff(), 1)) # one quick retry, not ten slow ones
and a breaker: after a Redis error, skip the cache entirely for 30 seconds instead of trying again on every call. With both, the requests took 1.2 seconds, the cost of the model step alone, and 39 of 40 cache calls were skipped without touching the network.
Retries are useful for a database you can't work without. For a cache you can work without, they turn a small outage into a big one.
12. When Redis is full
Memory is finite. When Redis reaches its limit (maxmemory), its eviction policy decides what happens next. I filled a 4 MB Redis with 10,000 answers of 1 KB each, once per policy:
noeviction writes ok: 2013 keys kept: 2013 evicted: 0 first error: command not allowed when used memory > 'maxmemory'.
allkeys-lru writes ok: 10000 keys kept: 2224 evicted: 7795 first error: None
noeviction
the default
2,013keys kept
2,013evicted
0
Write 2,014 failed: “command not allowed when used memory > 'maxmemory'”. New answers can no longer be cached.
allkeys-lru
what a cache wants
10,000keys kept
2,224evicted
7,795
Every write succeeded. Redis dropped the least recently used keys to make room.
noeviction is the default on a plain Redis server (my Docker container reported maxmemory-policy noeviction). It suits Redis used as a database, where losing data is worse than refusing writes. For a cache it's the wrong choice: once memory is full, new answers can't be cached at all.
allkeys-lru drops the least recently used keys to make room. Old, unpopular answers go first, and popular ones stay. That's exactly what a cache wants. Azure Managed Redis lets you choose the policy when you create the database. The Microsoft Learn lab for this topic uses AllKeysLRU.
13. The same code on Azure Managed Redis
Azure Managed Redis is Redis run by Microsoft: no servers to patch, with high availability and scaling built in. The commands are the same, so the cache code doesn't change. Only the connection does:
| Local (the lab) | Azure Managed Redis | |
|---|---|---|
| Host | localhost | <name>.<region>.redis.azure.net |
| Port | 6379 (6390 in the lab) | 10000 |
| Encryption | none | TLS, required |
| Login | none | Microsoft Entra ID token (or an access key) |
from redis_entraid.cred_provider import create_from_default_azure_credential
provider = create_from_default_azure_credential(("https://redis.azure.com/.default",))
r = redis.Redis(host=os.environ["REDIS_HOST"], port=10000, ssl=True,
credential_provider=provider,
socket_timeout=0.5, socket_connect_timeout=0.5,
retry=Retry(NoBackoff(), 1))
The credential provider signs in with whoever is logged in (az login on a laptop, or the app's managed identity in Azure) and refreshes the token before it expires. No password goes in your code. Your identity needs a data access policy on the database.
Two things to settle when you create it. Put Redis in the same region as your app, because every cache call crosses the network and a hit should stay a fraction of a millisecond away. And choose the eviction policy from section 12.
I didn't run this lab on Azure for this article. The connection details are from the Microsoft Learn module "Implement data operations in Azure Managed Redis" (AI-200), checked in October 2026.
14. Common mistakes
- No TTL. Every key should expire. Without a TTL, a missed invalidation lasts forever.
- Invalidating one layer. Section 8: delete the answer, and the retrieval cache rebuilds it stale. Clear every layer built from the page, or bump the version.
- Keys without the model name. Switch embedding models and old vectors look valid but live on a different map. Put the model in the embedding key.
- Caching per-user answers under a shared key. If an answer depends on who's asking (their orders, their plan), the key must include the user or tenant, or one customer sees another's answer.
- Default retries on a cache. Section 11: about a minute per request while Redis was down.
noevictionon a cache. Section 12: the cache stops accepting new answers when it's full.- Using
KEYSto find keys. It walks every key in one blocking call. UseSCAN, which walks them in small batches, or keep dependency sets so you never need to search. - Measuring only the hit rate. A high hit rate on wrong answers is worse than a miss. Re-run your test questions after every content change.
The whole guide in ten lines
- A RAG request embeds, retrieves and writes; the model step is the slow, expensive one.
- Redis keeps values in memory under keys and returns them in under a millisecond.
- Cache-aside: check Redis, and on a miss do the work and store a copy.
- Give every copy a TTL, and build keys from layer, document version and the normalised question.
- Cache three layers: answers for speed, chunks and question vectors for the misses.
- In my run, the answer cache cut a request from 1,262 ms to 0.6 ms.
- When a page changes, every layer built from it is stale. Deleting only the answer wasn't enough.
- Invalidate precisely with dependency sets, or bluntly by bumping a version in every key.
- Semantic caching reuses answers for similar questions, and similar questions can need different answers.
- A dead cache must be cheap (few retries, a breaker), and a full cache must evict (allkeys-lru).
Next: how the "this page changed" message reaches every copy of your app, and how re-embedding work gets done exactly once, even when a worker crashes. That's Redis pub/sub and Streams.
Checked against Redis 7.4.11 (Docker image redis:7.4), redis-py 8.0.1, pgvector 0.8.7 on PostgreSQL 17.11, sentence-transformers 6.1.0 and all-MiniLM-L6-v2, as of October 2026. All timings, keys, similarity scores and eviction counts are real output from my laptop on 7 October 2026. The "write" step is a fixed 1.2-second stand-in, not a language model, so the totals show the shape of the savings, not a benchmark. The redis-py retry defaults were read from its source (redis/_defaults.py). The Azure section was not run for this article and follows the Microsoft Learn AI-200 module. The Parcelo help centre is invented.