Aryan Tripathi — Writing
← All writing

October 7, 2026 · 19 min read

Caching a RAG App with Redis: What to Cache, When It Goes Stale, and What Breaks

A plain-English guide to putting Redis in front of a RAG app. Learn cache-aside and TTLs, cache the three slow steps of a RAG request, watch a cached answer go stale when a help page changes, and see what happens when Redis is down or full. Every number is measured.

#redis#caching#rag#ai-engineering#azure#performance

Most RAG apps answer the same questions again and again, and pay for the model every time. A cache fixes that in an afternoon. It also brings new ways to be wrong: answers that go stale, a "similar question" that isn't, and a dead cache that makes the app slower than having no cache at all. This guide builds the cache step by step on a small help centre, and every one of those failures is shown with real output.

You don't need to know Redis. Each idea is explained the first time it appears, and every section ends with a one-line Remember. There's a ten-line summary at the end.

The running example is Parcelo, the same invented online shop from the RAG debugging guide. Its help centre has 12 pages, and customers ask questions like "Can I get my money back for the yearly Plus membership?"

The guide has four parts. Jump to the one you need:

  • Part 1 · Understand it: why RAG requests are slow, what Redis is, cache-aside and TTLs, sections 1–4.
  • Part 2 · Cache a RAG app: three cache layers and what they saved, sections 5–7.
  • Part 3 · When the cache lies: stale answers and "similar" questions, sections 8–10.
  • Part 4 · Run it for real: Redis down, Redis full, Azure Managed Redis, sections 11–14.

Part 1

Understand it

What makes a RAG request slow, and the one pattern that fixes it.

1. A RAG request pays three times

In the RAG guide we split a RAG app into a clerk and a writer. The clerk finds the right pages. The writer answers from them. Every request makes them do three jobs:

  1. Embed the question: turn it into a list of numbers (a vector) so it can be compared with the pages.
  2. Retrieve: find the 5 chunks of text closest to that vector.
  3. Write: send those chunks and the question to a language model, which writes the answer.

When I ran these on my laptop, embedding took about 49 ms and retrieval about 4 ms. The model is the slow and expensive part. A hosted model usually takes a second or more and is billed per token. My lab uses a stand-in that waits a fixed 1.2 seconds, so the totals below are realistic in shape, not a benchmark of any model.

Now notice what a help centre looks like from the inside. The same few questions arrive all day: refunds, delivery times, lost parcels. Each one goes through all three jobs, and gets the same answer as last time.

RememberEvery RAG request embeds, retrieves and writes. The writing step costs the most, and repeated questions pay it again for nothing.

2. Redis is a fast notebook

Imagine the clerk keeps a notebook on the desk. Before going to the archive, they check the notebook: "Did someone ask this already?" If yes, they read out the answer. If not, they do the work and write the answer down for next time.

Redis is that notebook. It's a database that keeps everything in memory (RAM) instead of on disk, so reading a value takes well under a millisecond. You store a value under a name, called a key, and read it back by the same name:

> SET greeting "hello"
OK
> GET greeting
"hello"
> GET something-else
(nil)

(nil) means "nothing stored under that key". In Python, with the redis package, the same thing reads r.set("greeting", "hello") and r.get("greeting"), and a missing key comes back as None.

Redis can hold more than plain text (hashes, lists, sets, streams), but for caching, a string per key covers almost everything.

RememberRedis keeps values in memory under a key. Reading one takes less than a millisecond.

3. Cache-aside: look first, then work, then write it down

The notebook routine has a name: cache-aside. Your app, not Redis, does all of it:

  1. Look in Redis for the key.
  2. Hit: it's there. Use it.
  3. Miss: it's not. Do the slow work, then store the result in Redis.

Watch two identical questions go through it:

Animatedcache-aside · two requests
Your appthe RAG serviceRedisin memory · sub-millisecondThe slow workembed · search · modelrequest 1 · 1,262 msrequest 2 · 0.6 ms

7. Answered in 0.6 ms. The model never ran.

GET ans:v1:9a66ab4fb01269b7 → (nil)
embed 49 ms · retrieve 4 ms · write 1,200 ms (stand-in)
SET ans:v1:9a66ab4fb01269b7 '{"answer": …}' EX 3600 → OK
GET ans:v1:9a66ab4fb01269b7 → '{"answer": …}'
step 7 of 7
Cache-aside: look in Redis first; on a miss, do the work and save a copy. The first request pays, every repeat is answered from memory until the copy expires.

In code it's five lines around the slow part:

cached = r.get(key)
if cached is not None:              # hit
    return json.loads(cached)
answer = do_the_slow_work(question)  # miss: embed, retrieve, write
r.set(key, json.dumps(answer), ex=3600)
return answer

Redis never fetches anything by itself. If the app forgets to store the result, or stores the wrong thing, the cache stays empty or wrong. That's why it's called cache-aside: the cache sits next to the real work, and the app keeps the two in step.

RememberCache-aside: check Redis, and on a miss do the work and save a copy. The app is responsible for both halves.

4. Every copy gets an expiry date

A copy in the notebook can go out of date. Parcelo might change its refund policy tomorrow. The simplest protection is to write an expiry date on every page of the notebook.

In Redis that's a TTL (time to live): a countdown attached to a key. When it reaches zero, Redis deletes the key by itself. You set it when you write, with ex= (seconds):

AnimatedSET … EX 10, then wait
ans:v1:9a66ab4fb01269b7deleted by Redis
> SET ans:v1:9a66ab4fb01269b7 "…" EX 10 → OK
> TTL ans:v1:9a66ab4fb01269b7 → (integer) -2
> GET ans:v1:9a66ab4fb01269b7 → (nil)

Three TTL answers to remember: a positive number is seconds left, -1 means the key exists but never expires, -2 means there is no such key.

step 12 of 12
A TTL is a countdown on the key. When it reaches zero, Redis deletes the key by itself. TTL then answers -2: the key no longer exists.

The TTL is a promise about how wrong you're willing to be. A one-hour TTL means that, without any other help, an answer can be up to an hour out of date. Section 8 shows how to do better than waiting.

The key name matters as much as the expiry. In the lab, every key has three parts:

Diagramone key, three parts
answhich layer
:
v1corpus version
:
9a66ab4fb01269b7hash of the normalised question
“Can I get my money back for the yearly Plus membership?”
“ can i get my money back for the yearly plus membership”
“Can I get my money back for the yearly Plus membership?? ”

All three normalise to the same text, so all three hit the same answer (0.5–0.6 ms each when I ran it). “Is the yearly Plus fee refundable?” means the same thing but is a different key: a miss.

A key names the layer, the version of the documents it was built from, and the question. Normalising first means “Hi??” and “hi” share one entry.

The hash is a short fingerprint of the question. A fingerprint keeps keys short, and the same question always produces the same fingerprint. Normalising the question first (lower case, single spaces, no trailing "?") lets small differences in typing share one entry.

RememberGive every cached value a TTL, and build its key from the layer, the document version and the normalised question.

Part 2

Cache a RAG app

Three cache layers on the Parcelo help centre, measured.

5. Three things worth caching

Each of the three jobs from section 1 can have its own cache. They are worth different amounts and go stale in different ways:

Diagramwhere each cache sits
Answer cache · checked firstans:v<N>:<hash> · TTL 1 h · a hit skips every box below (0.6 ms)Questionnormalise itEmbedding cacheemb:minilm:<hash>TTL 24 hEmbedquestion → vector49 msRetrieval cacheret:v<N>:<hash>:k5TTL 1 hRetrievetop 5 chunks4 msWritethe model≈1,200 ms*Answerto the usercold request, median of 10 questions, when I ran it
Three layers. The answer cache is checked first and skips everything. On a miss, the embedding and retrieval caches still save their own steps. *The write step is a 1,200 ms stand-in, not a measurement.
LayerKeyStoresTTL in the labGoes stale when
Answer cacheans:v1:<hash>the final answer and its source page1 hourany page it used changes
Retrieval cacheret:v1:<hash>:k5the top 5 chunks1 hourany of those pages changes
Embedding cacheemb:minilm:<hash>the question's vector (1,536 bytes)24 hoursonly if you change the embedding model

The answer cache is checked first. A hit skips everything. On a miss, the two lower layers still save their own steps.

The embedding cache is the safest of the three. A question's vector depends only on the question and the model, so it can't go stale when pages change. That's why its key names the model (minilm) instead of the document version.

RememberCache the answer for speed, the retrieved chunks and the question vector for the misses. Only the vector cache is safe from page changes.

6. The lab

Here's the setup. Everything was run on 7 October 2026.

  • Redis 7.4 in Docker (redis:7.4), and PostgreSQL 17 with pgvector in Docker for the chunks.
  • Python 3.14 with redis 8.0.1 (the client), sentence-transformers 6.1.0 and the all-MiniLM-L6-v2 embedding model.
  • The 12-page Parcelo help centre, cut at headings with the page title added to each chunk (the fix from the RAG guide).
  • No language model. The "write" step picks the retrieved sentence that best matches the question, then waits a fixed 1.2 seconds to stand in for a model call. The 1.2 seconds is my choice, not a measurement. Everything else is measured.

The heart of it is one function per layer, each following the routine from section 3. Here is the embedding layer:

def embed(cache, q):
    key = f"emb:minilm:{h(normalise(q))}"
    if (hit := cache.get(key)) is not None:
        return np.frombuffer(hit, dtype=np.float32), "hit"
    vec = model.encode([q], normalize_embeddings=True)[0].astype(np.float32)
    cache.set(key, TTL_EMBEDDING, vec.tobytes())   # 384 floats = 1,536 bytes
    return vec, "miss"

The retrieval and answer layers look the same, with one addition: they put the corpus version in the key, and they record which pages they were built from. Section 9 explains why.

To run it yourself:

./run.sh                               # Redis on 6390, Postgres on 5545, venv
./.venv/bin/python cached_rag.py layers   # also: keys, stale, semantic, down
RememberThe lab caches all three layers with cache-aside. Only the model step is a stand-in.

7. What the cache saved

I asked all 10 test questions three times: once with nothing cached, once with only the lower two layers warm, and once with the answer cache on.

median per request                 cold (all miss)   warm, no answer cache
embed the question                         49.1 ms                  0.8 ms
retrieve top 5 chunks                       4.4 ms                  1.4 ms
whole request                            1262.4 ms               1209.0 ms

warm, answer cache on: whole request      0.6 ms
Real outputmedian of 10 questions · log scale · when I ran it
0.1ms1ms10ms100ms1s10sNothing cachedembed + retrieve + write1,262.4 msEmbedding + retrieval cachedonly the write step runs1,209 msAnswer cachedone Redis GET0.6 ms
embed the question: 49.1 ms → 0.8 ms
retrieve top 5: 4.4 ms → 1.4 ms
Caching the cheap steps barely moves the total, because the model step dominates. Caching the answer removes it: 1,262 ms became 0.6 ms. (The write step is a fixed 1,200 ms stand-in.)

Two things stand out.

First, the lower layers barely move the total. Saving 48 ms on embedding is real, but next to a model call it's 4% of the request. Those layers matter on misses: a question that's new as an answer can still reuse its vector and chunks.

Second, the answer cache removes the model call completely. 1,262 ms became 0.6 ms. If 40% of your questions are repeats, 40% of your model bill and latency goes away.

Here's what the notebook held after one question, from the keys demo:

ans:v1:9a66ab4fb01269b7                 string    232 bytes   TTL 3600 s
corpus:version                          string     56 bytes   TTL -1 s
deps:membership                         set       128 bytes   TTL 86400 s
deps:returns                            set       120 bytes   TTL 86400 s
emb:minilm:9a66ab4fb01269b7             string   1864 bytes   TTL 86399 s
ret:v1:9a66ab4fb01269b7:k5              string   1608 bytes   TTL 3599 s

The two deps: sets are the record of which pages this answer was built from. You'll need them in a moment.

RememberThe answer cache gives the big win (1,262 ms to 0.6 ms). The lower layers help the questions the answer cache misses.

Part 3

When the cache lies

Stale answers, and why "similar" isn't "same".

8. A help page changes

Parcelo changes its policy: Plus membership fees are now refundable for 30 days instead of 14. Someone edits the returns page, and the app re-embeds it. The database is correct.

The cache isn't. Step through what happened in my run:

Animatedthe returns page changes · when I ran it

5. Or: bump the version

INCR corpus:version changes every key name from v1 to v2. Nothing old can be found any more. The 2 old keys wait in memory until their TTL ends.

Answer cacheans:v2:9a66ab4f…
“refundable within 30 days”
Retrieval cacheret:v2:9a66ab4f…:k5
“refundable within 30 days”
Database (source of truth)chunks table
“refundable within 30 days”

User sees: “… refundable only within 30 days …” ✓ correct

step 5 of 5
A cache layer built on top of another cache layer inherits its staleness. Invalidate every layer that was built from the changed page, or change the key names for all of them at once.

Step 3 is the trap. Deleting the cached answer feels like enough. But the answer was rebuilt from the retrieval cache, which still held the old chunks, so the "fresh" answer said 14 days again. This is the real output:

2) no invalidation              answer hit  -> … refundable only within 14 days …
3) delete the answer key only   answer miss retrieval hit -> … refundable only within 14 days …
4) delete every key built from 'returns' (2 keys)  answer miss retrieval miss -> … within 30 days …

A TTL would have fixed it eventually, after up to an hour of wrong answers. For a refund policy, an hour of wrong answers means real support tickets.

RememberWhen a page changes, every cache layer built from it is stale, including layers that other layers read from.

9. Two ways to invalidate

Invalidation means removing copies you know are wrong, instead of waiting for them to expire. The lab shows two ways.

Precise: delete what depended on the page. Each time the app stores an answer or a retrieval result, it adds that key to a set named after every page it used:

for doc in {c["doc"] for c in chunks}:
    r.sadd(f"deps:{doc}", key)          # deps:returns -> {ans:v1:…, ret:v1:…}

When the returns page changes, read deps:returns and delete every key in it. Only answers that used that page are dropped, and everything else stays cached.

Blunt: change every key name at once. The corpus version is part of every retrieval and answer key. After any page changes, increase it:

> INCR corpus:version
(integer) 2

Now the app builds keys starting ans:v2: and ret:v2:. Nothing from v1 can be found any more, so every request rebuilds from the database. The old keys aren't deleted. They sit in memory until their TTL ends (2 keys in my run, thousands in a real one).

Delete dependantsBump the version
What's droppedonly answers that used the changed pageeverything except embeddings
Extra worka set per page, written on every missone INCR
Old keysdeleted at onceleft until their TTL ends
Easy to get wrongyes: forget one layer and it's stale (section 8)hard to get wrong
Good forfrequent small editsrare bulk re-indexing

Both need one more thing: something has to tell the app that a page changed. If you run three copies of the app, all three need to hear it. That's a job for Redis pub/sub, and it's the subject of the next guide.

RememberDelete the keys that used a page for precise invalidation, or bump a version in every key to drop everything at once.

10. "Similar" is not "same"

Exact-key caching only helps when people type the same question. "Is the yearly Plus fee refundable?" means the same as the question we cached, but it's a different key, so it's a miss (1,225 ms in my run).

A semantic cache tries to fix that. It embeds each new question and reuses a cached answer if an old question is similar enough: if the cosine similarity of the two vectors is above a threshold. It sounds like a free upgrade. Here's what the similarity scores looked like with the same embedding model:

0.800  same question, new words   'Can I get a refund on my Plus membership?' vs 'Is the Plus membership fee refundable?'
0.671  different answer           'Can I get a refund on my Plus membership?' vs 'Can I get a refund on my order?'
0.786  different answer           'How long does standard delivery take?' vs 'How long does express delivery take?'
0.909  different answer           'What does error E-4013 mean?' vs 'What does error E-4012 mean?'

Try the threshold yourself:

Interactiveall-MiniLM-L6-v2 · cosine · when I ran it
0.800
“What does error E-4013 mean?” vs “What does error E-4012 mean?”0.909
Different question, different answer. Served the wrong cached answer.
“Can I get a refund on my Plus membership?” vs “Is the Plus membership fee refundable?”0.800
Same question, new words. Served from cache, correctly.
“How long does standard delivery take?” vs “How long does express delivery take?”0.786
Different question, different answer. Not reused: goes to the model.
“Can I get a refund on my Plus membership?” vs “Can I get a refund on my order?”0.671
Different question, different answer. Not reused: goes to the model.

At 0.800: 2 reused · 1 wrong answer

A “semantic cache” reuses an answer when a new question is similar enough to an old one. Move the threshold: any setting that catches the real paraphrase also serves the E-4012 answer to an E-4013 question.

The real paraphrase scored 0.800. Two questions about different error codes scored 0.909, because the sentences differ by one character. Any threshold low enough to catch the paraphrase also gives the E-4012 answer to someone asking about E-4013. That answer would be confidently and specifically wrong.

This is one small model and four pairs, so the exact numbers don't generalise. The lesson does: similarity measures how alike the words are, not whether the answers are the same. If you add a semantic cache, keep it for questions where a near-miss answer is harmless, and test it on pairs that differ by one important detail.

RememberA semantic cache reuses answers for similar questions, and similar questions can need different answers. Exact keys are slower to hit but never wrong this way.

Part 4

Run it for real

Redis down, Redis full, and the same code on Azure Managed Redis.

11. When Redis is down

A cache is an optimisation. If Redis is unreachable, the app should still answer, just slower. My lab wraps every Redis call so an error counts as a miss. That part worked: every request still answered.

The timing didn't. I pointed the app at a port where nothing was listening and asked three questions:

redis-py 8.0.1 defaults                    50313 ms    59951 ms    58370 ms   errors 40, calls skipped 0
tuned: 1 retry + skip cache for 30 s        1249 ms     1223 ms     1221 ms   errors 1, calls skipped 39
Real outputRedis unreachable · 3 requests · when I ran it
0 s20 s40 s60 sredis-py 8.0.1 defaults50.3 s60.0 s58.4 s1 retry + skip the cache for 30 s1.2 s1.2 s1.2 s
With redis-py's defaults, every cache call retried 10 times before giving up. Each request made about 13 cache calls, so a dead cache turned a 1.2-second answer into a one-minute answer.

With the default settings, redis-py 8.0.1 retries a failed connection up to 10 times, waiting a little longer each time (I checked the defaults in its source: 10 retries, 10 ms base backoff, 1 second cap). One GET against the dead port took 3.3 seconds on its own. A request makes about 13 cache calls, so a request with no cache took about a minute.

The fix has two parts:

from redis.backoff import NoBackoff
from redis.retry import Retry

r = redis.Redis(host=..., socket_timeout=0.5, socket_connect_timeout=0.5,
                retry=Retry(NoBackoff(), 1))     # one quick retry, not ten slow ones

and a breaker: after a Redis error, skip the cache entirely for 30 seconds instead of trying again on every call. With both, the requests took 1.2 seconds, the cost of the model step alone, and 39 of 40 cache calls were skipped without touching the network.

Retries are useful for a database you can't work without. For a cache you can work without, they turn a small outage into a big one.

RememberA dead cache must cost almost nothing. Use short timeouts, few retries, and stop calling Redis for a while after it fails.

12. When Redis is full

Memory is finite. When Redis reaches its limit (maxmemory), its eviction policy decides what happens next. I filled a 4 MB Redis with 10,000 answers of 1 KB each, once per policy:

noeviction   writes ok:  2013   keys kept:  2013   evicted:     0   first error: command not allowed when used memory > 'maxmemory'.
allkeys-lru  writes ok: 10000   keys kept:  2224   evicted:  7795   first error: None
Real outputmaxmemory 4 MB · 10,000 × 1 KB writes · when I ran it

noeviction

the default

writes ok
2,013
keys kept
2,013
evicted
0

Write 2,014 failed: “command not allowed when used memory > 'maxmemory'”. New answers can no longer be cached.

allkeys-lru

what a cache wants

writes ok
10,000
keys kept
2,224
evicted
7,795

Every write succeeded. Redis dropped the least recently used keys to make room.

When memory is full, the eviction policy decides what happens. For a cache, losing old entries is fine; refusing new ones is not.

noeviction is the default on a plain Redis server (my Docker container reported maxmemory-policy noeviction). It suits Redis used as a database, where losing data is worse than refusing writes. For a cache it's the wrong choice: once memory is full, new answers can't be cached at all.

allkeys-lru drops the least recently used keys to make room. Old, unpopular answers go first, and popular ones stay. That's exactly what a cache wants. Azure Managed Redis lets you choose the policy when you create the database. The Microsoft Learn lab for this topic uses AllKeysLRU.

RememberSet a memory limit and an eviction policy such as allkeys-lru, so a full cache forgets old answers instead of refusing new ones.

13. The same code on Azure Managed Redis

Azure Managed Redis is Redis run by Microsoft: no servers to patch, with high availability and scaling built in. The commands are the same, so the cache code doesn't change. Only the connection does:

Local (the lab)Azure Managed Redis
Hostlocalhost<name>.<region>.redis.azure.net
Port6379 (6390 in the lab)10000
EncryptionnoneTLS, required
LoginnoneMicrosoft Entra ID token (or an access key)
from redis_entraid.cred_provider import create_from_default_azure_credential

provider = create_from_default_azure_credential(("https://redis.azure.com/.default",))
r = redis.Redis(host=os.environ["REDIS_HOST"], port=10000, ssl=True,
                credential_provider=provider,
                socket_timeout=0.5, socket_connect_timeout=0.5,
                retry=Retry(NoBackoff(), 1))

The credential provider signs in with whoever is logged in (az login on a laptop, or the app's managed identity in Azure) and refreshes the token before it expires. No password goes in your code. Your identity needs a data access policy on the database.

Two things to settle when you create it. Put Redis in the same region as your app, because every cache call crosses the network and a hit should stay a fraction of a millisecond away. And choose the eviction policy from section 12.

I didn't run this lab on Azure for this article. The connection details are from the Microsoft Learn module "Implement data operations in Azure Managed Redis" (AI-200), checked in October 2026.

RememberOn Azure Managed Redis, only the connection changes: port 10000, TLS, and an Entra ID token. Keep it in the app's region.

14. Common mistakes

  • No TTL. Every key should expire. Without a TTL, a missed invalidation lasts forever.
  • Invalidating one layer. Section 8: delete the answer, and the retrieval cache rebuilds it stale. Clear every layer built from the page, or bump the version.
  • Keys without the model name. Switch embedding models and old vectors look valid but live on a different map. Put the model in the embedding key.
  • Caching per-user answers under a shared key. If an answer depends on who's asking (their orders, their plan), the key must include the user or tenant, or one customer sees another's answer.
  • Default retries on a cache. Section 11: about a minute per request while Redis was down.
  • noeviction on a cache. Section 12: the cache stops accepting new answers when it's full.
  • Using KEYS to find keys. It walks every key in one blocking call. Use SCAN, which walks them in small batches, or keep dependency sets so you never need to search.
  • Measuring only the hit rate. A high hit rate on wrong answers is worse than a miss. Re-run your test questions after every content change.
RememberExpire everything, invalidate every layer, and make sure a broken cache is slower, not wrong.

The whole guide in ten lines

  1. A RAG request embeds, retrieves and writes; the model step is the slow, expensive one.
  2. Redis keeps values in memory under keys and returns them in under a millisecond.
  3. Cache-aside: check Redis, and on a miss do the work and store a copy.
  4. Give every copy a TTL, and build keys from layer, document version and the normalised question.
  5. Cache three layers: answers for speed, chunks and question vectors for the misses.
  6. In my run, the answer cache cut a request from 1,262 ms to 0.6 ms.
  7. When a page changes, every layer built from it is stale. Deleting only the answer wasn't enough.
  8. Invalidate precisely with dependency sets, or bluntly by bumping a version in every key.
  9. Semantic caching reuses answers for similar questions, and similar questions can need different answers.
  10. A dead cache must be cheap (few retries, a breaker), and a full cache must evict (allkeys-lru).

Next: how the "this page changed" message reaches every copy of your app, and how re-embedding work gets done exactly once, even when a worker crashes. That's Redis pub/sub and Streams.


Checked against Redis 7.4.11 (Docker image redis:7.4), redis-py 8.0.1, pgvector 0.8.7 on PostgreSQL 17.11, sentence-transformers 6.1.0 and all-MiniLM-L6-v2, as of October 2026. All timings, keys, similarity scores and eviction counts are real output from my laptop on 7 October 2026. The "write" step is a fixed 1.2-second stand-in, not a language model, so the totals show the shape of the savings, not a benchmark. The redis-py retry defaults were read from its source (redis/_defaults.py). The Azure section was not run for this article and follows the Microsoft Learn AI-200 module. The Parcelo help centre is invented.