October 5, 2026 · 14 min read
Debugging RAG: Check Retrieval Before You Touch the Model
A plain-English guide to why RAG answers go wrong. Split the system into a clerk (retrieval) and a writer (the model), run a five-minute test to find out which one failed, then try it on a small help centre with pgvector. Every number is measured.
When a RAG answer is wrong, the usual reaction is to rewrite the prompt or try a bigger model. Often neither helps, because the model never saw the right passage. This guide gives you a five-minute test that tells you which half of the system failed, then runs it on a small help centre so you can see what each kind of failure looks like.
You don't need to know how embeddings work. The vector database guide covers them, and each idea here is explained again in a line. Every section ends with a one-line Remember, and there's a nine-line summary at the end.
The guide has three parts. If you only want one thing, jump straight to it:
- Part 1 · Understand it: what a RAG system is made of, and the test, sections 1–3.
- Part 2 · Run the test: a small lab with real output, sections 4–7.
- Part 3 · Use it well: what to check when retrieval is fine, and how to keep the test, sections 8–10.
Part 1
Understand it
Three short steps: what RAG is, where it can fail, and the test that tells you where.
1. A RAG system is a clerk and a writer
A language model doesn't know your documents. RAG (retrieval-augmented generation) fixes that at question time. The system first looks up the passages most likely to hold the answer, pastes them into the prompt, and asks the model to answer from them.
Picture two people. A clerk goes to the archive and puts a few pages on a desk. A writer then answers the customer using only the pages on the desk. That picture runs through the whole guide.
The running example is a made-up online shop called Parcelo, with a help centre of 12 pages. Customers ask things like "Can I get my money back for the yearly membership?" and the system has to answer from those pages.
Stages 2 and 3 are the clerk's job. Stage 4 is the writer's. A wrong answer can come from either, and they need different fixes.
2. The writer can only use what's on the desk
If the right page isn't on the desk, the writer can't use it. A better prompt won't put it there, and neither will a bigger model. The model may fill the gap with a plausible guess, and that guess is what people call a hallucination.
The reverse also holds. If the right page is on the desk and the answer is still wrong, the clerk did their job and the problem is in the writing step.
So the first question for any wrong answer is not "how do I improve the prompt?" It is:
3. The five-minute test
Take one question the system got wrong, and do this:
- Get the exact chunks that were retrieved for it. Log them or print them. Don't guess.
- Read the top five.
- Ask: is the answer in there?
The two branches send you to different places:
| The answer is… | It's a… | Look at |
|---|---|---|
| Not in the top five | Retrieval problem | How pages were cut into chunks, the embedding model, metadata filters, the ranking, how many chunks you pass on |
| In the top five | Generation problem | The prompt, the order of the chunks, irrelevant chunks around it, two pages that disagree, the model |
The test takes minutes, and it stops you from spending a week on the wrong half.
Part 2
Run the test
A small lab with real output. You can run it yourself.
4. The lab
Here's the setup. Everything below was run on 5 October 2026.
- The pages. 12 invented help-centre pages for Parcelo: returns, delivery, tracking, lost parcels, checkout error codes, gift cards and so on.
- The questions. 10 questions a customer might ask. For each one I wrote down a phrase that must appear in the retrieved text for any model to answer correctly, such as
refundable only within 14 days. - The check. For each question, retrieve the top chunks and ask whether that phrase is inside one chunk. That is the five-minute test, automated.
- The search. Each chunk is turned into 384 numbers by a small local embedding model (
all-MiniLM-L6-v2) and stored in Postgres with pgvector. With only 25 chunks, the search is exact and there's no index to tune.
The search itself is one query. <=> is pgvector's cosine distance, and the smallest distance is the closest meaning:
SELECT doc_id, content, embedding <=> $1 AS distance
FROM chunks
ORDER BY distance
LIMIT 5;
The lab is four small files: corpus.py, retrieval_check.py, run.sh and requirements.txt. You need Docker and Python 3.
./run.sh # start pgvector, install packages
python retrieval_check.py --strategy fixed # check all 10 questions
python retrieval_check.py --strategy fixed --show q4 # print the top 5 for one question
Before any search, each page has to be cut into chunks, and there are many ways to cut. I tried two:
- Fixed: cut every 300 characters. It's the simplest thing that works, and it ignores the page's structure.
- Structure: cut at each heading, and put the page title and heading in front of the chunk.
5. Failure 1: the chunker cut the answer in half
Start with the fixed chunker. This is what I got for all 10 questions when I ran it:
strategy=fixed chunks=25 model=sentence-transformers/all-MiniLM-L6-v2
q first rank with the answer question
q1 2 Can I get my money back for the yearly Plus membership?
q2 2 My parcel hasn't arrived and the date has passed. What now?
q3 1 When does a missing parcel count as lost?
q4 not in top 10 What does error E-4013 mean?
q5 1 Can I change my address after the order has shipped?
q6 1 How long do I have to use my gift card?
q7 1 Do you keep my invoices after I delete my account?
q8 3 Is express delivery free for members?
q9 1 How do I download everything you store about me?
q10 1 I paid but got no order confirmation. What should I do?
answer in top-1: 6/10 top-3: 9/10 top-5: 9/10
Nine of ten questions have the answer somewhere in the top five. Question 4 never gets it, even in the top ten. Run the five-minute test on it:
python retrieval_check.py --strategy fixed --show q4
Q: What does error E-4013 mean?
Needed in context: 'e-4013: the billing address'
[1] errors (distance 0.406)
…E-4012: The card was declined by the issuing bank. Try another card or contact your bank.
E-4013: Th
[2] errors (distance 0.555)
e billing address does not match the address your bank holds. Update the billing address and retry.
E-4021: The card does not support online payments. …
The clerk found the right page. The chunker had already cut the line for E-4013 in the middle of the word "The". Neither chunk holds the whole sentence.
Fixed 300-character chunks
E-4013: Th
✂ cut here
E-4021: The card does not support…
The needed sentence is in no single chunk. Result: not in the top 10.
Cut at the page structure, with the title added
Checkout shows a short code when a payment cannot be completed…
E-4012: The card was declined by the issuing bank…
E-4013: The billing address does not match the address your bank holds. Update the billing address and retry.
E-4021: The card does not support online payments…
The whole sentence sits in one chunk. Result: rank 1.
This is a retrieval problem, and no prompt change would have found it. The fix is in how the page is cut. Cutting at headings, and keeping the error-code list together with its title, gave this on the same 10 questions:
strategy=sections chunks=25 model=sentence-transformers/all-MiniLM-L6-v2
...
q4 1 What does error E-4013 mean?
...
answer in top-1: 8/10 top-3: 10/10 top-5: 10/10
6. Failure 2: the right chunk is there, but not first
Look at question 1 with the fixed chunker. The needed sentence is in the top five, but at rank 2. The top chunk holds the second half of the same paragraph:
[1] returns (distance 0.555)
urchase, and only if you have not used free delivery. After 14 days the fee is not refundable, …
[2] ✓ returns (distance 0.571)
Items that cannot be returned … ## Membership fees
Parcelo Plus membership fees are refundable only within 14 days of p
Cutting at headings fixes the split, but question 1 still doesn't get rank 1. The structure chunker puts a different chunk first, the "Cancelling" section of the membership page:
[1] membership (distance 0.559)
Parcelo Plus membership > Cancelling
Cancel any time from Account settings. …
[2] ✓ returns (distance 0.566)
Returns and refunds > Membership fees
Parcelo Plus membership fees are refundable only within 14 days of purchase, …
Both chunks are about the yearly membership, and the distances differ by 0.007. If you only pass the top chunk to the writer, it has the wrong page. If you pass the top three, it has the right one.
That is what the k setting does: how many chunks go onto the desk. Here is how often the needed phrase was inside the top 1, 3 and 5 chunks:
| Top 1 | Top 3 | Top 5 | |
|---|---|---|---|
| Fixed 300-character chunks | 6/10 | 9/10 | 9/10 |
| Cut at page structure | 8/10 | 10/10 | 10/10 |
| Structure, plus keyword search merged in | 8/10 | 10/10 | 10/10 |
A larger k recovers misses, and it has a price. Every extra chunk costs tokens and adds text the writer has to ignore. A common fix is reranking: fetch a longer list, say 20, then let a slower, more accurate model put the best few on top. I did not run a reranker here.
7. What these numbers can't tell you
These results are real, and they are small. Read them as an illustration of the method.
- The corpus is tiny and invented. 12 pages and 10 questions. I wrote the questions after writing the pages, so they are friendlier than real customer questions.
- The check is crude. It looks for one phrase inside one chunk. An answer can be possible without that exact phrase, and a chunk can hold the phrase and still mislead.
- There is no index. With 25 chunks, search is exact, so these numbers say nothing about the recall of an approximate index such as HNSW.
- One embedding model. A different model would rank differently.
- Keyword search made no difference here. Merging keyword results in gave 10/10 at top 5, the same as vector search alone. On a bigger set of error codes, exact matches may matter more. This lab can't show it, and yours might.
- I didn't run a language model. The lab stops at retrieval. Part 3 is a checklist, not a measurement.
Part 3
Use it well
What to check when retrieval is fine, and how to keep the test.
8. When the answer is on the desk
Suppose the five-minute test says the answer was in the top five and the answer is still wrong. Now the writing step is the suspect. These are the things to check, in the order I'd try them:
- Noise. Are four irrelevant chunks sitting around the right one? Try a smaller k, or rerank.
- Order. Some models use the start and end of a long prompt better than the middle. Move the best chunk and see if the answer changes. Test it on your model.
- Conflicting pages. An old policy and a new policy, both in the context. The writer can't know which is current unless the chunks say so. Put dates or versions in the chunk text.
- Instructions. Tell the model to answer only from the provided pages, and to say so when the pages don't contain the answer.
- The model. Only now is it worth trying a different or larger one.
I didn't measure any of this. It's a list of places to look, in an order that costs the least.
9. Make the test permanent
The five-minute test works best as a habit, and then as a unit test.
- Write 20 to 30 real questions. Take them from support tickets or search logs, not from your own head.
- Write the phrase each one needs. Do it once. Fix the phrase if the page changes.
- Re-run the check after every change to chunking, the embedding model, filters or k. If the hit rate drops, you know before users do.
- Log the retrieved chunk IDs in production. Then any wrong answer can be tested in five minutes, with the exact chunks the model saw.
10. What trips people up
- Judging from one example. One failing question tells you where to look. It doesn't tell you what's typical.
- Changing the embedding model without re-embedding. Old and new vectors live on different maps, so results look random. Re-embed everything, and store which model made each vector.
- Chunks that are too small or too big. Too small cuts answers apart, as in section 5. Too big blurs several topics into one position. Test a few sizes on your own questions.
- A metadata filter that hides the right page. Filters for tenant, region or permissions run before the ranking. A wrong filter silently removes the page, and it looks exactly like a retrieval miss.
- Fixing the prompt first. It's the easiest thing to change, and it only helps when the answer was already on the desk.
- Reading distances as certainty. A distance is only meaningful next to the other distances for the same question. Don't pick one cutoff for every question.
The whole guide in nine lines
- RAG is a search step (the clerk) followed by a writing step (the writer).
- The writer can only use the pages the clerk puts on the desk.
- For a wrong answer, print the top five retrieved chunks and ask whether the answer is in them.
- Not in them: a retrieval problem. In them: a generation problem.
- Chunking can cut an answer in half. Cut along the page's structure and keep the title with each chunk.
- The right chunk must also rank high enough to be passed on. k and reranking control that.
- In my small run, structure-based chunks raised top-1 from 6/10 to 8/10 and top-5 from 9/10 to 10/10.
- When the answer is on the desk, check noise, order, conflicts, instructions, then the model.
- Keep 20 to 30 real questions and re-run the check after every change.
Checked against pgvector 0.8.7 on PostgreSQL 17.11 (Docker image pgvector/pgvector:pg17), with sentence-transformers 6.1.0 and the all-MiniLM-L6-v2 model, as of October 2026. The retrieval tables and printed chunks are real output from my run on 5 October 2026 (chunk text is shortened with "…" where marked). The Parcelo help centre is invented. The checklist in section 8 was not measured, and no language model was run.