October 9, 2026 · 8 min read
Is Your Vector Search Slow? 5 Fixes, Measured
Five common reasons vector search in Postgres feels slow, each tested on 5,000 real passages and 300 real questions: no index, fetching the vectors back, a new connection per question, the ef_search knob, and big vectors. Plain English, Indian memes, real numbers.
Vector search is the "find similar text" step inside AI apps, like a chatbot that answers from your documents. When it feels slow, the search itself is often not the problem. I tested five common fixes on the same data and timed each one. Some were obvious. One cost ten times more than the search.
Here's the short version as a video. Turn the sound on.
Every section ends with a one-line Remember, and there's a summary at the end.
The setup
- Data: 5,000 passages from the Natural Questions dataset, and 300 real questions whose answers are in those passages.
- Embeddings: Qwen3-Embedding-4B, 2,560 numbers per passage.
- Database: PostgreSQL 17 with pgvector 0.8.7, in Docker on my laptop.
- How I timed it: each setting ran all 300 questions after 30 warm-up questions, and I report the median.
- How I checked quality: for each question I know which passage answers it. "Right passage in top 10" is how often it came back.
Part 1
The five fixes
One at a time, each measured.
1. Stop checking every guest
Imagine finding your friend at a shaadi with 5,000 guests. Without a plan, you check every face, one by one. That's a search without an index: Postgres compares your question with every row, then sorts them all.
An index is the aunty who knows everyone. You ask her, she points you to someone who knows someone closer, and a few hops later you've found your friend. pgvector's HNSW index works like that: it's a graph of neighbours, and the search hops through it instead of reading every row.
The index answered about 14 times sooner, and found the right passage just as often (99.3%).
-- for vectors up to 2,000 numbers
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);
My vectors have 2,560 numbers, which is over the limit for a normal index, so I indexed a half-precision copy instead. The big embeddings guide explains that limit and the fix.
no index (reads all 5,000 rows) median 8.63 ms · right passage in top 10 99.3%
HNSW index median 0.60 ms · right passage in top 10 99.3%
The no-index time grows with every row you add, because it reads every row. In an earlier run on the same data it was 11.3 ms, so treat these as "about 9 to 11 ms", not a fixed number.
2. Ask for the samosa, not the whole kitchen
Each passage's vector is 2,560 numbers. If your query asks for the vector back, every answer carries 10 × 2,560 numbers that your app almost never uses. That's ordering one samosa and having the waiter bring the whole kitchen.
-- slow: sends every vector back
SELECT id, embedding FROM docs ORDER BY embedding <=> $1 LIMIT 10;
-- fast: sends back only what you'll show
SELECT id, title FROM docs ORDER BY embedding <=> $1 LIMIT 10;
SELECT id median 0.60 ms
SELECT id, embedding median 2.74 ms (4.5x slower)
Watch out for SELECT * too. It quietly includes the vector column.
3. Stop pressing "1 for English" every time
Opening a database connection is like calling customer care: before anything useful happens, you sit through "Press 1 for English". If your app opens a new connection for every question, it pays that cost every time.
new connection for every question median 6.09 ms per question (connect + set up + search)
one connection, reused median 0.60 ms per question
connection pool median 0.82 ms per question
A connection pool keeps a few connections open and lends them out. Here's the one I timed:
from pgvector.psycopg import register_vector
from psycopg_pool import ConnectionPool
pool = ConnectionPool(DATABASE_URL, min_size=2, max_size=5, configure=register_vector)
def search(question_vector):
with pool.connection() as conn:
return conn.execute(
"SELECT id FROM docs ORDER BY embedding <=> %s LIMIT 10", (question_vector,)
).fetchall()
configure=register_vector runs once per connection, not once per question. My test ran on one laptop with no network in between. Over a real network, with TLS, a new connection usually costs more, but I didn't measure that.
4. Don't ask every uncle
HNSW has a knob called ef_search: how many candidates it keeps while it hops through the graph. Think of it as how many uncles you ask before you decide. Ask too few and you might miss your friend. Ask too many and you're there all evening.
time per search
0.61 ms
right passage in top 10
99.3%
same top 10 as exact search
99.0%
The default. Here it already found the right passage as often as exact search did.
ef_search 10 median 0.35 ms · right passage in top 10 96.3%
ef_search 20 median 0.50 ms · right passage in top 10 98.7%
ef_search 40 median 0.61 ms · right passage in top 10 99.3% <- default
ef_search 100 median 0.91 ms · right passage in top 10 99.3%
ef_search 200 median 1.41 ms · right passage in top 10 99.3%
ef_search 400 median 1.98 ms · right passage in top 10 99.3%
On this data the default was the sweet spot. Lower lost answers; higher cost time and found the right passage no more often.
If you change it, change it per search, not for the whole connection, especially with a pool:
BEGIN;
SET LOCAL hnsw.ef_search = 100; -- only for this transaction
SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10;
COMMIT;
A plain SET stays on the pooled connection and quietly affects the next request that borrows it.
5. Pack light
Smaller vectors mean a smaller index and less work per hop. Some embedding models are trained so you can keep only the first part of the vector. With those, I tried 512 numbers instead of 2,560:
2,560 numbers (half precision) index 39 MB · median 0.60 ms · right passage in top 10 99.3%
1,024 numbers index 39 MB · median 0.59 ms · right passage in top 10 99.7%
512 numbers index 13 MB · median 0.45 ms · right passage in top 10 99.0%
Kam saaman, safar aasaan. Only do this with a model that supports it (Matryoshka models, like OpenAI's text-embedding-3 or Qwen3-Embedding). Cutting an ordinary model's vector throws away meaning. The big embeddings guide has the full test, including why the 1,024 index is no smaller.
Part 2
The twist
Where the time actually goes.
6. The search was the fastest part
Once the index was in, the search took 0.6 ms. The two mistakes around it cost more than the search itself:
opening a new connection for every question
fetching the vectors back with the results
the search itself (HNSW index)
So before you tune the index, check the boring parts: one connection per question, and columns you never use. They're easier to fix, and here they were worth more.
7. What I didn't test
- One laptop, no network. Postgres ran in Docker on the same machine, one question at a time. Real servers, real networks and many users at once will give different numbers.
- 5,000 rows. The no-index scan gets slower as the table grows; the 14x gap is for this size only.
- One model, one dataset. Qwen3-Embedding-4B on Natural Questions.
- No filters. A
WHEREclause next to a vector search changes the picture, and that's a test for another day. - Default index build settings.
mandef_constructionstayed at pgvector's defaults.
The whole post in seven lines
- Without an index, every question reads every row: 8.6 ms here, and growing with the table.
- An HNSW index answered in 0.6 ms with the same right-answer rate (99.3%).
- Asking for the vectors back made each search 4.5x slower. Ask only for what you show.
- A new connection per question cost 6.1 ms in total. A pool did it in 0.82 ms.
ef_search40 (the default) was the sweet spot. Change it per transaction withSET LOCAL.- With a model that supports it, 512-number vectors cut the index to a third and kept 99% of right answers.
- Once the index is in, check connections and columns before tuning the search.
Checked against pgvector 0.8.7 on PostgreSQL 17.11 (Docker image pgvector/pgvector:pg17), psycopg 3.3.6, psycopg_pool 3.3.3 and pgvector-python 0.5.0, as of October 2026. All timings and rates are real output from my run on 9 October 2026: medians of 300 questions after 30 warm-up questions (100 questions for the new-connection test). Data: 5,000 passages and 300 questions from sentence-transformers/natural-questions, embedded with Qwen3-Embedding-4B. The video's network drawing is an illustration; its numbers are from this run.