Aryan Tripathi — Writing
← All writing

October 9, 2026 · 8 min read

Is Your Vector Search Slow? 5 Fixes, Measured

Five common reasons vector search in Postgres feels slow, each tested on 5,000 real passages and 300 real questions: no index, fetching the vectors back, a new connection per question, the ef_search knob, and big vectors. Plain English, Indian memes, real numbers.

#pgvector#postgresql#vector-search#performance#rag#experiment

Vector search is the "find similar text" step inside AI apps, like a chatbot that answers from your documents. When it feels slow, the search itself is often not the problem. I tested five common fixes on the same data and timed each one. Some were obvious. One cost ten times more than the search.

Here's the short version as a video. Turn the sound on.

Video37 s · turn the sound on
The five fixes as a meme video. Every number in it comes from the measurements in this post. The network drawing in fix 1 is an illustration, not the real index.

Every section ends with a one-line Remember, and there's a summary at the end.

The setup

  • Data: 5,000 passages from the Natural Questions dataset, and 300 real questions whose answers are in those passages.
  • Embeddings: Qwen3-Embedding-4B, 2,560 numbers per passage.
  • Database: PostgreSQL 17 with pgvector 0.8.7, in Docker on my laptop.
  • How I timed it: each setting ran all 300 questions after 30 warm-up questions, and I report the median.
  • How I checked quality: for each question I know which passage answers it. "Right passage in top 10" is how often it came back.
RememberFive fixes, one dataset, one laptop: same questions every time, so the numbers compare fairly.

Part 1

The five fixes

One at a time, each measured.

1. Stop checking every guest

Imagine finding your friend at a shaadi with 5,000 guests. Without a plan, you check every face, one by one. That's a search without an index: Postgres compares your question with every row, then sorts them all.

An index is the aunty who knows everyone. You ask her, she points you to someone who knows someone closer, and a few hops later you've found your friend. pgvector's HNSW index works like that: it's a graph of neighbours, and the search hops through it instead of reading every row.

Animatedmedian of 300 questions · 5,000 rows · when I ran it
No index: check every row8.6 ms
HNSW index: hop between neighbours0.6 ms

The index answered about 14 times sooner, and found the right passage just as often (99.3%).

step 2 of 2
Both searches start together. The bar length is the time each one took. The scan has to read every row, so it gets slower as the table grows; the index doesn't.
-- for vectors up to 2,000 numbers
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);

My vectors have 2,560 numbers, which is over the limit for a normal index, so I indexed a half-precision copy instead. The big embeddings guide explains that limit and the fix.

no index (reads all 5,000 rows)   median 8.63 ms · right passage in top 10 99.3%
HNSW index                        median 0.60 ms · right passage in top 10 99.3%

The no-index time grows with every row you add, because it reads every row. In an earlier run on the same data it was 11.3 ms, so treat these as "about 9 to 11 ms", not a fixed number.

RememberNo index means reading every row on every question. Add one before you tune anything else.

2. Ask for the samosa, not the whole kitchen

Each passage's vector is 2,560 numbers. If your query asks for the vector back, every answer carries 10 × 2,560 numbers that your app almost never uses. That's ordering one samosa and having the waiter bring the whole kitchen.

-- slow: sends every vector back
SELECT id, embedding FROM docs ORDER BY embedding <=> $1 LIMIT 10;

-- fast: sends back only what you'll show
SELECT id, title FROM docs ORDER BY embedding <=> $1 LIMIT 10;
SELECT id              median 0.60 ms
SELECT id, embedding   median 2.74 ms   (4.5x slower)

Watch out for SELECT * too. It quietly includes the vector column.

RememberAsk only for the columns you'll show. Never send the vectors back unless you need them.

3. Stop pressing "1 for English" every time

Opening a database connection is like calling customer care: before anything useful happens, you sit through "Press 1 for English". If your app opens a new connection for every question, it pays that cost every time.

new connection for every question   median 6.09 ms per question (connect + set up + search)
one connection, reused              median 0.60 ms per question
connection pool                     median 0.82 ms per question

A connection pool keeps a few connections open and lends them out. Here's the one I timed:

from pgvector.psycopg import register_vector
from psycopg_pool import ConnectionPool

pool = ConnectionPool(DATABASE_URL, min_size=2, max_size=5, configure=register_vector)

def search(question_vector):
    with pool.connection() as conn:
        return conn.execute(
            "SELECT id FROM docs ORDER BY embedding <=> %s LIMIT 10", (question_vector,)
        ).fetchall()

configure=register_vector runs once per connection, not once per question. My test ran on one laptop with no network in between. Over a real network, with TLS, a new connection usually costs more, but I didn't measure that.

RememberOpen connections once and reuse them. A pool did the job in 0.82 ms instead of 6.1 ms.

4. Don't ask every uncle

HNSW has a knob called ef_search: how many candidates it keeps while it hops through the graph. Think of it as how many uncles you ask before you decide. Ask too few and you might miss your friend. Ask too many and you're there all evening.

Interactivehnsw.ef_search · 300 questions · when I ran it
102040100200400

time per search

0.61 ms

right passage in top 10

99.3%

same top 10 as exact search

99.0%

The default. Here it already found the right passage as often as exact search did.

ef_search is how many candidates the index keeps while it hops through the graph. More candidates means more work. Exact search on the same data found the right passage 99.3% of the time.
ef_search 10    median 0.35 ms · right passage in top 10 96.3%
ef_search 20    median 0.50 ms · right passage in top 10 98.7%
ef_search 40    median 0.61 ms · right passage in top 10 99.3%   <- default
ef_search 100   median 0.91 ms · right passage in top 10 99.3%
ef_search 200   median 1.41 ms · right passage in top 10 99.3%
ef_search 400   median 1.98 ms · right passage in top 10 99.3%

On this data the default was the sweet spot. Lower lost answers; higher cost time and found the right passage no more often.

If you change it, change it per search, not for the whole connection, especially with a pool:

BEGIN;
SET LOCAL hnsw.ef_search = 100;  -- only for this transaction
SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10;
COMMIT;

A plain SET stays on the pooled connection and quietly affects the next request that borrows it.

RememberLeave ef_search at 40 until your own questions show a reason to move it.

5. Pack light

Smaller vectors mean a smaller index and less work per hop. Some embedding models are trained so you can keep only the first part of the vector. With those, I tried 512 numbers instead of 2,560:

2,560 numbers (half precision)   index 39 MB · median 0.60 ms · right passage in top 10 99.3%
1,024 numbers                    index 39 MB · median 0.59 ms · right passage in top 10 99.7%
  512 numbers                    index 13 MB · median 0.45 ms · right passage in top 10 99.0%

Kam saaman, safar aasaan. Only do this with a model that supports it (Matryoshka models, like OpenAI's text-embedding-3 or Qwen3-Embedding). Cutting an ordinary model's vector throws away meaning. The big embeddings guide has the full test, including why the 1,024 index is no smaller.

RememberIf your model supports shorter vectors, 512 numbers cut the index to a third here and still found the right passage 99% of the time.

Part 2

The twist

Where the time actually goes.

6. The search was the fastest part

Once the index was in, the search took 0.6 ms. The two mistakes around it cost more than the search itself:

Real outputextra time per question · when I ran it

opening a new connection for every question

+5.5 ms

fetching the vectors back with the results

+2.1 ms

the search itself (HNSW index)

0.6 ms
Each mistake was measured on its own and compared with the same 0.6 ms indexed search. These are separate tests, not slices of one request, but they show where the time goes once the index is in.

So before you tune the index, check the boring parts: one connection per question, and columns you never use. They're easier to fix, and here they were worth more.

RememberAfter the index, look at connections and columns before you touch the search settings.

7. What I didn't test

  • One laptop, no network. Postgres ran in Docker on the same machine, one question at a time. Real servers, real networks and many users at once will give different numbers.
  • 5,000 rows. The no-index scan gets slower as the table grows; the 14x gap is for this size only.
  • One model, one dataset. Qwen3-Embedding-4B on Natural Questions.
  • No filters. A WHERE clause next to a vector search changes the picture, and that's a test for another day.
  • Default index build settings. m and ef_construction stayed at pgvector's defaults.
RememberCopy the method, not the numbers: time your own questions on your own setup.

The whole post in seven lines

  1. Without an index, every question reads every row: 8.6 ms here, and growing with the table.
  2. An HNSW index answered in 0.6 ms with the same right-answer rate (99.3%).
  3. Asking for the vectors back made each search 4.5x slower. Ask only for what you show.
  4. A new connection per question cost 6.1 ms in total. A pool did it in 0.82 ms.
  5. ef_search 40 (the default) was the sweet spot. Change it per transaction with SET LOCAL.
  6. With a model that supports it, 512-number vectors cut the index to a third and kept 99% of right answers.
  7. Once the index is in, check connections and columns before tuning the search.

Checked against pgvector 0.8.7 on PostgreSQL 17.11 (Docker image pgvector/pgvector:pg17), psycopg 3.3.6, psycopg_pool 3.3.3 and pgvector-python 0.5.0, as of October 2026. All timings and rates are real output from my run on 9 October 2026: medians of 300 questions after 30 warm-up questions (100 questions for the new-connection test). Data: 5,000 passages and 300 questions from sentence-transformers/natural-questions, embedded with Qwen3-Embedding-4B. The video's network drawing is an illustration; its numbers are from this run.