Aryan Tripathi — Writing
← All writing

October 9, 2026 · 12 min read

Big Embeddings in pgvector: The 2,000-Dimension Limit, Two Fixes, and What They Cost

Many embedding models now output more than 2,000 numbers per text, and pgvector can store them but not index them. A plain-English experiment report: the error, why it exists, half-precision indexes, shorter embeddings, and the cast that silently turns your index off. Measured with real embeddings and real questions.

#pgvector#postgresql#embeddings#rag#ai-engineering#experiment

Newer embedding models give you more numbers per text: 2,560, 3,072, even 4,096. pgvector stores vectors that big without complaint. Then you try to add an index, and it refuses. This guide shows the error, explains where the limit comes from, and measures the two standard fixes with real embeddings and real questions, including one detail that makes the first fix silently do nothing.

You don't need to know how vector indexes work inside. The vector database guide covers that if you want it. Every section ends with a one-line Remember, and there's a short summary at the end.

  • Part 1 · The wall: the error, which models hit it, and why it exists, sections 1–3.
  • Part 2 · Two fixes: a half-precision index and shorter embeddings, each measured, sections 4–7.
  • Part 3 · Choose: side-by-side results, how to pick, and what I didn't test, sections 8–10.

Part 1

The wall

What breaks, for which models, and why.

1. Storing worked. Indexing didn't.

I created a table for 3,072-dimensional embeddings, the size OpenAI's text-embedding-3-large returns by default, and loaded 10,000 of them. That worked. Then I added the index every RAG tutorial adds:

=# CREATE INDEX ON items3072 USING hnsw (embedding vector_cosine_ops);
ERROR:  column cannot have more than 2000 dimensions for hnsw index

=# CREATE INDEX ON items3072 USING ivfflat (embedding vector_cosine_ops);
ERROR:  column cannot have more than 2000 dimensions for ivfflat index

Both of pgvector's index types refuse. Then I did the same with real embeddings: 5,000 passages from the Natural Questions dataset, embedded with Qwen3-Embedding-4B, which outputs 2,560 numbers per text. Same error.

These are the limits, each checked on pgvector 0.8.7:

TypeBytes per numberCan store up toCan index (HNSW, IVFFlat) up to
vector416,000 dims2,000 dims
halfvec216,000 dims4,000 dims
bit (binary)1/8not checked64,000 dims

Pick a model to see where it lands:

Interactivepgvector 0.8.7 · HNSW and IVFFlat
01,0002,0003,0004,000vector index ≤ 2,000halfvec index ≤ 4,0003,072 dims

Too big for a vector index. A halfvec index fits, or shorten the output.

pgvector stores vectors of up to 16,000 dimensions, but its indexes stop much earlier: 2,000 for vector, 4,000 for halfvec. Pick a model to see which index it can use as-is.
Rememberpgvector stores vectors up to 16,000 dimensions but indexes plain vectors only up to 2,000.

2. Why the limit exists

The numbers give it away. Postgres stores tables and indexes in pages of 8 KB. pgvector's indexes need each vector to fit in one page with a little room to spare:

  • vector: 2,000 numbers × 4 bytes = 8,000 bytes.
  • halfvec: 4,000 numbers × 2 bytes = 8,000 bytes.

So the limit isn't a bug or an arbitrary setting. It's the page size. You can't raise it with a configuration flag; you have to make each vector take fewer bytes. There are two ways: use fewer bytes per number (Fix A), or use fewer numbers (Fix B).

RememberThe limit comes from Postgres's 8 KB page. To get under it, use fewer bytes per number or fewer numbers.

3. Do you even need the index?

Without an index, pgvector still answers. It compares the question with every row and sorts them: exact, but the work grows with the table. On the 5,000 real passages:

No index: exact scan of all 5,000 rows (vector(2560)): median 11.3 ms per query
Fix A with an HNSW index:                              median 0.59 ms per query

11 ms is fine for a demo. But a scan reads every row, so it grows with the table. At 500,000 rows you'd expect roughly 100 times the work, and the index's time grows far more slowly. (I didn't run 500,000 rows; that's the shape of the cost, not a measurement.) If your table will stay at a few thousand rows, you may not need an index at all.

RememberWithout an index, every query reads every row. Fine for thousands of rows, painful for millions.

Part 2

Two fixes

Fewer bytes per number, or fewer numbers. Both measured.

4. Fix A: index a half-precision copy

halfvec stores each number in 2 bytes instead of 4. You don't have to change your table. Keep the full vector column and build the index on a half-precision copy of it, using an expression index:

CREATE INDEX ON nq_full
  USING hnsw ((embedding::halfvec(2560)) halfvec_cosine_ops);

embedding::halfvec(2560) means "convert to halfvec". The index stores the converted values; the table keeps the originals.

The obvious worry: 2 bytes keeps about 3 significant digits, so every number moves slightly. Does that change the answers? Here are real values from one embedding, before and after:

Animatedreal values from one embedding · numpy float16

vector · float32 · 4 bytes per number

halfvec · float16 · 2 bytes per number

float32→ halfvecmoved by
0.041141970.041137704.3e-6
0.027908010.027908333.2e-7
-0.00008509-0.000085123.0e-8
0.007534860.007534038.3e-7
0.026013580.026016242.7e-6
step 3 of 3
halfvec keeps each number in 2 bytes instead of 4. It keeps about 3 significant digits, so every number moves a tiny bit. The question is whether those tiny moves change which passages come back.

To test whether those moves matter, I embedded the same 5,000 passages with a model that produces full 32-bit numbers (all-MiniLM-L6-v2), then searched with and without rounding:

top-10 overlap float32 vs halfvec: 100.0% · identical order: 99.0%
answer in top 10: float32 99.3% · halfvec 99.3%
rounding error per number: median 4.5e-06 · max 1.1e-04

Every question got the same ten passages. For 3 of the 300 questions they came back in a slightly different order. The answer rate didn't move.

On the Qwen embeddings, the HNSW index on halfvec found 99.0% of the exact top 10, and the right answer was in the top 10 for 99.3% of questions, the same as exact search. (One honest note: I ran Qwen in 16-bit mode to fit it on my laptop, so its numbers were already half-precision. That's why the MiniLM check above was needed.)

RememberA halfvec index halves the bytes per number. In my tests the rounding didn't change which passages came back.

5. The trap: the index only works with the same cast

An expression index only helps queries that use the same expression. So the query must convert in exactly the same way:

AnimatedEXPLAIN on the real table · when I ran it
SELECT id FROM nq_full
ORDER BY embedding <=> $1
LIMIT 10;
→ Sort → Seq Scan on nq_full (all 5,000 rows)

4. It computes the distance to every row and sorts them all: 11.3 ms per query on 5,000 rows. No error, no warning. The answer is even slightly more exact, which hides the problem.

step 4 of 4
An expression index only helps queries that use the same expression. Check every vector query with EXPLAIN once, and look for “Index Scan”.
with cast
   Limit  (cost=35.82..76.23 rows=10 width=12)
     ->  Index Scan using nq_half on nq_full
without cast
   Limit  (cost=202.55..202.57 rows=10 width=12)
     ->  Sort  (cost=202.55..215.05 rows=5000 width=12)
           ->  Seq Scan on nq_full

The version without the cast doesn't fail. It returns correct results (slightly more exact, even) and gets slower as the table grows. Nothing tells you the index is being ignored. The fix is a habit: run EXPLAIN on every vector query once, and look for Index Scan.

In application code, put the cast in one place, a single search function, so nobody writes the query by hand.

RememberQuery with exactly the cast the index was built on, and confirm "Index Scan" with EXPLAIN.

6. Fix B: ask the model for fewer numbers

Some models are trained so that the first numbers in the vector carry the most meaning. You can keep the first 1,024 of 2,560, re-scale the vector to length 1, and still have a useful embedding. These are called Matryoshka models, after the nesting dolls. OpenAI's text-embedding-3 models take a dimensions parameter for this, and Qwen3-Embedding supports any size from 32 to 2,560.

def shorten(x, d):
    x = x[:, :d]                                     # keep the first d numbers
    return x / np.linalg.norm(x, axis=1, keepdims=True)   # back to length 1

Then store vector(1024) and use a normal index. No casts, no special query. How much does it lose? Exact search on the real passages:

AnimatedQwen3-Embedding-4B · exact search · when I ran it
2,560 dimsanswer in top 10: 99.3%1,024 dimsanswer in top 10: 99.7%512 dimsanswer in top 10: 99.0%256 dimsanswer in top 10: 98.0%
step 4 of 4
Models trained for it (Matryoshka models) put the most important information in the first numbers. Keep the first N, re-normalise, and you get a smaller vector that still finds most of the right answers.
exact, float32, 2560 dims    recall@10 vs exact 100.0% · answer at #1 89.3% · answer in top 10 99.3%
exact, shortened to 1024     recall@10 vs exact  81.4% · answer at #1 89.0% · answer in top 10 99.7%
exact, shortened to 512      recall@10 vs exact  71.2% · answer at #1 89.3% · answer in top 10 99.0%
exact, shortened to 256      recall@10 vs exact  60.2% · answer at #1 86.0% · answer in top 10 98.0%

Only cut vectors from a model that was trained for it. Cutting an ordinary model's vector throws away meaning at random.

RememberWith a Matryoshka model, keep the first N numbers and re-normalise. You get a normal index and no cast to forget.

7. Two ways to measure "lost nothing"

Look at the 1,024 row again. Compared with the full vector's top 10, only 81% of the passages were the same. That sounds bad. But the passage that actually answers the question was in the top 10 99.7% of the time, a little more often than with the full vector.

Both numbers are real. They measure different things:

  • Recall against the full vector asks: did I get the same neighbours? My reading: shortening shifts every distance a little, so passages that were nearly tied swap in and out of the top 10, and this number drops. (I didn't check which ranks changed.)
  • Answer rate asks: did I get the right passage? That's what a RAG app cares about, and it barely moved.

So when you compare settings, measure against questions with known answers, not only against another index. Twenty or thirty real questions from your own users are enough to start.

Remember"Same neighbours as before" and "found the right answer" are different tests. For RAG, measure the second.

Part 3

Choose

Side by side, how to pick, and the limits of this test.

8. The results side by side

All three on the same 5,000 passages and 300 questions, HNSW with default settings (m = 16, ef_construction = 64, ef_search = 40):

Real output5,000 passages · 300 questions · HNSW ef_search 40 · when I ran it

halfvec(2560)

Fix A · expression index

build
2.0 s
index
39 MB
query
0.59 ms
recall@10
99.0%
answer in top 1099.3%

vector(1024)

Fix B · shortened

build
1.7 s
index
39 MB
query
0.50 ms
recall@10
81.3%
answer in top 1099.7%

vector(512)

Fix B · shortened more

build
1.0 s
index
13 MB
query
0.42 ms
recall@10
70.9%
answer in top 1099.0%
Recall is how many of the exact top 10 the index found. “Answer in top 10” is how often the passage that really answers the question came back. Exact search on the full vectors got the answer into the top 10 for 99.3% of questions.
Fix A: HNSW on embedding::halfvec(2560)
   build 2.0 s · index 39 MB
   ef_search 40   median  0.59 ms · recall@10 vs exact  99.0% · answer at #1 89.3% · answer in top 10 99.3%
Fix B: HNSW on vector(1024), model output shortened
   build 1.7 s · index 39 MB
   ef_search 40   median  0.50 ms · recall@10 vs exact  81.3% · answer at #1 89.0% · answer in top 10 99.7%
Fix B: HNSW on vector(512), model output shortened
   build 1.0 s · index 13 MB
   ef_search 40   median  0.42 ms · recall@10 vs exact  70.9% · answer at #1 89.3% · answer in top 10 99.0%

One surprise: halfvec(2560) and vector(1024) produced the same index size, 39 MB, even though one entry is 5,120 bytes and the other 4,096. Both are over half a page, so it looks like only one fits in each 8 KB page, and the index is about one page per row. At 512 dimensions (2,048 bytes), several fit per page and the index dropped to 13 MB. The 10,000 random 3,072-dimension vectors from section 1 showed the same thing: 78 MB for both fixes. Index size moves in page-sized steps, not smoothly.

Raising ef_search to 100 (search a wider part of the graph) brought Fix A's recall from 99.0% to 99.6% for about 0.25 ms more per query. It didn't change the answer rate for any fix.

RememberOn these questions, all three fixes found the right answer about as often as exact search, in about half a millisecond.

9. How to pick

Your situationPickWhy
You can't change the model output (already stored, or a model without Matryoshka training)Fix A, halfvec expression indexKeeps every number; no re-embedding
You're starting fresh with a Matryoshka modelFix B, ask for ~1,024 dimsSmaller table, normal index, no cast to forget
Over 4,000 dims (for example Qwen3-Embedding-8B at 4,096)Fix B, or binary quantizationhalfvec stops at 4,000
A few thousand rows that won't grow muchMaybe no indexAn exact scan was 11 ms here
Storage and memory matter mostFix B with fewer dims, or both together512 dims cut the index to a third

The two fixes also combine: shorten to 1,024 and store halfvec, and each vector is 2 KB instead of 10 KB.

Whatever you pick, measure it on your own questions before you trust it. My numbers come from one model and one dataset.

RememberCan't re-embed: index a halfvec copy. Starting fresh with a Matryoshka model: ask for fewer dimensions.

10. What I didn't test

  • One model, one dataset. Qwen3-Embedding-4B on 5,000 Natural Questions passages, with the questions' own answer passages as the right answers. Other models and domains can behave differently, especially when shortened.
  • Small scale. 5,000 and 10,000 rows. Build times and query times at millions of rows will be much higher; the relative picture should hold, but I didn't measure it.
  • Default HNSW settings. m and ef_construction stayed at their defaults.
  • Qwen ran in 16-bit on Apple's GPU to fit in 24 GB of memory, so the halfvec rounding check used a separate 32-bit model (section 4).
  • The 3,072-dimension run used random vectors. It shows the error and the sizes, not search quality.
RememberTreat these numbers as a method to copy, not results to quote for your system.

The whole guide in nine lines

  1. pgvector stores vectors up to 16,000 dimensions but indexes plain vectors only up to 2,000.
  2. The limit is Postgres's 8 KB page: 2,000 × 4 bytes and 4,000 × 2 bytes both come to 8,000 bytes.
  3. Without an index every query reads every row: 11 ms for 5,000 rows here, growing with the table.
  4. Fix A: an HNSW index on embedding::halfvec(N) keeps all the numbers at 2 bytes each.
  5. Query with exactly the same cast, or Postgres silently ignores the index. Check with EXPLAIN.
  6. Fix B: with a Matryoshka model, keep the first 1,024 numbers, re-normalise, and use a normal index.
  7. Shortened vectors shared only 81% of the full top 10, but found the right answer 99.7% of the time.
  8. All fixes answered in about half a millisecond and matched exact search on the answer rate.
  9. Measure with questions that have known answers, on your own data.

Checked against pgvector 0.8.7 on PostgreSQL 17.11 (Docker image pgvector/pgvector:pg17), sentence-transformers 6.1.0, Qwen3-Embedding-4B (run in float16 on an Apple M5 with 24 GB) and all-MiniLM-L6-v2, as of October 2026. All errors, plans, sizes, timings and rates are real output from my runs on 9 October 2026. The 3,072-dimension run in sections 1 and 8 used 10,000 random unit vectors; everything else used 5,000 passages and 300 questions from the Natural Questions dataset (sentence-transformers/natural-questions). The dimension limits were checked by creating indexes one dimension over and at each limit.