pg_turbovec

Open-source vector similarity search for PostgreSQL, backed by Google Research’s TurboQuant algorithm via the turbovec Rust crate.

Store compact vector indexes alongside your PostgreSQL data: 1-bit sign-BQ with full-precision reranking, or 2/¾-bit TurboQuant. The original vectors stay in the table for reranking. Supports:

  • approximate nearest-neighbour search with full-precision candidate reranking; exact search through PostgreSQL’s distance operators and a sequential scan
  • 1-bit + rerank, including IVF: WITH (bit_width = 1, lists = N); see the 1-bit example
  • three index kinds: a flat quantized scan (exact-recall- capable), an opt-in IVF layer (WITH (lists = N)) that is out-of-core end-to-end so a larger-than-RAM index can be built and queried (a Vamana navigable-graph kind, WITH (graph = true), also exists but is deprecated — see below)
  • single-precision (f32) vectors with on-disk compression to ~16× smaller than pgvector at 4-bit
  • L2 distance (<->), inner product (<#>), cosine distance (<=>), L1 distance (<+>)
  • filtered / hybrid ANN — partial index, in-kernel allowlist (a selective filter gets cheaper, not more expensive), or iterative scan (guide)
  • multivector & dense+sparse hybrid — ColBERT-style MaxSim re-rank (max_sim) + reciprocal rank fusion (rrf_score) + named-vector schema pattern (guide)
  • pgvector-compatible function names (to_vector, array_to_vector, subvector, vector_dims, vector_norm, inner_product, l2_distance, cosine_distance, l1_distance)
  • any language with a Postgres client

Plus ACID compliance, point-in- time recovery, JOINs, GUCs, parallel-safe aggregates, and all of the other great features of Postgres.

Rust 1.96+ PostgreSQL 13-19 Apache 2.0

Status: v2.11.0 - built on upstream turbovec 1.1.1 (wire format v8; staged 2/4-bit search on aarch64 and AVX-512 VBMI+VNNI hosts, 1.14-1.17x faster 4-bit end-to-end on Graviton4 at 1M x 1024-d – see CHANGELOG). The full #[pg_test] suite passes against PostgreSQL 13, 14, 15, 16, 17, and 18 (and 19beta1, experimentally). v2.0.0 is a MAJOR wire-format break (v7 → v8): upgrading from any 1.x requires ALTER EXTENSION ... UPDATE then a one-time REINDEX INDEX per turbovec index (a pre-v8 index ERRORs at first scan with a REINDEX hint — never silent). It is materially faster than the 1.29 line at the same storage. See docs/UPGRADING.md for the migration and docs/PARITY_GAPS.md for the honest scoreboard vs pgvector.

Why pg_turbovec?

On 1 M × 1536-d real OpenAI embeddings, pg_turbovec matches pgvector HNSW’s recall at ~10–20× less on-disk storage, with exact re-ranking against the heap. Storage efficiency and exact/near-exact recall are what pg_turbovec does better than anything else — not raw latency (see the honest latency note below).

Head-to-head, warm cache, release build. Storage + recall (measured on real embeddings):

metric (1 M × 1536-d, real OpenAI) pg_turbovec 2-bit pgvector HNSW
On-disk index / vector ≈ 412 B (measured) ≈ 8 192 B
Storage @ 1 M ≈ 412 MB ≈ 8 GB (~20× larger)
Build (1 M) ~minutes ~5 min
Recall@10 (IVF, tuned) ~0.90 (R@100 → 0.99) 0.96–0.99
Exact re-ranking vs heap ✓ (xs_recheckorderby) ✓

Common objections, answered with measurements

Four things evaluators say about pg_turbovec. Two are misunderstandings we caused, one is true-and-fixed, one is true-and-stands.

“pg_turbovec doesn’t support ANN”

It does. The default (lists = 0) scans all quantized codes, selects a candidate set, and reranks candidates using the original vectors. Quantization can exclude a true neighbour from that set, so even this flat index is an approximate search. IVF additionally restricts candidate search to selected cells:

-- Flat quantized search with full-precision candidate reranking (DEFAULT).
CREATE INDEX ON items USING turbovec (embedding vec_cosine_ops);

-- Approximate, cell-pruned IVF scan. This is the ANN path.
CREATE INDEX ON items USING turbovec (embedding vec_cosine_ops)
  WITH (lists = 1000);

lists defaults to 0, which means flat. Nothing in a default build tells you IVF exists, which is a documentation failure on our side, not a misreading on yours. Since v2.10.2 a flat build over 100 k rows emits a NOTICE pointing at the ANN option. See Choosing lists for how to pick a value — and note that on our own measurements, flat is the right choice more often than you would expect (that section says when).

“HNSW has a high memory footprint and is slow to build”

Agreed — that is why we did not build on HNSW. Measured head-to-head on the same host, 10 M × 1536-d real embeddings (docs/RECALL.md):

pgvector HNSW (m=16, ef_construction=64) pg_turbovec 4-bit
Index size 65.5 GiB 14.9 GiB (4.4× smaller)
Build time 3 h 38 min 1 h 24 min (2.6× faster)

Our build log also shows HNSW slowing super-linearly past 5 M rows even with the whole corpus in page cache, which matches the common complaint.

“pg_turbovec’s own build memory was worse than HNSW’s”

This was true, and it is now fixed. The same 10 M × 1536-d benchmark measured pg_turbovec’s build at 121 GiB peak + 60 GiB swap against HNSW’s 16.9 GiB — we were ~7× worse. That figure is from v1.5.1 (May 2026) and is still quoted in docs/RECALL.md for the run it belongs to.

Root cause (found and fixed in v2.10.1): vector is stored as CBOR, so FromDatum palloc’d a decoded buffer per row, and the build callback had no memory-context management at all — every decoded row accumulated for the whole scan. The spill was streaming the corpus to disk while PostgreSQL held a decoded copy of all of it in RAM.

Measured after the fix:

build before after
2 M × 1024-d peak 12.16 GiB 3.45 GiB (3.5× less, 16 % faster)
10 M × 1024-d on a 61 GiB host OOM-killed (anon-rss 58.5 GiB) completes at 11.20 GiB in 70 min

Scope, stated plainly: the post-fix numbers are 1024-d. We have not re-measured 10 M × 1536-d, which is what the 121 GiB figure was, so treat that specific number as superseded-in-mechanism but not yet re-measured at its original dimension.

“HNSW is faster on query latency”

True, and we don’t dispute it. A flat pg_turbovec scan is O(n·dim); HNSW is sublinear. If p50 latency is your binding constraint and your corpus fits in RAM, pgvector HNSW is the better choice today — see the honest latency note immediately below, which we keep at the top of this README on purpose.

Where we are measurably better: storage (4.4–20× smaller), build time (2.6× faster), larger-than-RAM corpora (IVF is out-of-core end-to-end), and exact recall when you need R@10 = 1.000 rather than 0.96.


Choosing lists — the ANN tuning knob

lists is the IVF coarse-cell count (nlist). The default is 0, which means a flat quantized scan, not guaranteed exact top-k. We keep it at 0 deliberately: on our own measurements, enabling IVF at the default bit_width = 4 makes things worse up to at least 1 M rows.

First: do you want IVF at all?

At 1 M × 1024-d with the default bit_width = 4 (docs/BQ_RECALL_BENCH.md § 0.6g):

config latency recall@10
flat (lists = 0, the default) 6.08 ms 1.000
lists = 1024 (≈ √n) 16.04 ms (62 % slower) capped at 0.959

That recall cap is the important half: at probes = 128, widening the rerank window from 32 to 2000 left recall at exactly 0.959 across all eight windows. It is a per-probe ceiling, it is CPU-independent, and no amount of tuning removes it. So IVF is not a free “make it faster” switch — it trades a recall ceiling for a smaller scan.

Our measured decision rule:

your situation use
bit_width ≥ 2 (incl. the default 4), n ≲ 1 M lists = 0 (flat) — faster and exact
bit_width = 1, n ≳ 1 M, target recall ≲ 0.95 lists ≈ √n — measured 38–47 % faster
bit_width = 1, target recall ≳ 0.98 lists = 0 — IVF cannot reach it
index larger than RAM lists ≈ √n — IVF is the only out-of-core query path
n ≫ 1 M with bit_width ≥ 2 measure it — we have no data above 1 M here

Second: if you do want it, pick lists ≈ √n

-- 1 M rows -> lists = 1000;  10 M -> 3162;  100 M -> 10000.
SELECT round(sqrt(count(*)))::int AS suggested_lists FROM items;

CREATE INDEX ON items USING turbovec (embedding vec_cosine_ops)
  WITH (lists = 1000);            -- substitute the value above

√n is a ceiling, not a target. At 1 M we measured lists = 4096 (4× √n) as worse than lists = 1024 on every axis: 11× the build time and ~50 % higher latency. Each cell holds proportionally fewer rows, so a given recall needs proportionally more probes — you pay twice. If in doubt, go under √n.

Third: tune probes at query time, not lists

lists is baked in at build time; probes is a per-session GUC, so this is the knob to sweep:

SET turbovec.probes = 16;          -- cells scanned per query (default 16)
SET turbovec.oversample = 2.0;     -- widen the exact-rerank window

Recall rises with probes and saturates at the per-probe ceiling described above. Sweep probes (8, 16, 32, 64, …) against your own ground truth and stop at the first value that meets your recall target; everything beyond that is latency you are paying for nothing.

Fourth: measure on your own data

Our 1 M and 250 k figures come from different corpora, so cross-scale deltas in our docs are suggestive rather than measured. Your corpus is the only authority for your workload. docs/PRODUCTION.md has a 20-minute experiment that settles flat-vs-IVF on your data, comparing at matched recall (comparing at matched probes instead is the classic way to get a flattering, wrong answer).

Honest latency note (read this)

v2.11.0 (turbovec 1.1.1): on aarch64 (Graviton, Ampere, Apple) and x86 with AVX-512 VBMI+VNNI, flat 4-bit queries are 1.14–1.17× faster end-to-end at 1M × 1024-d (Graviton4: 7.88 → 6.72 ms at search_k=32), recall unchanged; the scan kernel itself is 3–6.5× faster, so the per-candidate heap recheck is now most of a query’s cost. AVX2-only x86 is unchanged. docs/BENCHMARKS.md. v2.11.0 also fixes a silent index-entry corruption under concurrent writes + VACUUM present since v1.29.1 – upgrade, then REINDEX indexes that took writes from long-lived connections. See CHANGELOG.

pg_turbovec’s flat kind is an O(n·dim) quantized full scan — at 1 M rows its warm p50 is ~2.5 s on AVX2, and pgvector HNSW (a sublinear graph) is ~490× faster on latency. We do not beat HNSW on latency, and any earlier README claim to the contrary was a pre-AVX2-fix measurement artifact and is retracted (see docs/PARITY_GAPS.md). The latency paths are:

  • IVF (WITH (lists = N)) — cell-pruned scan, out-of-core capable. Measured ~17–24 ms warm p50 at 1 M × 1536-d (R@10 0.84–0.89), the practical latency config.
  • Vamana graph (WITH (graph = true)) — DEPRECATED, scheduled for removal. It was added to chase HNSW’s query latency while keeping TurboQuant’s compression, but measured at matched recall it never delivers: its apparent sublinearity holds only at iso-beam, and once recall is held equal the curves diverge rather than cross. Use the default flat index below ~1 M vectors, or WITH (lists = N) (IVF) at scale.

    corpus / target flat IVF graph
    SIFT-1M/128d, R@10 ≥0.95 0.98 ms, qps@8 1380 1.8 ms, qps@8 2039 26.2 ms, qps@8 299
    GIST-1M/960d, R@10 ≥0.95 5.88 ms, qps@8 279 11.3 ms, qps@8 480 unreachable
    GIST-10M/960d, R@10 ≥0.98 34.2 ms, qps@8 31 28.4 ms, qps@8 161 unreachable

    It also loses on build time (57–90×), storage, and has no out-of-core path. It is IVF, not the graph, that beats flat’s O(n) wall.

Use pg_turbovec when your workload is cosine / inner-product semantic search that is storage-constrained and you want full-precision re-ranking and Postgres-native ACID/joins — and you use the IVF kind (WITH (lists = N), not a bare flat scan) for latency at scale. If you need raw HNSW latency or ANN over halfvec/sparse representations, pgvector + HNSW is the right pick today.

Feature breakdown:

Feature pg_turbovec pgvector + HNSW
Storage / 1536-dim row (2-bit) ≈ 412 B (measured) 8 192 B (measured)
Latency at 1 M+ (raw) slower (flat O(n); IVF/graph prune) faster (sublinear HNSW)
Index kinds flat, IVF (out-of-core), ColBERT, Vamana graph HNSW, IVFFlat
Filtered search In-kernel SIMD allowlist Post-filter
Index AM lifecycle CREATE / CIC / aminsert / ambulkdelete / VACUUM / REINDEX same
Distance ops indexed <#> <=> <-> <+> (turbovec kernel) <-> <#> <=> <+> (HNSW + IVF)
L2 / L1 distance ANN indexed (vec_l2_ops / vec_l1_ops) indexed via HNSW
Halfvec, sparsevec, bitvec types ✓ (exact ops; not AM-indexable) ✓
License Apache-2.0 PostgreSQL

Methodology: recall numbers use exact top-k ground truth; see docs/RECALL.md. Latency numbers are contention- controlled AVX2 warm p50 (docs/PARITY_GAPS.md, docs/BENCHMARKS.md).

Why pg_turbovec instead of binary_quantize() + bit_hamming_ops?

If you’ve reached for pgvector’s binary_quantize() + bitvec + bit_hamming_ops HNSW index, the reason is almost always memory pressure: “I have 100 M × 1536-dim embeddings, FP32 doesn’t fit, I’ll trade recall for 32× compression.”

At the same byte budget, pg_turbovec’s 2-bit mode wins on recall. Measured numbers from benches/results/recall_dbpedia_1M_2026_05_24.json on 1 M × 1536-d OpenAI ada-002 embeddings; the 1-bit Hamming line is the upper bound from the upstream pgvector docs since we don’t have a direct 1 M-row measurement on the same corpus.

Approach Bytes / 1536-dim row R@10 (real OpenAI ada-002 embeddings)
FP32 (raw vector) 6 144 1.00 (ground truth)
FP16 (halfvec) 3 072 ≈ 1.00
TurboQuant 4-bit (turbovec index, default) 780 (measured payload / 1 M rows) 1.000 (search_k = 100)
TurboQuant 2-bit (turbovec index) 396 (measured payload / 1 M rows) 1.000 (search_k = 100)
1-bit + Hamming HNSW (pgvector bit_hamming_ops) 192 ≈ 0.65-0.75 (literature)

The 4-bit / 2-bit numbers come from the 50-query head-to-head sweep in docs/RECALL.md § 2.2; the synthetic random-vector measurements in docs/RECALL.md § 2.1 are deliberately pessimistic because random points have no clustering structure to exploit - the dbpedia run shows what real embedding geometry buys you.

Why does Lloyd-Max scalar quantization beat 1-bit thresholding at the same byte count? TurboQuant first rotates the input by a fixed orthogonal matrix so that, after rotation, each coordinate independently follows a known Beta distribution that converges to N(0, 1/d). It then assigns buckets via Lloyd-Max scalar quantization - provably the distortion-rate-optimal scalar code for that distribution - and packs them at 2, 3, or 4 bits per coordinate. 1-bit thresholding (pgvector’s binary_quantize()) is the same idea pinned to bit_width = 1: it keeps the sign and throws the magnitude away. At 2 bits, Lloyd-Max with 4 reconstruction levels lands materially closer to the Shannon distortion-rate lower bound than a 2-bucket sign threshold can - so pg_turbovec at bit_width = 2 occupies essentially the same byte budget as 1-bit Hamming with strictly higher recall. See the TurboQuant paper, arXiv:2504.19874 for the full distortion analysis.

When is bit_hamming_ops still the right tool? When the bit vector is the data, not a compression of an f32 vector - i.e. native binary embeddings (Cohere’s binary mode), perceptual / image fingerprints (pHash, dHash for near-duplicate detection), and SimHash / MinHash signatures over text shingles. For those workloads pgvector’s HNSW on bit_hamming_ops is the right tool and there is no reason to use pg_turbovec instead.

Choose your bit_width

1-bit search with full-precision reranking

pg_turbovec supports 1-bit binary quantization followed by reranking. bit_width = 1 uses centered sign-BQ in pg_turbovec, separate from upstream TurboQuant’s 2/¾-bit codec. Hamming distance selects candidates; PostgreSQL fetches their original vectors from the heap and recomputes the requested distance. Reranking is automatic on the index-AM query path: no separate binary-vector column or manual reranking CTE is required.

-- Assuming items(id, embedding) contains turbovec.vector values.
CREATE INDEX items_embedding_bq ON items
  USING turbovec (embedding turbovec.vec_cosine_ops)
  WITH (bit_width = 1, lists = 0);

BEGIN;
SET LOCAL turbovec.search_k = 800;  -- example candidate budget; tune for your data
SET LOCAL turbovec.oversample = 1.0;
SELECT id FROM items
ORDER BY embedding OPERATOR(turbovec.<=>)
  (SELECT embedding FROM items WHERE id = 42)
LIMIT 10;
COMMIT;

Use EXPLAIN to confirm index use. search_k controls the initial candidate budget, oversample multiplies it, and hi_dim_rerank may raise its floor. Exact candidate distances do not guarantee exact top-k recall: reranking cannot recover a neighbour omitted by Hamming selection.

WITH (bit_width = 1, lists = 1000) also supports IVF with reranking. Tune probes for cell coverage as well as the candidate budget; increasing only search_k cannot recover neighbours in unprobed cells. Keep lists = 0 unless measurements justify IVF. See the storage/recall trade-offs below.

Workload Recommended Storage / 1536-dim (measured) R@10 (1 M dbpedia)
Want pgvector-equivalent recall, halve storage halfvec (no quantization) 3 072 B ≈ 1.0
RAG / semantic search, R@10 ≥ 0.95 acceptable bit_width = 4 (default) 780 B 1.000
Memory pressure dominates, R@10 ≥ 0.85 acceptable bit_width = 2 396 B 1.000
Replacing binary_quantize() + bit_hamming_ops bit_width = 2 (strictly better) 396 B 1.000 (vs 0.65-0.75)
Absolute minimum storage, latency has slack bit_width = 1 (sign BQ, opt-in) 192 B 0.967 @ w=256 (measured, 1024-d)

bit_width = 1 (new in v2.6.0) is a different scheme from the 2/¾-bit TurboQuant path: it keeps only each coordinate’s sign (after subtracting the corpus mean), scores by Hamming, and relies on the exact re-rank for accuracy. It is the smallest option — dim/8 bytes and no per-vector scale — and deliberately opt-in, never a default: it is lossy enough that turbovec.hi_dim_rerank widens the exact-rerank window for it at any dimension. A corpus whose vectors all share one sign pattern even after centering is rejected at build rather than silently returning arbitrary rows. Measured on 250k x 1024-d Cohere-wiki (AVX2 host, 100 held-out queries, exact ground truth): storage is 3.98x smaller than 4-bit and 2.02x smaller than 2-bit, but at matched recall it costs 2.7-6.1x the latency and needs a 25x wider exact-rerank window than 2-bit to clear R@10 >= 0.99 (window 800 vs 32). Use it where storage is the binding constraint and latency has slack – never as a default.

1-bit is a HIGH-DIMENSION technique. A dim sweep (256/512/1024-d, same corpus) shows the penalty collapsing as dimension rises: the rerank window it needs versus 2-bit for R@10 >= 0.95 goes 125x (256-d) -> 25x (512-d) -> 8x (1024-d), and its storage edge improves too (1.90x -> 1.97x vs 2-bit). At 256-d it needs to rerank 6.4% of the corpus to reach R@10 >= 0.99, which makes it effectively unusable at 256-d and below. Prefer it at 768-d and up.

The storage, recall and latency ratios are all confirmed: the original timings were taken on a host that could not reach the harness’s load gate, so the sweep was re-run after that was fixed. Recall reproduced exactly, the contended p50s proved uniformly 14-16% pessimistic, and the ratios held to two decimal places (2.70 -> 2.75x and 6.13 -> 6.09x). Absolute milliseconds quoted below are therefore ~15% conservative. Full curve, the resolution, and the pre-registered predictions: docs/BQ_RECALL_BENCH.md § 0. It composes with lists = N (IVF): WITH (lists = N, bit_width = 1) stores the sign codes cell-contiguous and probes only turbovec.probes cells, combining the storage win with the scan win.

When IVF actually pays for 1-bit (measured on a real 1M x 1024-d Cohere corpus, AVX-512): at 1M rows IVF beats flat by 47% at R@10 >= 0.90 and 38% at R@10 >= 0.95 – but it CANNOT reach R@10 >= 0.99 at all, because cell-restricted search caps recall per probe count (0.986 max at probes=128). At 250k rows flat wins at every target. And for bit_width >= 2 flat wins everywhere up to 1M, because 2-bit only needs a 32-wide rerank window so its full scan is already cheap. So: 1-bit + n >= ~1M + target <= ~0.95 -> use lists = N; otherwise flat.

That is guidance about when IVF is worth enabling, not about what is supported. WITH (lists = N) composes with every bit_width and always has – 4-bit IVF is the original IVF path, out-of-core end-to-end since v1.13.0. The only combination the code rejects is bit_width = 1 with graph = true. If you are on 4-bit and want IVF, nothing is blocking you; see docs/BQ_RECALL_BENCH.md 0.6g for what to expect and what has actually been measured. Note an INSERT into an existing IVF+BQ index appends rather than placing the row in its cell, which degrades that index to a flat Hamming scan until the next REINDEX — reportable via turbovec.index_is_degraded().

Measured storage and recall come from the head-to-head sweep on 1 M × 1536-d OpenAI ada-002 embeddings; methodology and the synthetic random-vector numbers (§ 2.1) live in docs/RECALL.md. For dimensions other than 1536, multiply storage through by dim / 1536.

Installation

pg_turbovec requires PostgreSQL 13–18 (19beta1 experimental) and a Rust toolchain ≥ 1.96 (pgrx 0.19’s MSRV; the default stable toolchain works). PostgreSQL 16 is the reference development platform; 13/14/15/17/18/19 are tested in CI.

# One-time setup.
cargo install --locked cargo-pgrx --version 0.19.1
cargo pgrx init                # bootstraps a private PostgreSQL cluster

# Build & install into the dev cluster.
git clone https://codeberg.org/gregburd/pg_turbovec
cd pg_turbovec
cargo pgrx install --release   # default features include the index AM

# Or build a stripped-down variant without the index AM:
cargo pgrx install --release --no-default-features --features pg16

Nix flake

The repo ships a flake with one package per supported PostgreSQL major (pg_turbovec_13 … pg_turbovec_19; default = PG18):

# Build the PG18 extension:
nix build github:gburd/pg_turbovec#pg_turbovec_18
# result/lib/pg_turbovec.so
# result/share/postgresql/extension/pg_turbovec--2.10.2.sql + .control

On PostgreSQL 18+ you can point a stock server at the store path without copying files (PG18 added extension_control_path):

extension_control_path = '$system:/nix/store/…-pg_turbovec-X.Y.Z/share/postgresql'
dynamic_library_path   = '$libdir:/nix/store/…-pg_turbovec-X.Y.Z/lib'

For PG13–17, copy (or symlink) lib/pg_turbovec.so into pg_config --pkglibdir and share/postgresql/extension/* into pg_config --sharedir/extension, or overlay the extension into nixpkgs' postgresql_NN.pkgs. A dev shell with the full toolchain (rust, cargo-pgrx 0.19.1, clang/bindgen, openblas, bison/flex/readline/icu) is available via nix develop.

For a Nix-based build (the dev environment for this project) see docs/BUILDING.md. For migrating from pgvector see docs/MIGRATING_FROM_PGVECTOR.md.

Getting Started

Enable the extension (do this once per database):

CREATE EXTENSION pg_turbovec;
SET search_path = public, turbovec;

Add pg_turbovec to shared_preload_libraries. The turbovec.* GUCs (probes, search_k, iterative_scan, out_of_core, …) are registered when the library loads. Without preloading, a session has none of them: SET turbovec.probes = 16 is silently accepted and does nothing, and every query runs at the defaults no matter what you tune.

ALTER SYSTEM SET shared_preload_libraries = 'pg_turbovec';  -- then restart

Verify (expect a non-zero count, not 0):

SELECT count(*) FROM pg_settings WHERE name LIKE 'turbovec.%';

Indexes and queries work without preloading — only the tuning knobs disappear, which makes the failure quiet and easy to misdiagnose. This cost a full benchmark round on 2026-09-22 before it was spotted.

Create a table with a vector column:

CREATE TABLE items (
    id        bigserial PRIMARY KEY,
    body      text,
    embedding vector
);

Insert vectors:

INSERT INTO items (body, embedding) VALUES
  ('hello',  '[0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8]'),
  ('world',  '[0.2, 0.1, 0.4, 0.3, 0.6, 0.5, 0.8, 0.7]');

-- Or via array cast:
INSERT INTO items (body, embedding)
VALUES ('greeting',
        ARRAY[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]::real[]::vector);

-- Or via the pgvector-style function:
INSERT INTO items (body, embedding)
VALUES ('hi',  to_vector('[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]', 8, false));

Get the nearest neighbours by cosine distance:

SELECT id, body, embedding <=> '[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]'::vector AS dist
FROM   items
ORDER  BY embedding <=> '[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]'::vector
LIMIT  5;

Also supports inner product (<#>), L2 (<->), and L1 (<+>). <#> returns the negative inner product so ORDER BY ... ASC returns most-similar-first - same convention as pgvector.

Storing

-- Variable dimension:
CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector);

-- Or with a runtime dim assertion via a CHECK constraint
-- (typmod-style enforcement without typmod plumbing):
CREATE TABLE items (
    id        bigserial PRIMARY KEY,
    embedding vector,
    CHECK (turbovec.vec_check_dim(embedding, 1536))
);

-- Or add to an existing table:
ALTER TABLE items ADD COLUMN embedding vector;

vector accepts dim 1..16000 (matching pgvector’s cap). The TurboQuant kernel additionally requires dim be a multiple of 8 when used with the index AM; pad your embeddings if your model emits something awkward (e.g. 384 = 48 × 8 ✓, 1536 = 192 × 8 ✓).

Insert vectors in bulk via COPY:

COPY items (embedding) FROM STDIN WITH (FORMAT TEXT);
[1,2,3,4,5,6,7,8]
[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]
\.

Upsert via ON CONFLICT:

INSERT INTO items (id, embedding) VALUES (1, '[1,2,3,4,5,6,7,8]')
    ON CONFLICT (id) DO UPDATE SET embedding = EXCLUDED.embedding;

Querying

Distance operators

Op Meaning Indexed (turbovec AM)?
<-> Euclidean (L2) distance yes - vec_l2_ops
<#> negative inner product yes - vec_ip_ops (default)
<=> cosine distance (1 - cos θ) yes - vec_cosine_ops
<+> taxicab (L1) distance yes - vec_l1_ops

Distances are returned as double precision. Distance accumulators are f64 internally - pg_turbovec’s avg(vector) and sum(vector) preserve more precision than pgvector’s f32 accumulators on million-row corpora.

Named functions (pgvector-compatible)

l2_distance(a vector, b vector)            RETURNS double precision
inner_product(a vector, b vector)          RETURNS double precision
cosine_distance(a vector, b vector)        RETURNS double precision
l1_distance(a vector, b vector)            RETURNS double precision
vector_dims(v vector)                       RETURNS integer
vector_norm(v vector)                       RETURNS double precision

to_vector(text)                          RETURNS vector
to_vector(text, integer, boolean)        RETURNS vector  -- with dim check
array_to_vector(real[])                  RETURNS vector
array_to_vector(real[], integer, boolean) RETURNS vector  -- with dim check

subvector(v vector, start integer, length integer) RETURNS vector
vec_normalize(vector)                   RETURNS vector
vec_zeros(integer)                       RETURNS vector
vec_check_dim(vector, integer)          RETURNS vector  -- assertion

Aggregates

SELECT avg(embedding) FROM items WHERE topic = 'cats';   -- centroid
SELECT sum(embedding) FROM items;

Both are PARALLEL SAFE and use f64 accumulators.

JSONB I/O

SELECT '[1, 2.5, -3]'::vector::jsonb;             -- → [1, 2.5, -3]
SELECT '[1, 2.5, -3]'::jsonb::vector;             -- → vector

Indexing

pg_turbovec ships an index access method named turbovec, included in the default build (the experimental_index_am and relfile_storage Cargo features were retired in v1.3.0; the relfile-resident storage path is now the only strategy).

-- Cosine-distance ordering (most common for semantic search):
CREATE INDEX items_emb_idx ON items USING turbovec (embedding vec_cosine_ops)
    WITH (bit_width = 4);

-- Inner-product ordering:
CREATE INDEX items_emb_ip_idx ON items USING turbovec (embedding vec_ip_ops)
    WITH (bit_width = 2);   -- 32× compression vs FP32; some recall loss

-- Online (non-blocking) build:
CREATE INDEX CONCURRENTLY items_emb_idx ON items
    USING turbovec (embedding vec_cosine_ops);

Reloptions:

Option Default Range Notes
bit_width turbovec.bit_width_default (4) 1, 2, 3, 4 Lower = smaller index, lower recall. 2/3/4 = TurboQuant. 1 = centered sign binary quantization with full-precision reranking (shipped in v2.6.0; see 1-bit search).

Index AM lifecycle

The turbovec AM supports the full PostgreSQL index lifecycle:

  • CREATE INDEX [CONCURRENTLY] (parallel build is roadmap)
  • INSERT → aminsert (idempotent on IdAlreadyPresent - handles CIC’s two-pass build and HOT updates)
  • DELETE + VACUUM → ambulkdelete (we track every live u64 id in a parallel Vec so dead rows are removed correctly)
  • REINDEX [CONCURRENTLY]
  • DROP INDEX
  • Order-by-op scans (ORDER BY emb <=> q LIMIT k) via the executor’s recheck-orderby path

Filtered (hybrid) search

Three patterns, one decision matrix — full guide in docs/FILTERING.md: a partial index (CREATE INDEX ... WHERE tenant_id = X) for known filter values, the in-kernel allowlist turbovec.knn(..., allowed) for selective per-query id sets, and iterative scan for the normal ORDER BY ... LIMIT ergonomics. The allowlist (shown below) pushes the id set into the SIMD scoring loop — a selective filter gets cheaper, not more expensive (the kernel short-circuits 32-vector blocks whose allowed-slot mask is empty before any LUT lookup; measured crossover at ~7–10% selectivity).

-- Top-10 nearest, restricted to a tenant or topic:
SELECT k.id, d.body
FROM   turbovec.knn(
         'items'::regclass,
         'id', 'embedding',
         '[...]'::vector,
         10, 4,
         ARRAY(SELECT id FROM items WHERE tenant_id = $1)::bigint[]
       ) k
JOIN   items d USING (id)
ORDER  BY k.score DESC;

The function-driven turbovec.knn(...) API and the turbovec index AM share the same TurboQuant kernel and the same backend-local cache; pick whichever fits your query shape (the AM integrates with ORDER BY ... LIMIT; knn(...) lets you pass an allowlist for hybrid retrieval).

Performance

Operations note: shared_buffers. As of v1.19.0, every byte of a pg_turbovec index is read through PostgreSQL’s buffer manager (ReadBufferExtended) — there is no relfile mmap. shared_buffers sizing matters: size it to hold the hot (compressed) index for best cold-fill latency. pg_turbovec’s 7–15× compression vs fp32 HNSW is what makes “the index fits shared_buffers” achievable at corpus sizes where an uncompressed index could not. A cold-fill on an index larger than shared_buffers pays the buffer-manager’s per-page pin/lock/copy cost the first time each backend touches it; warm (already-cached) scans are unaffected by index size.

Out-of-core (>RAM) serving for IVF indexes doesn’t need shared_buffers to hold the whole index either — turbovec. out_of_core (default auto) keeps the per-backend resident set at O(probes·cell_size) by gathering only the probed cells' pages through the buffer manager.

Architecture: docs/ARCHITECTURE.md § 8.1 “Index AM · buffer-cache-only reads”. Design rationale for the (now-shipped) buffer-cache-only read path: docs/BUFFER_CACHE_ONLY_DESIGN.md. Historical mmap-era benchmarks (v1.5.0–v1.18.x, since removed): docs/RECALL.md § 2.5–2.6.

Performance methodology. The headline numbers in the table at the top of this README come from a real head-to-head against pgvector 0.8.0 HNSW on the dbpedia-entities-openai-1M corpus (1 M Wikipedia/DBpedia entities × 1536-d OpenAI text-embedding-ada-002 embeddings) running on a single Intel i9-12900H box with 32 GiB RAM and PG 17.9 from the pgrx-managed install tree. Full methodology, query set, ground-truth generation, and reproduction scripts in docs/RECALL.md § 2.2 and benches/scripts/. The synthetic-uniform tables below are from a pure-Rust kernel bench and are kept for historical comparison - they understate real-world recall because uniform-random vectors have no clustering structure for quantization to exploit.

Recall (synthetic, 1 000 random unit-norm vectors, 50 queries)

dim bit_width R@1 R@10 R@100
128 2 0.40 0.65 0.76
128 4 0.80 0.89 0.93
384 2 0.34 0.62 0.76
384 4 0.78 0.89 0.93
768 2 0.50 0.62 0.76
768 4 0.82 0.88 0.92

Random vectors have no clustering structure for the quantiser to exploit - real embeddings (GloVe, OpenAI ada-002) recall meaningfully better. Reproduction:

cargo bench --bench recall --no-default-features --features pg16

Real-world fixtures via the TURBOVEC_FIXTURE_PATH env var; format documented in docs/RECALL.md § 6.1.

Compression (from the TurboQuant paper)

dim FP32 / vector TurboQuant 4-bit / vector TurboQuant 2-bit / vector
128 512 B 68 B 36 B
384 1 536 B 196 B 100 B
768 3 072 B 388 B 196 B
1536 6 144 B 772 B 388 B
3072 12 288 B 1 540 B 772 B

A 10 M-row × 1536-dim corpus that needs ~62 GiB of RAM as FP32 fits in ~7.7 GiB at 4-bit and ~3.9 GiB at 2-bit - without any data-dependent codebook training.

Search speed (from the TurboQuant paper, x86 AVX-512BW)

100 K vectors, 1 K queries, k=64, single-threaded:

  • TurboQuant matches or beats FAISS IndexPQFastScan at every 4-bit configuration tested (d=384, 768, 1536, 3072).
  • TurboQuant runs within ±1% of FAISS at 2-bit single-threaded.
  • On ARM (Apple M3 Max), TurboQuant beats FAISS by 12-20% at every config the paper measured.

We have not yet run pg_turbovec end-to-end against pgvector + HNSW or pgvectorscale + StreamingDiskANN. That comparison is the next item on the v1.0.0 roadmap; see docs/RECALL.md.

How it works

TurboQuant compresses each vector to 2/¾ bits per coordinate using:

  1. Normalize - strip the L2 norm; store as a single f32 scale.
  2. Random rotation - multiply by a fixed orthogonal matrix so each coordinate independently follows a known Beta distribution.
  3. Lloyd-Max scalar quantisation - bucket each coordinate into 2/¾-bit codes optimal for the known distribution.
  4. Bit-pack - dim coordinates → dim * bit_width / 8 bytes.
  5. Length-renormalised scoring - one extra scalar per vector removes the inner-product downward bias the quantiser introduces.

No codebook training, no data passes - adding vectors is O(dim) per vector with no rebuild as the corpus grows. Search rotates the query once and scores directly against the bit-packed codes via SIMD nibble-LUT kernels (NEON, AVX2, AVX-512BW).

Configuration

pg_turbovec exposes 20 GUCs under the turbovec.* namespace (all USERSET — settable per session). The full reference with tuning guidance is in docs/PRODUCTION.md; the most commonly-tuned ones:

GUC Type Default Range
turbovec.bit_width_default int 4 2..=4
turbovec.probes int 16 1..=65536
turbovec.search_k int 32 1..=100000
turbovec.oversample float 1.0 1.0..=100.0
turbovec.hi_dim_rerank enum auto off, auto, on
turbovec.iterative_scan enum off off, relaxed_order
turbovec.max_probes int 64 1..=65536
turbovec.max_scan_tuples int (see docs)
turbovec.out_of_core enum auto off, auto, on
turbovec.coarse_graph enum auto off, auto, on
turbovec.graph_ef int 0 0 auto (=512) / 1..=1000000
turbovec.graph_build_partitions int -1 -1 auto / 0-1 single-pass / N
turbovec.build_parallelism int 0 0..=128
turbovec.scan_parallelism int 0 0..=128
turbovec.cache_size_mb int 256 0..=65536
turbovec.normalize_on_insert bool true -
turbovec.warn_on_rebuild bool true -
turbovec.allowlist str "" CSV of heap-TID bigints
-- Compress harder during this session:
SET turbovec.bit_width_default = 2;

-- Disable the backend-local index cache:
SET turbovec.cache_size_mb = 0;

Migrating from pgvector

pg_turbovec and pgvector coexist cleanly - different schema, type name, and operator-dispatch table. See docs/MIGRATING_FROM_PGVECTOR.md for the full cookbook. TL;DR:

ALTER TABLE docs ADD COLUMN embedding_tv turbovec.vector;

UPDATE docs SET embedding_tv = embedding::real[]::turbovec.vector;

CREATE INDEX CONCURRENTLY docs_emb_tv_idx
    ON docs USING turbovec (embedding_tv vec_cosine_ops)
    WITH (bit_width = 4);

A binary-compatible vector varlena layout (zero-copy cast to/from pgvector’s vector) is on the v1.0 roadmap; until then the real[] bridge is the supported interop path.

Reference

  • Type: vector (variable dimension, f32 coordinates, 1..16000)
  • Schema: turbovec (set on the search_path or fully qualify)
  • Operator classes: vec_ip_ops (default, <#>), vec_cosine_ops (<=>), vec_l2_ops (<->), vec_l1_ops (<+>), and vec_colbert_ops (multivector MaxSim)
  • Index AM: turbovec (build with WITH (bit_width = 1|2|3|4))
  • Aggregates: avg(vector), sum(vector)
  • Full surface listing: docs/USAGE.md and the generated sql/pg_turbovec--<version>.sql after cargo pgrx schema.

Documentation

  • docs/ARCHITECTURE.md - module map, type / operator / aggregate signatures, index AM contract, GUC semantics, phased roadmap.
  • docs/USAGE.md - cookbook covering install, exact + ANN search, aggregates, arithmetic, tuning.
  • docs/PRODUCTION.md - deployment guide: install, GUC tuning, replication, monitoring, troubleshooting.
  • docs/MIGRATING_FROM_PGVECTOR.md
    • hands-on migration with query rewrite tables and a feature comparison.
  • docs/FILTERING.md - filtering & hybrid search guide: partial index vs in-kernel allowlist knn() vs iterative scan, a decision matrix, and the measured allowlist selectivity crossover.
  • docs/HYBRID_SEARCH.md - multivector & hybrid breadth guide: ColBERT MaxSim re-rank (max_sim), dense+sparse reciprocal rank fusion (rrf_score + CTE recipe), and the named-vector multi-column pattern.
  • docs/INDEXAM.md - implementation guide for the turbovec index access method.
  • docs/RECALL.md - recall benchmark methodology and the latest measured numbers.
  • docs/PARITY_GAPS.md - feature-by-feature comparison against pgvector.
  • docs/PARTITIONED_SCALE.md - scaling past PostgreSQL’s single-table ceiling to billions of vectors via hash-partitioning + per-partition indexes + native Merge Append.
  • docs/DEPLOYING_ON_MANAGED_POSTGRES.md
    • durability/replication, restricted-superuser compatibility, read replicas, cancellation, and the memory / insert-cost model for managed or hosted PostgreSQL.
  • docs/PG_VERSION_SUPPORT.md - the per-major support matrix (PG 13-18, 19beta1 experimental).
  • docs/BENCHMARKS.md - the published head-to-head benchmark (Cohere-wiki 1M vs pgvector HNSW), incl. the AVX2 latency frontier and the honest “flat-scan loses on latency, wins on storage + exact recall” finding.
  • docs/BUILDING.md - Nix-specific build recipe (writable pg_config wrapper, BINDGEN_EXTRA_CLANG_ARGS, RUSTFLAGS for openblas).
  • RELEASING.md - release process, version-bump checklist, Codeberg release flow.
  • CHANGELOG.md - phase-by-phase release notes.
  • tests/ - psql regression scripts you can run yourself (cargo pgrx run pg16, then \i tests/03_full_demo.sql).

FAQ

Is pg_turbovec a drop-in replacement for pgvector?

No, by design. We coexist: type name vector, schema turbovec, operator dispatch by argument type. Pgvector users have years of vector(1536) columns and tooling - pretending to be a drop-in would silently change semantics around normalisation and recall. The migration cookbook shows the explicit real[] bridge.

What about halfvec, sparsevec, bitvec?

The types and their distance/arithmetic operators exist (halfvec, sparsevec, bitvec with <->, <#>, <=>, <+>, and Hamming/Jaccard <~>/<%> for bitvec) and coexist with pgvector’s. What they are not is indexable by the turbovec index AM: the AM quantises full-precision f32 input, and half-precision, sparse, and bit representations don’t map onto the TurboQuant kernel. Use them as column types and with the exact operators; for ANN over those representations, pgvector’s HNSW is the right tool.

What about L2 / L1 ANN?

Both are indexable by the turbovec AM — vec_l2_ops (<->) and vec_l1_ops (<+>) drive the kernel and rerank exactly, the same as vec_ip_ops / vec_cosine_ops. l2_distance / l1_distance are also available as exact functions.

What’s not in 1.0?

The two items most likely to be asked about:

  • Binary-compatible varlena layout for vector. We use a CBOR-derived varlena rather than pgvector’s [vl_len_, dim, unused, f32[dim]] byte layout. The cross-extension migration via ::real[]::vector is one-shot and finishes in seconds on a million rows; the 16× quantization savings dominate the per-row layout overhead, so binary-compat is a nice-to-have, not a 1.0 blocker.
  • Indexed Hamming / Jaccard ANN on bitvec. The TurboQuant kernel is a scalar quantizer for dense f32 vectors; it doesn’t fit Hamming-space ANN. And the workload that motivates bit_hamming_ops (memory-pressured semantic search) is already covered better by bit_width = 2 - same byte budget, materially higher recall.

Why two crates: pg_turbovec and turbovec?

turbovec is the upstream TurboQuant implementation in pure Rust by Ryan Codrai. pg_turbovec is the PostgreSQL extension built with pgrx on top of it. pg_turbovec currently pins a small fork at the upstream 0.9.0 line plus an additive integration layer (a borrowed/skip-prepare reconstruction path used to rebuild a search-ready index from PostgreSQL’s own relation pages without a turbovec file). Upstream has since released 1.0.0, which stabilises the on-disk format and reorganises the encode/rotation API (the v5 block-Hadamard rotation changed every encoded byte); adopting it is a planned, wire-format-aware major migration on the pg_turbovec side rather than a routine bump. The additive pieces we carry are tracked for upstreaming (turbovec issue

70’s from_parts + in-memory I/O asks have already landed upstream).

Ecosystem

Vector search on PostgreSQL has three serious open-source options. The honest comparison:

  • pgvector - production- tested at scale, larger feature surface (HNSW for L2 and inner product and L1, plus halfvec, sparsevec, bitvec), and an ecosystem of clients that already speak its types and operators. Choose pgvector if you don’t have memory pressure and you value maturity.
  • pgvectorscale - SOTA published latency on 50 M+ row corpora via StreamingDiskANN, layered on top of pgvector. Choose pgvectorscale if your corpus is in the tens-of-millions-of-rows range and you can run TimescaleDB.
  • pg_turbovec - smallest on-disk footprint, in-kernel filtered ANN (selective WHERE clauses make scans cheaper, not more expensive), zero codebook training. Choose pg_turbovec if memory dominates your cost equation.

The three coexist cleanly in the same database - separate schemas, separate type oids, separate operator dispatch. You can A/B them on your own data without committing to one.

Contributing

Issues and patches: https://codeberg.org/gregburd/pg_turbovec.

# Run the full test suite (boots a private PG cluster):
cargo pgrx test pg16

# Run pure-Rust kernel + recall benches (no Postgres):
cargo bench --bench distance --no-default-features --features pg16
cargo bench --bench recall   --no-default-features --features pg16

# Lints + format:
cargo clippy --features pg16 --tests -- -D warnings
cargo fmt --all -- --check

See CONTRIBUTING.md.

Acknowledgements

  • Ryan Codrai for the turbovec Rust crate and the SIMD kernels that do all the actual work.
  • Google Research for TurboQuant (ICLR 2026) - the algorithm.
  • The pgvector authors for setting the API conventions (<-> <#> <=> <+>, to_vector, array_to_vector, subvector, vector_dims, vector_norm) we mirror.
  • The pgrx maintainers for making PostgreSQL extension development in Rust possible.

License

Apache-2.0 © Greg Burd. See LICENSE.