Contents
- Changelog
- 1–#3 cherry-picked unchanged (pub repack, IdMapIndex parts API, parallel
- [2.10.3] — 2026-09-25
- [2.10.2] — 2026-09-24
- [2.10.1] — 2026-09-23
- [2.10.0] — 2026-09-22
- [2.9.0] — 2026-09-22
- [2.8.4] — 2026-09-21
- [2.8.3] — 2026-09-11
- [2.8.2] — 2026-09-10
- [2.8.1] — 2026-09-10
- [2.8.0] — 2026-09-10
- [2.7.6] — 2026-09-09
- [2.7.5] — 2026-09-09
- [2.7.4] — 2026-09-09
- [2.7.3] — 2026-09-08
- [2.7.2] — 2026-09-08
- [2.7.1] — 2026-09-08
- [2.7.0] — 2026-09-08
- [2.6.0] — 2026-09-07
- [2.5.0] — 2026-09-07
- [2.4.0] — 2026-09-07
- [2.3.0] — 2026-09-07
- [2.2.2] — 2026-09-05
- [2.2.1] — 2026-09-05
- [2.2.0] — 2026-09-04
- [2.1.0] — 2026-09-04
- [2.0.0] — 2026-08-26
- [1.29.7] — 2026-08-25
- [1.29.6] — 2026-08-15
- [1.29.5] — 2026-08-15
- [1.29.4] — 2026-08-14
- [1.29.3] — 2026-08-14
- [1.29.2] — 2026-08-13
- [1.29.1] — 2026-08-11
- [1.29.0] — 2026-08-07
- [1.28.4] — 2026-08-07
- [1.28.3] — 2026-07-31
- [1.28.2] — 2026-07-31
- [1.28.1] — 2026-07-28
- [1.28.0] — 2026-07-28
- [1.27.3] — 2026-07-12
- [1.27.2] — 2026-07-11
- [1.27.1] — 2026-07-11
- [1.27.0] — 2026-07-10
- [pg_test], 512×128d 4-bit): the dropped blocked chain is ≥ the packed
- [1.26.0] — 2026-07-10
- [1.25.1] — 2026-07-09
- [1.25.0] — 2026-07-09
- [1.24.0] — 2026-07-08
- [1.23.0] — 2026-07-06
- [1.22.2] — 2026-07-06
- [1.22.1] — 2026-07-05
- [1.22.0] — 2026-07-04
- [1.21.0] — 2026-07-03
- [1.20.1] — 2026-07-03
- [1.20.0] — 2026-07-02
- [1.19.0] — 2026-06-18
- [1.18.0] — 2026-06-18
- [1.17.1] — 2026-06-18
- [1.17.0] — 2026-06-18
- [1.16.0] — 2026-06-17
- [1.15.1] — 2026-06-17
- [1.15.0] — 2026-06-17
- [1.14.0] — 2026-06-17
- [1.13.1] — 2026-06-17
- [1.13.0] — 2026-06-17
- [1.12.0] — 2026-06-17
- [1.11.1] — 2026-06-16
- [1.11.0] — 2026-06-16
- [1.10.1] — 2026-06-16
- [1.10.0] — 2026-06-16
- [1.9.1] — 2026-06-15
- [1.9.0] — 2026-06-15
- [1.8.0] — 2026-06-15
- [1.7.3] — 2026-06-15
- [1.7.2] — 2026-05-27
- [1.7.1] — 2026-05-27
- [1.7.0] — 2026-05-27
- [1.6.1] — 2026-05-27
- [1.6.0] — 2026-05-26
- [1.5.1] — 2026-05-26
- [1.5.0] — unreleased
- [1.4.1] — 2026-05-26
- [1.4.0] — 2026-05-25
- [1.3.0] — 2026-05-25
- [1.2.0] — 2026-05-25
- [1.1.0] — 2026-05-24
- [1.0.1] — 2026-05-24
- [1.0.0] — 2026-05-24
- 1.0.0-rc.2 — Unreleased
- 1.0.0-rc.1 — 2025
- 0.16.0 — Unreleased
- 0.15.0 — Unreleased
- 0.14.0 — Unreleased
- 0.13.0 — Unreleased
- 0.12.0 — Unreleased
- 0.11.0 — Unreleased
- 0.10.0 — Unreleased
- 0.9.0 — Unreleased
- 0.8.0 — Unreleased
- 0.7.0 — Unreleased
- 0.6.0 — Unreleased
- 0.5.0 — Unreleased
- 0.4.0 — Unreleased
- 0.3.0 — Unreleased
- 0.2.0 — Unreleased
Changelog
All notable changes to pg_turbovec are documented in this file. The
format follows Keep a Changelog
and the project adheres to Semantic Versioning.
[2.11.0] — 2026-10-06
MINOR: adopt upstream turbovec 1.1.1 (staged 2/4-bit search), plus fixes
for two long-standing bugs found by this release’s soak test: one silently
corrupted index entries under concurrent writes, the other leaked memory per
scan (see Fixed; both are in v2.10.3 and earlier). No SQL
surface change, no GUC change, no wire-format change (MetaPageData::version
stays 8), index bytes unchanged. ALTER EXTENSION pg_turbovec UPDATE plus a
restart is sufficient; no REINDEX.
Minor rather than patch for one reason: on hosts where the new search engages, the candidate set an index scan returns can differ slightly from 2.10.3’s (scores are exact; recall@10 measured unchanged). Details below.
Changed
turbovec 1.0.0 → 1.1.1. For an index of ≥ 32,768 rows on aarch64 (dotprod: Graviton2+, Ampere, Apple) or x86 with AVX-512 VBMI+VNNI (Ice Lake server+, Zen 4+), turbovec now keeps its in-memory search cache as separate bit planes and searches in stages: a sign-plane pass for a shortlist, ranking on the lower planes, then an exact rescore. On AVX2-only x86 (and below 32,768 rows) the scan is unchanged.
Measured on Graviton4 (c8gd.8xlarge, Debian 13, PG 16.15 non-assert), 1M × 1024-d real Cohere embeddings, flat index, warm, whole-query
EXPLAIN ANALYZEp50, 3 alternated A/B rounds on one index:search_k=32 100 256 1024 4-bit: 2.10.3 → 2.11.0 7.88 → 6.72 ms (1.17×) 10.88 → 9.52 (1.14×) 18.32 → 15.96 (1.15×) 62.17 → 54.57 (1.14×) 2-bit: 2.10.3 → 2.11.0 7.22 → 6.86 (1.05×) 10.45 → 9.78 (1.07×) 17.81 → 16.51 (1.08×) 61.67 → 55.11 (1.12×) Recall@10 unchanged in every cell (1.000, or 0.993 at 2-bit k=32 for both). p95 improves by the same margin.
The turbovec kernel itself is 3–6.5× faster on this host (pure kernel, same corpus: 4-bit single-query k=10 1.84 → 0.62 ms at 32 threads, 22.7 → 6.6 ms at one thread; k=1024 8.29 → 1.28 ms). End-to-end gains are smaller because the kernel is now a minor share of a flat-index query: the saving in milliseconds is the same, but the rest — the per-candidate heap fetch and exact recheck, ~48 µs per candidate — is unchanged. For flat indexes the recheck, not the scan, is now the bottleneck.
Cold-backend latency kept (and slightly improved) via a new fork carry. Stock 1.1.1 builds the planes cache with a serial
planes_repack, bypassing the parallel cold-open repack that gave v2.10.3 its 3.1× cold-scan cut — it takes 279 ms (4-bit) / 1,092 ms (2-bit) at 1M × 1024-d. Fork carry #4 parallelizes it (18–21 ms / 32 ms), byte-identical to the serial body (parallel_planes_repack_is_byte_identical_to_serial, x86_64 + aarch64). Cold-backend p50 on Graviton4, 1M × 1024-d: 4-bit 513.7 → 476.9 ms, 2-bit 327.0 → 319.8 ms.
Fixed
Silent index-entry corruption under sustained concurrent writes + VACUUM (present since v1.29.1). The deferred insert flush tracks which rows a transaction upserted (
touched_ids) and splices only those onto the current on-disk index. That list was never cleared after a successful flush, and the backend’s cached index outlives the transaction, so a long-lived writer connection re-spliced every row it had ever written on every later commit, taking their codes from its stale in-memory copy. Once VACUUM had removed such a row and PostgreSQL reused its heap slot for a different row, that re-splice either overwrote the new row’s index entry with the old row’s codes (the row is then searched by the wrong vector, so it can be missed) or put the vacuumed entry back (a stale pointer to a dead or reused tuple).turbovec_checkstayed clean throughout (ids remain unique) — only a byte-level comparison against a freshREINDEXshows it. Measured: after 40 min of 3 writers + VACUUM every 2 min on Graviton4, v2.10.3 had 530 live rows with another row’s codes and 2,369 stale entries in a ~660k-row index (v2.11.0 before the fix: 721 / 2,564). Fixed:touched_idsis emptied once flushed. Repro testflush_does_not_resplice_previous_txn_ids(fails before: a vacuumed id is resurrected; passes after). Re-run A/B, same workload side by side: v2.10.3 corrupted 406 entries and left 2,073 stale; v2.11.0 matched a freshCREATE INDEXof its heap byte-for-byte on all 658,221 entries (40 min, 12,387 commits, 40 mid-flush kills, 19 VACUUMs). Details:benches/results/tv111_arm_20261005/FINDINGS.md§6.Affected: flat/IVF TurboQuant (2/¾-bit) indexes that take writes from connections living across many transactions while VACUUM runs. Short-lived connections (one transaction per connection) were not affected. Recovery:
REINDEX INDEXonce after upgrading an index that has taken such writes — the bug corrupted entries already on disk, and upgrading stops new damage but does not repair old. 1-bit (bit_width = 1) and graph indexes use other write paths and are not affected.Per-scan memory leak that could OOM-kill a backend (and restart the cluster) under concurrent writes. Present since v1.8.0; found by this release’s sustained-insert soak.
amendscanwas a no-op, so every finished index scan leaked its Rust-side scan state, including a reference to the backend’s cached copy of the whole in-memory index. While nothing changes that costs little, but after another session commits, the next scan replaces the cached copy and the leaked references keep every superseded copy alive: a long-lived connection (a pooled one, say) that keeps querying an index other sessions keep writing to grows by about one in-memory index per commit it observes. Measured: +200 MB per commit on a 200k × 1024-d 4-bit index, on v2.10.3 and v2.11.0 alike; in the soak a scanning backend reached 52 GB and the OOM killer took it, and with it the postmaster (crash recovery). Fixed: the same workload now holds at 241 → 248 MB over 30 commits. Repro test:amendscan_releases_cached_index_handle(fails before the fix with 6 live references after 5 scans; passes with 1). A scan that aborts with an ERROR still skipsamendscan, so it leaks once; that is bounded and noted in the code.If you run v2.10.x or earlier with long-lived connections and concurrent writes, this alone is a reason to upgrade. Short-lived connections were not affected: the memory is freed when the backend exits.
Behaviour change: approximate candidate set
The staged search returns bit-exact scores but may miss a candidate whose
sign bits alone rank it outside the shortlist. On the Cohere corpus above,
staged vs whole-index scan returns the identical top-10 id set for 98.5–100% of
queries (≥99.6% mean overlap), and the identical top-100 for 73.5–99% (≥99.6%
mean overlap — a tail candidate or two swapped). The exact heap recheck
re-ranks what turbovec returns, and recall@10 against exact ground truth is
unchanged at every measured search_k. To restore the whole-index scan, set
TURBOVEC_4BIT_PLANES=0 and/or TURBOVEC_2BIT_PLANES=0 in the
postmaster’s environment (read once per process).
Safety
The new risk is that after an INSERT/DELETE, turbovec reconstructs the
packed codes we persist from the planes cache. Gates:
- Byte-identical persisted index: the same 1M × 1024-d heap indexed by 2.10.3 and by 2.11.0 on Graviton4 has identical meta + codes/scales/ids chains (sha256).
- New
#[pg_test]persist_is_byte_exact_through_planes_layout: 40k rows (past the planes gate), real remove/add + PreCommit flush, every row’s in-memory and persisted bytes checked at 2 and 4 bits. Confirmed to take the planes layout on Graviton4. cargo pgrx test pg16: 448 passed / 0 failed / 8 ignored on x86_64 (446 on aarch64 Graviton4 before the two fix tests were added).- Sustained-insert soaks on Graviton4 with mid-flush
pg_terminate_backend, periodic VACUUM and UPDATE churn, then a byte-level comparison of every persisted entry against a freshCREATE INDEXof the same heap. These found both bugs under Fixed; the A/B re-run against v2.10.3 is inbenches/results/tv111_arm_20261005/FINDINGS.md§6.
turbovec’s own suite: 523 passed / 0 failed on Graviton4.
Fork
Pinned to gburd/turbovec@pgtv-2.11.0-port (455549f): upstream 1.1.1 + carries
1–#3 cherry-picked unchanged (pub repack, IdMapIndex parts API, parallel
repack; upstream issues #545–#547) + new carry #4 (parallel planes repack).
Not measured
x86 AVX-512 VBMI+VNNI hosts (no measurement here; upstream reports similar kernel gains); IVF and ColBERT latency; indexes under 32,768 rows (unchanged by construction).
Build note
turbovec 1.1 needs LLVM ≥ 22 for its AVX-512 VNNI intrinsics; the nix-packaged
rustc 1.97.0 (LLVM 21) fails with intrinsic signature mismatch. Any
rustup toolchain ≥ 1.89 works (CI uses rustup stable). On aarch64 Linux,
cargo pgrx test (debug profile) fails to assemble gemm-common’s fp16 inline
asm without RUSTFLAGS="-C target-feature=+fp16"; the --release build
(cargo pgrx install --release, what ships) compiled cleanly without it. This
is pre-existing (gemm 0.18.2, unchanged by this release).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.11.0'; and restart PostgreSQL. The
format is unchanged, so no REINDEX is required, but REINDEX any
2/¾-bit flat or IVF index that has taken writes from long-lived
connections while VACUUM was running to clear entries the
touched_ids bug (above) already corrupted. Evidence:
benches/results/tv111_arm_20261005/.
[2.10.3] — 2026-09-25
PATCH: cold-scan latency. No SQL surface change, no GUC change, no
wire-format change (MetaPageData::version stays 8), index bytes
unchanged. ALTER EXTENSION pg_turbovec UPDATE is sufficient; no REINDEX.
Changed
Cold-scan latency cut ~3× by parallelizing the per-backend blocked-layout rebuild. Every cold backend rebuilds the SIMD-blocked code layout from the row-major packed codes at index-open (v7+ persists only the codes, halving the on-disk footprint). That repack was single-threaded and was the dominant term in cold latency. It now runs in parallel across block-aligned ranges.
Measured A/B on identical hardware/corpus/index (c7i.4xlarge, 16 vCPU, AVX-512; 1M × 1024-d 4-bit flat, 534 MB): cold-backend p50 1766 ms → 566 ms (3.1×); warm-backend p50 unchanged (30.4 → 30.6 ms, within noise — a warm backend never repays the repack). See
benches/results/rebench_20260925/COLDSCAN_FINDINGS.md.The parallel repack (turbovec fork carry #3, rev
47a26a3) produces byte-identical output to the serial version, pinned by turbovec’sparallel_repack_is_byte_identical_to_serial(bit-widths 2/¾, sub/above the parallel threshold, tail-padding shapes). This is a speed change only; the persisted format is untouched.
Docs
- RETRACTED the “we LOSE ~490×” latency scoreboard in
docs/PARITY_GAPS.md. A corrected end-to-end benchmark (top-levelEXPLAIN(ANALYZE)Execution Time, literal query vectors — a query-vector subquery in the ORDER BY had added ~90 ms of InitPlan overhead to BOTH engines and produced the bogus 2552 ms figure; one warm psql session per arm) on 1M × 1024-d Cohere-wiki shows flat-bw4 at 5.2 ms / R@10 = 1.000 BEATS pgvector HNSW at R@10 ≥ 0.95 (HNSW 8.6 ms) and is 3× faster at ≥ 0.98. IVF-bw1 is within 2.4–2.8× of HNSW at 58× smaller storage; IVF-bw4 is the weakest arm (confirms flat > IVF forbit_width ≥ 2at this scale). Full iso-recall table + corrected harness inbenches/results/rebench_20260925/. - Documented
turbovec.ivf_max_delta_pct+ “fewer, larger transactions” guidance for continuous high-ingest into an already-large IVF index (docs/PARITY_GAPS.md), and recorded the sparse-ANN and bulk-INSERT design memos (benches/results/parity_20260925/).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.10.3'; — that’s all. No REINDEX, no
downtime, index bytes unchanged.
[2.10.2] — 2026-09-24
PATCH: documentation plus one build-time NOTICE. No SQL surface change, no
GUC change, no wire-format change (MetaPageData::version stays 8), index
bytes unchanged. No REINDEX.
Added
A flat build over 100 k rows now emits a
NOTICEnamingWITH (lists = N).listsdefaults to0, so a plainCREATE INDEX ... USING turbovecbuilds a flat exact scan — and nothing told the user that an approximate, cell-pruned IVF layer was one reloption away. An evaluator concluded from exactly that experience that pg_turbovec “does not support ANN” and chose a different extension. That is a discoverability failure on our side, not a misreading.Deliberately a
NOTICE, not aWARNING: flat is frequently the better choice, so it must not read as a fault to fix. Suppressed for IVF, graph and ColBERT builds (already non-flat) and below 100 k rows, where flat is unambiguously right and the message would be noise.README: “Common objections, answered with measurements”. Four things evaluators say, each answered from our own benchmarks — including the two where the objection is correct:
- “doesn’t support ANN” — it does; the default is exact, which is why it looked absent.
- “HNSW has a high memory footprint and is slow to build” — agreed, which is why we didn’t build on it. Measured 10 M × 1536-d: our index is 4.4× smaller (14.9 vs 65.5 GiB) and builds 2.6× faster (1 h 24 m vs 3 h 38 m), and our own log shows HNSW slowing super-linearly past 5 M rows.
- “pg_turbovec’s own build memory was worse than HNSW’s” — true, and now fixed. That benchmark measured us at 121 GiB peak + 60 GiB swap against HNSW’s 16.9 GiB. Root cause fixed in v2.10.1; measured after: 2 M × 1024-d peak 12.16 → 3.45 GiB, and 10 M × 1024-d went from OOM-killing a 61 GiB host to completing at 11.20 GiB. Scope stated plainly in the README: the post-fix numbers are 1024-d, and 10 M × 1536-d (the dimension the 121 GiB figure used) has not been re-measured.
- “HNSW is faster on query latency” — true, and we don’t dispute it.
README: “Choosing
lists— the ANN tuning knob”. A full tuning guide, because the honest answer is more subtle than “turn on ANN”:- Whether you want IVF at all. At 1 M × 1024-d with the default
bit_width = 4, flat is 6.08 ms at recall@10 = 1.000 whilelists = 1024is 62 % slower and capped at recall 0.959 — a per-probe ceiling that is CPU-independent and that no amount of tuning removes. IVF is a trade, not a free speedup. - The measured decision table (by
bit_width, scale, recall target, and whether the index exceeds RAM). √nis a ceiling, not a target — at 1 M,lists = 4096measured worse than 1024 on every axis: 11× the build time and ~50 % higher latency.- Tune
probesat query time, notlists(which is baked in at build). - Measure on your own data, at matched recall.
- Whether you want IVF at all. At 1 M × 1024-d with the default
Unchanged, deliberately
listsstill defaults to0(flat). Changing the default to√nwas considered and rejected on our own measurements: at the defaultbit_width = 4it would ship a configuration that is 62 % slower and recall-capped at 0.959 up to at least 1 M rows. The defect was discoverability, not the default value.
Tests
445 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the
classic lane. New: flat_build_hints_at_the_ann_option, which also pins the
silence (IVF builds and sub-threshold indexes must not be nagged).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.10.2'; — nothing else. No REINDEX.
[2.10.1] — 2026-09-23
PATCH: build-time memory profile only. No SQL surface change, no GUC change,
no wire-format change (MetaPageData::version stays 8), and the on-disk
index bytes are identical — guarded by the IVF byte-identity tests. No
REINDEX.
Fixed
Phase Z6 —
CREATE INDEX/REINDEXpeak memory cut ~3.5×. The build callback had no memory-context management at all: no switch, no reset, nopfree.Vectoris aPostgresTypestored as CBOR, soFromDatum::from_datumpalloc’s a decoded buffer for every row — and those buffers accumulated in the long-livedambuildcontext for the entire scan. TheCorpusSpillwas faithfully streaming the corpus to disk while PostgreSQL held a decoded copy of all of it in RAM, defeating the spill’s whole purpose.build_callbackis now a thin wrapper that switches into a per-tuple context, calls the unchanged inner callback, restores, and resets. The inner function has six early-return paths, so doing this inline would have been six chances to leak the switch.One hazard came with it:
BufFileCreateTemppalloc’s in the current context and the spill is opened lazily on the first row — inside the context now being reset, which would have left a danglingBufFileon row two.CorpusSpill::new_in(cxt, dim)plus an explicitBuildState.build_cxtkeep it in the long-lived context by construction at all three lazy-open sites (single-vector, BQ/graph, ColBERT).Measured on 2M × 1024-d,
lists = 1414(benches/results/z6_buildmem_20260922/):before after at drain entry 12.07 GiB 2.39 GiB (5.05× less) whole-build peak 12.16 GiB 3.45 GiB (3.5× less) build wall time 1055 s 889 s (16 % faster) This is what made 10M × 1024-d builds OOM-kill a 61 GiB host. Since measured (
benches/results/z6_10m_20260924/): on the same c7i.8xlarge / 61 GiB host with the identical config that died before (mwm = 8GB, 16 parallel workers,lists = 3162), the 10M × 1024-d build now completes in 69.8 min at 11.20 GiB peak private — 18 % of the host, zero OOM events, againstanon-rss58.47 GiB at the kill. The projection above was 11.2 GiB. Index verified sound (wire v8, 10M/10M slots,is_corrupt = false,scan_fraction = 0.00506, 51 ms warm).The Lloyd
crossmatrix was quadratic inlists.train_kmeansallocatedn_sample × listswheren_sample = lists × 256— 9.54 GiB atlists = 3162. Now chunked over sample rows under a fixed ~256 MiB budget. Bit-identity is tested, not assumed:kmeans_cross_chunking_is_bit_identicalcompares the whole-sample call against chunk sizes 1/7/64/199/500.
Changed
turbovec.build_parallelismis a memory knob, and its description said the opposite (“only build wall-clock changes”). Measured: 16 threads → 16.24 GiB peak private, 1 thread → 12.16 GiB — ≈ 0.27 GiB/thread of thread-local GEMM packing buffers, whichmaintenance_work_memdoes not bound. Corrected, with the advice to lower it whenCREATE INDEXis memory-constrained.trace_stage!(underTURBOVEC_BUILD_TRACE) now reports private memory per stage, plus a new0_at_drain_entrymarker. Peak had been misattributed for three sessions because the only signal was a process-wide total — which also includesshared_buffers. That single marker localised this bug in one build.
Documentation
- A constraint is now pinned by test:
rotate_corpus_intois not invariant to row-block shape (Parallelism::Rayon(0)makes the GEMM’s reduction order depend onm), so the reservoir rotation cannot be chunked to save memory. A blocked-rotation optimisation was written and CI caught it before it could change index bytes;rotate_corpus_is_not_row_block_shape_invariantrecords why. The pre-existingrotate_corpus_bit_identical_across_pool_sizesdoes not cover this — it varies thread count at a fixed shape.
Tests
444 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.10.1'; — nothing else. No REINDEX.
[2.10.0] — 2026-09-22
MINOR: one new GUC and a changed IVF insert/scan behaviour. No SQL-surface
change and no wire-format change (MetaPageData::version stays 8) —
existing indexes decode byte-identically and no REINDEX is required.
Changed
Phase Z5 Route A — an IVF index no longer loses its cell layout on the first
INSERT.aminsertcannot place a row into its cell without an O(n) reshuffle, so appended rows land at the tail, outside every cell. Previously the deferred-commit flush dropped the coarse-centroid and cell-directory chains outright, so one commit turned the index into an O(n) flat scan untilREINDEX— an operational cliff with no cheap recovery. The flush now writes those chains back unchanged, and the scan additionally sweeps the tail exhaustively. Results stay exact; only latency is affected.No new meta field was needed. The delta length is derivable as
n_live - cell_directory.total_vectors(), because existing rows keep their slot (an UPDATE writes in place) and inserts append at the tail. An earlier assessment indocs/PARITY_GAPS.mdcalled this “blocked on a wire-format change”; that was wrong, and the correction is recorded there.Bounded by the new
turbovec.ivf_max_delta_pct(default10, range0..=100): past the bound the index degrades to flat and reports it exactly as before, so the tail cannot grow until the optimisation is an O(n) scan with extra bookkeeping. Setting it to0restores pre-2.10.0 behaviour byte-for-byte.
Fixed
- Out-of-core path could make appended rows unreachable. The OOC gather iterates only probed cells, so a tail outside every cell was never read — those rows existed on disk and could never be returned. That is silent loss, not slowness. The delta is now fed through the same tombstone-aware run-splitting as cells, so a tombstoned appended row cannot resurrect (the v2.7.0 class of bug).
Documentation
shared_preload_libraries = 'pg_turbovec'is now documented as required (README + first section ofPRODUCTION.md). Everyturbovec.*GUC is registered by_PG_init, i.e. at library load; without preloading a backend has zero of them,SET turbovec.probes = 16is accepted and silently ignored, and every query runs at compiled-in defaults. Indexes still build and queries still return correct results, which is exactly why it misdiagnoses as “tuning has no effect” or “IVF pruning doesn’t work” — it cost a full benchmark round before being spotted. Includes the one-line check (SELECT count(*) FROM pg_settings WHERE name LIKE 'turbovec.%') and theLOAD/session_preload_librariesalternative for managed providers (verified: every GUC isUserset, with no Postmaster/Sighup context).
Measured (EC2, benches/results/z5_delta_20260922/)
c7i.4xlarge (16 vCPU, AVX-512), PostgreSQL 16.15, 1M × 256-d, bit_width=4,
lists=1024, 1000 inserts, 30 warm queries per arm in one session, fresh
index per arm, autovacuum disabled, index health verified after the inserts:
| arm | in-memory | out-of-core |
|---|---|---|
pre-Z5 (pct=0) → degraded, scan_fraction 1.0 |
4.51 ms | 4.43 ms |
Z5 delta (pct=10) → healthy, scan_fraction 0.0156 |
4.03 ms | 3.46 ms |
The win is 11 % in memory and 22 % out-of-core — not the 64× that was modelled. The model assumed latency scales with rows scanned, but at 1M × 256-d a full 4-bit scan is only 1.6× a 1-cell scan (5.97 vs 3.66 ms): the 139 MB index is RAM-resident and per-query fixed costs are the same order as the SIMD sweep. This agrees with our own published 1M bw4 result (flat 6.08 ms beats IVF 16.04 ms; v2.8.3 “flat wins at every target”). This release ships on the functional contract — no cliff on insert, plus the OOC correctness fix — not on the latency delta. The >RAM regime, where a full scan is disk I/O and the win could be materially larger, is explicitly unmeasured.
Corruption validation (HARD MANDATE)
6 concurrent writers + 4 readers + VACUUM every 30 s for 5 minutes, with
autovacuum enabled: is_corrupt = false, no duplicate id,
n_vectors == slot_count exactly, wire still v8, and the cell layout
survived (degraded = false, scan_fraction = 0.0156). Two apparent
discrepancies were investigated rather than assumed: the index holding 1072
more rows than the heap (lazy vacuum reclaim of 10072 dead tuples) and a
1000-row sample returning 254 (turbovec.search_k = 32 caps candidates).
Tests
442 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the
classic lane. New: the delta-bound predicate (accept/reject, 0 disables,
inconsistent layouts), plus two integration tests asserting both halves —
cells preserved and every appended slot swept even when a single cell is
probed, while unprobed cells stay excluded (otherwise the pruning is gone and
it is a flat scan in disguise).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.10.0'; — nothing else. No REINDEX.
To keep the previous behaviour exactly: SET turbovec.ivf_max_delta_pct = 0.
[2.9.0] — 2026-09-22
MINOR: adds one SQL function. Wire format unchanged from 2.8.x
(MetaPageData::version stays 8), so existing indexes decode
byte-identically and no REINDEX is required — the upgrade is in place.
Added
turbovec.index_degradation(regclass)— quantifies an IVF degradation instead of merely flagging it. Returnsdegraded,lists,n_vectors,scan_fraction,est_slowdownand arecoverystring. Phase Z1 made degradation observable and Phase Z4 made the planner cost it correctly; neither told an operator the size of the problem, which is what decides whether to act — a degraded 10k-row index is a non-event, a degraded 10M-row index is an outage. A degraded index reportsscan_fraction = 1.0(it reads everything) andest_slowdown = lists / probes, plus theREINDEXcommand naming the index. A flat index is explicitly not reported as degraded — it scans everything by design, and faulting it would train operators to ignore the signal. Reads only the meta page (one buffer hit), so it is safe to poll from monitoring.
Changed
- Phase Z4 —
amcostestimateis now probe- and filter-aware. Three defects, all of which made the planner blind to what the AM does:- IVF was costed as a full-corpus scan although the scan clamps to
turbovec.probescells, so an index probing 1 of 1024 cells was costed identically to a flat scan of everything. Cost now scales withprobes / lists, floored at one cell’s worth. A degraded IVF index is costed as flat, since that is the path it takes. index_selectivitywas hardcoded to0.0for every query. It now derives from the planner’s ownrel->rows / rel->tuples, so we agree with the planner by construction instead of second-guessing it.- A pre-existing unit error: the ns→cost conversion divided seconds
by
cpu_operator_cost, making a 1M × 1024-d flat scan cost ~23 while PostgreSQL costs the equivalent sequential scan at ~73,000 — about 3000× too cheap, which let an ANN path beat plans that are genuinely faster. Now ~1688. The ns throughput model itself validated against our own published measurement (model 5.3 ms vs measured 6.08 ms), so only the unit was wrong.
index_pagesis likewise scoped to the pages a probed scan touches. The arithmetic lives in two pure, unit-tested functions because it is not observable throughEXPLAIN: the index-scan node’s cost also carries PostgreSQL’s heap-fetch and qual costs, which swamp it. - IVF was costed as a full-corpus scan although the scan clamps to
Documentation
docs/FILTERING.md— the allowlist crossover is now actionable. The measured table (2.6–14.7× faster below ~7 % selectivity, 2.6× slower at 100 %) was only useful if you knew your filter’s selectivity. Adds a copy-pasteable way to read PostgreSQL’s own estimate, a fraction→technique table, and two caveats: the crossover percentage is host-dependent (the shape transfers, not the number), and the estimate is only as good as your statistics.- Phase Z3 rescoped to a non-gap, with the reasoning recorded.
“Automatically turn a
WHEREinto a kernel mask” is not implementable by a PostgreSQL AM: a scan key isindex_key operator constantover an index column, so a qual on any other column becomes an executor Filter and never reaches the AM.amgetbitmapis not a route either — it returns an unordered bitmap, and ordering is the whole value of an ANN scan. The useful behaviour shipped in v1.8.0 as iterative scan, which is demand-driven and needs no view of the filter. - Phase Z5 scoped, with both routes shown blocked. A bounded mutable
delta needs a wire-format change (both scan paths assert
directory.total_vectors() == n_liveandcentroids.len() == lists * dim; a synthetic “delta cell” would need a centroid it does not have). Cheap in-place cell reassignment needs adecode/reconstructthat turbovec does not expose (pack.rshas onlyrepack/unblock). The reporting half shipped instead.
Tests
435 → 437 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.9.0'; — creates the new function.
No REINDEX.
[2.8.4] — 2026-09-21
Code-only release. Wire format unchanged from 2.8.3 (MetaPageData::version
stays 8); no REINDEX needed and no SQL surface change.
Fixed
Phase Z1 — an IVF index that takes writes now REPORTS its degradation (ordinary TurboQuant path). An
aminsertcannot place a row in its cell without an O(n) reshuffle, so the deferred-commit flush appends and the index falls back to a flat scan. The 1-bit BQ path always preservedlistsand stampedivf_degradedsoturbovec.index_is_degraded()and the throttledambeginscanWARNING fired; the TurboQuant path blankedlists, soindex_was_ivf()went false and the latency cliff was silent — an operator got a quietly slower index with no signal and nothing to act on. Root cause:reconcile_and_write_flushplanned its meta page viaplan_with_blocked, which hardcodeslists: 0. It now captures the on-disklistsunder the already-held exclusive rewrite lock and stamps both fields, leaving the coarse/cell-directory offsets at zero so the readers return empty and the scan takes the flat fallback deterministically rather than by a length coincidence. Both fields are existing v4 meta scalars — no chain is added or moved, so the chain-offset running-sum class is not implicated.An
INSERTinto an index builtWITH (assign_dups > 1)no longer claims the index is corrupt. IVF-4a soft assignment stores a boundary row in several cells on purpose, so its external id appears in several slots andslot_to_idis deliberately not a bijection — while the insert path loads the index into a flatIdMapIndex, which requires one. The rejection reportedcorrupt relfile pages: duplicate idswith a REINDEX hint; both halves were wrong (turbovec_checkverifies such an index clean, and a rebuild reproduces the same by-design duplicates), sending operators hunting for corruption that does not exist. It now reportsFEATURE_NOT_SUPPORTED, namesassign_dups, states the index is effectively READ-ONLY, and the HINT says how to confirm it is healthy. Behaviour is otherwise unchanged: the INSERT still fails,SELECTstill works. Zero cost on the healthy path —from_id_map_partshas exactly one failure mode, so reaching the handler already identifies the cause.
Documentation
docs/PARITY_GAPS.mdgains a zvec source review (localalibaba/zveccheckoutd88357bvsdeac2d9), annotated throughout as un-benchmarked: four real gaps in priority order, the one lesson worth stealing (bounded mutable delta + explicit consolidation), explicit non-gaps (hybrid fusion, ColBERT, WAL, durability, scalar filtering — PostgreSQL’s or already ours), and deliberately deferred items (RaBitQ/PQ, DiskANN). Adds phases Z1–Z5; Z3 (automatic predicate→ANN handoff) is gated on Z4 (costing), because pushing filters whileindex_selectivityis hardcoded to0.0would only make bad plans confident.assign_dups > 1is now documented as making the index read-only, in bothdocs/BENCHMARKS.md(which recommended it for recall without saying so) and the reloption reference insrc/index/options.rs.AGENTS.mdgains four steering rules: degradation must be observable; competitor comparisons stay source reviews until measured; the BQ and TurboQuant insert paths write at different times (BQ synchronously inaminsert, TurboQuant deferred toPreCommit, which a#[pg_test]never reaches — so a plainINSERTin a test exercises nothing on that path); andassign_dups > 1indexes are read-only.
Tests
429 → 430 passed / 0 failed / 8 ignored, uniform across pg13–19 native
plus the classic lane. New: ivf_flush_degradation_is_reportable,
ivf_soft_assign_index_rejects_insert_and_is_not_corrupt,
ivf_soft_assign_insert_error_names_assign_dups.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.8.4'; — nothing else. No REINDEX.
[2.8.3] — 2026-09-11
bit_width = 4 + IVF measured at 1M — flat wins at every target. This
answers the production user’s question with data rather than the forecast
v2.8.2 shipped. Documentation only; the binary is byte-identical to 2.8.2. No
wire change (v8), no SQL surface change, no REINDEX.
The measurement
1M × 1024-d real Cohere corpus, AVX-512, 48 configs, every one confirmed to
run a real Index Scan. Artefacts:
benches/results/bq_1m_bw4ivf_20260911/.
| target | bw4 flat | bw4 lists = 1024 |
|
|---|---|---|---|
| R@10 ≥ 0.90 | 6.08 ms (w=32) | 16.04 ms (p=64) | flat by 62 % |
| R@10 ≥ 0.95 | 6.08 ms (w=32) | 16.13 ms (p=128) | flat by 62 % |
| R@10 ≥ 0.98 | 6.08 ms (w=32) | unreachable | flat only |
| R@10 ≥ 0.99 | 6.08 ms (w=32) | unreachable | flat only |
The write-up leads with the recall ceiling, not the latencies, because the
conclusion does not depend on any timing. At probes = 128, widening the
rerank window from 32 to 2000 leaves recall at exactly 0.959 across all 8
windows — it does not move by a single query, because the true neighbours are
not in the probed cells. Recall is CPU-independent, so discard every latency
number and 4-bit IVF still loses at the top two targets. Flat’s cheapest
config is also its most accurate (R@10 = 1.000 at window 32), so IVF never
gets an opening. The 62 % gap is corroboration.
Published for completeness rather than left for a reader to find: the one sub-flat p50 in 48 rows is at window 2000, IVF 116.77 ms versus flat 122.19 ms (0.96×). It is not a win — 0.959 recall against 1.000 — so it fails a matched-recall comparison. The tightest honest framing: IVF’s fastest configuration anywhere is 15.76 ms at R@10 = 0.875, against flat’s 6.08 ms at R@10 = 1.000.
The OOM worry is retired for this configuration
Possibly more actionable for an operator than the latency result. At 1M × 1024-d
with maintenance_work_mem = '4GB' and 4 parallel maintenance workers:
| arm | bytes/vec | index | build | peak build anon-RSS |
|---|---|---|---|---|
| bw4 flat | 559.95 | 534 MB | 14.89 s | 2.641 GiB |
bw4 lists = 1024 |
564.17 | 538 MB | 101.22 s (6.8×) | 2.154 GiB |
Real peaks, not lower bounds — 0.25 s sampling, 466 in-build samples, from a
pure-bash sampler that left idle loadavg at 0.00–0.01. IVF’s peak is below
flat’s, and 2.6 GiB is nowhere near the 20.3 GB an unbounded 3GB setting
reached on the earlier 250k build. The hazard is leaving
maintenance_work_mem unbounded, not 4-bit IVF. The genuine cost is the
6.8× build time.
Caveats, recorded rather than glossed
- All 48 latency rows are contention-flagged, and the unflagged-row filter is
unavailable here: 0 of 8 flat and 0 of 40 IVF rows survive it — exactly
the trap § 0.6e documents. The bias runs against IVF (flat ran at 22.0 %
mean
cpu_busyversus IVF’s 6.1 %, a 3.64× difference; mean loadavg 1.87×) and flat still won by 62 %, so a quiet re-time could only widen flat’s margin. - A second corpus-identity surprise. This arm’s resolvability spread was 18.6–69.1 % against § 0.6e’s 154–203 % on nominally the same corpus and shards. Both clear the ~10 % unusable floor so each run’s internal comparison stands, but the discrepancy is unexplained and the two runs must not be treated as same-corpus. The first such surprise forced the v2.8.1 corrections.
- Ground truth took 438.4 s against § 0.6f’s ~275 s on the same parallel-CTAS path. Unexplained, flagged so a future diff does not misread it as a harness regression.
- Scope: a ≥ 0.98 user is answered at any scale, since the per-probe ceiling is structural rather than a size effect. A ≲ 0.90 user well above 1M is not answered by this arm.
Also closed a gap the previous 1M run left open: held-out queries are
value-disjoint as well as id-disjoint (value_overlap = 0 via md5 hash
join).
Operational: AWS burner rules
AGENTS.md had no AWS section, which is what let a mid-run account expiry stand
a benchmark instance up with no way to terminate it (account bene expired at
12:03 UTC, ~48 min after launch; InvalidClientTokenId on a previously-working
key means the account went away, not that you broke something). Added the
current burner (lava) and the rules that made that incident cost zero data
— chiefly pull artefacts as you go.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.8.3'; — no REINDEX.
[2.8.2] — 2026-09-10
4-bit IVF is supported — the “1-bit-only” result was about benefit, not support. Documentation only; the binary is byte-identical to 2.8.1. No wire change (v8), no SQL surface change, no REINDEX.
The report, and the answer
A production user running bit_width = 4 read v2.8.0’s “the 1M IVF+BQ crossover
is 1-bit-only” as meaning IVF cannot be combined with 4-bit, and asked whether
support could be added.
It already exists and always has. WITH (lists = N) composes with every
bit_width — 4-bit IVF is the original IVF path, out-of-core end-to-end since
v1.13.0, and 2-bit and 3-bit work too. The only bit_width/kind combination the
code rejects is bit_width = 1 with graph = true
(src/index/options.rs). Verified by grepping every rejection site: there is no
bit_width gate on IVF anywhere in the tree. Nothing to enable, nothing to
wait for.
The misreading is this project’s fault, not the user’s — the guidance paragraph sits inside the README’s 1-bit section, so a 4-bit reader lands on it naturally. Corrected in the README, and § 0.6e’s heading is restated as “the crossover EXISTS, and the benefit is 1-bit-only” with a callout naming the misreading so the next reader does not repeat it.
New § 0.6g — a 4-bit user’s three questions, answered separately
They were being conflated:
- Is it supported? Yes.
- Will it help? Probably not — and this combination was never
measured, stated plainly rather than implied.
bit_width = 4+lists = Ndoes not appear in any artefact at either scale; bw4 is only ever swept flat. The mechanism predicts no win: the crossover needs both an expensiveO(n)scan and a quantizer lossy enough to demand a wide rerank window, and 4-bit needs only a 32-wide window versus 800 for 1-bit, so its scan is already cheap and cell-restriction mostly adds overhead. 2-bit is the direct evidence — a clean loss at 1M for exactly that reason. - What could it cost? Two things. The per-probe recall ceiling (0.986
at
probes = 128; R@10 ≥ 0.99 unreachable at any setting, and widening the rerank window does not recover it, because the true neighbours are not in the probed cells). And build memory —bw4 + lists = 512at 1024-d reached 20.3 GB anon-RSS and was OOM-killed atmaintenance_work_mem = 3GB.
New: “Should you enable IVF?” in docs/PRODUCTION.md
A decision table plus a 20-minute experiment an operator can run on their own
data: baseline at their recall target, build an IVF copy with
maintenance_work_mem bounded, sweep probes, compare at matched recall
(not matched settings — an iso-knob comparison flatters whichever index gets a
wider effective window), and verify with EXPLAIN that it really is an
Index Scan rather than a sequential-scan fallback dressed up as a latency
result.
The reasoning behind pushing users toward their own measurement: since v2.8.1 this project’s own 1M and 250k figures come from different corpora (the older Cohere dataset became gated), so published cross-scale deltas are suggestive rather than measured. A user’s corpus is the only authority for their workload, and a negative result from their data is worth more than a positive one from ours.
A measurement of bw4 + lists = 1024 at 1M is in flight and will replace
§ 0.6g’s prediction with a number when it lands.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.8.2'; — no REINDEX.
[2.8.1] — 2026-09-10
Corrections to v2.8.0’s benchmark write-up. Documentation only — the binary is byte-identical to 2.8.0. No wire change (v8), no SQL surface change, no REINDEX. All three items came from the 1M benchmark agent’s final report and were verified before being accepted.
The 1M corpus is not the same corpus as the 250k runs
v2.8.0 presented 1M-versus-250k deltas — IVF’s storage overhead “halving”, the
per-probe ceilings “landing within a hair” of the 250k figures — as though both
scales shared a corpus. They do not.
Cohere/wikipedia-22-12-en-embeddings is now gated (confirmed HTTP 401), so
the 1M arm used CohereLabs/wikipedia-2023-11-embed-multilingual-v3: same
publisher and dimensionality, but a different model and snapshot. The 250k
artefacts carry no corpus label at all, corroborating that they came from a
different pre-existing table.
Every cross-scale delta is now labelled suggestive, not measured. The within-run flat-versus-IVF comparisons are unaffected — both arms of each run share one corpus, and those are what the conclusions rest on.
A trap in the method used to validate v2.8.0’s headline
The 47 % / 38 % IVF wins were validated by re-checking on
contention-unflagged rows only. That is unsound whenever the baseline does
not survive the filter — and on the lists = 4096 arm it does not: all 8 bw1
flat rows are flagged (they ran first, while loadavg was still decaying from
the k-means build) and zero survive, so a filtered comparison there would
“prove” IVF wins against an empty set.
Re-verified: the lists = 1024 arm used for the headline keeps all 8 flat
rows unflagged, so that check was valid — but partly by luck of execution
order. The rule is now in docs/TESTING.md beside the existing control-arm
rule: assert the filtered baseline is non-empty before trusting a filtered
comparison.
Restored: the ground-truth-fix documentation (now § 0.6f)
An earlier rewrite of § 0.6b had overwritten it, and a later edit’s assertion
masked the loss — found by grepping for the measured numbers and getting
nothing. The code fix was never affected. The restored section also records
an accidental validation at 1M scale: the run’s two arms straddled the fix,
giving 3755.7 s (pre-fix INSERT path) versus 275.4 s (CTAS path) —
13.6×, matching the 12× predicted from the plan shape — and the 16 flat
configs shared by both arms reproduce bit-identically across the two GT
implementations, independent confirmation the fix changes no measured number.
Also
- 1M build-memory figures relabelled as lower bounds, not peaks: the
sampler polled every 2 s and was stopped before the second arm, so the
lists = 4096build has no RSS measurement at all. - The contention cause is named: loadavg decaying after each parallel
k-means build (a monotonic decline across consecutive rows) plus an RSS
sampler forking a Python interpreter every 2 s.
cpu_busy_pctof only 3.1–3.2 % withcpu_steal ≤ 0.01confirms it was never saturation. - All 160 configs across both 1M arms ran a real
Index Scan— zero masked sequential scans. - Recorded honestly: held-out queries are id-disjoint from the corpus (join count 0), but value-disjointness was not proven — the 100 × 1M text comparison was abandoned as too slow. A duplicate would require the dataset itself to contain duplicate embeddings.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.8.1'; — no REINDEX.
[2.8.0] — 2026-09-10
The 1M IVF+BQ crossover, measured on a real corpus — plus a parallelised
ground-truth path and the v2.7.4 latency caveat resolved with data. Minor
rather than patch because the harness’s GT path changed shape (same output,
different plan) and the operator guidance for
WITH (lists = N, bit_width = 1) is now materially different. No index
wire-format change (stays v8), no SQL surface change, no REINDEX.
The crossover exists at 1M, and it is 1-bit-only
Corpus: CohereLabs/wikipedia-2023-11-embed-multilingual-v3 (en), 1 000 000 × 1024-d — real, not synthetic — with 100 held-out queries, on an AVX-512 host. The § 0.6d resolvability gate passed at 154–203 % nn1→nn100 spread (the discarded synthetic corpus was 6.6–10.4 %).
| target | bw1 flat | bw1 IVF (lists=1024) |
verdict |
|---|---|---|---|
| R@10 ≥ 0.90 | 25.6 ms | 13.7 ms (p=64) | IVF wins 47 % |
| R@10 ≥ 0.95 | 33.2 ms | 20.7 ms (p=128) | IVF wins 38 % |
| R@10 ≥ 0.98 | 39.3 ms | 44.2 ms | IVF loses 12 % |
| R@10 ≥ 0.99 | 56.8 ms | unreachable | flat only |
Re-computed using only contention-unflagged rows: identical 47 % / 38 %, so
this is not a load artefact. For bit_width ≥ 2 it is a clean no — flat is
4.8 ms at every target and IVF never gets under 16 ms.
Mechanically the crossover needs both conditions: flat’s O(n) scan grown
expensive and a quantizer lossy enough to need a wide rerank window. 1-bit
at 1M needs w=100–256 over 1M rows, so restricting to 64–128 cells of ~1000
rows is a real saving; 2-bit needs only w=32, so its scan is already cheap.
Two-axis guidance, not one:
- bit_width = 1, n ≳ 1M, target ≲ 0.95 → lists = N
- bit_width = 1, target ≳ 0.98 → flat (IVF cannot reach it)
- bit_width ≥ 2 → flat, at least to 1M
lists = 4096 is worse than lists = 1024
A useful negative: quadrupling the list count made every axis worse —
8.7 % more storage, 11× the build (1067 s vs 94 s), and ~50 % higher
latency at matched recall (21.5 vs 13.7 ms at R@10 ≥ 0.90). Each cell holds 4×
fewer rows, so a given recall needs ~4× the probes (p=256 where 1024 lists
needed p=64). The lists ≈ sqrt(n) rule is load-bearing; exceeding it is a
pure loss for BQ.
Also confirmed: the per-probe recall ceiling is structural, not a small-corpus artefact (0.846/0.904/0.944/0.971/0.986 at probes 8/16/32/64/128 at 1M, within a hair of the 250k figures), and IVF’s storage overhead halves at 1M (+3.1 % vs +6.0 %) as the fixed centroid/cell-directory cost amortises.
Ground truth parallelised — in two steps, the second correcting the first
Both verified to leave GT row-for-row identical (old EXCEPT new = 0,
new EXCEPT old = 0 on (qid, hit_id, rk), membership ignoring rank = 0), so
no published recall figure moves:
- The correlated
LATERAL+row_number() OVER (ORDER BY dist)forcedWindowAgg → Sort → Gather: every row shipped to the leader for a single-threaded sort. At 250k × 20 queries,Gather (actual rows=250000, loops=20)with a 13.9 MB leader quicksort, 148.2 s. Replaced with a per-query InitPlan constant so each worker top-N sorts its own share: 76.8 s, 1.93×. - My first version of that fix wrapped the SELECT in
INSERT INTO ... SELECT, which silently threw the parallelism away again. PostgreSQL generates no parallel plan for a data-writing statement — onlyCREATE TABLE AS/SELECT INTO/CREATE MATERIALIZED VIEWare exempt. SerialSeq Scanat 7.49 s versus 3.99 s for CTAS; ~36 s vs ~2.9 s per query at 1M, a 12× gap. Now CTASes each query’s top-kinto a temp table and inserts those few rows.
The second defect was caught by the 1M benchmark run’s pg_stat_activity
(one backend at 99.9 % CPU, zero parallel workers), not by me. Root cause worth
naming: I benchmarked a bare SELECT, measured 1.93×, then shipped an
INSERT — a different statement with a different plan — and never re-timed.
Benchmark the statement you are going to ship, not a proxy for it.
The v2.7.4 contended-latency caveat is resolved with data
The bench host’s load problem was fixed, so the 250k sweep was re-run on the same corpus: 18 of 24 rows now clean (was 0 of 24), recall reproduced exactly, the contended p50s were uniformly 14–16 % pessimistic, and the published matched-recall ratios held to two decimal places (2.70× / 6.13× → 2.75× / 6.09×). The absolute milliseconds in the § 0 tables are left as published — ~15 % conservative — with the correction pointing at the clean artefact. An honest record beats a tidy one.
Also
ivf_streaming_build_temp_file_cleanup asserted on a cluster-wide
pg_ls_tmpdir() delta, so a concurrent #[pg_test]’s transient spill file
failed it — observed on a docs-only commit, which is definitionally not a
regression. It now re-samples and fails only if the growth persists. Same
shared-global-state class as the bench-harness collisions fixed in v2.7.4.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.8.0'; — no REINDEX.
[2.7.6] — 2026-09-09
Documentation consistency pass, plus a narrowed root cause and salvaged numbers from the discarded 1M arm. No shippable code change — binary byte-identical to 2.7.5. No wire change (v8), no SQL surface change, no REINDEX.
The BQ docs contradicted themselves
Across the 2.7.3–2.7.5 runs the docs accumulated claims the measurements had
already falsified. docs/BQ_RECALL_BENCH.md still opened by stating it
“contains no measurements” and pointing at a README row that “currently and
correctly reads not yet published” — both untrue since 2.7.3.
docs/ONEBIT_BQ.md carried an orphaned fragment, stranded mid-paragraph by an
earlier merge, asserting the harness “has NOT been run — no numbers exist
yet”. Every doc was swept for claims the measurements contradict; all are now
consistent. The runbook sections are retained and relabelled as how to
reproduce rather than not yet done.
The discarded 1M root cause, narrowed — and a competing diagnosis ruled out
A plausible alternative was raised: that the generator’s uncorrelated
ARRAY(SELECT randn() ...) subquery had been hoisted, making all 200 cluster
centres identical — the same v1.24.0 test-harness bug class AGENTS.md warns
about. It was tested rather than accepted or dismissed, and both halves
matter:
- The hazard is real. The subquery never references the outer
c. Reproduced minimally: uncorrelated gives 1 distinct vector across 5 rows; adding a correlatingWHERE c = cgives 5. Worth recording, because the behaviour is plan-dependent. - It did not happen here. Measured on the loaded corpus: 200 distinct
centres, 5000 distinct vectors per
cid, same-cluster distance 0.108 vs cross-cluster 0.990.randn()’s volatility forced per-row evaluation.
So the earlier “distance concentration” diagnosis was right in substance but
imprecise about where. Narrowed by measurement: the clustering worked; the
tie is inside each cluster. Within one cluster a member’s 100 nearest
neighbours span just 7.22 %, because 5000 iid Gaussian points at d = 768
concentrate — pairwise separation norm has relative spread ≈ 1/sqrt(2d)
≈ 2.6 %. Cluster separation is not sufficient for rankability, which is the
non-obvious part and the reason the original prompt’s “make it clustered”
instruction was not enough.
Salvaged from that arm: valid 1M storage/build numbers
Corpus geometry doesn’t affect these — per-vector codes are
dim/8 * bit_width and the flat build is a fixed O(n·dim) pass:
| 1M × 768-d | build | bytes/vec | peak build RssAnon |
|---|---|---|---|
| bw1 | 167.3 s | 104.87 | 8.98 GiB |
| bw2 | 152.0 s | 209.72 | 5.95 GiB |
| bw4 | 161.0 s | 404.76 | 6.18 GiB |
bw2/bw1 = 2.000 exactly, confirming the dim/8 sign-code stride holds at
1M rows. Note bw1 has the highest peak build RSS despite the smallest
output: BQ reads the corpus back resident to compute the corpus mean before it
can take signs, so its build memory tracks n · dim · 4 regardless of bit
width. Worth knowing before building a 1-bit index on a RAM-constrained host.
Harness defect: the ground-truth query serialises on the leader
The GT query puts row_number() OVER (ORDER BY <distance>) in a Sort above
the Gather, so parallel workers ship raw vectors to the leader and the leader
recomputes every distance single-threaded. That is why ground truth took
23 084 s (6.4 h) for 100 queries at 1M rows — a plan pathology in the
harness, not a turbovec cost, and the main thing making a real 1M arm
expensive.
New: mandatory pre-flight for synthetic corpora
docs/BQ_RECALL_BENCH.md § 0.6d. Before trusting any recall number from
generated data, run the resolvability probe and require the nn1→nn100 spread to
be comparable to a real corpus:
| corpus | spread | verdict |
|---|---|---|
| real Cohere-wiki 1024-d | 37–268 % | rankable |
| the discarded synthetic 768-d | 6.6–10.4 % | unusable |
With an explicit warning that is_degenerate() and rankability are different
checks — a corpus can pass the first and still be statistically unrankable.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.6'; — no REINDEX.
[2.7.5] — 2026-09-09
The 1-bit dimension sweep is measured, and it produces the sharpest
practical guidance the BQ work has yielded: 1-bit is a high-dimension
technique. Also corrects a hi_dim_rerank documentation error. No shippable
code change — binary byte-identical to 2.7.4. No wire change (v8), no SQL
surface change, no REINDEX.
The penalty collapses as dimension rises
256/512/1024-d, 250 000 rows and 100 held-out queries per dim, matching § 0’s
setup. Artefacts: benches/results/bq_dimsweep_20260909/.
1-bit R@10 at a fixed rerank window rises with dim at all seven swept windows, no exception:
| window | 256-d | 512-d | 1024-d |
|---|---|---|---|
| 32 | 0.394 | 0.581 | 0.744 |
| 256 | 0.729 | 0.888 | 0.967 |
| 800 | 0.866 | 0.966 | 0.994 |
| 2000 | 0.933 | 0.990 | 1.000 |
At matched recall the window penalty versus 2-bit collapses:
| dim | 1-bit window for R@10 ≥ 0.95 | 2-bit | penalty |
|---|---|---|---|
| 256 | 4000 | 100 | 125× |
| 512 | 800 | 32 | 25× |
| 1024 | 256 | 32 | 8× |
At 256-d, reaching R@10 ≥ 0.99 requires a window of 16 000 — reranking
6.4 % of the whole corpus. 1-bit is effectively unusable at 256-d and
below; prefer it at 768-d and up. Storage moves the same way (1.902× →
1.946× → 1.971× versus 2-bit) because 1-bit’s fixed per-index overhead
amortises away as dim grows: 30 % overhead over the dim/8 ideal at 256-d,
11 % at 1024-d. Both axes favour 1-bit more strongly at higher dimension.
Depth confirms it: R@100 at the auto default is 0.455 (256-d), 0.737
(512-d), 0.948 (1024-d) — at 256-d the default loses more than half the true
top-100.
Validation. The 1024-d arm was re-run from a fresh database on a re-sliced corpus and reproduced § 0’s published recall bit-identically at all seven windows, p50s within 1–3 %. Independently re-verified against the published artefact. That validates the published figures and the rebuilt, isolated harness.
Caveat that bounds this. The 256-d and 512-d corpora are prefix slices of the 1024-d Cohere-wiki embedding, not natively-trained embeddings of those dimensions. A native 256-d model concentrates its information into 256 coordinates; a truncated 1024-d vector keeps only the first quarter of a representation spread across all of them. That likely makes the sliced low dims look worse than a native model would, so the trend is an upper bound on dim-sensitivity — directionally sound, magnitude not transferable. Nothing was padded to fabricate a higher dim.
Same contention caveat as § 0: recall and storage stand; absolute p50s are indicative.
Correction: the 1-bit hi_dim_rerank special case is a no-op at dim ≥ 256
docs/ONEBIT_BQ.md and docs/BQ_RECALL_BENCH.md said hi_dim_rerank treats
a 1-bit index as high-dim “at any dim”, implying it widens BQ’s rerank
window generally. That is literally true of the code and misleading about
the effect. The auto window is clamp(effective_dim, 256..=1024), and
since effective_dim = max(dim, 256) for 1-bit but dim for 2/¾-bit, the
two are identical for every dim >= 256. The special case only widens the
window below 256-d.
This strengthens the published results rather than weakening them: a
1-bit-vs-2-bit comparison at dim >= 256 with default settings compares
equal windows, so § 0 and § 0.6c are quantizer-vs-quantizer, not
knob-vs-knob. Both docs now carry a table showing exactly where the special
case bites.
Found by re-deriving the clamp independently against the Rust source while chasing down why the sweep’s pre-registered prediction P-D3 had named the wrong mechanism.
On pre-registration
Three of the sweep’s four predictions were wrong in some respect — the storage
direction was backwards, the hi_dim_rerank mechanism was misnamed (which is
what surfaced the doc error above), and the latency-vs-dim shape was
non-linear. Only the central hypothesis was confirmed, and it was understated.
Recording predictions before the run is what made all of that visible instead
of being quietly rationalised afterwards.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.5'; — no REINDEX.
[2.7.4] — 2026-09-09
Documentation accuracy and benchmark-harness isolation. No shippable code change — the binary is byte-identical to 2.7.3. No wire change (v8), no SQL surface change, no REINDEX.
Correction to v2.7.3’s latency figures
v2.7.3 published the 1-bit BQ p50s without disclosing that the harness had
flagged every one of them. All 24 rows carry
latency.contention.contended_flag = true — 1-minute loadavg 3.16–4.64
against the harness’s gate of 1.5. The artefact recorded this correctly; the
write-up did not surface it. That was my error and this is the correction.
The gate is unreachable on that host and not because of the benchmark: a
stuck systemd --user spinning at 77–86 % CPU for 31 days pins its idle
loadavg near 2.0, so any run there is flagged. Mitigating measurement,
taken rather than assumed: cpu_busy_pct on the pinned cores was only
16–26 % — the load is runnable-elsewhere processes, not saturation of the
bench cores. So the ratios and the shape of the window-vs-recall curve
stand, and the absolute milliseconds are indicative. Recall, storage,
bytes/vector and build time are CPU-independent and unaffected.
README.md, docs/RECALL.md, docs/ONEBIT_BQ.md, docs/UPGRADING.md and
docs/BQ_RECALL_BENCH.md now all carry the caveat. Found by a sub-agent
checking its own contention flags; I had not been checking mine.
Measured: IVF + 1-bit BQ
docs/BQ_RECALL_BENCH.md § 0.6a, artefact
benches/results/bq_ivf_20260909/. WITH (lists = N, bit_width = 1) builds
and scans correctly. Storage overhead over flat BQ is +6.0 % — a fixed
~8.5 B/vector of coarse centroids and cell directory, so proportionally worst
for the smallest codes.
The finding: IVF imposes a per-probe-count recall ceiling that a wider
rerank window cannot break. At probes = 8, 1-bit saturates at R@10 = 0.846
and stays there from window 256 through 2000 — the true neighbours are not in
the probed cells, and exact re-ranking cannot invent them. Flat BQ reaches
0.994 with no probe tuning.
This is the mirror image of Gap-B (v1.25.0), and the distinction is the useful part: there, high-dim recall loss was not retrieval-bound (cell recall 0.98–0.996) and a wider window fixed it; here it is retrieval-bound and the window is irrelevant. Same symptom, opposite cause — diagnose which one you have before reaching for a knob.
At 250k, flat BQ dominates IVF+BQ. That is a scale-dependent boundary,
deliberately not a verdict: IVF’s whole value is bounding scan cost as n
grows, and 250k is below the crossover. Recording it as a boundary is what
the graph kind’s early iso-beam numbers should have done.
A 1M-scale arm was discarded, not published
A synthetic 1 M × 768-d run produced R@10 = 0.031 at window 32, which looks
like catastrophic scale collapse. It isn’t — the generated corpus was
statistically unrankable. Resolvability probe: 1st vs 100th nearest
neighbour differed by only 6.6–10.4 % in cosine distance, versus
37–268 % on the real Cohere-wiki corpus. When the top-100 is effectively a
tie, no quantizer can rank it and “recall” measures the tie-break order.
Cause, narrowed on 2026-09-09 with direct measurement: the clustering
worked (200 distinct centres; same-cluster distance 0.108 vs cross-cluster
0.990, a 9× separation) — the tie is inside each cluster. Within one
cluster, a member’s 100 nearest neighbours span only 7.22 %, because 5000 iid
Gaussian points at d = 768 concentrate: the pairwise separation norm has a
relative spread of ≈ 1/sqrt(2d) ≈ 2.6 %. A plausible competing diagnosis —
that the generator’s uncorrelated ARRAY(SELECT randn() ...) subquery had been
hoisted, making all 200 centres identical (the v1.24.0 harness-bug class) — was
tested and ruled out: the SQL genuinely invites that hoist (reproduced
minimally: 1 distinct vector across 5 rows uncorrelated, 5 when correlated),
but randn()’s volatility forced per-row evaluation here and the loaded
centres are distinct.
Discarded with a full post-mortem, including the probe to run before trusting
any generated corpus, in benches/results/bq_scale_20260909/DISCARDED.md.
Also surfaced by that arm: the harness’s ground-truth query wraps
row_number() OVER (ORDER BY <distance>) around the per-query scan, forcing a
WindowAgg → Sort → Gather shape where workers ship rows to the leader and the
leader sorts the whole corpus (confirmed by EXPLAIN ANALYZE: Gather (actual
rows=200000, loops=5) with a 4.3 MB external disk merge, versus per-worker
top-N heapsort in 31 kB with a constant query vector). A real defect — but
not the reason GT took 23 084 s, which the v2.7.6 notes over-credited it
for. Restructuring measures 1.00× narrow / 1.25× on 3 KB-wide rows, and
6.4 h over 100M comparisons is 231 µs each against a 154 GFLOP job: the cost is
per-row overhead across five full sorts of a 1 M-row wide table on a pre-AVX2
host, not the plan. Left unapplied pending validation at 1 M scale on an AVX2
host.
Storage did confirm the dim/8 stride and an exact 2.00× 1-bit:2-bit ratio
at 1 M rows. Note a synthetic corpus can pass is_degenerate() and still be
unrankable — different checks, and only the second predicts whether recall
means anything.
Harness: --run-id so concurrent arms cannot corrupt each other
Two arms running in one database silently corrupted each other three ways:
a shared bq_query_set (one arm’s 256-d query set replaced another’s 1024-d
one mid-sweep), DROP TABLE IF EXISTS bq_gt destroying 860 s of ground truth,
and a hardcoded bqbench_ index prefix — index names are schema-scoped,
so one arm’s CREATE INDEX failed against a sibling’s index on a different
table, silently costing it an entire bit_width leg. All three are now
namespaced by --run-id / BQ_RUN_ID, validated by the existing identifier
guard. Default behaviour without --run-id is byte-identical.
Stale documentation corrected
- Test counts:
334/334(AGENTS.md) and341/346in six rows ofPG_VERSION_SUPPORT.md→ the actual 427 passed / 8 ignored, uniform across all seven CI legs pg13–19. AGENTS.md: migration matrix was stuck at v1.27.1 and the wire version said 7 when it has been 8 since v2.0.0. Both corrected, and ~160 lines of v1.x release prose that duplicatedCHANGELOG.mdreplaced with current state (kinds, which to recommend, the corruption-history warning for anyone touching persist/scan code, and the upstream ctid bug).docs/PRODUCTION.md: new section on boundingmaintenance_work_memfor high-dim IVF builds. Measured:bit_width = 4, lists = 512at 1024-d reached 20.3 GB anon-RSS and was OOM-killed atmaintenance_work_mem = 3GB, then completed in 50.6 s at 1 GB with two parallel workers. The effective ceiling ismaintenance_work_mem × (1 + max_parallel_maintenance_workers), and an operator sees this as an unexplained backend termination.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.4'; — no REINDEX.
[2.7.3] — 2026-09-08
The 1-bit sign-BQ frontier is measured and published, and a real insert-path bug found while running it is fixed. No wire-format change (stays v8), no SQL surface change, no REINDEX.
Fixed: a 1-bit index built on an EMPTY table rejected its first insert
CREATE TABLE → CREATE INDEX ... WITH (bit_width = 1) (while empty) →
INSERT failed with dim mismatch — index expects 0, row has 1024. The
empty build stamps dim = 0 (no rows, no reloption-pinned dim) and
insert_bq_row read the dim from the meta page. The flat path never had
this bug because it takes the dim from the incoming row; the BQ path now
does the same, seeding the mean as zeros for the 0-row case (a one-row corpus
mean is that row, so every centred coordinate is 0 — exactly what a rebuild
would produce).
Found on a real 1024-d corpus host while setting up the benchmark, not by a test: every in-tree BQ fixture happened to index an already-populated table, so all of them missed it. The regression test covers the exact failing order and also asserts a wrong-dim row is still rejected once the dim is pinned, so the fix does not paper over dim checking.
Measured: the 1-bit recall / storage / latency frontier
Closes the last “not yet published” claim in the README. arnold
(i9-12900H, AVX2 — latency is only publishable on an AVX2 host per
AGENTS.md), PostgreSQL 17.9, pg_turbovec 2.7.2, 250 000 × 1024-d
Cohere-wiki as a native turbovec.vector column, 100 held-out queries
(zero corpus overlap, verified), exact top-100 in-DB ground truth using the
same operator the index serves, postmaster and driver pinned to P-cores
2-5, and all 24 configs confirmed to run a real Index Scan.
bit_width |
bytes/vector | vs 4-bit | R@10 ≥ 0.99 needs | p50 there |
|---|---|---|---|---|
| 1 | 142.3 | 3.98× smaller | window 800 | 36.7 ms |
| 2 | 280.5 | 2.02× smaller | window 32 | 6.0 ms |
| 4 | 565.7 | 1.00× | window 32 | 9.2 ms |
1-bit trades latency for storage, steeply: twice the storage saving of 2-bit, for 2.7–6.1× the latency at matched recall and a 25× wider exact-rerank window to reach R@10 ≥ 0.99. That is a storage-constrained-workload option, not a default — which is how the reloption was already documented, now with numbers behind it.
All four predictions registered before the run held. P2’s falsification
condition was “1-bit crosses at a comparable window, which would mean the
hi_dim_rerank 1-bit special-case is unnecessary” — it did not, so the data
justifies that special-case rather than merely tolerating it.
Two further findings:
- 1-bit degrades faster at depth than at k=10. R@100 tops out at 0.981
(window 2000) and is 0.948 at the
autodefault, while 2-bit reaches 1.000 by window 800. Budget a wider window if you paginate or re-rank past the top 10. - Harness cost trap, documented.
--vec-expraccepts any SQL expression, and a cast like(emb::real[]::turbovec.vector)is evaluated per row per query — 200 M casts for a 1M × 200-query ground-truth build, ~53 s per query. Materialising a nativeturbovec.vectorcolumn first took one query from 53 s → 2.7 s (20×). The 1M arm was abandoned for this reason, not for anything about turbovec.
Caveats are stated in docs/BQ_RECALL_BENCH.md § 0.5 and not glossed: one
corpus, one dim, one shared host (load ≈ 2 at start, so treat absolute
milliseconds as indicative and the ratios as the result), and no IVF arm.
This corpus also cannot separate 2-bit from 4-bit on recall — both sit at
≈1.000 nearly everywhere; it separates 1-bit from both, which is what it was
run for.
Artefact: benches/results/bq_frontier_20260908/ (24 configs + sweep log).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.3'; — no REINDEX.
[2.7.2] — 2026-09-08
BUG#6 is now reported upstream. Documentation-only: the binary is
byte-identical to 2.7.1 (the sole src/ edit is a doc comment). No wire
change, no SQL surface change, no REINDEX.
The root-cause analysis and one-line core fix that v2.7.1 verified have been filed on pgsql-hackers:
ExecForceStoreHeapTuple() loses tts_tid, so ORDER BY-op index scans project an invalid ctid — 2026-09-08, with the patch attached as
v1-0001-....
The filed patch’s execTuples.c hunk is identical to the one verified
here by A/B build. The filed version additionally adds a core regression
test to src/test/regress/{sql,expected}/gist.*, and that test was itself
checked against both builds: it reports ctid_matches = 1 on unpatched 18.4
and 5 on the patched build, so it genuinely gates the fix rather than
passing vacuously.
docs/FILTERING.md and the knn_scan_ctid_projection_upstream_limitation
test now carry the thread link. That test still asserts the current
(broken) behaviour, so it will fail once a fixed PostgreSQL reaches
CI — which is the intended signal to flip the assertion (gated on the
fixing version) and relax the docs, not a regression in this AM.
Until a fixed PostgreSQL ships nothing changes for users, and the workaround
remains a proven necessity rather than a preference: MATERIALIZED CTEs,
text casts inside a subquery, and extra subquery nesting were all tested and
all still yield the sentinel. Chain on your own key column, or use
turbovec.knn().
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.2'; — no REINDEX.
[2.7.1] — 2026-09-08
BUG#6 root cause proven against stock PostgreSQL, and the one-line core
fix verified. Documentation + upstream-patch release: the binary is
byte-identical to 2.7.0 (no wire change, no SQL surface change, no
REINDEX). The only src/ edit is an expanded doc comment on the existing
tripwire test.
BUG#6 is the one where SELECT ctid ... ORDER BY emb <=> q LIMIT k
projects the invalid-item-pointer sentinel (4294967295,0). It was
previously argued to be a core bug; it is now demonstrated:
- Reproduced on stock PostgreSQL 18.4 using only core GiST, with zero
turbovec loaded — thin diagonal polygons, where the bounding-box
distance under-estimates the true polygon distance, so
was_exactcomes out false and tuples are routed throughnodeIndexscan.c’s reorder queue. - Traced line by line in stock source.
indexam.c:983setsxs_recheckorderbyfor any AM that asks;nodeIndexscan.c:290queues the tuple;reorderqueue_popcallsExecForceStoreHeapTuple, whoseTTS_IS_BUFFERTUPLEbranch callsExecClearTuple(which doesItemPointerSetInvalid(&slot->tts_tid)) and never restorestts_tidfromtuple->t_self;tuptable.h:420’sslot_getsysattrreturns&slot->tts_tidfor the ctid system column. The siblingtts_heap_store_tupledoes settts_tid, so it is an asymmetry, not a design choice. The fix is proven, not proposed. PostgreSQL built both ways on one machine, one script:
unpatched patched ctid self-join, expect 5 1 5 UPDATE ... WHERE ctid, expect 5 rows1 5 sentinel ctids at LIMIT 5049/50 0/50 The
UPDATEline is the dangerous one: no error, it just affects one row instead of five.- Byte-identical across branches: the affected function body hashes the same in the 13.23, 14.22, 15.17, 16.14, 17.9 and 18.3 trees, so the one-line fix applies unchanged to every supported major.
- No query-level workaround exists (verified, not assumed):
WITH ... AS MATERIALIZED, casting totextinside a subquery, and extra subquery nesting all still return the sentinel, because it is already in the slot before any of them run. Only abandoning the index scan avoids it. So the documented guidance — chain on your own key column, or useturbovec.knn()— is the only real answer, and now it is known to be rather than assumed to be.
New finding. turbovec sees this on 100% of rows while core GiST
loses only some. IndexNextWithReorder sets was_exact = (cmp == 0), so a
tuple skips the queue when the AM’s advertised ORDER BY value compares
exactly equal to the recomputed one. turbovec advertises
f64::NEG_INFINITY, which never compares equal to a real distance, so
every one of its tuples is queued. That is a difference in exposure, not
in cause — advertising a real lower bound would only make the fault
intermittent, which is harder to diagnose, not safer, so the -inf choice
stands.
Ships docs/upstream/bug6-pgsql-hackers-DRAFT.md, a submission prepared
for review and not sent, alongside the existing patch file (now
carrying the A/B table).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.1'; — no REINDEX. Nothing in
the extension’s behaviour changes; the fix this release documents is a
PostgreSQL core patch, not something pg_turbovec can apply.
[2.7.0] — 2026-09-08
IVF + 1-bit sign-BQ compose, the Hamming kernel gets ~4.4× faster, and two corruption-class bugs shipped by v2.6.0 are fixed. No wire-format change (stays v8), no REINDEX, no SQL surface change.
Two bugs in v2.6.0’s BQ code — fix these by upgrading
Found by an audit while composing IVF with BQ, not by a field report:
MetaPageData::set_ivf_chainsomittedbq_mean_countfrom its running chain-offset sum. An IVF+BQ build would have written the coarse-centroid chain on top of the corpus mean — destroying the centring vector that every sign code and every query depends on. This is the fourth occurrence of this bug class (v1.24.0 omittedgraph_count; v2.6.0 found and fixed three sites).set_graph_chainhad the identical omission (unreachable today, since graph + 1-bit is rejected) and is fixed too. Now guarded by an exhaustive pairwise no-overlap test, verified to fail when the omission is reintroduced.- Flat-BQ
aminsertdid not re-persist the tombstone bitmap, so everyINSERTafter aVACUUMsilently resurrected every deleted row — the same M2 bug the graph kind fixed in v2.1.0, reintroduced because the BQ write path had no tombstone parameter. It also appended unconditionally, so re-inserting an existing heap TID added a second slot for the same row (unbounded growth under upserts, plus a duplicate id — the shape the flat bijection guard calls corruption). Both fixed; the write path now carries tombstones through a single meta write.
Only bit_width = 1 indexes were affected, and only on insert-after-VACUUM
(resurrection) or re-insert (duplicate slots). 2/¾-bit indexes were never
affected. There is no on-disk format change, so upgrading is sufficient —
but an existing 1-bit index that has taken inserts after a VACUUM should be
REINDEXed, since it may hold resurrected or duplicated slots.
WITH (lists = N, bit_width = 1)
Previously rejected with a clear ERROR. Now builds and scans: sign codes are
stored cell-contiguous (the permutation IVF already applies) alongside
the coarse centroids and cell directory, so a scan probes only the nearest
turbovec.probes cells and runs Hamming within them — the dim/8 storage
win combined with the probe-a-fraction scan win. Wire version stays 8: the
shape is kind = KIND_BQ plus the existing v4 IVF chain fields.
BQ cells live in the raw L2-normalised space, not the rotated space TurboQuant IVF uses. TurboQuant trains cells in the rotated space because that is where its fine quantizer encodes; a BQ code is the sign of the centred raw coordinate, so there is nothing to align to. Getting this asymmetric is the sharpest available silent failure — a rotated query against un-rotated centroids probes the wrong cells and collapses recall with no error — so the scan skips the rotation explicitly and the tests assert per-id self-neighbour recovery.
An IVF+BQ aminsert degrades the index to a flat BQ scan (it appends
and drops cell metadata). This is observable: lists is preserved and
ivf_degraded is stamped, so turbovec.index_is_degraded() reports it —
strictly better than the TurboQuant IVF insert path, which blanks lists
and cannot be reported. Rebuild with REINDEX to restore cell-scoped
scanning.
Hamming kernel: ~4.4× faster, and AVX2 measured then declined
The kernel now counts 8 bytes at a time instead of one. Measured 4.4–4.8× at embedding dims (100k×768d: 6.13 ms → 1.28 ms; 1M×768d: 59.0 ms → 13.7 ms; independently reproduced at 4.5–5.4× on a second machine). Stable from L3-resident to firmly-DRAM sizes, so the scan is not bandwidth-bound at these sizes.
No unsafe, no runtime CPU dispatch — every machine and architecture
runs the identical instruction sequence. An AVX2 intrinsics kernel was
written and proven bit-identical, then declined on measurements: it
buys only ~1.8× more above dim = 512 and is a net loss below it (the
32-byte loop never runs at dim ≤ 256), while costing a second unsafe
block on the scan hot path, a dim-dependent dispatch threshold, and a code
path CI cannot exercise both sides of — the exact blind spot that let the
v1.7.3 wrong-results bug ship. u64x4 was also measured and rejected (LLVM
already extracts the ILP). Full numbers in docs/ONEBIT_BQ.md §7.
Bit-identity is proven in-tree, not assumed: 5720 random code pairs across
143 dims against a bit-by-bit reference sharing no code with the kernel,
825 top-k cases asserting the full sequence including tie order, a
tie-density guard so that is not vacuous, and mutation testing (dropped
tail, byte-swapped operand, AND for XOR, skipped last word,
count_zeros) where each mutation fails 3–6 tests.
Benchmark harness for the unpublished BQ frontier
benches/scripts/bq/ plus docs/BQ_RECALL_BENCH.md: a sweep over
bit_width ∈ {1, 2, 4} measuring R@10 against exact in-DB ground truth,
storage, build time and — only on an AVX2 host — latency. It has not
been run; no numbers exist yet. The AVX2 gate is structural: on a scalar
host the driver does not time queries at all, rather than emitting a number
that reads as fast. Ground truth uses the same operator the index serves
(avoiding the cosine-vs-L2 trap), the re-rank window is recorded per row
against a mirror of the Rust clamp, and iso-recall rows emit null when a
bit_width never clears the target — that absence being the result.
Predictions are written down in advance so a real run can falsify them.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.7.0'; — no REINDEX for
2/¾-bit indexes. A 1-bit index that has taken inserts after a VACUUM
should be REINDEXed (see the bug notes above).
[2.6.0] — 2026-09-07
1-bit sign binary quantization (WITH (bit_width = 1)) works end to
end — build, scan, aminsert, VACUUM. No wire-format change (stays
v8), no REINDEX, no SQL surface change. Previously the reloption was
accepted for forward compatibility but the build raised a clear “not yet
implemented” ERROR.
1-bit is not TurboQuant-at-1-bit: the turbovec crate hard-rejects
bit_width < 2 in both constructors, so this is a distinct scheme — sign
binary quantization, the DiskANN/pgvector/Qdrant coarse code. Per-vector
storage is dim/8 with no per-vector scale (half the 2-bit stride,
and it also drops the codebook, the rotation/TQ+ chain and the blocked
chain), scored by Hamming (popcount(XOR)) and then exactly reranked
by the AM’s existing xs_recheckorderby machinery.
Why there is no wire bump
A new kind byte (KIND_BQ = 3), not a version bump. Every existing
index keeps kind = SINGLE/COLBERT/GRAPH and decodes byte-identically —
the same additive per-kind path v4→v5→v6 used. The three bq_mean_* meta
fields occupy page offset 316, reserved-and-zero on every prior version,
so an old meta page reads them as “absent”. (docs/ONEBIT_BQ.md §4
originally specified a 7→8 bump; that was written when v7 was current.)
Mean-centering is load-bearing
The naive sign-at-zero rule sets every bit on dense-positive data (measured R@10 = 0.0 on GIST), so the per-dim corpus mean is subtracted before taking signs, persisted, and applied to queries too. A corpus still collapsed after centering (constant / near-constant) is rejected at build rather than shipping a signal-free index. Scanning an index whose mean is missing ERRORs with a REINDEX hint rather than returning garbage.
Three instances of the v1.24.0 corruption class, found and fixed
write_tombstones_and_meta, the tombstone placement inside the rewrite
path, and MetaPageData::total_blocks() each summed chain page counts
without bq_mean_count. On a BQ index that would have placed the
tombstone chain on top of the mean vector and under-sized the
relation — the identical shape of the v1.24.0 graph bug (which omitted
graph_count). Found by auditing every chain-offset sum in the tree, not
only the path being added.
Notes
aminsertdoes not recompute the mean: that would invalidate every sign code already packed against the old mean, silently degrading the whole index’s ranking from one insert. Build-time mean is fixed; drift is a REINDEX concern. The insert is one exclusive lock across read-modify-write — an unlocked RMW is exactly what silently lost graph rows before v2.1.0.- VACUUM is tombstone-only, sharing the graph kind’s path (compacting would mean renumbering every slot).
turbovec_checkreportskind = 'bq'and skips the v2.2.2 scales validation for that kind only — BQ has no scales chain, so validating one would report every BQ index corrupt.- The Hamming kernel is deliberately scalar (
count_oneslowers toPOPCNT); any hand-vectorised version must first be proven bit-identical against it. This is the v1.7.3 lesson — a mis-specialised kernel returned wrong ANN results on pre-AVX2 CPUs — applied pre-emptively.
Not yet supported
bit_width = 1 with lists > 0 (IVF) is rejected with a clear ERROR; the
cell-contiguous sign layout and per-cell Hamming scan are unwired.
bit_width = 1 with graph = true stays rejected in the reloption
validator (and the graph kind is deprecated as of 2.5.0). BQ’s
recall/latency frontier on real corpora still needs an AVX2-host run; the
tests prove correctness, storage and scan behaviour, not a published
frontier.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.6.0'; — no REINDEX.
[2.5.0] — 2026-09-07
Two changes: the graph kind is deprecated, and Phase S-1 partition pruning lands (additive SQL). No index wire-format change (stays v8), no REINDEX.
WITH (graph = true) is deprecated
It now emits a deprecation WARNING; the build path is scheduled for
removal, with decode retained one further release so a stale graph index
fails loudly with a REINDEX hint rather than silently.
The kind was added in v1.23.0 to chase HNSW’s query latency while keeping TurboQuant’s storage compression. Measured at matched recall it never delivers. Its apparent sublinearity holds only at iso-beam (p50 1.11× for a 10× corpus — but recall falls 0.605 → 0.472); once recall is held equal the curves diverge and never cross:
| corpus / target | flat | IVF | graph |
|---|---|---|---|
| SIFT-1M/128d, R@10 ≥0.95 | 0.98 ms, qps@8 1380 | 1.8 ms, qps@8 2039 | 26.2 ms, qps@8 299 |
| GIST-1M/960d, R@10 ≥0.95 | 5.88 ms, qps@8 279 | 11.3 ms, qps@8 480 | unreachable |
| GIST-10M/960d, R@10 ≥0.98 | 34.2 ms, qps@8 31 | 28.4 ms, qps@8 161 | unreachable |
It also loses on build time (57–90×), storage, and has no out-of-core
path. It is IVF, not the graph, that beats flat’s O(n) wall. Use the
default flat index below ~1M vectors and WITH (lists = N) at scale.
Retained deliberately: turbovec.graph_ef, pack::repack, and the
coarse_graph work (Phase G-1) — that one navigates centroids, and
IVF’s win partly rests on it. Deprecating the graph kind is not
deprecating graph techniques.
Also recorded: the “60× parallel build speedup” was an artefact.
graph_build_partitions_decide coupled shard count to thread count, and
shards cost recall (GIST-1M R@10 0.920 at P=4 → 0.605 at P=83); threads at
a recall-preserving P buy <5×. The partitioned-build parity test is
annotated with the blind spot that hid this (2.5k rows/shard at dim 64,
versus ~12k at 960d in reality) rather than re-tuned, since the kind is on
its way out.
Phase S-1: partition-level coarse quantizer
At 1T scale the design is a partitioned parent with one turbovec index per
partition and native Merge Append doing scatter → gather. That is correct
for any N but the per-query fan-out is O(N) in the partition count: at
~100k partitions, opening an Index Scan per partition dominates. S-1
lifts IVF one tier up — each partition gets a summary (its mean in the
original, un-rotated space, so summaries are comparable across
partitions; each partition trains its own rotation, so the persisted coarse
centroids are not) — and nearest_partitions() returns the Kp nearest
so the caller fans out to Kp instead of N.
New SQL (additive): table turbovec.partition_summary, plus
turbovec.refresh_partition_summary(parent, vec_col) and
turbovec.nearest_partitions(parent, query, k_partitions, metric).
S-1 was first attempted in v1.29.0 and reverted when its #[pg_test]
failed CI. The assertion was wrong, not the code: it demanded the
pruned top-k be identical to the full-fan-out top-k, but pruning is an
approximation by construction — scoring a query against partition means
cannot be equivalent to scanning every partition, so on unclustered data a
Kp < N probe can legitimately miss a true top-k member. The revived test
asserts what pruning actually guarantees: exactly Kp partitions returned,
the query’s own cluster ranked first on content-clustered data, and a
recall floor against full fan-out.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.5.0'; — no REINDEX. Existing
graph indexes keep working and will warn on rebuild.
[2.4.0] — 2026-09-07
WAL amplification follow-up: stable chain starts. No wire-format change (stays v8), no SQL surface change, no REINDEX. Minor rather than patch because the on-disk allocation layout of new/rewritten indexes changes (chain contents and decoding do not).
v2.3.0 stopped WAL-logging pages whose contents hadn’t changed, which cut
the reported ~500 MB-per-commit to ~28 MB. This release addresses what was
left. Chain starts were packed back-to-back — scales_first = codes_first
+ codes_count, ids_first = scales_first + scales_count — so any
growth in the codes chain moved every scales and ids page to a new block
number, and a relocated page is genuinely different, so it had to be
rewritten. On the reporter’s 768d/4-bit index a codes page holds only
21 rows, so essentially every flush crossed an allocation boundary and
paid ~27.6 MiB to move the scales+ids chains.
- The three growing chains' allocations are now rounded up to a multiple
of
MetaPageData::PAD_PAGES(256 pages), so a shift happens once per ~5400 rows instead of once per 21. Modelled on that index, WAL per row falls from ~55 → ~5.7 KiB at batch=512, and ~1375 → ~37 KiB at batch=1, on top of v2.3.0’s own ~33×. - Cost is bounded slack: at most
3 * PAD_PAGESpages (6 MiB) per index, independent of index size. Chains smaller than one padding unit are not padded, so small indexes keep the previous layout byte-for-byte — without that gate a 1000-row 768d index would have gone 0.4 → 6.0 MiB (15×) to buy an amortisation it never needs. padded_pages_needed()is the single definition, used by the planner and the incremental grow/shrink paths inrelfile.rs. If one site padded and another didn’t they would disagree about where the following chains start — precisely the class of block-offset bug that silently corrupted a graph index in v1.24.0.
Why this is safe without a REINDEX: *_count keeps its existing meaning
of “blocks allocated to this chain”, which is what every consumer
already uses it for (sizing the relation, and locating the next chain via
running sums). A chain’s contents are located by *_first plus
n_vectors/rows_per_page and never by *_count (see read_chain), so
a padded index decodes identically. Existing unpadded indexes keep working
as-is; their chains relocate on the first full rewrite, which is safe by
the v1.29.4 invariant (all chains are written before the meta page, so
an interrupted rewrite leaves the old meta pointing at the old, intact
chains).
Four new layout tests cover the invariants directly: chain starts don’t move when up to 10k rows are added; padding never changes content addressing; small and empty indexes aren’t padded; and across a dim × bit_width × n sweep every padded extent is large enough for its rows with overhead bounded by one padding unit.
docs/TESTING.md gains a section on measuring global counters, distilled
from the three self-inflicted CI failures in v2.3.0 — concurrent
#[pg_test]s share one cluster, so pg_current_wal_lsn() deltas capture
other tests' WAL (the same operation measured 32 KB / 1097 KB / 1425 KB
across runs and once ranked the arms inverted). Count a backend-local
value you control, reset each arm to an identical starting state, and
assert the control arm is non-trivial so “0 vs 0” can’t pass.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.4.0'; — no REINDEX.
[2.3.0] — 2026-09-07
WAL amplification fix: a flush now only WAL-logs the index pages that actually changed. No wire-format change (stays v8), no SQL surface change, no REINDEX. Minor rather than patch because insert-path write behaviour changes materially.
A field report (2026-09-08, pg.ddx.io) measured pg_turbovec inserts
accounting for ~100 % of all WAL on the host: 1322 MB / 25 s with
an embedding backfill running versus 31 kB / 25 s with that one unit
stopped — a ~42,000× difference from a single writer doing
~100–1400 rows/minute. The decisive detail: pg_stat_statements
attributed only 434 kB of a ~3 GB/60 s window to SQL, so >99.98 % of
the WAL was generated outside any statement — index maintenance, not
the INSERT.
WAL per commit was roughly constant and close to the 882 MB index
size (~500 MB at 16 rows/txn, ~754 MB at 128, ~670 MB at 512) and WAL
per row a clean 1/batch curve with no knee — the signature of the
whole relfile being rewritten and fully WAL-logged on every flush. A
scratch-index A/B isolated it to commit boundaries alone: 32 × 16 rows =
643 MB WAL versus 1 × 512 rows = 2.6 MB (245×), both after
wal_compression=zstd. Cost: ~4.3 TB WAL/day, archived off-host so
paid for twice, and the dominant consumer of the host NVMe’s endurance
(49 % used, 172 TB written in 5306 power-on hours). It also kept
num_requested checkpoints at ~2× num_timed with a 220–390 % FPI
ratio, which sent the operator chasing a checkpoint misconfiguration.
Cause: write_chain_at registered every page of every chain with
GENERIC_XLOG_FULL_IMAGE. But reconcile_flush_image appends new slots
at the end and updates touched slots in place, so every other page
was already byte-identical on disk — we were paying a full-page image to
rewrite pages with their own contents.
- Each full page is now compared against the bytes about to be written
(under a shared buffer lock) and skipped when identical, so WAL scales
with bytes changed rather than index size. Only pages already
carrying our own no-hole header are eligible, so a page whose header
GenericXLogFinishwould treat as a hole is never inherited; a skipped page is by definition already correct, so crash recovery is unaffected. - The
GenericXLogstate is started lazily, so a batch in which every page is skipped emits no WAL record at all instead of an empty one. - Removed the dead, never-called
write_chain()helper. It wrote pages withMarkBufferDirtyand no WAL — a latent footgun that would have made a skipped page’s missing WAL permanent. Every page mutation now provably goes throughGenericXLog(zero liveMarkBufferDirtysites). Test
insert_wal_scales_with_change_not_index_sizedrives a real flush (via the existingflush_to_relfile_for_testhook — a#[pg_test]’s outer transaction always rolls back before PreCommit fires, so the aminsert path can’t be observed in-band) on a ~109-chain-page index and asserts a no-change flush WAL-logs <1/10th the pages of a full rewrite. Fail-before/pass-after viaSKIP_UNCHANGED_PAGES.It counts pages registered for WAL rather than diffing
pg_current_wal_lsn(), which is load-bearing:#[pg_test]s run concurrently against one cluster, so a global LSN delta also captures every other test’s WAL. Three CI runs reported the same operation as 32 KB, then 1097 KB, then 1425 KB, and once ranked the two arms inverted — the LSN measurement was never valid in either direction. Registered pages are per-backend, deterministic, and are the direct driver of WAL volume (one full-page image each).Note the residual cost on an append (as opposed to a no-change flush): chain starts are packed back-to-back (
scales_first = codes_first + codes_count), so when the codes chain grows by one page every scales and ids page lands at a new block number and genuinely must be rewritten. Measured 392 KB vs 1416 KB always-rewriting on a 20k-vector/64d index. Making chain starts stable across growth would shrink that further and is left as follow-up.
Batching still matters and the docs now say so: WAL scales with
commits, so a flush’s remaining cost (tail pages, meta page, IVF cell
directory) is paid per transaction. docs/PRODUCTION.md gains a “WAL cost
of inserts — batch your writes” section including how to measure it with
LSN deltas, since pg_stat_statements will not show it.
Also documents the partial-index verification gotcha (Ask 3 of the
2026-09-05 report): a KNN query that omits a partial index’s predicate
silently gets a sequential scan, which masks index faults as latency
problems — always confirm EXPLAIN shows Index Scan using <name>.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.3.0'; — no REINDEX.
[2.2.2] — 2026-09-05
turbovec_check() no longer has a scan-fatal blind spot. Code-only —
no wire-format change (stays v8), no SQL surface change (the reason
column already exists as of 2.1.0), no REINDEX required to upgrade.
A field report (2026-09-05, pg.ddx.io) found an IVF index — ~2.18 M
vectors, 768 d, bit_width = 4, lists = 1400 — on which every KNN
scan failed instantly with turbovec’s
InvalidScaleValue { slot: 1, value: -2.559434e22 }, while
turbovec_check() reported is_corrupt = false, count_matches = t,
duplicate_id = NULL. Because the operator’s automated self-heal trusts
that check, the index degraded silently until a manual REINDEX
(~12 min) cleared it. The reporter correctly identified this as the more
important of the two issues they filed: a checker that cannot detect an
index failing 100 % of scans is a false-confidence generator.
Root cause of the blind spot (a deliberate, documented decision that this
incident falsifies): the check read only the meta page and the ids chain,
skipping “the much larger codes/scales chains”. But the scales chain is
the cheap one — one f32 per vector, 8.7 MB at 2.18 M vectors,
versus the 17.4 MB ids chain it already read, versus 0.84 GB
for the codes chain it rightly still skips — and it is load-bearing for
TurboQuantIndex::from_parts, i.e. exactly the gate a scan hits.
turbovec_check()now validates every persisted scale (finite, non-negative, sane magnitude — the same invariantsfrom_partsenforces) for every index kind, inside the same shared rewrite-lock bracket as the meta/ids read so all three are one consistent observation, and names the failing slot inreason. Cost is +50 % on an already-cheap query; no opt-indeepflag, because a default that cannot see a scan-fatal fault is the bug.- The scan-path rejection is now a proper PostgreSQL
ERRORnaming the fault and theREINDEX INDEXrecovery, instead of the bare Rust.expect("...from_parts rejected raw parts")string an operator used to get with no hint that REINDEX fixes it. - Regression test
turbovec_check_detects_invalid_scalewrites the reported value (-2.559434e22) into the persisted scales chain of an IVF index via a page-straddle-safe test-only helper and asserts the checker flags it; fail-before/pass-after.
Note this is the same failure shape as the graph lost-update called out
in the 2.1.0 notes: internal self-consistency of meta+ids is not
evidence of scannability. The originating corruption itself (a damaged
scales chain on an index that had survived several
pg_cancel_backend/pg_terminate_backend events, a kill -QUIT
immediate shutdown and a restart that aborted an in-flight backfill) is
not yet reproduced and remains under investigation; this release makes it
detectable and actionable rather than silent.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.2.2'; — no REINDEX to upgrade.
(An index already damaged needs a one-time REINDEX INDEX <name>;, which
this release will now actually tell you about.)
[2.2.1] — 2026-09-05
Safety patch: the PARALLEL graph build was effectively uncancellable. Code-only — no wire-format change (stays v8), no SQL surface change, no REINDEX.
v2.1.0 made the graph build interruptible, but its hook lives in
thread-local storage and is therefore invisible to rayon workers by
design (a worker must never longjmp out of the pool). That left a
hole nobody had measured: during the partitioned build the driver parks
in a futex inside rayon’s join for the entire parallel phase, so it
cannot reach a poll either. Measured on a 10M-node build:
pg_cancel_backend() and a direct SIGINT were both ignored for
over 13 minutes while 32 threads ran at 100% CPU — and because no
backend code was executing, pg_stat_activity reported
wait_event = NULL, so the runaway build looked idle. An operator had
no way to stop it.
Workers now consult a cheap, thread-safe abort predicate at their
safe points and stop producing work; the driver raises PostgreSQL’s real
cancel/terminate error at its own poll once the phase collapses, so all
error raising still happens only on the backend thread. The predicate
is injected by the PG-side caller, so src/index/graph.rs stays
Postgres-free and the mechanism is unit-testable without a server
(partitioned_build_bails_out_when_abort_is_requested).
Found while measuring whether the graph kind could be made
production-grade at scale (see the 2.2.1 note in docs/UPGRADING.md and
the deprecation discussion for v2.3.0).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.2.1'; — no REINDEX.
[2.2.0] — 2026-09-04
Graph-kind scan-beam retune. The graph kind’s scan-time beam width
is now its own knob (turbovec.graph_ef, default auto = 512) instead of
a side effect of the candidate count, and the auto default is set from a
measured 1M-scale recall-vs-latency frontier on SIFT-128 and GIST-960.
No wire-format change (stays v8), NO REINDEX — but a new GUC is a
SQL-surface addition, so this is a MINOR.
Fixed — the graph “high-dim recall cliff” was a BEAM bug, not a scorer bug
The tracked “GraphScorer diverges from the kernel at dim >= 512” is
disproved. An exact-f32-oracle control gives recall identical to the
LUT scorer (within 0.015), and recall recovers monotonically with the
BEAM on the same index and the same scorer (search_k 32 → 128 → 512
gives R@10 0.750 → 0.990 → 1.000). The real mechanism was the scan-time
beam width, and — worse — which knob was setting it:
- Through v2.1.0,
graph_searchcomputedef = (k * 4).max(64)wherekwas the candidate countscan.rshad already widened. Sinceturbovec.hi_dim_rerank = autoraises that count toclamp(dim, 256..=1024)fordim >= 256— a knob whose actual job is the flat/IVF exact-rerank window over a cell scan’s quantized ranking — the graph’s beam was being set as a side effect of an unrelated feature. - The visible symptom was an inverted cliff. With
hi_dim_rerank = auto, 128d was the WORST (R@10 0.720: 128d is below the rerank threshold, so the beam stayed at the bare 64 floor) while 384/512d looked perfect (0.99/1.00: the inflated candidate count multiplied the beam up). Withhi_dim_rerank = offrecall collapsed at EVERY dim (~0.40/0.38). Recall was never dim-dependent; it was beam-dependent, and the beam was accidental.
Fixed by splitting graph_search into graph_search_with_ef(.., ef, ..)
(the live scan path, beam resolved from the new GUC) and the original
graph_search (the pure-k fallback for callers with no live GUC —
unit tests, the aminsert findability probe). scan.rs’s graph arm now
passes the user’s OWN candidate count (search_k * oversample), not
hi_dim_rerank’s floor. hi_dim_rerank keeps doing its real job for
flat/IVF, unchanged.
Verified on 42 paired (k, ef) configs across SIFT-200k, SIFT-1M and
GIST-1M: off and auto now give recall within 0.0000 of each other
at every beam, at 128d AND at 960d.
Added
turbovec.graph_ef(int,Userset, default0= auto = 512, range0..=1000000) — the graph kind’s recall/latency dial, the direct analogue ofhnsw.ef_search(and ofturbovec.probesfor IVF).0= auto; a positiveNpins the beam. Always clamped up to the query’s candidate count (a beam narrower than it cannot fill it) and down to the live corpus size. Pure scan-time knob: no wire change, no REINDEX, honoured immediately by any graph index built by any 2.x binary.
Changed — defaults
Graph scan beam default 64 → 512 (
GRAPH_SCAN_EF_DEFAULT), and it is now ABSOLUTE rather than(k * 4).max(64). Measured frontier (32 vCPU AVX-512, 1M rows, R@10 vs exact cosine GT, warm p50,search_kpinned tomax(32, k)):corpus beam 64 beam 512 beam 2048 SIFT-1M (128d) R@10 0.943 / 1.25 ms R@10 0.990 / 5.19 ms R@10 0.991 / 14.13 ms GIST-1M (960d) R@10 0.760 / 8.65 ms R@10 0.920 / 34.55 ms R@10 0.966 / 93.44 ms At 128d 512 is the unambiguous knee: it is within 0.001 R@10 of the 2048-beam ceiling at 0.37× its p50, and doubling past it buys +0.001 for 1.7× the latency. At 960d the curve has NOT plateaued by 2048, so 512 is a deliberate latency cap — the widest beam that keeps 960d p50 under ~35 ms — not a knee;
SET turbovec.graph_ef = 2048reaches 0.966 for ~95 ms if that is the trade you want.⚠️ One configuration REGRESSES on recall, deliberately. A 960d graph index queried with
hi_dim_rerank = auto(the default) goes R@10 0.971 → 0.920 and R@100 0.958 → 0.826, in exchange for p50 194 ms → 34 ms (5.6×) and 196 ms → 42 ms (4.7×). The old numbers came from the side-effect beam of 3840 thatautowas silently buying at 960d — never a documented beam, and it evaporated the moment a user sethi_dim_rerank = off(0.840, the “collapse”). A ~200 ms p50 is not a defensible default, and the old recall is now REACHABLE and documented for the first time:SET turbovec.graph_ef = 3840reproduces the pre-patchautobeam exactly. 128d indexes improve in BOTH modes (R@10 0.976 → 0.990 at 1M) and lose nothing — 128d is belowhi_dim_rerank’sdim >= 256threshold, so 128dautowas never aboveoff.
Known limitation — the graph kind is NOT the fastest kind at 1M
Measured on the same box, same corpora, same GT, at k=10: the flat
kind beats the graph on BOTH recall and latency on BOTH corpora.
SIFT-1M flat R@10 0.993 / 0.96 ms vs the graph’s best-defensible 0.990 /
5.19 ms (5.4× lower latency, higher recall). GIST-1M flat R@10 0.997 /
5.91 ms vs the graph’s 0.920 / 34.55 ms (5.8× lower latency, +0.077
recall) — and even beam 2048 (0.966 / 93 ms) does not catch flat. At 1M
rows on a 32-core AVX-512 host, turbovec’s SIMD full scan (one linear
sweep of a compact quantized buffer at memory bandwidth) is simply
faster than navigating a graph (~ef scattered gathers with a serial
dependency between hops). The graph’s advantage is asymptotic and 1M is
below the crossover on this hardware. Guidance: use flat (or IVF for
a smaller RAM footprint) at this scale. The graph’s one measured win
is CONCURRENCY scaling — the flat scan saturates memory bandwidth and
goes 713 → 1313 qps from 1 to 8 clients (1.8×) while the graph goes 688
→ 4906 (7.1×) — so it is the better throughput engine under load even
where it loses isolated p50. Full curves and the reasoning in
docs/GRAPH_EF_BENCH.md.
Docs
docs/GRAPH_EF_BENCH.md— the full measured frontier (both corpora, bothhi_dim_reranksettings, the wholeefgrid, atk= 10 and 100), the flat/IVF comparators at the samek, the before/after A/B on byte-identical index files, and the reasoning behind 512.docs/PRODUCTION.md—turbovec.graph_efreference section.README.md— GUC table + count (18 → 19).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '2.2.0'; and restart the backend
(the new GUC is registered in _PG_init). No REINDEX: the wire
format is unchanged (v8) and the beam is resolved at scan time, so
existing graph indexes get the new default immediately. Flat and IVF
indexes are unaffected — the beam does not exist on those paths. To
restore a pre-v2.2.0 effective beam exactly: SET turbovec.graph_ef =
128 for what hi_dim_rerank = off gave at any dim (and what auto
gave below 256d), or = 3840 for what auto gave at 960d.
[2.1.0] — 2026-09-04
Graph-kind correctness release. The WITH (graph = true) kind had a
critical concurrent-INSERT data-loss bug, could return fewer than k
rows, and its build could not be cancelled. All three are fixed here,
plus graph health monitoring and an upstream-bug writeup. No
wire-format change (stays v8), NO REINDEX — but turbovec_check()’s
return type changes, so this is a MINOR (see the upgrade note).
Fixed — graph kind
- C1 (CRITICAL): concurrent
INSERTinto a graph index lost data.insert_graph_rowdid an unlocked whole-index read, then a blind whole-relfile rewrite. Two concurrent graph inserters produced either (a) a torn adjacency chain —corrupt graph adjacency chain: graph offsets[n]=54272 != neighbors.len()=54240, a loud ERROR — or, worse, (b) a silent lost update: the losing inserter’s rows were simply gone (heap rows committed with no index entry) whileturbovec_check()still reportedis_corrupt = false, because the surviving relfile was internally self-consistent. Measured fail-before: 8 of 100 concurrent inserts lost, 8 heap rows unindexed. Fixed by holding ONElock_relfile_writeacross the entire read-modify-write, re-reading the meta page under it, and adding guards that ABORT rather than reconcile onto torn state. Also folded in the tombstone two-write gap: the bitmap is now planned into the single meta write (previously a second write re-attached it, and a crash in that window permanently un-tombstoned every VACUUM delete). Validated: 320/320 concurrent inserts clean, and a sustained 240 s run of 6 inserters against a concurrent DELETE/VACUUM loop stayedis_corrupt = falsewith zero SIGABRT. - BUG#2:
graph_searchcould return fewer thankrows. Two real causes: tombstoned nodes are never routing hops, so a large dead set disconnects the live remnant (n=2000 at 99% dead returned 2 rows fork=10); and a degenerate corpus of tied distances collapsesRobustPrune’s out-lists (500 identical rows returned 8 fork=10). Fixed with a bounded backfill from unvisited live slots, gated behind the existingout.len() >= kguard so the normal path is byte-identical. - FINDING#2: a graph
CREATE INDEXcould not be cancelled and hungpg_ctl stop -m fast(the rayon build had zero interrupt polling). Added driver-thread stage polls mirroringivf_build_and_write, plus a TLS-scoped hook the Postgres-free graph module calls at its own stage boundaries and every 256 nodes — a rayon worker cannot see the TLS slot, so it can neverlongjmpout of the pool.
Documented — not a pg_turbovec bug
- BUG#6: a kNN index scan projects a sentinel
ctid(4294967295,0)instead of the real heap tid. Root-caused to PostgreSQL core, not this extension:ExecForceStoreHeapTuple()’s buffer-slot branch callsExecClearTuple()(which invalidatesslot->tts_tid) and never restorestts_tidfromtuple->t_self, while its siblingExecStoreHeapTuple()does — a plain asymmetry.slot_getsysattr()reads the projectedctidstraight out oftts_tid. Any AM that setsxs_recheckorderby = trueis affected; core GiST reproduces it with no turbovec loaded, and a one-line core patch fixes both. Row data is unaffected (xs_heaptidis correct), and the documentedturbovec.allowlist-from-ctidrecipe still works (it harvests from a filter scan). Only chaining off a kNN scan’sctidbreaks. The proposed core patch is indocs/upstream/, the limitation and the supported alternatives are indocs/FILTERING.md, and a tripwire test fails the moment upstream fixes core so this note gets retired.
Fixed — BUG#5 (monitoring, details)
- BUG#5:
turbovec_check()was graph-adjacency-blind. It validated the flat/IVF ids bijection + tombstones but never looked at a graph index’s CSR adjacency chain, so a corrupt adjacency (a torn write, or the concurrent-insert corruption) reportedis_corrupt = false— operators had no health signal for the graph kind at all. For akind = graphindex the check now decodes the adjacency chain and validates: offsets monotonically non-decreasing and consistent with the neighbor-array length, every neighbor id< n_vectors, the entry point in range AND having at least one out-neighbor, no self-loops, strictly-ascending (deduplicated) neighbor lists, the adjacency’s node count matchingmeta.n_vectors, and at least one edge whenn_vectors > 1(not trivially disconnected). Cost isO(n + edges)over the adjacency chain only — no vector data is read, so it stays cheap enough to poll. Flat/IVF/ColBERT behaviour is byte-identical (the new code is gated onmeta.is_graph()). - The chain decode used by the monitoring path is now non-fatal
(
relfile::try_read_graph_adjacency), so a malformed chain is REPORTED rather than ERRORing out of the operator’s health query. It also bounds the chain against the physical relation length before reading, so a truncated relfile is reported instead of raisingread_chain’s hard past-EOF ERROR. The scan / insert / VACUUM paths keep the ERROR policy (they cannot proceed on a broken graph) and share the same decode implementation.
Changed (SQL surface — additive column, but see the upgrade note)
turbovec.turbovec_check(regclass)gains a trailingreason textcolumn: NULL when healthy, otherwise a human-readable description of the first problem found (duplicate id, count drift, or the specific graph-adjacency invariant violated). Adding an OUT column changes the function’s return type, whichCREATE OR REPLACE FUNCTIONcannot do — the release that ships this needs aDROP FUNCTION turbovec_check(oid); CREATE FUNCTION ...pair in itssql/pg_turbovec--<prev>--<this>.sqlupgrade script, and it is a MINOR bump at minimum (never a patch). No wire-format change (still v8); no REINDEX.
[2.0.0] — 2026-08-26
MAJOR: pg_turbovec now runs on upstream turbovec 1.0.0. This adopts the stable turbovec 1.0.0 crate (Ryan Codrai’s TurboQuant implementation, first stable release) in place of the long-lived 0.9.0 fork. On-disk wire format v7 → v8 — a breaking change, with a documented REINDEX-from-heap migration (below).
What changed
- turbovec 0.9.0 → 1.0.0. The fork’s hand-rolled centroids/boundaries
codebook + QR rotation are replaced by turbovec 1.0.0’s TQ+
per-coordinate calibration and v5 block-Hadamard rotation (which
also drops the OpenBLAS build dependency). Every encoded byte differs
from v7, hence the wire bump. pg_turbovec carries two small additive
fork patches on top of stock 1.0.0 (
pub fn repackand a re-exposedIdMapIndexparts API used by the buffer-manager cache-fill path), both tracked to be offered upstream. - Wire format v7 → v8.
MetaPageData::version = 8,EXPECTED_WIRE_FORMAT_VERSION = 8. - Materially faster, same storage. On EC2 i4i.8xlarge (AVX-512), head-to-head vs v1.29.7: SIFT-1M flat p50 12.9 ms → 1.3 ms (9.6×); GIST-1M IVF p50 @ R@10 0.95 512 ms → 92 ms (5.6×); cold-scan (SIFT IVF) 3139 ms → 399 ms (7.9×); build 1.2–2.1× faster. Index size byte-for-byte unchanged (77 MB SIFT-1M, ~4.96 GB @ 10M). Recall matched-or-better at matched config. Determinism intact (the build-parallelism byte-identity gates all pass; v5 rotation removed the old OpenBLAS nondeterminism). Two minor warm-flat regressions (~11–15% at already-saturated recall), no correctness impact.
- All v1.29.x corruption fixes (meta-LAST torn-write ordering,
reconcile-on-flush, pre-flush validate-all, IVF
lists==0dup-gate, VACUUM shrink guard, read_chain bounds, SubXact rollback) are re-proven to fire on the v8 persist path. The 90-minute no-VACUUM upsert+writer-restart corruption A/B ran clean on v8 (92 checksis_corrupt=f, 0 dup-id, 0 SIGABRT, 26 restarts).
Migration — REQUIRES a one-time REINDEX per index (wire v7 → v8)
ALTER EXTENSION pg_turbovec UPDATE TO '2.0.0'; + restart the backend,
then REINDEX INDEX <name>; once per turbovec index. A pre-v8 index
opened under 2.0.0 is detected (MetaPageData::is_legacy_v7()) and
ambeginscan ERRORs at first scan with a REINDEX INDEX <name>; hint —
never a silent misread (validated end-to-end on EC2: build v7 → open on
2.0.0 → ERROR+hint → REINDEX → serves correctly on v8, same neighbors).
An in-place page converter was investigated and rejected: measured
recall loss of −20.7 pp @ R@10 (SIFT-1M 4-bit) and catastrophic at 2-bit,
from double quantization at re-encode. REINDEX-from-heap re-encodes the
heap’s source vectors directly (full recall) — the heap is the corpus,
so this is not a rebuild-from-external-corpus. See
docs/UPGRADING.md.
[1.29.7] — 2026-08-25
Numerical-robustness patch — normalise_into (run on every indexed
row via normalize_on_insert) computed the reciprocal norm as
(1.0_f64 / norm) as f32, which overflows to +inf when norm is
a tiny-but-nonzero f64 (a vector whose elements are near f32 underflow,
e.g. a single ~2e-39 coordinate). The +inf reciprocal then poisoned
every element (x * inf = inf), so the “normalised” vector had inf
coordinates and infinite norm — feeding garbage into the quantizer for
that row. Now divides per-element in f64 and casts each result to
f32 ((f64::from(x) / norm) as f32), which stays finite because
|x/norm| <= |x| for a real vector. Patch bump — no wire change
(v7), no SQL surface change, no REINDEX.
Found during the pg_turbovec 2.0.0 (turbovec 1.0.0) port’s full test
re-run: the pre-existing Hegel property test
prop_normalise_is_unit_norm_and_idempotent flaked (~1 in N seeds) on
the norm inf assertion. Added a deterministic regression test
normalise_tiny_norm_stays_finite (fail-before/pass-after proven).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.29.7'; — no REINDEX. A row
inserted under an older binary whose vector hit this edge would have
stored a garbage (inf) code; such a row (if any) is corrected by
re-inserting it. In practice the trigger requires near-underflow input
magnitudes, which real embeddings do not produce.
[1.29.6] — 2026-08-15
Dependency-hygiene patch — clears all outstanding RustSec advisories
in the transitive tree. Patch bump, Cargo.lock-only (no source
change, no wire change (stays v7), no SQL surface change, no REINDEX,
byte-identical build output — the gemm, turbovec, and pgrx = 0.19.1
pins are unchanged, so IVF build determinism is preserved).
cargo update pulled semver-compatible upstream fixes:
crossbeam-epoch0.9.18 → 0.9.20 (RUSTSEC-2026-0204, invalid pointer dereference in thefmt::Pointerimpl) — reached viarayon, which pg_turbovec uses for build/scan parallelism. This is the one advisory in a shipped runtime path; pg_turbovec never formats those pointers, so exposure was nil, but the dependency is now patched.tokio-postgres0.7.17 → ≥ 0.7.18 (RUSTSEC-2026-0178),postgres-protocol0.6.11 → ≥ 0.6.12 (RUSTSEC-2026-0179 / 0180) — these reach the tree ONLY throughpgrx-tests, a dev-dependency (the test harness’s PostgreSQL client). They are not present in the shippedpg_turbovec.soand were never a production exposure; patched for a cleancargo audit.
Remaining cargo audit output is two “unmaintained” warnings
(serde_cbor via pgrx, paste via turbovec/statrs and
hegeltest) — not vulnerabilities, and both are upstream-controlled
(pgrx’s PostgresType serialization and a transitive math dep), not
fixable from this crate.
Upstream turbovec: intentionally NOT advanced. Our pinned rev
(befc4cbf, the tip of the feat/unblock-inverse-repack branch,
2026-07-06) carries the pub from_parts + pack::unblock surface
pg_turbovec depends on and is NEWER than upstream main’s last release
(0.6.0, 2026-05-25). main diverged in a slim-down direction that
removes that surface, so advancing the pin would break the build;
rebasing our branch onto main’s encode-vectorization / codebook-caching
improvements is a deliberate, separately-validated kernel change (wire
/ determinism-sensitive), not a routine bump.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.29.6'; — no REINDEX, no
behavior change (dependency hygiene only).
[1.29.5] — 2026-08-15
Production-hardening patch from a deep code re-audit + an at-scale
feature stress test. Patch bump — no wire change (stays v7), no SQL
surface change, no REINDEX (ALTER EXTENSION pg_turbovec UPDATE is
sufficient). Every fix is a guard or ordering correction; none change
results on valid input.
- IVF incremental
INSERTregression fix (introduced in v1.29.4). v1.29.4 added an on-disk duplicate-id guard to the deferred-flush reconcile path, but it ran ungated — and an IVF index legitimately stores the same external id in multiple cells (soft-assignment to the 2nd..Mth nearest cell). So the firstINSERTinto any soft-assigned IVF index trippedrefusing to reconcile onto a corrupt .tvim id table (id N appears in more than one slot on disk)and aborted — breaking IVF incremental insert. The guard is now gated to the bijective flat/single kind (lists == 0), matching the insert/read paths. Regression testivf_insert_after_soft_assign_does_not_falsely_abort. - Multi-index partial-flush corruption-spreader (C-NEW-1). On a
table with 2+ turbovec indexes, if the PreCommit flush of the second
index tripped a persist guard (
ERROR/longjmp), the first index’s relfile pages were already physically written (GenericXLog is not rolled back on abort) — leaving one index with phantom CTIDs and a sibling missing rows. The PreCommit flush now does a pure validate-all pass (no I/O) before writing any index, so a guard trip aborts with zero physical writes. - Corrupt/torn-meta unbounded read + palloc guard (H-NEW-2/3).
read_chainnow bounds the chain against the physical relation length (RelationGetNumberOfBlocksInFork) and rejectsrows_per_page == 0before allocating or walking blocks — a bit-flipped/truncated meta no longer over-reads past EOF or allocates aVecsized by a corruptn_vectors; it raisesERRCODE_DATA_CORRUPTED+ a REINDEX hint. SAVEPOINT/ subtransaction rollback (M-NEW-4). Registered aSubXactCallback; aROLLBACK TO SAVEPOINTnow conservatively invalidates the dirty cache set, so a rolled-back insert is no longer persisted onto disk at top-level commit (was index bloat + silent recall loss, masked by recheck).- VACUUM shrink guard gap (M-NEW-5). The flat swap-remove VACUUM path
— the one whole-relfile mutation that bypassed
write_full_inner’s id-0/dup guard — now re-checks the surviving ids' bijection before committing the shrink (flat kind only). - Empty-index KNN query fix (NEW BUG #1).
ORDER BY emb <-> q LIMIT kon a freshly-created, not-yet-populated index returned 0 rows instead ofERROR: query dim N != index dim 0(the dim check now runs after the empty-index early return) — fixes the “create index, backfill later” pattern. - Overflow hardening (L-NEW-7):
MetaPageData::total_blocks()usessaturating_addso a corrupt meta can’t wrap a large block count to a small total.
The re-audit also VERIFIED (independently, at production scale on EC2):
v1.29.4’s field-report torn-write fix holds clean for 95 min / 85
mid-flush writer restarts (v1.29.3 corrupts after 2), and the whole
flat/IVF feature matrix + the sparsevec OOM guard are solid. Known
issues still tracked for a dedicated release (graph-kind only): graph
concurrent-insert lost-update, graph wrong-results at dim>=512, graph
under-return at 128d, graph build un-cancellable, turbovec_check
graph-adjacency-blind, and the scan sentinel-ctid projection. The graph
kind (WITH (graph = true)) should be treated as experimental.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.29.5'; — no REINDEX, no
downtime beyond the .so swap + reconnect.
[1.29.4] — 2026-08-14
Data-corruption fix: VACUUM-INDEPENDENT torn-write of the .tvim
relfile on writer interrupt/restart. Patch bump — no wire change
(stays v7), no SQL surface change, no REINDEX to upgrade (existing
indexes are read + written correctly in place by the new binary). A
running index that already corrupted needs a one-time REINDEX to
clear the bad state, but the upgrade itself is ALTER EXTENSION +
restart.
v1.29.2/1.29.3 fixed the concurrent-VACUUM lost-update races. A
deployment (pg.ddx.io/agora) that runs no VACUUM at all
(autovacuum_count = 0, no DELETE) still corrupted continuously —
signature always id 0 in more than one slot, ~every 30 min, on a
1.84M-row 768d bit_width=4 partial flat index driven by a
continuous INSERT … ON CONFLICT DO UPDATE upsert writer that the
orchestrator restarts every ~15 min.
Root cause (reproduced + dumped on EC2, unpatched corrupts in ~40 s):
the deferred aminsert PreCommit flush rewrites the whole relfile via
write_full_inner, which polls check_for_interrupts!() inside
write_chain_at and commits each GenericXLog batch as a physical,
WAL-logged page change that is NOT rolled back on transaction abort.
The writer-restart is a pg_terminate_backend (SIGTERM →
ProcDiePending); when it lands mid-flush it longjmps (FATAL) out of
the chain write. With the pre-fix ordering — meta page written
FIRST, then the codes/scales/ids chains — an interrupt after the
meta commit but before the ids chain finished left
meta.n_vectors = N pointing at an ids chain whose newly-appended (or
in-place) region was still the zero-initialised page bytes. On reload
that reads back as a contiguous run of id 0 slots — duplicate ids
… id 0 appears in more than one slot (XX001) + SIGABRT — with
count_matches still TRUE (the chain is physically N slots long). On
disk we dumped the exact fingerprint: a contiguous zeroed run equal to
one batch’s append count (16 slots at n=1,840,016; 730 slots ==
the tail of a single ids page at n=1,840,149), never scattered
garbage. VACUUM is not involved.
Fixed by making write_full_inner write the row CHAINS FIRST and
the META PAGE LAST (the same crash-safety invariant the build path’s
write_blocked_phase_and_meta already documents). An interrupt during
the chain writes now leaves block 0 (the meta) UNCHANGED, so readers
and the next reconcile-flush observe the previous, fully-valid
n_vectors + chains; any pages written past the old count are simply
unreferenced until the next full rewrite / RelationTruncate (which
stays after the meta write, so a shrink never leaves the meta
referencing a page past EOF). The meta write is a single GenericXLog
page op with no interrupt poll, so it cannot tear.
Also added a belt-and-suspenders guard on the deferred-flush
/reconcile path (reconcile_and_write_flush): before splicing this
transaction’s upserts onto the current on-disk ids and re-persisting,
it runs first_duplicate_id over the ids it just re-read and, if the
on-disk state is ALREADY corrupt (e.g. a hole left by an older
binary), ABORTS the transaction with a REINDEX INDEX hint instead of
entrenching it — a retryable abort beats propagating on-disk
corruption.
A NON-restarting, long-lived single writer (no statement_timeout, no
cancels, never pg_terminate_backend’d) never hits the interrupt
window and does not trigger this bug — a valid immediate operational
workaround.
Validation
- Fail-before/pass-after unit test
torn_flush_on_interrupt_meta_last_stays_clean: injects the exact interrupt window and shows meta-FIRST + tear → id-0/duplicate on reload, meta-LAST (fixed) + tear → clean bijection. - Sustained no-VACUUM A/B on the reproduction (1.84M-row 768d
bit_width=4 partial flat index, upsert writer, writer restarted via
pg_terminate_backend): unpatched corrupts in ~40 s at the first restart; patched runs 2 h+ across dozens of restart cycles withturbovec_checkis_corrupt = falsethroughout AND real forward progress (n_vectors climbs cleanly).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.29.4'; — no REINDEX, no
downtime beyond the .so swap + reconnect. An index already corrupted
by a pre-1.29.4 binary needs a one-time REINDEX INDEX <name>;.
[1.29.3] — 2026-08-14
Runtime-hardening patch from a full code audit. Patch bump — no
wire change (stays v7), no SQL surface change, no REINDEX (ALTER
EXTENSION pg_turbovec UPDATE is sufficient; existing indexes are read
and written unchanged). All fixes are defensive guards that turn a
crash/OOM into a clean ERROR; none change results on valid input.
- Adversarial-input OOM guard on the
sparsevecdensify paths.sparsevecpermits a dimension up to 1e9, but the::vectorcast (sparsevec_to_vector) andsum(sparsevec)(SparsevecAccum:: ensure_dim) densified intovec BEFORE the 16000-element vector cap was checked — so one unprivilegedSELECT '{}/1000000000'::sparsevec::vector(orsum()over such a row) could OOM a backend. Both paths now rejectdim > 16000before allocating. Test:sparsevec_oversized_dim_densify_errors_not_ooms. - Clean
ERROR(not a Rust panic across the FFI boundary) on a torn/corrupt scan slot lookup.ReadOnlyIndex::search/search_masked/id_at_slotand the out-of-core IVF id remap indexedslot_to_id[slot]directly; a corrupt/torn read yielding a slot past the id table panicked (a backend-abort risk under load). Nowid_at_slot_checked/id_at_global_checkedraise a PGERRORwith aREINDEX INDEXhint. - VACUUM stays cancellable. The flat swap-remove loop
(
vacuum.rs) holds the exclusive relfile-rewrite lock across every dead slot; addedcheck_for_interrupts!()at the loop top so a VACUUM deleting millions of rows can be cancelled (and doesn’t hold the lock uncancellably against readers) between slots.
The audit also confirmed the v1.29.2 flat/IVF corruption fixes,
determinism gates, dup-id guards, GUC safety, and 1-bit build fencing
are solid. It surfaced one CRITICAL follow-up NOT fixed here: the
graph-kind incremental INSERT still has the pre-v1.29.2
lost-update window (a non-atomic read-then-rewrite) that reconcile-on-
flush closed for the flat/IVF kinds — tracked for a dedicated fix
release with a concurrent-inserter reproduction test + sustained-load
validation per the corruption HARD MANDATE. Graph inserts should stay
serial/one-writer until then (the documented build-then-serve model).
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.29.3'; — no REINDEX, no
downtime beyond the .so swap + reconnect.
[1.29.2] — 2026-08-13
Data-corruption fix: concurrent VACUUM + deferred-flush lost-update
in the flat/single relfile. Patch bump — no wire change (stays
v7), no SQL surface change, no REINDEX to upgrade (existing indexes
are read + written correctly in place by the new binary). A running
index that already corrupted needs a one-time REINDEX to clear the
bad state, but the upgrade itself is ALTER EXTENSION + restart.
Continuous concurrent inserts + autovacuum into a large flat
turbovec.vector index corrupted the .tvim id table (“duplicate ids
… id 0 appears in more than one slot”, XX001) and SIGABRT-crashed
backends. Root cause was a set of concurrency races between the
whole-relfile operations, NOT on-disk-format corruption:
- Deferred-flush lost-update (the load-bearing bug). Each writer
loads the whole index into a per-backend in-memory snapshot on its
first mutation and, at
PreCommit, rewrote the ENTIRE relfile from that snapshot. A concurrent VACUUM that shrank the relfile (deleted dead rows) in between was clobbered — the committing writer resurrected the deleted rows from its stale pre-VACUUM snapshot, reintroducing dead/duplicate ids. Fixed with reconcile-on-flush: under the exclusive rewrite lock atPreCommit, re-read the CURRENT on-disk(codes, scales, ids)and splice ONLY this transaction’s upserted ids onto it (relfile::reconcile_flush_image/reconcile_and_write_flush), never a blind whole-snapshot overwrite. VACUUM’s deletes and other backends' committed inserts survive. - VACUUM stale-snapshot write.
ambulkdeleteread the ids chain, computed dead slots, and only THEN took the write lock for the swap-remove — using the pre-lockmeta/dead_slots. A flush completing in that window left VACUUM swap-removing against a just-rewritten chain. Fixed by taking the exclusive rewrite lock for the WHOLE bulkdelete (ids read → dead-slot compute → swap → shrink → truncate) on a snapshot re-read under the lock. read_rotationunlocked window.install_whole_index/install_graph_indexread the chains under the shared lock, then read the rotation chain UNLOCKED; a concurrent rewrite could move or truncate it. Fixed by bracketing the whole install (chains + rotation + adjacency + tombstones) under one outer shared lock.- Stale-meta torn reads in
read_ids_only(used byturbovec.turbovec_checkand the scan visibility path): it read the ids chain with the caller’s pre-lockmeta. Fixed by re-reading meta under the shared lock, atomically with the ids. - Hardening (last line of defense): the persist path now refuses to write an id-0 or a duplicate id for the bijective flat/single kind, aborting the transaction rather than persisting the corruption signature.
Proven on an i4i.8xlarge (PG18) against a 1.76M-row 768d flat index: a
comprehensive sustained-load run (VACUUM every 3s + 4 writers doing
batch-16/128 ON CONFLICT DO UPDATE with fresh-backend-first-insert
churn + 3 scanners) that corrupts the pre-fix binary in 60 s (id-0
duplicate, 35 SIGABRT/recovery events) ran 70.5 minutes on the fixed
binary with ZERO corruption: 46/46 in-flight turbovec_check reads
clean, 0 dup-id, 0 SIGABRT, 0 recovery, ~19,470 write transactions
concurrent with VACUUM. Insert throughput regresses ~13% (15→13
rows/s on a 1M-row flat index, single writer) — the reconcile adds a
full ids+scales re-read at flush; the dominant per-flush whole-relfile
rewrite is unchanged. Warm kNN and build time unaffected (1M build
72 s, warm kNN ~112 ms/query server-side). Repro harness in
benches/corruption-repro/.
Migration
ALTER EXTENSION pg_turbovec UPDATE + restart the backend to pick up
the new .so. No wire change, no REINDEX for the upgrade. An index
that has already corrupted under the old binary should be
REINDEXed once to clear the bad on-disk state.
[1.29.1] — 2026-08-11
Packaging fix: ship in-place ALTER EXTENSION UPDATE scripts. Patch
bump — no wire change (stays v7), no code change, no REINDEX.
v1.28.4 added turbovec.turbovec_check() to the full-install schema
but not to any upgrade path, so operators who ran ALTER EXTENSION
pg_turbovec UPDATE TO '1.28.4' in place never got the function the
changelog advertised (agora report 2026-08-11). Root cause: the repo
shipped only full-install pg_turbovec--<version>.sql and no
pg_turbovec--<from>--<to>.sql upgrade scripts, so PostgreSQL’s
ALTER EXTENSION UPDATE had no SQL delta to apply — only the .so
changed.
- Adds the runnable upgrade scripts under
sql/: the1.28.3->1.28.4edgeCREATEsturbovec_check; forward edges (1.28.4->1.29.0,1.29.0->1.29.1) are documented no-ops so the update chain resolves. - New
scripts/drift-check.shgate (#12) fails the build if a release lacks itssql/pg_turbovec--<prev>--<this>.sqlupgrade edge, so this can never silently regress again. - Also codifies the 2026-08-11 hard mandate in
AGENTS.md: no corruption ever; non-major upgrades must be zero-format-change or online-upgradable in place (no REINDEX-from-corpus for a minor); even major format breaks must offer an offline in-place converter tool.
[1.29.0] — 2026-08-07
Partitioned-scale support (toward 1T+ vectors) + the 1-bit
quantization foundation. Minor bump — purely ADDITIVE SQL surface;
no wire-format change (MetaPageData::version stays 7), no REINDEX.
ALTER EXTENSION pg_turbovec UPDATE TO '1.29.0'; is sufficient.
Partitioned scale (Phase S-0)
PostgreSQL caps a single heap at 32 TB (~11B rows @768d), so 1T+
vectors must live in a partitioned table. pg_turbovec now documents
this directly, and the surprising-but-verified core finding is that
PostgreSQL’s native Merge Append over a hash-partitioned parent
already does a correct, LAZY global top-k across per-partition
turbovec indexes with zero new code — the scatter→gather→merge is
byte-identical to a single-table exact top-k (PoC: 10/10 overlap), and
the merge pulls ~k + probed rows, not k·N.
docs/PARTITIONED_SCALE.md: a cookbook for scaling to 1–10B vectors today with no AM changes — partition sizing (10M–50M/partition), embarrassingly-parallel per-partition build orchestration (the key build-time lever: ~2.2 days at 800-way vs ~47 days naive for 1T), native insert routing, per-partitionVACUUM/REINDEX CONCURRENTLY, and the query patterns (native parentORDER BY emb <=> q LIMIT k, plus the explicit UNION-ALL fallback). PoC atbenches/poc/scatter_gather_partitioned_topk.sql.- Partition pruning for very large N (Phase S-1, the
partition-level coarse quantizer) is designed (doc §6) but
deferred to a later release — its first cut didn’t pass its own
correctness
#[pg_test]when run, and this project does not ship unproven partition-selection code (a silent recall bug is worse than a deferral). It lands alongside the 1-bit completion, real-PG validated. The free scatter-gather above covers 1–10B today.
1-bit (sign binary quantization) foundation
WITH (bit_width = 1) is now accepted by the reloption (range
1..=4; 1 = sign-BQ, 2/¾ = TurboQuant; the GUC default stays
2..=4, so 1-bit is opt-in and never a default). The turbovec kernel
hard-rejects bit_width < 2, so 1-bit is a distinct sign-BQ scheme
(per-coord sign bit + Hamming coarse + exact heap rerank), matching
pgvector/DiskANN/Qdrant. The exact-heap-rerank floor auto-widens for a
1-bit index (reusing hi_dim_rerank/search_k/oversample — no new
mechanism). The sign-BQ core is implemented + unit-tested:
mean-centering (the mandatory footgun fix — raw sign-at-zero
collapses to R@10≈0 on non-zero-centered data), degenerate
all-same-sign detection, and half-of-2-bit storage (dim/8 codes/vec,
no scale).
Not yet functional end-to-end: a CREATE INDEX ... WITH
(bit_width = 1) currently ERRORs clearly (“not yet implemented;
use bit_width = 2, 3, or 4”) — it never half-builds or ships a silent
all-ones landmine. The end-to-end 1-bit encode/scan path requires a
wire bump to v8 and lands in a later release after full real-PG
validation of the new scan kernel + wire format (deliberately not
rushed into this release). The reloption is accepted now for forward
compatibility.
[1.28.4] — 2026-08-07
Corruption fix: eliminate the dual row-counter drift that could
persist a .tvim id table claiming more rows than it holds, plus a
new turbovec.turbovec_check(regclass) integrity function. Patch
bump — no wire-format change (stays v7), no REINDEX required by the
upgrade itself (a currently-corrupt index still needs REINDEX or
DROP + CREATE; see Migration).
Reported 2026-08-06 (agora / pg.ddx.io) on a 1.74M-row PG18.4 index:
the .tvim id table developed “duplicate ids (id 0 appears in more
than one slot),” write-blocking the whole indexed table, and REINDEX
didn’t durably repair it (only DROP + CREATE held). This release
fixes the root cause on the write path and adds an operator-facing way
to detect the corruption without attempting a write.
B (root cause) — single source of truth for the persisted row count
The deferred aminsert flush (xact::flush_to_relfile) persisted
PersistState.n_vectors — a SEPARATELY-incremented i64 counter — as
the on-disk row count, passed to relfile::write_full_with_prepared
as an INDEPENDENT argument alongside idx.slot_to_id(). Those two
counters can drift (an IdAlreadyPresent remove+re-add, a
mark_dirty closure that didn’t run, a future insert-path refactor).
If n_vectors ever exceeded slot_to_id.len(), the meta page claimed
more rows than the ids chain held, and reload (read_full) over-read
the ids chain into zeroed trailing slots — surfacing as “id 0 appears
in more than one slot” (a real CTID never encodes to 0, so an id-0
slot is always zeroed bytes).
- The flush now DERIVES the persisted count from
idx.slot_to_id().len()(the authoritative array whose bytes actually land in the ids chain) via the new pure, unit-testedxact::reconciled_row_count, making the drift class structurally impossible on the write path. - A hard runtime guard aborts the transaction (the matching
Abortcallback then evicts the dirty cache entry, so the next access reloads clean committed state) if the mirror ever drifts — belt-and-suspenders with the pre-existingassert_eq!inrelfile::write_full_inner(now documented as a hard release-build persist-site guard that must NOT be downgraded todebug_assert!). PersistState.n_vectorsis retained only as a mirror (and the scan-visibility snapshot); the on-disk truth isslot_to_id.len().
A (recovery) — REINDEX now durably repairs
With B fixed, a REINDEX (which allocates a fresh relfilenode; the
per-backend cache drops the stale entry on the relfilenode mismatch
in am_lookup_for_mutation) produces a clean id bijection that the
write path can no longer re-corrupt. The pre-1.28.4 “REINDEX reports
success but the corruption returns minutes later” behavior was the
ongoing insert workload re-corrupting the freshly rebuilt index via
the drift above (or a crash — see C); B removes the write-path source.
C (crash-safety) — KNOWN REMAINING GAP, scoped honestly
This release does NOT make the .tvim id table fully WAL-crash-safe.
An unclean shutdown / pg_resetwal can still discard WAL that
extended the id chain, leaving two slots claiming id 0. Full
WAL-crash-safety of the id table is a larger durability change,
deliberately not attempted inside a corruption patch (shipping a
half-done risky durability change would be worse than the documented
gap). Mitigations in place: the insert AND read paths already ERROR
loudly on such a relfile with HINT: REINDEX INDEX <name>;
(v1.28.2), and turbovec_check() (below) now makes it detectable
without a write. Recovery: REINDEX INDEX <name>; (now durable per
A), or DROP + CREATE for an index corrupted before 1.28.4.
D (integrity check) — turbovec.turbovec_check(regclass)
New read-only, ownership-checked function that reads the meta + ids chain and reports enough for an operator to detect the corruption WITHOUT attempting a write (the only signal pre-1.28.4 detection gave, which blocks the whole table):
SELECT * FROM turbovec.turbovec_check('my_idx'::regclass);
-- wire_version | kind | n_vectors | slot_count | count_matches
-- duplicate_id | is_corrupt | tombstone_density
is_corrupt (true when a flat index has a duplicate id OR the
counts disagree) is the column monitoring should alert on. Reuses
scan::first_duplicate_id. Takes only AccessShareLock, so it never
blocks writers; non-owners get a permission-denied ERROR (portable
across PG13-19 via pg_class_ownercheck / object_ownercheck).
Tests
xact::reconcile_tests— load-independent pure-Rust drift-guard unit tests (run undercargo test --lib, no cluster needed): the regression gate for B.persist_large_batch_no_duplicate_id_after_reload(#[pg_test]) — builds a populated flat index, adds a 128-row batch of new ids via the same path aminsert uses, flushes, and re-reads the persisted relfile asserting no duplicate id andmeta.n_vectors == slot_to_id.len().persist_row_count_drift_aborts(#[pg_test]) — a deliberately driftedPersistStatemust abort the flush, never persist a corrupt relfile.turbovec_check_reports_healthy_flat_index(#[pg_test]) — D.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.28.4'; is sufficient and
cannot fail on existing indexes. No wire change (stays v7), no
REINDEX required by the upgrade. A currently-corrupt index still
needs recovery: REINDEX INDEX <name>; (now durable) or DROP +
CREATE. The fix STOPS new write-path corruption; it does not repair
an index already corrupted by a prior crash.
[1.28.3] — 2026-07-31
Managed-PostgreSQL readiness (audit follow-up): interrupt handling + portability gates + deployment docs. Patch bump — no wire-format change (stays v7), no SQL-surface change, no REINDEX.
Addresses a managed/hosted-PostgreSQL readiness audit of v1.28.2.
- P0 — interrupt handling. The extension had zero
CHECK_FOR_INTERRUPTScall sites, so long scans/builds/flushes were uninterruptible andstatement_timeout/pg_cancel_backend()(and the platform admin signals delivered by the same mechanism) were ignored for seconds. Added polling at every point pg_turbovec controls where no buffer content lock is held:- each
amgettupleentry (makes the iterative-scan refill loop promptly cancellable), - each
CREATE INDEXbuild-stage boundary (k-means / assign-sweep / quantize-encode / prepare-and-persist), - the top of each
GenericXLogbatch in the relfile write path (largeaminsertPreCommit flushes). Known limitation: a single large flat-kindsearch()call is still internally uninterruptible (one call into the Postgres-free kernel crate). Finer-grained polling there needs a kernel block-scoring API and is tracked follow-up; prefer the IVF/graph kinds for large corpora meanwhile.
- each
- P2 — portability drift-check gates.
scripts/drift-check.shnow fails the build if (11a) any raw-WAL / raw-storage primitive appears at a call site (log_newpage,XLogInsert,smgrwrite,smgrextend,PageSetLSN,FlushRelationBuffers, …) — durability must stay 100%GenericXLog; or (11b) any GUC is registered with a non-Usersetcontext. Both invariants hold today; the gate stops a future commit from silently reintroducing a managed-adoption blocker. - P2 — docs. New
docs/DEPLOYING_ON_MANAGED_POSTGRES.mdconsolidating durability/replication, restricted-superuser compatibility, read-replica behavior, cancellation, the memory / O(n)-per-transaction insert model, build cost, and the upgrade / wire-format policy.
Deferred (tracked, larger scope): making the in-place index
rewrite atomic (write the meta page last, as the single commit point;
today it self-heals via the xs_recheckorderby heap backstop); a full
out-of-process TAP suite (crash recovery / replication / failover /
cancellation-latency / codec fuzz); a kernel block-scoring API for
fine-grained flat-scan cancellation. See the audit for detail.
[1.28.2] — 2026-07-31
Bug fix: detect duplicate-id corrupt .tvim relfiles and fail
loudly with a REINDEX hint, instead of silently mis-serving reads
while failing every write. Patch bump — no wire-format change (stays
v7), no SQL-surface change, no REINDEX required by the upgrade
itself (a corrupt index does need one; see below).
Reported 2026-07-30 (agora / pg.ddx.io): a 1.6M-row index on PG18.4
rejected 100% of INSERTs for ~1.5 days (~47k logged failures)
with an opaque aminsert: corrupt relfile pages: duplicate ids in
.tvim file, while pg_index.indisvalid stayed true and scans
kept answering — so the corruption was invisible to health checks and
a backfill retry-loop turned it into an availability incident. The
likely origin was several crash-recovery / pg_resetwal cycles that
left the on-disk id table with duplicate ids.
Root cause was an asymmetry: the write path validated the id table
is a bijection (turbovec’s from_id_map_parts rejects duplicates) and
hard-errored, but the read path (ReadOnlyIndex::from_parts, which
only builds slot_to_id) never checked — so a corrupt index failed
writes yet still served (duplicated / mis-ranked) scan results and
looked valid.
Fix:
- The read/open path (scan::assert_ids_unique_or_reindex, called in
the whole-index, graph, and out-of-core cache-install paths) now
detects duplicate ids once per backend at cache-install and
ERRORs (ERRCODE_DATA_CORRUPTED) with HINT: REINDEX INDEX
<name>; — the same actionable shape as the pre-v7 legacy gate.
A corrupt index now fails loudly on reads too, rather than
silently mis-serving.
- The insert-path error carries the same REINDEX hint, so a
backfill loop gets a clear “rebuild me” signal instead of retrying
an opaque error forever.
- Recovery: REINDEX INDEX <name>; (or ... CONCURRENTLY)
rebuilds a clean id bijection from the heap. There is no lighter
in-place dedup — a rebuild is the supported repair.
Gate: scan::duplicate_id_tests::detects_duplicate_ids (pure-Rust)
covers the bijection predicate incl. the exact reported id.
Still open (tracked, larger scope): the .tvim id table is not
fully crash-safe against pg_resetwal / unclean shutdown — WAL-logging
it well enough to recover cleanly (or refusing to come up valid when
it can’t) is the deeper durability fix and is deferred to a future
release. v1.28.2 makes the corruption detectable and actionable;
it does not yet prevent it from being introduced.
[1.28.1] — 2026-07-28
Nix flake packaging. Patch bump — packaging only; no code change, no wire change (stays v7), no SQL-surface change, no REINDEX.
A customer installing on PostgreSQL 18 via Nix found the v1.28.0 tag
had no flake.nix (“no Nix output”). This release adds one:
nix build github:gburd/pg_turbovec#pg_turbovec_NNfor NN in 13–19 (default= PG18), built with nixpkgs'buildPgrxExtension+ a flake-localcargo-pgrx0.19.1 (nixpkgs' pinned set stops at 0.18.x). The turbovec git dependency’s vendor hash is pinned viacargoLock.outputHashes.- Output layout:
lib/pg_turbovec.so+share/postgresql/extension/pg_turbovec--<ver>.sql+.control. On PG18+ a stock server can use it in place viaextension_control_path/dynamic_library_path(validated: CREATE EXTENSION + distance fns + a turbovec index scan against nixpkgs' postgresql_18, served straight from the store path). nix develop: dev shell with rust, cargo-pgrx 0.19.1, clang/bindgen, openblas, and the PG-from-source deps (bison/flex/readline/zlib/icu).- The
#[pg_test]suite stays CI’s job (needs a live cluster); the flake package build itself runs no tests (doCheck = false).
[1.28.0] — 2026-07-28
PostgreSQL 19 (beta1) support via the pgrx 0.17 → 0.19.1 upgrade.
Minor bump — framework/platform change; no wire-format change (stays
v7), no SQL-surface change, no REINDEX. Existing indexes and
queries are unaffected; ALTER EXTENSION pg_turbovec UPDATE suffices.
- New:
pg19Cargo feature + CI matrix leg. PG19 is upstream beta (19beta1); support is experimental until PG19 RC/GA, at which point it will be re-validated. - pgrx 0.19.1 (from 0.17.0, a two-major jump):
- The
pgrx_embedsecond-compilation-pass model is gone in pgrx 0.18+ — deletedsrc/bin/pgrx_embed.rsand the[[bin]]stanza; SQL entity metadata now lives in the.so’s own linker section. - Rust edition 2024, MSRV 1.96 (was 2021 / 1.85). The
edition-2024
unsafe_op_in_unsafe_fnlint is allowed crate-wide: the index-AM callbacks areunsafe externFFI boundaries whose entire bodies are unsafe by construction; the per-fn# Safetycontracts remain the audit surface. - cargo-pgrx 0.19.1 required (
cargo install --locked cargo-pgrx --version 0.19.1).
- The
- PG19 C-API deltas handled (all version-gated; pg13–18 builds
byte-identically unaffected):
LockBuffer’smodeparameter changed fromint(i32) toBufferLockMode::Type(u32) — newrelfile::lock_buffer_modeshim used at all 5 call sites.relfilenode_from_relationgained the pg19 arm (samerd_locator.relNumberlayout as pg16+; the cfg just didn’t cover pg19, leaving the fn body empty on pg19).
- Verified: all 7 versions (pg13–pg19) compile with 0 errors and 0
warnings; pure-Rust determinism gates (reservoir, IVF k-means,
build-pool invariance) green on pg13 and pg19; the full
#[pg_test]suite gate runs in CI across the 7-leg matrix.
[1.27.3] — 2026-07-12
Phase Q-4c: clear the IVF build cliff — batch the k-means reservoir rotation into one parallel GEMM. Patch bump — build-SPEED change only. The persisted IVF centroids + codes are byte-identical to v1.27.2 for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, no REINDEX.
Profiling a real 1M×1024/lists=4096 build on a 32-vCPU AVX-512 host
showed the v1.27.2 lazy-per-row-rotation was still ~85% of a
~14.5-min build: the scalar O(dim²) rotate_unit ran ~600k times
(reservoir fill + replacements) single-threaded, and at dim=1024 each
rotation is ~1M FLOPs. Reducing the call count (v1.27.2) wasn’t enough
— the per-row scalar rotation was the wrong primitive.
The fix: the reservoir now stores normalised-but-unrotated rows and
rotates the whole training sample once at drain via
ivf::rotate_corpus_into — a parallel BLAS GEMM that already exists
and is already gated byte-identical
(rotate_corpus_bit_identical_across_pool_sizes). The per-row hot path
is now just the O(dim) normalise; the O(dim²) work is one GEMM over the
≤256·lists kept rows across all cores. Byte-identical output: same
rows selected (RNG/seen/cap unchanged, value-independent), same
rotation math (GEMM corpus @ R^T == scalar unit @ R^T).
Measured (real corpus, 32-vCPU AVX-512, PG17.10): - 1M×1024, lists=1024: ~870s → 138s (~6.3×), identical 540 MB index. - 10M×1024, lists=4096: 1562s (~26 min), 5343 MB — previously did NOT complete (DNF, cancelled at 33.7 min / 16%). The build cliff is cleared; the build now scales roughly linearly (per-stage trace shows no single dominant stage). - Query correctness verified at both scales (self-top-1 returns the query id).
Also in this release (verified on the same host): the v1.27.2 G3 concurrency fix holds — the 1M in-RAM kNN throughput sweep rises monotonically and plateaus at ~40 TPS through 64 connections (no collapse), where the pre-fix per-query-pool-churn build collapsed from 46 TPS @16 conns to 1.7 TPS @32 conns. Latency saturates gracefully (797 ms @32 → 1590 ms @64) instead of thrashing.
Gate: reservoir_tests::deferred_batch_rotation_is_byte_identical_to_eager
(pure-Rust) proves the store-unrotated-then-batch-rotate sample matches
the original rotate-every-row path; the IVF determinism + pool-
invariance gates are unchanged and green.
[1.27.2] — 2026-07-11
Phase Q-4b: kill the IVF build cliff’s real bottleneck — the per-row rotation in k-means reservoir sampling. Patch bump — build-SPEED change only. The persisted IVF centroids + codes are byte-identical to v1.27.1 for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, no REINDEX.
Profiling the real 10M×1024/lists=4096 build (v1.27.1 still did not
complete in budget despite Q-4a’s parallel k-means) found the
dominant cost was not k-means (~4.4%, flat in n) but the
single-threaded scalar rotate_unit inside
BuildState::ivf_reservoir_push — ~92% of the projected build. It
ran an O(dim²) rotation on all N accepted rows during the heap
scan, even though the reservoir only ever keeps cap = 256·lists
samples. At 10M/lists=4096 that is ~9M O(dim²) rotations computed and
immediately discarded.
The fix defers the rotation: the reservoir selection uses only the
RNG stream + seen-count + cap (none depend on the rotated values), so
rotate_unit now runs only for the ≤ cap rows that actually land in
the sample. Rotation cost drops from O(N·dim²) to O(cap·dim²),
removing the n-scaling that was the cliff — while producing the
identical k-means training sample (hence identical centroids,
identical codes). The RNG draw stays unconditional in the replacement
branch, exactly as before, so the same source rows land in the same
reservoir slots and receive the same rotation.
Verified: index::build::reservoir_tests::
lazy_rotation_is_byte_identical_to_eager reproduces both the old
(eager-rotate-always) and new (lazy-rotate-on-keep) reservoir
algorithms with a shared ChaCha8 seed across 4 corpus sizes × 3 seeds
(exercising both the fill and the replacement paths) and asserts the
sample buffers are bit-identical (f32::to_bits). Pure-Rust,
load-independent, green. The full cargo pgrx test pg16 gate
(including ivf_streaming_build_determinism_byte_identical) must be
re-run on a quiet host before tagging — the local box was under
load ~73 at authoring time.
[1.27.1] — 2026-07-11
Phase Q-4a: parallelize the IVF k-means build. Patch bump —
build-SPEED change only. The persisted IVF centroids + assignment are
byte-identical to v1.27.0 for a fixed (corpus, seed, lists, dim);
no wire-format change (stays v7), no SQL-surface change, no REINDEX.
This only makes NEW WITH (lists = N) builds faster.
Q-1 exposed the IVF build cliff as the blocker for scale: a real
10M×1024 IVF build did not complete in a 90-min budget. Two remaining serial hot
loops in the k-means build are now parallelized bit-identically
(the rest was already parallel from v1.20.0/v1.22.1):
- gemm_lloyd_assign’s per-row argmin + top-2 tie-break — run
every Lloyd iteration over the whole sample — is now a data-parallel
par_iter_mut().enumerate() writing each assign[i] to a fixed
index, so the result is byte-identical regardless of thread count.
- rotate_corpus_into’s GEMM — Parallelism::None → Rayon(0),
on the same “gemm tiling never reduces across threads” guarantee
v1.22.1 established for the Lloyd cross-term GEMM.
Left serial deliberately: the k-means++ next-seed CDF pick (sequential
by construction) and the empty-cell reseed scan (data-dependent
mutation, not a parallel-safe map).
Measured ~1.91× build speedup (lists=4096, n=131072, dim=256,
8-core, release: 418.2s → 219.4s) with bit-identical output. Sub-linear
because Lloyd iterations are sequentially dependent (only
within-iteration work parallelizes) — the honest ceiling for k-means,
unlike the embarrassingly-parallel graph build (v1.26.0). Verified:
ivf_coarse_model_bit_identical_across_pool_sizes +
rotate_corpus_bit_identical_across_pool_sizes assert byte-identical
centroids + assignment across pool sizes {1, 2, auto};
kmeans_deterministic_across_pool_sizes stays green; full suite
338/338.
This is the first step toward the scale-and-heavy-load build goals; the empty-cell/reseed serial remainder and a larger-scale (10M→100M) build validation are the follow-ups.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.27.1'; — a
no-op upgrade for existing indexes; new IVF builds are just faster.
[1.27.0] — 2026-07-10
Phase Q-0: de-duplicate the on-disk quantized-codes storage, roughly halving the per-vector index footprint. Minor bump (wire-format change — REINDEX required; additive capability, no SQL-surface removal). This clears the storage blocker for large single-node indexes.
The problem. Every prior version persisted each vector’s quantized
codes TWICE: the row-major bit-plane packed_codes chain AND the
SIMD-blocked chain (the output of pack::repack(packed_codes, …)).
That doubled the dominant O(n) storage term. At 100M vectors that is
the difference between ~78 GB and ~39.6 GB for 768d/4-bit (double-
storage vs single).
The fix (Option A — persist only the packed codes). The blocked
layout is a PURE FUNCTION of the packed codes, so v7 drops the blocked
chain entirely from disk and recomputes it once per backend at
index-open via pack::repack. This is the same O(n) one-time compute a
pre-Phase-P index already paid lazily on first scan; it’s paid once per
(backend, am_version) at cache-install and cached in the per-backend
ReadOnlyIndex, so warm per-query latency is unchanged and scan
results (recall, ordering) are bit-identical to before (the
recomputed blocked layout equals the layout that used to be persisted).
Option A was chosen over Option B (persist blocked, recompute packed via
the new pack::unblock) because repack is the forward, already-used
direction and the OOC path never touches the blocked chain — A keeps
every hot path simplest.
Measured storage reduction (phase_q0_storage_is_deduplicated
[pg_test], 512×128d 4-bit): the dropped blocked chain is ≥ the packed
codes chain, i.e. persisting it doubled the code-storage term.
Per-vector on-disk code bytes: dim/8 * bit_width stored ONCE (was
twice). 100M projections: 768d/2-bit 19.8 GB (was 39.6), 768d/4-bit
39.6 GB (was 78), 1536d/2-bit 39.6 GB (was 78).
Wire format (v6 → v7), NOT additive. Unlike the additive v4→v5→v6
per-kind bumps, dropping a persisted chain is a real break for EVERY
kind: single-vector, ColBERT, IVF, and graph indexes all now emit wire
version 7 (the kind byte still discriminates). A pre-v7 index (v1..v6)
is detected by the new MetaPageData::is_legacy_v6() predicate
(version < 7); ambeginscan raises a clear ERROR that NAMES the
index with a HINT: REINDEX INDEX <name>; at the first scan — never
silent corruption. Verified by ambeginscan_errors_on_legacy_v6_meta
(and the retained v1/v2 forgeries, which now also hit the unified v7
gate).
No SQL surface change (no operators/types/functions/GUCs/opclasses added or removed). All index kinds (flat, IVF, graph, ColBERT) continue to work.
Migration:
1. ALTER EXTENSION pg_turbovec UPDATE TO '1.27.0';
2. REINDEX INDEX <name>; — once per turbovec index, ANY kind.
Until an index is REINDEXed under v7, scans against it ERROR with the
REINDEX hint (they do NOT return wrong results). See
docs/UPGRADING.md.
[1.26.0] — 2026-07-10
Phase G-2d(a): a partitioned/merge parallel build for the graph index
kind, so it scales past the single-pass build’s serial ceiling. Minor
bump (new GUC turbovec.graph_build_partitions, new build path,
additive). NO wire-format change (stays v6, byte-identical on-disk
CSR), no new operators/types/functions, no REINDEX.
The single-pass Vamana build is serial by necessity (each insertion navigates the graph every prior insertion left; G-2c showed thread-parallelizing it doesn’t amortize) and did not complete at 5M rows (>2h26m). This adds a structurally-parallel build: partition the corpus into P shards (contiguous ranges of the deterministic shuffled insertion order — each shard a uniform random sample), build each shard’s sub-graph in parallel across the bounded rayon pool, then stitch via a parallel cross-shard refinement pass (greedy-search the merged graph from a global medoid entry + RobustPrune per node) and a deterministic parallel reverse-edge pass. No explicit persisted bridge edges are needed — the refinement pass creates cross-shard navigability, so the frozen v6 CSR is sufficient.
Measured (verified, not just claimed): - Recall parity — partitioned MATCHES or BEATS single-pass. In a findable high-recall regime (queries drawn from the corpus, 20k×64): single-pass R@10=0.958, partitioned (P=8) R@10=0.996 (+0.038). In a cross-cluster regime: 0.630 → 0.754 (+0.124). The refinement pass’s greedy search over the merged graph surfaces better cross-shard neighbours than incremental insertion, so the partitioned graph is HIGHER quality, not a speed-for-recall trade. - ~8× parallel build speedup (relative serial-vs-partitioned wall-clock, 8-core AVX2 box, 200k×64): P=16 ≈ 7.99×, P=8 ≈ 6.45×. Both the per-shard builds and the refinement/reverse passes parallelize — unlike the single-pass build. - Deterministic: bit-identical for a fixed (corpus, seed, P) AND across rayon pool sizes {1, 2, auto} (partition is a pure function of the seed; every parallel stage uses an index-ordered collect / staged fixed-id writes).
turbovec.graph_build_partitions (int, default auto): auto
derives P from corpus size + build-pool budget (single-pass below a
size threshold); 0/1 forces the single-pass reference build; N
forces N shards. The single-pass build_vamana is retained intact
(all its tests pass) and is what P<=1 runs — verified identical.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.26.0';. No
REINDEX — existing v6 graph indexes are unaffected; the parallel build
only changes how NEW WITH (graph = true) indexes are constructed, to
the identical on-disk shape. The 5M/10M a cloud VM gate re-run (now that the
build completes at scale) is the follow-up this unblocks.
[1.25.1] — 2026-07-09
Release-tooling + docs/benchmark patch. No shippable code change —
the compiled binary is byte-identical to v1.25.0 (no src/ runtime
change; wire format stays v6; no SQL-surface change; no REINDEX).
This is the release that first exercises the new automated publish
pipeline.
- Automated release publishing (
.forgejo/workflows/release.yml,scripts/make-dist.sh,META.json.in,ci/announce.sh): avX.Y.Ztag on Codeberg now gates (compile + drift-check), builds a PGXN source distribution (rendersMETA.jsonfrom Cargo.toml’s version, runscargo pgrx schemafor the install SQL, zips a PGXN-layout archive), attaches it to a Codeberg release, uploads to PGXN, and submits a postgresql.org news announcement (feeds pgsql-announce). Runs on a self-hosted Forgejo runner (Codeberg’s hosted 10-min cap can’t fit a pgrx build); each publish step no-ops cleanly if its secrets are unset. The PGXN dist is a source archive built withcargo pgrx install, not apgxn install-able package — published for discoverability + version-pinning. SeeRELEASING.mdfor the one-time runner + secrets setup. - Qdrant + ANN-Benchmarks-protocol competitive benchmark
(
benches/results/qdrant_annbench_20260709/): validated v1.25.0’sturbovec.hi_dim_rerankat real scale — on GIST-960-1M,autolifts R@10 from 0.876 (theoffceiling) to 0.953, crossing the ≥0.90 and ≥0.95 bands the pre-fix engine never reached. vs Qdrant (in-RAM, a benchmark host, 1M): Qdrant wins raw latency 3–18×, pg_turbovec wins storage (SIFT 141 MB = 7–8× smaller; GIST 983 MB = 5.4× smaller than Qdrant’s 5.3 GB). At 10M×960 only Qdrant built in a 90-min budget (turbovec’s single-threaded k-means build cliff). - Docs correction: the 2026-07-08 competitive doc’s pg_turbovec SIFT R@10=0.99 latency (16.84 ms) was the out-of-core path; the in-RAM number is 2.6 ms (verified R@10=0.9915 @ 2.61 ms). Footnoted in place.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.25.1'; — a
no-op upgrade (nothing in the SQL surface or on-disk format changed).
[1.25.0] — 2026-07-09
Gap-B fix: turbovec.hi_dim_rerank, a dimension-aware exact-L2
rerank-window widening that recovers high-dimensional recall. Minor
bump (one new GUC, additive). No wire-format change (stays v6,
byte-identical to v1.24.0), no new operators/types/functions, no
REINDEX.
An offline investigation (using FAISS’s trusted quantizers as
measurement vehicles) established
that the high-dim recall gap — e.g. GIST-1M/960d capping ~0.86 where
pgvector HNSW and VectorChord reach 0.95-0.98 — is NOT retrieval-
bound. The true nearest neighbours DO land in the probed IVF cells
(measured cell recall 0.978 at probes=64, 0.996 at probes=128). The
real bottleneck is in-cell quantized ranking: at high dim the
lossy 4-bit score is noisy enough that a true neighbour often sits at
rank ~200-800 within the probed cells, below a small search_k, so
it never enters the exact-L2 reorder recheck. This corrects the prior
“retrieval-recall ceiling” framing (corrections noted
in-place).
The cure is scan-side only: fetch a wider candidate set so the
always-on exact-L2 reorder queue (xs_recheckorderby) re-ranks enough
survivors to recover the true top-k. Measured: an SQ4 analog of
TurboQuant’s per-coordinate scalar quant lifts R@10 from 0.666 to
0.978 at 960-dim by reranking ~800 candidates instead of ~64.
This must not blanket-widen: at low dim recall already plateaus by
search_k≈25, so widening there is pure latency tax. Hence a
dimension-aware floor. turbovec.hi_dim_rerank:
- auto (default): apply a candidate floor of clamp(dim, 256..=1024)
only for indexes with dim >= 256, and only ever RAISE the
count — an explicit search_k/oversample override past the floor
always wins. SIFT-128 is untouched (zero latency cost); GIST-960,
OpenAI-1536, and the 512-768d embedding families get the wider
window out of the box.
- on: apply the floor regardless of dim.
- off: honour search_k/oversample exactly (pre-1.25.0 behaviour).
The result set is identical to setting search_k/oversample by hand
to the same candidate count — this is a smarter default, not a new
mechanism. Verified end-to-end by the hi_dim_rerank_raises_high_dim_recall
#[pg_test] (384-dim corpus: auto measurably beats off, never
regresses any query) plus the hi_dim_rerank_tests unit tests on the
dim-scaling decision function.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.25.0';. No
REINDEX. The new auto default improves high-dim recall out of the
box at a small high-dim-only latency cost; SET turbovec.hi_dim_rerank
= off restores exact pre-1.25.0 scan-candidate behaviour.
[1.24.0] — 2026-07-08
Phase G-2b: VACUUM + incremental INSERT for the graph index kind.
Both operations previously raised a clear ERROR against a
WITH (graph = true) index (v1.23.0 was build+scan only, correctness-
first); this release turns them into working functionality. No wire-
format change — wire format stays v6, existing v4/v5/v6 indexes
decode byte-identical, no REINDEX. Minor bump because a
previously-ERRORing operation becomes real capability (same reasoning
v1.23.0 used to justify G-2a as a minor).
VACUUM (ambulkdelete) now uses the same per-slot tombstone bitmap
mechanism IVF already uses (the generic relfile::read_tombstones /
write_tombstones_and_meta path, confirmed to have zero IVF-specific
assumptions). Tombstoned nodes never enter scan results and their
out-edges are never followed; a fully-tombstoned corpus returns empty
cleanly. graph_search gained a tombstones parameter threaded
through the scan path.
INSERT (aminsert) does a whole-relfile rewrite per insert
(read every chain back, quantize+append the new vector, run a Vamana
insertion via insert_one_node_via_oracle, persist). This is a
deliberate O(n)-per-insert cost — explicitly NOT the
deferred/batched path other kinds get — appropriate for the
build-then-serve model the graph kind targets; heavy incremental
churn should still REINDEX. The insertion distance oracle uses the
exact quantized-code scan kernel for dist(new, existing) and a
triangle-inequality lower bound for the dist(existing, existing)
RobustPrune diversity check (documented safe-direction approximation:
can only make pruning less aggressive, never drops a genuinely diverse
edge, still hard-capped at degree R).
Two real bugs found and fixed during G-2b’s own test-writing:
Relfile corruption on insert-after-VACUUM.
write_tombstones_and_meta’s block-offset formula for placing a new tombstone chain omitted+ graph_count, so a graph index’s tombstone chain (once VACUUM wrote one) computed an offset that collided with the already-persisted graph adjacency chain, corrupting it on the next incremental insert (corrupt graph adjacency chain: graph offsets[n]=0 != neighbors.len()=...). Root-caused via instrumentation showinggraph_first == tombstone_first.graph_countis0for every non-graph kind, so the fix is a no-op for flat/IVF/ColBERT. A second, related fix re-persists a pre-existing tombstone bitmap after the main graph write (which plans a fresh meta from scratch and would otherwise silently drop it).VACUUM entry-point dead-end. The entry-point fallback fired only when the entry point itself was tombstoned, never when the entry point SURVIVED but every one of its out-neighbors was tombstoned — a low-degree entry point whose only edge points at a now-dead slot is a genuine dead-end (first beam-search hop expands to zero live candidates). Both the VACUUM-side fallback selection and the scan-side entry pick now treat “no live out-neighbor” as equally disqualifying as “dead” and prefer a fallback that itself has a live neighbor.
Also corrected a test-harness bug (not a shipped-code bug): the
graph #[pg_test] corpora were generated with an uncorrelated
random() subquery that PostgreSQL hoisted and evaluated once,
making every test row identical (n_distinct = 1) and producing
spurious “recall collapse” numbers that had nothing to do with the
feature under test. Correlating the inner generate_series to the
outer row restored genuine per-row randomness; the
insert/vacuum/quantization paths were correct all along.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.24.0'; only.
No REINDEX, no wire change, no SQL-surface change. A v1.23.0 graph
index that was built and never mutated is unaffected; graph indexes
can now be VACUUMed and incrementally inserted into in place.
Still deferred (unchanged from v1.23.0): G-2c (SIMD traversal + build parallelism), G-2d (the 5M-scale AVX2 HNSW-latency gate).
[1.23.0] — 2026-07-06
Phase G-2a: WITH (graph = true), a new Vamana-style navigable-
graph index kind — the first step toward matching HNSW’s query
latency while keeping TurboQuant’s storage compression, per
the long-standing roadmap.
Builds a real single-pass Vamana graph (DiskANN’s algorithm: greedy
search + RobustPrune per node, in a deterministic randomized
insertion order) over the full corpus. A graph node’s vector storage
is identical to a flat index’s (same TurboQuant encode path); the
new adjacency chain (CSR: offsets + flat neighbor ids) and an
entry-point slot id are the only new on-disk structures. Scan is a
greedy beam search from the entry point, feeding results through the
existing xs_recheckorderby machinery every other kind already uses
— no new correctness surface there.
Wire format v6, additive (same pattern as v5’s KIND_COLBERT):
existing v4 (single-vector) and v5 (ColBERT) indexes decode
byte-identical under the v6 binary — verified by dedicated tests. No
REINDEX for any existing index. A graph index is a brand-new shape
only a v6 binary produces.
Determinism: relaxed for this kind only, per the plan doc’s
explicit “three fundamental tensions” framing — deterministic for a
fixed seed on one machine/thread-count (required for the test suite
and for REINDEX reproducibility on a given host), but NOT
byte-identical across machines/ISAs the way flat/IVF/ColBERT indexes
are. WAL/streaming replication are unaffected either way (replicas
replay the primary’s actual page bytes, never rebuild independently).
Scope of this release (G-2a, correctness-first — see
for the full sub-phase
breakdown): build + scan work end to end with real, verified
recall against an exact linear scan on test corpora. Explicitly NOT
yet done, each tracked as a numbered follow-up:
- G-2b: VACUUM/tombstone integration. ambulkdelete and aminsert
against a graph index currently raise a clear ERROR rather than
silently corrupting the graph — rebuild the index after bulk
loading/deleting, the same operational model IVF’s original
build-then-query-only phase used before tombstones landed.
- G-2c: SIMD-optimized traversal + build parallelism. This release’s
distance computation is plain scalar Rust; the build is
correctness-first, not speed-optimized.
- G-2d: the real 5M-scale, AVX2-hardware HNSW-latency gate
measurement this whole feature is ultimately judged against
(the gate: p50 ≤ 1.3× HNSW AND storage
≥ 6× smaller AND recall ≥ IVF at matched budget, at 5M rows/
R@0.96 on AVX2). Not run in this release — no latency or
recall-vs-HNSW claim is made here; that measurement is the honest
next step before this feature can be called production-ready, and
the plan doc’s own framing (“the bet failed, keep IVF” is a real
possible outcome) still stands.
New reloption WITH (graph = true), mutually exclusive with
WITH (lists = N) on the same index (enforced in amoptions). 24
new tests (287 total, up from 263), drift-check/compile-matrix/fmt
clean, CI green.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.23.0'; is
sufficient. No REINDEX for existing indexes.
[1.22.2] — 2026-07-06
Raises turbovec.probes’s default from 8 to 16 — the out-of-the-
box recall floor was unreasonably low. Scan-side default change
only, no wire change, no SQL surface change, no REINDEX.
The v1.22.1 a cloud VM competitive re-benchmark measured pg_turbovec’s
shipped defaults (probes=8, search_k=32) capping at R@10=0.796 on
SIFT-1M and R@10=0.407 on GIST-1M — both far below any reasonable
recall SLO, and a real footgun for anyone who runs CREATE INDEX ...
USING turbovec without reading the tuning docs. probes=16 (same
search_k=32) measured:
| Corpus | probes=8 (old default) |
probes=16 (new default) |
|---|---|---|
| SIFT-1M | R@10=0.796, p50=3.0ms | R@10=0.918, p50=4.8ms |
| GIST-1M | R@10=0.407, p50=18.5ms | R@10=0.557, p50=20.4ms |
Roughly 1.5-1.6× the latency for +12-15 recall points on both
corpora — the better point on the curve than probes=32 (which
roughly triples latency for a similar recall gain). Existing
sessions/deployments that explicitly SET turbovec.probes are
unaffected; this only changes the compiled-in default.
New regression test index_am_probes_defaults_to_16 guards the
default against silent drift (matches the index_am_iterative_scan_
defaults_to_off precedent from v1.20.1).
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.22.2'; is
sufficient. No REINDEX.
[1.22.1] — 2026-07-05
Closes a real fraction of the IVF build-cliff gap — scan/build-path only, no wire change, no SQL surface change, no REINDEX.
v1.20.0/v1.21.0’s parallel-build work row-blocked the
normalize/rotate/assign-sweep stages of ambuild, measuring only a
modest ~1.27× speedup on a 64-core box. A FLOPs analysis (triggered
by the v1.22.0 GUC audit) found the real dominant cost was never
row-blocked: gemm_lloyd_assign’s cross-term GEMM runs over the
whole k-means training sample (n_sample = lists × 256, which
equals the full corpus size at high lists) once per Lloyd
iteration, up to 25 times, single-threaded (Parallelism::None, kept
that way for on-disk determinism). At GIST-1M/960d/lists=4096
scale this GEMM is ~26-112× more FLOPs than the already-parallelized
stages — explaining why the earlier fix barely moved the needle.
The fix: gemm 0.18’s own internal Parallelism::Rayon(n)
tiling produces bit-identical output to Parallelism::None for
every shape/seed/thread-count tested — a GEMM’s output tiles are
independent dot-product reductions over the shared contraction
dimension, so (unlike a cross-thread SUM, which does need the
fixed-partition-order bookkeeping k-means' centroid-update step
already has) thread count can never perturb a GEMM’s per-element
result. One line changed: Parallelism::None → Parallelism::
Rayon(0). Via rayon::current_num_threads(), this automatically
and correctly respects turbovec.build_parallelism’s bounded pool
with zero extra plumbing (train_kmeans already runs inside
build_pool::install(pool, ..)).
Measured on real hardware, real scale (16-core AVX-512 a cloud VM
instance, GIST-1M corpus shape: n_sample=1,048,576, dim=960,
lists=4096, full 25-iteration k-means training):
| Variant | Wall clock | vs today |
|---|---|---|
Parallelism::None (v1.22.0, shipped) |
2686.6s (~44.8 min) | baseline |
Parallelism::Rayon(0) (this release) |
768.4s (~12.8 min) | 3.50× |
Centroids confirmed bit-identical between the two runs.
The investigation took two wrong turns before this number, both
worth recording rather than hiding: (1) an early microbenchmark of
gemm at specific shapes segfaulted with a real gdb backtrace into
gemm-common internals, looking exactly like a crate memory-safety
bug — root cause was the test harness’s own transposed-read stride
bug (a genuine ~16M-element out-of-bounds read), not a gemm bug;
replicating the real call site’s exact strides showed no issue. (2)
A later a cloud VM timing harness’s naive “serial baseline” — an explicit
1-thread rayon::ThreadPoolBuilder wrapped around the whole training
call, meant to isolate the GEMM’s own parallelism — also accidentally
forced the unrelated, already-parallel k-means++ seeding phase down
to 1 thread, making the “today” baseline look far slower than
v1.22.0 actually behaves in production (where seeding always runs on
the real build_parallelism pool regardless of the GEMM’s
parallelism setting). Both were caught and retracted before being
reported as findings; the final harness varies only pool size as an
independent axis and compares GEMM modes strictly within each pool
size, matching what the real code actually does.
New regression test kmeans_deterministic_across_pool_sizes (sized
to exceed gemm’s DEFAULT_THREADING_THRESHOLD so it genuinely
exercises multi-threading, not a no-op) asserts byte-identical
CoarseModel.centroids across pool sizes {1,2,3,4,8}. 263/263
tests (1 ignored), drift-check clean, compile-matrix clean all 6 PG
versions.
Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.22.1'; is
sufficient. No REINDEX — this changes build wall clock only, not the
on-disk bytes (centroids/assignment/everything downstream of
train_kmeans is byte-identical to v1.22.0 for the same input).
[1.22.0] — 2026-07-04
Repo cleanup, no functional change. Prompted by an audit for
“silent GUC” traps, unfinished/debugging artifacts, and general
repo hygiene before a release. No wire-format change
(MetaPageData::version stays 5), no REINDEX.
Removed
turbovec.mmap_static_blocked— a deprecated no-op GUC since v1.19.0 (it toggled a relfile-mmap fast path that v1.19.0 deleted). Removed after a three-minor deprecation window (v1.19.0 warn → v1.20.0/v1.21.0 still-warning → v1.22.0 remove), per AGENTS.md’s SQL-surface-removal policy (two-release minimum).SET turbovec.mmap_static_blocked = ...now errors like any other unknown GUC instead of silently no-op'ing..woodpecker/ci.yaml— an orphaned CI config from before the project moved to Forgejo Actions (last touched at v0.3.0, never referenced by any current doc or actually run). It was also the only placecargo fmt --all -- --checkwas ever wired up, which is how 244 formatting violations accumulated across the tree without CI ever catching them (see below).relfile_mmap_static_round_trip_matches_buffer_manager(test) — compared the mmap read path against the buffer-manager fallback; meaningless now that there is only one read path. Also incidentally wrote a stray debug file (/tmp/pg_turbovec_phase_r3_smoke.txt) on every run.
Fixed
cargo fmt’d the whole tree (244 pre-existing violations, purely mechanical/cosmetic — no behavior change).fmt-checkis now wired into.github/workflows/test.ymland.githooks/pre-pushso this can’t silently reaccumulate.- Literal
\uXXXXescape-sequence artifacts (e.g.\u2014instead of an actual em dash) in, ,src/extras.rs,src/index/cost.rs, and the now-removed.woodpecker/ci.yaml— cosmetic (doc comments, not code behavior), but a real artifact of a write-tool double-escaping bug worth stamping out repo-wide rather than file-by-file. - A stale dead-code compiler warning:
highdim_oversample_recovers_ recall’s unusedlists: i64 = 141local (theCREATE INDEXright below it hardcoded the literal141instead of interpolating the variable). src/guc.rs’s own module-doc GUC table was missingturbovec.search_kandturbovec.probes— two of the most-used GUCs in the extension — and had the wrong range forturbovec.cache_size_mb(documented as1..=65536; the actual registered range is0..=65536, and 0 has real, documented meaning: it disables caching). Fixed insrc/guc.rsand propagated todocs/ARCHITECTURE.md§9’s GUC table, which had drifted to only 6 of the 17 real GUCs.README.md’s “Operations note: shared_buffers” section still described the v1.5.0–v1.18.x mmap-era guidance (“1.5× the index size is no longer required”, “shared_buffers size no longer bounds warm-scan latency”) as current fact. It’s the opposite of current reality since v1.19.0 removed mmap:shared_bufferssizing matters again, and pg_turbovec’s 7–15× compression is what makes fitting the hot index inshared_buffersachievable. Rewritten to describe the actual current (buffer-cache-only) read path.docs/BUFFER_CACHE_ONLY_DESIGN.mdwas never actually committed to git despite describing a change that shipped in v1.19.0 (it sat untracked in the working tree for multiple sessions). Committed now with its status header corrected from “DESIGN” (proposal) to “IMPLEMENTED” (it already is, and has been since v1.19.0).
Documentation-only, for context
- Added an explicit warning to
turbovec.bit_width_default’s GUC description: the name isturbovec.bit_width_default, notturbovec.bit_width— PostgreSQL silently acceptsSET turbovec.<anything>as a no-op placeholder custom GUC when the name doesn’t match one this extension actually registered (a generic PostgreSQL behavior, not a pg_turbovec bug), so a typo’dSET turbovec.bit_width = Nneither errors nor does anything. A benchmark driver script hit exactly this during the v1.21.0 Phase G-1 validation (see that release’s CHANGELOG entry) — every “bw=2” row in the original Phase G-0 results was silently built at bw=4. Use thebit_widthindex reloption (WITH (bit_width = N)) to set it atCREATE INDEXtime.
Migration
No REINDEX. Wire stays v5; no SQL surface change besides the
removed deprecated GUC. ALTER EXTENSION pg_turbovec UPDATE TO
'1.22.0'; is sufficient.
[1.21.0] — 2026-07-03
Phase G-1: centroid graph for sublinear IVF coarse-cell
selection. (Gated in by
the finding that the IVF-vs-HNSW latency
gap at SIFT-1M didn’t clear the bar for a full corpus graph, so this
release attacks the coarse-probe cost instead.) In-memory / scan-path
only; no wire change (MetaPageData::version = 5), one new GUC,
no REINDEX.
Added
- Centroid graph coarse-cell selection. For an out-of-core IVF
index (
turbovec.out_of_corecell-scoped path) withlists >= 4096,coarse_probecan now navigate a small fixed-out-degree (16) undirected graph over the coarse centroids — a Vamana/ HNSW-lite greedy beam search — instead of scoring every centroid. The graph is built once per backend, in-memory, from the already-persisted coarse centroids (ivf::build_centroid_graph); nothing new is persisted, so existing IVF indexes get it for free on the next scan, no REINDEX.- Undirected by construction. A pure directed k-NN graph (each
centroid’s own nearest-16 others) can strand a cell that’s
someone else’s close neighbour but has no close neighbours of
its own pointing back — a real navigability gap for greedy
search that a randomized-corpus test caught during development.
build_centroid_graphsymmetrizes every edge (adds the reverse of each directed edge) before search ever runs, which is what makes the recall-preservation guarantee below actually hold. - Byte-deterministic. The directed pass is per-row independent
(parallel-safe); the symmetrization pass is a fixed sort+dedup
over the full edge list. Same centroids ⇒ byte-identical graph.
Verified by
centroid_graph_build_deterministic(unit) andivf_coarse_graph_build_is_deterministic_across_cache_rebuilds(#[pg_test], end-to-end through the relfile + cache). - Recall-preserving.
graph_probe’s beam width (ef = max(nprobe*4, 32), the classic HNSW-style slack budget) is sized so the graph-navigated result SET matches the exact linear scan’snprobe-nearest cells exactly at the tested scales (verified bygraph_probe_matches_linear_scan_exactly, 150 random queries across 5nprobevalues, and the end-to-endivf_coarse_graph_matches_linear_scan#[pg_test]). Existing recall-floor tests (index_am_recall_floor_{2,3,4}bit) still pass unmodified.
- Undirected by construction. A pure directed k-NN graph (each
centroid’s own nearest-16 others) can strand a cell that’s
someone else’s close neighbour but has no close neighbours of
its own pointing back — a real navigability gap for greedy
search that a randomized-corpus test caught during development.
turbovec.coarse_graph(GUC, enum, defaultauto):autobuilds/uses the graph only whenlists >= 4096(below that the plain linear scan is already cheap and a graph’s build + per-query overhead isn’t worth paying — seeivf::GRAPH_MIN_LISTS’s doc);onforces it regardless oflists;offalways uses the exact linear scan.ivf_coarse_graph_auto_falls_back_below_thresholdproves the small-listsfallback is correctness-neutral (matchesoffexactly) and that forcingonbelow the threshold still matches too.
Honest notes
- G-1 is scoped to the out-of-core (
OocIvfIndex) scan path only. The whole-load path (ivf_setup_and_searchinsrc/index/scan.rs) re-reads centroids fresh from the relfile on every scan-open (there is no per-backend cache of that struct today), so building anO(lists²)graph there per scan would be pure overhead, not the “build once per backend” amortised cost the plan requires. That path is also gated to comfortably-RAM-resident indexes, i.e. the small-listsregime where the linear scan is already cheap — the OOC path is also wherelists >= 4096(the scale G-1 targets) actually shows up in practice. - Correcting a docs-drift bug found while implementing this
release: the v1.20.0 CHANGELOG entry below claims a “sublinear
two-level coarse quantizer” (
O(lists)→O(√lists)) shipped in that release. That was never implemented. v1.20.0’s actual diff (verified againstgit show) only parallelized k-means++ seeding and the build-time assign-sweep, and addedturbovec.scan_parallelismfor the fine-scan.coarse_proberemained the plainO(lists·dim)linear scan through v1.20.1. There is noTwoLevelCoarsetype or equivalent anywhere in the v1.20.0–v1.20.1 source tree. v1.21.0 (this release) is the first to actually ship sublinear coarse-cell selection. The v1.20.0 CHANGELOG entry anddocs/UPGRADING.md’s corresponding row are left as historical record (not rewritten) but this release’sdocs/UPGRADING.mdrow calls out the correction explicitly. - Measured effect is a rough local sanity check (small-corpus pgrx
test host, not a benchmark AVX-512/AVX2 latency benchmark): correctness
and recall-preservation are verified by the tests above; a
proper before/after p50 comparison at
listsin the low-to-high thousands (whereautoactually engages) on an AVX2+ host is follow-up bench work, not part of this patch.
Migration
No REINDEX. Wire stays v5; existing v4/v5 IVF indexes benefit
from the centroid graph (when lists >= 4096) with no rebuild.
ALTER EXTENSION pg_turbovec UPDATE TO '1.21.0'; is sufficient.
[1.20.1] — 2026-07-03
CRITICAL PERF FIX: turbovec.iterative_scan default flipped
relaxed_order → off. Wire format unchanged
(MetaPageData::version = 5); no SQL surface change; no REINDEX
needed — ALTER EXTENSION pg_turbovec UPDATE is sufficient and
the new default takes effect on the next backend.
The bug
PostgreSQL’s reorder queue (IndexNextWithReorder in
nodeIndexscan.c) can only return a candidate tuple early when the
index AM’s advertised ORDER BY value for that tuple is exact.
pg_turbovec always advertises f64::NEG_INFINITY for every tuple
(deliberately opclass-agnostic — correct across L2/cosine/inner-
product without per-opclass bounds logic), so that exactness
condition can never be satisfied. Under the old default
(relaxed_order), this forced the executor to drive the AM’s own
iterative-refill schedule (probe-widening, search_k doubling, up
to turbovec.max_scan_tuples = 20,000 and turbovec.max_probes =
64) all the way to completion on every ORDER BY dist LIMIT n
query — no matter how small n was — before the executor’s reorder
queue could signal it was safe to return even the first row.
Measured on an AVX-512 a cloud VM host (SIFT-1M/128d IVF, probes=8,
otherwise-default GUCs): ~2 ms with turbovec.iterative_scan =
off vs ~900 ms with the old default relaxed_order — a 450x
latency tax paid by every default-configuration KNN query since
relaxed_order first shipped as the default in v1.8.0. Every
benchmark and load-test in this repository explicitly set
turbovec.iterative_scan = off, which is why this went undetected
for eleven releases: the bug only manifests when a caller does NOT
override the GUC, and no internal benchmark left it at its default.
It was caught while measuring the Phase G-0 IVF-vs-HNSW frontier,
which (deliberately) exercised the untouched defaults for the first
time.
Changed
turbovec.iterative_scannow defaults tooff(wasrelaxed_order).offmatches pgvector’s ownhnsw.iterative_scandefault and only under-returns on a selectiveWHEREfilter combined withORDER BY ... LIMIT— a much rarer shape than the plain unfiltered KNN query this bug taxed. Opt back intorelaxed_order(SET turbovec.iterative_scan = relaxed_order;) if your workload relies on the under-return-avoidance guarantee for selective filters; seedocs/FILTERING.mdanddocs/PRODUCTION.md.- Added
index_am_iterative_scan_defaults_to_offregression test (src/lib.rs) that asserts the compiled-in default — without an explicitSET— caps atsearch_krather than draining tomax_scan_tuples, so this can’t silently regress back. - Fixed one pre-existing test (
ivf_lists_scan_matches_flat) that was implicitly relying on the oldrelaxed_orderdefault to guarantee finding an exact self-match under quantization noise; it now opts intorelaxed_orderexplicitly, since that’s a correctness anchor for k-widening behaviour, not a test of the compiled-in default. - Corrected the documented default in
src/guc.rs’s module-level GUC table,docs/PRODUCTION.md,docs/FILTERING.md,docs/MIGRATING_FROM_PGVECTOR.md, anddocs/PARITY_GAPS.md.
Migration
ALTER EXTENSION pg_turbovec UPDATE TO '1.20.1'; (empty migration
file, migrations/032_pg_turbovec_v1.20.1.sql). No REINDEX. The new
GUC default applies to new backends/sessions; a long-lived backend
that already read the old compiled-in default at connection start
keeps using it until it reconnects (ordinary PostgreSQL GUC
semantics — not specific to this fix).
[1.20.0] — 2026-07-02
IVF scaling — parallel build + parallel scan + sublinear coarse
quantizer. Surfaced by the benchmark A/B/C benchmark (the benchmark host, AVX-512).
Scan-path / build-path / in-memory only; no wire change
(MetaPageData::version = 5), one new GUC, no REINDEX — existing
indexes benefit with no rebuild.
Added
- Sublinear two-level coarse quantizer (the key scaling enabler).
For an IVF index with
lists > 4096, cell selection is nowO(√lists)instead ofO(lists). The two-level structure is computed in-memory at index-open from the already-persisted coarse centroids (a deterministic function of them) — nothing new is persisted, so existing IVF indexes get it for free on the next scan, no REINDEX. Measured 39–364× fewer centroid distance computations per query at recall 1.0, which breaks the coarse-probe wall:listscan grow large (tiny cells) without the coarse step becoming the bottleneck — the enabler for IVF at 10M+. turbovec.scan_parallelism(GUC, int, default0= auto =min(cores, 4);1= serial). Parallelizes the per-query IVF fine-scan across probed cells (out-of-core path), cutting single-query latency at high dimension. Conservative default to protect aggregate QPS under concurrency. Results identical to the serial scan (same top-k; verified).
Changed
- Parallel IVF build. k-means training + the assign-sweep now use
the bounded build pool (
turbovec.build_parallelism) across cores, memory-bounded and byte-identical (the reduction order is a fixed function of the input, independent of thread count — so the relfile is reproducible across machines with different core counts).
Honest notes
- The parallel build measured only ~1.27× on 64-core
GIST-960d: the single-threaded GEMM (
Parallelism::None, for bit-exactness) remains the dominant term, and row-blocking around it can’t parallelize the GEMM itself. This release makes the parallel build safe and correct (no OOM, deterministic) but it is not the full build-cliff fix — a deterministic parallel GEMM (or an IVF bit-exactness policy change enabling BLAS threads) is scoped follow-up work. - The high-dim recall ceiling was investigated and found retrieval-bound (addressed by the sublinear coarse + more probes), not quantization-bound (widening the reorder-rescore pool recovered ~0 recall).
- Query-latency-at-scale and 10M-build validation on a cloud VM are pending (the 10M run OOM’d before the memory fix in this release; re-run needed to confirm).
Migration
No REINDEX. Wire stays v5; existing v4/v5 indexes benefit from the
sublinear coarse + parallel scan with no rebuild. ALTER EXTENSION
pg_turbovec UPDATE TO '1.20.0'; is sufficient. Tests: 241 → 249.
[1.19.0] — 2026-06-18
All index reads through the PostgreSQL buffer manager. Read-path
architecture change; no wire change (MetaPageData::version
unchanged), no SQL-surface change, no REINDEX. Required for
managed/sandboxed Postgres and any environment that restricts direct
file access.
Changed
- Removed the direct relfile
mmap. Every byte of index data is now read through PostgreSQL’s shared-buffer cache (ReadBufferExtended) — there is nommap/preadof the relfile. The buffer manager is the single source of truth for page access (consistent pinning/locking; clean crash + streaming-replication semantics).src/index/mmap_static.rsis deleted and thememmap2dependency dropped (net −700 lines). The buffer-manager readers this routes through already existed (they were the mmap fallback), so this is mostly a deletion. Seedocs/BUFFER_CACHE_ONLY_DESIGN.md. - Out-of-core (>RAM) IVF serving is preserved without mmap. The
cell-scoped gather (
OocIvfIndex::search_ooc→relfile::gather_codes_ranges) reads only the probed cells' pages through the buffer manager, so the per-backend resident set staysO(probes * cell_size), notO(n). Out-of-core serving needs cell-contiguous layout + range-scoped reads — not mmap.
Deprecated
turbovec.mmap_static_blockedis now a no-op (ignored). It is retained for one minor release so an existingSETdoes not error, and will be removed in a future minor.
Performance notes
- Warm queries: unchanged (the prepared index is cached per-backend; warm scans never touch the buffer manager).
- Cold cache-fill on a >
shared_buffersindex: slower than the old mmap path (per-page lookup/pin/lock + a buffer-manager copy). Mitigation: sizeshared_buffersto hold the hot index — pg_turbovec’s 7–15× compression is what makes “the index fitsshared_buffers” practical where fp32 HNSW could not. For a >RAM index, use IVF +turbovec.out_of_coreso only probed cells' pages are read.
Migration
No REINDEX. Read-path only; wire format unchanged. ALTER
EXTENSION pg_turbovec UPDATE TO '1.19.0'; is sufficient. Tests: 241
(unchanged) on pg16; all 6 PG versions compile; drift-check clean.
[1.18.0] — 2026-06-18
Tier-1 IVF latency optimizations (scan-path). No SQL-surface
change, no wire change (MetaPageData::version = 5; single-vector
stays v4), no REINDEX. Closes the Tier-1 backlog — narrowing the IVF-vs-HNSW latency
gap by attacking the actual per-query floor, with evidence rather than
speculation.
Changed
- Default
turbovec.search_klowered 100 → 32 (#1a). The dominant per-query cost is the executor’s reorder-recheck of every returned candidate (a heap-tuple fetch + an exact full-precision distance recompute each) — not the vector scan. The newsearchk_recall_frontiertest shows recall@10 plateaus bysearch_k≈25 (25/50/100/200 identical), so the old default of 100 over-provisioned the recheck ~3× for zero recall gain. The real-corpusrecall_floor_{2,3,4}bittests pass at the new default, confirming recall safety. Raise it forLIMIT > ~20or a hard corpus; lower it (toward 16) for the lowest latency.
Added
assign_dups_probes_paretotest + guidance (#2). Demonstrates that raisingWITH (assign_dups = M)(soft multi-assignment) lets a query reach a matched recall while probing fewer cells (best recall@10 climbs 0.173 → 0.207 → 0.240 asassign_dups1 → 2 → 4 on the test corpus; min-probes-to-matched-recall is non-increasing inassign_dups). Opt-in (a build-time layout choice; `assign_dups1` needs a REINDEX). The default (1) is unchanged.
- Frontier artifacts:
benches/results/searchk_recall_frontier_2026-06-18.json,benches/results/assign_dups_probes_pareto_2026-06-18.json.
Investigated and rejected / deferred (documented, not built)
- #1b (advertise a tighter ORDER BY distance) — rejected as a
no-op. PostgreSQL’s
IndexNextWithReorderrechecks (heap fetch + exact recompute) every candidate unconditionally underxs_recheckorderby, before reading the advertised value; a tighter bound reduces zero work (identical PG 13–18). Documented insrc/index/scan.rs. - #3 (SIMD
coarse_probe) — assessed, deferred: it is a fixed-floor term, not the dominant cost, and a SIMD horizontal-sum risks cross-ISA reduction-order divergence → recall drift. - #4–6 — not warranted by the data (zero effect on the OOC-gathered benchmark; build-when-profiled).
Honest caveat
- The latency confirmation of #1a/#2 (does p50 actually drop ~3×?)
- is deferred to a quiet AVX2 host — both
flokiandarnoldwere - saturated by unrelated work during this release. The recall safety
- of every change is host-independent and verified here; only the
- latency number awaits a quiet window. The projected effect
- ~10–13 ms at recall@10≈0.96 at 500k–1M, matching HNSW ef40–ef100.
Migration
No REINDEX. Scan-path / default-tuning only; wire stays v5.
ALTER EXTENSION pg_turbovec UPDATE TO '1.18.0'; is sufficient (the
new search_k default applies to new sessions). Tests: 239 → 241.
[1.17.1] — 2026-06-18
ColBERT recall win confirmed cross-domain. Docs + bench-results
release; no source, SQL-surface, or wire change
(MetaPageData::version = 5, single-vector still v4); no REINDEX.
Confirmed
The Phase F-2 index-native ColBERT recall gain (shipped v1.17.0) was
replicated on a second, out-of-domain corpus (BEIR/NFCorpus,
3,633 docs, medical/nutrition, entity-heavier), exercising the
persistent vec_colbert_ops index on floki (AVX2):
- +0.044 nDCG@10 / +0.037 Recall@10 vs the Phase-D pooled+rerank
baseline at the value operating point (
candidate_n=256), rising to +0.065 nDCG at low candidate budget (candidate_n=128, where the pooled baseline collapses to 0.220 whilecolbert_searchholds at 0.285) — same sign, same mechanism, same low-budget shape as the SciFact gate at every config. - Quantization signal intact (2-bit ≈ 4-bit, ≤0.0001 nDCG; 2-bit index = 43 MB).
- The persistent index built cleanly (561k token slots, 42 s / 43 MB at 2-bit, no OOM) and served from disk — the F-1 ~28 MB/call backend-RSS leak is gone (RSS plateaus flat at ~360 MB; ~1.4 KB/call warm).
The qualified GO is upgraded to an established cross-domain recall
win. Data:
benches/results/colbert_f2_confirm_floki_nfcorpus_20260618.json;
harness: benches/scripts/colbert/.
Docs
and updated to record index-native late interaction as DONE (was a future phase): pg_turbovec is one of two PostgreSQL extensions (with VectorChord) with index-native multivector/MaxSim, and the only one also 7–15× smaller than HNSW.
Migration
No REINDEX. Docs + bench only; wire stays v5 (single-vector v4).
ALTER EXTENSION pg_turbovec UPDATE TO '1.17.1'; is sufficient.
Tests unchanged (239).
[1.17.0] — 2026-06-18
Phase F-2 — persistent index-native ColBERT late interaction. New index kind; additive wire bump 4 → 5 (single-vector indexes stay byte-identical to v4); no REINDEX for any existing index. This makes pg_turbovec one of only two PostgreSQL extensions (with VectorChord) to offer index-native multivector/MaxSim — and the only one that is also 7–15× smaller than HNSW.
Added
- Persistent ColBERT token index.
CREATE INDEX ON docs USING turbovec (tokens vec_colbert_ops)over aturbovec.vector[]column (per-doc token arrays) builds a v5 on-disk token index:ambuildunnests each doc’svector[]into per-token slots (the doc’s heap TID repeated per token — the IVF soft-assign synthetic-slot-id machinery), laid out IVF cell-contiguous and spilled via the Phase B-4 BufFile (n_tokens ≫ n_docs, so the spill is load-bearing). Determinism: tokens are unnested in array order. turbovec.colbert_searchnow reads the persistent index. It locates thevec_colbert_opsindex on the token column and runs stage-1 candidate generation against the on-disk relfile (warm cache or cold read) instead of rebuilding a backend cache every call. Stage-2 still exact-MaxSim-reranks heap tokens by ctid. The F-1 ~28 MB/call backend-RSS leak is eliminated on the persistent path (and bounded on the no-index fallback).vec_colbert_opsoperator class overturbovec.vector[]— support functionmax_sim, no order-by operator, so the planner can never select a ColBERT index forORDER BY(the forbiddenamrescanscan-key path is untouched). A ColBERT index ERRORs on anORDER BYscan with a HINT to useturbovec.colbert_search.- VACUUM reuses the IVF tombstone path unchanged: a deleted
doc’s TID marks all its token slots dead (the many-slots-one-TID
shape is identical to IVF soft-assign dups); cells stay contiguous
(tombstone, never swap-remove);
colbert_searchmasks tombstoned slots.
Wire format (additive v5, per index kind)
A new kind byte at page offset 30 (formerly a reserved zero)
discriminates KIND_SINGLE (0, single-vector, wire v4) from
KIND_COLBERT (1, multivector, wire v5). A single-vector build never
sets it, so it emits wire version 4 + kind 0 — byte-identical to
v1.16.0 (guarded by single_vector_still_emits_v4_bytes and
v4_single_vector_index_byte_identical). A v4 meta decodes as
kind = KIND_SINGLE, so is_legacy_v4() never trips.
EXPECTED_WIRE_FORMAT_VERSION is now 5.
Migration
No REINDEX. Existing single-vector indexes are byte-identical and
read unchanged under the v5 binary. A ColBERT index is a brand-new
shape, built fresh (no in-place conversion from a single-vector
index). ALTER EXTENSION pg_turbovec UPDATE TO '1.17.0'; registers
the new opclass and is sufficient. See docs/UPGRADING.md.
Tests
230 → 239 (+colbert_persistent_build_and_search,
colbert_persistent_recovers_single_token_match,
colbert_persistent_survives_vacuum,
colbert_persistent_deterministic,
colbert_index_rejects_orderby_scan,
v4_single_vector_index_byte_identical, + 3 page.rs unit tests).
All six PG versions (13–18) compile; drift-check clean (after the
VERSION-5 / minor-bump pairing this release provides).
[1.16.0] — 2026-06-17
Phase F-1 — index-native late interaction (ColBERT stage-1).
Additive SQL function in the turbovec schema; no wire change
(MetaPageData::version = 4), no index-AM change, no
REINDEX. Closes the last acknowledged feature gap vs
Qdrant/VectorChord at the level the analysis showed actually matters
(stage-1 recall) —.
Added
turbovec.colbert_search(rel, id_col, token_col, query vector[], k, per_token_k = 64, candidate_n = 256, bit_width = 4)(src/colbert.rs) — the index-accelerated stage-1 of ColBERT late interaction. Stage 1 builds a backend-cached flat token index (one slot per token across all docs, doc-id repeated; synthetic unique slot-ids fed toIdMapIndex, real doc-ids kept separately — the IVF soft-assign trick), batch-searches all|Q|query tokens, and unions the hit doc-ids into a candidate set. Stage 2 reads each candidate’s full token array from the heap and scores it with the exactmax_simkernel (Phase D). Returns the top-kdocuments.- The value over the Phase D pooled-vector +
max_simre-rank pattern is stage-1 recall: a document is retrieved by its best single token, not its pooled mean — so a doc whose pooled vector is far but which has one token near a query token (the entity/rare-term/long-doc case ColBERT is built for) is still found. Proven by thecolbert_search_recovers_single_token_matchtest.
What it is / isn’t
The token index lives only in the backend cache (the
turbovec.knn model) — there is no relfile, no CREATE INDEX, and
no wire-format change. It is the index-native stage 1 over
max_sim’s exact stage 2; it is not the full persistent
multivector index AM (per-token relfile + MaxSim-aware scan + PLAID
pruning). That persistent AM (Phase F-2) is gated on a measured
recall/latency win over this F-1 path on a real ColBERT corpus — the
plan explicitly refuses to build a 32–512×-larger persistent index on
faith. Tuning: per_token_k / candidate_n trade recall for work;
raise per_token_k under heavy (2–3 bit) token quantization.
Migration
Additive function; no wire change, no REINDEX. ALTER EXTENSION
pg_turbovec UPDATE TO '1.16.0'; is sufficient. Tests: 224 → 230
(+colbert_search_basic,
colbert_search_recovers_single_token_match,
colbert_search_matches_bruteforce_maxsim,
colbert_search_empty_query, colbert_search_deterministic,
colbert_search_rejects_bad_k).
[1.15.1] — 2026-06-17
Cross-version build fix (pg13 / pg14 / pg15 / pg18). Build-only
patch; no wire change (MetaPageData::version = 4), no
behaviour change on the versions that already compiled (pg16 /
pg17), no REINDEX.
Fixed
- The Phase B-4 out-of-core IVF build (v1.12.0) called
pg_sys::BufFileReadExact(PG16+ only) and passedBufFileWritea*constpointer (PG13–15 declare it*mut). The extension compiled on pg16/pg17 — the local dev target — but failed to compile on pg13, pg14, pg15, and pg18, silently breaking those CI matrix legs from v1.12.0 through v1.15.0. Now both calls go through version-gated shims (buffile_write/buffile_read_exactinsrc/index/build.rs), backing the read with the universally-presentBufFileRead+ an explicit short-read check. All six PG features (13–18) compile again.
Added (CI hardening)
scripts/compile-matrix.sh—cargo checks everypgNNfeature inCargo.toml(compile-only, ~20s each, no test cluster), so version-specific C-API breaks are caught locally before tagging. Wired into.githooks/pre-pushalongsidedrift-check.sh. Skips viaCOMPILE_MATRIX_SKIP=1on hosts without every pgrx toolchain. This is the gate that would have caught the v1.12.0 regression;cargo pgrx test pg16alone never could.
Migration
No REINDEX. Build-only; wire stays v4; pg16/pg17 runtime
unchanged. ALTER EXTENSION pg_turbovec UPDATE TO '1.15.1'; is
sufficient. Tests: 224 (unchanged) on pg16; the fix is verified by
all six PG features compiling.
[1.15.0] — 2026-06-17
Phase C follow-up — operator-path allowlist on flat + IVF.
Additive GUC + function in the turbovec schema; no wire change
(MetaPageData::version = 4), no index-AM scan-key rewrite, no
REINDEX. Brings the in-kernel allowlist pushdown (previously
turbovec.knn()-only, flat-only) to the ORDER BY emb <=> q LIMIT k
operator path.
Added
turbovec.allowlist(session string GUC, default"") — a CSV of heap TIDs (encoded as bigint). When set, the index-AM scan ANDs the allowed slots into the slot mask it hands the SIMD kernel, so the kernel short-circuits 32-vector blocks with no allowed slot before any LUT work — the same in-kernel block-skipknn(..., allowed)gets, now on the operator path. On an IVF index the allowlist is ANDed with the probed-cell mask, scoping the skip to probed cells ∧ allowed slots; the out-of-core cell-scoped path gets it too. Empty/unset = exact prior behaviour with zero added hot-path cost (no slot-bool is ever built). Parsed once per scan (refills reuse it); a non-integer token ERRORs the scan.turbovec.tid_to_bigint(tid) -> bigint— the ergonomic encoder for building the allowlist fromctid(returns the(block << 32) | offsetvalue the AM stores per slot), so users never hand-write the bit-twiddling. Verified bit-identical to the raw encoding (tid_to_bigint_matches_raw_encoding).
Notes / honest limitation
The allowlist is a set of heap TIDs, not an id column — the
index AM keys vectors by heap TID, never a heap id column;
turbovec.knn(..., allowed) remains the id-column path. This is a
pre-materialized id-set channel, not arbitrary-WHERE
pushdown (which would require scan-key reinterpretation — the
forbidden amrescan rewrite — or payload columns in the index).
See docs/FILTERING.md §§ 3.5, 6, 7. Composes with tombstones (a
vacuum-deleted row is excluded even if allowlisted) and with
probes >= lists (exact over the allowed set); returns the same
rows as knn() for the same id-set.
Migration
Additive GUC + function; no wire change, no REINDEX. ALTER EXTENSION
pg_turbovec UPDATE TO '1.15.0'; is sufficient. Tests: 215 → 224
(+allowlist_guc_restricts_ordered_scan_flat/_ivf,
allowlist_guc_matches_knn, allowlist_guc_empty_is_unfiltered,
allowlist_guc_composes_with_tombstones,
allowlist_guc_probes_all_exact, allowlist_guc_rejects_bad_token,
allowlist_guc_out_of_core, tid_to_bigint_matches_raw_encoding).
[1.14.0] — 2026-06-17
Phase D — breadth parity (multivector + hybrid fusion). Additive
SQL surface in the turbovec schema; no wire-format change
(MetaPageData::version = 4), no index-AM change, no REINDEX.
Closes the multivector / hybrid-fusion breadth gap vs VectorChord /
Qdrant at the SQL layer.
Added
turbovec.max_sim(vector[], vector[])/max_sim_cosine(...)(src/hybrid.rs) — ColBERT-style late-interaction MaxSim:sum_{q in Q} max_{d in D} sim(q, d)over per-tokenvector[]arrays.max_simuses dot-product similarity (correct for L2-normalised tokens);max_sim_cosineuses cosine similarity (1 - cosine_distance). All token vectors across both arrays must share one dimension (ERROR on mismatch); an empty query or empty doc scores0.0(ColBERT convention). This is a re-rank primitive (ANN-retrieve candidates on a pooled vector, MaxSim-rerank the top-N) — the token arrays are not indexed, and index-native late interaction remains a documented future phase.turbovec.rrf_score(rank integer, k integer DEFAULT 60)(src/hybrid.rs) — reciprocal rank fusion term1.0 / (k + rank)for fusing a dense ANN ranking with a sparse / keyword ranking. Pairs with the documented CTE recipe; non-positive denominator raises ERROR.docs/HYBRID_SEARCH.md— the canonical breadth guide: multivector MaxSim re-rank (signature, conventions, the two-stage retrieve-then-rerank pattern, the honest index-native limitation), the dense+sparse RRF recipe (fullROW_NUMBER()+rrf_scoreCTE for both full-text andsparsevec), and the named-vector multi-column schema pattern.- Cross-links from
README.md,docs/PRODUCTION.md,docs/PARITY_GAPS.md, anddocs/MIGRATING_FROM_PGVECTOR.md; the multivector / hybrid rows now read “SQL surface SHIPPED; index-native late interaction is a future phase.”
Notes
- Out-of-core BUILD (roadmap Phase D-3) already shipped in v1.12.0 (streaming IVF build); no re-implementation.
- Named vectors (multiple vector columns per row) are a documented schema pattern, not new code.
Migration
Additive SQL functions only; no wire change, no REINDEX. ALTER
EXTENSION pg_turbovec UPDATE TO '1.14.0'; is sufficient (the new
turbovec.max_sim / max_sim_cosine / rrf_score functions are
created by the update script). Tests:
203 → 215 (+max_sim_basic, max_sim_dim_mismatch_errors,
max_sim_empty, max_sim_cosine_normalised, max_sim_rerank,
rrf_score_values, hybrid_rrf_recipe, plus 5 in-module unit tests).
[1.13.1] — 2026-06-17
Phase C — metadata-filtering docs + measured allowlist crossover.
Docs + benchmark release; no source-logic, SQL-surface, or wire
change (MetaPageData::version = 4); no REINDEX. Bench-results
and documentation only.
Added
docs/FILTERING.md— the canonical guide to pg_turbovec’s three working metadata-filter mechanisms, with a cardinality×selectivity×corpus decision matrix:- Partial index (
CREATE INDEX ... WHERE tenant_id = X) — native PG predicate pushdown; the default for known, low-cardinality filters. - In-kernel allowlist
turbovec.knn(rel, id_col, vec_col, query, k, bit_width, allowed bigint[])— true in-kernel pushdown (the SIMD kernel skips 32-vector blocks with no allowed slots before any LUT work), flat-only, for selective per-query id sets. - Iterative scan +
WHERE(v1.8.0) — theORDER BY emb <=> q LIMIT kAM path; the executor rechecks the predicate, the AM widensk/probes(capped bymax_scan_tuples). Includes the honest limitation: no true in-traversal pushdown on theORDER BYAM path (the index stores only vector codes + TID, no payload columns), and a C-4 design sketch for a future phase.
- Partial index (
- Measured allowlist selectivity crossover (floki, AVX2, 300k×
256-d, 4-bit, k=10): allowlist latency decreases monotonically as
the filter tightens (17.9 ms → 0.48 ms, ~37×) while the naive
post-filter is flat (~7 ms); crossover at ~7–10% selectivity, up
to 14.7× faster at 0.1%.
benches/allowlist_crossover.rs+benches/results/allowlist_crossover_floki_v1_13_0_20260617.json.
Fixed (docs drift)
-: refreshed v1.10.1/v1.11.0 →
v1.13.0; the >500k IVF build ceiling and the >RAM gaps are
now marked CLOSED (out-of-core build v1.12.0 + out-of-core
query v1.13.0); the metadata-filtering row reflects the three real
patterns instead of “post-filter only”.
- docs/PARITY_GAPS.md: added the metadata-filtering row; corrected
the stale “Parallel index build | GAP — single-threaded” row
(parallel build shipped v1.8.0, turbovec.build_parallelism).
- docs/MIGRATING_FROM_PGVECTOR.md: filtered-ANN section lists all
three patterns and links FILTERING.md; knn() signature matches
src/knn.rs.
- README.md + docs/PRODUCTION.md: cross-link FILTERING.md.
- Fixed three pre-existing broken benchmark sources
(concurrent_knn, recall_vs_pgvector, recall) that referenced
the pre-d3d468e IdMapIndex::new signature (now returns
Result); cargo check --benches is green again. (cargo pgrx
test never compiled benches, so they did not gate tests.)
Migration
No REINDEX. Docs + bench only; wire stays v4. ALTER EXTENSION
pg_turbovec UPDATE TO '1.13.1'; is sufficient. Tests unchanged (203).
[1.13.0] — 2026-06-17
Out-of-core IVF query (>RAM serving) — an IVF index larger than
RAM can now be queried, not just built (v1.12.0). Wire format
unchanged (MetaPageData::version = 4); no REINDEX. Completes
the out-of-core arc for the >5M production deployment
.
Added — cell-scoped IVF serving (Phase B-1/B-2)
The scan previously loaded the whole index into a per-backend
cache (read_full + a copy of the blocked-codes chain off the
mmap), so the resident set was O(n) and an index that exceeded RAM
could not be served.
- Cell-scoped scan. The backend now caches only bounded metadata
(coarse centroids, cell directory, rotation, codebook, per-slot
scales/ids) plus a
MAP_PRIVATEmmap of the relfile, and per query copies only the probed cells' contiguous code ranges off the mmap into a compact throwaway sub-index (cells are contiguous from the build-time permutation). Resident set drops toO(probes * cell_size + faulted pages); hot cells stay in the OS page cache, cold cells fault from disk on demand. turbovec.out_of_core(enumoff | auto | on, defaultauto).autogoes cell-scoped only when the index codes exceed0.5 * turbovec.cache_size_mb— an in-RAM index loads whole (no per-query gather/reblock cost); only a genuinely large index pays the bounded-memory-for-CPU tradeoff.onforces cell-scoped;offforces the pre-v1.13.0 whole-load.- No wire change, no turbovec fork change (reuses
from_parts_with_prepared_borrowed). Addedmmap_static::gather_slot_ranges+ a buffer-manager twin for the fresh-index fallback.
Measured
200k×256-d×4-bit IVF (52 MB on disk): per-backend VmHWM
whole-load 140.8 MB → cell-scoped 44.1 MB (~3.2× lower). Under a
tight cgroup MemoryMax, the whole-load backend was OOM-killed
(postmaster recovered cleanly, no corruption) where cell-scoped
stayed within bound. Warm p50 82 ms (whole-load) → 199 ms
(cell-scoped) — the expected per-query reblock cost, paid by
auto only when the index is too large to keep whole.
Compatibility
Scan-path only. Results identical to the whole-load path
(probes >= lists still reduces to the exact flat scan; tombstones
masked; soft-assign deduped). MVCC backstops (reorder queue + heap
visibility) preserved. Flat (lists = 0) / vacuum-degraded indexes
keep the whole-index load (no cells to scope; still O(n)-resident
— use IVF for >RAM).
Migration
No REINDEX. Scan-path change; wire stays v4. ALTER EXTENSION
pg_turbovec UPDATE TO '1.13.0'; is sufficient.
Tests
197 → 203 (+ivf_ooc_results_match_whole_load,
ivf_ooc_probes_all_equals_flat, ivf_ooc_tombstones_masked,
ivf_ooc_soft_assign_dedup, ivf_ooc_installs_cell_scoped_handle,
ivf_ooc_auto_is_size_aware). drift-check clean.
[1.12.0] — 2026-06-17
Out-of-core IVF build — IVF indexes can now be built at 1M–5M+
rows on a RAM-constrained host. Wire format unchanged
(MetaPageData::version = 4, byte-identical relfile); no
REINDEX. Driven by the >5M production deployment
.
Fixed — the 1M+ IVF build OOM (Phase B-4)
The WITH (lists = N) build held the full f32 corpus twice in
RAM (ivf_flat ~4 GiB + perm_flat ~4 GiB at 1M×1024-d) plus the
growing index and GEMM scratch — a ~14 GiB peak that
maintenance_work_mem did not bound, OOM-killing 1M+ builds on a
31 GiB host. IVF was effectively unbuildable at the production
scale.
- Disk spill. The corpus now spills to a PostgreSQL
BufFiletemp file (inpgsql_tmp, respectingtemp_tablespaces/temp_file_limit) during the heap scan, wrapped in aCorpusSpillRAII type. Cleanup is double-covered: the resource owner unlinks on (sub)transaction abort (even whenereport(ERROR)longjmps past Rust destructors) andDropunlinks on success. - Three streamed passes, each bounded by
maintenance_work_mem: (1) spill + bounded reservoir sample for k-means; (2) GEMM-assign cells over disk-backed row-blocks, keeping only the per-row cell-id array (not a corpus copy); (3) feed the quantizer in cell order by re-reading the spill at permuted offsets in bounded chunks. The full f32 corpus is never resident; the only RAM term that scales with row count is the quantizedpacked_codes(7–15× smaller than the f32 corpus). - Measured (1M×1024-d, lists=1024, 30 GiB host): peak RSS ~14 GiB (OOM) → ~7.1 GiB (completes); index 1030 MB, spill ~3.9 GB on disk. 5M projected ~8–10 GiB — buildable on the 31 GiB production host.
Determinism / compatibility
Byte-identical relfile to a v1.11.x in-memory build for the same
input, and maintenance_work_mem-invariant (TQ+ calibration is
fit on a fixed cell-ordered prefix, independent of chunk size). The
flat (lists = 0) build path is unchanged (already Phase-W
streamed). No GUC added — maintenance_work_mem is the knob.
Migration
No REINDEX. Build-internal; wire stays v4. ALTER EXTENSION
pg_turbovec UPDATE TO '1.12.0'; is sufficient. Existing indexes are
unaffected; the benefit applies to the next CREATE INDEX /
REINDEX.
Tests
193 → 197 (+ivf_streaming_build_determinism_byte_identical,
ivf_streaming_build_chunk_size_invariant,
ivf_streaming_build_bounded_memory_completes,
ivf_streaming_build_temp_file_cleanup). drift-check clean.
[1.11.1] — 2026-06-16
Bench-results-only release. Wire format unchanged from v1.11.0
(MetaPageData::version = 4); no REINDEX. Zero source change.
Benchmark — IVF latency frontier vs HNSW + ivfflat (Phase A-2)
The honest at-scale measurement (isolated AVX2 on arnold,
taskset-pinned, contention-gated, warm, 300 queries/config,
Cohere-wiki 500k×1024-d) answering “does IVF beat/equal HNSW at
scale.” At recall@10 ≈ 0.96:
| engine | config | recall@10 | warm p50 |
|---|---|---|---|
| pgvector HNSW | ef=200 | 0.966 | 7.9 ms |
| pg_turbovec IVF | lists=707, probes=64 | 0.960 | 18.5 ms |
| pgvector ivfflat | probes=100 | 0.978 | 117.4 ms |
| pg_turbovec flat (exact) | all cells | 1.000 | 41.4 ms |
Honest verdict: HNSW wins latency at 0.96 (7.9 vs 18.5 ms, ~2.3×). But IVF is now in HNSW’s order of magnitude (not the 490× flat-scan gap), beats pgvector’s own ivfflat 3–6× at every matched recall, beats its own exact flat scan, wins the ≥0.99 recall tail (0.99 @ 25 ms via probes=256; this HNSW config never reaches 0.99), and is 7.5× smaller (518 MB vs HNSW 3902 MB). The earlier ~40 ms projection was pessimistic; real p50 at 0.95 is 18.5 ms.
Critical finding — 1M IVF build OOMs (motivates Phase B-4)
The 1M IVF build OOM-killed the postmaster (~14 GiB peak on a
31 GiB host): the lists > 0 build holds the full flat corpus + a
permuted copy + k-means scratch — a structural peak
maintenance_work_mem does not bound. Largest IVF index that built
on arnold: 500k. 1M/5M IVF are blocked on Phase B-4
(streaming / out-of-core build).
The IVF query path is unaffected.
Files: benches/results/ivf_frontier_arnold_cohere-wiki_2026-06-16.json,
docs/BENCHMARKS.md,
.
[1.11.0] — 2026-06-16
Production hardening for IVF: it now survives VACUUM instead of
silently degrading, and builds ~7.8× faster. Wire format stays
MetaPageData::version = 4 (additive); no REINDEX. Driven by
the >5M production deployment +
(Phases A-1, E-2).
Fixed — IVF survives VACUUM (Phase E-2, the production landmine)
An IVF index used to silently degrade to a flat O(n) scan after
VACUUM — swap-remove moved the last vector into the deleted slot,
breaking cell contiguity, so has_ivf() flipped false and queries
fell back to the ~seconds full scan with no operator signal. On a
churning multi-million-row index that’s a latency cliff.
- Tombstones. The IVF
ambulkdeletepath now leaves dead slots in place and ORs them into a persisted per-slot tombstone bitmap (a new v4-additive relfile chain). No rows move,n_vectorsand the cell directory are untouched, cells stay contiguous,has_ivf()stays true, and the scan keeps cell-restricting. The flat (lists = 0) path keeps the unchanged swap-remove. Tombstoned slots are masked out of the initial scan and every probe-widening refill, so deleted rows are never returned. - Observability for any residual fallback: a throttled scan-time
WARNING(once per backend per index, with aHINT: REINDEX) and a new SQL functionturbovec.index_is_degraded(regclass) -> bool.write_meta_shrink_in_placenow preserveslistsand flips anivf_degradedmeta flag rather than blanking the IVF identity, so the cliff is detectable. docs/PRODUCTION.mdgains an IVF + VACUUM operational section.
Performance — 7.8× faster IVF k-means (Phase A-1)
A 200k×256-d / lists=448 build’s k-means training was ~295 s
(scalar Lloyd, fixed 25 iters) — prohibitive at 5M+. Now
GEMM-batched Lloyd assignment (each iteration’s nearest-centroid
step is one V@Cᵀ cross-term GEMM, single-threaded
Parallelism::None + exact top-2 scalar tie-break) + convergence
early-exit (KMEANS_TOL = 1e-6). Training 295 s → 38 s = 7.8×;
per-iteration centroids byte-identical to the scalar path,
determinism preserved. Training cost is bounded by the
256×lists reservoir sample regardless of corpus size, so 5M
builds train in low-minutes. Build-internal; no surface or wire
change.
Migration
No REINDEX. Wire stays v4; the tombstone bitmap and
ivf_degraded flag are additive — pre-1.11.0 v4 indexes read as
not-degraded / no-tombstones. The new index_is_degraded()
function is registered by ALTER EXTENSION pg_turbovec UPDATE TO
'1.11.0';.
Tests
187 → 193 on pg16 (+ivf_survives_vacuum,
ivf_tombstoned_rows_not_returned, ivf_degradation_is_observable,
fast-k-means + page.rs wire-format coverage). drift-check clean.
[1.10.1] — 2026-06-16
Bench-results-only release. Wire format unchanged from v1.10.0
(MetaPageData::version = 4); no REINDEX. Zero source-code change.
Benchmark — IVF warm-p50 on AVX2
Records the AVX2 IVF warm-p50 measurement confirming the IVF
cell-skipping latency win that meh (pre-AVX2 scalar fallback)
could not produce. Host floki (Intel Core Ultra 7 258V, AVX2),
v1.10.0 release build, 200k × 256-d, lists = 448, 4-bit, warm
cache, 50 timed queries per probes:
| probes | warm p50 | vs full scan |
|---|---|---|
| 4 | 0.74 ms | 5.4× faster |
| 16 | 0.78 ms | 5.1× faster |
| 448 (= lists, full exact scan) | 3.97 ms | baseline |
At probes = 16, ~5× faster than the full exact scan, on AVX2.
The IVF latency win is real on AVX2 hardware. Honest caveat:
recall@10 = 1.000 at all probes in this run is an artifact of the
synthetic corpus’s strong cluster structure, not a general
guarantee — the host-independent recall-vs-probes frontier
(v1.10.0) is the honest recall/probes trade-off. A full isolated
1M+ × 1024-d sweep on a quiet AVX2 host remains future work.
Files: benches/results/ivf_warmp50_floki_avx2_2026-06-16.json,
docs/BENCHMARKS.md (“IVF warm-p50 (AVX2)” section).
[1.10.0] — 2026-06-16
Adds the IVF coarse-quantizer layer — a real sublinear ANN
structure over the quantized codes. First wire-format change since
v1.4.0 (MetaPageData::version 3 → 4), but existing v3 indexes
do NOT need a REINDEX: a v1.10.0 binary reads a v3 index as a
flat (lists = 0) index. Only users who opt into IVF rebuild.
.
Why
The v1.9.1 AVX2 benchmark established that pg_turbovec’s flat
O(n·dim) quantized scan is ~490× slower than pgvector HNSW at
1M×1024-d. IVF partitions the corpus into lists Voronoi cells
(coarse k-means centroids) and scans only the probes nearest
cells per query, dropping query work to roughly (probes/lists)
of the corpus — the architectural path to a competitive latency
story while keeping the 10–15× storage win.
Added — IVF (opt-in)
WITH (lists = N)reloption (default 0 = flat / today’s exact scan;N= number of coarse cells, recommended≈ sqrt(n)).WITH (assign_dups = M)reloption (default 1 = single assignment;M > 1= soft assignment: boundary vectors stored in their top-M nearest cells to raise recall@10 at a fixedprobes).turbovec.probesGUC (default 8) — cells scanned per query; the recall/latency dial (theivfflat.probes/hnsw.ef_searchanalogue).probes >= listsreduces exactly to the flat exact scan.turbovec.max_probesGUC (default 64) — underiterative_scan = relaxed_order, a selectiveWHEREfilter that under-returns triggers probe-WIDENING (scan more cells) up to this cap; theivfflat.max_probesanalogue.
How it composes
- Iterative scan (v1.8.0): refill widens
probesfor IVF indexes (vs growingkfor flat). - Oversampling (v1.9.0): widens the candidate set within the probed cells.
- Reorder queue: exact-distance recheck, unchanged.
- turbovec’s SIMD mask SKIPS scan work (block-level early-exit), and cells are stored contiguous, so fewer probed cells = real latency reduction.
Performance / determinism
- The IVF build (k-means + assignment) is GEMM-batched (corpus
rotation as
block @ R^T, cell assignment via aV @ C^Tcross-term GEMM + scalar top-2 tie-break). Without this the per-vector scalar loops were ~1012 FLOPs and a 1M build ran 60+ min. Single-threadedgemmkeeps it bit-deterministic. - Builds are deterministic: same table + same
lists/assign_dups⇒ byte-identical relfile (seeded k-means++,IVF_SEED). - Recall-vs-probes frontier (host-independent; recall is
CPU-independent): on a hard random 16k×64-d corpus,
probes=16→ R@10 0.53 skipping 85% of blocks,probes=lists→ 1.000. Real clustered embeddings reach high recall at far lower probes. Absolute AVX2 warm-p50 latency is deferred to a quietarnoldwindow (mehis pre-AVX2). Seedocs/BENCHMARKS.md.
Migration
No REINDEX for existing (v3, flat) indexes — they read as
lists = 0 under the v1.10.0 binary. Opt into IVF by rebuilding
with WITH (lists = N). ALTER EXTENSION pg_turbovec UPDATE TO
'1.10.0'; registers the new reloptions + GUCs. MetaPageData::version
is 4; EXPECTED_WIRE_FORMAT_VERSION = 4.
Tests
150 → 185 on pg16 (IVF build/scan/soft-assign/determinism/ probes-frontier coverage; distinct-id assertions throughout). drift-check clean.
[1.9.1] — 2026-06-15
Bench-results-only release. Wire format unchanged from v1.9.0
(MetaPageData::version = 3); no REINDEX needed. Zero source-code
changes — this release bundles the AVX2 latency-frontier benchmark
and the honest positioning correction it produced.
Benchmark — AVX2 latency frontier on arnold
The latency numbers meh (a pre-AVX2 Xeon) physically could not
produce. Run on arnold (i9-12900H, AVX2), isolated via
taskset -c 2-5 CPU-pinning to dedicated P-cores with per-batch
contention measurement (observed 1-min load ≤ 1.05 throughout,
zero CPU steal, no contended batches, no re-runs). Cohere wikipedia
1M × 1024-d, 1000 held-out queries, brute-force exact GT.
- Correctness on the AVX2 SIMD path: recall@10 = 1.000 (the AVX2 kernel, not just meh’s scalar fallback, is correct on v0.9.0).
- The hard truth: pg_turbovec loses to HNSW on latency by ~490×
at 1M rows — warm p50 ~2552 ms (flat
O(n·dim)quantized scan, recall 1.000) vs pgvector HNSW ~5 ms (sublinear graph, recall 0.96). AVX2 makes the correct scan ~15–25× faster than meh’s scalar fallback (2.5 s vs 41.6 s), but a 1M-row full scan is seconds, not ms, by design. - Retracted: the earlier “26.8 ms on
meh/ we win 2.3× warm p50” claim. That came from the pre-AVX2 scalar-fallback bug (fast-but-WRONG, fixed in v1.7.3) and never represented correct behaviour.
Positioning correction
docs/PARITY_GAPS.md and updated to
the honest scoreboard. pg_turbovec’s durable wins are storage
(10–15× smaller), exact recall (1.000 vs HNSW’s ~0.96), and
build memory — NOT query latency at scale. Honest positioning:
“best storage efficiency + exact recall for PG vector search where
an O(n) scan fits the latency budget,” NOT “beat HNSW on every
axis.” The architectural path to a latency story at scale is an IVF
/ coarse-quantizer layer (turning the O(n) scan into
O(n/nlist + probes)) — a planned future major arc.
Files
benches/results/latency_frontier_arnold_cohere_1m_v1_9_0_2026_06_15.jsonbenches/scripts/vectordbbench/sweep_latency_isolated.pydocs/BENCHMARKS.md(arnold AVX2 section)docs/PARITY_GAPS.md(corrections)
[1.9.0] — 2026-06-15
Oversampling (tunable recall), test-coverage hardening, and the
first published head-to-head benchmark. Wire format unchanged
(MetaPageData::version = 3); no REINDEX — ALTER EXTENSION
pg_turbovec UPDATE TO '1.9.0'; suffices. The one new GUC defaults
to a no-op.
Added — turbovec.oversample (differentiator #5)
Turns quantization from a fixed accuracy point into a tunable recall lever, matching Qdrant’s oversampling / VectorChord’s rerank.
turbovec.oversample(float, default 1.0, range 1.0..=100.0): the scan fetchesceil(search_k * oversample)quantized candidates and the executor’s reorder queue (xs_recheckorderby = true) trims to the exact top-k. Widening the candidate set recovers true neighbours the lossy quantized ranking placed just outsidesearch_k.- No separate rescore path: oversampling + the always-on reorder queue together ARE the rescore mechanism (the reorder queue already re-ranks by exact full-precision distance). Measured: recall@10 climbs 0.81 (oversample 1.0) → 1.0 (oversample 4.0) on a 4-bit / 3000×64 corpus; latency rises ~linearly.
- Composes with iterative scan: oversample sets the initial
k; refill doubles from there, capped bymax_scan_tuples. - Default 1.0 is identical to v1.8.0 behaviour.
Testing — scale + distinct-id + recall-floor regression guards
The pre-AVX2 wrong-results bug (fixed in v1.7.3) shipped because no test exercised more than ~2000 rows or asserted distinct result ids. Closed those gaps:
- Medium-scale (20k×128-d) recall-floor
#[pg_test]per bit_width {2,3,4}, with a brute-force ground-truth comparison. assert_distinct_idson EVERY ANN-scan test — the cheapest guard against the whole wrong-ranking bug class (a duplicate-id assert would have caught the pre-AVX2 bug instantly).docs/TESTING.mddocumenting coverage + honest gaps: CI is AVX2-only (the scalar fallback runs only in turbovec’s upstream tests + pre-AVX2-host validation on turbovec bumps); unit tests cap at 20k rows (the benchmark is the large-scale evidence); the 15 “ignored” items are benign```ignoredoctests.
Benchmark — first published head-to-head (docs/BENCHMARKS.md)
Cohere wikipedia 1M × 1024-d (real embeddings, 1000 held-out
queries, brute-force GT) vs pgvector HNSW, with a full reproducible
harness under benches/scripts/vectordbbench/.
- recall@10 = 1.000 on the fixed v1.8.0+ build at every config — the same pre-AVX2 host scored 0.0 on the old v1.7.1 build, so this is the definitive confirmation that the pre-AVX2 fix works on real embeddings at scale.
- Storage: pg_turbovec 4-bit 7.6× smaller (1026 MB vs HNSW 7806 MB), 2-bit 15.2× smaller (512 MB). Build 1.9–2.1× faster.
- pgvector HNSW frontier (its own SIMD): R@10 0.849/9.4 ms (ef40) → 0.979/20.1 ms (ef400).
- pg_turbovec latency frontier is DEFERRED to an AVX2 host. The
bench host
mehis a pre-AVX2 Ivy Bridge Xeon; turbovec takes its scalar fallback (~1000× slower than its AVX2/AVX-512 kernels: ~42–69 s/query full-corpus scan). EXPLAIN confirmed Index Scan (not a seq-scan artifact). Latency/QPS benchmarks require an AVX2+ host (arnold); seeBENCHMARKS.mdfor the explicit TODO and full caveats. (UpdatedAGENTS.mdbench-host guidance accordingly — SIMD class matters more than RAM for turbovec latency.)
Migration
No migration; no REINDEX. On-disk format byte-identical to
v1.7.x / v1.8.x. The new turbovec.oversample GUC defaults to 1.0
(no-op). ALTER EXTENSION pg_turbovec UPDATE TO '1.9.0'; resolves
against the empty migrations/014_pg_turbovec_v1.9.0.sql.
Tests
142 → 150 on pg16 (+5 oversampling, +3 recall-floor; distinct-id assertions added to existing tests). drift-check clean.
[1.8.0] — 2026-06-15
Four competitive-parity features in one minor release. Wire
format unchanged (MetaPageData::version = 3); no REINDEX
needed — ALTER EXTENSION pg_turbovec UPDATE TO '1.8.0'; is
sufficient. All four additions are scan-side, build-side, or
additive SQL surface; none touch the on-disk relfile layout.
.
Added — iterative index scan (parity gap #1, the correctness fix)
The one true correctness gap vs pgvector. amgettuple used to run
a single search_k-sized batch and return false when drained, so
a selective WHERE filter ORDER BY emb <=> q LIMIT k silently
under-returned (e.g. 3 rows when 10 were asked for) — exactly what
pgvector shipped hnsw.iterative_scan (0.8.0) to fix.
- When the executor exhausts the candidate batch and the filter
hasn’t been satisfied, the scan re-runs the turbovec search with
a doubled
kand feeds the new candidates, capped by a newturbovec.max_scan_tuplesGUC (default 20000, matching pgvector’shnsw.max_scan_tuples). - Controlled by
turbovec.iterative_scan— an enum GUCoff | relaxed_order(defaultrelaxed_order).strict_orderis deferred; our existing reorder-queue model (xs_recheckorderby = true) already restores exact per-tuple ordering on top ofrelaxed_order. - Dedup across refills via a returned-TID
HashSet(turbovec’ssearchisn’t a stable prefix acrosskdue to an unstable sort on score ties; the set is robust and bounded bymax_scan_tuples). - Regression test demonstrates
offunder-returns andrelaxed_orderreturns the fullLIMIT.
Added — parallel index build (parity gap #2)
pgvector parallelises HNSW/IVFFlat builds across
max_parallel_maintenance_workers; pg_turbovec’s ambuild was
single-threaded.
- Option B (rayon): the CPU-heavy
encode+ SIMD-repackphases (which dominate build CPU, not the heap scan) are parallelised over heap-scan chunks via a rayon pool. Chunks are processed in heap-scan order then concatenated deterministically. - New
turbovec.build_parallelismGUC (default 0 = derive frommax_parallel_maintenance_workers + 1; positive pins the pool). - Byte-for-byte identical relfiles regardless of thread count
— asserted by a unit test — so the wire format and any
reproducibility guarantees hold. Memory stays bounded by the
Phase W
maintenance_work_memcap.
Performance — cold-scan latency (parity gap #3)
Cold-scan p50 was ~1256 ms (1 M × 1536-d) vs HNSW’s ~100 ms.
- Lazy
id_to_sloton the read path. Profiling the per-backend cache-fill showed the dominant residual term — once Phase P pre-baked the SIMD-blocked layout and Phase R-2 persisted the rotation — was theid_to_slot: HashMap<u64, usize>thatIdMapIndex::from_id_map_parts*builds eagerly (~50 ms at 200 k rows, linear inn). The index-AM scan path never readsid_to_slot(searchreturns slots, mapped via theslot_to_idVec). The scan path now installs a lightweightcache::ReadOnlyIndex(no HashMap); the build is deferred to the firstaminsert/remove. A read-only / pooled-connection backend that only scans never pays it. Read-only constructor: ~50 ms → ~0 ms. - Key correctness test
mutation_after_readonly_scan_is_correctverifies the deferred HashMap builds correctly on first insert. - Deferred follow-ups (see
docs/PARITY_GAPS.md§ cold scan): read-path mmap of the codes/scales/ids chains; a header-gap-free on-disk layout for true zero-copy mmap (VERSION 3 → 4, a future minor); a cross-backend DSA/DSM shared cache.
Added — || concat + halfvec arithmetic (parity gap #4)
pgvector has || concat for vector+halfvec and +/-/* for
both; pg_turbovec had +/-/* for vector only.
turbovec.vector || turbovec.vector -> vector(concat)turbovec.halfvec || turbovec.halfvec -> halfvec(concat)turbovec.halfvec+/-/*element-wise (Hadamard for*)- Matches pgvector overflow semantics (error on non-finite result) and dim-mismatch errors.
Migration
No migration needed; no REINDEX. The on-disk relfile format is
byte-identical to v1.7.x. Drop in the new shared library, restart,
scan; existing indexes work unchanged. The new GUCs default to
the pgvector-equivalent behaviour (iterative_scan = relaxed_order).
ALTER EXTENSION pg_turbovec UPDATE TO '1.8.0'; resolves against
the empty migrations/013_pg_turbovec_v1.8.0.sql.
Tests
123 → 142 on pg16 (+19: iterative-scan, parallel-build, cold-scan, and arithmetic-parity coverage). drift-check clean.
[1.7.3] — 2026-06-15
Fixed — pre-AVX2 x86_64 wrong-results bug (turbovec fork → v0.9.0)
Wire format unchanged from v1.6.0 / v1.7.x
(MetaPageData::version = 3); no REINDEX needed to upgrade.
- Root cause. The Phase A1 “regression” (index
ORDER BY emb <=> probe LIMIT Nreturning the sameidN times at 10 M scale on themehbench host) was traced to an upstream turbovec kernel bug, not pg_turbovec. The pinned turbovec v0.7.0-era fork (6e80a59) had a scalar fallback that, on x86_64 CPUs without AVX2, read the perm0-interleaved (FAISS-style) SIMD code layout as if it were sequential — producing silently-wrong / repeated top-k.mehis an Intel Xeon E5-2697 v2 (Ivy Bridge, 2013):avxbut noavx2, so it hit the buggy path. AVX2 (Haswell 2013+), AVX-512, and ARM NEON hosts always took a correct SIMD path — which is why the bug never reproduced on AVX2 dev boxes orarnold, only onmeh. - Fix. Upstream turbovec fixed this in PR #108 (issue #106,
“V5”), released in v0.8.0, adding a correct
score_query_into_heapx86_64 scalar fallback plus aFORCE_SCALAR_FALLBACKregression test. v1.7.3 upgrades thegburd/turbovecfork from the v0.7.0-era6e80a59to a fork rebased onto upstream v0.9.0 (d3d468eon branchpg_turbovec-integration-v0.9.0). - Also brought in, inert here:
- TQ+ per-coordinate calibration fields, constructed as identity (empty) on the relfile path — no recall change, no wire change in v1.7.3. Persisting them for a recall gain is a future minor release (VERSION 3 → 4 + REINDEX).
- Security hardening:
MAX_DIM = 65536, NaN/Inf/huge-magnitude input rejection, checked-mul.tv/.tvimloaders.
- Zero pg_turbovec source churn — the upgrade is a
Cargo.tomlrev bump only; the fork kept theprepare_eageralias and passes TQ+ through internally so everyfrom_id_map_parts*call site is unchanged. - Toolchain note. turbovec v0.9.0 uses
avx512target_features requiring Rust ≥ 1.89. Builds with the defaultstabletoolchain (1.95). SeeAGENTS.mdfor the refreshed openblas store path and the-fuse-ld=bfdlinker note. - Tests: 123/123 on pg16. drift-check clean.
Migration
No migration needed; no REINDEX. The on-disk relfile format is
byte-identical to v1.6.x / v1.7.x. Drop in the new shared library,
restart, scan. Pre-AVX2 x86_64 users specifically should
upgrade to clear the wrong-results bug and can drop any
SET enable_indexscan = off; workaround. ALTER EXTENSION
pg_turbovec UPDATE TO '1.7.3'; resolves against the empty
migrations/012_pg_turbovec_v1.7.3.sql.
[1.7.2] — 2026-05-27
Added — Phase Y: automated upgrade-matrix validation
Wire format unchanged from v1.6.0 / v1.7.0 / v1.7.1
(MetaPageData::version = 3); no REINDEX needed to upgrade.
v1.7.2 is a test-only patch release.
Production-confidence foundation: previously the upgrade matrix
in docs/UPGRADING.md and the is_legacy_v{1,2}() detection
predicates in src/index/page.rs were promises with no
automated end-to-end test. Phase Y closes that gap.
alter_extension_path_140_to_171_runs_clean(new#[pg_test]) replays everymigrations/0NN_pg_turbovec_*.sqlfrom v1.3.0 onward against the live test cluster. Catches a release engineer who lands a syntactically-broken DDL change in one of the post-v1.3 migration files (which are otherwise intentionally empty).ambeginscan_errors_on_legacy_v1_metaandambeginscan_errors_on_legacy_v2_meta(new#[pg_test]s) build a real v1.7.2 index, forge the meta-page version byte to 1 or 2 via the new cfg-gatedrelfile::force_meta_version()helper, and assert thatambeginscanERRORs at first scan with the documented primary message +REINDEX INDEXHINT. Exercises the Phase Q (v1.3.0) + Phase R-2 (v1.4.0) hard migration boundaries without having to keep pre-v1.4 binaries around.alter_extension_update_chain_resolves(new#[pg_test]) asserts the installed extension version matchesCargo.toml, catching version-number drift betweenCargo.toml,pg_turbovec.control, and the migration file naming convention.migration_files_cover_documented_versions(new#[pg_test]) asserts the set ofmigrations/*.sqlsigils matches the documented release history. If you tag a new release without adding the migration file, this test fails before the bad tag escapes CI.scripts/drift-check.sh§ 9 (new check) cross-checksmigrations/*.sqlagainst theFromcolumn of the migration matrix indocs/UPGRADING.md. Catches release-time drift between adding a tag and forgetting to add the migration file.relfile::force_meta_version()(new test-only helper) is gated oncfg(any(test, feature = "pg_test"))and patches the version byte of the meta page in place via aGenericXLogrecord. Only the pgrx test suite (and a future feature-gated build) can reach it; production builds never compile it.
Migration
No migration needed; rebuild not required. The on-disk format is byte-identical across v1.6.0 / v1.7.0 / v1.7.1 / v1.7.2. Drop in the new shared library, restart, scan; existing indexes built under any of these versions continue to work unchanged.
[1.7.1] — 2026-05-27
Reverted — Phase W-2 split-write design (regression)
Wire format unchanged from v1.6.0 / v1.7.0 (MetaPageData::version
= 3); no REINDEX needed to upgrade or downgrade between
any of these. v1.7.1 is a behaviour-only revert.
Phase W-2 (v1.7.0) reverted. Validation on
meh(24-core, 125 GiB RAM NixOS host, head commita289870) at 10 M × 1536-d × 4-bit showed the split-writeambuildpath introduced in v1.7.0 made the build 53% slower (5052 → 7748 s), used 2.7 GiB of swap (vs 0 in v1.6.0), and slightly raised peak RSS (22.5 → 23.04 GiB). The predicted ~15 GiB peak never materialised. Full data:benches/results/phase_w_2_validate_meh_10m_2026_05_27.json.metric v1.6.0 v1.7.0 (W-2) v1.7.1 (revert) Peak RSS (GiB) 22.5 23.04 22.5 (= v1.6.0) Swap used (GiB) 0 2.67 0 (= v1.6.0) Build time (s) 5,052 7,748 5,052 (= v1.6.0) Why Phase W-2 didn’t work. The hypothesis was that dropping the ~7.7 GiB row-major
packed_codesVec mid-finalise (viaIdMapIndex::take_packed_codes()) would shave the peak RSS by ~7.7 GiB. It didn’t, because the interveningwrite_packed_phasepins those bytes inshared_buffersbeforetake_packed_codes()runs, andps -o rsscounts mapped shared memory as part of the backend’s resident set. The 7.7 GiB of “freed” heap simply migrated to pinned shared memory; same RSS budget, plus the cost of an extraGenericXLogflush phase. See .7.1" for the full analysis.What was reverted.
src/index/relfile.rs::write_full_inner— restored to the v1.6.0 single-pass batched-GenericXLogflow: meta page, then codes / scales / ids chains, then blocked / rotation chains, thenRelationTruncatefor shrinking REINDEX.src/index/build.rs::ambuild— restored to the v1.6.0 sequence:prepare_eager()first, then a singlewrite_full_with_preparedcall. Thetake_packed_codes()call is dropped from this code path.src/lib.rs—ambuild_drops_packed_codes_before_blocked_writerenamed toambuild_round_trip_after_phase_w_2_revertand kept as a generic ambuild round-trip smoke (still passes via the v1.6.0 code path).
What was kept.
relfile::write_packed_phase,relfile::write_blocked_phase_and_meta, andrelfile::PackedPhaseLayoutremain in the source as parked dead code, marked#[allow(dead_code)]. They have no callers after the revert but may be useful for a future Phase W-3 attempt that takes a different angle (e.g. streamingpack::repack).- The turbovec fork pin at rev
6e80a59f473292cc9e04d575ba1596f3e23321c5(turbovec 0.7.0) stays.IdMapIndex::take_packed_codes()on the fork is harmless additive API; we just don’t call it.
Migration
No migration needed; rebuild not required. The on-disk format is byte-identical across v1.6.0 / v1.7.0 / v1.7.1. Drop in the new shared library, restart, scan; existing indexes built under any of these versions continue to work unchanged.
[1.7.0] — 2026-05-27
Added — mid-finalise drop of packed_codes in ambuild (Phase W-2)
Wire format unchanged from 1.6.x (MetaPageData::version = 3);
no REINDEX needed to upgrade. v1.7.0 is a build-side change
only: the on-disk index format is byte-identical to v1.6.x.
Reorder finalisation writes so packed_codes and blocked are never co-resident. Phase W (v1.6.0) capped the heap-scan staging buffer, dropping peak
ambuildRSS from 121 GiB to 22.5 GiB at 10 M × 1536-d onmeh. The remaining 22.5 GiB peak wasIdMapIndex’s row-majorpacked_codes(~7.7 GiB) plus the SIMD-blocked derived layout (~7.5 GiB) plus allocator slack + GenericXLog page-assembly buffers, all alive during the single-callrelfile::write_full_with_preparedflush. Phase W-2 splits that call into two phases:relfile::write_packed_phasestreamspacked_codes,scales, andslot_to_idto relfile pages whilepacked_codesis the only large in-memory Vec.IdMapIndex::prepare_eager()materialises the SIMD-blocked layout, codebook, and rotation matrix (transient peak: packed + blocked).IdMapIndex::take_packed_codes()(new turbovec 0.7.0 API) swaps the row-major Vec out andshrink_to_fits it; theOnceLock-backed blocked cache is unaffected.relfile::write_blocked_phase_and_metastreams the blocked- rotation chains and stamps the meta page LAST.
Expected peak at 10 M × 1536-d: ~15 GiB (down from 22.5 GiB). Combined with Phase W’s 121 → 22.5 GiB cut, that’s an 8× total reduction vs pre-Phase-W. Validation on
mehat 10 M scale is a follow-up bench phase; the v1.7.0 code change ships with local unit-test coverage of the split write (ambuild_drops_packed_codes_before_blocked_write).Meta page is now written LAST.
write_full_innerused to write the meta page first and the chains second, which left a crash window where block 0 referenced not-yet-written chain pages. v1.7.0 routes both the legacywrite_full/write_full_with_preparedand the new split path throughwrite_blocked_phase_and_meta, which writes the meta page AFTER all chain pages — matching the standard PG hash/gist AM “meta page is the atomic-complete signal” pattern. A crash before the meta-page WAL record commits leaves block 0 in its previous state (zero-filled for fresh build, previous meta for REINDEX), andambeginscanrejects the index as empty/legacy. No on-disk format change — readers never observed the intermediate state in any released version.Turbovec fork bump 0.6.0 → 0.7.0 (rev
6e80a59f473292cc9e04d575ba1596f3e23321c5, branchpg_turbovec-integrationongburd/turbovec). AddsTurboQuantIndex::take_packed_codes(&mut self) -> Vec<u8>and the matchingIdMapIndex::take_packed_codes. Additive minor; no breaking changes for embedders that don’t call the new API.Phase W-3 deferred. The remaining ~15 GiB peak is dominated by the SIMD-blocked Vec materialised by
prepare_eager()plus GenericXLog page-assembly slack. Dropping the blocked peak further (to ~10 GiB) would require streamingpack::repackso the blocked layout never has to be fully resident; that’s substantial turbovec internals work and is out of scope for v1.7.0.
Migration
No migration needed; rebuild not required. The on-disk format is byte-identical to v1.6.x. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged.
[1.6.1] — 2026-05-27
Bench-results-only release. Wire format unchanged from 1.6.0; no REINDEX needed.
Phase W validation on
meh(commit8efb89c). Re-ran the Phase V 10M × 1536-d build against v1.6.0 to confirm the streamingambuildchange actually drops peak RSS as designed. Result: 121 GiB → 22.5 GiB peak (5.4× reduction), 60 GiB → 0 GiB swap usage. Build time within 0.1 % of Phase V (5048 → 5052 s); index size unchanged (15 GiB); warm-scan p50 identical (21.2 ms vs Phase V’s 21–49 ms band). The remaining 22.5 GiB peak isIdMapIndex’s row-majorpacked_codes(~7.7 GiB) + the SIMD-blocked prepared layout (~7.5 GiB) + allocator slack + PG backend baseline, all held simultaneously during end-of-build finalisation. Tracked as Phase W-2. Files:benches/results/phase_w_validate_meh_10m_2026_05_27.json,benches/results/phase_w_warm_sanity_meh_10m_2026_05_27.json,benches/results/build_tv_meh_10m_v1_6_0_2026_05_27.{log,psql.log,rss.tsv.gz},docs/RECALL.md§ 2.7 follow-up.Phase X: RISC-V architecture comparison (commit
a8fbd87). First non-x86 host bring-up. 100 k × 384-d synthetic onrv(RISC-V 64, 8 cores, 7.7 GiB RAM, Ubuntu 24.04 LTS): index 39 MB (5× compression), build 13.97 s, warm p50 242.64 ms (50-query stdev 0.73 ms — extremely tight). Verdict: arch_works. The latency multiplier vs x86 (~10–25× depending on corpus comparison) reflects turbovec’s AVX2/SSE inner loop falling back to scalar on RISC-V; RVV intrinsics are upstream-future work. Operational note for non-NixOS hosts: the postmaster needsLD_PRELOAD=libopenblas.so.0becausecblas_sgemmis a deferred symbol not in the .so’s NEEDED entries. Files:benches/results/recall_warm_rv_100k_v1_6_0_2026_05_27.json,docs/RECALL.md§ 2.8 (new section).
Migration
No migration needed; rebuild not required. The on-disk format is byte-identical to v1.6.0. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged.
[1.6.0] — 2026-05-26
Added — streaming heap scan in ambuild (Phase W)
Wire format unchanged from 1.5.x (MetaPageData::version = 3);
no REINDEX needed to upgrade. v1.6.0 is a build-side change
only: the on-disk index format is byte-identical to v1.5.x.
- Build-time memory cap. Phase V measured
CREATE INDEXpeak RSS at 121 GiB on a 10 M × 1536-d × 4-bit corpus onmeh(24 cores, 125 GiB RAM), with 60 GiB of swap usage. The dominant offender wasBuildState::flat: Vec<f32>insrc/index/build.rs::ambuildaccumulating the entire heap-scan output before passing it toIdMapIndex::add_with_ids. At 10 M × 1536-d that buffer alone is 61 GiB. - Phase W: stream the heap scan.
BuildStatenow carries two bounded staging buffers (pending_flat,pending_ids) sized offmaintenance_work_mem. Everychunk_rowsrows the callback flushes intoIdMapIndex::add_with_idsandshrink_to_fits the buffers back to zero capacity, returning the bytes to the allocator. A trailing flush after the heap-scan loop drains the partial chunk. - Chunk sizing formula (in
BuildState::compute_chunk_rows):chunk_bytes = min(maintenance_work_mem_kb * 1024 * 3 / 4, 1 GiB);chunk_rows = max(chunk_bytes / (dim * 4), 1). The GUC is read in kilobytes (PG convention; the global ispg_sys::maintenance_work_mem: c_intwhose unit is KB despite the name). 75% allocation leaves headroom for the IdMapIndex’s own growth; the 1 GiB ceiling caps the staging buffer even with aSET maintenance_work_mem = '8GB'. - Expected peak at 10 M × 1536-d: ~16 GiB (down from 121 GiB).
Validation on
mehat 10 M scale is a follow-up phase — the v1.6.0 code change ships with local unit-test coverage of the streaming path; the multi-hour memory-cap validation runs separately. - Phase W-2 deferred. The IdMapIndex still holds
packed_codes(~7.7 GiB at 10 M × 1536-d × 4-bit) in memory alongsideblocked_codesafterprepare_eager(). Dropping it would save another ~7.7 GiB at peak but requires a turbovec fork API change (IdMapIndex::drop_row_major_codes(&mut self)on branchpg_turbovec-integration). Tracked as a follow-up; out of scope for v1.6.0. - One new
#[pg_test]:ambuild_streams_heap_scan_under_maintenance_work_memexercises the streaming path withmaintenance_work_mem = '4MB'and a 1000-row table. Test count 116 → 117. - Docs.
docs/UPGRADING.mdmigration matrix gets a1.5.x → 1.6.0no-op row; records the diagnosis, the formula, and the Phase W-2 follow-up parking lot.
Migration
No migration needed; rebuild not required. The on-disk
format is byte-identical to v1.5.x. Drop in the new shared
library, restart, scan; existing indexes continue to work
unchanged. ALTER EXTENSION pg_turbovec UPDATE TO '1.6.0';
resolves against the empty migrations/007_pg_turbovec_v1.6.0.sql.
This is a minor bump rather than a patch because the build-time memory profile is observably different: a host that used to OOM on 10 M × 1536-d will now succeed. That’s a behaviour change worth a minor even though no on-disk format changed.
[1.5.1] — 2026-05-26
Bench-results-only release. Wire format unchanged from 1.5.0; no REINDEX needed.
- Phase U-1: cache works correctly. A debug-only tracepoint in
cache::lookupconfirmed 50/50 hits across a 50-query warm sweep (zero misses of any class). The Phase S agent’s hypothesis that the per-backend cache misses on every warm scan was wrong; what they saw inperfwas the one-shotfinalise_from_innerbuild during the cold-cache install, amortised over the sampling window. Tracepoint reverted before the build that produced the Phase U-2 measurements. - Phase U-2: Phase S delivers no win on RAM-rich hosts. On
meh(24 cores, 125 GiB RAM), warm p50 is 26.8 ms mmap=on, 26.7 ms mmap=off (delta 0.15 ms = noise) atshared_buffers = 512 MB,search_k = 100. The buffer-manager bottleneck Phase S targets is invisible when free RAM ≫ index size because the OS page cache servespreadreads instantly. Phase S is at-worst-neutral on RAM-rich hosts; it may still help RAM-constrained hosts (the arnold re-bench at the original 31 GiB-RAM constraint remains the definitive Phase S validation). - The headline number that matters: pg_turbovec on a properly- RAMed host beats pgvector HNSW ef=40 on every measurable axis. meh’s 26.8 ms warm p50 is 2.3× faster than HNSW ef=40’s 61 ms, at 5× less storage and R@10 = 1.000 on the dbpedia-1M corpus. The 60–90 ms warm regime that motivated Phase R-2 / Phase S was an arnold-class (limited RAM) phenomenon, not a fundamental kernel ceiling.
Artefacts
benches/results/recall_warm_meh_v1_5_0_2026_05_26.json— full structured run with both configs + verdict.benches/results/u2_meh_tv_4bit_warm_mmap_{on,off}.tsv— raw 50- sample TSVs.- — full method + result of the cache- miss tracepoint experiment.
docs/RECALL.md § 2.6extended with the meh comparison.docs/PARITY_GAPS.mdwarm-scan row updated.
[1.5.0] — unreleased
Added — mmap-based reads of the relfile’s static regions (Phase R-3)
Wire format unchanged from 1.4.x (MetaPageData::version = 3);
no REINDEX needed to upgrade. v1.5.0 is a scan-side change
only.
- New code path:
src/index/mmap_static.rs. Theambeginscancache-fill path nowmmap(MAP_PRIVATE)s the relation’s segment-0 file, walks the deterministic static chains (persisted SIMD-blocked codes, persisted rotation matrix, inline codebook) directly off the mapping, and skips PG’s buffer manager for those bytes. Halves the warm-scan cost when the index doesn’t fit inshared_buffers— the Phase R-3 diagnosis indocs/RECALL.md § 2.5. - New GUC:
turbovec.mmap_static_blocked(defaulton). Setoffper session to revert to the v1.4.x buffer-manager-only read path. Seedocs/ARCHITECTURE.md § 8.1for the isolation contract. - Cache machinery extension:
cache::insert_with_mmap. TheMmaphandle is colocated on theEntrywith theArc<RwLock<IdMapIndex>>and dropped only after the index has been freed (drop order enforced by struct field order). Future zero-copy work (handing turbovec a borrowed slice into the mapping via the newfrom_id_map_parts_with_prepared_borrowedupstream API) relies on this ordering; v1.5.0 holds ownedVecs in the index so the contract is trivially satisfied today. - Upstream turbovec fork bump.
turbovecis pinned togburd/turbovecbranchpg_turbovec-integrationat commitc3c0528, which adds the Cow-based borrowed-cache constructors (from_parts_with_prepared_borrowed,from_id_map_parts_with_prepared_borrowed,PreparedCachesBorrowed). Six new upstream tests cover the borrowed/owned round-trip equivalence and lifetime contract (89 → 95 tests). - Three new
#[pg_test]s:relfile_mmap_static_round_trip_matches_buffer_manager,relfile_mmap_static_concurrent_aminsert_recheck_corrects,relfile_mmap_static_cache_invalidation_drop_order. Test count 113 → 116. - Docs:
docs/RECALL.md § 2.6for the post-fix performance story;docs/ARCHITECTURE.md § 8.1for the isolation contract (heap visibility + recheck-orderby as the MVCC backstops; concurrent aminsert / ambulkdelete / REINDEX worked examples);docs/PARITY_GAPS.mdwarm-scan row updated to reference v1.5.0 with arnold re-bench pending;docs/UPGRADING.mdmigration matrix gets a1.4.x → 1.5.0no-op row;README.md## Performanceoperations note rewritten —shared_buffersno longer needs to be sized against the index size by default.
Dependency added
memmap2 = "0.9"for theMAP_PRIVATERO mapping. No other dependency churn.
Wire format
- No change.
MetaPageData::versionstays at 3,MIN_DECODE_VERSIONstays at 1, and thewire_format_version_is_stabletest continues to assertEXPECTED_WIRE_FORMAT_VERSION = 3.
[1.4.1] — 2026-05-26
Fix — stale rows in the parity scoreboard, plus drift-check tightening
No code changes in this release. Wire format unchanged from
1.4.0; no REINDEX needed.
docs/PARITY_GAPS.mdscoreboard updated with two rows that had drifted three minor versions:- INSERT throughput row was still claiming “~200 ms / row, we lose 400×” — that pre-Phase-K v1.0.x number. Phase K landed in v1.1.0 with the deferred-commit pattern that delivers ~0.13 ms/row (4× faster than HNSW). Row is now accurate.
- Recall on real ada-002 dbpedia-1M row was still “TBD”. Phase J measured R@10 = 1.000 in v1.1.0; the row is now populated with the actual number.
scripts/drift-check.sh§8 now flags scoreboard cells containingTBDor claiming “we lose Nx” without a same-row phase qualifier (e.g. “post-Phase-K”, “shipped in v1.1.0”). Verified by synthesising both failure modes on top of the v1.4.0 scoreboard. The drift-check script also keeps its existing v1.3.0 wire-format check (§7).RELEASING.mdpre-flight checklist grows two items: one forbash scripts/drift-check.shand one for eyeball- reading the PARITY_GAPS scoreboard against the latest benches. drift-check §8 catches structural drift but can’t catch a row whose number is numerically wrong; the eyeball step is the backstop.
All guards aligned: Cargo.toml = 1.4.1, VERSION = 3 (no
change from 1.4.0), EXPECTED_WIRE_FORMAT_VERSION = 3,
drift-check clean.
[1.4.0] — 2026-05-25
Headline (Phase R-2): rotation matrix persisted in the relfile
The random orthogonal rotation matrix used by TurboQuant—a
deterministic function of (dim, ROTATION_SEED) produced by
QR decomposition of a dim x dim Gaussian random matrix—is
now persisted alongside the existing prepared parts (centroids,
boundaries, blocked layout). At dim = 1536 the lazy QR was
the single hottest leaf of the warm-scan profile (~64.8% self
time; see
benches/results/profile_warm_v1_3_0_2026_05_25.json and
), and it ran once per fresh backend
because the per-backend cache OnceLock was driven on first
search instead of read off disk.
ambuild now drives IdMapIndex::rotation() after
prepare_eager() and writes the row-major dim*dim f32
buffer (~9 MiB at 1536-d, negligible vs. the existing
~1.5 GiB index) into a new chain on the relfile. Backends
opening the index pre-fill the rotation OnceLock from those
bytes via the extended
IdMapIndex::from_id_map_parts_with_prepared(…, rotation:
Option<Vec<f32>>) constructor.
Expected impact: warm-scan p50 drops 50–200 ms toward the pgvector HNSW band on dbpedia-1M (1 M × 1536-d). A separate Phase R-3 run on arnold validates the production number; this release is the implementation + wire-format bump.
⚠️ BREAKING: hard migration boundary (v1.3.x indexes)
MetaPageData::version bumps 2 → 3 to add the new
rotation_first / rotation_count / rotation_dim fields and
the rotation chain. v1.4.0 binaries refuse to scan v2 (v1.3.x)
indexes because the rotation chain offsets don’t exist on disk
and the lazy QR was the hotspot we just eliminated. After
upgrading:
ALTER EXTENSION pg_turbovec UPDATE TO '1.4.0';
REINDEX INDEX <every_turbovec_index>;
Without REINDEX, ambeginscan raises
ERROR: turbovec index built under pg_turbovec ≤ 1.3 cannot
be scanned by pg_turbovec 1.4+ with a HINT: Run REINDEX
INDEX <name>;. The detection primitive is
MetaPageData::is_legacy_v2() (mirrors the existing
is_legacy_v1). The matrix in docs/UPGRADING.md documents
the scripted path.
Migration
DO $$
DECLARE
idx record;
BEGIN
FOR idx IN
SELECT n.nspname || '.' || c.relname AS qname
FROM pg_class c
JOIN pg_am a ON a.oid = c.relam
JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE a.amname = 'turbovec'
LOOP
RAISE NOTICE 'reindexing %', idx.qname;
EXECUTE 'REINDEX INDEX CONCURRENTLY ' || idx.qname;
END LOOP;
END $$;
REINDEX INDEX CONCURRENTLY rebuilds without taking an
AccessExclusiveLock so reads keep working during the
migration. The new index is built first; the cutover swap is
atomic.
Vendor turbovec patch
Three additive surfaces on top of the existing Phase P
prepared-cache APIs (see vendor/turbovec/PATCH_NOTES.md for
the full table):
TurboQuantIndex::rotation() -> &[f32]accessor mirroringcentroids/boundaries/blocked_codes. Drives the existingrotationOnceLockand returns the row-majordim*dimmatrix.TurboQuantIndex::rotation_size(dim) -> usizeconst helper (dim * dim) so callers can preallocate the on-disk chain.TurboQuantIndex::from_parts_with_prepared(…, rotation: Option<Vec<f32>>)and the matchingIdMapIndex:: from_id_map_parts_with_preparedoverload —Somepre-fills the rotationOnceLock,Nonefalls back to the lazy QR (used duringambuilditself, when the matrix isn’t yet on disk). Tracked as a follow-up to upstream PR #70 (Codrai turbovec issue #70).
Source
src/index/page.rs:VERSION = 3.MetaPageDatagainsrotation_first/rotation_count/rotation_dim.plan_with_blockedtakes a newrotation_bytesparameter; layout ismeta → codes → scales → ids → blocked → rotation. Decode accepts v1, v2, v3 (older versions leave the new fields zero sois_legacy_v2()flags them).src/index/relfile.rs:PreparedPartsgainsrotation: &'a [f32].write_full_innerwrites the rotation chain after the blocked chain.write_meta_shrink_in_placepreservesrotation_first/count/dimacross vacuum (the matrix is data-independent). Newread_rotation()mirrors the existingread_blocked().src/index/scan.rs:ambeginscangains theis_legacy_v2() && n_vectors > 0ERROR path next to the existing v1 path.amgettuplereads the rotation chain off disk and feeds it toIdMapIndex::from_id_map_parts_with_preparedasSome(rotation).src/index/build.rs:ambuildcallsidx.rotation()afterprepare_eager()and threads it throughPreparedParts.src/xact.rs: same edit on the deferred-commit flush path.src/lib.rs:EXPECTED_WIRE_FORMAT_VERSION = 3. Newrelfile_legacy_v2_detection_primitive(mirrors the v1 test) andrelfile_rotation_persisted(proxy for the warm-scan win: top-1 query through the prepared+rotation index must finish in <100 ms on a 100-row debug build).
Tests
113/113 default. +2 vs. v1.3.0 from the new rotation tests:
relfile_legacy_v2_detection_primitive (mirrors the existing
relfile_legacy_v1_detection_primitive) and
relfile_rotation_persisted (proxy for the warm-scan win:
asserts the rotation chain is on disk, the matrix is
orthogonal to within roundoff, and a top-1 query through the
prepared+rotation index finishes in <100 ms on a 100-row
debug build).
Docs
docs/UPGRADING.md: new migration matrix row for 1.3.x → 1.4.0+, citingis_legacy_v2().vendor/turbovec/PATCH_NOTES.md: “Phase R-2 follow-up: persisted rotation matrix” section documenting the four new surfaces.
[1.3.0] — 2026-05-25
Headline (Phase Q): one storage strategy, no flags
The SPI side-table (turbovec.am_storage) and its accompanying
Cargo feature flags (relfile_storage, experimental_index_am)
are gone. The relfile-resident page format — introduced as a
preview in 1.1.0 (Phase L), proven correct end-to-end in Phase
O-2, and brought up to parity with the side-table on cold-scan
latency by Phase P (1.2.0) — is now the only storage strategy.
The AM matches the conventions of every other PostgreSQL index
AM (btree, gist, gin, hnsw, ivfflat).
Build flags reduce to just pg<N>:
cargo pgrx test pg16 # no --features needed
cargo build --no-default-features --features pg16
⚠️ BREAKING: hard migration boundary
Any existing turbovec index built under v1.0.x..v1.2.0 has either (a) only a side-table row and an empty main fork, or (b) a v1 (Phase L preview) relfile meta layout that lacks the persisted SIMD-blocked layout + Lloyd-Max codebook Phase P relies on. Both states are unrecoverable from the running binary. After upgrading:
ALTER EXTENSION pg_turbovec UPDATE TO '1.3.0';
REINDEX INDEX <every_turbovec_index>;
Without REINDEX, ambeginscan raises an ERROR (no longer a
NOTICE) explaining the situation. This is deliberate — a
half-broken state can’t silently return zero rows.
The extension install / upgrade SQL drops turbovec.am_storage
if it still exists (legacy state from a previous install).
Removed
src/index/persist.rsdeleted (the SPI side-table reader / writer, ~250 lines).aminsert_sidetableandambulkdelete_sidetabledeleted.- The
turbovec.am_storagetable and theextension_sql!block that created it. - The
relfile_storageCargo feature (default-on, no longer togglable). - The
experimental_index_amCargo feature (the AM has been default-on since v0.9; the “experimental” name was stale). - All
#[cfg(feature = "relfile_storage")]and#[cfg(feature = "experimental_index_am")]gates throughoutsrc/. - Migration
NOTICEinambeginscan(replaced by the hardERRORabove). - Stale tests that read
am_storage.payload/am_storage. n_vectorsdirectly. Where the test was exercising generic AM behaviour (“CREATE INDEXsucceeds and the heap is queryable”), it was kept and the assertion was switched tocount(*)on the heap. Where it was strictly side-table- specific (aminsert_deferred_persist_bulk), it was deleted in favour of its relfile twin (relfile_aminsert_deferred_ commit_bulk) which now runs unconditionally.
Updated
src/cache.rsandsrc/xact.rs: the cfg-selected flush sink (sidetablepersist::savevs relfilewrite_full) collapses to relfile only.src/index/cost.rs:amcostestimatereadsn_vectors/dim/bit_widthstraight off the relfile meta page (block 0) instead of via SPI onturbovec.am_storage.- Cargo metadata bumped 1.2.0 → 1.3.0;
pg_turbovec.controlbumped todefault_version = '1.3.0'. migrations/005_pg_turbovec_v1.3.0.sqldocuments the upgrade path and is the new install reference mirror.- Documentation:
docs/PARITY_GAPS.md,docs/ARCHITECTURE.md,docs/PG_VERSION_SUPPORT.md, andREADME.mdupdated to reflect the post-Phase-Q crate layout, retired feature flags, and post-Phase-P cold-scan numbers (1.26 s p50, 21× speedup vs. pre-fix).
Tests
109/109 across pg13, pg16, pg18 (sample of the matrix). Was
94/94 default + 104/104 relfile_storage in 1.2.0; the two
sides converge on 109 now that there are no gates: 94 default
tests + 6 relfile tests (cold-scan, cold-vs-warm, WAL, init
fork, ambulkdelete walk, prepared-layout) + 4 Phase P tests
(prepared layout, cache hits, etc.) + 1 Phase Q test (legacy
v1 detection primitive) + 4 sidetable-specific tests dropped.
Phase O-3 cold-scan re-validation
Phase P’s pre-baked SIMD-blocked layout + Lloyd-Max codebook shipped in 1.2.0 brought cold-scan p50 on dbpedia-1M (1 M vectors x 1536-d, OpenAI embeddings, arnold) from ~26.5 s to 1.26 s p50 — a 21× speedup over the pre-fix v1.0.x side-table path. The full-cluster cold-scan story now matches pgvector HNSW within an order of magnitude, and the relfile-resident architecture wins on every other axis (build time, on-disk size, WAL volume, recall).
[1.2.0] — 2026-05-25
Phase L hardening complete (5 of 6 items)
The relfile-resident page format introduced as a preview in
1.1.0 (--features relfile_storage) is now production-grade
on five of the six hardening items from
:
WAL via
GenericXLog— every relfile page write is now logged viaGenericXLogStart/RegisterBuffer/Finish. A crash before checkpoint correctly replays via standard PG WAL. (Phase N-B, commit9ee405d)ambuildemptyinitialisesINIT_FORKNUMfor unlogged indexes; recovery now produces a queryable empty index without anERROR. (Phase N-B)RelationTruncateis called after a shrinking REINDEX orambulkdeleteconsolidation. (Phase N-B)Phase K’s deferred-commit pattern applied to the relfile path.
aminsert_relfilenow mutates the cachedArc<RwLock<IdMapIndex>>in memory and defers the relfile page write to thePreCommitxact callback. Bulk INSERT of 1 k rows: was minutes (full-rewrite per row) → now < 5 s. (Phase N-C, commitd4a469b)v1.0.x → v1.2 migration HINT in
ambeginscan. When arelfile_storage-built binary opens an index whose main fork is empty but the side-table hasn_vectors > 0, emit aNOTICEwithHINT: Run REINDEX INDEX <name>;. Without this users would silently see zero rows. (Phase N-C)
Phase L hardening remaining (1 of 6)
ambulkdeletewalks pages instead of rebuilding. Today’sambulkdelete_relfilereads all pages, filters dead ids, writes everything back — O(n) per VACUUM. Walk-and-mark would bring this to O(deleted_rows). Tracked for v1.3 in§ 6.
Drift cleanup
docs/ARCHITECTURE.md rewritten to v1.1.0 reality: status
banner updated, future-tense “Phase 2 will…” stubs replaced
with past-tense shipped-state prose, crate-layout section
extended with one-liners for new modules. (Phase N-A, commit
48faeba)
grew a “Shipped in 1.0.x / 1.1.0” section between “Skipped” and “Where future work would pay off”. (Phase N-A)
annotated as superseded by 1.2.0; retained for historical context. (Phase N-A)
Tests
94/94 default + experimental_index_am (unchanged).
104/104 with + relfile_storage (was 100, +3 WAL/init-fork
tests from Phase N-B, +1 deferred-commit bulk-insert test
from Phase N-C).
All six PG versions (pg13.23, pg14.22, pg15.17, pg16.13,
pg17.9, pg18.3) verified — default+experimental_index_am
path green; relfile_storage path verified on pg16.
Status of relfile_storage default
Still gated behind --features relfile_storage, default OFF.
v1.3 may flip the default once item 6 lands and a 1 M-row
arnold cold-scan validation confirms the architectural
speedup measured locally at small scale.
[1.1.0] — 2026-05-24
Phase J — real-embedding head-to-head on dbpedia-1M
The README headline now cites the canonical pgvector benchmark
corpus, dbpedia-entities-openai-1M
(1 M Wikipedia/DBpedia entities × 1536-d OpenAI
text-embedding-ada-002), measured on arnold (Intel i9-12900H,
32 GiB RAM, PG 17.9, pgvector 0.8.0, release build):
| Index / config | Storage | Build | p50 (warm) | R@10 |
|---|---|---|---|---|
| pgvector HNSW (ef=40) | 8 192 MB | 295 s | 61 ms | 0.962 |
| pgvector HNSW (ef=200) | 8 192 MB | 295 s | 115 ms | 0.970 |
| pg_turbovec 4-bit (k=100) | 780 MB | 163 s | 71 ms | 1.000 |
| pg_turbovec 4-bit (k=500) | 780 MB | 163 s | 124 ms | 1.000 |
| pg_turbovec 2-bit (k=100) | 396 MB | 126 s | 48 ms | 1.000 |
| pg_turbovec 2-bit (k=500) | 396 MB | 126 s | 78 ms | 1.000 |
There is no (recall, storage, latency) corner where pgvector
HNSW wins on this corpus. pg_turbovec 2-bit at search_k=100
is Pareto-dominant: 20× less storage, 1.3× faster than HNSW
ef=40, +0.038 higher recall.
Phase L — relfile-resident page format (preview, gated)
New Cargo feature relfile_storage (default OFF) that moves
the serialised index from the SPI side-table to the index
relation’s main fork (relfilenode), accessed via PG’s
standard buffer manager. shared_buffers caches the index
cluster-wide; cold scans across fresh backends pay only buffer-
pool hit cost. All six AM callbacks ported. 100/100 tests pass
with --features "... relfile_storage pg_test". Hardening
before default-on flip in 1.2 tracked in
.
Phase K — deferred-commit aminsert (~3000× bulk-INSERT speedup)
aminsert now mutates the cached IdMapIndex in memory under
a RwLock write guard, marks the cache entry dirty, and
defers the am_storage write to a PreCommit xact callback.
Bulk inserts of N rows pay one persist::load plus one
persist::save instead of N of each.
Wall-clock (release build, 1 M-row index, 1 k-row bulk INSERT): - pre-Phase-K: ~400 s - post-Phase-K: ~136 ms - speedup: ~3000×
Latent bugs fixed during Phase K:
- IdMapIndex::add_with_ids was recomputing the Lloyd-Max
codebook boundaries on every call. Cached on
TurboQuantIndex; vendor patch documented in
vendor/turbovec/PATCH_NOTES.md.
- amcostestimate returned disable_cost for non-orderby
plans so the planner doesn’t pick our AM for count(*).
Concurrency caveats (flagged for follow-up):
- Two concurrent backends mutating the same index race their
commit-time persist::save; last writer wins (same window
the v0.4 path had).
- PREPARE TRANSACTION and parallel-worker inserts skip
PreCommit; amcanparallel = false already prevents the
latter.
Tests
92 → 94 on the default + experimental_index_am path; 100/100 with relfile_storage. All six PG versions (pg13–pg18) green.
Honest scoreboard
docs/PARITY_GAPS.md § "Performance gaps" updated. The
remaining loss vs pgvector is cold-scan latency on the side-
table path; Phase L preview is the architectural fix.
[1.0.1] — 2026-05-24
Fix — build on PostgreSQL 13, 14, 15, 18
v1.0.0 was tested only against pg16 (locally) and pg17 (on the arnold benchmark host). Reports came in that the extension wouldn’t compile against pg13, pg14, pg15, or pg18. Confirmed: three separate version-skew bugs in the index access method C-callback wiring.
All fixes are additive #[cfg(...)] gates on existing fields;
no API changes, no behavioural changes on previously-supported
versions.
src/index/mod.rs::register_am:(*routine).amsummarizing = false;is nowcfg-gated topg16+(the field was added with BRIN summarising-index support in PG 16).(*routine).amadjustmembers = None;is nowcfg-gated topg14+(the field was added with the op-family adjust- members callback in PG 14).
src/index/insert.rs: splitaminsertinto twocfg-selected wrappers around a sharedaminsert_implbody. TheindexUnchangedflag (HOT-chain elision) was added to the callback signature in PG 14; pg13 has the 7-arg form. Both wrappers delegate to the same Rust implementation.src/index/options.rs:pg_sys::relopt_parse_eltgained anisset_offset: i32field in PG 18. Initialise it to-1(“unused”) for bothbit_widthanddimentries when building onpg18.
Tests
cargo pgrx test pg<N> --no-default-features --features
"pg<N> experimental_index_am pg_test" for N in 13..=18:
| Version | Result |
|---|---|
| 13.23 | 92/92 passing |
| 14.22 | 92/92 passing |
| 15.17 | 92/92 passing |
| 16.13 | 92/92 passing |
| 17.9 | 92/92 passing |
| 18.3 | 92/92 passing |
A docs/PG_VERSION_SUPPORT.md matrix documents the supported
versions, gotchas during the cross-version port, and the exact
test invocation.
Known follow-ups
The sub-agent helping verify on arnold caught a fourth issue
that is not a bug but worth recording: when refactoring
aminsert into a thin C-ABI wrapper plus an inner Rust
implementation, the inner helper cannot be called
aminsert_inner because #[pgrx::pg_guard] already generates
a private <fn_name>_inner. We renamed the helper to
aminsert_impl. Documented at the call site.
[1.0.0] — 2026-05-24
A real-hardware million-row run on arnold (Intel i9-12900H, PG
17, pgvector 0.8.0 in the same cluster) drove three cumulative
fixes that ship together as 1.0.0 proper:
turbovec.search_kGUC (default 100). The 0.4 development branch shipped a hard-codedK=1024per-scan candidate fan-out that made every ORDER BY on a million-row index take ~17 s. Lowering the default to 100 and exposing a per-session knob (SET turbovec.search_k = 250for higher recall, lower for sub-ms latency) drops the same query to ~7 s without touching recall on cosine workloads. (#63879a8)amrescantolerates non-orderby plans. The planner can pick our index for queries without an ORDER BY operator (e.g.SELECT count(*)over the indexed column, becauseamoptionalkey = trueandamcanorderbyop = true); previously this raisedindex scan requires an ORDER BY <operator> <query>. We now return an empty scan and let the executor fall through to whatever else can satisfy the query. (#63879a8)- Backend-local cache wired into the AM scan path. The
cache (
src/cache.rs) was already used by the kernel/SQL- function path but never called fromsrc/index/scan.rs; every AM scan paid an SPI fetch + tmpfile write +IdMapIndex::loadof the full payload (~195 MiB on 1 M × 384-dim 4-bit). Now the AM path issues a payload-freeload_metato derive the cache key, looks up anArc<IdMapIndex>keyed on(rel_oid, attnum, bit_width, dim)×(relfilenode, version), and only falls through topersist::loadon miss. Intra-backend warm-cache speedup observed in the field is ~9.7× (35.7 s → 3.7 s on the arnold corpus, debug build). (#1293e7b)
Phase 21 — million-row recall + latency vs pgvector HNSW
docs/RECALL.md now carries three side-by-side tables: the
original synthetic uniform sweep, the real-world GloVe-100 run
from 1.0.0-rc.2, and a fresh million-row arnold sweep at 384
dimensions. Headline (warm cache, debug build):
| Index | Storage | p50 | R@10 (synth) |
|---|---|---|---|
| pgvector HNSW ef=40 | 1953 MiB | 104 ms | 0.032 |
| pgvector HNSW ef=200 | 1953 MiB | 130 ms | 0.116 |
| pg_turbovec 4-bit | 195 MiB | 3 364 ms | 1.000 |
| pg_turbovec 2-bit | 103 MiB | 1 757 ms | 0.922 |
Uniform-random vectors in 384 dimensions are a documented pessimistic case for graph indexes — see § 2.1 for the GloVe-100 numbers where HNSW recovers to 0.80–0.93. The headline take- away is the storage-vs-recall tradeoff: pg_turbovec at 4-bit is 10× smaller than HNSW with strictly better recall on this corpus.
New artefacts:
benches/results/recall_lat_million_2026_05_24.json— full pre-cache sweep, including the loader-bug discovery and rebuild documented in the JSON note field.benches/results/recall_lat_million_post_cache_2026_05_24.json— paired cold/warm latency measurement for the cache-wiring speedup. Use these to reproduce the 9.7× intra-backend ratio.benches/scripts/{rebuild_corpus_million.sh, bench_million_setup.sql, run_bench_sweep_million.sh, MILLION_ROW_BENCH.md}— reproduction harness.
Tests
88 → 92 #[pg_test] cases. Two added with the cache wiring
(index_am_cache_hits_on_second_query,
index_am_cache_invalidates_on_insert); two added with the GUC
(search_k_guc_round_trip, index_am_count_star_does_not_error).
All green on PostgreSQL 16 and 17.
Known follow-ups (not blocking 1.0)
- Cold-cache p50 on a fresh backend is still dominated by
IdMapIndex::loadgoing through a tmpfile because the upstream crate’s deserialiser only reads from a path. An in-memory load inturbovec(or a relfile-resident page format here) would drop first-query latency from ~32 s to ~tens of ms on a million-row 4-bit index. - The post-cache warm p50 of 3.4 s on debug is debug-build cost,
not algorithm cost; a
--releaserebuild on the same corpus is expected to drop us into the tens-of-ms range.
1.0.0-rc.2 — Unreleased
Phase 20 — real-embedding recall benchmark vs pgvector
The synthetic-only recall numbers in docs/RECALL.md § 2.1 are
now joined by a real-world fixture run against
ann-benchmarks‘ GloVe-100 dataset
(100 000 corpus rows, 1 000 query rows, exact ground truth
recomputed against the subset). Two new bench drivers:
benches/recall_vs_pgvector.rs: a pure-Rust harness that loads a binary fixture (corpus.bin / queries.bin / ground_truth.bin), buildsturbovec::IdMapIndexat bit_width 4 and 2, and reports R@1 / R@10 / R@100, p50/p95/p99 latency, and bytes/row of the serialised index. Drives the kernel directly — no Postgres.benches/scripts/run_recall_vs_pgvector.py: an end-to-end SQL driver that loads pgvector + pg_turbovec into the same cluster, builds an HNSW index and the pg_turbovec index, and runs the same query workload through both. Sweepshnsw.ef_searchto produce a recall-latency curve.benches/scripts/prepare_glove_fixture.py: converts an ann-benchmarks HDF5 file into the binary format that both drivers consume.
Results committed under benches/results/ and the headline table
is published in docs/RECALL.md § 2.1.1. Headline at
bit_width=4 on GloVe-100, 100 000 corpus, 1 000 queries: kernel
R@10 = 0.862 at 744 µs/query (8.4× faster than brute force at
6.25× less storage); SQL R@10 = 1.000 at 315 ms/query (re-rank
fan-out dominates latency — documented as a known cost of the
v1.0 index AM).
Phase 18 — fix munmap_chunk() abort on forced index scan
The forced-index-scan path (SET enable_seqscan = off; SELECT ...
ORDER BY emb <=> q LIMIT k) had been crashing the backend with
munmap_chunk(): invalid pointer (or SIGSEGV) since v0.4. The
crash was tracked as Phase 12’s “known issue” and gated the
index_am_forced_index_scan #[pg_test] case as #[ignore]d
through v1.0.0-rc.1.
Root cause: amrescan passed nkeys * size_of::<ScanKeyData>()
as the count argument to
std::ptr::copy_nonoverlapping::<ScanKeyData>. Rust’s
copy_nonoverlapping<T> takes count in elements of T, not
bytes — so for norderbys = 1 we copied
sizeof(ScanKeyData) (≈ 88) ScanKeyData elements into a slot
sized for one, smashing the IndexScanDesc and adjacent heap
chunks. The crash surfaced lazily, only when glibc later walked
the affected arena. The other 39 tests dodged it because the
planner kept small-table queries on a sequential scan, never
calling amrescan with norderbys > 0.
Secondary fix: with xs_orderbyvals now correctly populated,
the executor’s IndexNextWithReorder path needs the AM to
advertise a lower bound on the recomputed orderby distance.
We now write f64::NEG_INFINITY into xs_orderbyvals[0] so
cmp_orderbyvals(recomputed, am_supplied) is always ≥ 0,
guaranteeing the executor never trips its “index returned tuples
in wrong order” assertion. Every tuple goes through the reorder
queue and is drained in exact order at end-of-scan; the cost is
negligible because we cap at k = 1024 results per scan.
Tests
- 40/40
#[pg_test]cases pass withexperimental_index_am, including the previously-#[ignore]dindex_am_forced_index_scan.
1.0.0-rc.1 — 2025
Phase 17 — release-candidate prep
First release-candidate. The default + experimental_index_am
builds are both green (39/39 #[pg_test] cases, 1 documented
#[ignore]); every public surface has at least one passing
test; user-facing docs are complete.
Cleanup
- Removed unused imports and
#[allow(dead_code)]-annotated the one remaining intentionally-unused constant (STRAT_ORDER_BY). - Default
cargo build --features pg16now produces zero warnings.
README
- Status banner reflects v1.0.0-rc1 reality: 39/39 tests, real cluster, documented limitations.
- New “Documentation” section linking every docs/ file from a single index.
What’s in the box
Stable user-facing API:
vectortype with text I/O, full operator suite (<-> <#> <=> <+>).- Distance functions, helpers, element-wise arithmetic.
avg(vector)/sum(vector)aggregates withf64accumulators.- Casts to/from
real[]/double precision[]/integer[]/jsonb. subvector,vec_normalize,vec_check_dim,vec_zeros,turbovec_self_score,vec_random_unit.turbovec.knn(rel, id_col, vec_col, query, k, bit_width, allowed)function-driven ANN with optionalbigint[]allowlist (in-kernel filter, not post-filter).turbovec.*GUC namespace.CREATE INDEX ... USING turbovecaccess method with operator classesvec_ip_ops(default,<#>) andvec_cosine_ops(<=>).CREATE INDEX CONCURRENTLYsupport.- aminsert / ambulkdelete via VACUUM / REINDEX all functional.
Known limitations:
- Forced index path (
SET enable_seqscan = off; ORDER BY emb <=> q LIMIT k) crashes withmunmap_chunk()in the executor’s recheck-orderby memory management. Workaround:turbovec.knn(). Tracking indocs/INDEXAM.md. - L2 / L1 distances are exact-only — no index acceleration.
- Halfvec / sparsevec types are not provided.
0.16.0 — Unreleased
Phase 16 — informed cost estimate + end-to-end demo script
Better amcostestimate. v0.4..v0.15 returned constants
(startup = 1.0, total = 10.0). v0.16 reads the actual
n_vectors, dim, and bit_width from turbovec.am_storage
and computes a SIMD throughput model:
- 8 ns per scored vector at d=1536, bit_width=4 (calibrated
against
cargo bench --bench distanceon AVX2). - Linear scaling with
dim * bit_width / (1536 * 4). - Startup cost =
1 + log2(n_vectors)to model the cache load. - Pages estimate =
n_vectors * (dim * bit_width / 8 + 4) / 8192.
The planner now has real numbers to compare our index against
Seq Scan / Sort plans. Falls back to (1000, 384, 4) if the
side-table row is missing (typical immediately after CREATE
INDEX before commit).
tests/03_full_demo.sql (NEW, 109 lines)
psql script exercising every public feature end-to-end:
- vector type literals + dims/norm/normalize
- All four distance operators with hand-checked numeric answers
- Element-wise arithmetic
- real[]/jsonb casts (both directions)
- subvector / vec_zeros / vec_check_dim
- avg/sum aggregates
- turbovec.knn() unfiltered + with bigint[] allowlist
- CREATE INDEX, aminsert via INSERT, ambulkdelete via DELETE+VACUUM, REINDEX — with side-table assertions
- GUC visibility
- Diagnostics (version, self-score)
Verified to run cleanly against the dev cluster with no ERRORs:
psql -d demo -f tests/03_full_demo.sql.
Verified
cargo pgrx test pg16 -> 39 ok / 0 failed / 1 ignored
psql -f tests/03_full_demo.sql -> all sections complete cleanly
0.15.0 — Unreleased
Phase 15 — functional ambulkdelete (39 tests pass)
v0.4..v0.14 had a stub ambulkdelete that did nothing — deleted
rows accumulated in the index until the user ran REINDEX.
v0.15 implements actual delete handling. We now track every live
u64 id in a parallel Vec<u64>, persisted as a new
live_ids bytea column on turbovec.am_storage. ambulkdelete
walks the live-ids list, calls the supplied bulk-delete callback
for each id (after decoding back to ItemPointerData), removes
those flagged dead from both the IdMapIndex and the live-ids
list, and persists the result.
Schema migration
am_storage gains a live_ids bytea NOT NULL DEFAULT ''::bytea
column, added via an IF NOT EXISTS DO $$ ... $$ block in
extension_sql!. Existing rows from v0.14 and earlier get an
empty live_ids, which means a single REINDEX repopulates the
list correctly.
Source
src/index/persist.rs:StoredIndexgainslive_ids: Vec<u64>.save()takes&[u64]for the live-ids and persists.load()reads the new column, decodes viadecode_live_ids(little-endianu64packing).encode_live_ids/decode_live_idshelpers.
src/index/build.rspasses&state.idstosave()afterindex_build_range_scancollects them.src/index/insert.rspushes the new id intostate.live_idson the success path; CIC-replace path leaves it unchanged.src/index/vacuum.rs(full rewrite): walkslive_ids, calls the callback per id, removes dead ones, persists. Reportstuples_removedin the IndexBulkDeleteResult.src/index/mod.rs: schema migration block adds thelive_idscolumn conditionally; bothpayloadandlive_idscolumns areSTORAGE EXTERNAL(no PGLZ).src/lib.rs:index_am_vacuum_removes_dead#[pg_test]verifies that DELETE + REINDEX leaves the side-table reflecting only the surviving rows.
Verified
cargo pgrx test pg16 -> 39 ok / 0 failed / 1 ignored
0.14.0 — Unreleased
Phase 14 — recall benchmark + pgvector migration cookbook
benches/recall.rs— pure-Rust recall harness usingcriterion. Generates 1 000 deterministic random unit-norm vectors per(dim, bit_width), builds aturbovec::IdMapIndex, runs 50 random queries, computes R@1, R@10, R@100 against a brute-force ground truth. Output is one JSON line per criterion sample for downstream tooling.benches/results/recall_2026_05_21.json— first run results. Headlines: 4-bit hits R@1 ≈ 0.80 across 128/384/768 dims; 2-bit costs ~40 R@1 points; R@100 reaches 0.93 at 4-bit. These are random corpus numbers — real embeddings recall better because they have clustering structure for the quantiser to exploit.docs/RECALL.md— “Latest results” table now populated.docs/MIGRATING_FROM_PGVECTOR.md(NEW, 200 lines) — cookbook covering: coexistence, single-column conversion viareal[]bridge (one-shot + batched), CIC build, query rewrite table (pgvector → pg_turbovec), filtered-ANN pattern that pushes the WHERE into the SIMD kernel, aggregates withf64accumulators, full feature comparison table, and “when not to migrate” honest section (halfvec/sparsevec gaps, L2-dominated workloads, real-embedding recall floor).
Verified
cargo bench --bench recall --no-default-features --features pg16 -> 6 configs run
cargo pgrx test pg16 -> 38 ok / 1 ignored
0.13.0 — Unreleased
Phase 13 — CREATE INDEX CONCURRENTLY support (38/38 pass)
CIC works end-to-end. The fix exposed a real bug in aminsert:
CIC’s two-pass build calls ambuild + validate, and validate
invokes aminsert for every in-snapshot row — some of which
ambuild already inserted. v0.12 raised
IdAlreadyPresent(1) and the index ended up INVALID.
Fix: aminsert is now idempotent. On IdAlreadyPresent it
removes the existing slot and re-adds, preserving n_vectors.
This also covers HOT updates that fire aminsert with the same
CTID more than once.
Source
src/index/insert.rs: catchIdAlreadyPresentfromIdMapIndex::add_with_ids, callIdMapIndex::remove(id), then re-add. n_vectors stays the same on replace.src/lib.rs:index_am_create_index_concurrently#[pg_test]exercises the CIC syntax inside the pgrx test framework’s enclosing transaction (where PG ERRORs SQLSTATE 25001 — we treat that as “syntax accepted” and verify the AM works under a normal CREATE INDEX in the same test).
Manual verification (psql, no transaction wrapper)
CREATE TABLE cic_demo (id bigint PRIMARY KEY, emb vector);
INSERT INTO cic_demo VALUES (1, '[1,0,0,0,0,0,0,0]'), ...;
CREATE INDEX CONCURRENTLY cic_demo_idx
ON cic_demo USING turbovec (emb vec_cosine_ops);
\d cic_demo
Indexes:
"cic_demo_idx" turbovec (emb vec_cosine_ops) -- valid, no INVALID marker
Before v0.13 this terminated with
ERROR: turbovec aminsert: add_with_ids failed: IdAlreadyPresent(1)
and left the index marked INVALID.
Verified
cargo pgrx test pg16 -> 38 ok / 0 failed / 1 ignored
0.12.0 — Unreleased
Phase 12 — forced-index-scan investigation
Added a stress test index_am_forced_index_scan that calls
SET enable_seqscan = off to force the planner onto our index
path. The test reliably crashes the backend with
munmap_chunk(): invalid pointer (glibc free abort) somewhere in
the executor’s recheck-orderby path. Marked the test
#[ignore] with a precise reproducer comment so Phase 13 can
pick it up.
During debugging:
- Allocated
xs_orderbyvals/xs_orderbynullsinambeginscan(PG core does NOT do this for AMs that advertiseamcanorderbyop = true). This fixed an earlier SIGSEGV in the projection path; it did not fix the forced-index-scan crash. - Tried
Box::leak-ing theStoredIndexreturned bypersist::load, in case turbovec’sIdMapIndex::Dropwas freeing memory across an allocator boundary. Did not help. - Tried setting
xs_recheck = truein addition toxs_recheckorderby = true. Did not help. - Confirmed the crash is not in our amgettuple body — a
stub returning
falsewith no result-vector writes still triggersmunmap_chunk().
Working theory: the executor’s recheck-orderby path frees a
Datum-pointed object the AM is supposed to manage. Phase 13 will
gdb the crash to identify the exact free() call site.
Workaround for users
The planner-picks-naturally path works (37/37 tests pass
including the AM). The index_am_create_and_query /
index_am_aminsert_path / index_am_recall_64_rows /
index_am_2bit_round_trip / index_am_realistic_dim_384 tests
all exercise small/medium tables where enable_seqscan = on
(the default) keeps the planner on seqscan and the AM is used
only via CREATE INDEX storage — not yet via query plans.
For larger corpora, recommend turbovec.knn() (same SIMD
kernel, no executor-recheck path).
Source
src/index/scan.rs:ambeginscanallocates the order-by arrays;amgettuplepopulates them. Net behaviour unchanged on the test path; remains broken underenable_seqscan = off.src/lib.rs:index_am_forced_index_scan#[pg_test],#[ignore]-d with a reproducer and link to the docs.docs/INDEXAM.md: “Phase 12 known issue” section documenting the crash, hypothesis, workaround, and Phase 13 plan.
Verified
cargo pgrx test pg16 -> 30 ok / 0 failed
cargo pgrx test pg16 --features experimental_index_am -> 37 ok / 1 ignored
0.11.0 — Unreleased
Phase 11 — realistic-scale tests + 2-bit round-trip + psql regression
Proves the index AM scales to real-world dimensionality and to the most-compressed bit_width.
New tests
index_am_realistic_dim_384— 200 deterministic 384-dim vectors (typical sentence-embedding dim). Asserts:am_storage.n_vectors = 200after CREATE INDEX.- Self-vector is rank 1 in
ORDER BY emb <=> q LIMIT 1. - Self-vector lands in top-10.
index_am_2bit_round_trip— 100 vectors at d=128 withWITH (bit_width = 2). Verifies the tightest TurboQuant mode works end-to-end and the side table recordsbit_width = 2. Self-recall in top-20 (relaxed from top-10 because 2-bit costs ~2 R@k points).
New psql regression script
tests/02_index_am.sql— walks through CREATE INDEX, EXPLAIN, aminsert via INSERT, REINDEX, DROP INDEX, then a hybrid retrieval example usingturbovec.knn(...)with a SQL-derived allowlist. Run viacargo pgrx run pg16then\i tests/02_index_am.sql.
Verified
cargo pgrx test pg16 -> 37 ok / 0 failed
0.10.0 — Unreleased
Phase 10 — filtered search via IdMapIndex::search_with_allowlist
The headline feature from upstream turbovec’s API is now wired
through to SQL. turbovec.knn() gains an optional allowed
bigint[] argument:
-- Restrict candidates to a tenant or topic without paying the
-- cost of a post-filter:
SELECT k.id
FROM turbovec.knn(
'docs'::regclass, 'id', 'embedding',
$1::vector, 10, 4,
ARRAY(SELECT id FROM docs WHERE tenant_id = $2)::bigint[]
) k
ORDER BY k.score DESC;
The SIMD kernel honours the allowlist at 32-vector block granularity — selective filters cost less, not more. With the allowlist passed inside the kernel, blocks containing zero allowed slots short-circuit before any LUT lookup.
SQL signature
turbovec.knn(
rel regclass,
id_col text,
vec_col text,
query vector,
k integer,
bit_width integer DEFAULT 4,
allowed bigint[] DEFAULT NULL
) RETURNS TABLE(id bigint, score double precision)
When allowed is NULL or omitted, behaviour is identical to v0.9
(unfiltered IdMapIndex::search). When non-NULL the function
sorts and dedupes the array, then calls
IdMapIndex::search_with_allowlist. Empty allowlist returns zero
rows.
Source
src/knn.rs: factored search dispatch into arun_search()helper used by both the cache-hit and miss paths. The dispatch picksIdMapIndex::search(unfiltered) orIdMapIndex::search_with_allowlist(query, k, Some(&buf))depending on whetherallowedwas passed.src/lib.rs:knn_filtered_allowlist#[pg_test]covers four sub-cases: unfiltered baseline, two-id allowlist, single-id allowlist, empty allowlist (returns 0 rows).
Verified
cargo pgrx test pg16 -> 35 ok / 0 failed
0.9.0 — Unreleased
Phase 9 — index AM promoted to default + AM scan path uses the cache
After v0.7’s hardening (32/32 AM tests) and v0.8’s cache work, the
turbovec index access method is promoted out of the experimental
feature gate and into the default build:
[features]
default = ["pg16", "experimental_index_am"]
A stripped-down build without the AM is still available via
cargo build --no-default-features --features pg16.
Source
src/index/scan.rs:amgettuplenow consults the sharedcrate::cachebefore falling back topersist::load. On cache hit the scan skips:- The
am_storagerow read (one PG round-trip). - The bytea ->
IdMapIndexdeserialization (TVIM file load via a tempfile dance — substantial cost on large indexes). Cache validity is the same as the function path: relfilenode - n_vectors, plus LRU under
turbovec.cache_size_mb.
Cache key uses
attnum = 0to distinguish the AM’s index relation fromturbovec.knn()’s heap-relation entries (which use the column attnum).- The
Cargo.toml:experimental_index_amadded to default features but kept as an opt-out feature.
Verified
cargo pgrx test pg16 -> 34 ok / 0 failed
cargo build --no-default-features --features pg16 -> builds clean
0.8.0 — Unreleased
Phase 8 — backend-local cache for turbovec.knn()
turbovec.knn(rel, id_col, vec_col, query, k, bit_width) previously
rebuilt the entire IdMapIndex from the heap on every call. v0.8
introduces a backend-local cache keyed by
(rel_oid, attnum, bit_width, dim):
- First call in a backend pays the build cost as before
(heap scan via SPI,
IdMapIndex::add_with_ids). - Subsequent calls with the same key, on a relation whose
pg_class.relfilenodeandcount(*)haven’t changed, skip rebuild and reuse the cachedArc<IdMapIndex>. - DML invalidates implicitly — INSERT / UPDATE / DELETE
changes
count(*); CLUSTER / VACUUM FULL / TRUNCATE / REINDEX changesrelfilenode. Either mismatch forces a rebuild on the next lookup. - LRU eviction keeps total cache bytes within
turbovec.cache_size_mb(default 256 MiB; setting to 0 disables caching entirely).
Source
src/cache.rs(NEW, 175 lines)CacheKey { rel_oid, attnum, bit_width, dim }.Entry { index: Arc<IdMapIndex>, bytes, relfilenode, n_rows, seq }.- Public API:
lookup,insert,invalidate,current_relfilenode,len. - LRU enforcement against
turbovec.cache_size_mb.
src/knn.rsrewired:- On entry, computes the cache key and
lookups. Hit fast-paths straight toIdMapIndex::searchon the cachedArc. - Miss path builds as before, then calls
cache::insertwith an estimated byte size (dim * bit_width / 8 + 4 + 64per vector) before returning.
- On entry, computes the cache key and
src/lib.rsmounts the cache module and adds two#[pg_test]cases:knn_cache_hit_after_first_call— second call returns the same answer;crate::cache::len() >= 1confirms the entry survives.knn_cache_invalidates_on_insert— INSERT a closer row after the warmup; the nextknn()call returns the new row (proving the cache detected thecount(*)change and rebuilt).
Verified
cargo pgrx test pg16 -> 29 ok / 0 failed
cargo pgrx test pg16 --features experimental_index_am -> 34 ok / 0 failed
0.7.0 — Unreleased
Phase 7 — hardened index AM, four new end-to-end tests, real bug fixes
The v0.6 index AM passed a single happy-path test. This release adds
four more #[pg_test] cases that uncovered — and fixed — four
real bugs in the AM:
index_am_aminsert_path— build, insert, query. Verifiesaminsertactually grows the side-table payload and that the newly inserted row is returned by subsequent ORDER BY queries.index_am_recall_64_rows— 64 deterministic 16-dim vectors, build, query the corpus’s own row-17 emb, assert it lands in the top-10. (Top-1 is too tight at 4-bit quantisation; top-10 is the recall floor we won’t ship below.)index_am_reindex—REINDEX INDEX foosucceeds and the side-table payload reflects the rebuild.index_am_rejects_bad_bit_width—WITH (bit_width = 5)raises ERROR cleanly without crashing the backend.
Bug fixes uncovered by the new tests
- Missing
#[pg_guard]on AM callbacks caused apgrx::error!insideamoptions(“bit_width must be in 2..=4”) to unwind across the FFI boundary, segfault the backend with signal 6, and cascade to every later test in the run. Everyextern "C-unwind"callback insrc/index/now wears#[pg_guard]. - SPI in
ambuildcouldn’t survive REINDEX — the planner inside SPI tried to AccessShareLock the very index being rebuilt, hittingcannot access index ... while it is being reindexed. Replaced with a direct call to the table AM’sindex_build_range_scancallback ((*heap_rel.rd_tableam) .index_build_range_scan) plus a freshbuild_callbackthat populates aBuildStatethread-locally. Same path the built-in btree / GIN / hash AMs use; no SPI lock surface. - Random-vector test data was identical across rows — PG
materialised
(SELECT random() FROM generate_series(1,16))once per query and reused it for every INSERT row, so the recall test was actually scoring 64 copies of the same vector (all distances zero, false negatives). Switched to ahashtext(i::text || ':' || k::text) % 2000 / 1000.0 - 1per-element formula that’s stable per(i,k)and varies across rows.
Source changes
src/index/build.rs: full rewrite ofambuildas aBuildState+index_build_range_scan+build_callbackpipeline (no SPI). The callback validates dim consistency, optionally L2-normalises, and accumulates(u64, Vec<f32>)rows into the per-build state.src/index/{build,cost,insert,options,scan,vacuum,validate}.rs: every AM callback now has#[pgrx::pg_guard].src/lib.rs:index_am_aminsert_path,index_am_recall_64_rows,index_am_reindex,index_am_rejects_bad_bit_width.
Verified
cargo pgrx test pg16 -> 27 passed; 0 failed
cargo pgrx test pg16 --features experimental_index_am -> 32 passed; 0 failed
This is the first release where aminsert and REINDEX are
actually proven to work.
0.6.0 — Unreleased
Phase 6 — validated against a real PostgreSQL 16 cluster
This is the first release where every #[pg_test] case has actually
been executed and passes. The default-feature build runs 28/28
tests green; the experimental_index_am-feature build also runs
28/28, including a new end-to-end index_am_create_and_query
test that:
CREATE TABLEs an 8-dimvectorcolumn,- inserts four rows,
CREATE INDEX ... USING turbovec (... vec_cosine_ops) WITH (bit_width = 4),- asserts the side-table row was created with
n_vectors = 4, - runs
ORDER BY emb <=> $1 LIMIT 1and asserts the right row is returned, DROP INDEXand verifies the heap is intact.
Fixes uncovered by running the suite
- Aggregate transition function was implicitly STRICT (pgrx
derives it from non-Option args), causing
CREATE EXTENSIONto fail withmust not omit initial value when transition function is strict and transition type is not compatible with input type. Bothvec_accumandvec_combinenow acceptOption<VecAccum>so pgrx generates non-strict SQL. trusted = trueinpg_turbovec.controlwas rejected by pgrx 0.17’s control-file parser asRedundantField. Removed.- Default
cargo pgrx test pg16build target — switched the Cargodefaultfeatures topg16so the local Nix-installed PostgreSQL 16 cluster is the one exercised. Runs against pg17 / pg18 still work via the matching feature flag. - build.rs propagates the
openblaslink directive fromturbovec(transitive dep) into ourcdylib’sDT_NEEDED, fixingLOAD 'pg_turbovec'failing withundefined symbol: cblas_sgemm. - Index AM scaffold compile errors against pg16 IndexAmRoutine:
amcanbuildparallelandaminsertcleanupare pg17+ only; feature-gated.pg_externcannot returnpg_sys::Datum; rewroteturbovec_index_handleras a hand-rolledextern "C-unwind"wrapper plus a manualpg_finfo_*companion (the same shape pgrx generates internally for#[pg_extern]functions).pg_sys::TupleDescAttrisn’t exposed as a Rust function in pgrx 0.17; rewroteresolve_indexed_attrto use(*tupdesc).attrs.as_slice(natts).(*indrel).indkey.values[0]doesn’t compile against an__IncompleteArrayField; replaced with.as_slice(nkey).Spi::connectexposes only&SpiClient; switched the write paths inpersist.rstoSpi::connect_mut.- Implicit autoref on
(*opaque).results[(*opaque).cursor]against a raw pointer; rewrote with explicit&(*opaque)borrow scope.
- Test fixture:
pg_testcases that use bare operator symbols nowSET search_path = turbovec, publicfirst.
Added
docs/BUILDING.mddocumenting the Nix-specific build dance (writable pg_config wrapper, libclang / glibc include flags, openblas RUSTFLAGS, ICU sidestep).index_am_create_and_query#[pg_test]case (gated by theexperimental_index_amCargo feature).
Changed
- Default Cargo
defaultfeatures set to["pg16"](was["pg17"]) to match the local development cluster.
0.5.0 — Unreleased
Added — Phase 5: pgvector-parity helpers
subvector(vector, start integer, length integer) -> vector— 1-indexed slice. Bounds-checked; raisesERRORon overrun.vec_to_jsonb(vector) -> jsonbandjsonb_to_vec(jsonb) -> vectorplus explicit casts in both directions. Useful for replication via JSONB columns, logging, and audit trails.vec_check_dim(vector, integer) -> vector— runtime dim assertion. Use as aCHECKconstraint when typmod-style enforcement is wanted without the full typmod plumbing.vec_zeros(integer) -> vector— zero-vector helper; identity forsum(vector)in extension queries.vec_to_text(vector) -> text— explicit text rendering callable from SQL (the type’s OUTPUT function as a regular function).
Tests
subvector_basic,subvector_out_of_bounds,jsonb_round_trip,check_dim_passes_and_fails,zeros_helper.
Changed
Cargo.toml/pg_turbovec.controlbump to0.5.0.migrations/004_pg_turbovec_v0.5.0.sqlreference mirror.
0.4.0 — Unreleased
Added — Phase 4: experimental turbovec index access method (opt-in)
A full IndexAmRoutine-based access method is now scaffolded under
src/index/, gated behind the experimental_index_am Cargo
feature. Default builds do not include it; the v0.3 surface
(type, operators, aggregates, turbovec.knn()) remains the only
stable user-facing API.
Build:
cargo pgrx install --release --features experimental_index_am
Use:
CREATE INDEX docs_emb_idx
ON docs USING turbovec (embedding vec_cosine_ops)
WITH (bit_width = 4);
SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10;
Source layout (src/index/)
mod.rs—IndexAmRoutinepopulator and theturbovec_index_handler(internal) RETURNS index_am_handlerSQL function. Also emits theCREATE ACCESS METHOD turbovec,CREATE OPERATOR CLASS vec_ip_ops, andCREATE OPERATOR CLASS vec_cosine_opsdeclarations viaextension_sql!.options.rs—bit_width(2…=4) anddim(0 = auto, else positive multiple of 8) reloption parsing under the AM-side callbackamoptions.persist.rs— SPI-backed read/write ofturbovec.am_storage (indexrelid, bit_width, dim, n_vectors, payload, version, updated_at).payloadisSTORAGE EXTERNAL(no PGLZ on already-quantised bytes).build.rs—ambuild(heap scan via SPI, buildsIdMapIndex, persists) andambuildempty(writes empty marker).insert.rs—aminsert(load-then-update; v0.5 will batch).scan.rs—ambeginscan/amrescan/amgettuple/amendscanwith aScanOpaquecarrying the query vector and cached result list. ORDER-BY-only scans are required.vacuum.rs—ambulkdelete/amvacuumcleanupstubs (Phase 5 needs an upstream way to enumerate live ids inIdMapIndex).cost.rs—amcostestimateconstant heuristic so the planner picks us over a full sort.validate.rs—amvalidatereturnstrue(Phase 5 will check opclass strategy numbers).
CTID encoding
We use pgrx’s canonical 32 / 16 packing (item_pointer_to_u64):
block number in the top 32 bits, offset number in the bottom 16,
upper 16 reserved for a future epoch. This gives IdMapIndex u64
ids natural ordering inside a relfile and lets amgettuple fill
xs_heaptid via u64_to_item_pointer directly.
Capability flags
amstrategies = 0
amsupport = 1
amcanorder = false
amcanorderbyop = true
amcanbackward = false
amcanunique = false
amcanmulticol = false
amoptionalkey = true
amstorage = true
amcanparallel = false // Phase 5
amcanbuildparallel = false // Phase 5
amusemaintenanceworkmem = true
Status
Untested against a running cluster. This release is the
complete scaffold ready for a Phase 5 session that has
cargo-pgrx and a Postgres dev cluster: cargo pgrx test pg17
--features experimental_index_am is the gate. Known follow-ups
are enumerated in docs/INDEXAM.md § “Test plan” and § “Known
risks”.
Added — docs
docs/INDEXAM.md— implementation guide for the index AM (callback responsibilities, side-table schema, test plan, known risks).migrations/003_pg_turbovec_v0.4.0.sql— reference mirror of the SQL surface that ships only when the feature is enabled.
Changed
Cargo.tomladdslibc = "0.2"(used bypersist.rsfor pid-stamped tempfile paths) and theexperimental_index_amCargo feature.pg_turbovec.controldefault_versionbumped to0.4.0.src/lib.rsmountsmod indexonly under#[cfg(feature = "experimental_index_am")].
0.3.0 — Unreleased
Added — Phase 3: kernels module, benches, CI, docs
src/kernels.rs— pure-Rust math kernels (dot,l2_sq,l1_abs,norm2,cosine_distance,normalise_into,normalise_to_vec). Distance and normalisation code indistance.rs/normalize.rsnow delegate to this module so the kernels are exercisable under plaincargo test --no-default-featureswithout booting Postgres.vec_random_unit(integer)— random unit-normvector, for benchmarking and recall scaffolding.benches/distance.rs—criterion-based micro-benchmarks of the distance kernels at d=128, 384, 768, 1536, 3072. Runs viacargo bench --bench distance --no-default-features.- Codeberg Woodpecker CI (
.woodpecker/ci.yaml) — three pipelines: pure-Rust unit tests + clippy on every push;cargo pgrx test pg17onmain/ release branches. docs/USAGE.md— cookbook with install, exact search, ANN viaturbovec.knn(), aggregates, arithmetic, GUCs, pgvector coexistence migration, diagnostics.docs/RECALL.md— recall/perf benchmark methodology, matched-bit-budget comparison plan against pgvector for v0.4.- Pure-Rust unit tests in
kernels::testscovering every kernel plus a precision regression (1 048 576-element sum of squares stays within 1 ppm of the f64 truth).
Changed
Cargo.tomladdsrand = "0.8",criterion = "0.5"(dev), declares[[bench]] name = "distance".pg_turbovec.controldefault_versionbumped to0.3.0.
0.2.0 — Unreleased
Added — Phase 2: function-driven ANN
turbovec.knn(rel regclass, id_col text, vec_col text, query vector, k int, bit_width int default 4)— function-driven ANN backed byturbovec::IdMapIndex. ReturnsTABLE(id bigint, score float8), ordered by score DESC for most-similar-first.- Optional unit-normalisation via
turbovec.normalize_on_insertGUC; constraintsk > 0,bit_width ∈ {2,3,4},dim % 8 == 0. migrations/002_pg_turbovec_v0.2.0.sqlreference mirror.#[pg_test]cases forknn_returns_nearest_firstandknn_rejects_bad_k.
Removed
src/phase2_knn.rsscaffold — promoted to mountedsrc/knn.rs.
Added — Phase 1: type, operators, functions, aggregates
vectortype — variable-dimensionf32vector, stored as a CBOR-serialised varlena viapgrx::PostgresType. Text I/O accepts'[1, 2, 3]'with whitespace tolerance and rejects NaN / ±Inf. Hard cap at 16 000 dimensions, matching pgvector.- Distance operators between
vectoroperands:<->Euclidean (L2)<#>negative inner product (soORDER BY a <#> bsorts most- similar-first under ASC, mirroring pgvector)<=>cosine distance (1 - cos θ, clamped to[0, 2])<+>taxicab (L1)
- Distance functions:
l2_distance,l2_squared_distance,inner_product,negative_inner_product,cosine_distance,l1_distance. - Helper functions:
vector_dims,vector_norm,vec_normalize. - Element-wise arithmetic:
vec_add(+),vec_sub(-),vec_mul(*). - Aggregates:
avg(vector)andsum(vector). Internal state usesf64accumulators to preserve precision on large corpora. Both arePARALLEL SAFE;combinefnmerges partial states. - Casts (explicit only):
real[]→vectordouble precision[]→vectorinteger[]→vectorvector→real[]
- GUCs under the
turbovec.*namespace:bit_width_default(int, default 4, range 2..=4)cache_size_mb(int, default 256, range 0..=65536)warn_on_rebuild(bool, default true)search_concurrency(int, default 1, range 1..=128)normalize_on_insert(bool, default true)
- Diagnostic:
turbovec_self_score(vector, bit_width)exercises the upstreamturbovec::IdMapIndexend-to-end and returns the self-score, used by the test suite as an integration tripwire.
Tests
#[pg_test]cases insrc/lib.rs::testscovering text I/O, every operator, dimension-mismatch ERROR, aggregates, casts, normalisation, and a turbovec round-trip.tests/01_type_basic.sql— psql-style regression script.
Project layout
pgrx = "=0.17.0"to match the cached toolchain.pg_turbovec.controldeclares schematurbovec,relocatable = false,trusted = true.migrations/001_pg_turbovec_v0.1.0.sqlmirrors the generated SQL surface (the authoritative file is generated bycargo pgrx schema).
Not yet shipped (Phase 2 / Phase 3)
- Index access method
turbovecand operator classesvec_ip_ops,vec_cosine_ops. A starter is checked in atsrc/phase2_knn.rs(not yet mounted bylib.rs). - Filtered search via
IdMapIndex::search_with_allowlist. - Binary-compatible varlena layout with pgvector’s
vector. - WAL-logged persistent index pages.