Contents
Testing coverage and known gaps
This document describes what the pg_turbovec test suite covers, what
it deliberately does not cover, and why. It exists because a
silent wrong-results regression (the “pre-AVX2 bug”, below) shipped
once already through a gap that looked, at a glance, like good
coverage. Read this before assuming a class of bug is tested.
How to run
cargo pgrx test pg16 # full suite on PG 16 (local loop)
bash scripts/drift-check.sh # project-level invariants (docs vs code)
CI runs cargo pgrx test pg<N> for N in 13..=18 (see docs/CI.md).
The bug this coverage exists to catch
A previous turbovec revision returned silently wrong ANN results on CPUs without AVX2: instead of the top-N distinct neighbours it returned the same TID N times. It shipped undetected because:
- No automated test exercised a corpus larger than ~2000 rows.
The bug only manifested at scale and non-trivial dimensionality on
a specific CPU class. Every
#[pg_test]used tiny corpora (often 64 rows, never more than 2000). - No recall-regression test compared against a real ground truth at meaningful scale. A 64-row “recall” test is far below where quantiser behaviour is observable.
- CI ran only unit tests on AVX2 hardware, so the scalar fallback
path was never exercised by
pg_turbovec’s own CI.
The single cheapest guard against the whole “wrong-ranking” class is a distinct-ids assertion on every ANN result: the duplicate-id bug fails it instantly, regardless of which SIMD path runs or how high recall happens to be.
What the suite covers
- Unit-scale ANN correctness —
#[pg_test]s that build a turbovec index on small corpora (8–384 dims, up to 2000 rows) and assert nearest-neighbour ordering, self-recall, and operator behaviour (<->,<#>,<=>, L1/L2/cosine opclasses). - Distinct-ids invariant — every ANN-scan
#[pg_test]that pulls back more than one id runs the result throughassert_distinct_ids. This is the regression guard for the pre-AVX2 bug class. The helper andfetch_idslive in thetestsmodule insrc/lib.rs. - Medium-scale recall floor —
index_am_recall_floor_{2,3,4}bitbuild a 20 000 × 128 corpus of distinct deterministic random vectors, compute brute-force exact top-10 for 20 held-out queries (forced seqscan, exact<=>), and assert the index’s recall@10 clears a per-bit-width floor and that every returned id is distinct. This is the test the pre-AVX2 bug would have failed: it is ~10× larger than the historical ceiling and asserts distinct ids.- Observed recall@10 = 1.000 at every bit width on this synthetic uniform-random corpus (the vectors are near-orthogonal in 128-d, so even 2-bit TurboQuant separates them cleanly). The floors (4-bit ≥ 0.95, 3-bit ≥ 0.90, 2-bit ≥ 0.80) are therefore catastrophic-collapse guards — they fire well before the ~0.1 recall the duplicate-id bug produced — not fine-grained quality gates. Fine-grained per-bit-width quality is measured by VectorDBBench on real embeddings, not by this unit test.
- In-harness corpus + query + index build is well under the ~30 s budget per test (the whole 3-test set, including compile, runs in ~60 s).
- SQL surface — operator/opclass registration, type round-trips
(
vector,halfvec,sparsevec,bitvec),knn()table function, filtered allowlist, reloption validation (bit_width2..=4). - Iterative scan — refill correctness,
max_scan_tuplesceiling,iterative_scan = offsingle-batch behaviour, and no duplicate TIDs across refill batches. - Index lifecycle — REINDEX, CREATE INDEX CONCURRENTLY, VACUUM /
ambulkdeletedead-tuple removal, parallel vs serial build equivalence (same ranked top-k, same relfile size). - Wire-format stability —
wire_format_version_is_stable(EXPECTED_WIRE_FORMAT_VERSION = 6) and the legacy-metaambeginscanerror paths (is_legacy_v{1,2}→ clearERRORwith aREINDEXhint;is_legacy_v{3,4,5}are deliberately always-false because v3→v4 IVF, v4→v5 ColBERT, and v5→v6 graph are all ADDITIVE per index kind — an older single-vector index decodes unchanged). - Upgrade matrix —
migration_files_cover_documented_versionscross-checksmigrations/against the documented release history.
What the suite does NOT cover (and why)
(a) Large-scale behaviour (1M+ rows)
Unit tests cap at the ~20 000-row recall-floor corpus. Building 1M+
rows in the in-process pgrx harness is too slow for a unit test and
the harness holds the whole corpus in one transaction. The
large-scale evidence is the VectorDBBench run (see benches/ and
docs/RECALL.md), not the unit suite. A regression that only appears
above ~20 000 rows would not be caught by cargo pgrx test; it is the
benchmark’s job.
(b) The pre-AVX2 scalar fallback path
CI runs on GitHub ubuntu-latest, which is AVX2-capable. turbovec
selects its SIMD kernel with a runtime is_x86_feature_detected!
check, so on AVX2 hardware the scalar path never runs — and you cannot
force it from pg_turbovec (turbovec’s FORCE_SCALAR_FALLBACK is
pub(crate)). Compile-time -C target-feature=-avx2 does not help:
it changes what the compiler emits, not what the runtime feature-detect
selects on an AVX2 machine, so a CI job built that way would still run
the AVX2 path and prove nothing. We deliberately do not add such a
job — it would be coverage theatre.
The real mitigations for a pre-AVX2 regression are:
- turbovec’s own upstream test —
x86_scalar_fallback_tests::scalar_fallback_matches_simd_topkflipsFORCE_SCALAR_FALLBACKand asserts the scalar top-k matches the SIMD top-k. That test owns this path. - Validate turbovec bumps on a pre-AVX2 host (or QEMU). When
bumping the
turbovecgit rev inCargo.toml, run the suite on a non-AVX2 target before tagging. Therv(riscv64) bench host exercises a non-x86 path; an x86 pre-AVX2 host orqemu-x86_64 -cpu Nehalemcovers the scalar x86 kernel.
A future turbovec bump that reintroduced the scalar-path bug would
not be caught by pg_turbovec CI. Treat the turbovec rev bump as
the trigger for the host/QEMU validation above.
© Cross-PG-version wire compatibility
The matrix tests each PG version independently (build-on-N,
read-on-N). It does not test build-on-16 / read-on-17. The
on-disk format is PG-version-independent and wire-format-stable
(MetaPageData::version = 6 since v1.23.0 — additive per kind: a
single-vector index still emits 4, ColBERT 5, graph 6; was 3 for
v1.4.0–v1.9.x), so this is low risk, but it
is not directly asserted.
(d) Concurrency / races
Beyond the existing CREATE INDEX CONCURRENTLY test and the
parallel-vs-serial build equivalence test, the suite does not stress
concurrent insert/scan/vacuum interleavings. The pgrx harness runs
each test in a single backend, so true multi-backend race coverage
would need an external harness (see benches/sql/).
The “ignored” tests are not skipped tests
cargo pgrx test reports a number of ignored entries. These are
```ignore doctests — SQL usage examples embedded in /// doc
comments that are not valid standalone Rust and so are marked
ignore. They are documentation, not disabled tests. Nothing in the
real test suite is being silently skipped; the ignored count is
purely these doc-comment SQL snippets.
Writing tests that measure a global counter (WAL, LSN, cluster stats)
#[pg_test]s run concurrently against ONE shared cluster. Any
assertion built on a cluster-wide counter therefore also observes every
other test’s activity, and will be nondeterministic in a way that looks
like a bug in the code under test.
This bit hard during the v2.3.0 WAL-amplification work. A regression test
diffed pg_current_wal_lsn() around a flush to prove WAL dropped. Across
three CI runs the same operation measured 32 KB, then 1097 KB, then
1425 KB, and one run ranked the fixed and unfixed arms inverted —
i.e. the measurement “disproved” a fix that was in fact working. The
numbers were other tests' WAL.
Rules that follow:
- Count the thing you control, not a global. The fix was a test-only
AtomicU64counting pages the code path registered for WAL. Per-backend, deterministic, and a direct proxy for the quantity of interest. - Beware sequencing when a flush mutates what the next one observes. An earlier version measured the two arms back-to-back; because the first flush left every page at its final contents, the second arm had nothing left to write. Reset to an identical starting state for each arm.
- Assert the control arm is non-trivial.
assert!(unfixed > 50)alongsideassert!(fixed * 10 < unfixed)— otherwise “0 vs 0” passes and proves nothing. An earlier version failed exactly this way, because a#[pg_test]’s outer transaction always rolls back beforePreCommitfires, so the flush under test never ran at all. Usexact::flush_to_relfile_for_testto drive a flush in-band. - A wildly varying “before” number is a broken harness, not a flaky fix. Investigate the measurement before relaxing the threshold; loosening it would have shipped a test that asserted nothing.
- If you filter rows to exclude noisy ones, assert the BASELINE survives the
filter. Discarding contended rows and re-checking a ratio is a good habit,
but it silently becomes a lie when the filter removes every baseline row —
you then “prove” the treatment wins against an empty set. This actually
happened: on one benchmark arm all eight
flatbaseline rows were flagged (they ran first, while load was still decaying) and zero survived, while the IVF rows ran later and survived. The comparison would have looked spectacular and meant nothing. Same shape as the control-arm rule above.
The same reasoning applies to pg_stat_* views, checkpoint counters, and
anything else that is cluster-global rather than backend-local.