Embedded mentat fixes (1.10.0): v1.9.0 vs 6fed515a

The scale suite (benchmarks/scale/run.sh), BACKENDS=embedded SCALES="s m", on the same EC2 c7i.8xlarge (us-east-2, AL2023 kernel 6.12, rustc 1.90), pinned to cores 0-23. Defaults everywhere: REPS=3, MIN_S=10, MAX_S=60, PROBE_S=20.

  • before/: mentat_git 72ee3765. The engine is identical to v1.9.0; it differs only in docs and lockfile lines. Run 2026-09-27T13:43Z.
  • after/: mentat_git 6fed515a (master, not pushed). Run 2026-09-28T01:08Z. 52 of 52 correctness checks pass (after/checks.txt), as before.
  • compare.txt: benchmarks/scale/compare.py before after (threshold 10%).

Default configuration throughout (AutoIndex::Schema, no environment knobs).

Headline (p50, 1 client unless noted)

scenario s before s after m before m after
Store::open (cold_vs_warm) 1621 ms 45 ms 16499 ms 50 ms
point_lookup (q1, unique email) 0.070 ms 0.019 ms 0.515 ms 0.019 ms
ref_traversal (q2) 42.8 ms 22.6 ms 556 ms 290 ms
aggregate (q3, (count ?i) by state) 102 ms 31.2 ms 1352 ms 372 ms
as_of (q2 as of t_mid) 20769 ms 88 ms >20 s (ceiling) 1040 ms
since 48.4 ms 4.2 ms 485 ms 52.5 ms
predicate_scan (q4) 25.6 ms 24.0 ms 280 ms 275 ms
pull 0.042 ms 0.036 ms 0.048 ms 0.044 ms
input_bindings (100-email :in collection) 0.378 ms 0.383 ms 2.03 ms 2.13 ms
write_mixed writer 3.87 ms 3.39 ms 4.17 ms 3.41 ms
write_mixed, 8 readers, reads/s 9.7 577 0.7 36.9

Concurrency sweep: read mix (q1/q2/pull), ops/s

clients 1 8 32 64 128
s before 29 6.9 6.8 ceiling not run
s after 137 996 2095 2093 2101
m before 1.6 ceiling not run not run not run
m after 10.6 72 157 161 160

(“ceiling”: the first call took over PROBE_S = 20 s, so the suite skipped the remaining points.)

What changed (commits on master)

commit fix
60c4a966 transactions of ≥5461 datoms no longer panic
cac32747 an interrupted query returns an error and leaves the Store usable
5072ef13 as-of/since use a covering history index (idx_transactions_aevt); Store::open reads the persisted partition marks and no longer scans the log; schema v2 upgrade
7feb42c1 readers no longer serialize on SQLite’s global page-cache mutex: .cargo/config.toml LIBSQLITE3_FLAGS=-USQLITE_ENABLE_MEMORY_MANAGEMENT -DSQLITE_DEFAULT_MEMSTATUS=0, plus per-connection mmap_size (MENTAT_MMAP_SIZE, default 1 GiB)
bf26af3e adaptive per-attribute value indexes (AutoIndex::Adaptive, Store::tune_indexes)
7ff0a0e0 (count ?x) drops the inner DISTINCT when it is provably redundant, else uses count(DISTINCT ?x); temp_store 2 → 1 (MENTAT_TEMP_STORE)
da2241e2 shared mentat::options_from_json (SQLite and DuckDB extensions, CLI)
266006f6 CLI: .q options, .pull, .eval, .tune, batch mode
8dd332b8 scripting: (q db query & inputs), plus model tests for history, cas and retractEntity
b08f3219 :db/unique attributes, and refs marked :db/index, get a usable value index in the default mode; schema v3 upgrade
6fed515a stale statistics on those value indexes are refreshed on open

Flagged by compare.py (16 rows) and why

  • write_mixed 8-reader read p50 (s 2.1 → 10.2 ms, m 4.2 → 6.3 ms). Throughput went up 59× (s) and 53× (m). Before, the readers were serialized by the page-cache mutex: few reads finished, and those were mostly cheap point lookups. After, all eight readers run the full q1/q2/pull mix concurrently with a writer, so the median read includes q2. p99 fell 97% (s) and 93% (m).
  • write_mixed m writer p99 (5.6 → 8.6 ms). p50 improved from 4.17 to 3.41 ms and writes/s rose 10%. The tail comes from the eight readers that now actually run alongside the writer.
  • cold_vs_warm rows (cold_aggregate m, warm_pull, a few p99s). These are one-shot or 30-sample measurements taken right after dropping the OS page cache. cold_aggregate at m includes paging in a 2.1 GB store (up from 1.7 GB) through mmap. The steady-state rows for the same queries all improved or held: pull is 0.042 → 0.036 ms (s) and 0.048 → 0.044 ms (m), and warm_aggregate is 1333 → 372 ms (m).
  • MISSING / NEW: the before run’s ceilings (as_of m; concurrency s ≥64, m ≥8) are now measured points.

Costs

  • Store size: s 166 → 213 MB, m 1.69 → 2.15 GB (+28%). The additions are the history index (as-of/since) and five schema value indexes (unique email/label name, assignee/reporter/label refs).
  • Bulk load: s 66 → 87 s, m 1482 → 2041 s (6646 → 4825 datoms/s). The same indexes are maintained on every insert.
  • One-time upgrade on first open of an older store: about 21 s at m.

Not changed / known gaps

  • The schema’s partial idx_datoms_avet / idx_datoms_unique_value still never match mentat’s SQL. They’re kept because they back the transactor’s uniqueness checks.
  • :db/index on a scalar (an enum or a number) gets no automatic index. At 1M datoms, q4 went 28 → 55 ms when SQLite started the join at such an index. AutoIndex::Adaptive decides these per workload.