Skip to content

Fix sampled planner statistics; add autoOptimize and Optimize() - #36

Open
arotolo wants to merge 11 commits into
mainfrom
feature/optimize-sampling
Open

arotolo wants to merge 11 commits into
mainfrom
feature/optimize-sampling

Conversation

@arotolo

@arotolo arotolo commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Fixes the sampled planner statistics that made indexed lookups on a large type read every row of the type (5.0.1–5.3.1), adds autoOptimize / Optimize() / OptimizeAsync() for deferring one-time upgrade maintenance, and keeps the planner able to rank a type's partial indexes against each other.

The bug

PRAGMA optimize on connect/disconnect (since 5.0.1) and the ANALYZE after CreateIndex (since 5.1.1) set analysis_limit = 400. The general (FullTypeName, Partition) index is ordered by type name and the types that sort first are tiny, so its sample credited every type with a handful of rows (81 on a 1.68M-row production store whose largest type held 586,986), while a partial index's sample sat inside one key and credited every key with hundreds (401). The planner preferred the general index and read the whole type on each indexed lookup: 82–366 ms per lookup on that store, under a millisecond with correct statistics. Results were never wrong; only speed.

The design

The branch's first design kept full statistics with PRAGMA optimize, and its second removed statistics entirely. Review found that the second one regresses two indexed filters ANDed — with no sqlite_stat1 the planner takes whichever partial index was created last — and that both suffer from the general index's own statistics row. What ships:

  • Full statistics at index creation. CreateIndex runs PRAGMA analysis_limit = 0; ANALYZE JsonValue; after it builds an index, and again when it re-declares an index that was built on an empty type and has rows now. Rows recorded for an empty type are not kept. Nothing samples; PRAGMA optimize is gone from connect, disconnect and dispose.
  • The general index's row is pinned to N N/2 N/2. ANALYZE writes it as the average rows per type, which is wrong for every type when one type dominates a store: it sends small types to a table scan when few types existed at ANALYZE time, and sends a big type's indexed lookups to the general index otherwise. Half the table keeps the general index cheaper than a scan and dearer than any partial index whose match is under half the rows, so the partial indexes' accurate rows decide between themselves. SQLite documents writing sqlite_stat1 for this ("Manual Control Of Query Plans", lang_analyze.html).
  • Upgraded stores are repaired once, gated on PRAGMA user_version. With autoOptimize on (default) the first connect replaces sampled rows with full ones — or discards them when the store has no expression index — and stamps the store; a stamped store pays one header read at connect and writes nothing (schema_version is unchanged across reconnects; a test pins it). With autoOptimize: false the store keeps its previous plans until Optimize() / OptimizeAsync(), which also drop the four redundant 4.x indexes.

Full detail, including the plan table for every query shape on both engines and why the other two designs were rejected: docs/planner-statistics-sampling.md.

Benchmark

dotnet run --project TychoDB.Benchmarks -c Release -- --filter '*PlannerStatistics*' (new class; 40 small types written first, one type with 40,000 rows over 200 keys, indexed on a unique Seq and a 200-value GroupId; built, closed and reopened across three launches like an app). Apple M4 Max, .NET 9.0.17, SQLite 3.50.3, medians of 16 × 1 launch.

Shape main (v5.3.1, sampled) 322be81 (no statistics) This PR
Indexed equality, 200 rows of 40,000 10,108 µs 570 µs 551 µs
Two indexed equalities ANDed, selective index created first 22.7 µs 102.6 µs 22.8 µs
Two indexed equalities ANDed, selective index created last 23.6 µs 22.0 µs 24.5 µs
Whole read of a 10-row type 28.6 µs 31.5 µs 26.1 µs

The first row is the bug (18x). The second is the no-statistics regression this PR avoids; it grows with the rows matching the less selective property (4.6x here at 200 per key, 200x at 10,000 per key, and larger again on SQLCipher, where every extra page is decrypted).

Public API

  • Tycho(..., bool autoOptimize = true) — source-compatible; assemblies that construct Tycho need recompiling.
  • void Optimize(), ValueTask<bool> OptimizeAsync(CancellationToken = default).

Minor release: 5.4.0. Directory.build.props updated to match (the tag still sets the package version).

Tests

379 passed, 0 failed, 4 skipped in both Debug (SQLite 3.50.3) and Encrypted (SQLCipher, SQLite 3.39.2). PlannerStatisticsTests asserts the plan for every shape TychoDB emits (equality, non-selective equality, range, sort with limit, two indexed equalities in both creation orders, unindexed filter, whole type large and tiny, count), the two 5.3.1 failure scenarios, repair with autoOptimize on and off, a store without indexes, a fresh store, and that a current store is not written at connect.

🤖 Generated with Claude Code

Sampled ANALYZE (analysis_limit = 400) on connect, disconnect and after
CreateIndex read only the first ~400 entries of each index. The general
(FullTypeName, Partition) index is ordered by type name, so its sample saw
only the smallest types and credited every type with a handful of rows,
while the sample of a per-type partial index sat inside one key and
credited every key with hundreds. The planner then preferred the general
index and parsed every row of the type on each lookup: 82-366 ms per
lookup on a 1.68M-row store, against 0-1 ms with full statistics. Results
were correct throughout.

CreateIndex now runs a full ANALYZE, and connect/disconnect run
PRAGMA optimize(0x10002), SQLite's form for long-lived connections, which
leaves out the temporary analysis limit the default mask applies.
PRAGMA optimize(0x10002) only re-analyzes a table whose indexes lack
statistics or whose row count moved tenfold, so the sampled rows 5.3.1
wrote survive it and the slow plan persists. A store without the
user_version stamp now gets one full ANALYZE at connect and is stamped;
a stamped store gets PRAGMA optimize(0x10002). The stamp is written only
when the stored value is lower, so a higher version is never lowered.
Connecting a store written by an earlier version drops the four
redundant indexes 4.x created and gathers planner statistics, which can
stall the first launch of a large upgraded store for 10-30 s on the
connecting thread. autoOptimize (default true) keeps that behavior; false
connects without the drops, the statistics step and the ANALYZE after
CreateIndex, leaving them to Optimize()/OptimizeAsync(), which do exactly
the work connect skipped: the drops, then a full ANALYZE on a store
written by an earlier release or PRAGMA optimize(0x10002) otherwise, so a
call with nothing to do costs ~0 ms. Disconnect is unaffected.
SQLite before 3.46 (the SQLCipher build is 3.39.2) ignores the 0x10000
bit and lets PRAGMA optimize consider only tables the connection's
planner has used, so a connection that had run no query yet analyzed
nothing at connect, at disconnect and in Optimize(). A probe read of
JsonValue now precedes the pragma so every build behaves the same.
A write-up for the team (symptom, mechanism, measurements, decisions,
limitations) and a self-contained sqlite3 CLI script that reproduces the
plan flip on synthetic data in about three seconds.
arotolo and others added 6 commits October 10, 2026 00:50
The stamp SQL hard-coded the value GatherStatistics compared against, so
bumping one without the other would have made every connect and Optimize
run a full ANALYZE. One literal now feeds both. Also note, above the
analyze step in ExecuteCreateIndex, that an opted-out index gets its
statistics at Optimize or at the next Disconnect, whichever comes first.
…optimize

PRAGMA optimize's tenfold check estimates a table's row count from the
cell counts down the leftmost path of its b-tree. A few large documents
at the lowest rowids leave that leaf nearly empty and the estimate an
order of magnitude low, so a 2.2 GB production store (1.84M rows,
estimated at 174,370) was re-analyzed on every connect, disconnect and
Optimize(), 20-30 s each, although every index had full, current
statistics. Reproduced with three 3 KB documents before 540k small rows.

Connect, disconnect and Optimize() now re-analyze only when an index of
JsonValue has no sqlite_stat1 row or the exact count(*) differs tenfold
from the row count recorded in the general index's row; the count is a
scan of the covering general index, 36 ms on 1.8M rows. This also makes
the probe read for SQLite 3.39.2 unnecessary.
Every query shape TychoDB emits chooses the right index on SQLite's default
heuristics: a partial index is credited with half the table, so the per-type
index beats the general one for the property it covers whatever the data
looks like. Statistics could only make that choice worse, and did in 5.3.1.
Connect, Disconnect, Dispose, CreateIndex and Optimize no longer run ANALYZE
or PRAGMA optimize, and the user_version stamp and the staleness check go
with them. autoOptimize and Optimize now cover only the legacy index drops.
A store written by 5.3.0 or 5.3.1 holds the sampled sqlite_stat1 rows that
make the planner read every row of a type on each indexed lookup. The connect
script now drops the statistics tables and reloads, so every connect, with or
without autoOptimize, leaves the planner on its default heuristics. The reload
(ANALYZE sqlite_master) leaves an empty sqlite_stat1 behind.
The write-up, CLI repro, README and CHANGELOG now describe the fix as
shipped: no ANALYZE anywhere, leftovers dropped at connect, autoOptimize
and Optimize covering only the legacy index drops. The repro script shows
the plan for every query shape with no statistics; the row-estimate repro
for PRAGMA optimize goes, since the pragma is no longer used.
…eral index row

Removing statistics altogether fixed the sampled-statistics regression but
left the planner unable to rank two partial indexes on one type: with a
filter on both properties it took whichever index was created last, 4.6 ms
(57 ms on SQLCipher) against 0.02 ms for a one-row result at 40,000 rows.

Statistics are now gathered in full, for JsonValue only, when CreateIndex
builds an index or re-declares one that was built on an empty type and has
rows now. Rows recorded for an empty type are not kept. The general index's
sqlite_stat1 row is pinned to "N N/2 N/2": the per-type average ANALYZE
writes there misleads the planner whenever one type dominates a store, into
a table scan for every type when few types existed at ANALYZE time and into
the general index over a big type's partial index otherwise. Half the table
keeps the general index cheaper than a scan and dearer than any partial index
whose match is under half the rows, so the partial indexes' accurate rows
rank against each other.

Upgraded stores are repaired once, gated on PRAGMA user_version: with
autoOptimize on, the first connect replaces sampled rows with full ones (or
discards them when the store has no expression index) and stamps the store;
a stamped store pays one header read at connect and writes nothing, so
schema_version no longer changes on every connect. With autoOptimize off,
Optimize()/OptimizeAsync() perform the repair. PRAGMA optimize stays removed.

Tests cover every query shape on both engines, both index creation orders,
the empty-type-then-bulk-load case, the repair with autoOptimize on and off,
a store without indexes, a fresh store, and that a current store is not
written at connect. A PlannerStatistics benchmark measures the three shapes
whose speed depends on statistics through the public API across launches.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@michaelstonis michaelstonis changed the title Optimizes database statistics management Fix sampled planner statistics; add autoOptimize and Optimize() Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants