Repository navigation
Conversation
Sampled ANALYZE (analysis_limit = 400) on connect, disconnect and after CreateIndex read only the first ~400 entries of each index. The general (FullTypeName, Partition) index is ordered by type name, so its sample saw only the smallest types and credited every type with a handful of rows, while the sample of a per-type partial index sat inside one key and credited every key with hundreds. The planner then preferred the general index and parsed every row of the type on each lookup: 82-366 ms per lookup on a 1.68M-row store, against 0-1 ms with full statistics. Results were correct throughout. CreateIndex now runs a full ANALYZE, and connect/disconnect run PRAGMA optimize(0x10002), SQLite's form for long-lived connections, which leaves out the temporary analysis limit the default mask applies.
PRAGMA optimize(0x10002) only re-analyzes a table whose indexes lack statistics or whose row count moved tenfold, so the sampled rows 5.3.1 wrote survive it and the slow plan persists. A store without the user_version stamp now gets one full ANALYZE at connect and is stamped; a stamped store gets PRAGMA optimize(0x10002). The stamp is written only when the stored value is lower, so a higher version is never lowered.
Connecting a store written by an earlier version drops the four redundant indexes 4.x created and gathers planner statistics, which can stall the first launch of a large upgraded store for 10-30 s on the connecting thread. autoOptimize (default true) keeps that behavior; false connects without the drops, the statistics step and the ANALYZE after CreateIndex, leaving them to Optimize()/OptimizeAsync(), which do exactly the work connect skipped: the drops, then a full ANALYZE on a store written by an earlier release or PRAGMA optimize(0x10002) otherwise, so a call with nothing to do costs ~0 ms. Disconnect is unaffected.
SQLite before 3.46 (the SQLCipher build is 3.39.2) ignores the 0x10000 bit and lets PRAGMA optimize consider only tables the connection's planner has used, so a connection that had run no query yet analyzed nothing at connect, at disconnect and in Optimize(). A probe read of JsonValue now precedes the pragma so every build behaves the same.
A write-up for the team (symptom, mechanism, measurements, decisions, limitations) and a self-contained sqlite3 CLI script that reproduces the plan flip on synthetic data in about three seconds.
The stamp SQL hard-coded the value GatherStatistics compared against, so bumping one without the other would have made every connect and Optimize run a full ANALYZE. One literal now feeds both. Also note, above the analyze step in ExecuteCreateIndex, that an opted-out index gets its statistics at Optimize or at the next Disconnect, whichever comes first.
…optimize PRAGMA optimize's tenfold check estimates a table's row count from the cell counts down the leftmost path of its b-tree. A few large documents at the lowest rowids leave that leaf nearly empty and the estimate an order of magnitude low, so a 2.2 GB production store (1.84M rows, estimated at 174,370) was re-analyzed on every connect, disconnect and Optimize(), 20-30 s each, although every index had full, current statistics. Reproduced with three 3 KB documents before 540k small rows. Connect, disconnect and Optimize() now re-analyze only when an index of JsonValue has no sqlite_stat1 row or the exact count(*) differs tenfold from the row count recorded in the general index's row; the count is a scan of the covering general index, 36 ms on 1.8M rows. This also makes the probe read for SQLite 3.39.2 unnecessary.
Every query shape TychoDB emits chooses the right index on SQLite's default heuristics: a partial index is credited with half the table, so the per-type index beats the general one for the property it covers whatever the data looks like. Statistics could only make that choice worse, and did in 5.3.1. Connect, Disconnect, Dispose, CreateIndex and Optimize no longer run ANALYZE or PRAGMA optimize, and the user_version stamp and the staleness check go with them. autoOptimize and Optimize now cover only the legacy index drops.
A store written by 5.3.0 or 5.3.1 holds the sampled sqlite_stat1 rows that make the planner read every row of a type on each indexed lookup. The connect script now drops the statistics tables and reloads, so every connect, with or without autoOptimize, leaves the planner on its default heuristics. The reload (ANALYZE sqlite_master) leaves an empty sqlite_stat1 behind.
The write-up, CLI repro, README and CHANGELOG now describe the fix as shipped: no ANALYZE anywhere, leftovers dropped at connect, autoOptimize and Optimize covering only the legacy index drops. The repro script shows the plan for every query shape with no statistics; the row-estimate repro for PRAGMA optimize goes, since the pragma is no longer used.
…eral index row Removing statistics altogether fixed the sampled-statistics regression but left the planner unable to rank two partial indexes on one type: with a filter on both properties it took whichever index was created last, 4.6 ms (57 ms on SQLCipher) against 0.02 ms for a one-row result at 40,000 rows. Statistics are now gathered in full, for JsonValue only, when CreateIndex builds an index or re-declares one that was built on an empty type and has rows now. Rows recorded for an empty type are not kept. The general index's sqlite_stat1 row is pinned to "N N/2 N/2": the per-type average ANALYZE writes there misleads the planner whenever one type dominates a store, into a table scan for every type when few types existed at ANALYZE time and into the general index over a big type's partial index otherwise. Half the table keeps the general index cheaper than a scan and dearer than any partial index whose match is under half the rows, so the partial indexes' accurate rows rank against each other. Upgraded stores are repaired once, gated on PRAGMA user_version: with autoOptimize on, the first connect replaces sampled rows with full ones (or discards them when the store has no expression index) and stamps the store; a stamped store pays one header read at connect and writes nothing, so schema_version no longer changes on every connect. With autoOptimize off, Optimize()/OptimizeAsync() perform the repair. PRAGMA optimize stays removed. Tests cover every query shape on both engines, both index creation orders, the empty-type-then-bulk-load case, the repair with autoOptimize on and off, a store without indexes, a fresh store, and that a current store is not written at connect. A PlannerStatistics benchmark measures the three shapes whose speed depends on statistics through the public API across launches. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the sampled planner statistics that made indexed lookups on a large type read every row of the type (5.0.1–5.3.1), adds
autoOptimize/Optimize()/OptimizeAsync()for deferring one-time upgrade maintenance, and keeps the planner able to rank a type's partial indexes against each other.The bug
PRAGMA optimizeon connect/disconnect (since 5.0.1) and theANALYZEafterCreateIndex(since 5.1.1) setanalysis_limit = 400. The general(FullTypeName, Partition)index is ordered by type name and the types that sort first are tiny, so its sample credited every type with a handful of rows (81on a 1.68M-row production store whose largest type held 586,986), while a partial index's sample sat inside one key and credited every key with hundreds (401). The planner preferred the general index and read the whole type on each indexed lookup: 82–366 ms per lookup on that store, under a millisecond with correct statistics. Results were never wrong; only speed.The design
The branch's first design kept full statistics with
PRAGMA optimize, and its second removed statistics entirely. Review found that the second one regresses two indexed filters ANDed — with nosqlite_stat1the planner takes whichever partial index was created last — and that both suffer from the general index's own statistics row. What ships:CreateIndexrunsPRAGMA analysis_limit = 0; ANALYZE JsonValue;after it builds an index, and again when it re-declares an index that was built on an empty type and has rows now. Rows recorded for an empty type are not kept. Nothing samples;PRAGMA optimizeis gone from connect, disconnect and dispose.N N/2 N/2.ANALYZEwrites it as the average rows per type, which is wrong for every type when one type dominates a store: it sends small types to a table scan when few types existed atANALYZEtime, and sends a big type's indexed lookups to the general index otherwise. Half the table keeps the general index cheaper than a scan and dearer than any partial index whose match is under half the rows, so the partial indexes' accurate rows decide between themselves. SQLite documents writingsqlite_stat1for this ("Manual Control Of Query Plans", lang_analyze.html).PRAGMA user_version. WithautoOptimizeon (default) the first connect replaces sampled rows with full ones — or discards them when the store has no expression index — and stamps the store; a stamped store pays one header read at connect and writes nothing (schema_versionis unchanged across reconnects; a test pins it). WithautoOptimize: falsethe store keeps its previous plans untilOptimize()/OptimizeAsync(), which also drop the four redundant 4.x indexes.Full detail, including the plan table for every query shape on both engines and why the other two designs were rejected:
docs/planner-statistics-sampling.md.Benchmark
dotnet run --project TychoDB.Benchmarks -c Release -- --filter '*PlannerStatistics*'(new class; 40 small types written first, one type with 40,000 rows over 200 keys, indexed on a uniqueSeqand a 200-valueGroupId; built, closed and reopened across three launches like an app). Apple M4 Max, .NET 9.0.17, SQLite 3.50.3, medians of 16 × 1 launch.main(v5.3.1, sampled)322be81(no statistics)The first row is the bug (18x). The second is the no-statistics regression this PR avoids; it grows with the rows matching the less selective property (4.6x here at 200 per key, 200x at 10,000 per key, and larger again on SQLCipher, where every extra page is decrypted).
Public API
Tycho(..., bool autoOptimize = true)— source-compatible; assemblies that constructTychoneed recompiling.void Optimize(),ValueTask<bool> OptimizeAsync(CancellationToken = default).Minor release:
5.4.0.Directory.build.propsupdated to match (the tag still sets the package version).Tests
379 passed, 0 failed, 4 skipped in both
Debug(SQLite 3.50.3) andEncrypted(SQLCipher, SQLite 3.39.2).PlannerStatisticsTestsasserts the plan for every shape TychoDB emits (equality, non-selective equality, range, sort with limit, two indexed equalities in both creation orders, unindexed filter, whole type large and tiny, count), the two 5.3.1 failure scenarios, repair withautoOptimizeon and off, a store without indexes, a fresh store, and that a current store is not written at connect.🤖 Generated with Claude Code