20 KiB
Milestone 36: Overview performance — combined endpoint, projections, cache
Replace the five per-panel stats endpoints with one GET /api/overview served
from materialized projections in querylog.db plus an in-memory response
cache, so Overview cost stops growing with query-log size.
Motivation (measured)
Today each Overview load runs five separate scans of every raw row in the
window, serialized on WebState.querylog_lock, re-polled every 30 s. Measured
x86 ReleaseSafe (bench at scratchpad statsbench2/, production-like skew;
Pi ≈ 3–3.5× slower):
| Rows | 30d, five scans (today) | 30d, one combined scan | 30d, projections |
|---|---|---|---|
| 1M | ~2.2 s | 264 ms | 41 ms |
| 3M | ~6.6 s | 793 ms | 38 ms |
| 5M | ~11.8 s | 1,346 ms | 40 ms |
Projection maintenance costs +10% per 100-row insert batch, and ~1.5 MB of disk in the bench — a size bounded by retained buckets × distinct client/type/route keys, independent of raw query volume. Production is on a ~100k rows/day growth curve (≈3M rows at 30-day retention), so the projection path is the design target, not a contingency. Both computations were cross-checked for identical output in the bench.
Design ruling (owner + Codex consultation, 2026-08-27): stay on SQLite;
projections live in the same file as the raw rows and are updated in the same
transaction, so SQLite's transaction is the coherence mechanism — no second
file, no epoch protocol. The DDL change re-fingerprints querylog.db; the
existing rename-aside path handles old files (one-time history reset,
disclosed in the changelog). No backfill migration.
Sessions
Four sessions. A first. B and C after A, in parallel (disjoint files). D after B.
- A: storage — projection schema, writer maintenance, retention, new read path.
A does NOT delete the five existing aggregate functions —
stats.zigstill calls them until B lands, and A must leavezig build testgreen. - B: web —
/api/overviewhandler, response cache, removal of the five old endpoints, OpenAPI/contract regeneration. - C: admin — one overview query, types, component/data plumbing, tests.
- D: storage cleanup — delete the five now-unreferenced aggregate functions.
Session A: storage
A.1 Schema (src/storage/querylog_schema.zig)
Append four projection tables to ddl. Grain: 30-minute buckets, bucket =
floor-to-grid of the row timestamp: @divFloor(timestamp, 1800) * 1800 in Zig
and the equivalent floor semantics in any SQL (SQLite integer / truncates
toward zero, which differs on negative timestamps — use floor everywhere, as
window() does). 1800 divides every serving
width ≥ 30 min (1800, 3600, 21600), which is what makes one grain serve the
24h, 7d and 30d windows exactly. The 1h window (60 s buckets) is NOT served
from projections (A.4).
CREATE TABLE bucket_totals (
bucket INTEGER PRIMARY KEY,
queries INTEGER NOT NULL,
blocked INTEGER NOT NULL,
cached INTEGER NOT NULL,
rt_sum INTEGER NOT NULL, -- sum(response_time_us) over timed rows
rt_count INTEGER NOT NULL -- count(response_time_us)
) WITHOUT ROWID;
CREATE TABLE bucket_clients (
bucket INTEGER NOT NULL,
client_ip TEXT NOT NULL,
queries INTEGER NOT NULL,
PRIMARY KEY (bucket, client_ip)
) WITHOUT ROWID;
CREATE TABLE bucket_types (
bucket INTEGER NOT NULL,
qtype INTEGER NOT NULL, -- -1 encodes a NULL qtype, losslessly
count INTEGER NOT NULL,
PRIMARY KEY (bucket, qtype)
) WITHOUT ROWID;
CREATE TABLE bucket_routes (
bucket INTEGER NOT NULL,
route_kind TEXT NOT NULL,
source_present INTEGER NOT NULL, -- 0: source NULL; 1: source = source_text
source_text TEXT NOT NULL, -- '' when source_present = 0
count INTEGER NOT NULL,
PRIMARY KEY (bucket, route_kind, source_present, source_text),
CHECK (source_present IN (0, 1)),
CHECK (source_present = 1 OR source_text = '')
) WITHOUT ROWID;
Column semantics match the existing aggregates exactly: blocked counts
blocked <> 0; cached counts cache_hit = 1; routes' source is the
existing CASE (upstream rows → upstream, forward_zone rows →
forward_zone, else NULL). The fingerprint moves automatically; do not touch
the fingerprint machinery.
A.2 Writer maintenance (src/storage/repositories/queries_repo.zig)
BatchWriter.writeBatch updates all four projections inside the same
transaction that inserts the raw rows:
- Aggregate the batch in Zig first, producing per-key deltas; then one UPSERT
per touched key:
INSERT ... ON CONFLICT(...) DO UPDATE SET queries = queries + excluded.queries, .... The aggregation must accept any slice length —writeBatch's API does not enforce the logger's 100-row batching, so no fixed-size arrays sized to it.BatchWritercurrently owns no allocator:initgains one, owned for the writer's life, used only for the per-batch delta maps; scratch is freed (or a retained map cleared) at the end of everywriteBatch, anderror.OutOfMemoryfails the batch before the transaction opens — no hidden global allocator, no implicit size cap, no quadratic rescanning. - No per-row SQL, no triggers.
- Failure contract: any failed projection statement rolls the whole
transaction back — raw rows and projections together — resets every
projection statement, and leaves the writer usable for the next batch
(
resetAlldiscipline as for the raw statements today). Fault-injection acceptance: a batch whose projection update fails leaves the database unchanged, and the next batch succeeds. - The bench measured this at 0.39 ms vs 0.34 ms per batch — acceptance is correctness, not speed.
A.3 Retention (src/storage/repositories/queries_repo.zig, prune path)
In the same transaction as pruneOlderThan(cutoff)'s raw delete:
- Delete projection rows with
bucket < floor(cutoff / 1800) * 1800from all four tables. - If
cutoffis not on a bucket boundary, recompute the straddling bucket (floor(cutoff/1800)*1800) from the remaining raw rows and replace its projection rows in all four tables. Never approximate.
Failure atomicity: a failure during the projection delete or the straddling-bucket replacement rolls back the raw delete, the watermark advance and every projection change together — one transaction, tested by fault injection.
Implementation note (Session A, recorded post-build): the bucket_totals
recompute carries HAVING count(*) > 0 — a bare SQL aggregate always yields
one row, and an emptied straddling bucket must disappear, not persist as
zeros.
A.4 Read path (src/storage/repositories/queries_repo.zig)
One function producing the whole Overview payload for a window, from one already-open read transaction (the caller owns transaction + lock, as today):
pub const Overview = struct {
totals: StatsTotals,
buckets: []const Bucket, // bucket_count entries, zero-filled
clients: ClientsBreakdown, // top-8 + other, as today
types: []const TypeCount, // sorted as stats_types_sql sorts
routes: []const RouteCount, // sorted as stats_routes_sql sorts
};
pub fn overview(
database: *db.Db,
arena: Allocator,
since: i64,
bucket_seconds: u32,
bucket_count: u32,
) db.Error!Overview
Storage owns these scalars — no import of any web module. until is derived
as since + bucket_seconds * bucket_count with the same overflow checks as
timeseries. Preconditions, checked before path selection and tested:
bucket_seconds != 0 and bucket_count != 0 (else error.Misuse, matching
the existing clients contract); on the projection path
(bucket_seconds >= 1800) additionally since a multiple of 1800 and
bucket_seconds % 1800 == 0, else error.Misuse. The handler's window()
guarantees all of them.
Two implementations behind one entry point, chosen by bucket_seconds:
bucket_seconds >= 1800(24h, 7d, 30d): read the four projection tables over[since, until), aggregating 30-min rows up to the serving width in Zig.distinct_clientscomes from groupingbucket_clientsbyclient_ipover the window — never from summing per-bucket counts.avg_response_time_us=sum(rt_sum) / sum(rt_count), null whenrt_countsums to 0. Top-8 clients ranked by window total desc, ties byclient_ipasc (BINARY), residual summed intoother— identical cut semantics tostatsClients.bucket_seconds < 1800(1h): one single pass over the raw rows in the window (one SELECT of the needed columns, stepped once), aggregating everything in Zig. Memory bound: O(distinct clients + distinct qtypes + distinct routes) in the window — explicitly permitted; this is a household LAN and the same bound the arena-returning aggregates already carry. This replaces today's five scans and the clients rank+bucket double scan. Same output contracts.
Sort orders and tie-breaks must reproduce the existing SQL orderings exactly
(types: count desc, null last within tie, qtype asc; routes: count desc,
route_kind asc, null source last, source asc) — the goldens' byte-stability
argument carries over. Note: RouteKind's enum declaration order is not
alphabetical; "route_kind asc" means the stored text's byte order, so any Zig
comparator orders by @tagName bytes, never by enum ordinal (Session A's
accumulator already does; mutation-tested).
A.5 Acceptance criteria
zig build testgreen.- Property test: after an arbitrary interleaving of batches and prunes
(including a prune cutoff off the bucket grid), every projection table
equals a from-scratch recomputation from
query_log. - Equivalence test:
overview()output (both paths) equals a test-only oracle over the same window on the same data — including empty windows, NULL qtype, NULL source on anupstreamrow, ties in ranking, and a window whose last bucket is in progress. The oracle is a copy of the five existing SQL aggregates living in the test file, so it survives Session D's deletion of the production functions. - Fingerprint test updated (table/index count assertions in querylog_schema tests).
Session B: web
B.1 Endpoint (src/web/handlers/overview.zig, replacing stats.zig's five)
GET /api/overview?period=1h|24h|30d|7d (same grammar, default 24h, same 400
text). One read transaction under WebState.querylog_lock covering the
aggregate and coverage.read — one snapshot, no cross-panel skew. Response:
{
"period": "24h", "since": ..., "until": ..., "bucket_seconds": 1800,
"totals": { "queries": n, "blocked": n, "clients": n, "avg_response_time_us": n|null },
"buckets": [ { "ts": ..., "queries": n, "blocked": n, "cached": n }, ... ],
"clients": [ { "client": "ip", "buckets": [n, ...] }, ... ],
"other": [n, ...],
"types": [ { "qtype": n|null, "count": n }, ... ],
"routes": [ { "route": "...", "source": "..."|null, "count": n }, ... ],
"coverage": { ... }
}
Field shapes and semantics are exactly today's five bodies merged; Period,
window(), max_buckets move to (or stay importable from) the new handler.
503 when the query log is unavailable; 500 logging unchanged. Route metadata
identical to the removed endpoints: same authentication (.session), same
rate-limit class (.counted), same authority policy (.read).
Remove GET /api/stats, /api/stats/timeseries, /api/stats/types,
/api/stats/routes, /api/stats/clients and their routes.
Contract surface (B owns all of it): add the new path and schema to the
OpenAPI document, remove the five old operations and their schemas, update
every drift guard that lists them, add a contract sample for
/api/overview, and regenerate admin/src/lib/contractSamples.gen.ts
(reserved for B — Session C must not touch it). Update the API listings in
docs/ and PLAN.md that name the five endpoints or the querylog layout.
B.2 Response cache (src/web/server.zig WebState + overview.zig)
Per-period cached response body, invalidated by data change or window roll.
Key: (period, window.until, data_version) where data_version is PRAGMA data_version on the web task's connection (it changes when any other
connection — logger, retention — commits).
The entire cache decision happens under querylog_lock; nothing touches the
shared connection or the slots outside it. Exact sequence per request:
- Acquire
querylog_lock— ONCE.server.QuerylogRead.openacquires this lock itself, so the overview handler must not call it after step 1: B refactors the scope into a lock-owning wrapper plus a locked-caller variant (for exampleQuerylogRead.openLocked, documented as requiring the lock), and the overview path uses the locked-caller variant for step 4. A literal "lock, then QuerylogRead.open" deadlocks. - Sample
PRAGMA data_version(inside the lock — the shared connection may otherwise have a foreign transaction open, and the slots need the mutual exclusion anyway). - Hit (
slot.period == period and slot.until == window.until and slot.data_version == sampled): copy the stored bytes into the request arena, release the lock, respond. The copy is what makes a concurrent rebuild's free-and-replace safe. - Miss: open the read transaction, build the body, commit. Publish to the
slot ONLY after a successful commit, keyed by the version sampled in
step 2 (a commit landing during the build bumps
data_version, so the next request rebuilds — stale-under-new-key is impossible). A failed commit or build publishes nothing and responds 500 as today. - Copy to the request arena, release the lock, write the socket. The lock never spans a socket write (existing discipline).
Because the check happens only under the lock, querylog_lock is the
single-flight: a second request for the same key waits and then hits.
Storage: one slot per period (4 slots) in WebState; body bytes allocated
from WebState.gpa, replaced on rebuild (free old, install new), freed in
deinit. No capacity limit beyond the allocator — a body is bounded by the
fixed bucket counts plus the household client/type/route cardinality.
No adaptive polling and no combined-endpoint staging: with projections + this cache a rebuild is ~40 ms x86 / ~0.13 s Pi, so the admin's existing 30 s cadence is fine.
B.3 Acceptance criteria
zig build testgreen; handler tests ported from stats.zig (period grammar, window math, one-snapshot behavior) plus: cache hit returns byte-identical body; a logger commit (data_version bump) invalidates; a window roll invalidates; a retention prune committed through another connection invalidates (both the aggregates and the cachedcoverage.available_sinceare replaced); a failed read-transaction commit neither installs nor replaces a cache entry.curl /api/overview?period=30don a seeded scratch instance returns all panels consistent (breakdowns sum to totals on a quiet database).- The five old routes return 404.
Session C: admin
C.1 Data layer
admin/src/lib/types.ts: oneOverviewtype mirroring B.1; remove the five per-panel response types.admin/src/lib/api.ts:getOverview(period); remove the five getters.admin/src/lib/queries.ts:overviewQuery(period)withrefetchInterval: 30_000and key["overview", period]; remove the five stats query factories and their keys.
C.2 Overview page
admin/src/features/overview/overviewWindow.ts and the chart components
consume the single query: one useQuery where five ran in parallel. Loading,
error and coverage handling collapse to one page-level surface (one spinner
state, one error state for the whole Overview); the per-panel shells,
layout, copy, chart dimensions and accessibility attributes stay exactly as
they are. Chart components (TimeseriesChart, ClientChart, Donut) keep their
props — adapt the mapping layer, not the charts. C must not touch
contractSamples.gen.ts (B owns its regeneration).
C.3 Acceptance criteria
npm testgreen in admin/ (mock the one endpoint; port the five-query tests).npm run typecheckgreen — the build alone does not run tsc. B and C are file-disjoint but type-coupled throughcontractSamples.gen.ts: C's typecheck/test/build gates run (or re-run) AFTER B has regenerated that file. Implementation may proceed in parallel; the green gate is sequenced.npm run buildgreen; bundle-size assertion still passes.- Manual: scratch instance renders all Overview panels from the new endpoint on all four periods.
Session D: storage cleanup
After B is merged and green: delete statsTotals, timeseries, statsTypes,
statsRoutes, statsClients and their SQL constants from
queries_repo.zig — nothing references them once stats.zig is gone. The
test-only oracle from A.5 stays. Acceptance: zig build test green, no dead
stats SQL remains in production code.
Module layout
src/web/handlers/overview.zig— new; replacessrc/web/handlers/stats.zig(deleted by B).src/storage/querylog_schema.zig— projection DDL appended.src/storage/repositories/queries_repo.zig— writer maintenance, retention integration,overview()read path (A); five old aggregates deleted (D).- Admin files per C.1/C.2.
File ownership
- A:
src/storage/*(old aggregates left in place), plus the mechanical allocator-plumbing at everyBatchWriter.initcall site outside storage (src/web/handlers/*,src/web/web_integration_test.zig, any test using the writer) — A runs before B, so this is sequential, not shared, ownership; A's gate is the fullzig build test. - B:
src/web/*, plus:src/app.zig(cache cleanup at the composition root),src/tests.zig(handler import swap), the OpenAPI document and its drift guards, contract samples includingadmin/src/lib/contractSamples.gen.ts,docs/**API listings,PLAN.mdstale sections, and the Overview/API sections ofspecs/ui-redesign.md(which still mandates the five endpoints and must be amended, not obeyed). - C:
admin/*EXCEPTadmin/src/lib/contractSamples.gen.ts. - D:
src/storage/repositories/queries_repo.zig(sequential, after B). - Orchestrator:
CHANGELOG.md(hand-written, per release process). B and C run in parallel; the one shared-tree exception above is reserved to B.
Acceptance criteria (milestone complete)
zig build testandzig build test -Dintegrationgreen (the integration suite carries the live route walk, contract-sample comparison, concurrent querylog reads and OpenAPI guards); adminnpm test/typecheck/buildgreen.- Scratch-instance smoke: seeded data + live digs; Overview correct on all periods; old endpoints gone.
- Changelog discloses: schema change resets query history (rename-aside),
five endpoints replaced by
/api/overview. - Release gate (
zig build cutfingerprint check) satisfied.
Anti-requirements
- No second database file, no epoch/validity protocol, no ATTACH.
- No backfill migration, no rebuild command, no catch-up cursor — projections are born with the file and maintained transactionally; that is the whole coherence story.
- No connection pool, no adaptive polling, no DuckDB.
- No new indexes on
query_log, no triggers, no per-row projection SQL. - Do not change the 1h/24h/7d/30d period grammar, bucket widths or counts.
- No HTTP-level caching of any kind: no ETag, no
Cache-Control, no stale-while-revalidate, no background refresh, and never cache a 500/503 body. The cache is exactly the in-process design of B.2. - No visual redesign, no chart-prop changes, no cache configuration knobs, no cache metrics. This milestone changes data acquisition and storage only.