# Milestone 36: Overview performance — combined endpoint, projections, cache Replace the five per-panel stats endpoints with one `GET /api/overview` served from materialized projections in `querylog.db` plus an in-memory response cache, so Overview cost stops growing with query-log size. ## Motivation (measured) Today each Overview load runs five separate scans of every raw row in the window, serialized on `WebState.querylog_lock`, re-polled every 30 s. Measured x86 ReleaseSafe (bench at scratchpad `statsbench2/`, production-like skew; Pi ≈ 3–3.5× slower): | Rows | 30d, five scans (today) | 30d, one combined scan | 30d, projections | |-----:|------------------------:|-----------------------:|-----------------:| | 1M | ~2.2 s | 264 ms | 41 ms | | 3M | ~6.6 s | 793 ms | 38 ms | | 5M | ~11.8 s | 1,346 ms | 40 ms | Projection maintenance costs +10% per 100-row insert batch, and ~1.5 MB of disk in the bench — a size bounded by retained buckets × distinct client/type/route keys, independent of raw query volume. Production is on a ~100k rows/day growth curve (≈3M rows at 30-day retention), so the projection path is the design target, not a contingency. Both computations were cross-checked for identical output in the bench. Design ruling (owner + Codex consultation, 2026-08-27): stay on SQLite; projections live in the same file as the raw rows and are updated in the same transaction, so SQLite's transaction is the coherence mechanism — no second file, no epoch protocol. The DDL change re-fingerprints `querylog.db`; the existing rename-aside path handles old files (one-time history reset, disclosed in the changelog). No backfill migration. ## Sessions Four sessions. A first. B and C after A, in parallel (disjoint files). D after B. - A: storage — projection schema, writer maintenance, retention, new read path. A does NOT delete the five existing aggregate functions — `stats.zig` still calls them until B lands, and A must leave `zig build test` green. - B: web — `/api/overview` handler, response cache, removal of the five old endpoints, OpenAPI/contract regeneration. - C: admin — one overview query, types, component/data plumbing, tests. - D: storage cleanup — delete the five now-unreferenced aggregate functions. --- ## Session A: storage ### A.1 Schema (src/storage/querylog_schema.zig) Append four projection tables to `ddl`. Grain: 30-minute buckets, `bucket` = floor-to-grid of the row timestamp: `@divFloor(timestamp, 1800) * 1800` in Zig and the equivalent floor semantics in any SQL (SQLite integer `/` truncates toward zero, which differs on negative timestamps — use floor everywhere, as `window()` does). 1800 divides every serving width ≥ 30 min (1800, 3600, 21600), which is what makes one grain serve the 24h, 7d and 30d windows exactly. The 1h window (60 s buckets) is NOT served from projections (A.4). ```sql CREATE TABLE bucket_totals ( bucket INTEGER PRIMARY KEY, queries INTEGER NOT NULL, blocked INTEGER NOT NULL, cached INTEGER NOT NULL, rt_sum INTEGER NOT NULL, -- sum(response_time_us) over timed rows rt_count INTEGER NOT NULL -- count(response_time_us) ) WITHOUT ROWID; CREATE TABLE bucket_clients ( bucket INTEGER NOT NULL, client_ip TEXT NOT NULL, queries INTEGER NOT NULL, PRIMARY KEY (bucket, client_ip) ) WITHOUT ROWID; CREATE TABLE bucket_types ( bucket INTEGER NOT NULL, qtype INTEGER NOT NULL, -- -1 encodes a NULL qtype, losslessly count INTEGER NOT NULL, PRIMARY KEY (bucket, qtype) ) WITHOUT ROWID; CREATE TABLE bucket_routes ( bucket INTEGER NOT NULL, route_kind TEXT NOT NULL, source_present INTEGER NOT NULL, -- 0: source NULL; 1: source = source_text source_text TEXT NOT NULL, -- '' when source_present = 0 count INTEGER NOT NULL, PRIMARY KEY (bucket, route_kind, source_present, source_text), CHECK (source_present IN (0, 1)), CHECK (source_present = 1 OR source_text = '') ) WITHOUT ROWID; ``` Column semantics match the existing aggregates exactly: `blocked` counts `blocked <> 0`; `cached` counts `cache_hit = 1`; routes' `source` is the existing CASE (`upstream` rows → `upstream`, `forward_zone` rows → `forward_zone`, else NULL). The fingerprint moves automatically; do not touch the fingerprint machinery. ### A.2 Writer maintenance (src/storage/repositories/queries_repo.zig) `BatchWriter.writeBatch` updates all four projections inside the same transaction that inserts the raw rows: - Aggregate the batch in Zig first, producing per-key deltas; then one UPSERT per touched key: `INSERT ... ON CONFLICT(...) DO UPDATE SET queries = queries + excluded.queries, ...`. The aggregation must accept any slice length — `writeBatch`'s API does not enforce the logger's 100-row batching, so no fixed-size arrays sized to it. `BatchWriter` currently owns no allocator: `init` gains one, owned for the writer's life, used only for the per-batch delta maps; scratch is freed (or a retained map cleared) at the end of every `writeBatch`, and `error.OutOfMemory` fails the batch before the transaction opens — no hidden global allocator, no implicit size cap, no quadratic rescanning. - No per-row SQL, no triggers. - Failure contract: any failed projection statement rolls the whole transaction back — raw rows and projections together — resets every projection statement, and leaves the writer usable for the next batch (`resetAll` discipline as for the raw statements today). Fault-injection acceptance: a batch whose projection update fails leaves the database unchanged, and the next batch succeeds. - The bench measured this at 0.39 ms vs 0.34 ms per batch — acceptance is correctness, not speed. ### A.3 Retention (src/storage/repositories/queries_repo.zig, prune path) In the same transaction as `pruneOlderThan(cutoff)`'s raw delete: 1. Delete projection rows with `bucket < floor(cutoff / 1800) * 1800` from all four tables. 2. If `cutoff` is not on a bucket boundary, recompute the straddling bucket (`floor(cutoff/1800)*1800`) from the remaining raw rows and replace its projection rows in all four tables. Never approximate. Failure atomicity: a failure during the projection delete or the straddling-bucket replacement rolls back the raw delete, the watermark advance and every projection change together — one transaction, tested by fault injection. Implementation note (Session A, recorded post-build): the bucket_totals recompute carries `HAVING count(*) > 0` — a bare SQL aggregate always yields one row, and an emptied straddling bucket must disappear, not persist as zeros. ### A.4 Read path (src/storage/repositories/queries_repo.zig) One function producing the whole Overview payload for a window, from one already-open read transaction (the caller owns transaction + lock, as today): ```zig pub const Overview = struct { totals: StatsTotals, buckets: []const Bucket, // bucket_count entries, zero-filled clients: ClientsBreakdown, // top-8 + other, as today types: []const TypeCount, // sorted as stats_types_sql sorts routes: []const RouteCount, // sorted as stats_routes_sql sorts }; pub fn overview( database: *db.Db, arena: Allocator, since: i64, bucket_seconds: u32, bucket_count: u32, ) db.Error!Overview ``` Storage owns these scalars — no import of any web module. `until` is derived as `since + bucket_seconds * bucket_count` with the same overflow checks as `timeseries`. Preconditions, checked before path selection and tested: `bucket_seconds != 0` and `bucket_count != 0` (else `error.Misuse`, matching the existing clients contract); on the projection path (`bucket_seconds >= 1800`) additionally `since` a multiple of 1800 and `bucket_seconds % 1800 == 0`, else `error.Misuse`. The handler's `window()` guarantees all of them. Two implementations behind one entry point, chosen by `bucket_seconds`: - `bucket_seconds >= 1800` (24h, 7d, 30d): read the four projection tables over `[since, until)`, aggregating 30-min rows up to the serving width in Zig. `distinct_clients` comes from grouping `bucket_clients` by `client_ip` over the window — never from summing per-bucket counts. `avg_response_time_us` = `sum(rt_sum) / sum(rt_count)`, null when `rt_count` sums to 0. Top-8 clients ranked by window total desc, ties by `client_ip` asc (BINARY), residual summed into `other` — identical cut semantics to `statsClients`. - `bucket_seconds < 1800` (1h): one single pass over the raw rows in the window (one SELECT of the needed columns, stepped once), aggregating everything in Zig. Memory bound: O(distinct clients + distinct qtypes + distinct routes) in the window — explicitly permitted; this is a household LAN and the same bound the arena-returning aggregates already carry. This replaces today's five scans and the clients rank+bucket double scan. Same output contracts. Sort orders and tie-breaks must reproduce the existing SQL orderings exactly (types: count desc, null last within tie, qtype asc; routes: count desc, route_kind asc, null source last, source asc) — the goldens' byte-stability argument carries over. Note: `RouteKind`'s enum declaration order is not alphabetical; "route_kind asc" means the stored text's byte order, so any Zig comparator orders by `@tagName` bytes, never by enum ordinal (Session A's accumulator already does; mutation-tested). ### A.5 Acceptance criteria - [ ] `zig build test` green. - [ ] Property test: after an arbitrary interleaving of batches and prunes (including a prune cutoff off the bucket grid), every projection table equals a from-scratch recomputation from `query_log`. - [ ] Equivalence test: `overview()` output (both paths) equals a test-only oracle over the same window on the same data — including empty windows, NULL qtype, NULL source on an `upstream` row, ties in ranking, and a window whose last bucket is in progress. The oracle is a copy of the five existing SQL aggregates living in the test file, so it survives Session D's deletion of the production functions. - [ ] Fingerprint test updated (table/index count assertions in querylog_schema tests). --- ## Session B: web ### B.1 Endpoint (src/web/handlers/overview.zig, replacing stats.zig's five) `GET /api/overview?period=1h|24h|30d|7d` (same grammar, default 24h, same 400 text). One read transaction under `WebState.querylog_lock` covering the aggregate and `coverage.read` — one snapshot, no cross-panel skew. Response: ```json { "period": "24h", "since": ..., "until": ..., "bucket_seconds": 1800, "totals": { "queries": n, "blocked": n, "clients": n, "avg_response_time_us": n|null }, "buckets": [ { "ts": ..., "queries": n, "blocked": n, "cached": n }, ... ], "clients": [ { "client": "ip", "buckets": [n, ...] }, ... ], "other": [n, ...], "types": [ { "qtype": n|null, "count": n }, ... ], "routes": [ { "route": "...", "source": "..."|null, "count": n }, ... ], "coverage": { ... } } ``` Field shapes and semantics are exactly today's five bodies merged; `Period`, `window()`, `max_buckets` move to (or stay importable from) the new handler. 503 when the query log is unavailable; 500 logging unchanged. Route metadata identical to the removed endpoints: same authentication (`.session`), same rate-limit class (`.counted`), same authority policy (`.read`). Remove `GET /api/stats`, `/api/stats/timeseries`, `/api/stats/types`, `/api/stats/routes`, `/api/stats/clients` and their routes. Contract surface (B owns all of it): add the new path and schema to the OpenAPI document, remove the five old operations and their schemas, update every drift guard that lists them, add a contract sample for `/api/overview`, and regenerate `admin/src/lib/contractSamples.gen.ts` (reserved for B — Session C must not touch it). Update the API listings in `docs/` and `PLAN.md` that name the five endpoints or the querylog layout. ### B.2 Response cache (src/web/server.zig WebState + overview.zig) Per-period cached response body, invalidated by data change or window roll. Key: `(period, window.until, data_version)` where `data_version` is `PRAGMA data_version` on the web task's connection (it changes when any other connection — logger, retention — commits). The entire cache decision happens under `querylog_lock`; nothing touches the shared connection or the slots outside it. Exact sequence per request: 1. Acquire `querylog_lock` — ONCE. `server.QuerylogRead.open` acquires this lock itself, so the overview handler must not call it after step 1: B refactors the scope into a lock-owning wrapper plus a locked-caller variant (for example `QuerylogRead.openLocked`, documented as requiring the lock), and the overview path uses the locked-caller variant for step 4. A literal "lock, then QuerylogRead.open" deadlocks. 2. Sample `PRAGMA data_version` (inside the lock — the shared connection may otherwise have a foreign transaction open, and the slots need the mutual exclusion anyway). 3. Hit (`slot.period == period and slot.until == window.until and slot.data_version == sampled`): copy the stored bytes into the request arena, release the lock, respond. The copy is what makes a concurrent rebuild's free-and-replace safe. 4. Miss: open the read transaction, build the body, commit. Publish to the slot ONLY after a successful commit, keyed by the version sampled in step 2 (a commit landing during the build bumps `data_version`, so the next request rebuilds — stale-under-new-key is impossible). A failed commit or build publishes nothing and responds 500 as today. 5. Copy to the request arena, release the lock, write the socket. The lock never spans a socket write (existing discipline). Because the check happens only under the lock, `querylog_lock` is the single-flight: a second request for the same key waits and then hits. Storage: one slot per period (4 slots) in `WebState`; body bytes allocated from `WebState.gpa`, replaced on rebuild (free old, install new), freed in `deinit`. No capacity limit beyond the allocator — a body is bounded by the fixed bucket counts plus the household client/type/route cardinality. No adaptive polling and no combined-endpoint staging: with projections + this cache a rebuild is ~40 ms x86 / ~0.13 s Pi, so the admin's existing 30 s cadence is fine. ### B.3 Acceptance criteria - [ ] `zig build test` green; handler tests ported from stats.zig (period grammar, window math, one-snapshot behavior) plus: cache hit returns byte-identical body; a logger commit (data_version bump) invalidates; a window roll invalidates; a retention prune committed through another connection invalidates (both the aggregates and the cached `coverage.available_since` are replaced); a failed read-transaction commit neither installs nor replaces a cache entry. - [ ] `curl /api/overview?period=30d` on a seeded scratch instance returns all panels consistent (breakdowns sum to totals on a quiet database). - [ ] The five old routes return 404. --- ## Session C: admin ### C.1 Data layer - `admin/src/lib/types.ts`: one `Overview` type mirroring B.1; remove the five per-panel response types. - `admin/src/lib/api.ts`: `getOverview(period)`; remove the five getters. - `admin/src/lib/queries.ts`: `overviewQuery(period)` with `refetchInterval: 30_000` and key `["overview", period]`; remove the five stats query factories and their keys. ### C.2 Overview page `admin/src/features/overview/overviewWindow.ts` and the chart components consume the single query: one `useQuery` where five ran in parallel. Loading, error and coverage handling collapse to one page-level surface (one spinner state, one error state for the whole Overview); the per-panel shells, layout, copy, chart dimensions and accessibility attributes stay exactly as they are. Chart components (TimeseriesChart, ClientChart, Donut) keep their props — adapt the mapping layer, not the charts. C must not touch `contractSamples.gen.ts` (B owns its regeneration). ### C.3 Acceptance criteria - [ ] `npm test` green in admin/ (mock the one endpoint; port the five-query tests). - [ ] `npm run typecheck` green — the build alone does not run tsc. B and C are file-disjoint but type-coupled through `contractSamples.gen.ts`: C's typecheck/test/build gates run (or re-run) AFTER B has regenerated that file. Implementation may proceed in parallel; the green gate is sequenced. - [ ] `npm run build` green; bundle-size assertion still passes. - [ ] Manual: scratch instance renders all Overview panels from the new endpoint on all four periods. --- ## Session D: storage cleanup After B is merged and green: delete `statsTotals`, `timeseries`, `statsTypes`, `statsRoutes`, `statsClients` and their SQL constants from `queries_repo.zig` — nothing references them once `stats.zig` is gone. The test-only oracle from A.5 stays. Acceptance: `zig build test` green, no dead stats SQL remains in production code. --- ## Module layout - `src/web/handlers/overview.zig` — new; replaces `src/web/handlers/stats.zig` (deleted by B). - `src/storage/querylog_schema.zig` — projection DDL appended. - `src/storage/repositories/queries_repo.zig` — writer maintenance, retention integration, `overview()` read path (A); five old aggregates deleted (D). - Admin files per C.1/C.2. ## File ownership - A: `src/storage/*` (old aggregates left in place), plus the mechanical allocator-plumbing at every `BatchWriter.init` call site outside storage (`src/web/handlers/*`, `src/web/web_integration_test.zig`, any test using the writer) — A runs before B, so this is sequential, not shared, ownership; A's gate is the full `zig build test`. - B: `src/web/*`, plus: `src/app.zig` (cache cleanup at the composition root), `src/tests.zig` (handler import swap), the OpenAPI document and its drift guards, contract samples including `admin/src/lib/contractSamples.gen.ts`, `docs/**` API listings, `PLAN.md` stale sections, and the Overview/API sections of `specs/ui-redesign.md` (which still mandates the five endpoints and must be amended, not obeyed). - C: `admin/*` EXCEPT `admin/src/lib/contractSamples.gen.ts`. - D: `src/storage/repositories/queries_repo.zig` (sequential, after B). - Orchestrator: `CHANGELOG.md` (hand-written, per release process). B and C run in parallel; the one shared-tree exception above is reserved to B. ## Acceptance criteria (milestone complete) - [ ] `zig build test` and `zig build test -Dintegration` green (the integration suite carries the live route walk, contract-sample comparison, concurrent querylog reads and OpenAPI guards); admin `npm test`/`typecheck`/`build` green. - [ ] Scratch-instance smoke: seeded data + live digs; Overview correct on all periods; old endpoints gone. - [ ] Changelog discloses: schema change resets query history (rename-aside), five endpoints replaced by `/api/overview`. - [ ] Release gate (`zig build cut` fingerprint check) satisfied. ## Anti-requirements - No second database file, no epoch/validity protocol, no ATTACH. - No backfill migration, no rebuild command, no catch-up cursor — projections are born with the file and maintained transactionally; that is the whole coherence story. - No connection pool, no adaptive polling, no DuckDB. - No new indexes on `query_log`, no triggers, no per-row projection SQL. - Do not change the 1h/24h/7d/30d period grammar, bucket widths or counts. - No HTTP-level caching of any kind: no ETag, no `Cache-Control`, no stale-while-revalidate, no background refresh, and never cache a 500/503 body. The cache is exactly the in-process design of B.2. - No visual redesign, no chart-prop changes, no cache configuration knobs, no cache metrics. This milestone changes data acquisition and storage only.