27 KiB
Milestone 30: Overview replacement
Redesign step 4 of specs/ui-redesign.md (§Overview, §API changes /api/health, §Deletions), amended by the owner's 2026-08-22 ruling: the dashboard becomes Overview with Pi-hole's dashboard layout (stat tiles, two charts, two donuts — no status or issues sections; see S3/S2). The health contract becomes explicit — nothing may degrade the rollup without appearing in the response. Codex design review folded in (thread 01a028b1); corrections marked where they changed the shape.
One amendment to the accepted design, made here and recorded in the changelog: ui-redesign.md:113/203 lists an upstream_history health condition. After this milestone deletes the dashboard's upstream table, the upstream-minute history subsystem has no product consumer at all — a writer whose only reader is its own failure signal. The no-versioning rule forbids keeping it as a stub, so this milestone deletes the subsystem (S1.2) and the health contract has five conditions, not six.
Sessions
S1 (done): Zig — health contract, subsystem deletion, contracts, docs. S3 (after S1): Zig — the three period aggregations. S2 (after S3): admin — the Pi-hole Overview, sidebar pause control, Diagnostics badge/strip, deletions, smoke. Sequential; the milestone lands as one atomic commit (Codex: S1 alone breaks the generated contract sample's TS assignment and every Health consumer — its red admin typecheck is an intra-milestone state, never a landed one).
Session S1: the explicit health contract
S1.1 GET /api/health (src/web/handlers/health.zig)
New body: status plus five condition objects. The degrading set is exactly (Codex: name it, don't gesture at "not ok"): protection unavailable, upstreams unavailable, query_history losing|failed, diagnostics unavailable, disk low|critical. Nothing else reaches degraded() — this kills the hidden history_flush_failing input (health.zig:65-73/177-178) along with its subsystem.
| object | shape | state rule |
|---|---|---|
protection |
{state: "active"|"paused"|"unavailable", until: ?i64} |
Precedence unavailable → paused → active (Codex). unavailable ⇔ no current filter snapshot exists (the condition under which handler.zig answers with snapshot_unavailable provenance) — pause state is irrelevant then, until null. Else paused ⇔ the pause is live now (indefinite → until null; timed and unexpired → until set). An expired timed pause is active (Codex). Paused does NOT degrade — operator choice, surfaced not alarmed; unavailable degrades. |
upstreams |
{state: "ok"|"unavailable", available, total} |
unavailable ⇔ available == 0. total = enabled routing candidates (the pool is built from enabled upstreams only — Codex); S2 labels it "of N enabled". |
query_history |
{state: "recording"|"losing"|"failed", dropped_total, last_drop_s: ?i64} |
failed ⇔ writer_failed. losing ⇔ a logger-owned gating-episode state (Codex critical: draining is a shutdown flag, false during normal disk gating): logger marks the episode active while the disk gate holds writes, flips to losing once a drop occurs during that episode, clears when the gate reopens. Event-driven; no time window. Else recording. Queue-overflow drops outside a gating episode do not flip the state (cumulative ≠ current); they surface via dropped_total + last_drop_s (atomic stamped per drop, null until one happens). |
diagnostics |
{state: "recording"|"unavailable", active_warnings, active_errors} |
current rule (!present or write_failed) |
disk |
{state: "ok"|"low"|"critical", free_bytes} |
disk_monitor states; warn→low renamed at the serialization boundary; db_bytes/log_bytes/sample_failures leave the body (Codex: sample failures already have a Prometheus counter and active disk.probe diagnostics — same ruling as refreshes_gated) |
All current top-level extras die: queries_dropped/writer_failed fold into query_history; refreshes_gated leaves health entirely (its Prometheus counter remains the record — Codex: table/prose disagreed, this is the ruling); snapshot_generation dies (no real admin consumer — Codex; generation stays in Prometheus, and protection unavailable is the UI-facing fact).
logger.zig work (S1-owned): the gating-episode state + last_drop_s atomic, stamped at the drop sites (logger.zig:687 area). Episode start must be race-free (Codex: runWriter fills a batch before checking the gate, so a naïve "mark on gate check" misclassifies gate-caused overflow drops as ordinary ones): the episode begins when the first pending batch observes the closed gate — before further timed filling — or via an equivalent race-free gate observation. Rollup tests rewritten per condition: each degrading state degrades alone and is visible in the body; paused does not degrade; expired pause reports active; protection unavailable wins over paused.
S1.2 Delete the upstream-minute history subsystem
Extent (verified by exploration; the build agent must hit every item):
- src/upstream/history.zig (whole file) and src/storage/repositories/upstream_history_repo.zig (whole file).
- querylog_schema.zig:
upstream_targets(:74),upstream_minute(:79),idx_upstream_minute_ts(:90), expected-object list :341-342. This changes the schema fingerprint → querylog.db resets via the m28 aside mechanism; changelog states it. - app.zig wiring: :52,75,549-553 init/diagnostics/pool.history, :639 reopen path, :800 web state, :925-939 shutdown flush, :954 run task.
- pool.zig:
historyfield :52,138-141 andrecordHistory:377-386 (history-only).entry.health.recordSuccess/recordFailure(:345,:364) andrecordDiagnostics(:390+) are live routing/backoff/diagnostics state and must survive; callsites :259,:269 keep those calls. - retention.zig: history prune :19,134-135 + tests :457-505.
- Handler src/web/handlers/upstream_health.zig, route routes.zig:77, server.zig:43, openapi.yaml :501 path + :2122 enum, goldens.
- health.zig history signal (dies with the rollup rewrite anyway).
- metrics.zig: the four
nxdns_upstream_history_*metrics :34,221,391-411 + tests :819-828. - events.zig: the emitter enum entry dies, but the read side keeps the legacy code (Codex critical: rows survive, the list endpoint passes stored codes through, and dropping the union member would make real responses violate the contract; unresolvable active legacy episodes would also pin
untracked_active_countand the SQL slow path forever). Contract:upstream_history.writestays in the openapi/TS event-code union and in eventCopy.ts, marked legacy; on store init/upgrade, any activeupstream_history.writeepisode is resolved once, with the open time as its resolution time (not oldlast_seen— Codex: else the episode can be pruned in the same init), and before the active-mirror load/count and pruning (or the mirror is recomputed after) sountracked_active_countand the SQL fast path settle correctly. Test: an m29 database with an active flush-failure episode opens with that episode resolved and listable,untracked_active_countat 0, and the fast path in use. - retention.zig: also
Retention.Stats.upstream_rows_pruned, its atomic counter, andnxdns_retention_upstream_rows_pruned_total(Codex: they would remain forever zero). - tests.zig :80,:84; affected web_integration_test.zig / resolver_integration_test.zig / pool.zig:937-980 tests.
stats.periodParamstays (stats uses it).- Sweep (S1.4 gate): case-insensitive grep over src/ docs/ PLAN.md for
upstream[_ -]history,upstream_minute,upstream_targets,upstream_rows_pruned,/api/upstream/health— known stale spots include src/cli.zig, querylog_schema.zig WAL commentary, src/upstream/health.zig, metrics comments (the read-side legacy event code is the one allowed survivor).
S1.3 Contracts and docs
openapi.yaml: Health schema rewritten (five objects, enums exact), upstream/health path+schema removed. Goldens regenerated; drift guards updated (route count −1; the health guard covers the nested objects). Docs are S1-owned (Codex): a prose sweep, not just cited lines — PLAN.md (upstream-history/API mentions :328,489-500,563 plus the summary and page/manual-check passages), docs/reference/api.md, docs/how-to/troubleshoot.md (old flat health JSON), docs/tutorial/first-run.md and docs/reference/configuration.md ("Dashboard" naming → Overview), stale comments in pool.zig. specs/ui-redesign.md is S1-owned too (Codex): amend §§event sources, API changes, recreation, deletion cost, and build sequence to the five-condition contract and the subsystem deletion, with a dated amendment note. CHANGELOG (S1 half): health contract break (pre-v0.1), subsystem + endpoint removal, querylog.db reset, warn→low, the design amendment.
Destructive-transition acceptance (Codex: generic fingerprint tests do not pin this): a test opens an m29-schema querylog.db, asserts it is renamed aside, recreated without the two tables, coverage restarts, and query_log.recreated is emitted. S2 updates the query_log.recreated eventCopy wording, which currently says the database "was unreadable" — false for a planned schema change; the copy must encompass both causes.
S1.4 Acceptance (S1)
- Rollup tests per S1.1, incl. gating-episode losing-state transitions (gate holds without a drop → recording; drop during episode → losing; gate reopens → recording; writer death → failed) and
last_drop_sstamping on a forced queue overflow. zig build test+-Dintegration0 failed; goldens green; fmt clean; the S1.2 case-insensitive sweep finds nothing except the allowed read-side legacy survivors: theupstream_history.writecode in the event union/openapi/read path and their tests (Codex: state the exception so the gate is passable). Admin is expectedly red until S2 — an intra-milestone state only; nothing is committed before S2's gates pass.
Session S3: period aggregations (Zig)
Owner ruling 2026-08-22: after reviewing the built status-row Overview on a wide monitor, the owner replaced the three-section design with Pi-hole's dashboard layout ("just copy pi-hole's dashboard's layout; get rid of the current status section"). S1's health contract is unchanged. S3 adds the aggregations; S2 describes the final admin state. Build order: S1 (done) → S3 → S2.
All three aggregations read the query log over the same [since,until) Span the stats endpoints use (stats.periodParam, floor-aligned buckets), return period/since/until/coverage like the existing stats bodies, and honor hide_domains/hide_client_ips transforms where fields are sensitive. Pure SQL over m28's provenance columns — no revival of the upstream-minute subsystem, no new writers, no new state.
Per-response atomicity (Codex): each response — the three new endpoints AND the existing totals/timeseries — reads its aggregation and its coverage inside one SQLite read transaction, so retention cannot prune between the aggregate and the watermark, and the clients endpoint ranks and buckets from one database state. Transaction ownership (Codex critical): WebState shares one querylog_db connection across concurrent HTTP tasks with no mutex — SQLite serialized mode protects single calls, not transactions; concurrent BEGINs fail and foreign reads can interleave inside a transaction. S3 adds a querylog_lock covering every web-layer access to that connection (or separate read connections — implementer's choice, stated in the report), and the read transaction uses deferred BEGIN, not the existing db.Tx BEGIN IMMEDIATE (which would block the logger/retention writers); if a read-transaction abstraction is added, db.zig joins S3 ownership. Tests: concurrent stats/query requests under load; a prune racing a response never yields pre-prune data tagged with a post-prune available_since.
S3.1 GET /api/stats/types — query-type breakdown
{period, since, until, coverage, types: [{qtype: ?u16, count: u64}]} — GROUP BY qtype. No name field (Codex: the only qtype-name mapping lives in admin qtype.ts; a Zig copy would drift — the UI labels codes with its existing qtypeName()). query_log.qtype is nullable: null groups into its own row (qtype: null), never silently dropped. Ordering: count DESC, then qtype ASC with null last (deterministic for goldens/colors). No zero rows.
S3.2 GET /api/stats/routes — how queries were answered
{period, since, until, coverage, routes: [{route: RouteKind-wire-string, source: ?string, count: u64}]}. Grouping key exact (Codex: source_name is blocklist provenance, NOT the answering upstream): upstream rows group by query_log.upstream; forward_zone rows by query_log.forward_zone; blocked/cache/local/rejected rows have source: null. A null upstream/forward_zone identity on those route kinds is its own source: null row (UI labels it "Unknown"). Ordering: count DESC, then route ASC (the stored wire string, i.e. alphabetical: blocked, cache, forward_zone, local, rejected, upstream — not enum declaration order), then source ASC nulls last. This feeds the "Upstream servers" donut: blocked/cache/local/rejected shares by route kind, each upstream and forward zone by name.
S3.3 GET /api/stats/clients — per-client timeseries
{period, since, until, bucket_seconds, coverage, clients: [{client: string, buckets: [u64]}], other: [u64]} — bucket_seconds, the established field name (Codex). Top 8 clients ranked by total in-window count DESC then client ASC (deterministic cut); everything else sums into other, which is always present and bucket-count-sized, including empty windows and ≤8-client windows. Client strings arrive as stored — redaction is write-time (logger substitutes the hidden marker before the row exists), so the read path has no transform; hidden rows aggregate as one client named hidden (as-built, tested). Bucket alignment identical to /api/stats/timeseries so the charts share an x-axis; zero-filled; every series length equals the bucket count.
S3.4 Contracts
Routes registered (session auth, stats rate-limit class), openapi paths + schemas, drift guards extended to the three bodies. Contract-sample seed traffic must exercise ≥2 qtypes, ≥2 route kinds and ≥2 clients. contractSamples.gen.ts regeneration happens in S3 (Codex: the committed golden is TypeScript and byte-compared by W10 during integration; S3 cannot pass its own gate without it) — the handwritten-TS typecheck stays red until S2, same intra-milestone rule as S1.
Handler tests per endpoint: empty window, populated matrix, redaction transforms, coverage flag, null-qtype row, null-source row, the top-8 boundary (9th client folds into other), tie-order determinism. Empty-window bodies are exact: types/routes return empty arrays, clients returns zero-filled buckets with an empty clients list. Conservation tests (Codex, controlled-state only — see S2.2's window-coherence limit): over one shared span, sum(types.count) == totals.queries, sum(routes.count) == totals.queries, and per bucket sum(named client series) + other == timeseries.queries — catches null-loss, route omission, and bad partitioning.
S3.5 Acceptance (S3)
zig build test+-Dintegration0 failed (W10 included); goldens green; fmt clean. Admin handwritten-TS red stays expected until S2.
Session S2 (revised): Overview as Pi-hole's dashboard (admin)
S2.1 Shell
- Nav: "Overview" first; root index redirects to
/overview. The header ProtectionIndicator is deleted — no protection display in the header. - PauseControl moves to the sidebar bottom, directly above the version label — in BOTH sidebar renderings (desktop rail and the mobile drawer, each of which has the version footer; Codex): same control, reading
useProtection(), hidden while protection is unavailable/unknown, menu closes when state leaves active, mutation invalidates healthQuery. Existing PauseControl tests move with it plus placement tests for both renderings; the 390px smoke exercises the drawer. The control states the pause with itself (live-smoke gap, 2026-08-22: with the indicator and the status row both deleted, a paused resolver left no trace in the DOM outside/diagnostics, and "Resume" names an action without naming the state it would end): a state line rendered with the control in both renderings —Paused until HH:MMwhenuntilis set,Pausedwhen it is null, nothing at all while active, since the button already says Pause. Text, never colour alone; the sameformatClockand the sameprotectionreading as the health strip, so the two cannot disagree, anduseProtection's expiry refetch retires both together. Tested for both paused shapes and for its absence while active. - Diagnostics nav item badge: from healthQuery —
active_warnings + active_errorsas a count; degraded healthstatuswith zero events shows "!" so no degraded state is invisible; a failed health poll shows a neutral "!" with accessible text "Health unavailable" (Codex: unknown must not be unbadged — only initial loading may be); hidden only when health data exists, the latest poll succeeded, and status is ok with zero events — "fresh" means exactly that, never TanStack staleness (Codex: healthQuery's staleTime is 0, soisStalewould badge every gap between polls). Count/shape+text, never color alone. Tests: 0 hidden, N shown, degraded-zero-events "!", failed-poll "!", initial-loading hidden. - Diagnostics page health strip: the five S1.1 conditions rendered compactly at the top; healthy quiet, degraded highlighted. Full load contract migrates from the deleted StatusSection (Codex): visible loading before first health response; on refetch failure an error row with Retry, cached conditions marked stale, never presented as current; recovery clears the staleness. Exact link matrix, restated (Codex: no dangling reference): protection unavailable →
/blocklists; upstreams unavailable →/upstreams; query history losing → in-page filtercomponent=disk; failed →component=query_log; disk low/critical →component=disk; diagnostics unavailable → no link, copy "Diagnostics are not being recorded. Check free disk space and the configuration database's permissions." In-page filter links setcomponentand clear incompatible active filters (severity/state stay default). The drops secondary text ("N queries dropped, last at HH:MM"; "N queries dropped" whenlast_drop_snull) lives on the query-history entry.
S2.2 The page — Pi-hole's layout with our data
Period is URL state: /overview?period=1h|24h|7d|30d validated in the route search, default 24h; the picker navigates (functional search update); deep-link and reload tests (Codex: the accepted time-scoping contract keeps shareable period state; the layout ruling did not revoke it).
Top to bottom, edge-to-edge grid, no status/issues sections:
- Four stat tiles, neutral chrome (owner ruling: no colored accents; emphasis via value typography): Queries, Blocked (count + %), Clients, Avg response. Footer links: Queries →
/activity?mode=history&since&until(returned bounds); Blocked → same +blocked=true; Clients →/clients; Avg response → none. Period picker in the section header governs every panel. - Queries over time — the existing stacked chart, full width.
- Client activity over time — new stacked chart from
/api/stats/clients: one series per named client + "other", same x-axis, legend with client names. Series are labelled by registered name where one exists (owner, 2026-08-22): the frontend resolves address → name at render time through the existinguseClientNames()/clientLabel()lookup the query tables already use — hand-typednamefirst, reverse-DNSlearned_namebehind it, the bare address when the clients list knows neither — and the legend and the hidden table carry the same resolved string, so the graphic and the accessible surface never name one client differently. Colour keys on the stored address regardless, so registering or renaming a client never repaints its series.clientsQueryjoins the loader's fire-and-forget set and polls on the same 30s cadence as the tables. Tests: named client renders its name, unregistered renders its address, and the swatch of a named series is still the colour of its address. - Query types and Upstream servers donuts from
/api/stats/types//api/stats/routes. Two-column row only above a named breakpoint (1280px); stacked below (Codex). Stacked, the ring and its legend are centred in the panel (owner, 2026-08-22: full-page width with the ring pinned left reads as a mistake); above the breakpoint each donut is one of a pair and stays left-anchored, in line with the panels above it. Labels: qtype via the existingqtypeName()(one home,TYPE<n>fallback, "Unknown" for null); routes via route-kind labels ("Blocked", "Cache", "Local", "Rejected") and source names ("Unknown" for null source on upstream/forward-zone rows). New reusable SVG donut, non-focusable SVG + visible legend + a visually-hidden table carrying every label/count (Codex: pin the pattern, no delegated a11y choice). Empty windows render "No queries in this period." in both donuts and the client chart — never a blank panel or a division by zero (tested). Legend/table identity and React keys are the composite(route, source)/ qtype value — display names may collide ("Unknown" twice, same source name on two kinds), so ambiguous entries carry secondary route-kind text. Series/slice colors key on semantic identity — qtype value, client string, and for routes: fixed colors for the source-less route-kind slices (Blocked/Cache/Local/Rejected) while upstream/forward-zone slices key on the full(route, source)pair (Codex: keying on route kind alone would merge adjacent upstream slices); "other" fixed — so a rank change between refreshes never recolors an entry. The donut SVG carriesaria-hidden="true"andfocusable="false"(Codex: non-focusable alone does not leave the accessibility tree; the legend + hidden table are the accessible surface).
Window coherence, five requests (totals, timeseries, clients, types, routes). This is window coherence, not data-snapshot coherence (Codex: matching identities cannot prove a common database state — live inserts between requests may shift counts slightly between panels, and that is accepted; conservation holds in controlled tests only). The page holds one coherent window identified by (period, since, until, coverage.available_since) (coverage joins the identity — retention advancing between requests must not mix pre/post-prune windows). Window selection comparator (Codex): newer until wins; for equal bounds, newer available_since; tested where a bucket boundary and a coverage advance cross. a response is a member only if all four fields match. Rendering is per-panel against the window: panels whose data matches render; a panel whose request is loading shows its own loading state; a panel whose request failed shows its own error+Retry — other panels keep rendering the coherent window (a failed donut never blanks the charts). Laggards: every response behind the window is refetched once per endpoint-keyed mismatch episode (episode key = endpoint + window identity); a retry token invalidates all stale completions, including across period changes; previous-period placeholder data never participates. If an endpoint stays behind after its retry, that panel shows the error state; the page never renders two windows at once. One CoverageNotice, derived from the window's watermark. Tests: one lagging endpoint refetched once then coherent; still-behind terminal error per panel; period change with in-flight stale completions; coverage-advance mismatch; per-panel failure isolation.
S2.3 Pause beside blocked queries — unchanged
RelatedActions on both detail surfaces, shared control, hidden when unavailable/unknown.
S2.4 Migration and smoke
- Delete: StatusSection, statusRows, IssuesSection, ProtectionIndicator (+ their tests — behaviors re-pinned on the sidebar control/badge/health strip or declared dead in the report; "unknown never renders active" re-pins on the sidebar PauseControl and badge). The
/overviewroute loader is replaced (Codex): drop the active-diagnostics prefetch, prefetch the five stats endpoints + health; sweep stale loader comments/imports and stale ownership comments in PauseControl, protection.ts, OverviewPage, ActivitySection. - types.ts for the three new endpoints (hand-written halves); healthQuery consumers updated (badge, strip, control).
- Tests per S2.1/S2.2 lists plus root redirect.
- Smoke on the real binary (smoke28 config, traffic across ≥2 clients and ≥2 qtypes): screenshots at 1280/1920/2560/390 — healthy Overview fully populated (tiles, both charts, both donuts); paused via the sidebar control (desktop + drawer); Diagnostics badge + health strip with a tripped blocklist-refresh diagnostic (a dead forward zone produces SERVFAIL but no diagnostic); the Queries tile link landing on the exact window;
/overview?period=1hdeep link. Prefix m30smoke-.
S2.5 Acceptance (S2)
- Admin gates clean;
zig build -Dadmin-dist=admin/distsucceeds; zig suites untouched-green. - Screenshots per S2.4.
File ownership
S1: done (src/, PLAN.md, docs/, specs/ui-redesign.md, CHANGELOG S1 half). S3: src/web/** (three handlers, routes.zig, openapi.yaml, goldens incl. contractSamples.gen.ts regeneration, web_integration_test.zig), src/storage/repositories/queries_repo.zig (aggregations), PLAN.md and docs/reference/api.md again (Codex: PLAN still says health is surfaced on Overview with status+issues; both need the three new routes and the Overview description fixed), CHANGELOG (S3 lines). S2: admin/src/**, CHANGELOG (S2 half), specs/ui-redesign.md re-amendment covering Overview, shell placement, API additions, time scoping, deletions, and build sequence (dated note; the Pi-hole ruling supersedes the three-section prose). Sequential S3 → S2; one commit.
Anti-requirements
- No new health inputs beyond the five objects; no polling-cadence changes.
- No revival of the upstream-minute subsystem — the routes breakdown is a query-log aggregation.
- No charting dependency; SVG only, donuts included.
- No colored tile accents (owner ruling); color carries state on values only.
- No Zig qtype-name mapping (the admin's qtype.ts stays the one home).
- No task-shaped configuration pages (step 5); no contract-closure sweep (step 6).
- Paused protection must not degrade
status.
Acceptance (milestone complete)
- All suites green (zig, integration, goldens, admin); one signed commit; screenshots per S2.4.
- CHANGELOG documents the health contract break, the subsystem removal + querylog.db reset, warn→low, the three new stats endpoints, the Pi-hole Overview, and the ui-redesign amendments.