milestone 30: overview as a dashboard, explicit health contract, period aggregations
Gates / frontend (push) Successful in 1m32s
Gates / test (push) Successful in 1m54s
Gates / package (push) Successful in 5m28s
Gates / container (push) Successful in 14s
Gates / test-aarch64 (push) Failing after 3h10m0s
CI / gates (push) Failing after 3h11m55s

This commit is contained in:
2026-08-22 16:45:15 +02:00
parent 17422fac21
commit 648d9b4496
89 changed files with 7222 additions and 4239 deletions
+27 -21
View File
@@ -2,6 +2,10 @@
Author: Codex (gpt-5.6-sol, extra-high effort), 2026-08-19. **Not accepted yet.** Untracked on purpose until Mokhtar rules on the open questions at the end.
**Amendment, 2026-08-22 (milestone 30).** Step 4 deletes the upstream-minute history subsystem outright rather than keeping its health condition. Once Overview loses the upstream table, the subsystem has no product consumer at all — a writer whose only reader is its own failure signal — and the no-versioning rule forbids leaving it as a stub. So `/api/health` carries **five** conditions, not six: there is no `upstream_history` object. The `upstream_history.write` event code survives as **legacy on the read side only**: stored rows keep their code, the list endpoint passes it through, and the event store resolves any still-open episode once at init. The passages below are amended in place.
**Second amendment, 2026-08-22 (milestone 30, owner ruling after reviewing the built page on a wide monitor).** Overview is **not** three sections. It takes Pi-hole's dashboard layout: four stat tiles, two full-width charts, two breakdown donuts, and no current-state readout at all. The status rows and the active-issues list are withdrawn from Overview — the five health conditions become a compact strip at the top of the **Diagnostics** page, where the episodes that explain them already live, and the open-episode count becomes a badge on the Diagnostics navigation item. The shell's protection indicator is deleted outright; the Pause control moves to the foot of the sidebar, in both the desktop rail and the mobile drawer. Three period aggregations are added to feed the new panels. The §Overview, §API changes, §Time scoping, §Deletions and §Build sequence passages below are amended in place; where the older three-section prose survives elsewhere, this note supersedes it.
Answers that shaped it: the server and API may change; the surface-ownership split is right; file-mode configuration pages are read-only; diagnostics are curated structured events in the vein of Pi-hole's; time scoping is per workflow; a past query must be explainable exactly; Query Log and Live merge.
## Navigation
@@ -44,32 +48,31 @@ Keeps primary navigation because identifying and naming unknown devices is an op
## Overview
Three sections, nothing else.
One question, answered over a period the reader chooses: what did the resolver do. No current-state readout — that moved to Diagnostics (§Diagnostics) — and no active-issues list.
**1. Current status.** Five current facts, each conveyed by text and icon as well as colour. Healthy rows stay quiet; degraded rows link to the diagnostic or configuration surface that explains them.
**Period is URL state.** `/overview?period=1h|24h|7d|30d`, validated in the route search, defaulting to 24h. The picker navigates, so a view of the page is a link.
| Status | Shows | Why it belongs |
| --- | --- | --- |
| Protection | Active, paused until a timestamp, or unavailable, with Pause/Resume | Confirms filtering is in force, and carries the valid runtime action |
| Upstreams | available / configured, now | Confirms DNS can leave the network |
| Query history | Recording, losing rows, or writer failed | Says whether Activity can be trusted |
| Diagnostics | Recording or unavailable | A failure reporter that cannot record failures must itself be visible |
| Storage | ok / low / critical, free bytes | Says whether writes are safe, and explains write gating |
Top to bottom, edge to edge:
The shell carries a small global "Protection active/paused" indicator linking back to Overview. The controls themselves stay on Overview and beside blocked-query details.
1. **Four stat tiles**, neutral chrome throughout — no coloured accents; emphasis is typographic. Queries, Blocked (count and rate), Clients, Average response. Each tile carries the way into the rows behind its number: Queries and Blocked open Activity for exactly the bounds the stats response returned, Clients opens the clients page, and Average response has nothing to open.
2. **Queries over time** — the existing query-volume timeline, split blocked/cached/other, full width.
3. **Client activity over time** — one stacked series per named client plus "other", on the same bucket alignment as the timeline so the two charts share an x-axis. A client registered under a name is labelled by it, with the same precedence the query tables apply and the address kept as the title; colour keys on the address, so naming a client never repaints its series.
4. **Query types** and **Upstream servers** — two donuts, side by side above 1280px and stacked below, with the ring and its legend centred in the panel while stacked and left-anchored once they are a pair. Types are labelled by the admin's own `qtypeName()`; routes by route-kind labels and by the answering resolver or zone. Each donut's SVG is decoration (`aria-hidden`, `focusable="false"`); a visible legend and a visually hidden table are the accessible surface. An empty window says "No queries in this period." rather than drawing nothing.
**2. Active issues.** Severity, short title, affected object, how long it has been active, link to the detail. When none exist, one restrained line: "No active operational issues." Resolved failures never appear here, and healthy subsystems never get permanent green cards.
Colours key on semantic identity — the qtype value, the client string, the `(route, source)` pair — so a rank change between two polls never repaints an entry. Charts stay lightweight SVG; no charting dependency.
**3. Activity over a period.** The existing 1h / 24h / 7d / 30d control. Every value uses exactly the returned `[since, until)` window: queries, blocked count and rate, distinct clients, average response time, one query-volume timeline split blocked/cached/other, and "Open activity for this period" carrying the exact bounds. The timeline stays the existing lightweight SVG; no charting dependency.
**Window coherence, five requests.** Totals, timeseries, clients, types and routes are separate calls, and the page holds one window identified by `(period, since, until, coverage.available_since)` — the watermark joins the identity because retention advancing mid-page changes what the same span can answer for. A response is a member only if all four fields match. Rendering is per panel: a member renders, a panel still in flight shows its own loading state, a panel whose request failed shows its own error and Retry, and the members keep rendering throughout — a failed donut never blanks the charts. A response behind the window is refetched once per endpoint-keyed episode and, if it stays behind, that panel alone shows an error. This is window coherence, not data-snapshot coherence: live inserts between requests may shift counts slightly between panels, and that is accepted. One coverage notice for the page, from the window's watermark.
If the selected period predates available data, the section says "Query history is available from …" rather than charting the missing span as zero.
**The shell.** The header carries no protection display at all. The Pause/Resume control sits at the foot of the sidebar, above the version label, in both the desktop rail and the mobile drawer; it is the only global runtime action, and it belongs to the resolver rather than to any page. It still appears beside the detail of a query that was blocked. The control states a pause with itself — "Paused until 14:05", or "Paused" when the pause has no end — because "Resume" names an action without naming the state it would end, and with the indicator and the status rows both gone the sidebar is the only place a page other than Diagnostics can carry that fact. An active resolver gets no line; the button says Pause, which is the whole message. The line and the health strip read one `protection` condition through one clock format, so they cannot disagree. The Diagnostics navigation item carries a badge: the open-episode count, or a neutral "!" when the rollup is degraded with nothing open and when the latest health poll failed — an unknown must never read as healthy. It is hidden only when health data exists, the latest poll succeeded, and the rollup is ok with nothing open.
Removed from Overview: the historical upstream table, the "last failure" text, per-upstream period rates, the database and log byte breakdown, the standalone cache card.
Removed from Overview: the historical upstream table, the "last failure" text, per-upstream period rates, the database and log byte breakdown, the standalone cache card, the status rows, the active-issues list.
## Diagnostics
Not a journald viewer, and it does not subscribe to `std.log`. Producers emit a finite set of typed events at the failure boundary.
**The health strip (amended 2026-08-22).** The page opens with the five `/api/health` conditions rendered compactly: Protection, Upstreams, Query history, Diagnostics, Storage, each stating its state in words and an icon as well as colour. A healthy condition is quiet; a degraded one is highlighted and offers the way out. The link matrix is exact — protection `unavailable` to Blocklists, upstreams `unavailable` to Upstreams, query history `losing` to this page filtered to `component=disk`, `failed` to `component=query_log`, disk `low` or `critical` to `component=disk`, and diagnostics `unavailable` to nothing at all, because the surface a link would filter is the thing that is broken. An in-page filter link sets `component` and clears the time bounds, which could otherwise hide the very episodes it points at; severity and state keep whatever the reader chose. Dropped-row counts ride the query-history entry. Loading: a visible state before the first reading; on a refetch failure an error row with Retry, with the conditions on screen marked as the last reading that arrived rather than the current state, cleared when a poll succeeds again. This is the surface the withdrawn Overview status rows became.
### Event model
One row is one failure episode.
@@ -110,7 +113,6 @@ If the store itself cannot write, an atomic `event_store_failed` state appears i
| Certificate stat or reload failure | `certificate.reload`, keyed by `doh`/`dot` | files readable and reload succeeds |
| Query writer init or batch failure | `query_log.write`, keyed by `writer`/`batch`/`queue` | writer starts, or a batch succeeds without drops |
| Query retention prune/checkpoint/vacuum failure | `query_log.maintenance`, keyed by operation | that operation succeeds |
| Upstream-history flush failure | `upstream_history.write`, singleton | next flush succeeds |
| Client-name selection or persistence failure | `client_names.storage`, keyed by operation | next pass succeeds |
| Client materialisation or pruning failure | `clients.storage`, keyed by operation | next pass succeeds |
| Upstream exchange failure | `upstream.exchange`, keyed by upstream URL | next successful exchange |
@@ -118,7 +120,7 @@ If the store itself cannot write, an atomic `event_store_failed` state appears i
| Query-log recreation | `query_log.recreated`, one-shot | inserted resolved |
| Configuration warning leaving a capability skipped | `configuration.load`, keyed by setting | clean load after restart |
Sixteen codes. Codex proposed a seventeenth, `api.storage`, for a storage failure that produced an HTTP 500 — cut on review, because it breaks this section's own exclusion rule: a 500 already answered its caller. Do not merge the remaining codes to shrink the count either. A merged code forces `subject_key` to carry what the code no longer says.
Fourteen codes emitted. A fifteenth, `upstream_history.write`, is legacy and read-side only: milestone 30 deleted its producer and its emitter enum member, and it survives in the documented wire union — fifteen values in all — so stored rows stay readable and in-contract. Codex proposed a sixteenth, `api.storage`, for a storage failure that produced an HTTP 500 — cut on review, because it breaks this section's own exclusion rule: a 500 already answered its caller. Do not merge the remaining codes to shrink the count either. A merged code forces `subject_key` to carry what the code no longer says.
An upstream event describes a consecutive failure episode, not one row per retry. One timeout followed by success is one resolved episode.
@@ -180,7 +182,7 @@ Live SSE events carry the same provenance shape without a persisted `id`. A froz
One contract: unix seconds UTC, `since` inclusive, `until` exclusive, point data qualifies on `since <= ts < until`, diagnostic episodes qualify when their active interval overlaps the range, current state is labelled "Now" and no historical selector touches it.
URLs: `/overview?period=24h` with the server returning the exact aligned bounds; `/activity?mode=history&since=…&until=…` with `domain`, `client`, `blocked` and the other filters in the URL; `/diagnostics?since=…&until=…&severity=…&component=…`; `/activity?mode=live` with Follow/Freeze as ephemeral UI state. Investigation links always carry absolute bounds, so a viewed incident does not drift as time passes. Timestamps display in the browser's timezone; URLs and APIs stay timezone-independent.
URLs: `/overview?period=24h`, validated in the route search and defaulting to 24h, with the server returning the exact aligned bounds and every panel of the page judged against one window identity that includes the coverage watermark; `/activity?mode=history&since=…&until=…` with `domain`, `client`, `blocked` and the other filters in the URL; `/diagnostics?since=…&until=…&severity=…&component=…`; `/activity?mode=live` with Follow/Freeze as ephemeral UI state. Investigation links always carry absolute bounds, so a viewed incident does not drift as time passes. Timestamps display in the browser's timezone; URLs and APIs stay timezone-independent.
## File mode
@@ -200,7 +202,7 @@ Database mode uses the same information architecture with real edit actions, plu
`GET /api/config/status``{authority, path, reconciled_at, restart_pending}`. `restart_pending` is process state: database-mode mutations that need a restart set it, a successful restart clears it.
`GET /api/health` becomes explicit about every condition that contributes to degradation — `protection`, `upstreams`, `query_history`, `upstream_history`, `diagnostics`, `disk`, each an object with its own state. The current hidden `history_flush_failing` contribution is eliminated: nothing may degrade the rollup without appearing in the response.
`GET /api/health` becomes explicit about every condition that contributes to degradation — `protection`, `upstreams`, `query_history`, `diagnostics`, `disk`, each an object with its own state. Nothing may degrade the rollup without appearing in the response, so the hidden `history_flush_failing` contribution goes, and so does the subsystem behind it. The degrading set is exactly: protection `unavailable`, upstreams `unavailable`, query history `losing` or `failed`, diagnostics `unavailable`, disk `low` or `critical`. A paused protection is surfaced, never alarmed. The disk monitor's `warn` is renamed `low` at the serialization boundary. `queries_dropped`, `writer_failed`, `refreshes_gated` and `snapshot_generation` leave the body; the first two fold into `query_history`, and the last two stay in Prometheus.
`GET /api/diagnostics?state=&severity=&component=&since=&until=&limit=&before=` returns `{events[], next_before, active:{warnings, errors}}`. `GET /api/diagnostics/{id}` returns one event or 404 after retention. No acknowledgement, dismissal, generic-action or raw-log endpoints.
@@ -208,6 +210,8 @@ Database mode uses the same information architecture with real edit actions, plu
`GET /api/stats` and `/api/stats/timeseries` add `complete` and `available_since`.
**Three period aggregations (added 2026-08-22)** to feed the new Overview panels, all taking the same `period` parameter and reporting over the same aligned window, and all reading their rows and their coverage watermark inside one deferred SQLite read transaction. `GET /api/stats/types``{period, since, until, coverage, types:[{qtype, count}]}`, the numeric type only — naming types stays the admin's job, and a second table in the server would drift out of agreement with it — with the rows that recorded no type kept as their own `null` group. `GET /api/stats/routes``{period, since, until, coverage, routes:[{route, source, count}]}`, grouping `upstream` rows by the answering resolver and `forward_zone` rows by the zone, with blocked, cache, local and rejected carrying no source. `GET /api/stats/clients``{period, since, until, bucket_seconds, coverage, clients:[{client, buckets}], other}`, bucketed exactly as `/api/stats/timeseries`, the eight busiest clients named and everything else summed into `other`, which is always present and always bucket-count-sized. No new writers and no new state: all three are pure reads over the query log's provenance columns.
Existing mutation endpoints stay specific. Diagnostics introduces no generic "perform remediation" endpoint; it invokes the existing blocklist-refresh and certificate-reload operations.
## Schema changes
@@ -249,7 +253,7 @@ No new provenance table, no key/value store — household retention makes nullab
`transformed()` must hide `matched`, `cname_target` and `safe_search_target` under `hide_domains`, not only `domain`.
On recreation: query rows, provenance and upstream-minute history reset together as today; the old file stays aside; `config.db` diagnostics survive; a resolved `query_log.recreated` event records the reason, the aside path and the new coverage start; Overview and Activity report the incomplete range instead of charting zero.
On recreation: query rows and provenance reset together; the old file stays aside; `config.db` diagnostics survive; a resolved `query_log.recreated` event records the reason, the aside path and the new coverage start; Overview and Activity report the incomplete range instead of charting zero.
## Deletions and their cost
@@ -258,7 +262,9 @@ On recreation: query rows, provenance and upstream-minute history reset together
| Separate Query Log and Live pages | separate bookmarks; both modes remain in Activity |
| Standalone Lookup page | a top-level bookmark; testing remains under Activity |
| Top-level Groups, Blocklists, Rules, Local DNS, Upstreams, Settings | direct resource navigation; all capabilities remain under task-shaped configuration |
| Historical upstream table on Overview | at-a-glance period rates; availability stays, failures move to Diagnostics |
| Historical upstream table on Overview, and the upstream-minute history subsystem behind it | at-a-glance period rates, and the ranged per-upstream counts entirely; availability stays on the Diagnostics health strip, failures are Diagnostics episodes. `querylog.db` resets on the schema change |
| Overview status rows and active-issues list (2026-08-22) | a current-state readout on the landing page; the five conditions move to the Diagnostics health strip and the open count to the Diagnostics nav badge |
| Shell protection indicator (2026-08-22) | protection stated on every page; the sidebar Pause control carries it, its label for the action and its state line for a pause, and the health strip states it in full |
| "last failure · 9h ago" text | nothing actionable; the episode becomes a diagnostic |
| Detailed DB/log byte gauges | exact component sizes stay in Prometheus; free space stays on Overview |
| Standalone cache card | one prominent number; cache stays in the timeline and metrics |
@@ -279,7 +285,7 @@ Each step leaves the app working and shippable, and updates its OpenAPI contract
Parsed-SERVFAIL logging reverses ruling 20: `handler.zig:507` counts today rather than logging. The `Context` exists at every `servFail` site that follows question parsing. Pre-parse failures correctly stay counters.
3. **Activity consolidation.** The unified History/Live surface, URL filters, freeze/follow, live detail, current-policy test, historical detail links. Query Log, Live and Lookup routes and code are removed in the same change. Route, SSE, accessibility, reconnect and bounded-buffer tests.
4. **Overview replacement.** Current status, active diagnostics, one coherent activity section. New health contract and completeness states. The upstream-history table, stale-failure text, detailed DiskCard and cache card go.
4. **Overview replacement.** New five-condition health contract and completeness states, the three period aggregations, and — per the 2026-08-22 ruling — Pi-hole's dashboard layout: four stat tiles, the query-volume and per-client charts, the types and routes donuts, all against one coherent window. The five conditions become the Diagnostics health strip and the nav badge; the shell indicator dies and Pause moves to the sidebar foot. The upstream table, stale-failure text, detailed DiskCard and cache card go, and the upstream-minute history subsystem goes with the table — `src/upstream/history.zig`, its repository, its two `querylog.db` tables, `GET /api/upstream/health` and its four Prometheus metrics. Deleting the tables changes the query-log fingerprint, so this step resets `querylog.db` the same way step 2 does.
5. **Task-shaped configuration and file mode.** `/api/config/status` and server-owned `restart_pending`. Protection, Resolution and System in both read-only and editable forms. Clients and its detail route redesigned. Old configuration routes replaced atomically; global banner and disabled forms removed.
6. **Contract closure.** Remove obsolete queries, types, stores, CSS, tests and route fixtures. Regenerate contract samples, update OpenAPI and reference docs, add cross-surface acceptance tests for investigation links, file authority, query-log recreation, active-event recovery and time bounds. Zig, frontend, integration, accessibility and byte-budget checks; no new dependency.