Gates / frontend (push) Successful in 1m36s
Gates / test (push) Successful in 1m56s
Gates / test-aarch64 (push) Successful in 7m37s
Gates / package (push) Successful in 9m12s
Gates / container (push) Successful in 13s
CI / gates (push) Successful in 19m4s
query rows gain qclass, rcode, group, policy action and reason, the matched rule or list entry with its source, cname and safe-search targets, route kind, forward zone, and the resolver that actually answered — the pool and local markers die. servfails are logged and name the resolver that lost; post-parse protocol refusals become rows. a detail page at /queries/:id renders the ordered explanation, and coverage watermarks distinguish an empty history from a missing one. the schema fingerprint changes: existing query history is recreated with the old file kept aside and the reset filed as a resolved diagnostic. fixes an oversized udp reply being rebuilt as noerror, which handed clients a truncated nxdomain as success.
148 lines
25 KiB
Markdown
148 lines
25 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to nxdns are recorded here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
|
|
Sections are written by hand. Nothing here is generated from commit messages: the point of the file is to say what changed for an operator, which a commit subject rarely does.
|
|
|
|
## [Unreleased]
|
|
|
|
Query provenance: every logged query becomes exactly explainable — what the policy decided, what matched, where the answer came from and what the client saw. The handler records all of it as the reply goes out, `query_log` stores it, and a detail page reads one query back in the order the pipeline decided it. Read the upgrade note below first: it resets your query history.
|
|
|
|
### Added
|
|
|
|
- **Every logged query has a detail page.** A row in the query log now links to `/queries/{id}`, which explains that one query in the order it was decided: the request, the group it was matched under, the policy verdict with the rule that produced it and the blocklist source that rule came from, any CNAME uncloaking or safe-search rewrite, the route the answer took — blocked, local, forward zone, upstream or cache — and what the client got back, RCODE and duration included. `GET /api/queries/{id}` serves the same object; an id that retention has already deleted is a 404. The live view carries the same provenance for the queries it streams, so a query is explainable as it happens as well as afterwards.
|
|
- **The query log and the dashboard say how far back the history goes.** `GET /api/queries`, `/api/stats` and `/api/stats/timeseries` each carry a `coverage` object: `available_since`, the first second the file can answer for, and `complete`, whether the window you asked for begins inside it. A period that starts before the query log does now says so instead of charting the missing part as zero — which is what a recreate, a retention pass or a fresh install would otherwise look like.
|
|
|
|
### Changed
|
|
|
|
- **`GET /api/queries` rows changed shape.** Each row gains `qclass`, `rcode`, `policy_action`, `policy_reason` and `route_kind`, and `block_reason` is gone: the reason a query was blocked is now one of a closed set of values rather than a formatted string, and the status column reads it from `policy_reason`. `blocked`, `cache_hit`, `upstream` and every other existing field are unchanged.
|
|
- **Upgrading resets your query history.** The `query_log` table gains the provenance columns below, and `querylog.db` is never migrated (it holds expendable log rows, so a schema change replaces the file instead of upgrading it). On the first start after the upgrade the old file is set aside as `querylog.db.schema-changed-<unix seconds>` and a fresh one is created. Nothing else is touched: `config.db` keeps your configuration and your diagnostics history. The recreate files a resolved `query_log.recreated` diagnostics entry naming the file that was kept and the timestamp the new history begins at, and a new `querylog_meta` table records that coverage start, so the dashboard can say "history is available from ..." instead of charting an empty range as zero. The set-aside file is a working SQLite database and can be deleted once you have decided you do not want it.
|
|
- **`logging.query_log_buffer_max` now accepts 1 to 37449, down from 1 to 1000000.** The queued entry carries every new provenance field by value and is about four times as wide as before — 1792 bytes against 432 — so the meaningful bound is bytes rather than entries. The ceiling is computed at compile time from the width of the entry so that the queue's worst case stays within 64 MiB, and it moves whenever that width does. The default of 10000 is unchanged and costs about 17 MiB. A configuration above the new ceiling is rejected at startup with the ceiling in the message.
|
|
- **Group and blocklist source names are now capped at 64 bytes.** Both are copied into every query-log row that mentions them, so an unbounded name was an unbounded cost per row. A longer name is rejected as `GroupNameTooLong` or `SourceNameTooLong`.
|
|
|
|
### Fixed
|
|
|
|
- **A UDP reply that has to be truncated keeps the answer's RCODE.** When an answer does not fit the client's UDP buffer, nxdns replaces it with an empty reply carrying the TC bit, which tells the client to retry over TCP. That replacement was always built as NOERROR, whatever the answer said — so an oversized NXDOMAIN reached the client as a success, and an EDNS extended RCODE above 15 lost the eight upper bits it needs an OPT record to carry. The truncated reply now carries the full twelve-bit code the answer had, split across the header and the reply's OPT record where the code needs it, and the query-log row records the code the client actually saw. The retry over TCP always returned the right RCODE; this was the UDP answer that preceded it.
|
|
|
|
## [0.0.8] - 2026-08-21
|
|
|
|
One constant, chosen from the 0.0.7 field numbers: the checkpoint cadence was the last first-order write cost on the Pi's SD card.
|
|
|
|
### Changed
|
|
|
|
- **The query log checkpoints its write-ahead log every 32 MiB instead of every 4 MiB.** Batching the writer in 0.0.7 took the deployed Pi from about 0.5 to 0.281 GiB of writes a day, and about 130 MiB of what is left is checkpoint writeback: SQLite's 1000-page default trips roughly every 40 minutes and rewrites the same hot index and interior pages into `querylog.db` each time. Every read-write connection to `querylog.db` now sets `wal_autocheckpoint` to 8192 pages, which stretches that to roughly five hours and cuts those in-place rewrites about eightfold, for an expected total near 190 MiB a day. The price is durability under power loss or a kernel panic. At `synchronous = NORMAL` a commit does not fsync, so the checkpoint is the only guaranteed durability boundary, and it now sits about five hours of query rows and upstream-history minutes back rather than 40 minutes. Kernel writeback normally makes the real loss far smaller than that, but nothing guarantees it. A process crash or a clean stop still loses nothing that was committed, and the database is never left inconsistent: recovery replays the longest valid prefix of the log. The `querylog.db-wal` file is expected to sit near 32 MiB rather than capped there, since a long-running reader can hold a checkpoint off and let it overshoot, and the daily retention pass still truncates it. `config.db` is unchanged.
|
|
|
|
## [0.0.7] - 2026-08-20
|
|
|
|
Operational failures get a page of their own, and the query log stops wearing out the disk it lives on: the deployed Pi was writing half a gigabyte a day to store two megabytes of query rows, one transaction per query. Both came out of running 0.0.6 on real hardware.
|
|
|
|
### Added
|
|
|
|
- **A diagnostics page.** Operational failures now land in one curated log instead of only journald: blocklist download failures, certificate reload failures, disk pressure, query-log writer and maintenance failures, upstream exchange and history failures, client tracking failures, listener and configuration problems at boot, and the query-log recreation an upgrade causes. One entry per failing subject — an entry opens on the first failure, counts repeats, and closes itself when the subject recovers; nothing needs dismissing. Each entry says what it means for the service and what to do about it. `GET /api/diagnostics` serves the log, `GET /api/health` reports the active counts and degrades while the diagnostics store itself cannot write, and `/metrics` gains `nxdns_diagnostics_active_warnings`, `nxdns_diagnostics_active_errors` and `nxdns_diagnostics_write_failures_total`. Resolved entries can be purged when you decide the history has served its purpose — one entry from its row or its detail page, or the whole resolved history at once with "Purge all resolved" (`DELETE /api/diagnostics/{id}` and `DELETE /api/diagnostics`). An entry that is still failing is the current state of the box, not history, so it has no purge action and the API answers 409.
|
|
|
|
### Changed
|
|
|
|
- **The query log commits once a minute instead of once a query.** The writer batched for 100 milliseconds, which at a household's query rate means almost every query got a transaction of its own — and a transaction costs the disk far more than the row it carries. On the deployed Pi that came to roughly 0.5 GiB of writes a day to store 2.3 MB of query rows, the kind of write volume that kills an SD card. The batch window is now `logging.query_log_flush_interval_s`: 60 seconds by default (the same minute Pi-hole's `DBinterval` defaults to, for the same reason), anything from 0 to 3600, editable on the settings page. Batches are still capped at 100 rows, so a burst is committed as soon as it fills one rather than waiting out the window, and the in-memory queue, its drop-oldest backpressure and retention are untouched. The price is two kinds of lag: a crash costs about one interval of query history — more if the writer was held back by a full disk or a slow write — and every query-log-backed view — the query-log page, the dashboard totals, the timeseries — is about one interval behind. The live page is not affected; it is fed before the queue. Set the key to `0` for the old write-immediately behavior.
|
|
|
|
### Fixed
|
|
|
|
- **Shutdown no longer races the last query rows to the disk.** The query-log writer was stopped by the same cancellation that stopped the DNS listeners, so whether the batch it was holding reached the database depended on which happened to land first, the cancellation or the queue closing. Shutdown now stops and joins the listeners and every other query producer first, then closes the queue, then waits for the writer to finish emptying it — the held batch and everything still queued get written. If free space is below the critical threshold and the disk monitor will not let that final write through, the rows are counted as dropped instead of holding the exit open indefinitely.
|
|
- **An upstream success rate no longer rounds up to 100.0% while failures stand.** One decimal place cannot hold 12,696 successes out of 12,698 attempts: it rounded to `100.0%`, so the row claimed perfect reliability next to a failure count of 2. Neither end of the scale is reachable by rounding any more — `100.0%` needs an actual absence of failures and `0.0%` an actual absence of successes, and a rate a hair off either end shows `99.9%` or `0.1%` instead.
|
|
- **A query log set aside by a schema change is no longer named `corrupt`.** Every recreate wrote the old file to `querylog.db.corrupt-<unix seconds>`, whatever sent it there — including the fingerprint mismatch an upgrade causes, where the file is a healthy database this build simply cannot read. The name is the only account of the reason that outlives the log line, so it read as an accusation and invited operators to delete an intact file. The name now says which of the four cases it hit: `querylog.db.corrupt-…`, `.not-a-database-…`, `.quick-check-failed-…` or `.schema-changed-…`. The 0.0.6 upgrade produces `schema-changed`. Nothing else about the recreate changed, and no existing aside file is renamed.
|
|
|
|
## [0.0.6] - 2026-08-17
|
|
|
|
The period picker now scopes the whole dashboard. The upstream table was the last widget that ignored it, and fixing that meant recording upstream outcomes over time instead of counting them since boot. Read the query-log note below before you upgrade.
|
|
|
|
### Added
|
|
|
|
- Four metrics for the new upstream-history recorder: `nxdns_upstream_history_flushes_total`, `nxdns_upstream_history_flush_failures_total`, `nxdns_upstream_history_rows_dropped_total` and the `nxdns_upstream_history_pending` gauge. While a flush to the database keeps failing, `GET /api/health` reports `degraded`; it recovers on the next flush that succeeds.
|
|
|
|
### Changed
|
|
|
|
- **Upstream health answers for the selected period.** The dashboard's upstream table used to print counters accumulated since process start beside a success rate taken over the last 32 exchanges, which is how "63 failures" and "100.0% success rate" ended up in the same row under a period picker that scoped nothing there. Every upstream outcome is now aggregated into its wall-clock minute and written to `querylog.db`, and `GET /api/upstream/health?period=…` serves the selected window: attempts, failures, success rate, and the last failure with its error name, all inside the period, from 31 days of history. A window with no attempts reports no success rate at all instead of a perfect one, and the table shows an em-dash. The in-memory health state that drives failover and backoff is unchanged, as are its `/metrics` series.
|
|
- **`GET /api/upstream/health` changed shape.** Gone from each upstream: `consecutive_failures`, `total_successes`, `total_failures`, the last-32 `success_rate`, `last_error` and `last_error_age_s`. Each upstream keeps `url`, `enabled` and `available` and gains a `period` object with the ranged numbers; the body gains `period`, `since`, `until` and a `complete` flag that says whether any outcome was known to be dropped inside the window. The removed counters are still exported by `/metrics` under their existing names. On the dashboard the "Right now" section is gone with them: the upstream table rejoined the ranged part of the page, and the disk card, the one live widget left, is titled "Storage now".
|
|
- **The query log is recreated on upgrade.** Recording upstream history added two tables to the `querylog.db` schema, and its fingerprint check refuses a database that does not match the shipped definition. On first start this version renames the existing `querylog.db` aside as `querylog.db.corrupt-<unix seconds>` in the data directory and creates a fresh one, so query history and stats restart empty. The renamed file is left in place rather than deleted, so removing it is your call. `config.db` is untouched: no configuration is lost.
|
|
|
|
## [0.0.5] - 2026-08-16
|
|
|
|
One rendering fix on the 0.0.4 feature, caught the day it shipped.
|
|
|
|
### Changed
|
|
|
|
- The query tables no longer repeat the *learned* tag on every row: in the live page and the query log a learned name is just muted, with the address still in the row's tooltip. The clients page keeps the tag, where it appears once per client and says something.
|
|
|
|
## [0.0.4] - 2026-08-16
|
|
|
|
The names learned in 0.0.3 now show up where queries do: the live page and the query log name each client instead of printing its address.
|
|
|
|
### Added
|
|
|
|
- **Client names in the query tables.** The live page and the query log show each query's client by name, with the same precedence as the clients page: a hand-typed name wins, else the learned name (muted, tagged *learned*), else the bare address. When a name replaces the address, the address stays readable as the row's tooltip. Devices that appear mid-stream show their address first and pick up their name within half a minute.
|
|
|
|
## [0.0.3] - 2026-08-15
|
|
|
|
Devices name themselves: the clients table asks the router over reverse DNS instead of waiting for the operator to type every name. The CI container gate also moved from workflow shell into a compiled, tested tool, which fixed a latent temp-directory bug shared with the release tool.
|
|
|
|
### Added
|
|
|
|
- **Client names learned over reverse DNS.** A client row that carries no hand-typed name gets one from the network: each tracker flush pass takes up to 16 unnamed rows, builds each address's reverse name, matches it against the declared `forward_zones`, and on a match sends one PTR query to that zone's resolver, storing the answer as a *learned* name. This requires a conditional forward zone covering the LAN's reverse space — for example `168.192.in-addr.arpa` pointed at the router; without one, nothing is sent anywhere. A hand-typed name always wins, learned names never appear in `nxdns export` and are never set by `nxdns import`, and each row refreshes once a day (an hour after a failure), so a rename can show stale for up to 24 hours. The API's `Client` object gains a `learned_name` field and the clients page shows it.
|
|
|
|
### Changed
|
|
|
|
- The container CI gate — image build, image-contents assertion against the packaged artifacts, and the startup/shutdown smoke test — moved from workflow shell into `tools/container_check.zig`, compiled and unit-tested by `zig build test` and runnable on a laptop against a local docker daemon. The health probe now runs under a real 60-second deadline (the shell loop's "30 seconds" could stretch past three minutes), and the gate's docker objects carry an ownership label so anything a dead runner leaks is discoverable. The version in CI is parsed from `build.zig.zon` through the zon grammar, once, instead of by two copies of a `sed` regex.
|
|
|
|
### Fixed
|
|
|
|
- The release tool's temporary-directory claim was not exclusive: the "create" it relied on succeeds on a directory that already exists, so a stale or concurrent directory could be silently adopted, written into, and deleted on exit. Both the release tool and the new container gate now claim their directories exclusively and retry on collision.
|
|
|
|
## [0.0.2] - 2026-08-14
|
|
|
|
Configuration can now be a file that every boot converges to, filtering gains regex rules and honors blocklist exception lines, and two refresh bugs that silently kept stale state are fixed. Note the three breaking changes below if you script against `nxdns import` or run with `--config`.
|
|
|
|
### Added
|
|
|
|
- **Declarative configuration for IaC.** `nxdns run --config=<file>` makes the file the sole source of configuration: every boot converges the database to it in one transaction, preserving blocklist downloads, compiled lists and client history, so an unchanged file costs zero downloads and zero writes. Bare `nxdns run` keeps the database (and the web UI) in charge, exactly as before. In file mode the web UI is read-only for configuration and says so; runtime actions (pause, blocklist refresh, certificate reload) stay live. `GET /api/settings` reports which authority governs the process.
|
|
- `nxdns import` now refuses a file whose application would delete configuration rows, names the tables and counts, and applies it only with the new `--allow-delete` flag. Additive and edit-in-place imports need no flag.
|
|
- **Regex rules.** Rules gain a third kind, `regex`, beside `exact` and `wildcard`, for per-group allow and block patterns such as `^ad[0-9]+-`. The engine is homegrown and linear-time by construction, so no pattern can make matching blow up; backreferences and lookaround do not exist, and a bad pattern is refused at insert time with the limit it hit. Matches appear in `/api/lookup` and the query log as `rule_allow_regex` / `rule_block_regex`. Regex still comes only from you: regex lines in downloaded lists stay counted and skipped.
|
|
- **Blocklist exception lines are honored.** An Adblock-Plus `@@||name^` line in a downloaded list now lifts that name — and its subdomains — out of what the attached lists block. Exceptions sit below every rule you wrote: a downloaded list can reopen only a hole another downloaded list dug, never override an operator decision. Each source reports how many it carried.
|
|
- **Browser-only lines are counted where you can see them.** Every source now reports how many of its lines nxdns skipped as syntax with no DNS meaning — cosmetic filters, `$`-modifier rules — beside the existing skipped-regex count. Both blocklist tables show the number and the UI explains the difference: a list whose skipped-unsupported count dwarfs its domain count is written for browser extensions, and its DNS or hosts variant will block more. Previously such a list compiled to almost nothing and looked clean.
|
|
|
|
### Changed
|
|
|
|
- **Breaking: `nxdns run --config <file>` changed meaning.** It used to seed the database once and then ignore the file; it now makes the file the authority on every boot, which deletes any configuration the file does not declare — including edits made through the web UI since the seed. Before upgrading a unit that carries `--config`: either drop the flag to keep the database in charge, or adopt file mode with the sequence in the upgrade guide. Order matters there: export the file with the NEW binary (stopped).
|
|
- **Breaking: 0.0.1 exports are refused by this version.** A 0.0.1 `nxdns export` writes both `.password = ""` and the stored `.password_hash`, and this version refuses a file that carries both. This bites any old export — an adoption file or a configuration backup fed to `nxdns import` alike. Fix an existing export by deleting its `.password = ""` line (keep the `.password_hash` line). Take fresh backups with the new binary.
|
|
- **Breaking: the offline password-change recipe changed.** Setting `.password = "new"` together with `.password_hash = ""` is now refused (empty `password_hash` is an explicit "disable authentication", and the two fields cannot both be present). To change the password in the file: set `.password` and delete the `.password_hash` line entirely.
|
|
- `nxdns import --force` is renamed `--allow-delete`.
|
|
- A fresh install no longer seeds from `/etc/nxdns/config.zon` by presence. Use `nxdns import` once, or run in file mode with `--config`.
|
|
- The admin UI's internals moved to TypeScript 7 and replaced Tailwind with StyleX and React Aria. The visible change is small: selects are real widgets with working keyboard focus; everything else renders as before.
|
|
- The `config.db` schema is a single baseline definition again; numbered migration steps start accumulating at v0.1.
|
|
|
|
### Fixed
|
|
|
|
- **A list switching a name between its exact and wildcard forms never took effect.** The compiled-list checksum hashed the exact and wildcard bodies as one unseparated byte stream, so a list carrying `a.example` and the same list carrying `*.a.example` produced the same digest, and the refresh kept the old compiled files. The checksum now separates the bodies. Every source recompiles once on its first refresh after the upgrade; no re-download of unchanged content is forced beyond the refresh's normal fetch.
|
|
- **A refresh could store stale skip counts.** When a refresh found the list content unchanged, it wrote the previously stored skip counters back to the database while showing the fresh ones in the UI, and the next restart reverted the numbers to the stale copy. All counters now persist from the fresh compile.
|
|
- An Adblock-Plus entry with embedded whitespace (`||good.example bad.example^`) compiled into an entry no query could ever match. Such lines are now counted as unsupported instead.
|
|
|
|
## [0.0.1] - 2026-08-09
|
|
|
|
First release. Everything below is new.
|
|
|
|
### Added
|
|
|
|
- **Forwarding DNS server.** UDP and TCP listeners with a wire-format parser and encoder written against RFC 1035 and EDNS(0), a bounded worker model, per-client rate limiting and a `pause` control that stops filtering without stopping resolution.
|
|
- **Encrypted upstreams.** DNS-over-HTTPS and DNS-over-TLS clients over a pool that tracks per-upstream health and fails over, with SNI and certificate verification driven by a per-upstream TLS name.
|
|
- **DoH and DoT endpoints.** nxdns also answers as an encrypted resolver, with a certificate store that reloads on disk changes and through the API, so renewals do not need a restart.
|
|
- **Blocklist filtering.** Subscriptions in hosts, plain-domain and Adblock-Plus-style formats, compiled into a compact matcher; per-group allow and block rules with wildcards; safe-search enforcement.
|
|
- **Per-client policy groups.** Clients are identified by address and assigned to groups, so the filtering a device gets depends on which device it is.
|
|
- **Local DNS.** Local A/AAAA/CNAME/PTR records and conditional forwarding of internal zones to another resolver.
|
|
- **Cache.** A bounded in-memory cache that respects upstream TTLs and expires entries rather than serving them stale.
|
|
- **Query log.** Queries land in SQLite under a retention policy in both rows and days, with disk-full self-protection that degrades instead of corrupting, and a live SSE stream of the same events.
|
|
- **Web UI and REST API.** A React single-page admin UI embedded in the binary, a REST API with a served OpenAPI document, session authentication, API rate limiting and Prometheus-style `/metrics`.
|
|
- **Configuration.** A ZON configuration file seeds the database on first boot; after that the database is the truth, and `nxdns export` / `nxdns import` move configuration in and out. `nxdns check` validates a file without starting.
|
|
- **CLI.** `run`, `check`, `export`, `import`, `version` and `help`.
|
|
- **Packaging.** A hardened systemd unit with a sysusers fragment, and a `FROM scratch` container image holding the binary, a CA bundle and the licence files, assembled by a builder stage pinned to `alpine:3.22` by digest. Nothing from Alpine ships in the published image except that CA bundle.
|
|
- **Releases.** Tags publish five assets — static musl tarballs for `x86_64-linux-musl` and `aarch64-linux-musl`, `IMAGE-DIGEST.txt` naming the multi-architecture container image by digest, `SHA256SUMS.txt` over those three, and `SHA256SUMS.txt.asc`, a detached signature over the checksum file. `zig build dist` and `zig build verify-dist` produce and check the same artifacts on a laptop.
|
|
- **Licensing.** EUPL-1.2, with a `THIRD-PARTY-NOTICES` file in every tarball and image assembled from a reviewed inventory of what the artifacts contain.
|
|
- **Documentation.** A Diátaxis split — tutorial, how-to, reference, explanation — with drift guards that fail the build when the reference pages fall behind the code.
|