Files
nxdns/docs/how-to/upgrade.md
T
mokhtar cc23c97218
Gates / frontend (push) Successful in 1m34s
Gates / test (push) Successful in 2m3s
Gates / test-aarch64 (push) Failing after 3h13m33s
Gates / package (push) Successful in 5m20s
Gates / container (push) Successful in 15s
CI / gates (push) Failing after 6h30m45s
milestone 33: contract closure — samples, file-authority enumeration, dead code, bundle ceiling
2026-08-22 23:31:37 +02:00

367 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Upgrade nxdns
Replaces a running nxdns with a newer release without losing its configuration. The database is migrated in place on the first start of the new binary.
The normal path is to download the new release, verify it, and swap the binary. Upgrading a build you made yourself is the last section of this page.
> Verification: the export, the migration behaviour and the `version`/`check`
> steps below were run on the machine that wrote this page, against a
> populated scratch data directory and with an explicit `--data-dir`, since
> that machine has no `/var/lib/nxdns`. The commands were not run exactly as
> printed — the page uses the defaults and placeholders a real operator would
> have (`/var/lib/nxdns`, `/some/backup`, a `target` host), and every block
> where the substitution matters, or which was not run at all, carries its own
> note. Nothing here was verified except where a note says so.
>
> Step 2 could not be run at all: no nxdns release is published yet, so every
> release URL and the `docker compose pull` on this page fail today. The URL
> shapes and the verification commands are covered by
> [Verify a release](verify-a-release.md), which says what was probed and how.
## Breaking change: `run --config` now means file authority
**Read this before upgrading if anything on your box passes `--config` to `nxdns run`** — a systemd drop-in, a wrapper script, or a `command:` in a compose file.
`run --config FILE` used to mean *seed once*: the file was read only while the database was still empty, and ignored on every start after that. It now means *the file is the configuration*: every start reconciles the database onto it.
For a box that was seeded once and then configured through the admin interface, the first start after the upgrade converges the database back to that old seed file. **Every change made through the UI since seeding is deleted.**
There are two ways out, and you pick before you restart:
- **Keep the database.** Drop the flag. `nxdns run` with no `--config` serves the database exactly as it did before, and nothing reads a file. This is the right answer if the UI is how you change things.
- **Adopt file mode cleanly.** Install the new binary, stop the service, export the current database over the file path, check it, then start with the flag. The first reconcile is then a no-op, because the file was rendered from the database it governs. **Install the new binary first** — see the order trap below. The full procedure is [Adopt file mode](install-with-systemd.md#adopt-file-mode-on-a-box-that-is-already-running).
`nxdns check --config FILE` is unchanged: it graded that file before and it grades that file now.
### The order trap: export with the new binary, not the old one
Take the export **after** you have replaced the binary, with the service stopped. Exporting first — the instinctive order, and the one step 1 of this page tells you to take for a backup — produces a file the new binary refuses.
A 0.0.1 `nxdns export` writes both fields:
```zon
.password = "",
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$…",
```
An empty `password_hash` used to mean "unset". It now means "disable authentication", so it is a *present* value — and a file that carries both fields states two different things about the password and is refused:
```
FAIL web.password: password and password_hash are both set; ambiguity in a security setting is refused
FAIL web.password: password is set to the empty string; omit the field to keep the stored password, or set password_hash = "" to disable authentication
nxdns run failed: PasswordAndHashBothSet
```
The old `nxdns check` passes that file, because the old binary agreed with the old rule. So the failure lands at the first start after the upgrade, with the resolver stopped and the unit refusing to retry it (`RestartPreventExitStatus=2 64`). A new `nxdns export` writes `.password = null` instead and has no such problem.
**If you already have an old export you want to adopt**, you do not need to redo it. Delete the empty-password line and the file is valid:
```sh
sed -i '/^ \.password = "",$/d' /etc/nxdns/config.zon
nxdns check --config /etc/nxdns/config.zon
```
Keep the `.password_hash` line — that is the password, and deleting it as well would leave the file saying nothing about authentication, which means "keep whatever is stored" rather than anything you would notice.
The same trap has nothing to do with file mode as such: it is any 0.0.1 export fed to the new binary, so it also applies to a restore through `nxdns import`. Backups taken with 0.0.1 need that one line removed before they will load.
> Verified on this host, with one substitution stated: the repository has no
> 0.0.1 binary to hand, so the old export was **simulated** by taking a current
> `nxdns export` and rewriting `.password = null` to `.password = ""`, which is
> the one byte-level difference between the two formats. Against that file,
> `nxdns check --config` printed both FAIL lines above and exited 2, and
> `nxdns run --config` printed them and failed `PasswordAndHashBothSet`. After
> the `sed` above, `nxdns check --config` exited 0 with `OK: no problems found`
> and `nxdns import` of the same file exited 0, keeping the hash. The claim
> about what 0.0.1's `export` emitted is read from that release's source —
> `git show v0.0.1:src/config/export.zig` line 71 is `cfg.web.password = "";` —
> not from running that binary.
A database-mode install that never passed `--config` needs nothing. Under Docker, a fresh database-mode install must either take the new compose file or run `import` once — see [Database mode in Docker](install-with-docker.md#database-mode-in-docker-instead).
Two smaller renames in the same release: `nxdns import --force` is now `--allow-delete`, and it is required only when the file's diff would delete rows rather than whenever the database is non-empty. `nxdns check` no longer falls back to a default file path when there is no database; it reports the absent database and names the two ways to get one.
There is no schema migration in this change.
## 1. Take an export first
There is no downgrade path, so the export is what you fall back to:
```sh
nxdns export --out /some/backup/nxdns-config.zon
```
`/some/backup` is a stand-in for a directory you keep backups in, and the command relies on the default `--data-dir /var/lib/nxdns` that a systemd install has.
This export is a fallback, not a file to deploy. If you are adopting file mode, take a *second* export after the binary swap and use that one — an export written by 0.0.1 carries a `.password = ""` line the new binary refuses, as [the order trap](#the-order-trap-export-with-the-new-binary-not-the-old-one) explains. The same line has to come out of this backup before the new binary will import it.
> Verified on this host with both paths substituted, since it has neither
> `/var/lib/nxdns` nor `/some/backup`. `SCRATCH` below is a scratch directory,
> and its `data/` was populated beforehand with `nxdns import`:
>
> ```
> $ nxdns export --data-dir $SCRATCH/data --out $SCRATCH/nxdns-config.zon
> wrote /…/scratchpad/nxdns-config.zon
> $ stat -c '%a %n' $SCRATCH/nxdns-config.zon
> 600 /…/scratchpad/nxdns-config.zon
> $ grep password $SCRATCH/nxdns-config.zon
> .password = "",
> .password_hash = "$argon2id$v=19$m=19456,t=2,p=1$mCNEo…$i3DMz…",
> ```
>
> The shell umask was 022, so the 0600 is `export` setting it, not the umask.
> Only the two paths differ from the command above.
The file is written atomically at mode 0600 and carries `web.password_hash`, so treat it as a secret. See [Back up and restore](back-up-and-restore.md) for the full backup story. The query log is deliberately not part of it.
## 2. Download and verify the new release
Read the release notes for the version you are moving to before you take it — the `CHANGELOG.md` section for that version is the release body.
```sh
BASE=https://git.mial.net/mokhtar/nxdns
VERSION=$(curl -fsS -o /dev/null -w '%{redirect_url}' "$BASE/releases/latest" |
sed 's#.*/releases/tag/v##')
mkdir -p ~/nxdns-release && cd ~/nxdns-release
curl -fLO "$BASE/releases/download/v$VERSION/nxdns-$VERSION-x86_64-linux-musl.tar.gz"
curl -fLO "$BASE/releases/download/v$VERSION/SHA256SUMS.txt"
curl -fLO "$BASE/releases/download/v$VERSION/SHA256SUMS.txt.asc"
gpg --verify SHA256SUMS.txt.asc SHA256SUMS.txt
sha256sum -c --ignore-missing SHA256SUMS.txt
tar -xzf "nxdns-$VERSION-x86_64-linux-musl.tar.gz"
```
The first line asks the server which release is current, so this block does not carry a version number that goes stale — Gitea redirects `releases/latest` to the newest published release's tag page. To move to a particular version rather than the newest, set `VERSION=<version>` yourself. Check it against what you are running (`nxdns version`) before you download anything.
Take `aarch64-linux-musl` for a Raspberry Pi 5. Verify every time, not only on the first install — an upgrade is a fresh download of a fresh artifact. [Verify a release](verify-a-release.md) is the full procedure.
Under Docker there is nothing to download: step 3 pulls the image, and the `IMAGE-DIGEST.txt` asset is what you verify instead.
> Not verified on this host: no release exists yet, so the `releases/latest`
> lookup returns 404 and leaves `VERSION` empty, and every `curl` below it is a
> 404 too. The lookup form was run against `gitea.com/gitea/tea` on Gitea
> `1.27.0+dev` and printed `0.15.1`.
## 3. Replace the binary
### systemd
Step 2 leaves the new binary in the extracted directory. Copy the one that matches the host — `aarch64-linux-musl` for a Raspberry Pi 5:
```sh
scp "nxdns-$VERSION-x86_64-linux-musl/nxdns" target:/tmp/nxdns
```
> Not run on this host: `target` is a placeholder for the machine running
> nxdns, and this host has no such second machine to copy to. There is also no
> release to have extracted.
The tarball also carries `nxdns.service` and `nxdns.conf`. An upgrade does not normally reinstall them, but compare them against what is on the target when the release notes say the unit changed.
Then, as root on the target:
```sh
install -m 0755 /tmp/nxdns /usr/local/bin/nxdns
systemctl restart nxdns
journalctl -u nxdns -f
```
> Not run on this host: all three lines need root and an installed service.
> The migration half of what a restart does is checkable without either, and
> was — see the note under
> [What happens to the database](#what-happens-to-the-database). The swap of
> an older binary for a newer one on a live service was not reproduced here.
### Docker
```sh
NXDNS_VERSION=$VERSION docker compose -f deploy/docker/compose.yaml pull
NXDNS_VERSION=$VERSION docker compose -f deploy/docker/compose.yaml up -d
```
Compose recreates the container against the same `nxdns-data` volume. What happens to the file in `etc-nxdns` depends on the `command:` in your compose file: with the shipped `run --config=/etc/nxdns/config.zon` the file is the configuration and the restart reconciles onto it; without it, the database in the volume is the configuration and the file is read by nothing.
Set `NXDNS_VERSION` on both lines, or export it. Without it the compose file falls back to `:latest`, and `pull` and `up` could then land on different images if a release happens between them.
> Not run on this host: `pull` needs a published image, and there is none.
> What was run is `docker compose -f deploy/docker/compose.yaml config`, which
> resolves the variables without contacting a registry: `NXDNS_VERSION=0.0.1`
> gave `image: git.mial.net/mokhtar/nxdns:0.0.1`, and no variable at all gave
> `:latest`.
## 4. Confirm the upgrade
```sh
nxdns version
nxdns export --out /tmp/after-upgrade.zon
nxdns check --config /tmp/after-upgrade.zon
dig @127.0.0.1 example.com A +short
```
The restart in step 3 is what migrated the database, so by now the schema is current and the service is answering. Confirming with `nxdns check` alone would not work here, and the reason is worth knowing: `check` opens `config.db` immutable so it can never write to it, and the migration you just performed is sitting in `config.db-wal` waiting to be checkpointed. Rather than read around the log and grade older settings, `check` reports it:
```
FAIL /var/lib/nxdns/config.db: uncheckpointed changes are waiting in /var/lib/nxdns/config.db-wal, and reading without writing would answer from the older settings in the main file; `nxdns run` applies them. A running nxdns normally holds this log, which is the usual reason to see this line.
```
`export` opens the database read/write and does see the log, so exporting and then checking the export validates what is actually in force:
```
checking configuration file /tmp/after-upgrade.zon
OK upstreams[0] https://cloudflare-dns.com
OK: no problems found
```
> Verified on this host for the first three commands, with `--data-dir`
> pointing at a scratch data directory instead of `/var/lib/nxdns` — that path
> is the only difference from the blocks above:
>
> ```
> $ nxdns version
> nxdns <version> (unknown)
> zig 0.16.0
> ```
>
> Against a running server whose database had just taken a configuration write
> through the API, `nxdns check --data-dir` printed the uncheckpointed-log line
> above and exited 2, while `nxdns export --out` followed by
> `nxdns check --config` on the result exited 0 with `OK: no problems found`.
> Both were re-run for this revision. The `dig` line was not run in this round:
> nothing is listening on 127.0.0.1:53 here, and port 53 needs root.
## What happens to the database
Migrations run at startup, and also before `export` and `import`, so whichever of those you run first performs the upgrade. `nxdns check` is the exception: it opens the database immutable and never migrates, so on a database still one version behind it reports the mismatch and exits 2 rather than fixing it:
```
FAIL /var/lib/nxdns/config.db: schema version 0, this nxdns expects 1; `nxdns run` migrates it, `check` will not
```
That line was reproduced here against a database stamped at version 0; the path and the version numbers are what vary.
A fresh database is created at the current schema version; an older one is stepped up to it. The log line names both versions:
```
info(migrations): config.db migrated from schema version 0 to 1
```
> Verified on this host: that exact line is what `nxdns import` printed when it
> created the scratch database used throughout this page. An empty data
> directory is schema version 0, which is why a first run reports a migration
> rather than nothing. Version 1 is the only schema nxdns has published, so an
> upgrade from a populated older one is not a case that exists yet.
Rolling back is the case that has no answer. A database stamped by a newer binary refuses to open, so an older binary against an upgraded data directory fails to start:
```
warning(migrations): config.db is at schema version 99; this nxdns binary supports 1
nxdns run failed: SchemaTooNew
```
> Reproduced on this host, with one substitution: no binary from the future was
> available, so the scratch database's `schema_version` row was set to 99 by
> hand and `nxdns run` was pointed at it. The two lines above are that run's
> output.
That run exits 1. Recovering means importing the export you took in step 1 into a fresh data directory with the older binary.
### Rolling back from file mode
Putting an older binary back needs no unit edit. The old binary accepts `run --config` — it just reads it as the old seed-once flag — and against a database that already holds configuration it ignores the file entirely and serves the last state the new binary reconciled. So the service comes back up on the configuration it was running.
The consequence is worth stating plainly: **file edits stop applying.** The old binary will not re-read the file, so every change made to `config.zon` after the rollback does nothing at all, silently, until the newer binary is back. If you have to stay on the old binary, use `nxdns import` to apply file changes, or drop the flag so the invocation matches what the binary actually does.
The schema note above still governs: a database stamped by a newer binary refuses to open, whatever mode either binary runs in.
## Changing settings, not the binary
How you change a setting depends on which authority the service runs under. `nxdns run` in `ExecStart` means the database; `nxdns run --config FILE` means the file. The start log names it either way:
```
info(nxdns): authority: database
info(nxdns): authority: file (/etc/nxdns/config.zon)
```
**In file mode**, edit the file, validate it, restart. The admin interface will refuse the change with a 403 naming the file, so there is nothing to get wrong:
```sh
$EDITOR /etc/nxdns/config.zon
nxdns check --config /etc/nxdns/config.zon
systemctl restart nxdns
```
**In database mode**, change settings through the admin interface, through the API, or with an exporteditimport cycle against a stopped server:
```sh
nxdns export --out config-backup.zon
$EDITOR config-backup.zon
systemctl stop nxdns
nxdns import config-backup.zon
systemctl start nxdns
```
`import` needs no flag to add rows or to edit them. It needs `--allow-delete` only when applying the file would delete rows the database holds — including the case where you renamed something, since changing a group's name or an upstream's URL is a delete and an insert to the engine, not an edit. The refusal names the tables and rolls back:
```
FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
import failed: DestructiveImport
```
Stop the server first either way. `import` rewrites configuration underneath a process that read it at startup, and a running server picks up only some of it.
> Verified on this host against a populated scratch data directory, with
> `--data-dir` pointing at it — that path is the only difference from the blocks
> above:
>
> ```
> $ nxdns import $SCRATCH/etc/config.zon --data-dir $SCRATCH/dbmode
> imported /…/config.zon
> (exit 0)
> $ nxdns import $SCRATCH/etc/smaller.zon --data-dir $SCRATCH/dbmode
> FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
> import failed: DestructiveImport
> (exit 2)
> $ nxdns import $SCRATCH/etc/smaller.zon --data-dir $SCRATCH/dbmode --allow-delete
> imported /…/smaller.zon
> (exit 0)
> ```
>
> The first of those three is the additive case that needs no flag; the second
> file replaced the upstream, which is an identity change and therefore a
> delete. The `systemctl stop`/`start` lines need root and an installed service
> and were not run; `$EDITOR` is yours to run.
## Upgrading to a build of your own
If you are running something you built rather than a release, step 2 is a build instead of a download:
```sh
(cd admin && npm ci && npm run build)
VERSION=$(sed -n 's/^[[:space:]]*\.version[[:space:]]*=[[:space:]]*"\([^"]*\)".*/\1/p' build.zig.zon)
zig build dist -Dversion-string="$VERSION" -Dgit-commit="$(git rev-parse HEAD)" \
-Dadmin-dist=admin/dist -Doptimize=ReleaseSafe
```
Rebuild `admin/dist` before the binary on every upgrade. The admin interface is embedded at build time, and an old bundle against a new API is a broken System page. `dist` refuses the `admin/dist-placeholder` default outright, so the only way to ship a stale bundle is to leave an old `admin/dist` in place.
The staged payload for each target is under `zig-out/dist/stage/nxdns-<version>-<triple>/`, and step 3 continues from there with that path in place of the extracted one. The version string has to equal `.version` in `build.zig.zon``verify-dist` asserts it, so a made-up one builds and then fails verification. What tells your build apart from the published release of the same version is `-Dgit-commit`, which `nxdns version` prints beside the version.
Under Docker, build the image and name it instead of pulling:
```sh
DOCKER_BUILDKIT=1 docker build -t nxdns -f deploy/docker/Dockerfile .
NXDNS_IMAGE=nxdns docker compose -f deploy/docker/compose.yaml up -d
```
See [Install with Docker](install-with-docker.md) for what that build needs.
> Verified on this host for the two build commands in this section — the
> `zig build dist` block above and the `docker build` here. `dist` was run to
> completion with the version read out of `build.zig.zon` and exited 0, and the
> image was built from the resulting `zig-out/dist` tree, also exiting 0. The
> `docker compose ... up -d` line was not run in this round — the run itself is
> covered in [Install with Docker](install-with-docker.md#3-run-it), where a
> host port had to be moved to do it. See
> [Install with systemd](install-with-systemd.md#build-from-source-instead) for
> the `dist` and `verify-dist` detail.