Files
nxdns/docs/how-to/upgrade.md
T

474 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Upgrade nxdns
Replaces a running nxdns with a newer release without losing its
configuration. The database is migrated in place on the first start of the new
binary.
The normal path is to download the new release, verify it, and swap the binary.
Upgrading a build you made yourself is the last section of this page.
> Verification: the export, the migration behaviour and the `version`/`check`
> steps below were run on the machine that wrote this page, against a
> populated scratch data directory and with an explicit `--data-dir`, since
> that machine has no `/var/lib/nxdns`. The commands were not run exactly as
> printed — the page uses the defaults and placeholders a real operator would
> have (`/var/lib/nxdns`, `/some/backup`, a `target` host), and every block
> where the substitution matters, or which was not run at all, carries its own
> note. Nothing here was verified except where a note says so.
>
> Step 2 could not be run at all: no nxdns release is published yet, so every
> release URL and the `docker compose pull` on this page fail today. The URL
> shapes and the verification commands are covered by
> [Verify a release](verify-a-release.md), which says what was probed and how.
## Breaking change: `run --config` now means file authority
**Read this before upgrading if anything on your box passes `--config` to
`nxdns run`** — a systemd drop-in, a wrapper script, or a `command:` in a
compose file.
`run --config FILE` used to mean *seed once*: the file was read only while the
database was still empty, and ignored on every start after that. It now means
*the file is the configuration*: every start reconciles the database onto it.
For a box that was seeded once and then configured through the admin interface,
the first start after the upgrade converges the database back to that old seed
file. **Every change made through the UI since seeding is deleted.**
There are two ways out, and you pick before you restart:
- **Keep the database.** Drop the flag. `nxdns run` with no `--config` serves
the database exactly as it did before, and nothing reads a file. This is the
right answer if the UI is how you change things.
- **Adopt file mode cleanly.** Install the new binary, stop the service, export
the current database over the file path, check it, then start with the flag.
The first reconcile is then a no-op, because the file was rendered from the
database it governs. **Install the new binary first** — see the order trap
below. The full procedure is
[Adopt file mode](install-with-systemd.md#adopt-file-mode-on-a-box-that-is-already-running).
`nxdns check --config FILE` is unchanged: it graded that file before and it
grades that file now.
### The order trap: export with the new binary, not the old one
Take the export **after** you have replaced the binary, with the service
stopped. Exporting first — the instinctive order, and the one step 1 of this
page tells you to take for a backup — produces a file the new binary refuses.
A 0.0.1 `nxdns export` writes both fields:
```zon
.password = "",
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$…",
```
An empty `password_hash` used to mean "unset". It now means "disable
authentication", so it is a *present* value — and a file that carries both
fields states two different things about the password and is refused:
```
FAIL web.password: password and password_hash are both set; ambiguity in a security setting is refused
FAIL web.password: password is set to the empty string; omit the field to keep the stored password, or set password_hash = "" to disable authentication
nxdns run failed: PasswordAndHashBothSet
```
The old `nxdns check` passes that file, because the old binary agreed with the
old rule. So the failure lands at the first start after the upgrade, with the
resolver stopped and the unit refusing to retry it
(`RestartPreventExitStatus=2 64`). A new `nxdns export` writes
`.password = null` instead and has no such problem.
**If you already have an old export you want to adopt**, you do not need to
redo it. Delete the empty-password line and the file is valid:
```sh
sed -i '/^ \.password = "",$/d' /etc/nxdns/config.zon
nxdns check --config /etc/nxdns/config.zon
```
Keep the `.password_hash` line — that is the password, and deleting it as well
would leave the file saying nothing about authentication, which means "keep
whatever is stored" rather than anything you would notice.
The same trap has nothing to do with file mode as such: it is any 0.0.1 export
fed to the new binary, so it also applies to a restore through `nxdns import`.
Backups taken with 0.0.1 need that one line removed before they will load.
> Verified on this host, with one substitution stated: the repository has no
> 0.0.1 binary to hand, so the old export was **simulated** by taking a current
> `nxdns export` and rewriting `.password = null` to `.password = ""`, which is
> the one byte-level difference between the two formats. Against that file,
> `nxdns check --config` printed both FAIL lines above and exited 2, and
> `nxdns run --config` printed them and failed `PasswordAndHashBothSet`. After
> the `sed` above, `nxdns check --config` exited 0 with `OK: no problems found`
> and `nxdns import` of the same file exited 0, keeping the hash. The claim
> about what 0.0.1's `export` emitted is read from that release's source —
> `git show v0.0.1:src/config/export.zig` line 71 is `cfg.web.password = "";` —
> not from running that binary.
A database-mode install that never passed `--config` needs nothing. Under
Docker, a fresh database-mode install must either take the new compose file or
run `import` once — see
[Database mode in Docker](install-with-docker.md#database-mode-in-docker-instead).
Two smaller renames in the same release: `nxdns import --force` is now
`--allow-delete`, and it is required only when the file's diff would delete
rows rather than whenever the database is non-empty. `nxdns check` no longer
falls back to a default file path when there is no database; it reports the
absent database and names the two ways to get one.
There is no schema migration in this change.
## 1. Take an export first
There is no downgrade path, so the export is what you fall back to:
```sh
nxdns export --out /some/backup/nxdns-config.zon
```
`/some/backup` is a stand-in for a directory you keep backups in, and the
command relies on the default `--data-dir /var/lib/nxdns` that a systemd
install has.
This export is a fallback, not a file to deploy. If you are adopting file mode,
take a *second* export after the binary swap and use that one — an export
written by 0.0.1 carries a `.password = ""` line the new binary refuses, as
[the order trap](#the-order-trap-export-with-the-new-binary-not-the-old-one)
explains. The same line has to come out of this backup before the new binary
will import it.
> Verified on this host with both paths substituted, since it has neither
> `/var/lib/nxdns` nor `/some/backup`. `SCRATCH` below is a scratch directory,
> and its `data/` was populated beforehand with `nxdns import`:
>
> ```
> $ nxdns export --data-dir $SCRATCH/data --out $SCRATCH/nxdns-config.zon
> wrote /…/scratchpad/nxdns-config.zon
> $ stat -c '%a %n' $SCRATCH/nxdns-config.zon
> 600 /…/scratchpad/nxdns-config.zon
> $ grep password $SCRATCH/nxdns-config.zon
> .password = "",
> .password_hash = "$argon2id$v=19$m=19456,t=2,p=1$mCNEo…$i3DMz…",
> ```
>
> The shell umask was 022, so the 0600 is `export` setting it, not the umask.
> Only the two paths differ from the command above.
The file is written atomically at mode 0600 and carries
`web.password_hash`, so treat it as a secret. See
[Back up and restore](back-up-and-restore.md) for the full backup story. The
query log is deliberately not part of it.
## 2. Download and verify the new release
Read the release notes for the version you are moving to before you take it —
the `CHANGELOG.md` section for that version is the release body.
```sh
BASE=https://git.mial.net/mokhtar/nxdns
VERSION=$(curl -fsS -o /dev/null -w '%{redirect_url}' "$BASE/releases/latest" |
sed 's#.*/releases/tag/v##')
mkdir -p ~/nxdns-release && cd ~/nxdns-release
curl -fLO "$BASE/releases/download/v$VERSION/nxdns-$VERSION-x86_64-linux-musl.tar.gz"
curl -fLO "$BASE/releases/download/v$VERSION/SHA256SUMS.txt"
curl -fLO "$BASE/releases/download/v$VERSION/SHA256SUMS.txt.asc"
gpg --verify SHA256SUMS.txt.asc SHA256SUMS.txt
sha256sum -c --ignore-missing SHA256SUMS.txt
tar -xzf "nxdns-$VERSION-x86_64-linux-musl.tar.gz"
```
The first line asks the server which release is current, so this block does not
carry a version number that goes stale — Gitea redirects `releases/latest` to
the newest published release's tag page. To move to a particular version rather
than the newest, set `VERSION=<version>` yourself. Check it against what you are
running (`nxdns version`) before you download anything.
Take `aarch64-linux-musl` for a Raspberry Pi 5. Verify every time, not only on
the first install — an upgrade is a fresh download of a fresh artifact.
[Verify a release](verify-a-release.md) is the full procedure.
Under Docker there is nothing to download: step 3 pulls the image, and the
`IMAGE-DIGEST.txt` asset is what you verify instead.
> Not verified on this host: no release exists yet, so the `releases/latest`
> lookup returns 404 and leaves `VERSION` empty, and every `curl` below it is a
> 404 too. The lookup form was run against `gitea.com/gitea/tea` on Gitea
> `1.27.0+dev` and printed `0.15.1`.
## 3. Replace the binary
### systemd
Step 2 leaves the new binary in the extracted directory. Copy the one that
matches the host — `aarch64-linux-musl` for a Raspberry Pi 5:
```sh
scp "nxdns-$VERSION-x86_64-linux-musl/nxdns" target:/tmp/nxdns
```
> Not run on this host: `target` is a placeholder for the machine running
> nxdns, and this host has no such second machine to copy to. There is also no
> release to have extracted.
The tarball also carries `nxdns.service` and `nxdns.conf`. An upgrade does not
normally reinstall them, but compare them against what is on the target when
the release notes say the unit changed.
Then, as root on the target:
```sh
install -m 0755 /tmp/nxdns /usr/local/bin/nxdns
systemctl restart nxdns
journalctl -u nxdns -f
```
> Not run on this host: all three lines need root and an installed service.
> The migration half of what a restart does is checkable without either, and
> was — see the note under
> [What happens to the database](#what-happens-to-the-database). The swap of
> an older binary for a newer one on a live service was not reproduced here.
### Docker
```sh
NXDNS_VERSION=$VERSION docker compose -f deploy/docker/compose.yaml pull
NXDNS_VERSION=$VERSION docker compose -f deploy/docker/compose.yaml up -d
```
Compose recreates the container against the same `nxdns-data` volume. What
happens to the file in `etc-nxdns` depends on the `command:` in your compose
file: with the shipped `run --config=/etc/nxdns/config.zon` the file is the
configuration and the restart reconciles onto it; without it, the database in
the volume is the configuration and the file is read by nothing.
Set `NXDNS_VERSION` on both lines, or export it. Without it the compose file
falls back to `:latest`, and `pull` and `up` could then land on different
images if a release happens between them.
> Not run on this host: `pull` needs a published image, and there is none.
> What was run is `docker compose -f deploy/docker/compose.yaml config`, which
> resolves the variables without contacting a registry: `NXDNS_VERSION=0.0.1`
> gave `image: git.mial.net/mokhtar/nxdns:0.0.1`, and no variable at all gave
> `:latest`.
## 4. Confirm the upgrade
```sh
nxdns version
nxdns export --out /tmp/after-upgrade.zon
nxdns check --config /tmp/after-upgrade.zon
dig @127.0.0.1 example.com A +short
```
The restart in step 3 is what migrated the database, so by now the schema is
current and the service is answering. Confirming with `nxdns check` alone would
not work here, and the reason is worth knowing: `check` opens `config.db`
immutable so it can never write to it, and the migration you just performed is
sitting in `config.db-wal` waiting to be checkpointed. Rather than read around
the log and grade older settings, `check` reports it:
```
FAIL /var/lib/nxdns/config.db: uncheckpointed changes are waiting in /var/lib/nxdns/config.db-wal, and reading without writing would answer from the older settings in the main file; `nxdns run` applies them. A running nxdns normally holds this log, which is the usual reason to see this line.
```
`export` opens the database read/write and does see the log, so exporting and
then checking the export validates what is actually in force:
```
checking configuration file /tmp/after-upgrade.zon
OK upstreams[0] https://cloudflare-dns.com
OK: no problems found
```
> Verified on this host for the first three commands, with `--data-dir`
> pointing at a scratch data directory instead of `/var/lib/nxdns` — that path
> is the only difference from the blocks above:
>
> ```
> $ nxdns version
> nxdns <version> (unknown)
> zig 0.16.0
> ```
>
> Against a running server whose database had just taken a configuration write
> through the API, `nxdns check --data-dir` printed the uncheckpointed-log line
> above and exited 2, while `nxdns export --out` followed by
> `nxdns check --config` on the result exited 0 with `OK: no problems found`.
> Both were re-run for this revision. The `dig` line was not run in this round:
> nothing is listening on 127.0.0.1:53 here, and port 53 needs root.
## What happens to the database
Migrations run at startup, and also before `export` and `import`, so whichever
of those you run first performs the upgrade. `nxdns check` is the exception: it
opens the database immutable and never migrates, so on a database still one
version behind it reports the mismatch and exits 2 rather than fixing it:
```
FAIL /var/lib/nxdns/config.db: schema version 0, this nxdns expects 2; `nxdns run` migrates it, `check` will not
```
That line was reproduced here against a database stamped at version 0; the path
and the version numbers are what vary.
A fresh database is created at the current schema version; an older one is
stepped up to it. The log line names both versions:
```
info(migrations): config.db migrated from schema version 0 to 2
```
> Verified on this host: that exact line is what `nxdns import` printed when it
> created the scratch database used throughout this page. An empty data
> directory is schema version 0, which is why a first run reports a migration
> rather than nothing. The step from a populated older schema to 2 was not
> reproduced here — it needs a database written by an older binary, which this
> host does not have.
Rolling back is the case that has no answer. A database stamped by a newer
binary refuses to open, so an older binary against an upgraded data directory
fails to start:
```
warning(migrations): config.db is at schema version 99; this nxdns binary supports 2
nxdns run failed: SchemaTooNew
```
> Not reproduced on this host: the same missing ingredient as above, a
> database at a schema version this binary does not support. The two lines
> are the messages `src/storage/migrations.zig` emits, not a run captured
> here.
That run exits 1. Recovering means importing the export you took in step 1 into
a fresh data directory with the older binary.
### Rolling back from file mode
Putting an older binary back needs no unit edit. The old binary accepts
`run --config` — it just reads it as the old seed-once flag — and against a
database that already holds configuration it ignores the file entirely and
serves the last state the new binary reconciled. So the service comes back up
on the configuration it was running.
The consequence is worth stating plainly: **file edits stop applying.** The old
binary will not re-read the file, so every change made to `config.zon` after the
rollback does nothing at all, silently, until the newer binary is back. If you
have to stay on the old binary, use `nxdns import` to apply file changes, or drop
the flag so the invocation matches what the binary actually does.
The schema note above still governs: a database stamped by a newer binary
refuses to open, whatever mode either binary runs in.
## Changing settings, not the binary
How you change a setting depends on which authority the service runs under.
`nxdns run` in `ExecStart` means the database; `nxdns run --config FILE` means
the file. The start log names it either way:
```
info(nxdns): authority: database
info(nxdns): authority: file (/etc/nxdns/config.zon)
```
**In file mode**, edit the file, validate it, restart. The admin interface will
refuse the change with a 403 naming the file, so there is nothing to get wrong:
```sh
$EDITOR /etc/nxdns/config.zon
nxdns check --config /etc/nxdns/config.zon
systemctl restart nxdns
```
**In database mode**, change settings through the admin interface, through the
API, or with an exporteditimport cycle against a stopped server:
```sh
nxdns export --out config-backup.zon
$EDITOR config-backup.zon
systemctl stop nxdns
nxdns import config-backup.zon
systemctl start nxdns
```
`import` needs no flag to add rows or to edit them. It needs `--allow-delete`
only when applying the file would delete rows the database holds — including the
case where you renamed something, since changing a group's name or an upstream's
URL is a delete and an insert to the engine, not an edit. The refusal names the
tables and rolls back:
```
FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
import failed: DestructiveImport
```
Stop the server first either way. `import` rewrites configuration underneath a
process that read it at startup, and a running server picks up only some of it.
> Verified on this host against a populated scratch data directory, with
> `--data-dir` pointing at it — that path is the only difference from the blocks
> above:
>
> ```
> $ nxdns import $SCRATCH/etc/config.zon --data-dir $SCRATCH/dbmode
> imported /…/config.zon
> (exit 0)
> $ nxdns import $SCRATCH/etc/smaller.zon --data-dir $SCRATCH/dbmode
> FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
> import failed: DestructiveImport
> (exit 2)
> $ nxdns import $SCRATCH/etc/smaller.zon --data-dir $SCRATCH/dbmode --allow-delete
> imported /…/smaller.zon
> (exit 0)
> ```
>
> The first of those three is the additive case that needs no flag; the second
> file replaced the upstream, which is an identity change and therefore a
> delete. The `systemctl stop`/`start` lines need root and an installed service
> and were not run; `$EDITOR` is yours to run.
## Upgrading to a build of your own
If you are running something you built rather than a release, step 2 is a
build instead of a download:
```sh
(cd web && npm ci && npm run build)
VERSION=$(sed -n 's/^[[:space:]]*\.version[[:space:]]*=[[:space:]]*"\([^"]*\)".*/\1/p' build.zig.zon)
zig build dist -Dversion-string="$VERSION" -Dgit-commit="$(git rev-parse HEAD)" \
-Dweb-dist=web/dist -Doptimize=ReleaseSafe
```
Rebuild `web/dist` before the binary on every upgrade. The admin interface is
embedded at build time, and an old bundle against a new API is a broken
settings page. `dist` refuses the `web/dist-placeholder` default outright, so
the only way to ship a stale bundle is to leave an old `web/dist` in place.
The staged payload for each target is under
`zig-out/dist/stage/nxdns-<version>-<triple>/`, and step 3 continues from there
with that path in place of the extracted one. The version string has to equal
`.version` in `build.zig.zon``verify-dist` asserts it, so a made-up one
builds and then fails verification. What tells your build apart from the
published release of the same version is `-Dgit-commit`, which `nxdns version`
prints beside the version.
Under Docker, build the image and name it instead of pulling:
```sh
DOCKER_BUILDKIT=1 docker build -t nxdns -f deploy/docker/Dockerfile .
NXDNS_IMAGE=nxdns docker compose -f deploy/docker/compose.yaml up -d
```
See [Install with Docker](install-with-docker.md) for what that build needs.
> Verified on this host for the two build commands in this section — the
> `zig build dist` block above and the `docker build` here. `dist` was run to
> completion with the version read out of `build.zig.zon` and exited 0, and the
> image was built from the resulting `zig-out/dist` tree, also exiting 0. The
> `docker compose ... up -d` line was not run in this round — the run itself is
> covered in [Install with Docker](install-with-docker.md#3-run-it), where a
> host port had to be moved to do it. See
> [Install with systemd](install-with-systemd.md#build-from-source-instead) for
> the `dist` and `verify-dist` detail.