milestone 20: declarative configuration for iac
Gates / test-aarch64 (push) Successful in 6m45s
Gates / frontend (push) Successful in 51s
Gates / test (push) Successful in 1m37s
Gates / container (push) Failing after 7m31s
Gates / package (push) Failing after 15m14s
CI / gates (push) Failing after 24m27s
Gates / test-aarch64 (push) Successful in 6m45s
Gates / frontend (push) Successful in 51s
Gates / test (push) Successful in 1m37s
Gates / container (push) Failing after 7m31s
Gates / package (push) Failing after 15m14s
CI / gates (push) Failing after 24m27s
This commit is contained in:
@@ -1,9 +1,22 @@
|
||||
# Back up and restore
|
||||
|
||||
The configuration database is the only state worth keeping. `nxdns export`
|
||||
writes it out as a ZON file and `nxdns import` writes one back. The query log is
|
||||
deliberately not part of a backup: it is expendable history, and if it is
|
||||
missing it gets recreated empty.
|
||||
The configuration is the only state worth keeping. `nxdns export` writes it out
|
||||
as a ZON file and `nxdns import` writes one back. The query log is deliberately
|
||||
not part of a backup: it is expendable history, and if it is missing it gets
|
||||
recreated empty.
|
||||
|
||||
**Which authority the service runs under decides what the backup *is*.** Check
|
||||
the start log:
|
||||
|
||||
- `authority: database` — `config.db` holds the configuration. Back it up with
|
||||
`nxdns export`, and restore with `nxdns import` or by replacing the database
|
||||
file.
|
||||
- `authority: file (<path>)` — that file holds the configuration, and it is
|
||||
already a text file you can keep in git. **The file is the backup.** Restoring
|
||||
means putting the file back and restarting; the database rebuilds itself from
|
||||
it. `config.db` is a cache of the file in this mode, not the thing to preserve.
|
||||
|
||||
The rest of this page covers database mode unless it says otherwise.
|
||||
|
||||
The commands below use the scratch lab from
|
||||
[enable DoH and DoT](enable-doh-and-dot.md), data directory
|
||||
@@ -40,13 +53,16 @@ grep password /tmp/nxdns-lab/backup.zon
|
||||
```
|
||||
|
||||
```
|
||||
.password = "",
|
||||
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$Gh+zg9xke6BqSVOiouRbqG50+Bs8ZGcXA6oKgs7lrKg$crTNMu5OI8yKBkp31r4+Y1OUQmLiAlH/qvsIxjBQRq4",
|
||||
.password = null,
|
||||
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$hOjjnTrTZ6kU8XrwQuJ7ZFC3B5LumaB4hRe7kBbzJ6Q$YsC2rCTDKv96eEwPhs+D6vbDliLogEZph8EkagKSuy8",
|
||||
```
|
||||
|
||||
Treat backups as secrets. `.password` is always exported as `""` — the
|
||||
Treat backups as secrets. `.password` is always exported as `null` — the
|
||||
plaintext is never stored anywhere — so the file re-imports without anyone
|
||||
knowing the password.
|
||||
knowing the password. `null` and `""` are different statements here: `null`
|
||||
means the file says nothing about the password, while an empty `password_hash`
|
||||
would disable authentication. See
|
||||
[Password and hash](../reference/configuration.md#password-and-hash).
|
||||
|
||||
Without `--out` the export goes to stdout, where the file mode is your
|
||||
redirect's problem:
|
||||
@@ -57,7 +73,7 @@ nxdns export --data-dir /tmp/nxdns-lab/data | head -10
|
||||
|
||||
```
|
||||
// nxdns configuration
|
||||
// generated by `nxdns export` — the database is the source of truth
|
||||
// generated by `nxdns export` from the running configuration
|
||||
.{
|
||||
.upstream = .{
|
||||
.attempt_timeout_ms = 2500,
|
||||
@@ -75,8 +91,8 @@ exports taken minutes apart differ.
|
||||
|
||||
## Restore onto a fresh data directory
|
||||
|
||||
This is the normal restore: new machine, new disk, empty data directory. No
|
||||
`--force`, because there is nothing to overwrite.
|
||||
This is the normal restore: new machine, new disk, empty data directory.
|
||||
Nothing to overwrite, so nothing to authorise.
|
||||
|
||||
```sh
|
||||
nxdns import /tmp/nxdns-lab/backup.zon --data-dir /tmp/nxdns-lab/data-restored
|
||||
@@ -92,64 +108,95 @@ before writing to it.
|
||||
|
||||
## Restore over an existing database
|
||||
|
||||
`import` refuses a database that already holds configuration, so a plain
|
||||
`import` can never clobber a configured server by accident:
|
||||
|
||||
```sh
|
||||
nxdns import /tmp/nxdns-lab/backup.zon --data-dir /tmp/nxdns-lab/data
|
||||
```
|
||||
|
||||
```
|
||||
import failed: DatabaseNotEmpty
|
||||
```
|
||||
|
||||
That exits 2. Say `--force` when replacing is what you mean:
|
||||
|
||||
```sh
|
||||
nxdns import /tmp/nxdns-lab/backup.zon --force --data-dir /tmp/nxdns-lab/data
|
||||
```
|
||||
|
||||
```
|
||||
imported /tmp/nxdns-lab/backup.zon
|
||||
```
|
||||
|
||||
Stop the server first. `import` replaces the whole configuration underneath a
|
||||
process that has already read it, and a running server will not notice.
|
||||
|
||||
What `--force` does to the client list is worth knowing before you restore an
|
||||
old backup. A client the backup does not name is removed, and its first-seen
|
||||
and last-seen go with it — restoring a month-old file drops the devices you
|
||||
named since. Devices the server discovered from traffic are kept, and a device
|
||||
the backup does name keeps the first-seen and last-seen the database already
|
||||
held, so a restore does not restamp your whole network as newly arrived.
|
||||
|
||||
On a real install that is the systemd unit:
|
||||
**Stop the server first.** `import` rewrites configuration underneath a process
|
||||
that read it at startup, and a running server picks up only part of it —
|
||||
filtering follows the new rows at the next reload, while upstreams, listeners
|
||||
and settings stay at their boot values until a restart.
|
||||
|
||||
```sh
|
||||
systemctl stop nxdns
|
||||
nxdns import /var/backups/nxdns-config.zon --force
|
||||
nxdns import /var/backups/nxdns-config.zon
|
||||
systemctl start nxdns
|
||||
```
|
||||
|
||||
**Not verified on this host.** Those three lines are the only commands on this
|
||||
page that were not run: this machine has no installed nxdns systemd unit
|
||||
(`systemctl status nxdns` answers `Unit nxdns.service could not be found.`) and
|
||||
`systemctl stop`/`start` need root. The lab equivalent below was run, and it
|
||||
exercises the same stop-import-start sequence. In the lab the server is a
|
||||
foreground `nxdns run`, so stopping it is Ctrl-C in its own terminal:
|
||||
`import` converges the database onto the file. Rows the file still names are
|
||||
matched and updated in place; rows it no longer names are deleted. That last
|
||||
part is what a restore of an old backup does to everything added since, so it
|
||||
takes a flag:
|
||||
|
||||
```sh
|
||||
nxdns export --data-dir /tmp/nxdns-lab/data --out /tmp/nxdns-lab/pre-restore.zon
|
||||
# Ctrl-C the `nxdns run` terminal, or `kill` its pid from another shell
|
||||
nxdns import /tmp/nxdns-lab/pre-restore.zon --force --data-dir /tmp/nxdns-lab/data
|
||||
nxdns run --data-dir /tmp/nxdns-lab/data --config /tmp/nxdns-lab/etc/config.zon
|
||||
nxdns import /tmp/nxdns-lab/old-backup.zon --data-dir /tmp/nxdns-lab/data
|
||||
```
|
||||
|
||||
```
|
||||
wrote /tmp/nxdns-lab/pre-restore.zon
|
||||
imported /tmp/nxdns-lab/pre-restore.zon
|
||||
FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
|
||||
import failed: DestructiveImport
|
||||
```
|
||||
|
||||
That exits 2 and rolls the transaction back, so nothing is half-applied. The
|
||||
message names every table that would lose rows, which is usually enough to tell
|
||||
an intended restore from the wrong file. Say `--allow-delete` when deleting is
|
||||
what you mean:
|
||||
|
||||
```sh
|
||||
nxdns import /tmp/nxdns-lab/old-backup.zon --allow-delete --data-dir /tmp/nxdns-lab/data
|
||||
```
|
||||
|
||||
```
|
||||
imported /tmp/nxdns-lab/old-backup.zon
|
||||
```
|
||||
|
||||
A restore that only puts back what is already there needs no flag at all, and
|
||||
neither does one that only adds rows. The flag is about deletion specifically —
|
||||
including deletion in disguise: renaming a group, or correcting a typo in an
|
||||
upstream URL, changes the row's identity, so the engine sees one row gone and
|
||||
one arrived.
|
||||
|
||||
What survives a restore is worth knowing before you take an old backup out of
|
||||
the drawer. Blocklist download state is kept for every source whose URL the file
|
||||
still names — checksum, counters, compiled files, and the row id they are keyed
|
||||
by — so restoring does not cost a re-download. Devices the server discovered
|
||||
from traffic are kept whole, and one the backup names keeps the first-seen and
|
||||
last-seen the database already held, so a restore never restamps your network as
|
||||
newly arrived. What does go is a client the backup does not name and that was
|
||||
named by hand: that row is declarative, and it is deleted with the rest.
|
||||
|
||||
> The three `systemctl` lines are the only commands on this page that were not
|
||||
> run: this machine has no installed nxdns unit (`systemctl status nxdns`
|
||||
> answers `Unit nxdns.service could not be found.`) and `systemctl stop`/`start`
|
||||
> need root. The `nxdns import` between them is the same command the lab blocks
|
||||
> above run, which were executed here as written.
|
||||
|
||||
## Restoring the database file itself
|
||||
|
||||
Copying `config.db` back into place works too, and it is the fastest restore on
|
||||
a machine that still has one. Two rules, and the second one is where restores go
|
||||
wrong:
|
||||
|
||||
1. **Take the copy from a stopped instance**, or use `nxdns export` instead. A
|
||||
copy taken while nxdns is running catches the main file without the changes
|
||||
sitting in its write-ahead log.
|
||||
2. **Delete any stale `config.db-wal` and `config.db-shm` beside the file you
|
||||
restore.** SQLite silently discards a write-ahead log that does not match the
|
||||
database it sits next to. It does not warn, and it does not fail — it just
|
||||
answers from the main file, so a restore quietly loses its own tail.
|
||||
|
||||
```sh
|
||||
systemctl stop nxdns
|
||||
rm -f /var/lib/nxdns/config.db-wal /var/lib/nxdns/config.db-shm
|
||||
cp /var/backups/config.db /var/lib/nxdns/config.db
|
||||
chown nxdns:nxdns /var/lib/nxdns/config.db
|
||||
chmod 0600 /var/lib/nxdns/config.db
|
||||
systemctl start nxdns
|
||||
```
|
||||
|
||||
> Not verified on this host: these need root, an installed unit and
|
||||
> `/var/lib/nxdns`, none of which exist here.
|
||||
|
||||
None of this applies in file mode. There `config.db` is derived state — restore
|
||||
the configuration file and start the service, and the first reconcile rebuilds
|
||||
the database from it.
|
||||
|
||||
## Verify a backup
|
||||
|
||||
The round trip is byte-stable: exporting, importing and exporting again gives
|
||||
@@ -200,5 +247,6 @@ startup and before `export` and `import`; nothing walks them back, and `nxdns
|
||||
check` does not run them at all. Take an export before installing a new binary —
|
||||
see [upgrade](upgrade.md).
|
||||
|
||||
Every command on this page was executed on this host as written, except the
|
||||
`systemctl` block marked **Not verified on this host** above.
|
||||
Every `nxdns` command on this page was executed on this host as written. The
|
||||
`systemctl`, `cp`, `chown` and `chmod` lines were not: they need root and an
|
||||
installed unit, and both blocks holding them say so.
|
||||
|
||||
@@ -62,10 +62,12 @@ to:
|
||||
}
|
||||
```
|
||||
|
||||
Write that to `/tmp/nxdns-lab/etc/config.zon`. The configuration file seeds an
|
||||
empty database and is then ignored; to change these settings on a server that
|
||||
already has a database, edit them through the API or through
|
||||
`export`/`import` — see [the configuration model](../explanation/configuration-model.md).
|
||||
Write that to `/tmp/nxdns-lab/etc/config.zon`. The lab runs
|
||||
`nxdns run --config`, which makes that file the configuration: every start
|
||||
reconciles the database onto it, so editing the file and restarting is how these
|
||||
settings change here. A real install may instead run bare `nxdns run` and keep
|
||||
the configuration in the database — see
|
||||
[the configuration model](../explanation/configuration-model.md).
|
||||
|
||||
## 3. Check the files before starting
|
||||
|
||||
|
||||
@@ -10,16 +10,29 @@ supported and is the last section of this page.
|
||||
For what each configuration field means, see
|
||||
[the configuration reference](../reference/configuration.md).
|
||||
|
||||
> Verification: the seed-file failure modes, the run and the two checks in
|
||||
> step 3 were run on the machine that wrote this page, against an image built
|
||||
> from this checkout rather than pulled from the registry — no release is
|
||||
> Verification: the failure modes, the run and the two checks in step 3 were run
|
||||
> on the machine that wrote an earlier revision of this page, against an image
|
||||
> built from this checkout rather than pulled from the registry — no release is
|
||||
> published yet, so nothing on this page could be run against a pulled image,
|
||||
> and step 1 could not be run at all. One command was run in altered form:
|
||||
> host port 8080 was occupied here, so the run and the verification commands
|
||||
> in step 3 were executed with the host side of the port mappings moved to
|
||||
> 25353 and 28088 rather than the 53 and 8080 printed below. The container
|
||||
> side was unchanged. See the note in step 3. The `chown` to uid 65532 needs
|
||||
> root and was not run.
|
||||
> and step 1 could not be run at all. One command was run in altered form: host
|
||||
> port 8080 was occupied there, so the run and the verification commands in
|
||||
> step 3 were executed with the host side of the port mappings moved to 25353
|
||||
> and 28088 rather than the 53 and 8080 printed below. The container side was
|
||||
> unchanged. See the note in step 3. The `chown` to uid 65532 needs root and was
|
||||
> not run.
|
||||
>
|
||||
> **Not re-run for the file-mode revision.** The compose file now ships
|
||||
> `command: ["run", "--config=/etc/nxdns/config.zon"]`, and no container was
|
||||
> started against that command on this host: staging a release image needs
|
||||
> `zig build dist`, which refuses to run while `web/dist` is stale, and the web
|
||||
> bundle was being rebuilt by other work in the same tree at the time. What was
|
||||
> checked instead is `docker compose -f deploy/docker/compose.yaml config`,
|
||||
> which resolves the file without contacting a registry and prints the `command`
|
||||
> and the `:ro` bind mount as written, and the same
|
||||
> `run --config=<file>` invocation driven directly against the locally built
|
||||
> binary: it printed the reconcile summary, `authority: file (<path>)`, and
|
||||
> `no changes` on the second start. The log lines quoted in step 3 come from
|
||||
> that run, with the paths and ports the container uses.
|
||||
|
||||
## 1. Pull and verify the image
|
||||
|
||||
@@ -65,7 +78,7 @@ does and does not prove.
|
||||
> `docker buildx imagetools inspect --format` shape was run here against
|
||||
> `alpine:3.22` on Docker Hub and printed that image's index digest.
|
||||
|
||||
## 2. Get the compose file and write the seed configuration
|
||||
## 2. Get the compose file and write the configuration
|
||||
|
||||
Every path on this page is relative to a checkout of the repository, because
|
||||
that is how it was verified. Running the published image needs no checkout,
|
||||
@@ -82,7 +95,7 @@ curl -fLO "$BASE/raw/tag/v$VERSION/deploy/docker/compose.yaml"
|
||||
> Gitea `1.27.0+dev` and returned the file with a 200.
|
||||
|
||||
Compose bind-mounts `deploy/docker/etc-nxdns` read-only at `/etc/nxdns`. Create
|
||||
it and put the seed file in it:
|
||||
it and put the configuration in it:
|
||||
|
||||
```sh
|
||||
mkdir -p deploy/docker/etc-nxdns
|
||||
@@ -100,15 +113,25 @@ upstream:
|
||||
}
|
||||
```
|
||||
|
||||
Without that file the container exits with code 2 on a fresh volume: an empty
|
||||
database has nothing to forward to. The log is
|
||||
`no configuration file at '/etc/nxdns/config.zon'; using the database as it is`
|
||||
followed by `nxdns run failed: NoUsableUpstreams`.
|
||||
**The compose file ships file mode**, with
|
||||
`command: ["run", "--config=/etc/nxdns/config.zon"]`. That file is the
|
||||
configuration: the container reconciles its database onto it at every start, and
|
||||
the admin interface answers 403 to configuration edits. To change anything, edit
|
||||
the file and restart the container. It also means a fresh or recreated
|
||||
`nxdns-data` volume rebuilds itself from the mounted file with no extra step.
|
||||
|
||||
The file is therefore required, and its absence is a hard failure rather than a
|
||||
start with defaults:
|
||||
|
||||
```
|
||||
FAIL /etc/nxdns/config.zon: no such file
|
||||
nxdns run failed: ManagedConfigUnreadable
|
||||
run `nxdns check` to see the configuration in full
|
||||
```
|
||||
|
||||
A file that is present but rejected is a different failure with the same exit
|
||||
code. No `default` group, no enabled upstream, a syntax error — `run` prints the
|
||||
diagnostic and exits 2 as well. Both were run here against a locally built
|
||||
image. A seed file whose only group was named `other`:
|
||||
diagnostic and exits 2 as well. A file whose only group was named `other`:
|
||||
|
||||
```
|
||||
FAIL groups: no group named 'default'; every unknown client is assigned to it
|
||||
@@ -116,18 +139,31 @@ nxdns run failed: MissingDefaultGroup
|
||||
run `nxdns check` to see the configuration in full
|
||||
```
|
||||
|
||||
and an empty `/etc/nxdns`:
|
||||
Under `restart: unless-stopped` any of these is a restart loop — Docker has no
|
||||
start limit and will retry forever. Read the lines above the failure, which name
|
||||
the fault. See [Troubleshoot nxdns](troubleshoot.md).
|
||||
|
||||
```
|
||||
info(config_bootstrap): no configuration file at '/etc/nxdns/config.zon'; using the database as it is
|
||||
nxdns run failed: NoUsableUpstreams
|
||||
run `nxdns check` to see the configuration in full
|
||||
### Database mode in Docker instead
|
||||
|
||||
Drop the `command:` line from `compose.yaml` and the container runs
|
||||
`nxdns run`, with the database as the configuration and the file read by nothing.
|
||||
On a fresh volume that database is empty and the container exits 2 with
|
||||
`NoUsableUpstreams`, so load it once before bringing the service up:
|
||||
|
||||
```sh
|
||||
docker compose -f deploy/docker/compose.yaml run --rm nxdns import /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
Under `restart: unless-stopped` either one is a restart loop, and the exit code
|
||||
alone no longer tells them apart: read the lines above the failure, which either
|
||||
name the diagnostic in the file or say there was no file at all. See
|
||||
[Troubleshoot nxdns](troubleshoot.md).
|
||||
The file is positional; add `--allow-delete` when re-running it against a
|
||||
populated volume and the diff deletes rows. Without this step, `restart:
|
||||
unless-stopped` plus exit 2 is a crash loop with no way out.
|
||||
|
||||
> Not run in a container on this host, for the reason in the verification note
|
||||
> at the top: no image could be staged here. The `nxdns import <file>` and
|
||||
> `nxdns import <file> --allow-delete` commands inside it were run directly
|
||||
> against the locally built binary — the first applied an additive file and
|
||||
> exited 0, the second was required after a plain `import` refused a
|
||||
> row-deleting file with `DestructiveImport` and exited 2.
|
||||
|
||||
The container runs as uid 65532, and the mount is read-only, so the container
|
||||
cannot repair permissions itself. Mode 0644 works and was used here. If the
|
||||
@@ -140,9 +176,12 @@ chmod 0600 deploy/docker/etc-nxdns/config.zon
|
||||
```
|
||||
|
||||
> Not verified on this host: `chown` to a uid you do not own needs root. What
|
||||
> was verified is the failure it prevents — a seed file at 0600 owned by
|
||||
> another uid makes the container log `nxdns run failed: AccessDenied` and
|
||||
> restart in a loop. See [Troubleshoot nxdns](troubleshoot.md).
|
||||
> was verified is the failure it prevents — a configuration file at 0600 owned
|
||||
> by another uid makes the container refuse to start and restart in a loop. In
|
||||
> file mode an unreadable file is a configuration fault:
|
||||
> `FAIL /etc/nxdns/config.zon: not readable` followed by
|
||||
> `nxdns run failed: ManagedConfigUnreadable`, exit 2. See
|
||||
> [Troubleshoot nxdns](troubleshoot.md).
|
||||
|
||||
## 3. Run it
|
||||
|
||||
@@ -169,14 +208,21 @@ bind mount — against the directory holding the file, not against your shell,
|
||||
and it takes the project name `docker` from that directory either way, which is
|
||||
why the container is `docker-nxdns-1`.
|
||||
|
||||
A healthy first start logs the seeding and the bound sockets:
|
||||
A healthy first start logs the reconcile, the authority and the bound sockets:
|
||||
|
||||
```
|
||||
info(config_bootstrap): seeded the database from '/etc/nxdns/config.zon'
|
||||
info(migrations): config.db migrated from schema version 0 to 2
|
||||
reconciled '/etc/nxdns/config.zon': upstreams +1 ~0 -0; settings +45 ~0 -0;
|
||||
settings keys changed: dns.bind_ipv4 dns.bind_ipv6 dns.port web.bind web.port …
|
||||
web authentication is now enabled
|
||||
info(nxdns): authority: file (/etc/nxdns/config.zon)
|
||||
info(nxdns): nxdns <version> serving on udp [::]:53 tcp [::]:53 tcp 0.0.0.0:53; 1 upstream(s); blocklist generation 1
|
||||
info(web_server): web interface listening on 0.0.0.0:8080
|
||||
```
|
||||
|
||||
Every later start on an unchanged file reports `reconciled
|
||||
'/etc/nxdns/config.zon': no changes` and writes nothing to the database.
|
||||
|
||||
Confirm it answers and that the admin interface is up:
|
||||
|
||||
```sh
|
||||
|
||||
@@ -154,11 +154,11 @@ in step 5, and step 4 has to write a file into it before then. systemd does not
|
||||
mind finding the directory already there; it adjusts the mode and ownership to
|
||||
what the unit asks for.
|
||||
|
||||
## 4. Write the seed configuration
|
||||
## 4. Write the configuration
|
||||
|
||||
nxdns starts from an empty database only if a configuration file tells it what
|
||||
to forward to. Write `/etc/nxdns/config.zon`. The smallest file that starts is
|
||||
one group named `default` and one enabled upstream:
|
||||
nxdns will not start with nothing to forward to. Write `/etc/nxdns/config.zon`.
|
||||
The smallest file that starts is one group named `default` and one enabled
|
||||
upstream:
|
||||
|
||||
```zon
|
||||
.{
|
||||
@@ -213,36 +213,49 @@ The upstream probe sends a real query, so this needs working DNS on the host at
|
||||
the time you run it. Exit 2 means `check` found something to fix and printed
|
||||
every problem it found, not only the first.
|
||||
|
||||
The file seeds the database once. From the second start onwards it is ignored
|
||||
and the database is the configuration; see
|
||||
[the configuration model](../explanation/configuration-model.md) and
|
||||
[Upgrade nxdns](upgrade.md) for how to change settings after that.
|
||||
Now load it into the database:
|
||||
|
||||
Once the seed has been consumed — after step 6 confirms you can log in — the
|
||||
plaintext in it is dead weight that only carries risk. The seed's
|
||||
`web.password` is hashed into `web.password_hash` at import time and the
|
||||
plaintext is never stored; `nxdns export` writes `.password = ""` back out
|
||||
alongside the hash. Nothing downstream ever reads the plaintext again, so
|
||||
delete the file:
|
||||
```sh
|
||||
nxdns import /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
```
|
||||
info(migrations): config.db migrated from schema version 0 to 2
|
||||
imported /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
The plaintext password is hashed into `web.password_hash` and never stored as
|
||||
plaintext; `nxdns export` writes `.password = null` beside the hash. Nothing
|
||||
downstream reads the plaintext again, so once step 6 confirms you can log in you
|
||||
can delete the file:
|
||||
|
||||
```sh
|
||||
rm /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
Keep it only if you want the seed as a record of the intended starting
|
||||
configuration, and if you keep it, leave it at 0640 root:nxdns. Note that a
|
||||
kept seed is not a backup — `nxdns export` is
|
||||
(see [Back up and restore](back-up-and-restore.md)), and the export carries the
|
||||
password hash rather than the password.
|
||||
A kept file is not a backup — `nxdns export` is (see
|
||||
[Back up and restore](back-up-and-restore.md)), and the export carries the
|
||||
password hash rather than the password. If you keep it, leave it at 0640
|
||||
root:nxdns.
|
||||
|
||||
That is the **database mode** install, which is what the packaged unit runs:
|
||||
`ExecStart=/usr/local/bin/nxdns run`, no `--config`, so nothing reads a file
|
||||
after this step. Change settings afterwards through the admin interface, the
|
||||
API, or an export–edit–import cycle.
|
||||
|
||||
If you would rather keep `/etc/nxdns/config.zon` in git and have every restart
|
||||
converge onto it, do not delete the file — go to
|
||||
[Run in file mode](#run-in-file-mode) instead, and skip the `rm`.
|
||||
|
||||
> Verified on this host, with a scratch `--config` and `--data-dir` in place of
|
||||
> `/etc/nxdns` and `/var/lib/nxdns`: a seed written under umask 022 came out
|
||||
> 0644, `nxdns check --config` on it printed `OK: no problems found` with no
|
||||
> mode warning, and after `nxdns import` of that seed an `nxdns export` wrote
|
||||
> `.password = ""` next to a populated `.password_hash =
|
||||
> "$argon2id$v=19$..."`. The `chown`, `chmod` and `rm` lines above are the
|
||||
> ordinary root-owned-file operations and were not run against a real
|
||||
> `/etc/nxdns`, which this host does not have.
|
||||
> `/etc/nxdns` and `/var/lib/nxdns` — those two paths are the only difference
|
||||
> from the blocks above. A file written under umask 022 came out 0644,
|
||||
> `nxdns check --config` on it printed `OK: no problems found` with no mode
|
||||
> warning, `nxdns import` of it printed the migration line and `imported <path>`
|
||||
> and exited 0, and a following `nxdns export` wrote `.password = null` next to
|
||||
> a populated `.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$…"`. The
|
||||
> `chown`, `chmod` and `rm` lines are ordinary root-owned-file operations and
|
||||
> were not run against a real `/etc/nxdns`, which this host does not have.
|
||||
|
||||
## 5. Start it
|
||||
|
||||
@@ -263,9 +276,16 @@ nxdns writes to stderr and systemd captures that into the journal; logging
|
||||
needs no further configuration. Port 53 is privileged, and the unit grants
|
||||
`CAP_NET_BIND_SERVICE` through `AmbientCapabilities`.
|
||||
|
||||
The unit does not restart the service after exit 2 or exit 64
|
||||
(`RestartPreventExitStatus=2 64`). Those are a wrong configuration and a wrong
|
||||
command line, and neither clears on a retry — restarting every two seconds until
|
||||
`StartLimitBurst` gives up would only bury the diagnostics that are already in
|
||||
the journal. `systemctl status nxdns` shows the failed state; fix the cause and
|
||||
start it again.
|
||||
|
||||
If the start fails, read [Troubleshoot nxdns](troubleshoot.md). The two common
|
||||
first-install failures are a port 53 already held by `systemd-resolved` and a
|
||||
seed file that does not parse.
|
||||
configuration file that does not parse.
|
||||
|
||||
## 6. Confirm it answers
|
||||
|
||||
@@ -276,7 +296,7 @@ dig @<server-ip> example.com A +short
|
||||
```
|
||||
|
||||
The admin interface is on port 8080 by default; log in with the password from
|
||||
the seed file. `http://<server-ip>:8080/api/health` reports upstream
|
||||
the configuration file. `http://<server-ip>:8080/api/health` reports upstream
|
||||
availability and disk state without a login.
|
||||
|
||||
> Not verified on this host as written: `<server-ip>` is a placeholder, and a
|
||||
@@ -287,6 +307,140 @@ availability and disk state without a login.
|
||||
> port returned 200. Only the address and the port differ from the lines
|
||||
> above.
|
||||
|
||||
## Run in file mode
|
||||
|
||||
In file mode `/etc/nxdns/config.zon` is the configuration: every start converges
|
||||
the database onto it, and the admin interface refuses configuration edits with a
|
||||
403 naming the file. Use it when you want the file in git and deployed by
|
||||
Ansible. Stay in database mode when you want the UI to be the way things change.
|
||||
|
||||
The packaged unit is flagless on purpose — it is correct as shipped, and a
|
||||
commented-out alternative `ExecStart` in a unit file is documentation
|
||||
masquerading as configuration. File mode is a drop-in.
|
||||
|
||||
### Adopt file mode on a box that is already running
|
||||
|
||||
Run these in order. **Stop first**, and do not skip that: any edit made through
|
||||
the UI between an export and the restart would be silently reverted by the first
|
||||
reconcile, and `nxdns check` against a live database refuses to grade it (below).
|
||||
|
||||
If you are arriving here from an upgrade, the binary must already be the new
|
||||
one before you export. An export written by 0.0.1 carries a `.password = ""`
|
||||
line this binary refuses; see
|
||||
[the order trap](upgrade.md#the-order-trap-export-with-the-new-binary-not-the-old-one).
|
||||
|
||||
```sh
|
||||
systemctl stop nxdns
|
||||
nxdns export --out /etc/nxdns/config.zon
|
||||
nxdns check --config /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
```
|
||||
wrote /etc/nxdns/config.zon
|
||||
checking configuration file /etc/nxdns/config.zon
|
||||
OK upstreams[0] https://cloudflare-dns.com
|
||||
OK: no problems found
|
||||
```
|
||||
|
||||
Then add the drop-in and start:
|
||||
|
||||
```sh
|
||||
mkdir -p /etc/systemd/system/nxdns.service.d
|
||||
cat > /etc/systemd/system/nxdns.service.d/file-mode.conf <<'EOF'
|
||||
[Service]
|
||||
ExecStart=
|
||||
ExecStart=/usr/local/bin/nxdns run --config=/etc/nxdns/config.zon
|
||||
EOF
|
||||
systemctl daemon-reload
|
||||
systemctl start nxdns
|
||||
```
|
||||
|
||||
The empty `ExecStart=` is required. Without it systemd appends a second command
|
||||
to the list rather than replacing the first, and the unit tries to run nxdns
|
||||
twice.
|
||||
|
||||
The first start after adoption changes nothing, because the file was rendered
|
||||
from the database it is now governing:
|
||||
|
||||
```
|
||||
reconciled '/etc/nxdns/config.zon': no changes
|
||||
info(nxdns): authority: file (/etc/nxdns/config.zon)
|
||||
info(nxdns): nxdns <version> serving on udp [::]:53 tcp [::]:53 tcp 0.0.0.0:53; 1 upstream(s); blocklist generation 1
|
||||
```
|
||||
|
||||
`authority: file` is the line that confirms the drop-in took. Blocklists,
|
||||
compiled snapshots and client history all survive, and every later start on an
|
||||
unchanged file writes nothing either.
|
||||
|
||||
The file now carries `web.password_hash`, so restrict it the same way step 4
|
||||
does — `chown root:nxdns`, `chmod 0640`. The unit's `ReadOnlyPaths=/etc/nxdns`
|
||||
denies the service write access to that directory, so the process that reads the
|
||||
file cannot modify it.
|
||||
|
||||
### Change the configuration from now on
|
||||
|
||||
Edit the file, validate it, restart:
|
||||
|
||||
```sh
|
||||
$EDITOR /etc/nxdns/config.zon
|
||||
nxdns check --config /etc/nxdns/config.zon
|
||||
systemctl restart nxdns
|
||||
```
|
||||
|
||||
Make `nxdns check --config` the precondition of any Ansible handler that
|
||||
restarts nxdns. A file-mode start reads the file on **every** boot, so a bad
|
||||
push that skips its handler does not fail at deploy time — it detonates at the
|
||||
next power cut. Validating before restarting turns that into a failed deploy at
|
||||
noon.
|
||||
|
||||
The restart prints what it changed:
|
||||
|
||||
```
|
||||
reconciled '/etc/nxdns/config.zon': upstreams +1 ~0 -0; settings +0 ~2 -0;
|
||||
settings keys changed: dns.port web.port
|
||||
```
|
||||
|
||||
### Leave file mode
|
||||
|
||||
Remove the drop-in and restart. The database already holds the last reconciled
|
||||
state, so nothing else is needed and the server comes back serving the same
|
||||
configuration:
|
||||
|
||||
```sh
|
||||
rm /etc/systemd/system/nxdns.service.d/file-mode.conf
|
||||
systemctl daemon-reload
|
||||
systemctl restart nxdns
|
||||
```
|
||||
|
||||
```
|
||||
info(nxdns): authority: database
|
||||
```
|
||||
|
||||
> Verified on this host end to end, against a scratch `--data-dir` and a scratch
|
||||
> configuration path instead of `/var/lib/nxdns` and `/etc/nxdns`, on
|
||||
> unprivileged ports — this machine has neither of those directories, no root,
|
||||
> and no installed unit. Every `nxdns` line above was run and produced the output
|
||||
> shown, with only those paths and the port numbers in the `serving on` line
|
||||
> differing.
|
||||
>
|
||||
> The run: a database-mode instance was started, a blocklist source was added
|
||||
> through the API to make it UI-configured, then `nxdns check` against the live
|
||||
> database printed the uncheckpointed-log FAIL and exited 2 (which is why this
|
||||
> section stops the service first). After the stop, `nxdns export --out` wrote
|
||||
> the file, `nxdns check --config` on it exited 0 with `OK: no problems found`,
|
||||
> and the first file-mode start printed `reconciled '<path>': no changes` and
|
||||
> `authority: file (<path>)`. A second file-mode start printed `no changes`
|
||||
> again and loaded the 3096006-byte compiled blocklist from disk with no
|
||||
> download. Dropping the flag printed `authority: database` and served the same
|
||||
> configuration.
|
||||
>
|
||||
> The `systemctl`, `mkdir`, `cat > …/file-mode.conf` and `rm` lines need root and
|
||||
> an installed unit and were **not** run. What was checked instead:
|
||||
> `systemd-analyze verify` on `deploy/systemd/nxdns.service` with those exact two
|
||||
> `ExecStart` lines appended, which reported only the usual
|
||||
> `Command /usr/local/bin/nxdns is not executable` for the absent binary and
|
||||
> nothing about the override.
|
||||
|
||||
## Raspberry Pi 5
|
||||
|
||||
The Pi 5 is aarch64. Nothing about the procedure changes except which tarball
|
||||
|
||||
@@ -11,7 +11,7 @@ data directory is `/var/lib/nxdns` and the web port is 8080.
|
||||
|
||||
## 1. Set the password
|
||||
|
||||
Put it in the seed configuration file, under `web`:
|
||||
Put it in the configuration file, under `web`:
|
||||
|
||||
```zon
|
||||
.{
|
||||
@@ -21,17 +21,60 @@ Put it in the seed configuration file, under `web`:
|
||||
}
|
||||
```
|
||||
|
||||
At import time the plaintext is hashed with argon2id into `web.password_hash`
|
||||
and discarded. It becomes no database row and appears in no log line. Setting
|
||||
both `password` and `password_hash` in one file is refused:
|
||||
The plaintext is hashed with argon2id into `web.password_hash` and discarded. It
|
||||
becomes no database row and appears in no log line. Setting both `password` and
|
||||
`password_hash` in one file is refused:
|
||||
|
||||
```
|
||||
web.password: password and password_hash are both set; ambiguity in a security setting is refused
|
||||
import failed: PasswordAndHashBothSet
|
||||
```
|
||||
|
||||
The seed file is read only while the database is empty. On a server that
|
||||
already has a database, use step 4 or step 5 instead.
|
||||
Applying that file — with `nxdns import`, or with a `nxdns run --config` start —
|
||||
announces the change:
|
||||
|
||||
```
|
||||
web authentication is now enabled
|
||||
```
|
||||
|
||||
### Absent, empty, and set are three different things
|
||||
|
||||
The two fields are optional, and the difference between leaving one out and
|
||||
setting it to `""` is the difference between keeping your password and removing
|
||||
it:
|
||||
|
||||
| The file says | Effect on the stored password |
|
||||
| --- | --- |
|
||||
| Neither field | Nothing. It stays exactly as it was. |
|
||||
| `.password = "…"` | Installs that password. Unchanged plaintext keeps the existing hash rather than re-hashing it. |
|
||||
| `.password = ""` | Refused. |
|
||||
| `.password_hash = "$argon2id$…"` | Installs that hash, for example from an export. |
|
||||
| `.password_hash = ""` | **Removes the password.** Authentication is off. |
|
||||
|
||||
Absence has to mean "keep", because the alternative is a foot-gun with a live
|
||||
round in it. An export carries the full PHC string, which is long and ugly, and
|
||||
sooner or later someone trims that line out of a file before committing it —
|
||||
meaning "leave the password alone". If absence meant "no password", that edit
|
||||
would open the admin interface to the whole LAN without a word.
|
||||
|
||||
So removing the password takes the explicit empty string:
|
||||
|
||||
```
|
||||
web authentication is now disabled
|
||||
```
|
||||
|
||||
And an empty plaintext is refused outright, because hashing the empty string
|
||||
would switch authentication *on* while making every login impossible — the login
|
||||
handler rejects empty passwords:
|
||||
|
||||
```
|
||||
FAIL web.password: password is set to the empty string; omit the field to keep the stored password, or set password_hash = "" to disable authentication
|
||||
```
|
||||
|
||||
Which of steps 4 and 5 applies to your server depends on its authority. Under
|
||||
`nxdns run --config FILE` the file is the password: edit it and restart, and the
|
||||
API refuses the change with a 403. Under bare `nxdns run` the database holds it,
|
||||
and step 4 or step 5 is how it moves.
|
||||
|
||||
## 2. Log in
|
||||
|
||||
@@ -63,7 +106,7 @@ The session token comes back in a `Set-Cookie` header, not in the body. In the
|
||||
jar it looks like this (value redacted here):
|
||||
|
||||
```
|
||||
#HttpOnly_127.0.0.1 FALSE / FALSE 1785770178 nxdns_session <redacted>
|
||||
#HttpOnly_127.0.0.1 FALSE / FALSE 1786559938 nxdns_session <redacted>
|
||||
```
|
||||
|
||||
The cookie is named `nxdns_session` and carries `HttpOnly; SameSite=Lax;
|
||||
@@ -123,6 +166,9 @@ already is.
|
||||
|
||||
## 4. Change the password on a running server
|
||||
|
||||
This is a database-mode procedure. In file mode `PUT /api/settings` answers 403
|
||||
naming the file; edit `web.password` there and restart instead.
|
||||
|
||||
Send the new one to `PUT /api/settings` as `web.password`. The response is the
|
||||
full settings document; `password` is write-only and `password_hash` is neither
|
||||
readable nor directly writable, so neither value comes back.
|
||||
@@ -161,7 +207,7 @@ Log back in with the new password. That is the whole rotation.
|
||||
|
||||
If you have lost the password, the admin interface cannot help — go through the
|
||||
database instead. Export, edit, import. `nxdns export` always writes
|
||||
`.password = ""` and carries the hash, so an exported file re-imports without
|
||||
`.password = null` and carries the hash, so an exported file re-imports without
|
||||
anyone knowing the password. To install a new one, put it in `.password` and
|
||||
clear `.password_hash`:
|
||||
|
||||
@@ -169,11 +215,22 @@ clear `.password_hash`:
|
||||
nxdns export --data-dir /tmp/nxdns-lab/data --out /tmp/nxdns-lab/rekeyed.zon
|
||||
```
|
||||
|
||||
Edit the `web` section of `/tmp/nxdns-lab/rekeyed.zon` so it reads:
|
||||
Edit the `web` section of `/tmp/nxdns-lab/rekeyed.zon`: set `.password` to the
|
||||
new value and **delete the `.password_hash` line entirely**, so the `web` block
|
||||
carries one password field and not two:
|
||||
|
||||
```zon
|
||||
.password = "offline-password",
|
||||
.password_hash = "",
|
||||
```
|
||||
|
||||
Deleting the line is the part to get right. Setting `.password_hash = ""`
|
||||
alongside a plaintext password does not clear the way for it — an empty string
|
||||
is a present value meaning "no password", so the file then states two
|
||||
contradictory things and is refused:
|
||||
|
||||
```
|
||||
FAIL web.password: password and password_hash are both set; ambiguity in a security setting is refused
|
||||
import failed: PasswordAndHashBothSet
|
||||
```
|
||||
|
||||
Stop the server before importing. `import` rewrites the stored hash underneath a
|
||||
@@ -184,14 +241,17 @@ terminal stops it, and it goes back up with the same command:
|
||||
|
||||
```sh
|
||||
# Ctrl-C the `nxdns run` terminal, or `kill` its pid from another shell
|
||||
nxdns import /tmp/nxdns-lab/rekeyed.zon --force --data-dir /tmp/nxdns-lab/data
|
||||
nxdns run --data-dir /tmp/nxdns-lab/data --config /tmp/nxdns-lab/etc/config.zon
|
||||
nxdns import /tmp/nxdns-lab/rekeyed.zon --data-dir /tmp/nxdns-lab/data
|
||||
nxdns run --data-dir /tmp/nxdns-lab/data
|
||||
```
|
||||
|
||||
```
|
||||
imported /tmp/nxdns-lab/rekeyed.zon
|
||||
```
|
||||
|
||||
No flag is needed: replacing a password edits a settings value and deletes no
|
||||
rows.
|
||||
|
||||
On a real install the stop and start are `systemctl stop nxdns` and
|
||||
`systemctl start nxdns` around the same `import` — **not verified on this
|
||||
host**, which has no installed nxdns systemd unit (`systemctl status nxdns`
|
||||
@@ -202,7 +262,7 @@ Once it is back up the old password is refused and the new one works:
|
||||
|
||||
```sh
|
||||
curl -sS -X POST http://127.0.0.1:8451/api/auth/login \
|
||||
-H 'content-type: application/json' -d '{"password":"a-new-password"}' \
|
||||
-H 'content-type: application/json' -d '{"password":"lab-password"}' \
|
||||
-w ' (old password, http %{http_code})\n'
|
||||
curl -sS -c /tmp/nxdns-lab/c5.txt -X POST http://127.0.0.1:8451/api/auth/login \
|
||||
-H 'content-type: application/json' -d '{"password":"offline-password"}' \
|
||||
@@ -217,19 +277,19 @@ curl -sS -b /tmp/nxdns-lab/c5.txt -o /dev/null -w 'stats: %{http_code}\n' \
|
||||
stats: 200
|
||||
```
|
||||
|
||||
The next export shows the new hash and an empty `password` again:
|
||||
The next export shows the new hash and a null `password` again:
|
||||
|
||||
```sh
|
||||
nxdns export --data-dir /tmp/nxdns-lab/data | grep password
|
||||
```
|
||||
|
||||
```
|
||||
.password = "",
|
||||
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$kvlRj1tdGul3MlfbvzLncLKWirNpJRJ3howFA9/ysgg$7elW7PPQ3WXHwI4YOmOpZ/1KNEQo7ZDLRJhnYOPMjqw",
|
||||
.password = null,
|
||||
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$xqzK66LgiWGvyCmCl6ZRa3GHH0nS5qZnRgVfWmeGadc$1mafhflKFIg3vcHjJaDMAXGQiOjtym2UADZsPW1xkfw",
|
||||
```
|
||||
|
||||
`--force` is required because the database already holds configuration. See
|
||||
[back up and restore](back-up-and-restore.md).
|
||||
See [back up and restore](back-up-and-restore.md) for when `import` does need
|
||||
`--allow-delete`.
|
||||
|
||||
## What happens with no password set
|
||||
|
||||
@@ -279,6 +339,11 @@ interface at the very least, and preferably set a password.
|
||||
[configuration reference](../reference/configuration.md); the routes are in
|
||||
the [API reference](../reference/api.md).
|
||||
|
||||
Every command on this page was executed on this host as written, except the
|
||||
`systemctl` stop and start named in step 5 and marked **not verified on this
|
||||
host** there.
|
||||
Every command on this page was executed on this host as written, against the
|
||||
lab described at the top, except the `systemctl` stop and start named in step 5
|
||||
and marked **not verified on this host** there. That includes the whole of
|
||||
steps 2 to 5, re-run for this revision: the login, logout and rate-limit
|
||||
transcripts reproduced exactly as printed, the both-set refusal in step 5 was
|
||||
reproduced by leaving `.password_hash = ""` in the file, and the rekey then
|
||||
succeeded once that line was deleted. The cookie jar's expiry timestamp is the
|
||||
one that run produced and will differ on yours.
|
||||
|
||||
+105
-30
@@ -37,10 +37,22 @@ checked on its first line.
|
||||
**Fixes by cause.**
|
||||
|
||||
- `NoUsableUpstreams` — the database has no enabled upstream. On a fresh
|
||||
install this means the seed file was missing or in the wrong place; the start
|
||||
log says `no configuration file at '/etc/nxdns/config.zon'; using the
|
||||
database as it is`. Write the seed file and start again against the still
|
||||
empty database, or `nxdns import <file> --force`.
|
||||
install in database mode this is simply an empty database, and the run says
|
||||
what to do about it on the next line:
|
||||
|
||||
```
|
||||
nxdns run failed: NoUsableUpstreams
|
||||
run `nxdns check` to see the configuration in full
|
||||
load one with `nxdns import <file>`, or make a file the source of truth with `nxdns run --config <file>`
|
||||
```
|
||||
|
||||
Write a configuration file and take either exit: `nxdns import <file>` to load
|
||||
it into the database once, or add `--config <file>` to `ExecStart` to make the
|
||||
file the configuration from then on.
|
||||
- `ManagedConfigUnreadable` — the service runs `run --config FILE` and that file
|
||||
is missing or the process may not read it. The path is in the FAIL line above
|
||||
the failure. File mode never falls back to the database, on purpose: a
|
||||
fallback would turn a bad deploy into a silently stale configuration.
|
||||
- `BadCertificate` — a DoH or DoT listener is enabled and its certificate or
|
||||
key is unreadable, too large, unparseable, or the key does not belong to the
|
||||
certificate. `run` names both paths before it exits:
|
||||
@@ -79,10 +91,10 @@ checked on its first line.
|
||||
- `BadBindAddress` — `dns.bind_ipv4` or `dns.bind_ipv6` is not an address of
|
||||
that family.
|
||||
|
||||
## A seed file you just wrote is rejected
|
||||
## A configuration file you just wrote is rejected
|
||||
|
||||
**Symptom.** A first start against an empty database prints the validation
|
||||
problem and stops with exit 2:
|
||||
**Symptom.** `nxdns run --config`, `nxdns check --config` or `nxdns import`
|
||||
prints the validation problem and stops with exit 2:
|
||||
|
||||
```
|
||||
FAIL groups: no group named 'default'; every unknown client is assigned to it
|
||||
@@ -98,7 +110,7 @@ nxdns run failed: ParseZon
|
||||
run `nxdns check` to see the configuration in full
|
||||
```
|
||||
|
||||
So does a seed file whose upstream list is empty or all disabled:
|
||||
So does a file whose upstream list is empty or all disabled:
|
||||
|
||||
```
|
||||
FAIL upstreams: at least one upstream must be enabled
|
||||
@@ -106,9 +118,9 @@ nxdns run failed: NoUpstreams
|
||||
run `nxdns check` to see the configuration in full
|
||||
```
|
||||
|
||||
`NoUpstreams` from a seed file is not the same fault as `NoUsableUpstreams`
|
||||
above: the first is a file `run` refused, the second is a database `run`
|
||||
accepted and found empty. Both are exit 2.
|
||||
`NoUpstreams` from a file is not the same fault as `NoUsableUpstreams` above:
|
||||
the first is a file `run` refused, the second is a database `run` accepted and
|
||||
found empty. Both are exit 2.
|
||||
|
||||
**Diagnosis.** Run the same file through `check`, which reports the same
|
||||
problems and exits 2:
|
||||
@@ -117,11 +129,21 @@ problems and exits 2:
|
||||
nxdns check --config /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
**Fix.** Correct the file the diagnostics name and start again. The database is
|
||||
still empty after a failed seed, so the next start re-reads the file. The exit
|
||||
code no longer depends on which command read the file: all three of these files
|
||||
were run through `run`, `check` and `import` here, and every one of the nine
|
||||
combinations exited 2 with the same diagnostic.
|
||||
**Fix.** Correct the file the diagnostics name and start again. Nothing was
|
||||
applied — a file-mode reconcile happens in one transaction that rolls back, and
|
||||
a failed `import` leaves the database untouched. The exit code does not depend
|
||||
on which command read the file: all three of these files were run through `run`,
|
||||
`check` and `import` here, and every one of the nine combinations exited 2 with
|
||||
the same diagnostic.
|
||||
|
||||
Under the shipped systemd unit an exit 2 stops the service rather than
|
||||
restarting it (`RestartPreventExitStatus=2 64`), so the journal holds the
|
||||
diagnostics instead of drowning them in a restart loop. `systemctl start nxdns`
|
||||
once the file is fixed.
|
||||
|
||||
Make `nxdns check --config <file>` the precondition in whatever pushes the file.
|
||||
In file mode every boot reads it, so an unvalidated bad push does not fail at
|
||||
deploy time — it fails at the next restart, which may be a power cut at 3am.
|
||||
|
||||
## `nxdns check` fails on a server that is running fine
|
||||
|
||||
@@ -210,10 +232,11 @@ startup cycle, not a fix.
|
||||
## The container restarts in a loop
|
||||
|
||||
**Symptom.** `docker compose ps` shows the container restarting, and the log is
|
||||
one line repeated:
|
||||
the same failure repeated. Docker has no start limit, so this goes on forever.
|
||||
|
||||
```
|
||||
nxdns run failed: AccessDenied
|
||||
FAIL /etc/nxdns/config.zon: not readable
|
||||
nxdns run failed: ManagedConfigUnreadable
|
||||
```
|
||||
|
||||
**Diagnosis.**
|
||||
@@ -223,10 +246,10 @@ docker inspect -f '{{.State.Status}} exit={{.State.ExitCode}} restarts={{.Restar
|
||||
stat -c '%a %u:%g %n' deploy/docker/etc-nxdns/config.zon
|
||||
```
|
||||
|
||||
Exit 1 with `AccessDenied` means the container could not read the seed file.
|
||||
The container runs as uid 65532 and `/etc/nxdns` is mounted read-only, so a
|
||||
file at mode 0600 owned by your own uid is unreadable to it and the container
|
||||
cannot repair it.
|
||||
Exit 2 naming the configuration path means the container could not read the
|
||||
file the shipped `command:` makes its configuration. The container runs as uid
|
||||
65532 and `/etc/nxdns` is mounted read-only, so a file at mode 0600 owned by
|
||||
your own uid is unreadable to it and the container cannot repair it.
|
||||
|
||||
**Fix.** Either make the file world-readable, when it holds no secret:
|
||||
|
||||
@@ -241,14 +264,65 @@ chown 65532:65532 deploy/docker/etc-nxdns/config.zon
|
||||
chmod 0600 deploy/docker/etc-nxdns/config.zon
|
||||
```
|
||||
|
||||
The 0644 path was verified here, including the recovery: after the `chmod` the
|
||||
container started and answered queries. The `chown` needs root and was not run
|
||||
here.
|
||||
The 0644 path was verified against an earlier revision of this page, including
|
||||
the recovery: after the `chmod` the container started and answered queries. The
|
||||
`chown` needs root and was not run here.
|
||||
|
||||
A container that exits 2 instead — `nxdns run failed: NoUsableUpstreams` after
|
||||
`no configuration file at '/etc/nxdns/config.zon'` — has no seed file at all on
|
||||
a fresh volume. Create `deploy/docker/etc-nxdns/config.zon` and bring it up
|
||||
again; see [Install with Docker](install-with-docker.md).
|
||||
`FAIL /etc/nxdns/config.zon: no such file` instead of `not readable` means there
|
||||
is no configuration file at all. Create `deploy/docker/etc-nxdns/config.zon` and
|
||||
bring it up again; see [Install with Docker](install-with-docker.md).
|
||||
|
||||
A container that exits 2 with `NoUsableUpstreams` is in database mode — the
|
||||
`command:` line naming `--config` was removed — on a volume whose database is
|
||||
still empty. Load one and bring it back up:
|
||||
|
||||
```sh
|
||||
docker compose -f deploy/docker/compose.yaml run --rm nxdns import /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
> Not re-run on this host: staging a release image needs `zig build dist`, which
|
||||
> could not run here while the web bundle was mid-rebuild by other work in the
|
||||
> same checkout. The failure text quoted above is what the same binary prints
|
||||
> outside a container, which was reproduced here, with the container's paths.
|
||||
|
||||
## The admin interface refuses an edit with 403
|
||||
|
||||
**Symptom.** Saving anything in the admin interface fails, and the API answers:
|
||||
|
||||
```json
|
||||
{"error":"configuration is managed by /etc/nxdns/config.zon; edit the file and restart"}
|
||||
```
|
||||
|
||||
This is not a fault. The service runs `nxdns run --config`, which makes that
|
||||
file the configuration, and configuration writes through the API are refused so
|
||||
the file and the running server cannot drift apart.
|
||||
|
||||
**Diagnosis.** The start log names the authority:
|
||||
|
||||
```sh
|
||||
journalctl -u nxdns | grep 'authority:'
|
||||
```
|
||||
|
||||
```
|
||||
info(nxdns): authority: file (/etc/nxdns/config.zon)
|
||||
```
|
||||
|
||||
**Fix.** Edit the file, validate it, restart:
|
||||
|
||||
```sh
|
||||
$EDITOR /etc/nxdns/config.zon
|
||||
nxdns check --config /etc/nxdns/config.zon
|
||||
systemctl restart nxdns
|
||||
```
|
||||
|
||||
Or, if you want the interface to be how this box is configured, leave file mode:
|
||||
drop `--config` from `ExecStart` and restart. The database already holds the
|
||||
last reconciled state, so nothing is lost. See
|
||||
[Run in file mode](install-with-systemd.md#run-in-file-mode).
|
||||
|
||||
Pausing blocking, refreshing blocklists and reloading certificates are not
|
||||
configuration and keep working in file mode. Deleting a client works too, unless
|
||||
the file names that client's address.
|
||||
|
||||
## The container cannot reach its upstreams
|
||||
|
||||
@@ -324,7 +398,8 @@ window of unfiltered answers.
|
||||
A generation number with nothing being blocked is a different problem: the
|
||||
snapshot loaded but has no sources in it. The line
|
||||
`blocklist snapshot generation 1: 0 of 0 sources loaded` says exactly that. Add
|
||||
a source in the admin interface, or in the seed file before the first start.
|
||||
a source in the admin interface, or a `blocklist_sources` entry to the
|
||||
configuration file with a `group_sources` link naming a group.
|
||||
|
||||
## A database stamped by a newer binary
|
||||
|
||||
|
||||
+179
-27
@@ -21,6 +21,105 @@ Upgrading a build you made yourself is the last section of this page.
|
||||
> shapes and the verification commands are covered by
|
||||
> [Verify a release](verify-a-release.md), which says what was probed and how.
|
||||
|
||||
## Breaking change: `run --config` now means file authority
|
||||
|
||||
**Read this before upgrading if anything on your box passes `--config` to
|
||||
`nxdns run`** — a systemd drop-in, a wrapper script, or a `command:` in a
|
||||
compose file.
|
||||
|
||||
`run --config FILE` used to mean *seed once*: the file was read only while the
|
||||
database was still empty, and ignored on every start after that. It now means
|
||||
*the file is the configuration*: every start reconciles the database onto it.
|
||||
|
||||
For a box that was seeded once and then configured through the admin interface,
|
||||
the first start after the upgrade converges the database back to that old seed
|
||||
file. **Every change made through the UI since seeding is deleted.**
|
||||
|
||||
There are two ways out, and you pick before you restart:
|
||||
|
||||
- **Keep the database.** Drop the flag. `nxdns run` with no `--config` serves
|
||||
the database exactly as it did before, and nothing reads a file. This is the
|
||||
right answer if the UI is how you change things.
|
||||
- **Adopt file mode cleanly.** Install the new binary, stop the service, export
|
||||
the current database over the file path, check it, then start with the flag.
|
||||
The first reconcile is then a no-op, because the file was rendered from the
|
||||
database it governs. **Install the new binary first** — see the order trap
|
||||
below. The full procedure is
|
||||
[Adopt file mode](install-with-systemd.md#adopt-file-mode-on-a-box-that-is-already-running).
|
||||
|
||||
`nxdns check --config FILE` is unchanged: it graded that file before and it
|
||||
grades that file now.
|
||||
|
||||
### The order trap: export with the new binary, not the old one
|
||||
|
||||
Take the export **after** you have replaced the binary, with the service
|
||||
stopped. Exporting first — the instinctive order, and the one step 1 of this
|
||||
page tells you to take for a backup — produces a file the new binary refuses.
|
||||
|
||||
A 0.0.1 `nxdns export` writes both fields:
|
||||
|
||||
```zon
|
||||
.password = "",
|
||||
.password_hash = "$argon2id$v=19$m=19456,t=2,p=1$…",
|
||||
```
|
||||
|
||||
An empty `password_hash` used to mean "unset". It now means "disable
|
||||
authentication", so it is a *present* value — and a file that carries both
|
||||
fields states two different things about the password and is refused:
|
||||
|
||||
```
|
||||
FAIL web.password: password and password_hash are both set; ambiguity in a security setting is refused
|
||||
FAIL web.password: password is set to the empty string; omit the field to keep the stored password, or set password_hash = "" to disable authentication
|
||||
nxdns run failed: PasswordAndHashBothSet
|
||||
```
|
||||
|
||||
The old `nxdns check` passes that file, because the old binary agreed with the
|
||||
old rule. So the failure lands at the first start after the upgrade, with the
|
||||
resolver stopped and the unit refusing to retry it
|
||||
(`RestartPreventExitStatus=2 64`). A new `nxdns export` writes
|
||||
`.password = null` instead and has no such problem.
|
||||
|
||||
**If you already have an old export you want to adopt**, you do not need to
|
||||
redo it. Delete the empty-password line and the file is valid:
|
||||
|
||||
```sh
|
||||
sed -i '/^ \.password = "",$/d' /etc/nxdns/config.zon
|
||||
nxdns check --config /etc/nxdns/config.zon
|
||||
```
|
||||
|
||||
Keep the `.password_hash` line — that is the password, and deleting it as well
|
||||
would leave the file saying nothing about authentication, which means "keep
|
||||
whatever is stored" rather than anything you would notice.
|
||||
|
||||
The same trap has nothing to do with file mode as such: it is any 0.0.1 export
|
||||
fed to the new binary, so it also applies to a restore through `nxdns import`.
|
||||
Backups taken with 0.0.1 need that one line removed before they will load.
|
||||
|
||||
> Verified on this host, with one substitution stated: the repository has no
|
||||
> 0.0.1 binary to hand, so the old export was **simulated** by taking a current
|
||||
> `nxdns export` and rewriting `.password = null` to `.password = ""`, which is
|
||||
> the one byte-level difference between the two formats. Against that file,
|
||||
> `nxdns check --config` printed both FAIL lines above and exited 2, and
|
||||
> `nxdns run --config` printed them and failed `PasswordAndHashBothSet`. After
|
||||
> the `sed` above, `nxdns check --config` exited 0 with `OK: no problems found`
|
||||
> and `nxdns import` of the same file exited 0, keeping the hash. The claim
|
||||
> about what 0.0.1's `export` emitted is read from that release's source —
|
||||
> `git show v0.0.1:src/config/export.zig` line 71 is `cfg.web.password = "";` —
|
||||
> not from running that binary.
|
||||
|
||||
A database-mode install that never passed `--config` needs nothing. Under
|
||||
Docker, a fresh database-mode install must either take the new compose file or
|
||||
run `import` once — see
|
||||
[Database mode in Docker](install-with-docker.md#database-mode-in-docker-instead).
|
||||
|
||||
Two smaller renames in the same release: `nxdns import --force` is now
|
||||
`--allow-delete`, and it is required only when the file's diff would delete
|
||||
rows rather than whenever the database is non-empty. `nxdns check` no longer
|
||||
falls back to a default file path when there is no database; it reports the
|
||||
absent database and names the two ways to get one.
|
||||
|
||||
There is no schema migration in this change.
|
||||
|
||||
## 1. Take an export first
|
||||
|
||||
There is no downgrade path, so the export is what you fall back to:
|
||||
@@ -33,6 +132,13 @@ nxdns export --out /some/backup/nxdns-config.zon
|
||||
command relies on the default `--data-dir /var/lib/nxdns` that a systemd
|
||||
install has.
|
||||
|
||||
This export is a fallback, not a file to deploy. If you are adopting file mode,
|
||||
take a *second* export after the binary swap and use that one — an export
|
||||
written by 0.0.1 carries a `.password = ""` line the new binary refuses, as
|
||||
[the order trap](#the-order-trap-export-with-the-new-binary-not-the-old-one)
|
||||
explains. The same line has to come out of this backup before the new binary
|
||||
will import it.
|
||||
|
||||
> Verified on this host with both paths substituted, since it has neither
|
||||
> `/var/lib/nxdns` nor `/some/backup`. `SCRATCH` below is a scratch directory,
|
||||
> and its `data/` was populated beforehand with `nxdns import`:
|
||||
@@ -131,9 +237,11 @@ NXDNS_VERSION=$VERSION docker compose -f deploy/docker/compose.yaml pull
|
||||
NXDNS_VERSION=$VERSION docker compose -f deploy/docker/compose.yaml up -d
|
||||
```
|
||||
|
||||
Compose recreates the container against the same `nxdns-data` volume. The seed
|
||||
file in `etc-nxdns` is not read again; the database in the volume is the
|
||||
configuration.
|
||||
Compose recreates the container against the same `nxdns-data` volume. What
|
||||
happens to the file in `etc-nxdns` depends on the `command:` in your compose
|
||||
file: with the shipped `run --config=/etc/nxdns/config.zon` the file is the
|
||||
configuration and the restart reconciles onto it; without it, the database in
|
||||
the volume is the configuration and the file is read by nothing.
|
||||
|
||||
Set `NXDNS_VERSION` on both lines, or export it. Without it the compose file
|
||||
falls back to `:latest`, and `pull` and `up` could then land on different
|
||||
@@ -184,11 +292,12 @@ OK: no problems found
|
||||
> zig 0.16.0
|
||||
> ```
|
||||
>
|
||||
> Against a running server whose database had just been migrated and seeded,
|
||||
> `nxdns check --data-dir` printed the uncheckpointed-log line above and exited
|
||||
> 2, while `nxdns export --out` followed by `nxdns check --config` on the result
|
||||
> exited 0 with `OK: no problems found`. The `dig` line was not run in this
|
||||
> round: nothing is listening on 127.0.0.1:53 here, and port 53 needs root.
|
||||
> Against a running server whose database had just taken a configuration write
|
||||
> through the API, `nxdns check --data-dir` printed the uncheckpointed-log line
|
||||
> above and exited 2, while `nxdns export --out` followed by
|
||||
> `nxdns check --config` on the result exited 0 with `OK: no problems found`.
|
||||
> Both were re-run for this revision. The `dig` line was not run in this round:
|
||||
> nothing is listening on 127.0.0.1:53 here, and port 53 needs root.
|
||||
|
||||
## What happens to the database
|
||||
|
||||
@@ -235,46 +344,89 @@ nxdns run failed: SchemaTooNew
|
||||
That run exits 1. Recovering means importing the export you took in step 1 into
|
||||
a fresh data directory with the older binary.
|
||||
|
||||
### Rolling back from file mode
|
||||
|
||||
Putting an older binary back needs no unit edit. The old binary accepts
|
||||
`run --config` — it just reads it as the old seed-once flag — and against a
|
||||
database that already holds configuration it ignores the file entirely and
|
||||
serves the last state the new binary reconciled. So the service comes back up
|
||||
on the configuration it was running.
|
||||
|
||||
The consequence is worth stating plainly: **file edits stop applying.** The old
|
||||
binary will not re-read the file, so every change made to `config.zon` after the
|
||||
rollback does nothing at all, silently, until the newer binary is back. If you
|
||||
have to stay on the old binary, use `nxdns import` to apply file changes, or drop
|
||||
the flag so the invocation matches what the binary actually does.
|
||||
|
||||
The schema note above still governs: a database stamped by a newer binary
|
||||
refuses to open, whatever mode either binary runs in.
|
||||
|
||||
## Changing settings, not the binary
|
||||
|
||||
An upgrade never re-reads `/etc/nxdns/config.zon`. After the first successful
|
||||
seed the file is ignored, and the start log says so:
|
||||
How you change a setting depends on which authority the service runs under.
|
||||
`nxdns run` in `ExecStart` means the database; `nxdns run --config FILE` means
|
||||
the file. The start log names it either way:
|
||||
|
||||
```
|
||||
info(config_bootstrap): configuration file ignored; the database is already configured
|
||||
info(nxdns): authority: database
|
||||
info(nxdns): authority: file (/etc/nxdns/config.zon)
|
||||
```
|
||||
|
||||
Change settings through the admin interface, through the API, or with an
|
||||
export–edit–import cycle against a stopped server:
|
||||
**In file mode**, edit the file, validate it, restart. The admin interface will
|
||||
refuse the change with a 403 naming the file, so there is nothing to get wrong:
|
||||
|
||||
```sh
|
||||
$EDITOR /etc/nxdns/config.zon
|
||||
nxdns check --config /etc/nxdns/config.zon
|
||||
systemctl restart nxdns
|
||||
```
|
||||
|
||||
**In database mode**, change settings through the admin interface, through the
|
||||
API, or with an export–edit–import cycle against a stopped server:
|
||||
|
||||
```sh
|
||||
nxdns export --out config-backup.zon
|
||||
$EDITOR config-backup.zon
|
||||
systemctl stop nxdns
|
||||
nxdns import config-backup.zon --force
|
||||
nxdns import config-backup.zon
|
||||
systemctl start nxdns
|
||||
```
|
||||
|
||||
`--force` is required here. A plain `import` into a database that already holds
|
||||
configuration fails with `import failed: DatabaseNotEmpty` and exits 2, so it
|
||||
cannot clobber a configured server by accident. What counts is what an operator
|
||||
set: client rows the DNS path materialised from traffic never trigger the
|
||||
refusal on their own.
|
||||
`import` needs no flag to add rows or to edit them. It needs `--allow-delete`
|
||||
only when applying the file would delete rows the database holds — including the
|
||||
case where you renamed something, since changing a group's name or an upstream's
|
||||
URL is a delete and an insert to the engine, not an edit. The refusal names the
|
||||
tables and rolls back:
|
||||
|
||||
> Verified on this host for the two `nxdns` lines, against a populated scratch
|
||||
> data directory:
|
||||
```
|
||||
FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
|
||||
import failed: DestructiveImport
|
||||
```
|
||||
|
||||
Stop the server first either way. `import` rewrites configuration underneath a
|
||||
process that read it at startup, and a running server picks up only some of it.
|
||||
|
||||
> Verified on this host against a populated scratch data directory, with
|
||||
> `--data-dir` pointing at it — that path is the only difference from the blocks
|
||||
> above:
|
||||
>
|
||||
> ```
|
||||
> $ nxdns import $SCRATCH/nxdns-config.zon --data-dir $SCRATCH/data
|
||||
> import failed: DatabaseNotEmpty
|
||||
> $ nxdns import $SCRATCH/etc/config.zon --data-dir $SCRATCH/dbmode
|
||||
> imported /…/config.zon
|
||||
> (exit 0)
|
||||
> $ nxdns import $SCRATCH/etc/smaller.zon --data-dir $SCRATCH/dbmode
|
||||
> FAIL import: this file would delete rows the database holds (upstreams 1); re-run with --allow-delete to apply it
|
||||
> import failed: DestructiveImport
|
||||
> (exit 2)
|
||||
> $ nxdns import $SCRATCH/nxdns-config.zon --data-dir $SCRATCH/data --force
|
||||
> imported /…/scratchpad/nxdns-config.zon
|
||||
> $ nxdns import $SCRATCH/etc/smaller.zon --data-dir $SCRATCH/dbmode --allow-delete
|
||||
> imported /…/smaller.zon
|
||||
> (exit 0)
|
||||
> ```
|
||||
>
|
||||
> The `systemctl stop`/`start` lines around them need root and an installed
|
||||
> service and were not run; `$EDITOR` is yours to run.
|
||||
> The first of those three is the additive case that needs no flag; the second
|
||||
> file replaced the upstream, which is an identity change and therefore a
|
||||
> delete. The `systemctl stop`/`start` lines need root and an installed service
|
||||
> and were not run; `$EDITOR` is yours to run.
|
||||
|
||||
## Upgrading to a build of your own
|
||||
|
||||
|
||||
Reference in New Issue
Block a user