341 lines
13 KiB
Markdown
341 lines
13 KiB
Markdown
# Troubleshoot nxdns
|
|
|
|
Symptoms an nxdns install actually produces, what to run to identify each one,
|
|
and what to change. Every symptom on this page was reproduced on the machine
|
|
that wrote it, and every diagnosis command was run there. Two details differ
|
|
from a real install and cannot be otherwise on that machine: it has no
|
|
installed service, so the log lines were read from a foreground run instead of
|
|
`journalctl -u nxdns`, and ports 53 and 8080 were occupied, so DNS and the API
|
|
were exercised on unprivileged ports. Fixes that need root are marked.
|
|
|
|
The exit codes themselves are listed in
|
|
[the CLI reference](../reference/cli.md).
|
|
|
|
## The service exits with code 2
|
|
|
|
**Symptom.** The process stops immediately. The last two lines are the error
|
|
and a pointer:
|
|
|
|
```
|
|
nxdns run failed: NoUsableUpstreams
|
|
run `nxdns check` to see the configuration in full
|
|
```
|
|
|
|
Exit 2 means the configuration is wrong and you can fix it. Every subcommand
|
|
uses the same definition, so a file `run` exits 2 on exits 2 from `check` and
|
|
`import` too.
|
|
|
|
**Diagnosis.**
|
|
|
|
```sh
|
|
nxdns check
|
|
```
|
|
|
|
`check` prints every problem it finds, not the first, and names the source it
|
|
checked on its first line.
|
|
|
|
**Fixes by cause.**
|
|
|
|
- `NoUsableUpstreams` — the database has no enabled upstream. On a fresh
|
|
install this means the seed file was missing or in the wrong place; the start
|
|
log says `no configuration file at '/etc/nxdns/config.zon'; using the
|
|
database as it is`. Write the seed file and start again against the still
|
|
empty database, or `nxdns import <file> --force`.
|
|
- `BadCertificate` — a DoH or DoT listener is enabled and its certificate or
|
|
key is unreadable, too large, unparseable, or the key does not belong to the
|
|
certificate. `run` names both paths before it exits:
|
|
`doh_server: '<cert>' + '<key>': certificate file is not readable`.
|
|
|
|
`check` catches this without starting a listener. It loads both PEM files and
|
|
tests the key against the certificate through the same code `run` uses, so it
|
|
fails on exactly what `run` would fail on. Reproduced here with a self-signed
|
|
pair and the key from a second, unrelated pair:
|
|
|
|
```
|
|
$ nxdns check --config config.zon
|
|
warning(tls_server): mbedtls_pk_check_pair failed: RSA - Key failed to pass the validity check of the library (-16896)
|
|
checking configuration file config.zon
|
|
FAIL doh_server.key_path: 'mismatched-key.pem': private key does not belong to the certificate
|
|
OK upstreams[0] https://cloudflare-dns.com
|
|
# exit 2
|
|
|
|
$ nxdns run --config config.zon --data-dir ./data
|
|
warning(tls_server): mbedtls_pk_check_pair failed: RSA - Key failed to pass the validity check of the library (-16896)
|
|
doh_server: 'cert.pem' + 'mismatched-key.pem': private key does not belong to the certificate
|
|
nxdns run failed: BadCertificate # exit 2
|
|
```
|
|
|
|
The `warning(tls_server)` line comes from mbedTLS on stderr and can appear
|
|
before the `checking` line, which is on stdout. A cert file containing
|
|
`not a certificate` fails the same way, with
|
|
`FAIL doh_server.cert_path: 'junk.pem': certificate PEM could not be parsed`.
|
|
An unreadable file reads
|
|
`FAIL doh_server.cert_path: '<path>': certificate file is not readable`.
|
|
|
|
Fix the path, the ownership, or the pair; see
|
|
[Enable DoH and DoT](enable-doh-and-dot.md).
|
|
- `BadRateLimit` — a rate limit or window is zero. `import` refuses such a
|
|
configuration, so this only reaches a database that was edited by hand.
|
|
- `BadBindAddress` — `dns.bind_ipv4` or `dns.bind_ipv6` is not an address of
|
|
that family.
|
|
|
|
## A seed file you just wrote is rejected
|
|
|
|
**Symptom.** A first start against an empty database prints the validation
|
|
problem and stops with exit 2:
|
|
|
|
```
|
|
FAIL groups: no group named 'default'; every unknown client is assigned to it
|
|
nxdns run failed: MissingDefaultGroup
|
|
run `nxdns check` to see the configuration in full
|
|
```
|
|
|
|
A syntax error behaves the same way:
|
|
|
|
```
|
|
FAIL config: 3:16: error: expected ',' after initializer
|
|
nxdns run failed: ParseZon
|
|
run `nxdns check` to see the configuration in full
|
|
```
|
|
|
|
So does a seed file whose upstream list is empty or all disabled:
|
|
|
|
```
|
|
FAIL upstreams: at least one upstream must be enabled
|
|
nxdns run failed: NoUpstreams
|
|
run `nxdns check` to see the configuration in full
|
|
```
|
|
|
|
`NoUpstreams` from a seed file is not the same fault as `NoUsableUpstreams`
|
|
above: the first is a file `run` refused, the second is a database `run`
|
|
accepted and found empty. Both are exit 2.
|
|
|
|
**Diagnosis.** Run the same file through `check`, which reports the same
|
|
problems and exits 2:
|
|
|
|
```sh
|
|
nxdns check --config /etc/nxdns/config.zon
|
|
```
|
|
|
|
**Fix.** Correct the file the diagnostics name and start again. The database is
|
|
still empty after a failed seed, so the next start re-reads the file. The exit
|
|
code no longer depends on which command read the file: all three of these files
|
|
were run through `run`, `check` and `import` here, and every one of the nine
|
|
combinations exited 2 with the same diagnostic.
|
|
|
|
## `nxdns check` fails on a server that is running fine
|
|
|
|
**Symptom.** The service is up and answering, but `nxdns check` on the same
|
|
machine exits 2 with one long line about a write-ahead log:
|
|
|
|
```
|
|
checking database /var/lib/nxdns/config.db
|
|
FAIL /var/lib/nxdns/config.db: uncheckpointed changes are waiting in /var/lib/nxdns/config.db-wal, and reading without writing would answer from the older settings in the main file; `nxdns run` applies them. A running nxdns normally holds this log, which is the usual reason to see this line.
|
|
```
|
|
|
|
Nothing is damaged. `check` opens `config.db` immutable so that it can never
|
|
write to it, and an immutable open ignores the write-ahead log. When that log
|
|
holds bytes, the newest settings are in it and the main file holds older ones,
|
|
so `check` refuses rather than grade stale values.
|
|
|
|
The log holds bytes after a configuration write that has not been checkpointed
|
|
yet, which on a running server means someone changed something through the web
|
|
interface or the API. A server that has only been answering queries has an empty
|
|
`config.db-wal` and `check` reads it normally — so this line comes and goes, and
|
|
its absence is not proof that nothing is running.
|
|
|
|
**Fix.** Check the exported configuration instead of the live file:
|
|
|
|
```sh
|
|
nxdns export --data-dir /var/lib/nxdns --out /tmp/current.zon
|
|
nxdns check --config /tmp/current.zon
|
|
```
|
|
|
|
`export` opens the database read/write and does see the log, so it renders the
|
|
settings that are actually in force. Stopping the service and checking again
|
|
works too: a clean shutdown checkpoints the log away.
|
|
|
|
> Reproduced here on a scratch data directory rather than `/var/lib/nxdns` —
|
|
> that path is the only substitution in the output above. nxdns was started on
|
|
> unprivileged ports; `config.db-wal` was 0 bytes and `check` exited 0; one
|
|
> `POST /api/blocklists` took it to 8272 bytes and `check` then printed the line
|
|
> above and exited 2; `export` from the same live directory succeeded and its
|
|
> output checked clean; and after a clean shutdown `check --data-dir` exited 0
|
|
> again.
|
|
|
|
## Port 53 is already taken
|
|
|
|
**Symptom.** The process exits 1, having named the socket it could not have:
|
|
|
|
```
|
|
cannot bind udp [::1]:53: AddressInUse
|
|
nxdns run failed: AddressInUse
|
|
```
|
|
|
|
A bind conflict is a runtime failure, not a configuration fault, so this is
|
|
exit 1 and `nxdns check` will not find it.
|
|
|
|
**Diagnosis.**
|
|
|
|
```sh
|
|
ss -lnup 'sport = :53'
|
|
ss -lntp 'sport = :53'
|
|
systemctl is-active systemd-resolved
|
|
```
|
|
|
|
On most systemd distributions the holder is `systemd-resolved`, which runs a
|
|
stub listener on `127.0.0.53:53` and on some setups binds `0.0.0.0:53`.
|
|
|
|
**Fix.** Turn off the stub listener and keep resolved for the host's own
|
|
lookups:
|
|
|
|
```sh
|
|
mkdir -p /etc/systemd/resolved.conf.d
|
|
printf '[Resolve]\nDNSStubListener=no\n' > /etc/systemd/resolved.conf.d/nxdns.conf
|
|
systemctl restart systemd-resolved
|
|
```
|
|
|
|
If `/etc/resolv.conf` is a symlink to `/run/systemd/resolve/stub-resolv.conf`,
|
|
repoint it at `/run/systemd/resolve/resolv.conf` so the host still resolves.
|
|
|
|
> Not verified on this host: this needs root, and `systemd-resolved` is
|
|
> inactive here with port 53 free, so the conflict could not be reproduced
|
|
> against it. The bind failure itself was reproduced by starting a second nxdns
|
|
> on a port the first already held, which is the same error path.
|
|
|
|
Do not fix this by pointing the host's `/etc/resolv.conf` at nxdns when that
|
|
host is where nxdns resolves its own upstream DoH and DoT hostnames. That is a
|
|
startup cycle, not a fix.
|
|
|
|
## The container restarts in a loop
|
|
|
|
**Symptom.** `docker compose ps` shows the container restarting, and the log is
|
|
one line repeated:
|
|
|
|
```
|
|
nxdns run failed: AccessDenied
|
|
```
|
|
|
|
**Diagnosis.**
|
|
|
|
```sh
|
|
docker inspect -f '{{.State.Status}} exit={{.State.ExitCode}} restarts={{.RestartCount}}' docker-nxdns-1
|
|
stat -c '%a %u:%g %n' deploy/docker/etc-nxdns/config.zon
|
|
```
|
|
|
|
Exit 1 with `AccessDenied` means the container could not read the seed file.
|
|
The container runs as uid 65532 and `/etc/nxdns` is mounted read-only, so a
|
|
file at mode 0600 owned by your own uid is unreadable to it and the container
|
|
cannot repair it.
|
|
|
|
**Fix.** Either make the file world-readable, when it holds no secret:
|
|
|
|
```sh
|
|
chmod 0644 deploy/docker/etc-nxdns/config.zon
|
|
```
|
|
|
|
or give it to the container's uid:
|
|
|
|
```sh
|
|
chown 65532:65532 deploy/docker/etc-nxdns/config.zon
|
|
chmod 0600 deploy/docker/etc-nxdns/config.zon
|
|
```
|
|
|
|
The 0644 path was verified here, including the recovery: after the `chmod` the
|
|
container started and answered queries. The `chown` needs root and was not run
|
|
here.
|
|
|
|
A container that exits 2 instead — `nxdns run failed: NoUsableUpstreams` after
|
|
`no configuration file at '/etc/nxdns/config.zon'` — has no seed file at all on
|
|
a fresh volume. Create `deploy/docker/etc-nxdns/config.zon` and bring it up
|
|
again; see [Install with Docker](install-with-docker.md).
|
|
|
|
## The container cannot reach its upstreams
|
|
|
|
**Symptom.** The container starts, but every query fails and `nxdns check`
|
|
inside it reports each upstream as unreachable.
|
|
|
|
**Diagnosis.** Look at what the host resolves with:
|
|
|
|
```sh
|
|
cat /etc/resolv.conf
|
|
```
|
|
|
|
**Fix.** If it points at the nxdns container, repoint it at a real resolver.
|
|
The container resolves its upstream DoH and DoT hostnames through the host's
|
|
DNS configuration, so pointing that at nxdns makes nxdns depend on itself to
|
|
start. LAN clients point at nxdns; the container's own host does not.
|
|
|
|
## The disk is filling up
|
|
|
|
**Symptom.** Writes stop but DNS keeps answering. The journal shows the
|
|
transition:
|
|
|
|
```
|
|
warning(disk_monitor): disk state ok -> critical: 33349095424 bytes free on /var/lib/nxdns
|
|
```
|
|
|
|
**Diagnosis.**
|
|
|
|
```sh
|
|
curl -s http://127.0.0.1:8080/api/health
|
|
```
|
|
|
|
`/api/health` needs no login and reports the state and what has been gated:
|
|
|
|
```json
|
|
{"status":"degraded","disk":{"state":"critical","free_bytes":33349079040,"db_bytes":180224,"log_bytes":0,"sample_failures":0},"upstreams":{"available":1,"total":1},"queries_dropped":0,"writer_failed":false,"refreshes_gated":1,"snapshot_generation":2}
|
|
```
|
|
|
|
`/metrics` carries the same free, database and log byte gauges as
|
|
`nxdns_disk_free_bytes`, `nxdns_disk_db_bytes` and `nxdns_disk_log_bytes`; the
|
|
state itself is on `/api/health`, not in the metrics output.
|
|
|
|
**What the state means.** The monitor samples free space and database sizes
|
|
once a minute. Below `disk.warn_free_mb` it logs the transition. Below
|
|
`disk.min_free_mb` it gates every non-essential write: the query logger holds
|
|
its batches, the client tracker stops persisting, and blocklist refreshes are
|
|
skipped and counted in `refreshes_gated`. Resolution never degrades because the
|
|
disk is full — this was verified by setting the thresholds above the free space
|
|
on the volume: the state went critical, a refresh was gated, and queries kept
|
|
being answered.
|
|
|
|
**Fix.** Recover space — lower `logging.retention_days`, or stop the service
|
|
and delete `querylog.db` — and writes resume on the next sample.
|
|
|
|
## Blocklists are not filtering
|
|
|
|
**Symptom.** Domains that should be blocked resolve normally.
|
|
|
|
**Diagnosis.** Read the startup line:
|
|
|
|
```sh
|
|
journalctl -u nxdns | grep 'serving on'
|
|
```
|
|
|
|
It ends in either `blocklist generation N` or
|
|
`unfiltered (no blocklist snapshot)`.
|
|
|
|
**Fix.** `unfiltered` means no snapshot loaded at all; the download or compile
|
|
warning that explains it is earlier in the same start. nxdns serves anyway on
|
|
purpose — a household loses more from DNS that refuses to start than from a
|
|
window of unfiltered answers.
|
|
|
|
A generation number with nothing being blocked is a different problem: the
|
|
snapshot loaded but has no sources in it. The line
|
|
`blocklist snapshot generation 1: 0 of 0 sources loaded` says exactly that. Add
|
|
a source in the admin interface, or in the seed file before the first start.
|
|
|
|
## A database stamped by a newer binary
|
|
|
|
**Symptom.** After putting an older binary back, it will not start:
|
|
|
|
```
|
|
warning(migrations): config.db is at schema version 99; this nxdns binary supports 2
|
|
nxdns run failed: SchemaTooNew
|
|
```
|
|
|
|
**Fix.** There is no downgrade. Import the export you took before upgrading
|
|
into a fresh data directory with the older binary; see
|
|
[Upgrade nxdns](upgrade.md).
|