6.2 KiB
Measure performance
tools/bench.zig measures the three things nxdns can measure in-process:
blocklist lookup latency, cache-hit latency, and blocklist compile throughput.
Sustained query rate is not one of them — that one is end-to-end and needs a
load generator pointed at a running server.
The numbers this project treats as targets, and the numbers measured so far, are in the performance reference. Why those targets exist and why CI does not gate on them is in performance and testing.
Run the whole bench
zig build bench -Doptimize=ReleaseFast
That runs all three suites with the defaults: 1,000,000 domains, 200,000
iterations per suite, seed 0x5eed. It takes minutes, most of it generating and
loading the million-domain list.
-Doptimize=ReleaseFast is not optional if you want the numbers to mean
anything. A Debug build says so before it prints:
warning: Debug build; run with -Doptimize=ReleaseFast for meaningful numbers
Run one suite, smaller
Everything after -- goes to the harness. A suite name selects one of
filter, cache, compile (the default is all), and --domains /
--iters shrink the load:
zig build bench -Doptimize=ReleaseFast -- filter --domains=100000 --iters=20000
nxdns bench suite=filter domains=100000 iters=20000 seed=0x5eed optimize=ReleaseFast
suite ops p50(us) p95(us) p99(us) max(us)
filter 20000 0.14 0.25 0.27 0.51
blocked 6670/20000, Snapshot.memoryBytes 3.0 MiB, VmRSS 5.4 MiB
target p95 < 1ms: PASS
target VmRSS < 100 MiB: PASS
A reduced run is good for checking the harness works and for a rough
regression signal. It is not a result: the memory figure scales with
--domains, so 100,000 domains says nothing about the 1,000,000-domain memory
target.
The other two suites:
zig build bench -Doptimize=ReleaseFast -- cache --iters=20000
suite ops p50(us) p95(us) p99(us) max(us)
cache 20000 0.12 0.21 0.23 0.54
hits 10000/20000, DnsCache.memoryBytes 4.3 MiB, VmRSS 6.2 MiB
target p95 < 5ms: PASS
zig build bench -Doptimize=ReleaseFast -- compile --domains=100000
suite ops p50(us) p95(us) p99(us) max(us)
compile 100000 wall 11.512ms, 8686215 lines/s, 100000 domains kept (informational)
--seed=N changes the generated domains and the query order; the default is
0x5eed, so two runs on the same machine are comparable. --domains caps at
4,000,000, and the compile suite additionally refuses more than 2,000,000 —
the compiler's own limit.
An argument the harness does not recognise stops it before any measuring:
error: unknown argument 'nosuch'
usage: zig build bench -Doptimize=ReleaseFast -- [filter|cache|compile|all] [--domains=N] [--iters=N] [--seed=N] [--assert]
Read the output
p50/p95/p99/maxare per-operation microseconds, nearest-rank over every iteration. What one operation means differs per suite: forfilterit is normalising a name plus evaluating it against the snapshot; forcacheit is building the key, getting the entry and stamping the response id.blocked N/Mandhits N/Mare sanity counters. The harness aborts if either is zero — a suite that never hits its own path measures nothing.- Two memory figures appear on purpose.
Snapshot.memoryBytesandDnsCache.memoryBytesare the in-repo accounting of those structures;VmRSSis what the kernel holds resident for the whole process, allocator slack and code included. The truth is between them, and the memory target is judged onVmRSS. target ...: PASS/FAILlines appear for the targets a suite covers. On a plain run they are informational and the exit code stays 0.
Fail the run when a target is missed
--assert turns those lines into an exit code — 1 when any target was
exceeded, 0 otherwise. This is meant for an acceptance run on hardware you
control, not for CI:
zig build bench -Doptimize=ReleaseFast -- --assert
The full-scale form is the one worth asserting on, because the memory target only means something at a million domains. On this development host the reduced form was used to check the flag itself:
zig build bench -Doptimize=ReleaseFast -- filter --domains=100000 --iters=20000 --assert
filter 20000 0.14 0.26 0.27 2.42
blocked 6670/20000, Snapshot.memoryBytes 3.0 MiB, VmRSS 5.4 MiB
target p95 < 1ms: PASS
target VmRSS < 100 MiB: PASS
Not verified on this host at full scale: the plain
zig build bench -Doptimize=ReleaseFast -- --assert above was not run during
the writing of this page — the default run takes minutes. The reduced runs
shown were all executed as written. The full-scale numbers already recorded for
this host are in the performance reference.
Measure sustained query rate
The bench harness cannot do this. Query rate is a property of the whole server — sockets, upstreams, the query log writer — so it has to be driven from outside, against the real binary, on the machine you care about.
Start nxdns with real blocklists configured, then drive it from another host on
the LAN with a DNS load generator such as dnsperf:
dnsperf -s 192.168.1.10 -p 53 -d queries.txt -c 20 -Q 200 -l 60
Read the client's own rate and the server's /metrics together: a load
generator that reports 200 qps while the server counts fewer has lost queries
somewhere, and that is the interesting number.
Not verified on this host: dnsperf is not installed here and the target
platform is a Raspberry Pi 5, not this development machine. The command above
is the shape of the measurement, not a transcript.
Where to run it
The target platform is a Raspberry Pi 5. Numbers from a development x86_64 box
do not transfer — the Pi's Cortex-A76 is far slower — so a passing run here is
evidence the harness works and a baseline for spotting regressions on the
machine development happens on, and nothing more. Run --assert on the Pi,
where the numbers mean something.