milestone 31: concurrent upstream exchanges, dot session reuse, queue metrics
Gates / frontend (push) Successful in 1m26s
Gates / test (push) Successful in 1m55s
Gates / test-aarch64 (push) Failing after 3h1m8s
Gates / package (push) Successful in 3m55s
Gates / container (push) Successful in 15s
CI / gates (push) Failing after 3h21m29s

This commit is contained in:
2026-08-22 19:54:02 +02:00
parent 648d9b4496
commit 025edbb093
14 changed files with 2099 additions and 188 deletions
+3 -1
View File
@@ -324,7 +324,9 @@ Per-group boolean. Rewrites known engine domains to their safe-search CNAME targ
## 9. Upstream Resolution
- Schemes: `https://…` → DoH, `tls://host:853` → DoT.
- Ordered by priority; sequential attempt; per-upstream failure counters; exponential backoff with jitter; success resets.
- Ordered by priority; one query tries upstreams sequentially; per-upstream failure counters; exponential backoff with jitter; success resets.
- Concurrency is per upstream, not per pool: each entry owns `slots_per_entry = 8` leaf clients behind a semaphore, so at most 8 exchanges are in flight against one upstream at a time and the rest wait for a slot rather than for the whole entry. The slot count is compiled, not configured. A task that finds every slot taken counts itself in `queued_total`/`queued_seconds_total` on both exits — acquisition and cancellation — so a query that burned its total budget waiting is visible on `/metrics` (`nxdns_upstream_queued_total`, `nxdns_upstream_in_flight`, `nxdns_upstream_slots`) instead of being an unexplained SERVFAIL. The counters are admission samples, not an exact queue length.
- DoT keeps its connection: a `DotClient` holds one TLS session open across exchanges rather than dialing and handshaking per query. Staleness is detected at use, never by a keepalive timer — a reused session that fails before the first response byte with a connection-lifecycle error is redialed once and retried, and that recovery counts in `nxdns_upstream_reuse_recoveries_total` without touching health or emitting a diagnostics episode. Every other failure closes the session and classifies as before.
- `UpstreamHealth` per upstream: last_success_at, last_error_at, last_error_message, rolling success rate, consecutive failures, backoff-until. This is routing state: it drives failover and backoff, and is exposed through `/metrics` and `nxdns check`. There is no per-upstream API surface: `/api/health` reports the pool as available-of-enabled, and a failing upstream is a Diagnostics episode (`upstream.exchange`).
- DoH client: `std.http.Client` with `content-type/accept: application/dns-message`; strict status + payload checks.
- `platform/tls_client.zig` enforces per-connection read/write deadlines, classifies TLS errors explicitly, retries with backoff. Integration tests cover timeout/hang scenarios so compiler upgrades can't silently regress them.