# Milestone 6: Cache, Rate Limiter, Query Logger, Retention, Disk Monitor, Log Sink PLAN Phase 6. Build every runtime component between the resolver core and the web layer: the TTL cache (§8), the per-client DNS rate limiter (§10), the async query logger with backpressure (§11.4), retention (§11.5), the disk monitor and the rotating log sink with upstream-error dedup (§11.6). Nothing is wired into the query path — composition is Phase 7. Every component is built complete, tested standalone, and exposed through accessors Phase 7 and Phase 8 will consume. ## Sessions Eight sessions. Wave 1 starts five in parallel: S1 (cache), S2 (rate limiter), S3 (queries repo), S5 (statfs + disk monitor), S6 (log sink). Wave 2 starts when S3 verifies: S4 (logger) and S7 (retention) in parallel — both consume S3's repo and nothing of each other. S8 (integration) runs last. The orchestrator — not any session — wires `src/tests.zig` and installs the log sink's `std_options` hook in `src/main.zig`. (`validate.zig:221` already rejects `rate_window_seconds` outside 1–`max_rate_window_seconds`; no rule is missing.) Every later session is written against **this spec**, not against another session's source. Where this spec and a built file disagree after a session verifies, the built file wins and the orchestrator records the difference in "## As built". ## Design invariants - **Pure core.** `src/cache/dns_cache.zig` and `src/server/rate_limiter.zig` take timestamps as parameters and hold no `std.Io`, no clock, no socket. `std.Io` appears only in `src/storage/logger.zig`, `src/storage/disk_monitor.zig`, `src/storage/retention.zig`, `src/platform/logging.zig`, `src/platform/statfs.zig` (its test), and the S8 test file. - **Neither the cache nor the limiter is thread-safe.** Phase 7 decides the locking when it wires them. Both files state that in their module doc comment. - **No allocation on the hit path.** `DnsCache.get` and `RateLimiter.check` allocate nothing. `DnsCache.put` allocates exactly one buffer for the stored response bytes. - **Every failure mode is counted.** Dropped log entries, refused queries, evictions, expired entries, deduplicated log lines, failed disk samples — a counter each, readable by Phase 8. - **No `std.log.err` in any file this milestone writes.** A condition returned as a typed error logs at `warn` at most. `err` stays reserved for swallowed failures, and this milestone swallows none. - **Timestamps follow the house style** (health.zig): `std.Io.Timestamp` parameters, arithmetic on `.nanoseconds`, wall seconds for DB rows via `Clock.real.now(io).toSeconds()`, monotonic decisions on `.awake`, long schedules on `.boot`. - **Background loops follow `manager.runScheduler`'s shape**: `Cancelable!void`, non-cancel failures log at `warn` and the loop continues, cancellation returns. Phase 7 starts them. ## Resolved PLAN ambiguities (rulings for this milestone) 1. **Client-row pruning is Phase 7.** §7.2's "retention drops `hand_edited=0` clients" deletes from `config.db`, which §3.6 walls off from retention churn. The rows it prunes are created by auto-materialization, which is Phase 7; the pruning ships with it. §11.5 retention here touches `querylog.db` only. 2. **Free space drives the thresholds; sizes are gauges.** The monitor samples `statvfs` free bytes for the warn/critical decision and records DB and log-directory sizes as gauges for Phase 8's health payload. Both are sampled in the same pass. 3. **Flush tuning is comptime.** `flush_batch = 100`, `flush_interval_ms = 100` as public constants in `logger.zig`. §12.1 defines no config keys for them and a household deployment never tunes them. 4. **The SSE seam is the transformed entry.** Privacy transforms run at enqueue, so every entry in the queue is already presentable. Phase 8 subscribes by wrapping `Logger.log` — no hook, no callback, no dead vtable. 5. **The positive-TTL sanity ceiling is 86 400 seconds**, the same `max_ttl_seconds` constant `validate.zig:196` already enforces on `blocking.ttl` and `cache.negative_ttl_max`. Comptime constant in `dns_cache.zig` with a comment naming that origin. No new config key. 6. **`negative_ttl_max = 0` disables negative caching** (the response is not stored), not the cap. `validate.zig` already caps the field at `max_ttl_seconds`, so "no cap" would be expressible as 86 400 anyway; "0 disables" only has meaning as an off switch. 7. **`cache.size` counts entries**, not bytes. 10 000 entries at the observed household response sizes is single-digit megabytes; a byte budget would add config surface for nothing. 8. **The DNS limiter is a fixed window.** §10 reserves the token bucket for the API (Phase 8); "1000 per 60 s" plus "stale-key windowed sweep" describes a counter that resets each window. A fixed window admits at most 2× the limit across a window boundary, which is acceptable for abuse protection at household scale and costs half the state of a sliding window. 9. **Upstream-error dedup lives in the log sink.** It is a property of how lines reach the operator, not of the upstream module: `platform/logging.zig` deduplicates warn-and-above lines from the upstream scopes at one per key per minute. No milestone-3 file is edited. 10. **Schedules: retention daily, checkpoint after every prune, VACUUM every 7th pass.** §11.5 names no interval. Daily matches the granularity of `retention_days`; `wal_checkpoint (TRUNCATE)` after each prune keeps the WAL from monopolizing the disk the monitor watches; a nightly full VACUUM would rewrite the whole file on an SD card every day, so VACUUM runs on every 7th retention pass. ## Verified 0.16.0 stdlib facts (see specs/research/zig-0.16-api-notes.md, "Phase 6 verifications") - `std.Io.Queue(Elem)` (Io.zig:2184): bounded ring over a caller array. `put(q, io, elems, 0)` never blocks and returns 0 when full (Io.zig:2218); `getOne` blocks; `close(io)` unblocks it with `error.Closed`. Elements are copied raw (Io.zig:2189) — **queue elements must be self-contained values**, no slices into caller memory. - `std.Io.Condition` has **no `timedWait`**. The timed primitive is `Event.waitTimeout` (Io.zig:1827). The logger's flush interval uses `Select` racing `getOne` against a `Clock.Duration` sleep, the `fetchWithin` pattern from manager.zig:722. - **No statvfs in std.** `src/platform/statfs.zig` declares `extern fn statvfs` against libc (every build already links libc for sqlite; glibc and musl share the 64-bit POSIX layout). Free bytes = `f_bavail * f_frsize`. - **No append mode.** Log sink append: `openFile(io, path, .{ .mode = .write_only })`, `file.writer(io, &buf)`, `w.seekTo(try file.length(io))`. `Dir.rename` + `Dir.deleteFile` rotate. One owner per writer — the sink serializes behind its own mutex. - **No fixed-capacity hash map in std.** Bounded structures use a fixed slot array plus `std.HashMapUnmanaged` from key to slot index with `ensureTotalCapacity` at init; the CLOCK hand walks the slot array, never the map. Map modification invalidates live iterators (hash_map.zig:496). - `packet.decrementTtls(bytes, elapsed_seconds) ParseError!?u32` (packet.zig:234) ages every record in place, skips OPT, validates before writing, returns the smallest resulting TTL; on error the buffer is partly aged and must be discarded. With `elapsed = 0` it is a pure validate-and-min-TTL pass. `packet.setId` (packet.zig:202) rewrites the transaction ID. - `db.zig`: `Stmt.reset` clears bindings — rebind everything each iteration. `Db`/`Stmt` are not thread-safe; the logger's writer task owns its `Db` handle exclusively. A batch is one `Tx.begin` + reset/rebind/step loop + `commit`. - `address.NetAddress.Key` is `[17]u8` covering both families; `fromIp` folds IPv4-mapped IPv6 to `.ip4`, so one client is one key. --- ## Session S1: DNS cache — `src/cache/dns_cache.zig` New directory `src/cache/`. Imports: `std`, `../dns/packet.zig`, `../dns/types.zig`, `../dns/record.zig`, `../config/model.zig`. No `std.Io`. ### S1.1 Key ```zig /// qname (lowercased ASCII, no trailing dot) ++ 0x00 ++ qtype BE ++ qclass BE ++ do byte /// ++ ecs_len byte ++ ecs bytes. ECS participates only when the caller passes it (Phase 7 /// passes the forwarded subnet only under ecs_mode=forward, PLAN §8). pub const max_key_len = 255 + 1 + 2 + 2 + 1 + 1 + 40; pub fn buildKey(buf: *[max_key_len]u8, qname: []const u8, qtype: u16, qclass: u16, do_bit: bool, ecs: ?[]const u8) []const u8; ``` `buildKey` lowercases `A`–`Z` while copying and strips one trailing dot. Asserts `qname.len <= 255` and `ecs.?.len <= 40` — the caller has already validated both through the packet parser. ### S1.2 Classification ```zig pub const Class = struct { ttl_seconds: u32, negative: bool }; /// Decides whether a response is cacheable and for how long. Null means do not cache. pub fn classify(response: []const u8, negative_ttl_max: u32) ?Class; ``` Rules, in order: parse the header — rcode NOERROR with at least one answer record is positive; rcode NXDOMAIN, or NOERROR with zero answers, is negative; every other rcode returns null. Positive: `ttl = min(packet.decrementTtls(copy, 0) result, max_ttl)` — since `decrementTtls` mutates, `classify` runs it on a stack copy (`[65535]u8` is too big for the stack; instead compute the min TTL by iterating records via the parser without mutation — `record` module iteration, skip OPT). A parse error returns null. Negative: min(SOA record TTL, SOA MINIMUM) from the authority section per RFC 2308 §5, capped by `negative_ttl_max`; no SOA → null; `negative_ttl_max == 0` → null (ruling 6). A TTL of 0 → null. A TTL with the top bit set is read as zero per RFC 2181 §8 (packet.zig:220 leaves that to the caller), on both paths — such a response is not cached. ### S1.3 Storage and API ```zig pub const Config = struct { size: u32, negative_ttl_max: u32 }; pub const Stats = struct { hits, misses, inserts, evictions, expirations, invalid_hits: u64 = 0 }; pub const DnsCache = struct { pub fn init(gpa: Allocator, config: Config) Allocator.Error!DnsCache; // allocates slots + index once pub fn deinit(self: *DnsCache) void; /// Stores a clone of `response`. A key already present is replaced in place. pub fn put(self: *, now_s: i64, key: []const u8, response: []const u8, class: Class) Allocator.Error!void; /// Copies the entry into `out`, rewrites nothing but the TTLs (caller does setId), /// returns null on miss/expiry. `out.len >= response len` is the caller's problem; /// a too-small `out` is a miss counted under `misses`. pub fn get(self: *, now_s: i64, key: []const u8, out: []u8) ?[]u8; /// Removes expired entries; returns how many. Phase 7 schedules it. pub fn sweep(self: *, now_s: i64) u32; pub fn len(self: *const) u32; pub fn memoryBytes(self: *const) usize; stats: Stats, }; ``` Slot: inline key `[max_key_len]u8` + `key_len` (no allocation), `bytes: []u8` (gpa-owned clone), `stored_at_s: i64`, `expires_at_s: i64`, `negative: bool`, `referenced: bool`, `occupied: bool`. Index: `HashMapUnmanaged` keyed by the slot's key slice (context hashes the slice) to `u32` slot index, `ensureTotalCapacity(config.size)` at init — no rehash ever. CLOCK: on `get` hit set `referenced`; when `put` finds no free slot, advance the hand clearing `referenced` until an unreferenced slot, evict it (`evictions += 1`). Expired-on-access is freed and counted under `expirations`, then treated as a miss/free slot. `get` computes `elapsed = now_s - stored_at_s`, copies, then `packet.decrementTtls(copy, elapsed)`. An error from it removes the entry, counts `invalid_hits`, returns null — a poisoned entry never serves twice. ### S1.4 Tests (in-file, ≥ 14) buildKey lowercases and strips the dot; distinct qtype/qclass/DO/ECS produce distinct keys; ECS absent vs present differ; classify: positive min-TTL over multiple answers; ceiling applied; NXDOMAIN uses SOA minimum; capped by negative_ttl_max; `negative_ttl_max = 0` returns null; SERVFAIL null; zero TTL null; put/get round-trip with TTL decrement asserted via re-parse; expiry is a miss and counted; CLOCK evicts the unreferenced entry, not the referenced one; replace-in-place does not grow `len`; sweep removes only expired; `checkAllAllocationFailures` on init+put. ### S1.5 Acceptance - [ ] `zig fmt --check` + `zig ast-check` clean; in-file tests pass standalone via a src/-level temp root. - [ ] `grep -n "std.Io" src/cache/dns_cache.zig` hits doc comments at most. - [ ] `get` performs zero allocations (no allocator parameter exists on it). --- ## Session S2: rate limiter — `src/server/rate_limiter.zig` Imports: `std`, `../platform/address.zig`. No `std.Io`. ```zig pub const Config = struct { limit: u32, window_seconds: u32 }; pub const Stats = struct { allowed, refused, untracked: u64 = 0 }; pub const RateLimiter = struct { pub fn init(gpa: Allocator, config: Config) Allocator.Error!RateLimiter; // capacity fixed at init pub fn deinit(self: *) void; /// True = process the query; false = answer REFUSED. Never errors, never allocates. pub fn check(self: *, now: std.Io.Timestamp, key: address.NetAddress.Key) bool; /// Drops entries whose window ended before the previous full window. Phase 7 schedules it. pub fn sweep(self: *, now: std.Io.Timestamp) u32; stats: Stats, }; pub const max_clients = 4096; ``` Fixed window per key: entry `{window_start_ns: i96, count: u32}` in a `HashMapUnmanaged` with capacity `max_clients` ensured at init. `check`: if the entry's window has ended, reset it to the current window. `count >= limit` → refused. Map full and key unknown: allow, count `untracked` — refusing unseen clients because the table is full would let 4096 attackers deny every new device, and a household LAN never has 4096 honest clients; the counter makes the state visible. `sweep` walks and removes stale entries (collect keys first — map modification invalidates iterators). `std.Io.Timestamp` is a plain value type (`{nanoseconds: i96}`) — using it keeps this file pure; callers pass `.awake` timestamps. Tests (≥ 8): allows up to limit, refuses limit+1; new window resets; v4 and v6 keys independent; v4-mapped v6 shares the v4 bucket (via `address.fromIp` in the test); sweep removes only stale; map-full behavior allows and counts; stats add up; window arithmetic at i96 scale (a timestamp far from zero). Acceptance: fmt/ast-check clean; no allocation in `check` (signature proves it); in-file tests pass standalone. --- ## Session S3: queries repo — `src/storage/repositories/queries_repo.zig` Imports: `std`, `../db.zig`. Pure DB code, no `std.Io`. Follows the milestone-4 repository idiom (free functions over `*db.Db` — read `groups_repo.zig` first), plus one long-lived-statement struct the flush loop owns (`db.zig:360` names this file as the reason no statement cache exists). ```zig pub const Row = struct { timestamp: i64, domain: []const u8, // already privacy-transformed by the logger client_ip: []const u8, // ditto qtype: ?u16, blocked: bool, block_reason: ?[]const u8, response_time_us: ?i64, cache_hit: ?bool, upstream: ?[]const u8, }; /// Owns the prepared statements of the flush loop. Init once, reuse per batch. pub const BatchWriter = struct { pub fn init(database: *db.Db) db.Error!BatchWriter; pub fn deinit(self: *) void; /// One transaction. Domains are interned via INSERT OR IGNORE + SELECT id. pub fn writeBatch(self: *, rows: []const Row) db.Error!void; }; pub fn pruneOlderThan(database: *db.Db, cutoff_ts: i64) db.Error!i64; // rows deleted pub fn checkpointTruncate(database: *db.Db) db.Error!void; // PRAGMA wal_checkpoint(TRUNCATE) pub fn vacuum(database: *db.Db) db.Error!void; pub fn countRows(database: *db.Db) db.Error!i64; pub fn countDomains(database: *db.Db) db.Error!i64; ``` `writeBatch` on an empty slice returns without opening a transaction. `pruneOlderThan` deletes `query_log` rows only — orphaned `domains` rows stay (they are a dimension table; re-interning is cheaper than referential garbage collection, and §11.3 sets no requirement). Statements: three in `BatchWriter` (insert-or-ignore domain, select domain id, insert row), reset+rebind per use. Tests (≥ 8, all against `:memory:` + `querylog_schema.ddl`): batch inserts rows and interns domains once; second batch reuses the domain id; nullable columns round-trip null and value; empty batch writes nothing; prune deletes strictly-older rows only; prune returns the count; checkpoint and vacuum execute without error on a file DB (use `:memory:` where legal, a tmp file where WAL is needed — vacuum works on both); countRows/countDomains agree with inserts. Acceptance: fmt/ast-check clean; tests pass standalone via src/-level temp root with `-lc -lsqlite3`. --- ## Session S4: async query logger — `src/storage/logger.zig` (needs S3) Imports: `std`, `db.zig`, `repositories/queries_repo.zig`, `../config/model.zig`, `disk_monitor.zig` (S5's file — wave 2 starts after wave 1 verifies, so it exists). ### S4.1 Entry — self-contained (Queue copies raw bytes) ```zig pub const Entry = struct { timestamp: i64, domain_buf: [253]u8, domain_len: u8, client_buf: [45]u8, client_len: u8, // RFC 5952 text, fits any v6 qtype: ?u16, blocked: bool, reason_buf: [32]u8, reason_len: u8, // "" = null in the DB response_time_us: ?i64, cache_hit: ?bool, upstream_buf: [64]u8, upstream_len: u8, // "" = null pub fn domain(self: *const) []const u8; // + client/reason/upstream accessors }; ``` ### S4.2 Logger ```zig pub const flush_batch = 100; pub const flush_interval_ms = 100; pub const hidden_marker = "hidden"; pub const Logger = struct { /// `queue_buf.len` is the backpressure cap: pass cfg.query_log_buffer_max entries. pub fn init(cfg: model.Logging, queue_buf: []Entry) Logger; /// Applies the privacy transforms, then enqueues without blocking. A full queue drops /// the OLDEST unflushed entry (PLAN §11.4) and increments `queries_dropped`. pub fn log(self: *, io: std.Io, entry: Entry) void; /// The writer task. Owns `database` for its whole life. Runs until `shutdown`. pub fn runWriter(self: *, io: std.Io, database: *db.Db, monitor: ?*disk_monitor.Monitor) std.Io.Cancelable!void; /// Closes the queue; runWriter drains what remains, flushes, and returns. pub fn shutdown(self: *, io: std.Io) void; queries_dropped: std.atomic.Value(u64), rows_written: std.atomic.Value(u64), batches_gated: std.atomic.Value(u64), }; ``` - Privacy (§11.4): `hide_domains` replaces the domain with `hidden_marker`; `hide_client_ips` likewise. Applied inside `log`, so the queue never holds the real value — the transform runs before persistence and before any future SSE fanout by construction (ruling 4). - Drop-oldest: `put(io, &.{entry}, 0)`; on 0, `get` one entry nonblocking (min 0), count it dropped, retry the put — until the put succeeds, one drop counted per failed attempt. A fixed attempt cap is wrong under contention (one call would pay for two old entries plus its own). The only early exit is a zero-capacity queue, which has nothing to drop. - `runWriter` loop: `getOne` (blocking; `error.Closed` → final drain + flush + return). Then accumulate up to `flush_batch` entries: nonblocking `get` first; while short of the batch and the interval clock (started at the first entry, `.awake`) has time left, race `getOne` against the remaining interval via `Select` (manager's `fetchWithin` shape). Flush: if `monitor != null and !monitor.?.writesAllowed()`, do not write — hold the batch, sleep 1 s (`.awake`), re-check; count each held cycle under `batches_gated`. §11.6: log flushes stop at critical, the queue keeps dropping oldest behind them. On `writeBatch` error: log at `warn`, drop the batch (it is expendable log data; blocking would fill the queue), continue. - Entries convert to `queries_repo.Row` at flush (slices point into the batch array — safe, the batch lives across the call). Tests (≥ 8): transforms applied at enqueue (inspect the queue copy); drop-oldest drops the oldest and counts; entry accessors round-trip; batch conversion maps "" to null; a `runWriter` smoke over `:memory:` — enqueue N, shutdown, all N in the DB (single-threaded: run the writer via `io.concurrent` in the test with a Threaded instance — this is the one in-file test that touches `std.Io`; keep it gated under a plain `test` since it needs no network); gating holds rows back and releases them (fake monitor state flip); writeBatch failure drops and continues (close the DB under the writer? if unreachable in-file, mark for S8). Acceptance: fmt/ast-check clean; `grep -c "std.log.err" src/storage/logger.zig` = 0. --- ## Session S5: statfs + disk monitor — `src/platform/statfs.zig`, `src/storage/disk_monitor.zig` ### S5.1 `src/platform/statfs.zig` ```zig pub const StatVfs = extern struct { f_bsize: c_ulong, f_frsize: c_ulong, f_blocks: u64, f_bfree: u64, f_bavail: u64, f_files: u64, f_ffree: u64, f_favail: u64, f_fsid: c_ulong, f_flag: c_ulong, f_namemax: c_ulong, __reserved: [6]c_int, }; extern fn statvfs(path: [*:0]const u8, buf: *StatVfs) c_int; pub fn freeBytes(path: [:0]const u8) error{StatFailed}!u64; // f_bavail * f_frsize ``` The 64-bit glibc and musl `struct statvfs` layouts agree field-for-field; both build targets are 64-bit. One test: `freeBytes(".")` returns a nonzero value (gated `-Dintegration`? no — it is hermetic and deterministic-enough: assert `> 0`; a full disk failing CI is a real signal). ### S5.2 `src/storage/disk_monitor.zig` ```zig pub const State = enum(u8) { ok, warn, critical }; pub const Gauges = struct { free_bytes, db_bytes, log_bytes: u64 }; pub fn classify(free_bytes: u64, cfg: model.Disk) State; // pure; < min → critical, < warn → warn pub const Monitor = struct { pub fn init(cfg: model.Disk, data_dir: std.Io.Dir, data_path: [:0]const u8, log_dir_path: ?[:0]const u8) Monitor; /// One sample: statvfs on data_path, sizes of *.db/-wal/-shm under data_dir and of the /// log dir. Stores state + gauges atomically. Failures count `sample_failures`, keep the /// previous state, and log at warn (deduped by the sink). pub fn sample(self: *, io: std.Io) void; pub fn state(self: *const) State; // atomic load pub fn writesAllowed(self: *const) bool; // state != .critical pub fn gauges(self: *const) Gauges; // atomics, individually consistent /// Loop: sample, sleep 60 s (.boot), repeat. Phase 7 starts it. pub fn run(self: *, io: std.Io) std.Io.Cancelable!void; sample_failures: std.atomic.Value(u64), }; pub const sample_interval_s = 60; ``` State is `std.atomic.Value(u8)`; a reader never sees a torn value. On the ok↔warn↔critical edges log one line at `warn` (state changes only, not every sample). The blocklist-update gate (§11.6 "stop non-essential writes") is consumed by the logger now and by the manager in Phase 7 — no milestone-5 file is edited here. Tests: `classify` boundary cases (exactly min, exactly warn, between, above); state-change logging is observable via state transitions (assert states, not logs); init leaves `.ok` with zero gauges; gauge arithmetic (a fixture dir with files of known sizes, using `std.Io.Dir` in a plain test with a Threaded instance). Acceptance: fmt/ast-check clean on both files; `classify` covered exhaustively. --- ## Session S6: log sink — `src/platform/logging.zig` The custom `std.log` sink: level filter, stderr or rotating file output, upstream-error dedup (§11.6). Imports: `std`, `../config/model.zig`. ```zig /// Matches std.Options.logFn. Before install(), passes through to stderr formatting. pub fn logFn(comptime level: std.log.Level, comptime scope: @Type(.enum_literal), comptime format: []const u8, args: anytype) void; /// Called once from main() after config parse. Not called by tests. pub fn install(io: std.Io, cfg: model.Logging) void; pub fn deinstall() void; // flushes and closes the file; for tests and shutdown pub const Stats = struct { lines_written, lines_deduped, rotations, sink_errors: u64 }; pub fn stats() Stats; ``` Verify the exact `std.Options.logFn` signature against lib/std/std.zig + log.zig before writing — do not trust this spec's sketch. - Global sink state behind a `std.Thread.Mutex`? **No** — `std.Io.Mutex` needs an io; `logFn` receives none. Store the installed `std.Io` by value in the global state (it is a plain interface value); guard the whole sink with `std.debug` -style spinlock? Verify what std.log's default lock does in 0.16 and mirror it. If std provides no lock for custom logFn, use a file-scope `std.Thread.Mutex` (std.Thread still exists for this — verify) or an atomic spin guard; state the choice in a comment. This is the one deliberately-open point in this spec: resolve it against the stdlib source and record the resolution for "As built". - Level: drop below `cfg.level` (mapping model's level enum to std.log.Level — write it). - Dedup: for `warn`+ lines whose scope is one of `.doh_client`, `.dot_client`, `.pool`, `.forward_client`: key = scope + formatted message truncated to 96 bytes; a key seen within the last 60 s (`.awake`) is dropped and counted. Fixed 64-entry table, oldest-stamp replacement. - File mode: append (seekTo end-of-file pattern), byte counter from `file.length` at open; a line that would cross `cfg.maxLogBytes()` triggers rotation first: delete `.{max_files-1}`, shift `.N`→`.N+1`, rename live → `.1`, reopen fresh. All under the sink lock. Rotation or write failure: fall back to stderr for that line, count `sink_errors`, keep trying the file next line. - stderr mode: format + single `write` per line (std.debug.lockStdErr equivalent — verify). Tests: the pure pieces — dedup table admits first, drops repeat inside 60 s, admits after; level mapping; rotation name arithmetic (`rotatedName(buf, path, n)`). File behavior lands in S8 (needs a real dir). `logFn` itself is NOT exercised in tests (installing a sink under the test runner would eat the harness's own logs). Acceptance: fmt/ast-check clean; no `std.log` call inside the sink itself (it would recurse); dedup + rotation arithmetic covered. --- ## Session S7: retention — `src/storage/retention.zig` (needs S3) Imports: `std`, `db.zig`, `repositories/queries_repo.zig`, `../config/model.zig`. ```zig pub const vacuum_every_passes = 7; pub const Stats = struct { passes, rows_pruned, checkpoints, vacuums: u64 = 0 }; pub const Retention = struct { pub fn init(cfg: model.Logging) Retention; /// One pass: prune rows older than now - retentionSeconds, checkpoint, and on every 7th /// pass vacuum. DB errors log at warn and the pass counts as done (next pass retries). pub fn runOnce(self: *, io: std.Io, database: *db.Db) void; /// Daily loop (.boot), first pass immediately. `database` must be a dedicated /// connection no other task uses while the loop runs; Phase 7 opens it. pub fn run(self: *, io: std.Io, database: *db.Db) std.Io.Cancelable!void; stats: Stats, }; pub const pass_interval_s = 86_400; ``` `runOnce` computes the cutoff from `Clock.real.now(io).toSeconds() - cfg.retentionSeconds()`. Sharing one handle between concurrent tasks is NOT safe: FULLMUTEX serializes single SQLite calls, but a transaction is connection state — a prune landing between another task's BEGIN and COMMIT joins that transaction. The contract therefore requires a dedicated connection; cross-connection isolation is SQLite's own (WAL + busy_timeout), and a pass losing a race sees Busy/Locked, logs at warn, and retries next interval. Tests (`:memory:` + ddl): a pass prunes exactly the old rows; the 7th pass vacuums (assert via `stats.vacuums` after 7 `runOnce` calls); a pass on an empty DB is a no-op that still counts; cutoff arithmetic honors `retention_days`. Acceptance: fmt/ast-check clean; tests standalone with `-lc -lsqlite3`. --- ## Session S8: integration — `src/storage/phase6_integration_test.zig` (needs all) Gated by `build_options.integration` exactly like `storage_integration_test.zig` (same Fixture idiom over `.zig-cache/tmp/`). Cases, named `"S8 case N: …"`: 1. logger end-to-end: real file `querylog.db` via `querylog_schema.open`, writer task under `io.concurrent`, 250 entries → shutdown → 250 rows, domains interned (fewer domain rows than query rows), fingerprint stamped. 2. flush on interval: one entry, no shutdown; poll until it lands; well under 10× the interval. 3. backpressure: queue capacity 8, enqueue 20 with the writer stalled (gated by a monitor stub at critical); `queries_dropped >= 12`, the survivors are the NEWEST entries, then un-gate and confirm the survivors landed. 4. privacy: hide both → every row's domain and client_ip read `hidden`. 5. retention: back-dated rows pruned, fresh rows kept, WAL truncated after checkpoint (assert `-wal` size 0 or absent). 6. disk thresholds end-to-end (§17 exit criterion): monitor with `min_free_mb` far above the real free space → `.critical`, `writesAllowed() == false`, logger gates (case-3 machinery), `batches_gated > 0`; then thresholds far below → `.ok` and flushing resumes. 7. log sink file mode: install to a fixture path with `max_size_mb` tiny (write the size gate via a test-only override — if `install` takes cfg, a 1 MB minimum makes this slow; add `pub fn installForTest(io, cfg, max_bytes_override)` if needed and mark it test-only), emit lines past the limit → `.1` exists, live file small, `max_files` honored, then `deinstall`. If overriding proves ugly, rotation is proven at the arithmetic level in S6 and this case shrinks to append+reopen round-trip — say which in the report. 8. cache with real packets: build a response via `ResponseBuilder`, `classify` + `put`, advance `now_s`, `get` → TTLs visibly decremented (re-parse), expiry at the boundary is a miss. 9. rate limiter + address integration: 1001 checks from one v6-mapped-v4 client inside one window → exactly one refused… (limit 1000); a second distinct client unaffected. Acceptance: `zig build test -Dintegration` exit 0 with all 9 passing; no case sleeps longer than 2 s of wall time; no network. --- ## Module Layout (new files) | File | Purpose | |---|---| | `src/cache/dns_cache.zig` | pure TTL cache: key, classify, CLOCK-bounded store | | `src/server/rate_limiter.zig` | pure fixed-window per-client limiter | | `src/storage/repositories/queries_repo.zig` | query_log/domains writes, prune, checkpoint, vacuum | | `src/storage/logger.zig` | Io.Queue async logger, privacy, backpressure, disk gate | | `src/platform/statfs.zig` | libc statvfs wrapper | | `src/storage/disk_monitor.zig` | free-space thresholds, gauges, 60 s sampler | | `src/platform/logging.zig` | std.log sink: level, rotation, upstream dedup | | `src/storage/retention.zig` | daily prune + checkpoint + weekly vacuum | | `src/storage/phase6_integration_test.zig` | S8 cases | ## File Ownership | Files | Owner | |---|---| | `src/cache/dns_cache.zig` | S1 | | `src/server/rate_limiter.zig` | S2 | | `src/storage/repositories/queries_repo.zig` | S3 (frozen after S3 verifies) | | `src/storage/logger.zig` | S4 | | `src/platform/statfs.zig`, `src/storage/disk_monitor.zig` | S5 (frozen after wave 1) | | `src/platform/logging.zig` | S6 | | `src/storage/retention.zig` | S7 | | `src/storage/phase6_integration_test.zig` | S8 | | `src/tests.zig`, `build.zig`, `src/main.zig` (std_options) | orchestrator | No session touches `src/dns/`, `src/upstream/`, `src/filter/`, `src/local/`, `src/server/handler.zig`, `src/server/udp_server.zig`, `src/server/tcp_server.zig`, `src/storage/db.zig`, any milestone 4–5 repository, `src/config/`, or `src/cli.zig`. A needed change there is reported, not made. ## Acceptance Criteria (Milestone 6 Complete) - [ ] `zig build test` exit 0 with all nine new files wired into `src/tests.zig`. - [ ] `zig build test -Dintegration` exit 0 including the 9 S8 cases and every prior milestone's. - [ ] `zig build cross` still produces two statically linked executables (statvfs links against musl). - [ ] `grep -rn "std.log.err" src/cache/ src/storage/logger.zig src/storage/disk_monitor.zig src/storage/retention.zig src/server/rate_limiter.zig src/platform/statfs.zig src/platform/logging.zig` → nothing. - [ ] `validate.zig` rejects `rate_window_seconds = 0` (rule exists at validate.zig:221; covered by milestone-4 tests). - [ ] `DnsCache.get`, `RateLimiter.check` have no allocator in reach — signatures prove the no-allocation hit path. - [ ] The disk-degradation integration case (S8 case 6) passes — PLAN §17's Phase 6 exit criterion. - [ ] `zig fmt --check` clean repo-wide; GPG-signed lowercase commit. ## Anti-Requirements - **No handler or server wiring.** `handler.zig` keeps its signature; nothing calls the cache, the limiter, or the logger from the query path. Phase 7. - **No client-row pruning and no auto-materialization.** Phase 7 (ruling 1). - **No web/API/SSE/metrics/health payloads.** Phase 8 reads the accessors this milestone exposes. - **No API token bucket.** §10's API limiter is Phase 8. - **No cache persistence, no cache serialization.** In-memory only (§8). - **No ECS parsing.** The cache key accepts ECS bytes the caller extracted; extracting them from OPT is the handler's job in Phase 7. - **No statement cache in db.zig.** `BatchWriter` owns its statements; `db.zig` stays as is. - **No new config keys and no schema change.** Every knob this milestone reads exists in `model.zig`; flush tuning is comptime (ruling 3). - **No log sink installation under the test runner.** `install` runs only from `main`. - **No epoll/timerfd/signalfd.** Loops sleep on `Clock.Duration`; Phase 7 owns lifecycle. ## As built (S1–S7 and orchestrator wiring) Deviations from the text above, recorded after the sessions verified. Where this section and the session text disagree, this section wins. **S1 dns_cache.** `Config` is an alias of `model.Cache`, not a duplicate struct. `classify` computes the positive min-TTL by iterating records (no mutation, no 64 KiB copy); NODATA (NOERROR, zero answers) takes the negative path with NXDOMAIN per RFC 2308 §2; a malformed SOA means "do not cache", not an error; `negative_ttl_max == 0` disables negatives only — positives still cache. A `put` on an existing key keeps the slot's CLOCK reference bit. A slot the CLOCK hand reclaims because it expired counts `expirations`, not `evictions`. `size == 0` turns caching off (`put` no-ops). A backwards clock ages by zero. `misses` covers lookup failure, expiry and a too-small `out`; a poisoned entry counts `invalid_hits` only. The slot's `negative` flag is stored but has no accessor yet — Phase 8 adds one if it wants the split. After review: the negative TTL is min(SOA record TTL, SOA MINIMUM) per RFC 2308 §5, and a top-bit-set TTL reads as zero per RFC 2181 §8 on every non-OPT record in any section (positive path) and both SOA fields (negative path) — a zero result means the response is not cached. When `put` stores a negative response it lowers the clone's first authority SOA record TTL to the entry's lifetime, so a hit at elapsed 0 serves an SOA TTL equal to the entry TTL and downstream caches do not hold the negative answer past it (RFC 2308 §3); the caller's buffer is not modified. On the positive path `put` caps every non-OPT record TTL in the clone at `max_ttl_seconds` for the same reason — the stored bytes never advertise a lifetime past the entry's ceiling (the cap is our ruling 5, not an RFC rule; unbound's cache-max-ttl serves the capped value the same way). A negative response's non-SOA records are deliberately not capped — the clamped SOA is the record that governs negative caching downstream. The poisoned-packet ageing test's fixture keeps its TTL under `max_ttl_seconds` on purpose, so `put` stores it unmodified and only the ageing inside `get` breaks it. **S2 rate_limiter.** Windows anchor at each client's first query (`.awake` origin is arbitrary). `sweep` drops entries older than two full windows (exactly two survives). `untracked` is a subset of `allowed`, so `allowed + refused` equals the number of `check` calls. `init` asserts `window_seconds != 0` (validate.zig:221 guards real configs). `limit == 0` refuses everything. Added: `trackedClients()` occupancy accessor. `sweep` collects stale keys into a buffer allocated at init (~68 KiB), not onto the stack. **S3 queries_repo.** Domain interning is INSERT OR IGNORE + SELECT (no `lastInsertRowid` — unset on ignore); a missing row right after the insert is `error.NotFound`. `select_domain` resets immediately after the id is read; `errdefer tx.rollback()` is declared before `errdefer resetAll()` so no cursor is open at ROLLBACK. `checkpointTruncate` treats a blocked checkpoint as success (SQLite reports it in a discarded row; retention retries next pass). **S4 logger.** `Entry.init(Entry.Fields)` builds an entry from borrowed slices, copying and truncating; `setDomain`/`setClientIp` are public for the privacy transform. **Spec correction:** the `fetchWithin` shape's `cancelDiscard` is wrong here — a `getOne` that loses the race has already removed an entry, and discarding its result loses the entry silently. `getWithin` drains the Select with a `while (race.cancel())` loop; an entry recovered on the cancellation path counts `queries_dropped`. A batch dropped by a `writeBatch` failure adds its length to `queries_dropped`. The interval deadline starts at the first entry of a batch. `gate_retry_s = 1` is public. Queue facts verified: `put` returns `Closed` even with space; `get(min=0)` on a closed non-empty queue returns elements and reports `Closed` only when empty — that is what makes the shutdown drain correct. A writer sitting in the disk gate at shutdown holds its batch until un-gated or canceled (documented on `shutdown`); `Logger` must not move after `init`. After review: `fill` reports its count through an out-parameter, so a canceled `fill` or `flush` counts every entry still held in the batch under `queries_dropped` (including the gate-blocked cancellation); a `BatchWriter.init` failure warns, sets `writer_failed: std.atomic.Value(bool)`, closes the queue and drains it with `Queue.getUncancelable` — uncancelable specifically so a cancellation racing the failure cannot leave buffered entries uncounted (no other consumer can reach them at that point); the writer never pretends to run. The zero-capacity early exit reads `queue.capacity()`. The prepare-failure path and the cancellation accounting are both covered in-file (schema-less `:memory:` DB; deterministic 2-entry batch canceled in the gate). **S5 statfs + disk_monitor.** The statvfs layout is pinned by `@offsetOf`/`@sizeOf` tests (112 bytes); musl verified from the zig-shipped headers, glibc from the system's. `init` requires `data_dir` opened with `.iterate = true` (documented). `log_dir_path` opens via `cwd()` per sample; null keeps `log_bytes` at 0. A failed statvfs leaves the state untouched; a failed size scan keeps that gauge's prior value while the state still publishes. A per-file `statFile` failure inside a size scan fails that whole scan — the gauge keeps its prior value and `sample_failures` increments once for the scan, not per file; `error.FileNotFound` is the one exception (a file deleted between `iterate` and `statFile` is normal on rotation or recreate — skipped, scan continues). `classify` checks min before warn, so warn-below-min misconfiguration reports the severer state. `run` samples first, then sleeps. Sizes come from `Dir.statFile`. **S6 logging sink.** **Spec corrections:** the `logFn` scope parameter is `comptime scope: @EnumLiteral()`, and `std.Thread.Mutex` does not exist in 0.16.0. The sink lock is `std.debug.lockStderr`/`unlockStderr` (the same lock `std.log.defaultLog` uses; documented recursive; needs no io). Added: `installForTest(io, cfg, max_bytes_override)`; public pure helpers `toStdLevel`, `enabled`, `isDedupScope`, `buildKey`, `DedupTable`, `rotatedName`. `max_files` counts the live file (generations `.1`…`.{max_files-1}`); `max_files < 2` truncates in place. `LogOutput.syslog` routes to stderr (journald captures it — the only supported deployment). File lines are ` (): \n`, per-line durable (fresh writer, seekTo, writeAll, flush); write failure closes the file and the next line reopens it; overlong paths fall back to stderr and count `sink_errors`; cancel protection wraps file work. After review: `` is escaped on BOTH output paths — `\\`, newline, carriage return and tab as two-character escapes, every other byte below 0x20 plus DEL as `\xNN` (a deliberate divergence from `std.log.defaultLog`, which writes raw; these lines are parsed by operators and journald); `max_escaped_message_bytes = max_message_bytes * 4` is public and the line buffer holds the worst case. A failed rotation step counts `sink_errors`, sends the line to stderr, and leaves the live file CLOSED until a later line completes the rotation — the oversized file is never reopened for append (`rotate_pending`); `rotations` counts completed rotations only. One exception: a single line longer than `max_bytes` is written to an empty file rather than rotated forever. `lines_written` counts only lines a sink accepted; each failed write attempt counts one `sink_errors` (file-fail + stderr-success = one of each). The scope name is truncated to `max_scope_name_bytes` (32) by `boundedScopeName`, applied once in `logFn` so the file and stderr paths render the same event identically; the public constant `max_line_header_bytes` (65) bounds the record header and the file line buffer is `max_escaped_message_bytes + max_line_header_bytes`. Record construction is `buildLine`, a pure function taking the unix seconds as a parameter — it returns an error instead of a partial record, and a failure counts one `sink_error` and sends the line to stderr, so a malformed record never reaches the file. Counting contract: every path where `prepareFileLocked` returns false has already counted exactly one `sink_error` at the failing step (the caller never counts again); `rotateLocked` counts its own failure. **S7 retention.** Prune, checkpoint and vacuum are independent within a pass — a failed prune does not skip the checkpoint (the WAL was filled by the logger, not this pass). `passes` increments at the top of `runOnce`; vacuum fires at `passes % 7 == 0`. The cutoff is strictly-less-than (a row exactly at the cutoff survives). `runOnce` opens no transaction; the contract on `run` requires a dedicated connection (concurrent sharing joins the other task's transaction — FULLMUTEX does not prevent that); Phase 7 opens it. Sequential use of one handle, as S8 case 5 does, is permitted. **Orchestrator wiring.** All eight new files (S8's included) imported individually by `src/tests.zig`. `src/main.zig` gained `pub const std_options: std.Options = .{ .logFn = logging.logFn };` — effective in the executable build, inert under the test runner (tests.zig is that root); `zig build` type-checks the sink through real call sites. The spec's claim that a `rate_window_seconds` rule was missing from validate.zig was wrong — the rule exists at validate.zig:221; nothing was added. **S8 integration.** Case 7 runs the full rotation via `installForTest(io, cfg, 256)` with `max_files = 3`, calling `logging.logFn` directly (the test runner owns `std_options`, so a `std.log` call would never reach the sink) and comparing counters as before/after deltas (`install`/`deinstall` do not reset stats); it also proves append-and-reopen byte-identity. Case 3 enqueues the full burst before the writer starts, making "12 dropped, newest 8 survive" exact, and starts the writer gated. Cases 3 and 6 un-gate before `shutdown` (the gate-hold documented in S4). Case 6 samples real `statvfs` against the fixture dir and flips state by swapping `monitor.cfg` between huge and zero thresholds. Case 5 asserts the WAL was non-empty before the checkpoint, so truncation is proven, not assumed. Shared-Db reads happen only while the writer sleeps in the gate or after its future is awaited. All cases use `testing.io`. Per-case wall time stays under 1.5 s; cases 3 and 6 pay `gate_retry_s`. With `integration = false`, all 9 skip.