1799 lines
95 KiB
Markdown
1799 lines
95 KiB
Markdown
# Milestone 5: Filtering + Local DNS
|
||
|
||
Goal (PLAN §16 Phase 5): the rule matcher (exact / parent-walk / wildcard); blocklist parsers and the
|
||
compiled file format; the fetcher and its scheduled update; the RCU snapshot swap; per-group
|
||
safe-search; blocked-response synthesis; local records; forward zones.
|
||
Exit: the §3.10 precedence table is validated by tests; local zone answers and conditional forwards
|
||
work.
|
||
|
||
Read first: `AGENTS.md` (values), `specs/research/zig-0.16-api-notes.md` (verified stdlib facts —
|
||
pre-0.16 knowledge is stale and MUST NOT be used), `specs/milestone-1.md`, `specs/milestone-2.md`,
|
||
`specs/milestone-3.md`, `specs/milestone-4.md` (module conventions, the "As built" notes, and the
|
||
binding logging policy restated below). The Zig source of truth is `/home/mokhtar/app/zig` at tag
|
||
`0.16.0`. PLAN §2.2, §3.8–§3.10, §5, §6.2, §6.4, §6.5, §7, §11.2, §12, §18 are the scope authority.
|
||
|
||
## What already exists (do not respecify, import it)
|
||
|
||
- `src/dns/*` — pure wire format. Used here: `name.Name`, `name.fromText`, `name.formatText`,
|
||
`name.eqlIgnoreCase`, `types.Type`, `types.Class`, `types.Rcode`, `types.max_name_len`,
|
||
`question.Question`, `header.Header`, `edns.OptRecord`, `packet.ResponseBuilder`
|
||
(`init`, `setRcode`, `addAnswer`, `addOptEcho`, `finish`). Do not write a second name parser and do
|
||
not add anything to `src/dns/`.
|
||
- `src/platform/address.zig` — `NetAddress.parse/format/key/eql`, `Prefix.parse/contains`,
|
||
`matchLongest`. Group assignment by client IP uses these; do not write a second IP parser.
|
||
- `src/config/model.zig` — `Config`, `Group`, `Rule`, `RuleKind`, `RuleAction`, `BlocklistSource`,
|
||
`GroupSource`, `LocalRecord`, `RecordType`, `ForwardZone`, `Client`, `ClientPrefix`, `Blocking`,
|
||
`BlockResponse`, `BlocklistUpdate`, `updateIntervalSeconds`.
|
||
- `src/config/validate.zig` — `parseResolver(text) ResolverError!Resolver`, `Resolver`,
|
||
`ResolverScheme`. Its doc comment already names this phase as the importer. `local/forward_zones.zig`
|
||
imports it; nobody writes a second resolver-URL parser.
|
||
- `src/storage/repositories/*` — `listGroups`/`freeGroups`, `listGroupSources`/`freeGroupSources`,
|
||
`listRules`/`freeRules`, `listBlocklistSources`/`freeBlocklistSources`, `listLocalRecords`,
|
||
`listForwardZones`, `listClients`, `listClientPrefixes`, and the `count*` functions. Every list
|
||
returns `std.ArrayList(model.X)` with heap-owned strings and a matching `freeX`. **The repositories
|
||
gain no new functions in this milestone except the two named in S8.1**, which are additive and
|
||
owned by S8.
|
||
- `src/storage/db.zig` — `Db`, `Stmt`, `Tx`, `Error`. `db.zig` takes no `std.Io` (milestone 4's
|
||
documented exception); everything else that touches the filesystem takes `io: std.Io`.
|
||
- `src/upstream/transport.zig` — `Client` (the `exchangeFn` vtable), `ExchangeError`, `PeerFault`,
|
||
`LocalResource`, `mapLocal`, `group`, `validateResponse`, `max_message_len`. The forward-zone
|
||
client implements this interface; it does not invent a second one.
|
||
- `build.zig` — `-Dintegration` (hermetic, loopback/tmpdir only, PR-blocking) and `-Dlive` (leaves
|
||
the machine, manual only) reach test files through `@import("build_options")`. A second test
|
||
artifact carries the fuzz targets with `dns` as a named module.
|
||
|
||
## Sessions
|
||
|
||
```
|
||
S1 (parsers + wildcard) S2 (domain_set + compiler) S3 (safesearch + response)
|
||
S4 (local records + forward_zones) S6 (fetcher) S7 (forward_client) [all parallel, no deps]
|
||
| |
|
||
+---------+----------+
|
||
v
|
||
S5 (rules + matcher) [needs S1, S2]
|
||
v
|
||
S8 (manager) [needs S2, S5, S6]
|
||
v
|
||
S9 (integration + fuzz) [needs everything]
|
||
```
|
||
|
||
Six sessions start together. Every later session is written against **this spec**, not against the
|
||
previous session's source. The orchestrator — not any session — wires `src/tests.zig` imports and
|
||
every `build.zig` change. A session that needs a build change reports the exact change in its
|
||
completion report; the one this milestone needs is stated in S9.3.
|
||
|
||
## Session verification protocol (read this before starting)
|
||
|
||
Unchanged from milestone 4, and it still binds:
|
||
|
||
- `zig test <file>` does not work here: every file imports across `src/` subdirectories or links
|
||
against `sqlite3`. Each session verifies its own work with `zig fmt --check <its files>` and
|
||
`zig ast-check <its files>`. `zig ast-check` reports only syntax and AST-level errors; it does
|
||
**not** type-check, so it cannot prove the code compiles.
|
||
- The orchestrator wires the files into `src/tests.zig` and runs `zig build test` (and
|
||
`zig build test -Dintegration` for S9). That run is the real gate.
|
||
- Every session states in its completion report that its tests have not been executed, and lists the
|
||
exact test names it wrote so the orchestrator can confirm they ran.
|
||
- No session edits `build.zig`, `build.zig.zon`, or `src/tests.zig`. No session edits a milestone
|
||
1–4 file; a needed change there is reported, not made.
|
||
|
||
## Design invariants (all sessions)
|
||
|
||
- **The pure core stays pure.** `filter/parsers.zig`, `parser_hosts.zig`, `parser_domains.zig`,
|
||
`parser_abp.zig`, `wildcard.zig`, `domain_set.zig`, `compiler.zig`, `rules.zig`, `matcher.zig`,
|
||
`safesearch.zig`, `response.zig`, `local/records.zig` and `local/forward_zones.zig` take **no
|
||
`std.Io` value**. They take bytes, structs, an allocator, and `*std.Io.Reader` / `*std.Io.Writer`
|
||
interfaces — which are not `Io` and carry no backend. Clocks, entropy, sockets and files arrive as
|
||
parameters or do not arrive at all.
|
||
- **`std.Io` lives in exactly three files**: `filter/fetcher.zig`, `filter/manager.zig`,
|
||
`local/forward_client.zig`. PLAN §5 marks `filter/` "pure" and then lists `fetcher.zig` and
|
||
`compiler.zig` inside it; the split above keeps that promise where it matters — the compiler is
|
||
pure over reader/writer interfaces, and only the fetcher and the manager touch the network and the
|
||
filesystem.
|
||
- **Logging policy (binding, repo-wide, milestone 4 "As built")**: a condition that is returned as a
|
||
typed error logs at `warn` at most. `err` is reserved for failures the code swallows. The zig test
|
||
runner fails any test that emits `err` logs. In this milestone the only legitimate `err` site is a
|
||
background refresh task that has nowhere to return its failure — and even that logs at `warn`,
|
||
because the failure is recorded in `SourceStatus` and surfaced (S8.5). **No `std.log.err` call is
|
||
written in this milestone.**
|
||
- **Every failure mode is counted, not dropped** (AGENTS.md). A skipped regex line, an unsupported
|
||
ABP modifier, an over-long line, a single-label name, an unparseable name, a failed download and a
|
||
checksum mismatch each increment a named counter that reaches `SourceStatus` and (for the three
|
||
columns that exist) `blocklist_sources`.
|
||
- **No regex, ever** (PLAN §2.2). Regex lines are recognized, counted, skipped. No engine, no
|
||
dependency, no "just a small subset".
|
||
- **Nothing in this milestone allocates on the query path.** `Snapshot.evaluate` performs no
|
||
allocation, opens no file, and takes no lock beyond the reader lock its caller already holds. It is
|
||
callable from a `std.Io` task with a stack buffer and nothing else.
|
||
- **Determinism**: a compiled list body is a pure function of the downloaded bytes and the format.
|
||
See the determinism contract in §Determinism below; S2 and S9 both test it.
|
||
- Unit tests live in-file. Tests that touch the filesystem, loopback sockets or real HTTP live in
|
||
`src/filter/filter_integration_test.zig`, guarded by
|
||
`if (!build_options.integration) return error.SkipZigTest;`. Tests that leave the machine are
|
||
guarded by `build_options.live`. Mirrors milestones 1, 3 and 4.
|
||
|
||
---
|
||
|
||
## Resolved PLAN ambiguities (read before writing code)
|
||
|
||
These were decided while writing this spec. They are the ruling for this milestone; the orchestrator
|
||
carries any of them that change PLAN text back into PLAN.
|
||
|
||
1. **Blocklist domain entries match exactly; only rules parent-walk.** PLAN §7.1 builds a candidate
|
||
chain and then consults rules (step 3), blocklist domains (step 4) and blocklist wildcards
|
||
(step 5). If blocklist domains were also matched against the whole chain, §3.10's separate
|
||
"blocklist wildcards" level would be unreachable and one `com`-shaped line in a bad list would
|
||
black-hole the internet. §3.9 ties parent-walk to the **rule** model ("parent-walk (implicit via
|
||
candidate chain)"). Ruling: exact and wildcard **rules** are evaluated against every candidate in
|
||
the chain; **blocklist `.list` entries** are matched against the query name only; **blocklist
|
||
`.wild` entries** are matched against every *proper parent* of the query name (that is what
|
||
`*.x.y` means). ABP `||x.y^` therefore emits both a `.list` entry `x.y` and a `.wild` entry `x.y`,
|
||
which together give the "domain and all subdomains" semantics the syntax promises.
|
||
2. **`*` matches one or more labels.** PLAN §3.9 gives both `*.doubleclick.net` and
|
||
`ads.*.example.com`. A one-label-only `*` would make the first pattern miss
|
||
`a.b.doubleclick.net`, which is not what an operator writing it means. One rule for both
|
||
positions: each `*` label matches one **or more** labels. Bounded by the 128-label ceiling a
|
||
255-byte name implies, so the backtracking match cannot blow up.
|
||
3. **`evaluate` takes no qtype.** PLAN §7.1 keys evaluation on `{domain, qtype, group_id}` but no
|
||
step in §7 or §3.10 reads the qtype. Ruling: the filtering decision is
|
||
`{domain, group}`; the qtype travels with the query for logging and for response synthesis, not
|
||
for matching. Adding an unused parameter would be generality nobody asked for (AGENTS.md).
|
||
4. **Group assignment by client IP is in scope; auto-materialization is not.** PLAN §7.2 belongs to
|
||
the filtering engine this phase builds, and the matching half is pure (`clients` exact match →
|
||
`client_prefixes` longest-prefix → `default`). The half that **writes** — inserting the unseen
|
||
client row, updating `last_seen`, retention of `hand_edited = 0` rows — needs a clock and a
|
||
database write on the query path and belongs with the handler pipeline (Phase 7) and retention
|
||
(Phase 6). `Snapshot.groupForClient` is built here; nothing inserts a row.
|
||
5. **A plain UDP/TCP resolver client is in scope.** PLAN Phase 5's exit criterion says conditional
|
||
forwards *work*, and §6.5 permits plain transports. No such client exists: `transport.Endpoint`
|
||
knows only `https://` and `tls://`, by design. Ruling: `src/local/forward_client.zig` implements
|
||
`transport.Client` over plain UDP with TCP fallback, configured from `validate.Resolver` rather
|
||
than from `transport.Endpoint`, so no milestone-3 file is edited. Wiring it into the query path is
|
||
Phase 7.
|
||
6. **Compressed transfer encoding is out of scope.** `std.http.Client` advertises gzip and deflate by
|
||
default (verified, Client.zig:831) and `Response.readerDecompressing` would be needed to read
|
||
them. The fetcher overrides `accept_encoding` to identity, exactly as `doh_client.zig` does. A
|
||
daily download of a few megabytes does not pay for a decompression buffer and a second failure
|
||
surface. A source URL that serves a `.gz` **body** (content, not encoding) fails the compile with
|
||
a high invalid-line count and a visible error, rather than silently compiling to nothing.
|
||
|
||
## Determinism
|
||
|
||
For a given input byte stream and format, `compiler.compile` must produce byte-identical `.list` and
|
||
`.wild` bodies on every run, on both target architectures. Concretely:
|
||
|
||
- Names are lowercased over ASCII only (`A`–`Z`); a byte ≥ 0x80 makes the candidate invalid.
|
||
- A single trailing dot is stripped; the name is stored without it.
|
||
- Entries are sorted ascending by `std.mem.order(u8, …)` over the full name and deduplicated.
|
||
- Every line ends with a single `\n`, including the last. An empty body is zero bytes.
|
||
- The body carries no header, no timestamp, no hostname, no counts. The manager writes the header
|
||
(S8.3); the checksum covers the two bodies only.
|
||
- Sorting is by byte value, not by locale, not by label. Two runs of the same input in different
|
||
input order (S9 shuffles a fixture) produce the same body.
|
||
|
||
The snapshot built from those files is deterministic in content but **not** in hash-table layout: the
|
||
open-addressing seed is per-snapshot randomness (S2.1). `contains` results do not depend on it, and
|
||
that is the property tests assert.
|
||
|
||
## Memory budget (PLAN §18: < 100 MB with ~1M blocked domains)
|
||
|
||
The compiled in-memory structure is an **exact-match** set — a flat arena of length-prefixed
|
||
lowercase names plus an open-addressed table of `u32` offsets into it. No Bloom filter, no
|
||
hash-only key set, no probabilistic structure: a false positive in a DNS sinkhole blocks a real
|
||
domain for a real household and is undebuggable from the outside. Every probe compares full bytes.
|
||
|
||
Sizing for one 1,000,000-entry source, mean name length 22 bytes:
|
||
|
||
| Part | Size |
|
||
|---|---|
|
||
| arena (`1 + len` bytes per name) | ≈ 23 MB |
|
||
| index: capacity `2^21` slots at load factor ≤ 0.75, 4 bytes per slot | 8 MiB |
|
||
| per-source struct overhead | < 1 KB |
|
||
| **total, one 1M-domain source** | **≈ 31 MB** |
|
||
|
||
A typical household set (three lists, ~150k entries each, heavy overlap) is under 15 MB. Sets are
|
||
**per source, shared between groups** — a group holds indices into `Snapshot.sources`, so assigning
|
||
one source to four groups costs four `u32`s, not four copies.
|
||
|
||
Refresh peak, which is what the budget has to survive:
|
||
|
||
1. Sources are compiled **one at a time**, each into its own arena, and that arena is freed before
|
||
the next source starts. Compile working set for a 1M-entry source: 23 MB of candidate bytes plus
|
||
4 MB of `u32` offsets ≈ 27 MB.
|
||
2. The new snapshot is built only after every compile has finished, so compile arenas and the new
|
||
snapshot never coexist.
|
||
3. Peak is therefore `max(old + compile, old + new)` ≈ 62 MB for a 1M-domain configuration, leaving
|
||
headroom under the 100 MB target for the servers, the pool and (Phase 6) the cache.
|
||
|
||
Hard caps, all of which produce a typed error and a counter rather than a slow death:
|
||
`compiler.max_domains = 2_000_000` per source, `fetcher.max_body_bytes = 64 * 1024 * 1024`,
|
||
`compiler.max_line_len = 4096`, `rules.max_wildcards_per_group = 4096`.
|
||
|
||
Lookup cost, against PLAN §18's "blocklist lookup p95 < 1 ms": a 5-label query walks at most 5
|
||
candidates; per candidate the matcher does 2 rule-set probes plus one probe per assigned source
|
||
(`.list` for the full name, `.wild` for parents). With 5 sources that is under 40 Wyhash + memcmp
|
||
probes on structures that fit in L2 for typical list sizes. No allocation, no lock beyond the
|
||
caller's reader lock.
|
||
|
||
---
|
||
|
||
## Verified 0.16.0 stdlib facts used by this milestone
|
||
|
||
Read from `/home/mokhtar/app/zig` at tag `0.16.0`. Anything not listed here or in
|
||
`specs/research/zig-0.16-api-notes.md` must be re-verified against the source before use.
|
||
|
||
### Locks — they are **not** in `std.Thread`
|
||
|
||
`lib/std/Thread.zig` declares no `Mutex` and no `RwLock` in 0.16.0. Both live under `std.Io`:
|
||
|
||
```zig
|
||
pub const RwLock = @import("Io/RwLock.zig"); // Io.zig:48
|
||
pub const init: RwLock = .{ ... }; // Io/RwLock.zig:15
|
||
pub fn lock(rl: *RwLock, io: Io) Io.Cancelable!void; // Io/RwLock.zig:51
|
||
pub fn lockUncancelable(rl: *RwLock, io: Io) void; // Io/RwLock.zig:42
|
||
pub fn unlock(rl: *RwLock, io: Io) void; // Io/RwLock.zig:70
|
||
pub fn lockShared(rl: *RwLock, io: Io) Io.Cancelable!void; // Io/RwLock.zig:115
|
||
pub fn lockSharedUncancelable(rl: *RwLock, io: Io) void; // Io/RwLock.zig:97
|
||
pub fn unlockShared(rl: *RwLock, io: Io) void; // Io/RwLock.zig:133
|
||
pub fn tryLockShared(rl: *RwLock, io: Io) bool; // Io/RwLock.zig:75
|
||
pub const Mutex = extern struct { ... }; // Io.zig:1587
|
||
```
|
||
|
||
### Timers
|
||
|
||
```zig
|
||
pub fn sleep(io: Io, duration: Duration, clock: Clock) Cancelable!void; // Io.zig:2397
|
||
pub fn sleep(duration: Clock.Duration, io: Io) Cancelable!void; // Io.zig:900
|
||
pub fn fromSeconds(x: i64) Duration; // Io.zig:988
|
||
pub fn fromMilliseconds(x: i64) Duration; // Io.zig:984
|
||
pub fn now(clock: Clock, io: Io) Io.Timestamp; // Io.zig:778
|
||
pub fn toSeconds(t: Timestamp) i64; // Io.zig:943
|
||
```
|
||
|
||
There is no timer wheel and no periodic-callback API. A scheduled refresh is a task that sleeps and
|
||
loops; cancellation arrives as `error.Canceled` from the sleep (milestone 3 convention).
|
||
|
||
### `std.Io.Reader` line iteration
|
||
|
||
```zig
|
||
pub fn takeDelimiter(r: *Reader, delimiter: u8) error{ ReadFailed, StreamTooLong }!?[]u8; // Reader.zig:895
|
||
pub fn discardDelimiterInclusive(r: *Reader, delimiter: u8) Error!usize; // Reader.zig:1037
|
||
pub fn discardRemaining(r: *Reader) ShortError!usize; // Reader.zig:270
|
||
```
|
||
|
||
`takeDelimiter` returns `null` at end of stream, excludes the delimiter, and treats end-of-stream as
|
||
a delimiter for a final unterminated line. `error.StreamTooLong` means the line did not fit the
|
||
reader's buffer and **leaves the stream unmodified** (doc comment, Reader.zig:885) — the compiler
|
||
must then `discardDelimiterInclusive('\n')` to step over that line and count it, or it will spin
|
||
forever on the same bytes. This is the single most likely infinite loop in this milestone.
|
||
|
||
### `std.http.Client` — GET download flow
|
||
|
||
```zig
|
||
pub fn request(client: *Client, method, uri, options: RequestOptions) RequestError!Request;
|
||
pub fn sendBodiless(r: *Request) Writer.Error!void; // Client.zig:912
|
||
pub fn receiveHead(r: *Request, redirect_buffer: []u8) ReceiveHeadError!Response; // Client.zig:1133
|
||
pub fn reader(response: *Response, transfer_buffer: []u8) *Reader; // Client.zig:736
|
||
pub fn readerDecompressing(...) *Reader; // Client.zig:752 (NOT used, ambiguity 6)
|
||
redirect_behavior: Request.RedirectBehavior = @enumFromInt(3) // Client.zig:1654
|
||
pub const default_accept_encoding = { gzip, deflate, identity } // Client.zig:831
|
||
```
|
||
|
||
`receiveHead` follows redirects itself and needs the `redirect_buffer` to outlive `Request.uri`;
|
||
RFC 9110 recommends ≥ 8000 bytes (doc comment, Client.zig:1128). An empty buffer is legal only when
|
||
`redirect_behavior = .not_allowed`. Blocklist URLs redirect in practice, so the fetcher allows
|
||
redirects and supplies a real buffer. `Response.reader` returns **compressed** bytes if a compressed
|
||
encoding was negotiated — which is why the fetcher pins `accept_encoding` to identity.
|
||
|
||
### Hashing and sorting
|
||
|
||
```zig
|
||
pub fn hash(seed: u64, input: []const u8) u64; // hash/wyhash.zig:178 → std.hash.Wyhash
|
||
pub const Sha256 = Sha2x32(iv256, 256); // crypto/sha2.zig:23
|
||
pub fn hash(b: []const u8, out: *[32]u8, options) void; // crypto/sha2.zig:108
|
||
pub fn bytesToHex(input: anytype, case: Case) [input.len * 2]u8; // fmt.zig:1156
|
||
pub fn sortUnstableContext(a: usize, b: usize, context: anytype) void; // mem.zig:648
|
||
```
|
||
|
||
### Filesystem (unchanged from milestone 4, reproduced for the two files that need it)
|
||
|
||
```zig
|
||
pub fn createFileAtomic(dir: Dir, io: Io, sub_path: []const u8, options: CreateFileAtomicOptions)
|
||
CreateFileAtomicError!File.Atomic; // Dir.zig:1924
|
||
pub fn replace(af: *Atomic, io: Io) ReplaceError!void; // File/Atomic.zig:77
|
||
pub fn deinit(af: *Atomic, io: Io) void; // File/Atomic.zig:23
|
||
pub fn readFileAlloc(dir, io, sub_path, gpa, limit) ReadFileAllocError![]u8; // Dir.zig:1326
|
||
pub fn createDirPathStatus(dir, io, sub_path, permissions) ...!CreatePathStatus; // Dir.zig:851
|
||
pub fn deleteFile(dir, io, sub_path) DeleteFileError!void; // Dir.zig:1004
|
||
pub fn writer(file: File, io: Io, buffer: []u8) Writer; // File.zig:600, BY VALUE
|
||
```
|
||
|
||
Plain `createDirPath` hardcodes `0o777` (Dir.zig:843) and must not be used; the blocklist directory
|
||
is created with `createDirPathStatus(io, path, .fromMode(0o700))`.
|
||
|
||
---
|
||
|
||
## Session S1: `src/filter/parsers.zig`, `parser_hosts.zig`, `parser_domains.zig`, `parser_abp.zig`, `wildcard.zig`
|
||
|
||
Pure, and **`std` is the only import in all five files**. No `@import("../dns/...")`, no config, no
|
||
allocator, no I/O. That constraint is not stylistic: `parsers.zig` is the root of the fuzz module
|
||
S9.3 adds, and a module root cannot import across its own directory boundary
|
||
(`error.ImportOutsideModulePath`, milestone 2). Domain **validity** is decided in S2's compiler
|
||
through `dns.name.fromText`; these files decide **format**.
|
||
|
||
### S1.1 `parsers.zig` — the shared vocabulary
|
||
|
||
```zig
|
||
pub const hosts = @import("parser_hosts.zig");
|
||
pub const domains = @import("parser_domains.zig");
|
||
pub const abp = @import("parser_abp.zig");
|
||
pub const wildcard = @import("wildcard.zig");
|
||
|
||
pub const Format = enum { hosts, domains, abp };
|
||
|
||
pub const Kind = enum {
|
||
/// Nothing on the line, or only a comment.
|
||
ignore,
|
||
/// `text` holds one or more whitespace-separated candidate names.
|
||
domain,
|
||
/// `text` holds one candidate suffix; every proper subdomain of it matches.
|
||
wildcard,
|
||
/// A regex rule. Counted, skipped, never compiled (PLAN §2.2).
|
||
regex,
|
||
/// Syntactically a rule of this format, but one nxdns cannot honour:
|
||
/// an ABP modifier list, an exception rule, element hiding, a scheme anchor.
|
||
unsupported,
|
||
};
|
||
|
||
pub const Line = struct {
|
||
kind: Kind,
|
||
/// Borrowed from the caller's line. Not lowercased, not validated.
|
||
text: []const u8 = "",
|
||
/// `.wildcard` only. ABP `||x^` covers `x` itself as well as its subdomains,
|
||
/// so the compiler emits an additional `.list` entry when this is set.
|
||
covers_apex: bool = false,
|
||
};
|
||
|
||
/// Dispatches to the format's parser. The line must not contain '\n' or '\r';
|
||
/// the caller strips them.
|
||
pub fn parseLine(format: Format, line: []const u8) Line;
|
||
|
||
/// Picks a format from the first `sample_lines` lines that are not blank and
|
||
/// not comments: an ABP marker (`||`, `@@`, `##`, `$`) anywhere wins `.abp`;
|
||
/// otherwise a majority of lines whose first field parses as an IP literal
|
||
/// wins `.hosts`; otherwise `.domains`.
|
||
pub const sample_lines = 64;
|
||
pub fn detectFormat(sample: []const u8) Format;
|
||
```
|
||
|
||
`detectFormat` uses its own minimal IPv4/IPv6-literal recognizer over the first field (digits, dots,
|
||
colons and hex only) rather than `NetAddress.parse` — importing `platform/` would break the module
|
||
constraint above, and a sniffing heuristic does not need a correct parser.
|
||
|
||
### S1.2 `parser_hosts.zig`
|
||
|
||
```zig
|
||
pub fn parseLine(line: []const u8) parsers.Line;
|
||
```
|
||
|
||
1. Truncate at the first `#`. Trim ASCII whitespace. Empty → `.ignore`.
|
||
2. A line starting with `/` → `.regex` (a regex smuggled into a hosts list).
|
||
3. Split on ASCII whitespace. If the first field looks like an IP literal **and** at least one field
|
||
follows, the remaining fields are the candidate names → `.domain` with `text` spanning them.
|
||
4. If the first field does not look like an IP literal, the whole line is the candidate name(s) →
|
||
`.domain`. Real "hosts" lists are frequently bare domain lists with a hosts extension.
|
||
5. A leading `*.` on any candidate is not stripped here; `parsers.parseLine` is per line, and the
|
||
compiler splits `text` and re-classifies each name (S2.3). This is the one place a `.domain` line
|
||
can yield wildcard entries, and it is spelled out in S2.3 rather than duplicated here.
|
||
|
||
The sink address is not checked against a list of "blocking" IPs: a hosts-format blocklist that maps
|
||
to `127.0.0.1`, `0.0.0.0` or `::` is the same instruction to nxdns, and a source that maps to a real
|
||
address is still a source of names the operator asked to block.
|
||
|
||
### S1.3 `parser_domains.zig`
|
||
|
||
```zig
|
||
pub fn parseLine(line: []const u8) parsers.Line;
|
||
```
|
||
|
||
One candidate per line. `#` and `!` start a comment (both appear in the wild); truncate and trim.
|
||
A leading `/` → `.regex`. Everything else → `.domain` with the trimmed text. A line containing
|
||
whitespace after trimming → `.unsupported` (a domains-format list with two fields is a
|
||
mis-detected hosts file, and guessing is worse than counting).
|
||
|
||
### S1.4 `parser_abp.zig`
|
||
|
||
```zig
|
||
pub fn parseLine(line: []const u8) parsers.Line;
|
||
```
|
||
|
||
| Input | Result |
|
||
|---|---|
|
||
| `! comment`, `[Adblock Plus 2.0]`, empty | `.ignore` |
|
||
| `\|\|example.com^` | `.wildcard`, `text = "example.com"`, `covers_apex = true` |
|
||
| `\|\|example.com^` with a trailing `$…` modifier list | `.unsupported` |
|
||
| `@@…` (exception) | `.unsupported` |
|
||
| `##…`, `#@#…`, `#?#…` (element hiding) | `.unsupported` |
|
||
| `\|http://…`, `\|https://…` (scheme anchor) | `.unsupported` |
|
||
| `/regex/` | `.regex` |
|
||
| `example.com` (bare) | `.domain` |
|
||
| anything containing `*`, `^` or `\|` outside the forms above | `.unsupported` |
|
||
|
||
`^` is a separator token in ABP and only the trailing `^` (or a trailing `^` before a modifier) is
|
||
meaningful for a domain rule; anything else carrying `^` is `.unsupported`. Exception rules are
|
||
`.unsupported` rather than an allow entry: PLAN's allow surface is the `rules` table, and a list
|
||
that could quietly allow a domain across every group is a policy hole the operator did not open.
|
||
|
||
### S1.5 `wildcard.zig`
|
||
|
||
```zig
|
||
pub const max_labels = 128;
|
||
|
||
pub const PatternError = error{
|
||
/// No label is exactly "*".
|
||
NoWildcard,
|
||
/// A label contains '*' but is not exactly "*". Partial-label globbing
|
||
/// (`ad*.example.com`) is out of scope: it is regex by another name, and
|
||
/// PLAN §3.9 defines the wildcard as a label pattern.
|
||
PartialWildcardLabel,
|
||
EmptyLabel,
|
||
LabelTooLong,
|
||
PatternTooLong,
|
||
TooManyLabels,
|
||
};
|
||
|
||
/// Syntax only. A valid pattern has at least one label that is exactly "*",
|
||
/// every other label is 1–63 bytes with no '*' inside it, and the whole
|
||
/// pattern is at most 253 bytes over at most `max_labels` labels.
|
||
pub fn validate(pattern: []const u8) PatternError!void;
|
||
|
||
/// `domain` is already normalized: lowercase, no trailing dot. `pattern` is
|
||
/// lowercase. Each "*" label matches ONE OR MORE labels (resolved ambiguity 2).
|
||
/// Allocation-free; the backtracking is bounded by `max_labels` on both sides.
|
||
pub fn matches(pattern: []const u8, domain: []const u8) bool;
|
||
```
|
||
|
||
Implementation shape: split both sides into label slices in two fixed `[max_labels][]const u8`
|
||
stack arrays, then run the classic two-pointer glob match with a single backtrack point (a `*`
|
||
consumes one label and may extend). No recursion, no allocation. A pattern with more than
|
||
`max_labels` labels cannot occur because `validate` rejects it and the matcher only holds validated
|
||
patterns; `matches` asserts the bound.
|
||
|
||
### S1.6 Tests (in-file)
|
||
|
||
`parsers.zig`:
|
||
- `detectFormat` on a hosts fixture, a domains fixture, an ABP fixture, and a file whose first 64
|
||
lines are all comments (→ `.domains`).
|
||
- `parseLine` dispatches to the format's parser (one case each).
|
||
|
||
`parser_hosts.zig`:
|
||
- `0.0.0.0 ads.example.com` → `.domain`, `"ads.example.com"`.
|
||
- `127.0.0.1 a.example.com b.example.com` → `.domain` whose `text` spans both names.
|
||
- `::1 ip6-localhost ip6-loopback` → `.domain` with both names (the compiler drops single-label
|
||
names, S2.3 — this test asserts the parser does not silently swallow them).
|
||
- `0.0.0.0 ads.example.com # tracker` → the comment is gone.
|
||
- `# whole line`, ``, ` ` → `.ignore`.
|
||
- `/ads\d+/` → `.regex`.
|
||
- `example.com` (no IP) → `.domain`.
|
||
|
||
`parser_domains.zig`: bare name; `! comment`; `# comment`; inline `example.com # x`; `/re/` →
|
||
`.regex`; `0.0.0.0 example.com` → `.unsupported`.
|
||
|
||
`parser_abp.zig`: one named test per row of the S1.4 table.
|
||
|
||
`wildcard.zig`:
|
||
- `validate` accepts `*.doubleclick.net` and `ads.*.example.com`; one named test per `PatternError`
|
||
member: `example.com` → `NoWildcard`, `a*b.com` → `PartialWildcardLabel`, `a..b` → `EmptyLabel`, a
|
||
64-byte label → `LabelTooLong`, a 300-byte pattern → `PatternTooLong`, a 200-label pattern →
|
||
`TooManyLabels`.
|
||
- `matches("*.doubleclick.net", "a.doubleclick.net")` true;
|
||
`("*.doubleclick.net", "a.b.doubleclick.net")` true (ambiguity 2);
|
||
`("*.doubleclick.net", "doubleclick.net")` false;
|
||
`("ads.*.example.com", "ads.eu.example.com")` true;
|
||
`("ads.*.example.com", "ads.eu.west.example.com")` true;
|
||
`("ads.*.example.com", "ads.example.com")` false;
|
||
`("*.example.com", "example.com.evil.net")` false.
|
||
- A pathological pattern `*.*.*.*.*.*.*.*.example.com` against a 100-label domain terminates (assert
|
||
it returns, which is the runnable form of "the backtracking is bounded").
|
||
|
||
### S1.7 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on all five files.
|
||
- [ ] `grep -n "@import" src/filter/parsers.zig src/filter/parser_*.zig src/filter/wildcard.zig`
|
||
shows only `std` and sibling files in `src/filter/`.
|
||
- [ ] `PatternError` includes `PartialWildcardLabel` and every member is produced by a named test.
|
||
- [ ] No allocator, no `std.Io` value, no clock appears in any of the five files.
|
||
- [ ] Every bullet in S1.6 exists as a named test.
|
||
|
||
---
|
||
|
||
## Session S2: `src/filter/domain_set.zig`, `src/filter/compiler.zig`
|
||
|
||
Pure: an allocator and reader/writer interfaces, no `Io`. `compiler.zig` may import
|
||
`../dns/name.zig` (it decides domain validity); `domain_set.zig` imports only `std`.
|
||
|
||
### S2.1 `domain_set.zig`
|
||
|
||
```zig
|
||
pub const DomainSet = struct {
|
||
/// Length-prefixed lowercase names, back to back: [len: u8][bytes]…
|
||
arena: []const u8,
|
||
/// Open-addressed table of offsets into `arena`; `empty` marks a hole.
|
||
/// Length is always a power of two.
|
||
index: []const u32,
|
||
count: u32,
|
||
seed: u64,
|
||
|
||
pub const empty_slot: u32 = std.math.maxInt(u32);
|
||
pub const max_count: u32 = 4_000_000;
|
||
pub const max_arena_bytes: usize = 1 << 31;
|
||
|
||
pub const Error = error{ OutOfMemory, TooManyDomains, SetTooLarge, NotSorted, NotLowercase };
|
||
|
||
/// Builds from a compiled body: LF-separated, lowercase, sorted ascending,
|
||
/// deduplicated, every line 1–255 bytes. The sortedness is VERIFIED, not
|
||
/// assumed — a hand-edited or truncated file must fail loudly at load
|
||
/// (`error.NotSorted`) rather than produce a set that silently misses
|
||
/// entries. Verification is one comparison per line and costs nothing.
|
||
///
|
||
/// `seed` randomizes the hash. Query names are attacker-supplied, so a
|
||
/// fixed seed would make probe-chain flooding computable offline; the seed
|
||
/// arrives as a parameter so this file stays free of entropy sources.
|
||
pub fn build(gpa: std.mem.Allocator, body: []const u8, seed: u64) Error!DomainSet;
|
||
|
||
/// An empty set that owns nothing. `contains` on it is always false.
|
||
pub const empty: DomainSet;
|
||
|
||
pub fn deinit(self: *DomainSet, gpa: std.mem.Allocator) void;
|
||
|
||
/// `domain` must be normalized (lowercase, no trailing dot). Allocation-free.
|
||
pub fn contains(self: *const DomainSet, domain: []const u8) bool;
|
||
|
||
/// Bytes held, for the memory report in `Snapshot.memoryBytes`.
|
||
pub fn memoryBytes(self: *const DomainSet) usize;
|
||
};
|
||
```
|
||
|
||
Rules:
|
||
|
||
- Capacity is the smallest power of two ≥ `count * 4 / 3`, minimum 16. Linear probing, insertion in
|
||
file order, no tombstones (the set is immutable after `build`).
|
||
- Hash is `std.hash.Wyhash.hash(self.seed, domain)`; the slot is `hash & (index.len - 1)`. A probe
|
||
compares the full stored bytes with `std.mem.eql` before reporting a hit. **There is no path in
|
||
which a hash collision produces a match.**
|
||
- `build` makes exactly two allocations (arena, index) and frees both on any error.
|
||
- A body of length 0 yields `empty`.
|
||
- `count > max_count` → `error.TooManyDomains`; body longer than `max_arena_bytes` →
|
||
`error.SetTooLarge` (the index stores `u32` offsets and must not silently truncate).
|
||
|
||
Tests (in-file):
|
||
- Build from a small sorted body; `contains` true for every member, false for four non-members
|
||
including a prefix, a suffix, an uppercase spelling and the empty string.
|
||
- A body that is not sorted → `error.NotSorted`; a body with an uppercase byte → `error.NotLowercase`;
|
||
a duplicate line → `error.NotSorted` (equal is not ascending, and a duplicate means the compiler
|
||
broke its contract).
|
||
- Two sets built from the same body with different seeds answer `contains` identically over a
|
||
50-name probe list — the property that matters, stated as a test.
|
||
- 10,000 generated names round-trip; `memoryBytes` is within 2× of the naive `body.len` bound.
|
||
- `build` under `std.testing.checkAllAllocationFailures` leaks nothing.
|
||
- `empty.contains("x") == false` and `deinit` on `empty` is a no-op.
|
||
|
||
### S2.2 `compiler.zig` — signatures
|
||
|
||
```zig
|
||
pub const max_domains: u32 = 2_000_000;
|
||
pub const max_line_len: usize = 4096;
|
||
|
||
pub const Counts = struct {
|
||
domains: u32 = 0,
|
||
wildcards: u32 = 0,
|
||
skipped_regex: u32 = 0,
|
||
skipped_unsupported: u32 = 0,
|
||
/// Not a valid domain name (`dns.name.fromText` rejected it, a non-ASCII
|
||
/// byte, or fewer than two labels).
|
||
invalid: u32 = 0,
|
||
/// Lines longer than `max_line_len`, skipped whole.
|
||
long_lines: u32 = 0,
|
||
/// Duplicates removed by the sort/unique pass.
|
||
duplicates: u32 = 0,
|
||
};
|
||
|
||
pub const Result = struct {
|
||
counts: Counts,
|
||
/// Lowercase hex sha256 over the `.list` body followed by the `.wild` body.
|
||
checksum: [64]u8,
|
||
};
|
||
|
||
pub const Error = error{ OutOfMemory, TooManyDomains, ReadFailed, WriteFailed };
|
||
|
||
/// Reads `r` to end of stream, writes the two compiled bodies. Nothing else is
|
||
/// written — headers belong to the caller (S8.3), so this function is a pure
|
||
/// function of (bytes, format) and is tested by comparing two runs.
|
||
pub fn compile(
|
||
gpa: std.mem.Allocator,
|
||
r: *std.Io.Reader,
|
||
format: parsers.Format,
|
||
list_w: *std.Io.Writer,
|
||
wild_w: *std.Io.Writer,
|
||
) Error!Result;
|
||
```
|
||
|
||
### S2.3 `compile` algorithm
|
||
|
||
1. Two `std.ArrayList(u8)` arenas (`list_bytes`, `wild_bytes`) and two `std.ArrayList(u32)` offset
|
||
lists, all from `gpa`. All four are freed before returning, on every path.
|
||
2. Loop with `r.takeDelimiter('\n')`:
|
||
- `null` → end.
|
||
- `error.StreamTooLong` → `counts.long_lines += 1`, then
|
||
`r.discardDelimiterInclusive('\n')` (tolerating `error.EndOfStream` as end of input) and
|
||
continue. **Not doing this discard is an infinite loop** — `takeDelimiter` leaves the stream
|
||
unmodified on `StreamTooLong` (verified, Reader.zig:885).
|
||
- Strip a trailing `\r` (CRLF files are common).
|
||
3. `parsers.parseLine(format, line)`; `.ignore` continues, `.regex` and `.unsupported` increment
|
||
their counters and continue.
|
||
4. For `.domain`, split `line.text` on ASCII whitespace and process each field; for `.wildcard`,
|
||
process `line.text` as one field with `covers_apex` remembered.
|
||
5. Per candidate field, in order:
|
||
a. A leading `*.` makes it a wildcard candidate over the remainder; a `*` anywhere else makes it
|
||
`invalid` (blocklist entries are suffixes, not patterns — patterns belong to the `rules`
|
||
table).
|
||
b. Strip one trailing `.`.
|
||
c. Reject any byte ≥ 0x80 or any ASCII control byte → `invalid`.
|
||
d. Lowercase ASCII into a `[types.max_name_len]u8` stack buffer.
|
||
e. `dns.name.fromText` must accept it → else `invalid`.
|
||
f. Fewer than two labels → `invalid`. This is what keeps `localhost`, `local`, `broadcasthost`
|
||
and the `ip6-*` names that every hosts list carries from black-holing the loopback names of
|
||
every client on the LAN. It is the single highest-consequence rule in this file.
|
||
g. Append to `list_bytes` (domain) or `wild_bytes` (wildcard), recording the offset. A
|
||
`.wildcard` with `covers_apex` appends to **both**.
|
||
h. Either list exceeding `max_domains` entries → `error.TooManyDomains`.
|
||
6. Sort each offset list with `std.mem.sortUnstableContext` comparing the stored names bytewise, then
|
||
write unique entries to the matching writer, each followed by `\n`, counting duplicates.
|
||
7. Hash both bodies as they are written (`std.crypto.hash.sha2.Sha256` streaming: `update` per
|
||
emitted line, `final`, `std.fmt.bytesToHex(digest, .lower)`), `.list` first then `.wild`.
|
||
8. Return `Result`. `counts.domains` and `counts.wildcards` are the **written, deduplicated** counts —
|
||
they are what `blocklist_sources.domain_count` and `wildcard_count` store, and what the UI shows.
|
||
|
||
### S2.4 Tests (in-file, `std.Io.Reader.fixed` / `std.Io.Writer.Allocating`)
|
||
|
||
- Hosts fixture with sink IPs, comments, a duplicate, `localhost`, an over-long line and a regex line:
|
||
every counter has the expected value and the body is exactly the expected sorted text.
|
||
- Domains fixture and ABP fixture likewise, the ABP one asserting `||x.com^` produces `x.com` in
|
||
**both** bodies.
|
||
- **Determinism**: compile the same fixture twice → identical bodies and identical checksums.
|
||
- **Order independence**: compile a shuffled permutation of the same fixture → identical bodies.
|
||
- Uppercase input compiles to lowercase output; a trailing-dot input compiles without the dot.
|
||
- A candidate with a non-ASCII byte, a single-label candidate, and a `a*b.com` candidate each land in
|
||
`counts.invalid` and appear in neither body.
|
||
- An input of 3,000 lines with a 5,000-byte line in the middle: `long_lines == 1`, the surrounding
|
||
lines are all present, and the call **terminates** (the `StreamTooLong` regression test).
|
||
- An empty input produces two empty bodies and the sha256 of the empty string.
|
||
- `compile` under `checkAllAllocationFailures` leaks nothing.
|
||
|
||
### S2.5 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on both files.
|
||
- [ ] Neither file takes a `std.Io` value, opens a file, or reads a clock.
|
||
- [ ] `domain_set.zig` imports only `std`.
|
||
- [ ] `contains` compares full bytes on every probe; no code path returns a hit from a hash
|
||
comparison alone.
|
||
- [ ] The `StreamTooLong` discard is implemented and covered by the 3,000-line test.
|
||
- [ ] The two-label minimum is implemented and covered by a `localhost` test.
|
||
- [ ] Determinism and order-independence tests pass.
|
||
|
||
---
|
||
|
||
## Session S3: `src/filter/safesearch.zig`, `src/filter/response.zig`
|
||
|
||
Pure. `safesearch.zig` imports `std` and `../dns/name.zig`; `response.zig` imports `std`, `../dns/*`
|
||
and `../config/model.zig`.
|
||
|
||
### S3.1 `safesearch.zig`
|
||
|
||
```zig
|
||
pub const Entry = struct { domain: []const u8, target: []const u8 };
|
||
|
||
/// PLAN §7.4. Sorted by `domain`, so the table is searchable and diffable, and
|
||
/// asserted sorted by a comptime block. Every target is a name the operator's
|
||
/// upstream can resolve; nxdns never hardcodes an address.
|
||
pub const table = [_]Entry{
|
||
.{ .domain = "bing.com", .target = "strict.bing.com" },
|
||
.{ .domain = "duckduckgo.com", .target = "safe.duckduckgo.com" },
|
||
.{ .domain = "google.com", .target = "forcesafesearch.google.com" },
|
||
.{ .domain = "pixabay.com", .target = "safesearch.pixabay.com" },
|
||
.{ .domain = "www.bing.com", .target = "strict.bing.com" },
|
||
.{ .domain = "www.duckduckgo.com", .target = "safe.duckduckgo.com" },
|
||
.{ .domain = "www.google.com", .target = "forcesafesearch.google.com" },
|
||
.{ .domain = "www.youtube.com", .target = "restrictmoderate.youtube.com" },
|
||
.{ .domain = "youtube.com", .target = "restrictmoderate.youtube.com" },
|
||
.{ .domain = "m.youtube.com", .target = "restrictmoderate.youtube.com" },
|
||
.{ .domain = "youtubei.googleapis.com", .target = "restrictmoderate.youtube.com" },
|
||
.{ .domain = "youtube.googleapis.com", .target = "restrictmoderate.youtube.com" },
|
||
.{ .domain = "www.youtube-nocookie.com", .target = "restrictmoderate.youtube.com" },
|
||
};
|
||
|
||
/// Exact match on a normalized name (lowercase, no trailing dot). Google's
|
||
/// country domains (`google.de`, …) are NOT enumerated: the list is unbounded,
|
||
/// it goes stale, and `forcesafesearch.google.com` is the documented target for
|
||
/// every one of them — an operator who needs a country domain adds a rule.
|
||
/// State that reasoning in the file.
|
||
pub fn lookup(domain: []const u8) ?[]const u8;
|
||
|
||
/// The rewritten question name for a matched query, as a `dns.name.Name`.
|
||
/// Applying it — sending the rewritten question upstream and prefixing the
|
||
/// answer with a CNAME from the original name to `target` — is the handler's
|
||
/// job (Phase 7). Nothing here builds a response.
|
||
pub fn rewrite(domain: []const u8) ?name.Name;
|
||
```
|
||
|
||
The table is asserted sorted and duplicate-free in a `comptime` block, and `lookup` is a binary
|
||
search. A test walks every entry and asserts `dns.name.fromText` accepts both sides.
|
||
|
||
### S3.2 `response.zig` — blocked-response synthesis (PLAN §6.2)
|
||
|
||
```zig
|
||
pub const Options = struct {
|
||
mode: model.BlockResponse, // .zero | .nxdomain
|
||
ttl: u32, // blocking.ttl
|
||
};
|
||
|
||
pub const Error = packet.ResponseBuilder.Error;
|
||
|
||
/// Writes a blocked reply for `q` into `buf` and returns a prefix of it.
|
||
///
|
||
/// `.zero`: A → 0.0.0.0, AAAA → ::, every other qtype → NOERROR with no answer
|
||
/// (NODATA). Synthesizing an address for a qtype that does not carry one is
|
||
/// not possible, and NXDOMAIN for, say, an MX query would tell the client the
|
||
/// name does not exist while an A query says it does.
|
||
/// `.nxdomain`: RCODE = NXDOMAIN, no answer, for every qtype.
|
||
///
|
||
/// No SOA is placed in the authority section: nxdns is not authoritative for
|
||
/// the name and a synthesized SOA would give resolvers a negative-caching TTL
|
||
/// nxdns cannot honour. Document that in the file.
|
||
///
|
||
/// `request_opt` echoes EDNS exactly as `handler.zig` does: when the query
|
||
/// carried an OPT record, the reply carries one with the same payload size and
|
||
/// the DO bit passed through.
|
||
pub fn writeBlocked(
|
||
buf: []u8,
|
||
request: header.Header,
|
||
q: question.Question,
|
||
request_opt: ?edns.OptRecord,
|
||
do_bit: bool,
|
||
options: Options,
|
||
) Error![]u8;
|
||
```
|
||
|
||
Only the question's class `IN` is answered with addresses; a non-`IN` class takes the NODATA path.
|
||
|
||
Tests (in-file): for each of the two modes × {A, AAAA, MX, HTTPS} × {with OPT, without OPT}, parse
|
||
the produced message back with `packet.parse` and assert the rcode, ancount, arcount, the echoed
|
||
question, the rdata bytes and the TTL. One test asserts a `buf` too small returns
|
||
`error.WriteFailed` rather than truncating.
|
||
|
||
### S3.3 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on both files.
|
||
- [ ] The safe-search table is comptime-asserted sorted and duplicate-free.
|
||
- [ ] Neither file takes a `std.Io` value or an allocator.
|
||
- [ ] Every blocked-response case in S3.2 is a named test that re-parses the output.
|
||
|
||
---
|
||
|
||
## Session S4: `src/local/records.zig`, `src/local/forward_zones.zig`
|
||
|
||
Pure. Both build immutable lookup structures from the model types the repositories already return.
|
||
|
||
### S4.1 `records.zig`
|
||
|
||
```zig
|
||
pub const Value = union(enum) { a: [4]u8, aaaa: [16]u8, cname: name.Name };
|
||
|
||
pub const Record = struct {
|
||
/// Normalized owner name: lowercase, no trailing dot.
|
||
owner: []const u8,
|
||
value: Value,
|
||
ttl: u32,
|
||
};
|
||
|
||
pub const Error = error{ OutOfMemory, BadRecordValue, BadRecordName, TooManyRecords };
|
||
pub const max_records: usize = 10_000;
|
||
|
||
pub const Records = struct {
|
||
/// Arena-owned, sorted by (owner, rtype) so lookup is a binary search and
|
||
/// the answer order for one name is stable across restarts.
|
||
items: []const Record,
|
||
|
||
pub const empty: Records;
|
||
|
||
/// `gpa` owns the result; `deinit` frees it. Values are parsed here, once:
|
||
/// an `a` value through `NetAddress.parse` (must be `.ip4`), `aaaa` (`.ip6`),
|
||
/// a `cname` through `dns.name.fromText`. A bad value is an error, not a
|
||
/// skipped row — `validate.zig` already rejects these, so reaching one here
|
||
/// means the database was edited behind nxdns's back and silence would make
|
||
/// a record vanish with no signal.
|
||
pub fn build(gpa: std.mem.Allocator, rows: []const model.LocalRecord) Error!Records;
|
||
pub fn deinit(self: *Records, gpa: std.mem.Allocator) void;
|
||
|
||
/// All records for `domain` whose type matches `qtype`, plus any CNAME
|
||
/// (RFC 1034 §3.6.2: a CNAME answers every qtype). Empty slice = no local
|
||
/// record. Allocation-free.
|
||
pub fn lookup(self: *const Records, domain: []const u8, qtype: types.Type) []const Record;
|
||
|
||
/// True when the name has any local record of any type. The handler needs
|
||
/// this to answer NODATA instead of forwarding a name nxdns owns.
|
||
pub fn hasName(self: *const Records, domain: []const u8) bool;
|
||
};
|
||
|
||
/// Writes `records` as answers into a builder the caller has already
|
||
/// initialized with the request header and question. Mechanism only.
|
||
pub fn writeAnswers(
|
||
b: *packet.ResponseBuilder,
|
||
owner: name.Name,
|
||
records: []const Record,
|
||
) packet.ResponseBuilder.Error!void;
|
||
```
|
||
|
||
Local records are group-independent (PLAN §6.4) and are matched **before** filtering. Resolving a
|
||
CNAME target through the pipeline is Phase 7; `writeAnswers` emits the CNAME record and stops.
|
||
|
||
### S4.2 `forward_zones.zig`
|
||
|
||
```zig
|
||
pub const Zone = struct {
|
||
/// Normalized: lowercase, no trailing dot.
|
||
zone: []const u8,
|
||
resolver: validate.Resolver,
|
||
};
|
||
|
||
pub const Error = error{ OutOfMemory, BadZone, BadResolver, TooManyZones };
|
||
pub const max_zones: usize = 1_000;
|
||
|
||
pub const Zones = struct {
|
||
/// Arena-owned, sorted by descending label count then by name, so the first
|
||
/// match found by a forward scan is the longest one.
|
||
items: []const Zone,
|
||
|
||
pub const empty: Zones;
|
||
|
||
pub fn build(gpa: std.mem.Allocator, rows: []const model.ForwardZone) Error!Zones;
|
||
pub fn deinit(self: *Zones, gpa: std.mem.Allocator) void;
|
||
|
||
/// Longest-suffix match on LABEL boundaries: `lan.home` matches `nas.lan.home`
|
||
/// and `lan.home`, and does NOT match `notlan.home`. `10.in-addr.arpa`
|
||
/// matches every reverse name under it. Allocation-free.
|
||
pub fn match(self: *const Zones, domain: []const u8) ?*const Zone;
|
||
};
|
||
```
|
||
|
||
`build` parses each `resolver` string with `validate.parseResolver` — the function whose doc comment
|
||
already names this file as its Phase 5 importer. There is no second resolver parser.
|
||
|
||
### S4.3 Tests (in-file)
|
||
|
||
`records.zig`:
|
||
- Build from A, AAAA and CNAME rows; `lookup("nas.lan", .a)` returns the A record; `.aaaa` returns
|
||
the AAAA; `.mx` returns nothing; with a CNAME present, every qtype returns the CNAME.
|
||
- Two A records for one name both come back, in a stable order across two builds.
|
||
- Uppercase and trailing-dot owner names normalize to the same key.
|
||
- A bad `a` value (`"::1"`), a bad `aaaa` value, a bad CNAME target and an unparseable owner each
|
||
produce their typed error.
|
||
- `writeAnswers` output re-parses with the expected ancount, types, TTLs and rdata.
|
||
- `build` under `checkAllAllocationFailures`.
|
||
|
||
`forward_zones.zig`:
|
||
- `lan.home` matches `nas.lan.home`, `a.b.lan.home` and `lan.home`; does not match `notlan.home`,
|
||
`home` or `lan.home.evil.net`.
|
||
- With both `home` and `lan.home` configured, `nas.lan.home` matches `lan.home` (longest wins).
|
||
- `10.in-addr.arpa` matches `5.4.3.10.in-addr.arpa`.
|
||
- A bad resolver URL → `error.BadResolver`; a bad zone → `error.BadZone`.
|
||
- `build` under `checkAllAllocationFailures`.
|
||
|
||
### S4.4 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on both files.
|
||
- [ ] Neither file takes a `std.Io` value or reads a clock.
|
||
- [ ] `forward_zones.zig` calls `validate.parseResolver`; no second resolver parser exists
|
||
(`grep -rn "udp://" src/local/` shows only test text).
|
||
- [ ] Suffix matching is on label boundaries, proven by the `notlan.home` test.
|
||
- [ ] Every bullet in S4.3 exists as a named test.
|
||
|
||
---
|
||
|
||
## Session S5: `src/filter/rules.zig`, `src/filter/matcher.zig`
|
||
|
||
Pure, and the heart of the milestone. Depends on S1's `wildcard.zig` and S2's `domain_set.zig`.
|
||
|
||
### S5.1 `rules.zig` — one group's explicit rules
|
||
|
||
```zig
|
||
pub const Error = error{ OutOfMemory, BadPattern, TooManyWildcards } || domain_set.DomainSet.Error;
|
||
pub const max_wildcards_per_group: usize = 4096;
|
||
|
||
pub const RuleSet = struct {
|
||
exact_allow: domain_set.DomainSet,
|
||
exact_block: domain_set.DomainSet,
|
||
/// Arena-owned, normalized, sorted for determinism. Scanned linearly:
|
||
/// these are operator-authored and few, and a linear scan over ≤ 4096
|
||
/// short patterns is far cheaper than any index that would have to be
|
||
/// rebuilt on every swap.
|
||
wildcard_allow: []const []const u8,
|
||
wildcard_block: []const []const u8,
|
||
|
||
pub const empty: RuleSet;
|
||
|
||
/// `rows` are this group's rules only (already filtered by the caller).
|
||
/// Patterns are normalized (lowercase, trailing dot stripped) and validated:
|
||
/// `.exact` through `dns.name.fromText`, `.wildcard` through
|
||
/// `wildcard.validate`. An invalid pattern is `error.BadPattern` — the
|
||
/// database went through `validate.zig`, so an invalid one here means the
|
||
/// rows were edited underneath nxdns.
|
||
pub fn build(gpa: std.mem.Allocator, rows: []const model.Rule, seed: u64) Error!RuleSet;
|
||
pub fn deinit(self: *RuleSet, gpa: std.mem.Allocator) void;
|
||
pub fn memoryBytes(self: *const RuleSet) usize;
|
||
};
|
||
```
|
||
|
||
`build` must be given only one group's rows; splitting `listRules` output by group is the caller's
|
||
job (S5.3's `Snapshot.build`), because only it holds the group table.
|
||
|
||
### S5.2 `matcher.zig` — normalization, candidates, decision
|
||
|
||
```zig
|
||
pub const Reason = enum {
|
||
none,
|
||
rule_allow_exact,
|
||
rule_block_exact,
|
||
rule_allow_wildcard,
|
||
rule_block_wildcard,
|
||
blocklist_domain,
|
||
blocklist_wildcard,
|
||
};
|
||
|
||
pub const Decision = struct {
|
||
blocked: bool,
|
||
reason: Reason,
|
||
/// The candidate (or pattern) that decided it; borrowed from the caller's
|
||
/// normalized buffer or from the snapshot. "" when `reason == .none`.
|
||
matched: []const u8,
|
||
/// `.blocklist_*` only: index into `Snapshot.sources`, for the block reason
|
||
/// the query log (Phase 6) and the UI (Phase 8) will show.
|
||
source: ?u32 = null,
|
||
};
|
||
|
||
/// Lowercase ASCII, trailing dot stripped, written into `buf`. Returns a slice
|
||
/// of `buf`. The root name normalizes to "".
|
||
pub fn normalize(qname: name.Name, buf: *[types.max_name_len]u8) []const u8;
|
||
|
||
/// Full name, then each parent, ending at the last two-label suffix. The TLD
|
||
/// alone is NOT a candidate: a rule or list entry on `com` is a configuration
|
||
/// mistake that would take the whole internet with it, and refusing to walk
|
||
/// that far costs nothing real.
|
||
pub const Candidates = struct {
|
||
rest: []const u8,
|
||
pub fn init(domain: []const u8) Candidates;
|
||
pub fn next(self: *Candidates) ?[]const u8;
|
||
};
|
||
```
|
||
|
||
### S5.3 `matcher.zig` — the snapshot
|
||
|
||
```zig
|
||
pub const SourceSets = struct {
|
||
/// Row id, so the manager can map a decision back to `blocklist_sources`.
|
||
id: i64,
|
||
/// Borrowed from the snapshot arena; the source's display name for the UI.
|
||
name: []const u8,
|
||
domains: domain_set.DomainSet,
|
||
wildcards: domain_set.DomainSet,
|
||
};
|
||
|
||
pub const Group = struct {
|
||
id: i64,
|
||
name: []const u8,
|
||
safe_search: bool,
|
||
rules: rules.RuleSet,
|
||
/// Indices into `Snapshot.sources`, ascending, deduplicated.
|
||
sources: []const u32,
|
||
};
|
||
|
||
pub const ClientEntry = struct { key: address.NetAddress.Key, group: u32 };
|
||
pub const PrefixEntry = struct { prefix: address.Prefix, group: u32, priority: i32 };
|
||
|
||
pub const Snapshot = struct {
|
||
arena: std.heap.ArenaAllocator,
|
||
groups: []Group,
|
||
sources: []SourceSets,
|
||
clients: []ClientEntry,
|
||
prefixes: []PrefixEntry,
|
||
/// Index into `groups` of the group named "default". Always valid:
|
||
/// `build` returns `error.MissingDefaultGroup` otherwise.
|
||
default_group: u32,
|
||
/// Monotonic, assigned by the manager. Logged on every swap so an operator
|
||
/// can tell which generation answered a query.
|
||
generation: u64,
|
||
|
||
pub const Input = struct {
|
||
groups: []const model.Group,
|
||
group_sources: []const model.GroupSource,
|
||
sources: []const model.BlocklistSource,
|
||
rules: []const model.Rule,
|
||
clients: []const model.Client,
|
||
prefixes: []const model.ClientPrefix,
|
||
/// One entry per `sources[i]`, in the same order: the compiled bodies
|
||
/// already read from disk with their headers stripped. An enabled source
|
||
/// whose bodies are absent is `error.MissingCompiledSource` — a silently
|
||
/// unenforced blocklist is exactly the failure PLAN §1.3 exists to
|
||
/// prevent.
|
||
compiled: []const Compiled,
|
||
seed: u64,
|
||
generation: u64,
|
||
};
|
||
|
||
pub const Compiled = struct { list_body: []const u8, wild_body: []const u8 };
|
||
|
||
pub const Error = error{
|
||
OutOfMemory, MissingDefaultGroup, UnknownGroup, UnknownSource,
|
||
MissingCompiledSource, BadClientIp, BadClientPrefix,
|
||
} || rules.Error;
|
||
|
||
/// Builds an immutable snapshot. Every string is copied into `arena`, so the
|
||
/// caller may free the repository lists immediately afterwards. Disabled
|
||
/// sources are skipped entirely — they cost no memory.
|
||
pub fn build(gpa: std.mem.Allocator, input: Input) Error!Snapshot;
|
||
pub fn deinit(self: *Snapshot) void;
|
||
|
||
/// PLAN §3.10 precedence, allow wins at equal specificity:
|
||
/// 1. exact/parent allow rules 2. exact/parent block rules
|
||
/// 3. wildcard allow rules 4. wildcard block rules
|
||
/// 5. blocklist domains 6. blocklist wildcards
|
||
/// `domain` is normalized (`normalize`). No allocation, no lock, no clock.
|
||
pub fn evaluate(self: *const Snapshot, group: u32, domain: []const u8) Decision;
|
||
|
||
/// PLAN §7.2 matching half: exact client row, else longest-prefix match
|
||
/// (ties broken by longer prefix then higher `priority`), else the default
|
||
/// group. Auto-materialization is Phase 7 (resolved ambiguity 4).
|
||
pub fn groupForClient(self: *const Snapshot, addr: address.NetAddress) u32;
|
||
|
||
pub fn groupIndexById(self: *const Snapshot, id: i64) ?u32;
|
||
pub fn groupIndexByName(self: *const Snapshot, name_text: []const u8) ?u32;
|
||
pub fn safeSearch(self: *const Snapshot, group: u32) bool;
|
||
pub fn memoryBytes(self: *const Snapshot) usize;
|
||
};
|
||
```
|
||
|
||
`evaluate` order is the specification, and it is level-by-level over the **whole candidate chain**,
|
||
not candidate-by-candidate over the levels: level 1 is checked against every candidate before level 2
|
||
is checked against any. That is what makes "an allow rule on the parent beats a block rule on the
|
||
child" true, which is the behaviour operators expect from an allow list.
|
||
|
||
Levels 5 and 6 iterate the group's sources in ascending index order and return the first hit, so the
|
||
reported `source` is stable for a given snapshot. Level 5 tests **only the full name** and level 6
|
||
tests **only proper parents** (resolved ambiguity 1).
|
||
|
||
### S5.4 Tests (in-file)
|
||
|
||
Precedence — one named test per row, each asserting `blocked`, `reason` and `matched`:
|
||
|
||
| Configuration | Query | Expected |
|
||
|---|---|---|
|
||
| block rule `ads.example.com`, nothing else | `ads.example.com` | blocked, `rule_block_exact` |
|
||
| block rule `example.com` | `ads.example.com` | blocked, `rule_block_exact`, matched `example.com` (parent walk) |
|
||
| block rule `example.com`, allow rule `ads.example.com` | `ads.example.com` | allowed, `rule_allow_exact` |
|
||
| block rule `ads.example.com`, allow rule `example.com` | `ads.example.com` | allowed (allow level runs first over the whole chain) |
|
||
| allow rule `*.example.com`, block rule `ads.example.com` | `ads.example.com` | blocked (exact block, level 2, beats wildcard allow, level 3) |
|
||
| allow wildcard `*.example.com`, block wildcard `*.example.com` | `a.example.com` | allowed (tie → allow wins) |
|
||
| list entry `tracker.net` | `tracker.net` | blocked, `blocklist_domain` |
|
||
| list entry `tracker.net` | `sub.tracker.net` | **allowed** (ambiguity 1) |
|
||
| wild entry `tracker.net` | `sub.tracker.net` | blocked, `blocklist_wildcard` |
|
||
| wild entry `tracker.net` | `tracker.net` | allowed (a `.wild` entry covers proper subdomains only) |
|
||
| allow rule `sub.tracker.net`, wild entry `tracker.net` | `sub.tracker.net` | allowed |
|
||
| nothing configured | `example.com` | allowed, `reason == .none` |
|
||
| source assigned to group A only | same query in group B | allowed in B, blocked in A |
|
||
| disabled source | its entry | allowed (disabled sources are not loaded) |
|
||
|
||
Also:
|
||
- `normalize` lowercases, strips one trailing dot, and returns `""` for the root.
|
||
- `Candidates` over `a.b.example.com` yields exactly `a.b.example.com`, `b.example.com`,
|
||
`example.com` — and not `com`.
|
||
- `groupForClient`: exact IPv4 hit; exact IPv6 hit through the canonical key; `/24` prefix hit;
|
||
overlapping `/16` and `/24` → the `/24` wins; equal-length prefixes → higher priority wins;
|
||
no match → default group.
|
||
- `Snapshot.build` without a `default` group → `error.MissingDefaultGroup`; with a `group_sources`
|
||
row naming an unknown source → `error.UnknownSource`; with an enabled source whose `compiled` entry
|
||
is absent → `error.MissingCompiledSource`.
|
||
- Two snapshots built from the same input with different seeds produce identical decisions over a
|
||
40-query table.
|
||
- `Snapshot.build` under `checkAllAllocationFailures` leaks nothing.
|
||
- `memoryBytes` on a snapshot with 10,000 synthetic domains stays under the S§Memory-budget bound
|
||
(`count * (avg_len + 1) * 2 + index`), asserted as an inequality so it is a real regression guard.
|
||
|
||
### S5.5 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on both files.
|
||
- [ ] Neither file takes a `std.Io` value, opens a file, or reads a clock.
|
||
- [ ] `evaluate` performs no allocation (no allocator is reachable from its signature).
|
||
- [ ] Every row of the S5.4 precedence table exists as a named test.
|
||
- [ ] The candidate chain stops before the TLD, proven by a test.
|
||
- [ ] `Snapshot.build` copies every string into its arena; a test frees the input lists before
|
||
calling `evaluate`.
|
||
|
||
---
|
||
|
||
## Session S6: `src/filter/fetcher.zig`
|
||
|
||
The first of the three `std.Io` files. Downloads one source; knows nothing about parsing, files or
|
||
the database.
|
||
|
||
### S6.1 Signatures
|
||
|
||
```zig
|
||
pub const max_body_bytes: usize = 64 * 1024 * 1024;
|
||
pub const redirect_buffer_len: usize = 8192; // RFC 9110 recommendation (Client.zig:1128)
|
||
pub const min_transfer_buf: usize = 16 * 1024;
|
||
|
||
pub const Error = error{
|
||
BadUrl, ConnectFailed, TlsFailed, SendFailed, ReceiveFailed, HttpStatus,
|
||
BodyTooLarge, Timeout, Canceled, OutOfMemory, SystemResources, Unexpected,
|
||
};
|
||
|
||
pub const Result = struct { bytes_read: u64, status: std.http.Status };
|
||
|
||
pub const Fetcher = struct {
|
||
http: *std.http.Client, // caller-owned, shared, pools connections
|
||
transfer_buf: []u8, // caller-owned, ≥ min_transfer_buf
|
||
redirect_buf: []u8, // caller-owned, ≥ redirect_buffer_len
|
||
|
||
/// GETs `url` and streams the body into `w`. Nothing is buffered whole:
|
||
/// a 64 MB list must not become a 64 MB allocation, and the caller is
|
||
/// writing into a temporary file anyway.
|
||
///
|
||
/// `budget` bounds the WHOLE exchange. `std.http.Client` has no per-request
|
||
/// deadline (verified), so the caller runs this under `io.concurrent` and
|
||
/// cancels the future — the same pattern `tls_client_integration_test.zig`
|
||
/// established. This function therefore takes no timeout parameter and
|
||
/// simply propagates `error.Canceled`; S8 owns the deadline.
|
||
pub fn fetch(self: *Fetcher, io: std.Io, url: []const u8, w: *std.Io.Writer) Error!Result;
|
||
};
|
||
```
|
||
|
||
Rules:
|
||
|
||
- `std.Uri.parse` failure → `error.BadUrl`. A scheme other than `http`/`https` → `error.BadUrl`.
|
||
- `accept_encoding` is overridden to identity (resolved ambiguity 6), matching `doh_client.zig`.
|
||
- `redirect_behavior` keeps the stdlib default of 3; `receiveHead(self.redirect_buf)` follows them.
|
||
`error.TooManyHttpRedirects` maps to `error.HttpStatus`.
|
||
- Status other than `.ok` → `error.HttpStatus`. An error return carries no `Result`, so the fetcher
|
||
stores the numeric status in a `last_status: ?std.http.Status` field, cleared at the start of each
|
||
`fetch` and set from every response head; the caller reads it after `error.HttpStatus`. `content-type` is **not** checked: blocklists are served as `text/plain`,
|
||
`application/octet-stream`, `text/html` and worse, and the compiler's invalid-line counters are the
|
||
honest signal about content.
|
||
- The body is copied to `w` in `transfer_buf`-sized chunks with a running total; exceeding
|
||
`max_body_bytes` → `error.BodyTooLarge` (and the caller discards its temporary file).
|
||
- Error mapping reuses `transport.mapLocal` first, then classifies by phase, exactly as
|
||
`doh_client.zig`'s `mapError` does. Do not invent a second classification vocabulary.
|
||
|
||
### S6.2 Tests
|
||
|
||
In-file (no sockets): URL rejection table (`ftp://x`, `x`, `https://`), and a `Fetcher` value
|
||
constructed against an undefined `std.http.Client` proving the URL check precedes any client use
|
||
(the `doh_client.zig` pattern).
|
||
|
||
Everything real belongs to S9: a loopback `std.http.Server` serving a fixture, a redirect chain, a
|
||
404, and a body that exceeds a lowered cap.
|
||
|
||
### S6.3 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean.
|
||
- [ ] `accept_encoding` is pinned to identity and the reason is commented.
|
||
- [ ] The body is streamed; no `allocRemaining`, no whole-body allocation anywhere in the file.
|
||
- [ ] No timeout parameter and no sleep: the deadline is the caller's (S8.4).
|
||
- [ ] `transport.mapLocal` is used before any phase classification.
|
||
|
||
---
|
||
|
||
## Session S7: `src/local/forward_client.zig`
|
||
|
||
Conditional forwarding's transport (resolved ambiguity 5). Implements `transport.Client` so the
|
||
Phase 7 handler treats a forward zone exactly like any other exchange.
|
||
|
||
### S7.1 Signatures
|
||
|
||
```zig
|
||
pub const ForwardClient = struct {
|
||
resolver: validate.Resolver,
|
||
/// Caller-owned scratch for the TCP length-prefixed path.
|
||
frame_buf: []u8,
|
||
read_timeout: std.Io.Clock.Duration,
|
||
stats: Stats = .{},
|
||
|
||
pub const Stats = struct {
|
||
queries: u64 = 0,
|
||
udp_truncated: u64 = 0, // TC=1 → retried over TCP
|
||
failures: u64 = 0,
|
||
};
|
||
|
||
pub fn init(resolver: validate.Resolver, frame_buf: []u8, read_timeout: std.Io.Clock.Duration) ForwardClient;
|
||
pub fn client(self: *ForwardClient) transport.Client;
|
||
|
||
/// UDP: send, `receiveTimeout`, validate. TC=1 → retry over TCP with the
|
||
/// RFC 1035 §4.2.2 two-byte length prefix. `.tcp` resolvers skip straight
|
||
/// to the TCP path.
|
||
pub fn exchange(self: *ForwardClient, io: std.Io, query: []const u8, response_buf: []u8)
|
||
transport.ExchangeError![]u8;
|
||
};
|
||
```
|
||
|
||
Rules:
|
||
|
||
- Every response passes `transport.validateResponse(query, response)` before it is returned — a
|
||
forward zone points at LAN infrastructure, which is not a reason to trust its framing.
|
||
- UDP uses `Socket.receiveTimeout` (verified implemented on the POSIX Threaded backend, milestone 3).
|
||
A datagram from an address other than the resolver's is **discarded and counted**, and the receive
|
||
is retried within the remaining budget.
|
||
- TCP has no read timeout in 0.16.0: run the exchange under `io.concurrent` and cancel the loser
|
||
against a `std.Io.Clock.Duration` sleep, exactly as milestone 3 does. Never set
|
||
`ConnectOptions.timeout` — the Threaded backend panics (Threaded.zig:12077).
|
||
- Failures map through `transport.mapLocal` first, then to the phase's `PeerFault`. Every socket is
|
||
closed on every path.
|
||
- No health tracking and no backoff: `upstream/health.zig` and `pool.zig` model the *upstream* pool,
|
||
and a forward zone has exactly one designated resolver with no failover partner. Say so in a
|
||
comment so its absence does not read as an oversight.
|
||
|
||
### S7.2 Tests
|
||
|
||
In-file: the `transport.Client` vtable instantiation check (milestone 3's pattern) and the stats
|
||
struct defaults. Everything with a socket is S9's: a loopback UDP responder answering an A query, a
|
||
loopback responder setting TC=1 followed by a TCP responder returning the full answer, a silent
|
||
responder proving the timeout produces `error.Timeout`, and a responder answering with a mismatched
|
||
ID proving `error.ResponseMismatch`.
|
||
|
||
### S7.3 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean.
|
||
- [ ] `ConnectOptions.timeout` appears nowhere.
|
||
- [ ] Every response goes through `transport.validateResponse`.
|
||
- [ ] A datagram from a foreign source address is counted and discarded, not returned.
|
||
- [ ] The file implements `transport.Client` and defines no second client interface.
|
||
|
||
---
|
||
|
||
## Session S8: `src/filter/manager.zig`
|
||
|
||
The `BlocklistManager` of PLAN §4: compiled-file layout, refresh, metadata, snapshot build and the
|
||
RCU swap. The only file in this milestone that touches both the database and the filesystem.
|
||
|
||
### S8.1 The two repository additions this session owns
|
||
|
||
`blocklist_sources` carries counters and a checksum that milestone 4 deliberately excluded from the
|
||
config model (runtime facts, not configuration). Refresh has to write them, so `sources_repo.zig`
|
||
gains exactly two functions — additive, no signature of an existing function changes:
|
||
|
||
```zig
|
||
pub const SourceRow = struct {
|
||
id: i64,
|
||
url: []const u8,
|
||
name: []const u8,
|
||
enabled: bool,
|
||
last_updated: ?i64,
|
||
domain_count: i64,
|
||
wildcard_count: i64,
|
||
skipped_regex_count: i64,
|
||
checksum: ?[]const u8,
|
||
};
|
||
|
||
/// Every source with its row id and runtime columns. Strings are heap copies.
|
||
/// ORDER BY url (matching `listBlocklistSources`, milestone 4 §S4.2).
|
||
pub fn listSourceRows(database: *db.Db, gpa: Allocator) db.Error!std.ArrayList(SourceRow);
|
||
pub fn freeSourceRows(gpa: Allocator, items: []const SourceRow) void;
|
||
|
||
/// Writes the runtime columns for one source after a successful compile.
|
||
pub fn updateSourceStats(database: *db.Db, id: i64, stats: SourceStats) db.Error!void;
|
||
```
|
||
|
||
This is the one exception to "no session edits a milestone 1–4 file", and it is scoped to appending
|
||
two functions and their tests to `sources_repo.zig`. S8 owns that file for this milestone. No
|
||
schema change and no migration: every column already exists (PLAN §11.2).
|
||
|
||
### S8.2 Compiled-file layout (PLAN §3.8, §3.13)
|
||
|
||
```
|
||
<data_dir>/blocklists/<source_id>.list
|
||
<data_dir>/blocklists/<source_id>.wild
|
||
```
|
||
|
||
Directory created with `createDirPathStatus(io, path, .fromMode(0o700))`; files written 0o600 through
|
||
`createFileAtomic` + `replace` (milestone 4's verified helper), so a crash mid-write can never leave
|
||
a half-list that would load as a valid, shorter blocklist. `<source_id>` is the `blocklist_sources`
|
||
row id, so a renamed URL keeps its files and a deleted source's files are removed by
|
||
`pruneOrphans` (below).
|
||
|
||
Header, written by the manager ahead of the body, every line prefixed `# `:
|
||
|
||
```
|
||
# nxdns blocklist
|
||
# url <url>
|
||
# format <hosts|domains|abp>
|
||
# fetched_at <unix seconds>
|
||
# domains <n>
|
||
# wildcards <n>
|
||
# skipped_regex <n>
|
||
# skipped_unsupported <n>
|
||
# invalid <n>
|
||
# sha256 <64 hex chars>
|
||
```
|
||
|
||
The loader strips every leading `#` line and hands the remainder to `DomainSet.build`. The `sha256`
|
||
covers the `.list` body followed by the `.wild` body — **not** the header, so the checksum is stable
|
||
across refetches of unchanged content while `fetched_at` moves.
|
||
|
||
### S8.3 Signatures
|
||
|
||
```zig
|
||
pub const Paths = struct {
|
||
dir: std.Io.Dir, // <data_dir>
|
||
subdir: []const u8 = "blocklists",
|
||
};
|
||
|
||
pub const SourceStatus = struct {
|
||
id: i64,
|
||
url: []const u8, // borrowed from the manager's arena
|
||
state: enum { ok, never_fetched, fetch_failed, compile_failed, load_failed },
|
||
last_attempt: i64 = 0,
|
||
last_success: i64 = 0,
|
||
counts: compiler.Counts = .{},
|
||
/// Fixed-size, no allocation on the failure path.
|
||
last_error: [128]u8 = @splat(0),
|
||
last_error_len: u8 = 0,
|
||
};
|
||
|
||
pub const Manager = struct {
|
||
gpa: std.mem.Allocator,
|
||
database: *db.Db,
|
||
paths: Paths,
|
||
fetcher: *fetcher.Fetcher,
|
||
update: model.BlocklistUpdate,
|
||
total_budget: std.Io.Clock.Duration,
|
||
|
||
lock: std.Io.RwLock,
|
||
current: ?*matcher.Snapshot,
|
||
generation: u64,
|
||
statuses: []SourceStatus,
|
||
|
||
pub const Error = error{ OutOfMemory, ... } || db.Error || matcher.Snapshot.Error;
|
||
|
||
pub fn init(gpa, database, paths, fetcher_ptr, update, total_budget) Error!Manager;
|
||
pub fn deinit(self: *Manager, io: std.Io) void;
|
||
|
||
/// Reads the database and every compiled file, builds a snapshot and swaps
|
||
/// it in. Called at startup and after any refresh. A source whose compiled
|
||
/// files are missing or whose checksum does not match is marked
|
||
/// `.load_failed` and REFRESHED, not silently skipped.
|
||
pub fn reload(self: *Manager, io: std.Io) Error!void;
|
||
|
||
/// Fetch + compile + atomically replace the compiled files for one source,
|
||
/// then update its row. Returns false when the content was unchanged
|
||
/// (checksum equal), in which case only `last_updated` moves.
|
||
pub fn refreshSource(self: *Manager, io: std.Io, row: sources_repo.SourceRow) Error!bool;
|
||
|
||
/// Every enabled source, one at a time, then one `reload`. Never stops at
|
||
/// the first failure: a broken source must not hide the rest.
|
||
pub fn refreshAll(self: *Manager, io: std.Io) Error!void;
|
||
|
||
/// Long-running task: sleeps `update.interval_hours`, refreshes, repeats.
|
||
/// Returns on `error.Canceled`. Started with `io.concurrent` by Phase 7's
|
||
/// wiring; nothing in this milestone starts it automatically.
|
||
pub fn runScheduler(self: *Manager, io: std.Io) std.Io.Cancelable!void;
|
||
|
||
/// Reader side of the RCU swap. The handle holds a shared lock; the caller
|
||
/// MUST release it and MUST NOT retain the snapshot pointer afterwards.
|
||
pub const Handle = struct {
|
||
snapshot: *const matcher.Snapshot,
|
||
manager: *Manager,
|
||
pub fn release(self: Handle, io: std.Io) void;
|
||
};
|
||
pub fn acquire(self: *Manager, io: std.Io) ?Handle;
|
||
|
||
/// Copies the status table for the API and `nxdns check`.
|
||
pub fn statusSnapshot(self: *Manager, io: std.Io, out: []SourceStatus) usize;
|
||
|
||
/// Deletes `<id>.list`/`<id>.wild` files whose id is not in the database.
|
||
pub fn pruneOrphans(self: *Manager, io: std.Io) Error!void;
|
||
};
|
||
```
|
||
|
||
**The swap is an `std.Io.RwLock`, not a lock-free pointer.** PLAN §7.3 says "readers lock-free", and
|
||
this is a deliberate, documented deviation. Freeing the old snapshot safely without a lock requires
|
||
epoch-based reclamation or hazard pointers — a class of code that is very hard to get right and
|
||
impossible to test convincingly, bought for a household resolver whose target is 100 qps. A shared
|
||
lock held for the microseconds of one `evaluate` costs an uncontended atomic pair; the writer takes
|
||
the exclusive lock only on a swap, which happens on refresh. The old snapshot is freed **after**
|
||
`unlock` returns, and the `Handle` API makes "do not retain the pointer" the only shape the caller
|
||
can write. State this reasoning in the file, and have the orchestrator amend PLAN §7.3.
|
||
|
||
### S8.4 `refreshSource` order — the order is the specification
|
||
|
||
1. `now = std.Io.Clock.real.now(io).toSeconds()`; `status.last_attempt = now`.
|
||
2. Create `<subdir>/<id>.list.tmp` and `<id>.wild.tmp` through `createFileAtomic`
|
||
(`.permissions = .fromMode(0o600)`, `.replace = true`).
|
||
3. Download into a third temporary file (`<id>.raw.tmp`) rather than memory — `max_body_bytes` is
|
||
64 MB and the budget in §Memory has no room for it alongside two snapshots. Run
|
||
`fetcher.fetch` under `io.concurrent` with a `total_budget` sleep racing it, and cancel the loser
|
||
(milestone 3's pattern). A fetch failure sets `.fetch_failed` with the error name, logs at `warn`,
|
||
and returns without touching the live compiled files.
|
||
4. Reopen the raw file, sniff the format with `parsers.detectFormat` over the first
|
||
`parsers.sample_lines` lines, then `compiler.compile` from a `File.Reader` into the two temporary
|
||
writers. Any compile error sets `.compile_failed` and returns; the live files are still untouched.
|
||
5. If `result.checksum` equals the stored `checksum`, `deinit` the atomics (which discards the
|
||
temporaries), update only `last_updated`, and return `false`. Recompiling identical content into
|
||
a new file would invalidate the snapshot for nothing.
|
||
6. Otherwise write the header, `replace` both atomics, delete the raw temporary, and
|
||
`updateSourceStats` with the counts, the checksum and `last_updated = now`.
|
||
7. `status.state = .ok`, `status.counts = result.counts`, `status.last_success = now`.
|
||
|
||
Every failure path deletes its temporaries. No path leaves a `.tmp` behind, and no path deletes a
|
||
good compiled file.
|
||
|
||
### S8.5 Failure visibility
|
||
|
||
- Every non-`ok` state records the error name into `SourceStatus.last_error` and logs once at `warn`
|
||
with the source URL and the state. No `err` level, per the binding logging policy: the condition is
|
||
recorded and surfaced, not swallowed.
|
||
- `refreshAll` returns successfully when at least one source failed; the failures live in
|
||
`statuses`. It returns an error only when *nothing* could be done (out of memory, the database is
|
||
unreachable). Phase 8 exposes the statuses at `GET /api/blocklists`; `nxdns check` gains nothing in
|
||
this milestone.
|
||
- `reload` with zero enabled sources is normal (a fresh install), produces a valid empty snapshot,
|
||
and logs at `info`.
|
||
|
||
### S8.6 Startup and scheduling policy
|
||
|
||
`runScheduler` refreshes a source at startup only when it needs it: no compiled file, a checksum
|
||
mismatch against the file on disk, or `last_updated` older than `interval_hours`. A cold restart of
|
||
a Pi must not re-download every list, and a boot loop must not turn into a download loop. After the
|
||
initial pass it sleeps `interval_hours` between full passes. `update.enabled == false` means
|
||
`runScheduler` returns immediately after the initial load; manual refresh through `refreshAll` still
|
||
works (Phase 8's `POST /api/blocklists/update`).
|
||
|
||
### S8.7 Tests
|
||
|
||
In-file (`:memory:` database, no filesystem): `init`/`deinit`; `acquire` before any `reload` returns
|
||
`null`; the header writer produces the exact expected text for a known `Counts`; the header stripper
|
||
returns the body for a header-only file and for a file with no header; `SourceStatus.last_error`
|
||
truncation at 128 bytes.
|
||
|
||
Everything with real files, real HTTP and real swaps is S9's.
|
||
|
||
### S8.8 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on `manager.zig` and `sources_repo.zig`.
|
||
- [ ] `sources_repo.zig` gains exactly two public functions plus `SourceRow`/`freeSourceRows`; no
|
||
existing signature changes; the new list has a deterministic `ORDER BY url`.
|
||
- [ ] No schema change, no migration step, no new column.
|
||
- [ ] Compiled files are written through `createFileAtomic` + `replace` at 0o600, in a 0o700
|
||
directory created with `createDirPathStatus`.
|
||
- [ ] A fetch or compile failure leaves the previously compiled files byte-identical (test in S9).
|
||
- [ ] The RwLock deviation from PLAN §7.3 is documented in the file with its reasoning.
|
||
- [ ] The old snapshot is freed after `unlock`, never while a reader could hold it.
|
||
- [ ] `refreshAll` does not stop at the first failing source.
|
||
- [ ] `grep -n "std.log.err" src/filter/ src/local/` returns nothing.
|
||
|
||
---
|
||
|
||
## Session S9: `src/filter/filter_integration_test.zig`, `tests/fuzz/blocklist_fuzz.zig`
|
||
|
||
### S9.1 Hermetic integration cases (`-Dintegration`, `std.testing.tmpDir`)
|
||
|
||
*compile → load → match, through real files*
|
||
|
||
1. Compile a 5,000-line hosts fixture to real `.list`/`.wild` files, build a snapshot from them, and
|
||
assert 20 sampled domains match and 20 non-members do not.
|
||
2. Recompile the same fixture into a second directory: both files are byte-identical, and so is the
|
||
checksum.
|
||
3. Truncate a `.list` file mid-line and reload → `error.NotSorted` or `MissingCompiledSource`, and
|
||
**the previous snapshot is still serving** (assert through `acquire`).
|
||
|
||
*fetcher against a loopback `std.http.Server`*
|
||
|
||
4. A 200 response with a hosts body → the compiled files exist, `blocklist_sources` counters and
|
||
checksum are written, and `acquire` sees the domains.
|
||
5. A 302 to a second path → followed, same result.
|
||
6. A 404 → `.fetch_failed`, the previously compiled files are byte-identical, and the snapshot still
|
||
blocks what it blocked before.
|
||
7. A body larger than a lowered `max_body_bytes` → `error.BodyTooLarge`, no `.tmp` file remains in
|
||
the blocklist directory (assert by listing it).
|
||
8. Refetching identical content → `refreshSource` returns `false`, the compiled files' mtimes are
|
||
unchanged, and `last_updated` moved.
|
||
|
||
*swap*
|
||
|
||
9. Under a snapshot acquired by one task, a concurrent `reload` completes and the holding task still
|
||
reads a consistent snapshot; after `release` and a re-`acquire`, the new generation is visible.
|
||
10. `pruneOrphans` deletes files for a deleted source and leaves live ones alone.
|
||
|
||
*local records and forward zones, end to end*
|
||
|
||
11. Build `Records` from a seeded database (through `local_repo.listLocalRecords`), answer an A query
|
||
for `nas.lan`, and re-parse the reply: one answer, correct rdata, correct TTL.
|
||
12. A CNAME local record answers an A query with the CNAME record and nothing else.
|
||
13. `ForwardClient` against a loopback UDP responder: an A query for `nas.lan.home` returns the
|
||
responder's answer, and `Zones.match` selected that resolver.
|
||
14. The loopback responder sets TC=1; the client retries over TCP against a loopback TCP responder
|
||
and returns the full answer, with `stats.udp_truncated == 1`.
|
||
15. A silent responder → `error.Timeout` inside the configured budget (assert the elapsed time is
|
||
below twice the budget).
|
||
16. A responder answering with a wrong ID → `error.ResponseMismatch`.
|
||
|
||
*blocked responses*
|
||
|
||
17. For each `blocking.response` mode, synthesize a reply for a blocked A and AAAA query and re-parse
|
||
it: `.zero` gives `0.0.0.0` / `::` with `blocking.ttl`; `.nxdomain` gives NXDOMAIN with no answer.
|
||
|
||
### S9.2 Fuzz target
|
||
|
||
`tests/fuzz/blocklist_fuzz.zig`, in the style of `tests/fuzz/dns_fuzz.zig` (milestone 2): a
|
||
`std.testing.fuzz` test per parser plus one for `wildcard.matches`, each filling a buffer through
|
||
`Smith.sliceWithHash` and feeding the prefix in. The acceptance property is "does not crash and does
|
||
not hang": every parser returns a `Line` for any byte string, and `wildcard.matches` terminates for
|
||
any pattern/domain pair (the pattern is first passed through `validate`, and only accepted patterns
|
||
reach `matches`, matching how production uses it). Corpus entries: one hosts line, one ABP line, one
|
||
regex line, one over-long line.
|
||
|
||
### S9.3 The build change (orchestrator, not this session)
|
||
|
||
```zig
|
||
const parsers_mod = b.createModule(.{
|
||
.root_source_file = b.path("src/filter/parsers.zig"),
|
||
.target = target,
|
||
.optimize = optimize,
|
||
});
|
||
const blocklist_fuzz_mod = b.createModule(.{
|
||
.root_source_file = b.path("tests/fuzz/blocklist_fuzz.zig"),
|
||
.target = target,
|
||
.optimize = optimize,
|
||
});
|
||
blocklist_fuzz_mod.addImport("parsers", parsers_mod);
|
||
const blocklist_fuzz_tests = b.addTest(.{
|
||
.name = "blocklist-fuzz",
|
||
.use_llvm = if (fuzz) true else null,
|
||
.root_module = blocklist_fuzz_mod,
|
||
});
|
||
test_step.dependOn(&b.addRunArtifact(blocklist_fuzz_tests).step);
|
||
```
|
||
|
||
This is exactly why S1's five files import nothing outside `src/filter/`: `parsers.zig` is a module
|
||
root here and in the aggregator, and an import that escaped its directory would fail with
|
||
`error.ImportOutsideModulePath` (milestone 2). The in-file tests of those five files still run in the
|
||
aggregator artifact; the fuzz artifact runs only `blocklist_fuzz.zig`'s tests.
|
||
|
||
### S9.4 Acceptance criteria
|
||
|
||
- [ ] `zig fmt --check` and `zig ast-check` clean on both files.
|
||
- [ ] All 17 hermetic cases exist as named tests and pass under `zig build test -Dintegration`.
|
||
- [ ] No hermetic test reaches a non-loopback address or leaves a file outside its `tmpDir`.
|
||
- [ ] The fuzz tests pass under plain `zig build test` (corpus replay + empty input).
|
||
|
||
---
|
||
|
||
## Module Layout
|
||
|
||
```
|
||
src/filter/parsers.zig S1 format vocabulary + dispatch + sniffing (std only)
|
||
src/filter/parser_hosts.zig S1 hosts-file lines (std only)
|
||
src/filter/parser_domains.zig S1 one-domain-per-line lists (std only)
|
||
src/filter/parser_abp.zig S1 ABP subset; modifiers/exceptions unsupported (std only)
|
||
src/filter/wildcard.zig S1 `*`-label pattern validate + match (std only)
|
||
src/filter/domain_set.zig S2 arena + open-addressed exact set
|
||
src/filter/compiler.zig S2 stream -> sorted, deduped .list/.wild bodies + sha256
|
||
src/filter/safesearch.zig S3 per-group safe-search table + rewrite
|
||
src/filter/response.zig S3 blocked-response synthesis (zero | nxdomain)
|
||
src/filter/rules.zig S5 one group's compiled allow/block rules
|
||
src/filter/matcher.zig S5 Snapshot, §3.10 precedence, group-for-client
|
||
src/filter/fetcher.zig S6 std.http.Client GET, streamed, capped
|
||
src/filter/manager.zig S8 compiled files, refresh, metadata, RCU swap, scheduler
|
||
src/filter/filter_integration_test.zig S9 -Dintegration: real files, loopback HTTP/UDP/TCP
|
||
src/local/records.zig S4 local A/AAAA/CNAME lookup + answer writing
|
||
src/local/forward_zones.zig S4 zone suffix match -> validate.Resolver
|
||
src/local/forward_client.zig S7 plain UDP/TCP resolver, transport.Client
|
||
src/storage/repositories/sources_repo.zig S8 += listSourceRows, updateSourceStats
|
||
tests/fuzz/blocklist_fuzz.zig S9 parser + wildcard fuzz targets
|
||
```
|
||
|
||
## File Ownership
|
||
|
||
| Files | Owner | Notes |
|
||
|---|---|---|
|
||
| `src/filter/parsers.zig`, `parser_hosts.zig`, `parser_domains.zig`, `parser_abp.zig`, `wildcard.zig` | S1 | frozen after S1 verifies; `std`-only imports |
|
||
| `src/filter/domain_set.zig`, `src/filter/compiler.zig` | S2 | frozen after S2 verifies |
|
||
| `src/filter/safesearch.zig`, `src/filter/response.zig` | S3 | |
|
||
| `src/local/records.zig`, `src/local/forward_zones.zig` | S4 | |
|
||
| `src/filter/rules.zig`, `src/filter/matcher.zig` | S5 | |
|
||
| `src/filter/fetcher.zig` | S6 | |
|
||
| `src/local/forward_client.zig` | S7 | |
|
||
| `src/filter/manager.zig`, `src/storage/repositories/sources_repo.zig` | S8 | the only milestone-4 file any session edits, and only additively (S8.1) |
|
||
| `src/filter/filter_integration_test.zig`, `tests/fuzz/blocklist_fuzz.zig` | S9 | |
|
||
| `build.zig`, `build.zig.zon`, `src/tests.zig` | orchestrator | no session edits these |
|
||
| `PLAN.md` §7.1, §7.3 | orchestrator | amended per resolved ambiguities 1 and 3, and per S8.3 |
|
||
|
||
No session touches `src/dns/`, `src/server/`, `src/upstream/`, `src/platform/`, `src/config/`,
|
||
`src/cli.zig` or `src/main.zig`. A needed change there is reported, not made.
|
||
|
||
## Acceptance Criteria (Milestone 5 Complete)
|
||
|
||
- [ ] `zig build test` exits 0 with every new file wired into `src/tests.zig`, including the new
|
||
`blocklist-fuzz` artifact.
|
||
- [ ] `zig build test -Dintegration` exits 0: milestone 1's loopback TLS echo, milestone 3's
|
||
listener and resolver tests, milestone 4's 21 storage cases, and all 17 filtering cases.
|
||
- [ ] `zig build cross` still produces two statically linked executables.
|
||
- [ ] The PLAN §3.10 precedence table is proven row by row by the S5.4 tests.
|
||
- [ ] Compiling a fixture twice produces byte-identical bodies and checksums; compiling a shuffled
|
||
permutation produces the same bodies.
|
||
- [ ] `grep -rn "std.log.err" src/filter/ src/local/` returns nothing.
|
||
- [ ] `grep -rln "std.Io," src/filter/ src/local/` names only `fetcher.zig`, `manager.zig`,
|
||
`forward_client.zig` and the S9 test file `filter_integration_test.zig` (reader/writer
|
||
interface parameters do not count; check the match).
|
||
- [ ] `grep -rn "regex" src/filter/` shows counting and skipping only — no matching engine.
|
||
- [ ] A hosts fixture containing `localhost`, `ip6-localhost` and `broadcasthost` compiles to a body
|
||
containing none of them.
|
||
- [ ] A snapshot over a 1,000,000-entry synthetic list reports `memoryBytes()` under 40 MB (S5.4's
|
||
inequality test, scaled) — the runnable form of the PLAN §18 budget.
|
||
- [ ] `zig fmt --check` clean repo-wide; GPG-signed lowercase commits.
|
||
|
||
## Anti-Requirements
|
||
|
||
- **No handler integration.** `src/server/handler.zig` is not edited. Nothing in this milestone is
|
||
called from the query path; composing local records → forward zones → filtering → safe-search →
|
||
cache → upstream is Phase 7. The pieces are built and tested standalone.
|
||
- **No CNAME uncloaking.** PLAN §6.3 is explicitly Phase 7 (§16). No chain walking, no depth-8
|
||
budget, no re-evaluation of answer-section names anywhere in this milestone.
|
||
- **No cache, no rate limiter, no query log, no disk monitor, no retention, no log rotation.**
|
||
Phase 6. A `Decision` carries a `reason` for the future query log; nothing writes one.
|
||
- **No pause/resume.** Phase 7.
|
||
- **No web API, no SSE, no `/metrics`, no auth.** Phase 8. `SourceStatus` exists for that phase to
|
||
read; no handler exists.
|
||
- **No regex engine, in any form, for any reason.** PLAN §2.2. Regex lines are counted and skipped.
|
||
- **No allow rules sourced from blocklists.** ABP exception rules (`@@`) are counted as unsupported.
|
||
Allow policy lives in the `rules` table where an operator can see it.
|
||
- **No blocklist storage in SQLite.** PLAN Decision A: compiled flat files under
|
||
`<data_dir>/blocklists/`, metadata columns only in `config.db`.
|
||
- **No schema change and no migration step.** Every column this milestone writes already exists.
|
||
- **No HTTP compression, no conditional requests (`ETag`/`If-Modified-Since`), no HTTP/2.**
|
||
Unchanged content is detected by comparing the compiled checksum, which also catches a source that
|
||
changes its headers without changing its content.
|
||
- **No client auto-materialization, no `last_seen` updates.** Phase 7 (resolved ambiguity 4).
|
||
- **No snapshot persistence.** The snapshot is rebuilt from the database and the compiled files at
|
||
startup; caching it on disk would add a fourth thing that can be stale.
|
||
- **No lock-free reclamation scheme.** The documented `RwLock` deviation (S8.3) is the design, not a
|
||
placeholder for a later epoch-based rewrite.
|
||
- **No third-party Zig packages.** stdlib plus the two pinned C libraries.
|
||
|
||
## As built (S1–S4, S6, S7 and orchestrator wiring)
|
||
|
||
Deviations from the text above, recorded after the first six sessions verified. Where this section
|
||
and the session text disagree, this section wins.
|
||
|
||
**S1 parsers + wildcard.** `validate` counts labels before the 253-byte length check, so
|
||
`TooManyLabels` is reachable. A hosts line with a sink address and no name is `.unsupported`.
|
||
`||example.com` without `^` is `.wildcard` with `covers_apex = true`. `detectFormat` treats `$` as a
|
||
marker only on non-comment lines; `||`, `@@` and element-hiding separators are matched anchored.
|
||
`looksLikeIpLiteral` lives in `parsers.zig` and is shared with `parser_hosts.zig`.
|
||
`wildcard.matches` is total on unvalidated input: the label splitter returns null above `max_labels`
|
||
and `matches` returns false — no assert, because `parsers.zig` is the S9.3 fuzz root.
|
||
|
||
**S2 domain_set + compiler.** The line-length cap is enforced both ways: `error.StreamTooLong` is
|
||
discarded with `discardDelimiterInclusive`, and a returned line over `max_line_len` is skipped;
|
||
both increment `long_lines`. `DomainSet.build` maps an empty line to `error.NotSorted` and a line
|
||
over 255 bytes to `error.SetTooLarge`; per line, the lowercase check runs before the order check.
|
||
`covers_apex` applies only to `.wildcard` lines: `*.x` in a hosts or domains list is a wildcard with
|
||
no apex entry.
|
||
|
||
**S3 safesearch + response.** The safe-search table as written above was not sorted; the built table
|
||
holds the same entry set sorted ascending by `std.mem.order` (`www.youtube-nocookie.com` <
|
||
`www.youtube.com`; `youtube.com` < `youtube.googleapis.com` < `youtubei.googleapis.com`). `rewrite`
|
||
uses `catch unreachable` on `name.fromText`, upheld by the every-entry-is-valid test. The
|
||
mode × qtype × OPT matrix is eight named tests, each looping over OPT presence internally; all
|
||
sixteen cases are exercised and re-parsed.
|
||
|
||
**S4 records + forward_zones.** A CNAME is exclusive at a name (RFC 1034 §3.6.2): records sort by
|
||
(owner, rtype) with rank A, AAAA, CNAME, and a name carrying a CNAME answers with the CNAME run
|
||
alone. `validate.zig` does not reject a CNAME coexisting with an A, so the mixed case is reachable
|
||
from a hand-edited database. `qtype == .any` returns the whole record run. The root name is
|
||
rejected (`BadRecordName` / `BadZone`), as is any byte ≥ 0x80 in an owner, CNAME target or zone.
|
||
Each table is `items` plus one flat byte block — two surviving allocations, freed by `deinit`.
|
||
|
||
**S6 fetcher.** `Fetcher.last_status: ?std.http.Status` carries the numeric status past
|
||
`error.HttpStatus` (see S6.1). `error.WriteFailed` from the output writer maps to
|
||
`error.Unexpected`; the caller owns the writer and reads the concrete failure there. A `narrowLocal`
|
||
step folds `transport.mapLocal`'s fd-quota members into `error.SystemResources` and the rest into
|
||
`error.Unexpected`. A declared `content-length` over `max_body_bytes` returns `error.BodyTooLarge`
|
||
before the body streams.
|
||
|
||
**S7 forward_client.** `read_timeout` is `std.Io.Clock.Duration` (S9 constructs it as
|
||
`.{ .raw = .fromMilliseconds(200), .clock = .awake }`). `Stats` has a fourth field,
|
||
`foreign_datagrams`. `failures` counts `.peer_fault` and `.local_resource` only — cancellations are
|
||
shutdown, not failure. The TC bit is read only after `validateResponse` matches the reply.
|
||
`frame_buf` is split in half between the stream writer and reader; `min_frame_buf = 1024` is
|
||
asserted in `init`.
|
||
|
||
**Orchestrator wiring.** All thirteen new files are imported by `src/tests.zig` individually
|
||
(`wildcard.zig` is unreachable through `parsers.zig` for test collection). The TCP framing helpers
|
||
(`prefix_len`, `framePrefix`, `parsePrefix`) moved to `upstream/transport.zig` as their single home;
|
||
`dot_client.zig`, `tcp_server.zig`, `forward_client.zig` and both server integration tests now use
|
||
the `transport.*` forms, and the canonical framing tests live in `transport.zig`. `zig build test`
|
||
exits 0 after the move.
|
||
|
||
**S5 rules + matcher.** `Snapshot.Input` gains `group_ids: []const i64` and `source_ids: []const
|
||
i64`, parallel to `groups` and `sources` (the model structs carry no row id); length mismatch is an
|
||
assert. `compiled` is `[]const ?Compiled` — `null` for an enabled source is
|
||
`error.MissingCompiledSource`; a disabled source needs no entry. On a prefix-length tie the *lower*
|
||
priority number wins, matching `address.matchLongest` and `pool.zig`. `max_wildcards_per_group`
|
||
counts both wildcard lists of a group combined, over rows before deduplication. Duplicate rule rows
|
||
are deduplicated before `DomainSet.build`. Level-6 nesting iterates parents outermost (most specific
|
||
first), sources innermost. A `group_sources` row naming a disabled source is legal; an unknown URL is
|
||
`error.UnknownSource`, an unknown group `error.UnknownGroup`. `RuleSet` carries a fifth field,
|
||
`wildcard_bytes`, backing both pattern lists. The snapshot builds every sub-structure from its own
|
||
arena, so `Snapshot.deinit` is one `arena.deinit()`. **S8 must fill `group_ids`, `source_ids` and
|
||
the optional `compiled` entries accordingly.**
|
||
|
||
**S8 manager + sources_repo.** The compile stage writes plain `<id>.list.tmp` / `<id>.wild.tmp`;
|
||
`publish` streams header + body into the final files via `createFileAtomic` + `sync` + `replace` at
|
||
0o600 — the header carries counts that exist only after the compile, so it cannot go into the same
|
||
atomic pass. Group row ids come from `groups_repo.groupId(database, name)` (added by the
|
||
orchestrator; a vanished name is `error.GroupSetChanged`) — the manager does not run raw SQL. A
|
||
source whose compiled files are missing, unreadable or checksum-mismatched is marked `.load_failed`
|
||
and excluded from the snapshot; the reload succeeds without it. A present, checksum-clean but
|
||
malformed body still fails the build and the previous snapshot keeps serving. `updateSourceStats`
|
||
takes the whole stats struct; the unchanged path re-writes the row and verifies the on-disk bodies
|
||
hash to the stored checksum before skipping (a corrupt file takes the rewrite path, which repairs
|
||
it). `SourceStatus.state` is a named `pub const State` with a `no_valid_entries` member: a download
|
||
that compiles to zero domains and zero wildcards with any nonzero invalid/unsupported/long-line
|
||
count fails the refresh, leaves the previous files serving, and records
|
||
`NoValidEntries invalid=N unsupported=N long_lines=N`. A failed refresh seeds its status from the
|
||
prior entry, so `last_success` and the counts of the still-serving files survive. A second lock,
|
||
`writer_lock` (plain `std.Io.Mutex`, separate from the RCU lock so readers never wait behind a
|
||
download), serializes `reload`, `refreshSource`, `refreshAll`, `pruneOrphans` and the startup pass.
|
||
Format sniffing collects `parsers.sample_lines` countable lines (never a flat byte window).
|
||
`SourceStatus` is a value type that borrows nothing: `url` is inline `[max_url_len]u8` + `url_len`
|
||
(`max_url_len = 255`, truncation in the status only; the full url lives on the source row), read via
|
||
`urlText()`; a copy made under the lock outlives the table it came from. `SourceStatus` also splits
|
||
the refresh fact from the load fact: `state` records the latest attempt to produce the files,
|
||
`loaded: bool` records whether they are filtering right now, and a load outcome never overwrites a
|
||
refresh failure (`State.isRefreshFailure` names the three refresh states) — after `refreshAll`, a
|
||
source can honestly read `.fetch_failed` with `loaded = true`. Refresh tmp cleanup defers are
|
||
installed before the calls that create the files, so cancellation and OutOfMemory paths leak no
|
||
`.tmp`. A reload collects per-source `LoadOutcome`s while it builds the candidate snapshot and
|
||
applies them to the status table (`applyLoadOutcomes`) inside the same exclusive-lock section as the
|
||
swap — a reload that fails before publishing leaves both the snapshot and the status table
|
||
describing the previous generation, and the table and snapshot change atomically for readers. A
|
||
disabled source is recorded as not loaded; its `state` keeps the record of how it last stood. The
|
||
reload builds its status table off to the side (`StatusTable`, `buildStatusTable`, `mergeStatuses`)
|
||
and installs it in the same critical section that swaps the snapshot, so a reload that fails
|
||
publishes neither. `refreshAll` still syncs the table up front — a refresh pass needs somewhere to
|
||
record per-source outcomes as it goes.
|
||
`Manager.Error` adds `Canceled`, `FileSystem`,
|
||
`GroupSetChanged`. The snapshot seed comes from `io.random` (`std.crypto.random` does not exist in
|
||
0.16). The daily interval sleeps on `.clock = .boot`; S7's short timeouts stay on `.awake`.
|
||
`sources_repo` additions: `SourceRow`, `SourceStats`, `listSourceRows`, `freeSourceRows`,
|
||
`updateSourceStats` — additive only.
|
||
|
||
**S9 note from S8:** a loopback `std.http.Server` answering `request.respond(body, .{})` deadlocks a
|
||
second fetch — the fetcher keeps the connection alive while a one-accept server waits for a new one.
|
||
Respond with `.{ .keep_alive = false }` or serve a keep-alive loop on the stream.
|
||
|
||
**S9 integration + fuzz.** Case 3 splits into the two real failure modes: a checksum-mismatched file
|
||
marks the source `.load_failed` and the reload succeeds without it (generation advances); only a
|
||
checksum-clean but malformed body fails `reload` and leaves the previous snapshot serving. Case 7
|
||
proves the cap from the response head: an explicit `content-length` of 100 MiB returns
|
||
`error.BodyTooLarge` before any body streams. Case 8 back-dates `last_updated` through
|
||
`updateSourceStats` and asserts inode and mtime of the compiled file are unchanged. Case 15 asserts
|
||
`budget/2 <= elapsed < 2*budget` (the POSIX backend wakes ~0.8 ms early on a 200 ms deadline). The
|
||
fuzz corpus is inline — `tests/fuzz/corpus.zig` imports the `dns` module, which the blocklist-fuzz
|
||
module does not have. Fuzz targets assert properties, not only absence of crashes: `Line.text`
|
||
windows the input (pointer containment), `covers_apex` only on `.wildcard`, `detectFormat`'s answer
|
||
survives `parseLine` over the same bytes, and `matches` is exercised on rejected patterns.
|
||
|
||
**Final wiring.** `build.zig` gained the `blocklist-fuzz` artifact (module import `parsers` →
|
||
`src/filter/parsers.zig`, LLVM backend under `-Dfuzz`), hung off `test_step` beside the dns fuzz
|
||
artifact. Evaluation: `zig build test`, `zig build test -Dintegration` and `zig build cross` all
|
||
exit 0; both cross executables are statically linked.
|