diff --git a/CHANGELOG.md b/CHANGELOG.md index b18b9e1..c8a96a5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,18 @@ All notable changes to nxdns are recorded here. The format follows [Keep a Chang Sections are written by hand. Nothing here is generated from commit messages: the point of the file is to say what changed for an operator, which a commit subject rarely does. +## [0.0.14] - 2026-08-28 + +Schema changes stop costing you your query history. querylog.db is now version-stamped and migrated in place; the server refuses to start rather than ever reset a healthy file, and the release tooling refuses to ship a schema change that is neither migratable nor explicitly disclosed with recovery steps. Three releases (0.0.6, 0.0.9, 0.0.12) each discarded the log on upgrade; this ends that. + +### Changed + +- **querylog.db is migrated in place.** The file now carries a schema version, and a release that changes the schema ships a migration that runs at startup: one consistent backup (`querylog.db.pre-migrate-`, mode 0600, only the most recent kept), then every step and the version stamp in a single transaction. A failure before the commit rolls back and leaves your file exactly as it was. +- **The server refuses instead of resetting.** A querylog.db it cannot use — newer than the binary, older than 0.0.12, or mid-migration failure — is left untouched and the server exits with a clear message instead of setting the file aside and starting an empty log. The exit code (2) tells systemd not to restart-loop a deliberate refusal. Corruption is the only case that still sets a file aside automatically. +- **The release gate now enforces the contract.** A schema change cannot be tagged unless it either ships a working migration (proven in CI against a frozen fixture of the previous schema, with shipped migration files locked byte-for-byte once released) or explicitly declares a break — which requires a version bump the server refuses on, a reset disclosure, and step-by-step restore instructions in this file. + +**One hazard to know when downgrading.** The first start under this release restamps querylog.db from the old fingerprint to version 1 (contents untouched). If you later downgrade to 0.0.13 or older, that binary treats the new stamp as a schema mismatch, moves your file aside as `querylog.db.schema-changed-`, and starts an empty log. To recover: return to 0.0.14 or newer, stop the server, move the empty `querylog.db` away and delete its `querylog.db-wal` and `querylog.db-shm` files (leaving them would corrupt the restored file), rename the `.schema-changed-` file back to `querylog.db`, and start. + ## [0.0.13] - 2026-08-27 The upstream query budget becomes one honest deadline. A busy network no longer blames a healthy standby for running out of time, and a query burst no longer queues invisibly until everything answers SERVFAIL at once. diff --git a/PLAN.md b/PLAN.md index cfadb37..b83a346 100644 --- a/PLAN.md +++ b/PLAN.md @@ -85,13 +85,13 @@ Verified: 0.16.0 ships `std.crypto.tls.Client` only. There is no server-side TLS Two SQLite files with opposite write profiles, isolated from each other: - **`config.db`** — small, precious, rarely written: groups, clients, prefixes, upstreams, blocklist source metadata, rules, local records, forward zones, settings, schema version. -- **`querylog.db`** — high-churn, large, expendable: query log + its own private `domains` dimension table. Client identity stored as **IP text**, not a FK into config — log rows are immutable facts and must not point at mutable config rows. If `querylog.db` is missing or corrupt at startup, rename aside, recreate, keep serving. Log loss is not an outage. +- **`querylog.db`** — high-churn, large, expendable: query log + its own private `domains` dimension table. Client identity stored as **IP text**, not a FK into config — log rows are immutable facts and must not point at mutable config rows. If `querylog.db` is missing or corrupt at startup, rename aside, recreate, keep serving — corruption only; a healthy file whose schema this build cannot use refuses the startup instead (§3.7). - No cross-DB references. Retention/VACUUM churn never touches `config.db`; config backup is a copy of a tiny file. ### 3.7 Upgrades: Auto-Migration (Decision J) - `config.db`: numbered, sequential SQL migration steps compiled into the binary. At startup: read schema version row, apply newer steps inside a transaction, continue. Operator upgrade = install binary, restart. Before v0.1 the list holds one step — the baseline of §11.2, edited in place — because nxdns has no installs and a step exists only to reconcile a database somebody already has. -- `querylog.db`: **no migrations.** On schema mismatch: rename aside, recreate fresh. +- `querylog.db`: a logical version in `PRAGMA user_version`, migrated **in place** at startup by the same shape of compiled step list, inside one transaction and behind one `querylog.db.pre-migrate-` backup (only the newest is kept). A healthy file is never renamed aside: a version this build cannot reach refuses the startup with instructions, and only corruption recreates. A deliberate break is still allowed, but it must be versioned, refused at startup, and disclosed in the changelog — the cut gate enforces that. See `docs/reference/query-log-lifecycle.md`. ### 3.8 Blocklist Storage (Decision A) @@ -694,5 +694,5 @@ The project publishes released binaries and container images from its own Gitea | G | SQLite vendored amalgamation + own thin wrapper | | H | Two DBs: `config.db` (precious) + `querylog.db` (expendable, self-contained, client IP as text) | | I | Frontend embedded in binary; dev flag serves from disk; static musl release builds | -| J | Auto-migration for `config.db` at startup; `querylog.db` recreated on mismatch | +| J | Auto-migration at startup for both databases; `querylog.db` recreated only when corrupt | | — | Safe-search per-group; Prometheus `/metrics` in scope; CI on self-hosted Gitea Actions | diff --git a/docs/README.md b/docs/README.md index 6a85a90..4c68197 100644 --- a/docs/README.md +++ b/docs/README.md @@ -37,6 +37,7 @@ Descriptions of what is there. No procedures, no advice. - [reference/api.md](reference/api.md) — every REST route, authentication and the event stream. - [reference/cli.md](reference/cli.md) — the six subcommands, every flag, every exit code. - [reference/files-and-directories.md](reference/files-and-directories.md) — the data directory layout and file modes. +- [reference/query-log-lifecycle.md](reference/query-log-lifecycle.md) — how `querylog.db` is versioned, migrated, backed up and, rarely, recreated. - [reference/performance.md](reference/performance.md) — the targets and the measured numbers. ## Explanation diff --git a/docs/explanation/architecture.md b/docs/explanation/architecture.md index 892c3a3..4104140 100644 --- a/docs/explanation/architecture.md +++ b/docs/explanation/architecture.md @@ -27,7 +27,7 @@ Directories: | `src/cache/` | `dns_cache.zig`: bounded in-memory TTL cache of whole response messages, keyed by the question. The clock arrives as a parameter. | | `src/upstream/` | Upstream resolution: shared vocabulary and the `Client` interface (`transport.zig`), DoH client (RFC 8484), DoT client (RFC 7858), per-endpoint health and backoff (`health.zig`), and `pool.zig` — priority-ordered failover that is itself a `transport.Client`, so the handler sees one interface. | | `src/server/` | The serving side: UDP/TCP/DoH/DoT listeners, `handler.zig` (the whole query pipeline), `cert_store.zig` (refcounted TLS cert holder), `rate_limiter.zig`, `pause.zig`, `clients.zig` (client auto-materialisation), `local_tables.zig` (published local-answer tables), `query_sink.zig` (log and SSE fanout), `shutdown.zig` (SIGINT/SIGTERM into one `std.Io.Event`). | -| `src/storage/` | SQLite ownership: `db.zig` is the only file that calls SQLite, `config_schema.zig` + `migrations.zig` for `config.db`, `querylog_schema.zig` (open-or-recreate), async query `logger.zig`, `retention.zig`, `disk_monitor.zig`, and one repository per table under `repositories/`. | +| `src/storage/` | SQLite ownership: `db.zig` is the only file that calls SQLite, `config_schema.zig` + `migrations.zig` for `config.db`, `querylog_schema.zig` + `querylog_versions.zig` + `querylog_migrations.zig` for `querylog.db`, async query `logger.zig`, `retention.zig`, `disk_monitor.zig`, and one repository per table under `repositories/`. | | `src/config/` | The one configuration model (`model.zig`), the pure validator (`validate.zig`), `import.zig`/`export.zig` (ZON to and from `config.db`, byte-stable round trip), `loader.zig` (read/parse/validate a named file, with the shared fault mapping), `reconcile.zig` (converge the database onto a parsed config by row identity). | | `src/web/` | The admin HTTP layer: `server.zig` (listener), `router.zig`/`routes.zig`, one file per resource under `handlers/`, `auth.zig` (sessions), `sse.zig` (live query fanout), `static.zig` (embedded SPA), `metrics.zig` (Prometheus), `openapi.zig` (served contract), `api_limiter.zig`, `http_util.zig`. | | `src/platform/` | OS and TLS edges: IP address values, the `std.log` sink (`logging.zig`), `statfs.zig` (free-space query via libc), client TLS over `std.crypto.tls` (`tls_client.zig`), server TLS over vendored Mbed TLS (`tls_server.zig`). | @@ -115,7 +115,7 @@ Two databases with opposite contracts, in one data directory (see [reference/fil Which of the file and the database is *authoritative* is chosen by the invocation, not by state: bare `nxdns run` serves the database, and `nxdns run --config FILE` makes the file authoritative and reconciles the database onto it at every start. `reconcile.zig` is that convergence, matching rows by identity and writing only differences, so runtime state — blocklist checksums, compiled snapshots, client history — survives. Why it works that way is [configuration-model.md](configuration-model.md). -**`querylog.db` is expendable.** It is never migrated. Its schema carries a fingerprint derived from the DDL text, and at open, a missing, corrupt, non-database, `quick_check`-failing or fingerprint-mismatched file is moved aside and recreated empty — the old file is kept under a new name rather than deleted, so an operator can still look at it. Retention deletes old rows daily and periodically rewrites the file to reclaim space. +**`querylog.db` is the expendable one, but its history is not thrown away.** It carries a logical version in `PRAGMA user_version`, and an older supported version is migrated in place at open: one transaction, behind one `querylog.db.pre-migrate-` backup, of which only the newest is kept. Only real damage recreates the file — missing, corrupt, not a database, or failing `quick_check` — and then the old file is kept under a new name rather than deleted, so an operator can still look at it. A healthy file this build cannot read is neither migrated nor moved aside: the startup refuses and says why, because losing months of history to a rollback is worse than a server that will not start. Retention deletes old rows daily and periodically rewrites the file to reclaim space. The whole contract is [reference/query-log-lifecycle.md](../reference/query-log-lifecycle.md). The split exists so that the churn of the second database can never endanger the first. Query logs are high-volume, disposable, and the thing most likely to be corrupted by a power cut on an SD card; configuration is small, irreplaceable, and the thing an operator would have to reconstruct by hand. Giving them one file would force the careful contract onto the noisy data or the loose contract onto the valuable data. diff --git a/docs/how-to/troubleshoot.md b/docs/how-to/troubleshoot.md index 1abefeb..16d90d3 100644 --- a/docs/how-to/troubleshoot.md +++ b/docs/how-to/troubleshoot.md @@ -339,3 +339,37 @@ nxdns run failed: SchemaTooNew ``` **Fix.** There is no downgrade. Import the export you took before upgrading into a fresh data directory with the older binary; see [Upgrade nxdns](upgrade.md). + +## The server refuses to start over querylog.db + +**Symptom.** The process stops at startup naming the query log, and the error is one of four names: + +``` +error(querylog_schema): refusing to open querylog database '/var/lib/nxdns/querylog.db': it is stamped 7, and this build supports schema versions 1 to 1 (SchemaTooNew). The file is left exactly as it is; see docs/how-to/troubleshoot.md, "The server refuses to start over querylog.db" +nxdns run failed: SchemaTooNew +``` + +This is a refusal, not damage. nxdns will not replace a healthy query log to get itself started, so the file is left exactly as it was — schema, rows, coverage watermark and version stamp all unchanged — and the startup fails instead. All four exit 2, the code that means an operator has to act, because none of them resolves on a retry — the shipped systemd unit's `RestartPreventExitStatus=2 64` stops the unit on the first refusal instead of restart-looping it. `systemctl status nxdns` shows the refusal. `nxdns check` does not grade the query log at all, so it will not reproduce any of these. + +[The query-log lifecycle](../reference/query-log-lifecycle.md) is the full contract behind this page. + +**Fixes by name.** + +- `SchemaTooNew` — the file was stamped by a newer nxdns than the one you are running, which normally means a binary was rolled back. Put the newer release back and start it: the file is exactly as that release left it. If you mean to stay on the older release, that release cannot read this file, so restore the `querylog.db.pre-migrate-` copy the upgrade left beside it — stop the server, move `querylog.db` and its `querylog.db-wal` and `querylog.db-shm` out of the way, rename the backup to `querylog.db`, and start. Starting empty is also an option: with the server stopped, move `querylog.db` and both sidecars aside and the next start creates a fresh log. +- `SchemaUnsupported` — the stamp is not a version this build can reach. Either the file predates 0.0.12, or it came from somewhere else, or a release since deliberately broke the schema; the changelog section for the release you are running says so when it is the third case. There is no migration path, by contract. Keep the file if the history matters — copy it somewhere and read it with the `sqlite3` shell — and if starting with an empty log is acceptable, stop the server, move `querylog.db`, `querylog.db-wal` and `querylog.db-shm` out of the data directory by hand, and start again. +- `MigrationBackupFailed` — a migration was due and the pre-migration backup could not be written, so nothing was migrated. The line above names the destination and the reason, which is almost always a full or read-only data directory. Free space or fix the permissions and start again. +- `MigrationFailed` — read the log line above it, because two different states wear this one name. + +**Which `MigrationFailed` you have.** The distinction is in the line the migration logged, and it decides whether you do anything at all: + +- Before the commit: `querylog migration 1 -> 2 failed before commit (...); the database is unchanged`. Nothing was applied. The file still carries its old version and every row, and this run's backup was deleted because the original is intact. Restarting will attempt the same migration and fail the same way, so this needs the underlying cause — the log line names it — or a report. +- After the commit: `querylog migration 1 -> 2 COMMITTED and the database IS at version 2, but the connection could not be restored: ...; the backup '...' is kept and the next start will open the migrated file normally`. The migration DID complete. Only that one startup is refused, the pre-migration backup is kept, and the next start opens the migrated file on the ordinary current-version path. Start the server again. + +In neither case does the server start with an empty log on its own. Recreating a query log automatically is reserved for real corruption; see [why a query log is moved aside](../reference/files-and-directories.md#why-a-query-log-is-moved-aside). + +> Not reproduced against a running service: the four refusals are covered by the +> test suite rather than by a hand-driven install, and the released chain has no +> migration step in it yet, so no upgrade produces a `pre-migrate` backup today. +> The messages above are the ones `src/storage/querylog_schema.zig` and +> `src/storage/querylog_migrations.zig` emit, with a data directory path and +> example version numbers filled in. diff --git a/docs/how-to/upgrade.md b/docs/how-to/upgrade.md index 12336dc..9ba627b 100644 --- a/docs/how-to/upgrade.md +++ b/docs/how-to/upgrade.md @@ -263,7 +263,7 @@ nxdns run failed: SchemaTooNew > hand and `nxdns run` was pointed at it. The two lines above are that run's > output. -That run exits 1. Recovering means importing the export you took in step 1 into a fresh data directory with the older binary. +That run exits 2, and the shipped unit stops rather than restart-loops it. Recovering means importing the export you took in step 1 into a fresh data directory with the older binary. ### Rolling back from file mode diff --git a/docs/reference/cli.md b/docs/reference/cli.md index abf9945..9ff9b4d 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -214,7 +214,7 @@ Prints the usage text to stdout and exits 0. `nxdns --help` and `nxdns -h` do th | --- | --- | | 0 | Success. | | 1 | Runtime failure — I/O, database, out of memory. A partial diagnostic report caused by an allocation failure is a runtime failure, not a verdict on the configuration. | -| 2 | A configuration problem the operator can fix, or a `check` that found one. | +| 2 | A configuration problem the operator can fix, a `check` that found one, or a deliberate refusal to run that no retry will clear. | | 64 | Usage error — unknown command or flag, a flag without its value, a missing or extra argument. | Code 2 means the same thing from every subcommand. `src/config/faults.zig` holds the one list of errors that mean "the configuration the operator supplied is wrong", and `run`, `check` and `import` all ask it, so a rejected file exits 2 whichever command read it. The list is every error the validator raises, plus `ParseZon`, `ConfigTooLarge`, `NoUsableUpstreams`, `BadCertificate` and `ManagedConfigUnreadable`. In practice that covers a file with a syntax error, one larger than 4 MiB, one with no `default` group (`MissingDefaultGroup`), one with no enabled upstream (`NoUpstreams`), a bad bind address, a bad rate limit, an unusable certificate, `password` and `password_hash` set together, and a `--config` path that is absent or unreadable. @@ -237,6 +237,8 @@ load one with `nxdns import `, or make a file the source of truth with `nx `import` exits 2 for those faults and for `DestructiveImport`. That last one is deliberately not a configuration fault — it reports what applying the file would delete, rather than anything wrong with its content — and `import` decides it for itself; the answer to it is `--allow-delete`, not an edit. +`run` also exits 2 on the four query-log schema refusals — `SchemaTooNew`, `SchemaUnsupported`, `MigrationFailed` and `MigrationBackupFailed` — which are in the same list. They are not a verdict on a file the operator wrote, but they share the property exit 2 exists to signal: the server is refusing on purpose, an operator has to act, and a restart will only repeat the refusal. Exit 1 would put them under the unit's `Restart=on-failure` and loop them. See [the server refuses to start over querylog.db](../how-to/troubleshoot.md#the-server-refuses-to-start-over-querylogdb). + `OutOfMemory` is exit 1 even when problems were recorded, because the report is then incomplete. Every other error is 1. Where an exit code sends you next: [troubleshoot](../how-to/troubleshoot.md). diff --git a/docs/reference/files-and-directories.md b/docs/reference/files-and-directories.md index 8ecd009..ae394dc 100644 --- a/docs/reference/files-and-directories.md +++ b/docs/reference/files-and-directories.md @@ -16,10 +16,12 @@ Default `/var/lib/nxdns`, overridable with `--data-dir DIR`. `nxdns run` and `nx | --- | --- | --- | | `config.db` | The configuration database, including `web.password_hash`. The source of truth in database mode; in file mode it is the runtime substrate the file is reconciled onto (see [the configuration file](#the-configuration-file)). | 0600 | | `config.db-wal`, `config.db-shm` | SQLite write-ahead log and shared-memory index for `config.db`. Created by `run`, `import` and `export` when WAL is enabled, inheriting the main file's permissions. `check` creates neither. | 0600 | -| `querylog.db` | The query log: every domain every client asked for. Expendable — if it is missing or unusable it is recreated empty. | 0600 | +| `querylog.db` | The query log: every domain every client asked for. Missing or damaged, it is recreated empty; an older schema is migrated in place, and a schema this build cannot use refuses the startup. See [the query-log lifecycle](query-log-lifecycle.md). | 0600 | | `querylog.db-wal`, `querylog.db-shm` | WAL sidecars for `querylog.db`. | 0600 | -| `querylog.db.-` | A `querylog.db` this build could not use, moved aside before an empty one was created in its place. Kept, never overwritten. `` is one of `corrupt`, `not-a-database`, `quick-check-failed` or `schema-changed`; see [why a query log is moved aside](#why-a-query-log-is-moved-aside). | Whatever the renamed file had — no chmod reaches it | +| `querylog.db.-` | A damaged `querylog.db`, moved aside before an empty one was created in its place. Kept, never overwritten. `` is one of `corrupt`, `not-a-database` or `quick-check-failed`; see [why a query log is moved aside](#why-a-query-log-is-moved-aside). | Whatever the renamed file had — no chmod reaches it | | `querylog.db.--` | The same, when the plain name is taken — `` counts from 1 and rises until the name is free. Two recreates within one second is the case it exists for. | The same | +| `querylog.db.pre-migrate-` | A complete copy of `querylog.db` taken immediately before a schema migration. Exactly one survives: a successful migration deletes every other one, and a later successful start retries that cleanup. Written by `VACUUM INTO`, so it holds the committed database including anything still only in the write-ahead log, and it needs no sidecars of its own. | 0600 | +| `querylog.db.pre-migrate--` | The same, when the plain name is taken — `` counts from 2. | The same | | `blocklists/` | Compiled blocklist snapshots, one subdirectory of the data directory. | 0700 | | `blocklists/.list` | Exact domains for blocklist source ``, one per line, behind a header. | 0600 | | `blocklists/.wild` | Wildcard entries for the same source. | 0600 | @@ -52,14 +54,15 @@ The temporaries of a source that still exists are cleaned by the refresh that ow ### Why a query log is moved aside -A `querylog.db` is moved aside when it is missing nothing but usability, and the name it is given says which of the four cases it hit: +A `querylog.db` is moved aside only when it is genuinely damaged, and the name it is given says which of the three cases it hit: | `` | What happened | | --- | --- | | `corrupt` | SQLite reported the file as damaged. | | `not-a-database` | The file is not a SQLite database at all. | | `quick-check-failed` | `PRAGMA quick_check` did not answer `ok`. | -| `schema-changed` | Nothing is wrong with the file. Its `user_version` fingerprint does not match this build's schema, so this build cannot read it. Upgrades that touch the query-log schema produce this one, and the file they set aside is a healthy database. | + +There is no fourth case. A healthy file carrying a schema version this build cannot use is neither migrated nor renamed: the startup refuses and the file stays where it is, which is [the query-log lifecycle](query-log-lifecycle.md). You can still find a `querylog.db.schema-changed-` in a data directory, because 0.0.13 and older produced one on any schema change — and still do, if you downgrade to one of them. This build never writes that name. Only the main file is renamed — its `-wal` and `-shm` are deleted, because a stale WAL would be replayed into the fresh database. A missing `querylog.db` is created without any aside file. The rename happens inside `querylog_schema.open`, before the 0600 chmod, and that chmod names `querylog.db` and its two sidecars only — so an aside file keeps the mode the file had at rename time, which for a `querylog.db` nxdns itself created is 0600 and for one an operator put there is whatever they left it at. Nothing prunes the aside files; they accumulate until an operator removes them, and each one holds the same browsing history the live query log holds. @@ -67,6 +70,10 @@ Only the main file is renamed — its `-wal` and `-shm` are deleted, because a s The 0600 modes are not cosmetic. `config.db` holds the argon2id password hash and `querylog.db` holds the browsing history of every client on the LAN, so both are as sensitive as each other, and a WAL file holds the same rows as the database it belongs to. SQLite creates the main database at `0644 & ~umask`; nxdns chmods it to 0600 before enabling WAL, so the sidecars inherit 0600 rather than being created world-readable. +A `pre-migrate` backup holds that same browsing history, so it is chmodded 0600 the way the live file is: SQLite's `VACUUM INTO` creates it at `0644 & ~umask` and nxdns restricts it immediately afterwards. A backup it cannot restrict is a failed backup — the partial file is deleted and the startup refuses with `MigrationBackupFailed`, rather than leaving a world-readable copy behind. + +The aside files are the exception: they get no chmod at all, and an aside keeps whatever mode it had at rename time. For a `querylog.db` nxdns itself created that is 0600; for one an operator put there it is whatever they left it at. + ## The configuration file There is no default path. `--config FILE` names the file, and without that flag no file is read at all — a `config.zon` sitting in `/etc/nxdns` that no invocation names is inert. `/etc/nxdns/config.zon` is a convention the packaging follows, not a location nxdns probes. diff --git a/docs/reference/query-log-lifecycle.md b/docs/reference/query-log-lifecycle.md new file mode 100644 index 0000000..4944988 --- /dev/null +++ b/docs/reference/query-log-lifecycle.md @@ -0,0 +1,70 @@ +# The query log's lifecycle + +What happens to `querylog.db` when nxdns opens it: how the file is versioned, when it is migrated, when the server refuses to start over it, and the one case in which it is still replaced. Source of truth: `src/storage/querylog_versions.zig` (the version metadata), `src/storage/querylog_schema.zig` (the open path) and `src/storage/querylog_migrations.zig` (the migration runner). + +The rule this page exists to state: **a healthy `querylog.db` is never replaced and never moved aside.** A schema this build cannot use refuses the startup instead. Your query history is not the server's to discard. + +## The version stamp + +Every `querylog.db` carries a logical schema version in SQLite's `PRAGMA user_version`. It is a small counter — 1 in this release — and not a hash of anything. A file created by this build is stamped as it is created. + +Two other values matter, both in `querylog_versions.zig`: + +| Constant | Today | What it means | +| --- | --- | --- | +| `current_version` | 1 | The version this build creates and reads. | +| `minimum_supported_version` | 1 | The oldest stamped version this build can migrate up to `current_version`. | +| `legacy_fingerprint` | 1975011655 | The `user_version` the 0.0.12 and 0.0.13 binaries wrote: a CRC32 of their schema text, under the older policy where a mismatch meant "replace the file". | + +`legacy_fingerprint` is frozen forever. Those two releases stamped a hash rather than a version, so this build recognises that one literal number as "version 1" and restamps the file as 1 on the first open. The restamp runs in its own transaction; if it fails, the old stamp and every row stay exactly as they were and the startup refuses. + +## What an open does + +nxdns opens `querylog.db` once at startup, before it serves anything, and no second process shares a data directory. On a file that is readable and passes `PRAGMA quick_check`, the stamp decides: + +| Stamp | What happens | +| --- | --- | +| `current_version` | Opens. Nothing is migrated. | +| `legacy_fingerprint` | Read as version 1: restamped to 1, then treated as version 1 by the rows above and below. | +| Between `minimum_supported_version` and `current_version` | Migrated in place, then opens. | +| Above `current_version`, up to 1000000 | REFUSE: `SchemaTooNew`. | +| Anything else — 0, a negative, another fingerprint, a version below the minimum | REFUSE: `SchemaUnsupported`. | + +A refusal changes nothing. The schema, the rows, the coverage watermark and the stamp are all left as they are, no file is set aside, no new file is created, and `nxdns run` exits. The log line names the path, the stamp it found, the range this build supports and [the troubleshooting section](../how-to/troubleshoot.md#the-server-refuses-to-start-over-querylogdb). + +## Migrating in place + +A migration is one backup and one transaction. + +1. **Back up.** `VACUUM INTO` writes a complete copy — including anything still only in the write-ahead log — to `querylog.db.pre-migrate-` beside the database. If that name is taken, `-2`, `-3` and so on are tried. A backup that cannot be written is `MigrationBackupFailed`, and the partial copy is deleted; an older backup beside it survives. +2. **Migrate.** `BEGIN IMMEDIATE`, re-read the stamp under the lock, run every step, run `PRAGMA foreign_key_check`, stamp the new version, `COMMIT`. One transaction covers the whole chain, so the file is either at the old version or at the new one and never in between. +3. **Clean up.** Every other `querylog.db.pre-migrate-*` beside the file is deleted. **One backup is kept**: the one this migration just took. A later successful start retries that cleanup if it failed. + +If a step fails before the commit, the transaction rolls back, this run's backup is deleted, and the startup refuses with `MigrationFailed`. The database keeps the version and the rows it had. + +If the commit succeeds and something after it fails, the log says so plainly — the migration DID complete and the file IS at the new version. The backup is kept, the startup still refuses with `MigrationFailed`, and the next start opens the migrated file normally. + +## Corruption is the only automatic recreate + +Four conditions still create a fresh, empty `querylog.db`: the file is missing, SQLite reports it as corrupt, it is not a SQLite database at all, or `PRAGMA quick_check` does not answer `ok`. Except for the missing case, the unusable file is renamed to `querylog.db.-` and kept. See [why a query log is moved aside](files-and-directories.md#why-a-query-log-is-moved-aside). + +Every other failure — a lock held elsewhere, a permission problem, a full disk, a version this build cannot reach — propagates and leaves the file alone. + +## Downgrading + +**Downgrading to 0.0.13 or older resets your query log.** Those binaries predate this contract: they compare `user_version` against a hash of their own schema text, find this build's version stamp instead, and treat that as a mismatch — so they rename `querylog.db` to `querylog.db.schema-changed-` and start an empty log. Nothing is destroyed, but the live log is empty until you put the aside file back, and the restamp that provoked it takes no backup of its own. + +To recover, go back to a migration-aware release, stop the server, then, in the data directory: + +1. Move the empty `querylog.db` the old binary created out of the way. +2. Delete its `querylog.db-wal` and `querylog.db-shm`. This is not optional: replaying the empty file's write-ahead log into the restored history would corrupt it. +3. Rename `querylog.db.schema-changed-` back to `querylog.db`. +4. Start the server. + +Downgrading between two migration-aware releases is safe in the sense that matters: a build that finds a stamp above its own `current_version` refuses to start with `SchemaTooNew` and touches nothing. Go forward again, or restore the `pre-migrate` backup the upgrade left. + +## Breaking the schema on purpose + +A release may still break the query-log schema outright rather than migrate it. That is allowed, and it is never silent. Such a release raises `current_version`, sets `minimum_supported_version` to the same value, and ships no migration step — so files from before the break classify as below the minimum and `open` refuses them with `SchemaUnsupported` rather than replacing them. The release notes carry the phrase `resets your query history` and a `Restoring your query history` section, and the cut gate refuses to build the release without both. + +So the contract is: a break is always versioned, always refused at startup with the file intact, and always disclosed in the changelog. diff --git a/specs/milestone-38.md b/specs/milestone-38.md new file mode 100644 index 0000000..a00c220 --- /dev/null +++ b/specs/milestone-38.md @@ -0,0 +1,192 @@ +# Milestone 38: querylog schema migrations + +Stop the recurring query-history loss: schema changes migrate querylog.db in place; the automatic reset survives only for real corruption; explicit breaks stay possible but must be versioned, refused by `open`, and ship recovery instructions. + +Owner rulings (2026-08-28): baseline is the 0.0.12/0.0.13 schema — nothing older is migratable; breaking changes remain allowed but must be explicit with clear changelog instructions; keep only the most recent pre-migration backup. + +## Sessions + +A (storage framework) first. B (cut gate) needs A's modules. C (docs) after A (documents A's behavior; shares no files with B). The orchestrator writes the changelog. + +--- + +## Session A: migration framework in storage + +### A.1 Version metadata module (pure, no SQLite) + +New file `src/storage/querylog_versions.zig` — importable by `tools/cut.zig` without linking SQLite. ONLY comptime data: + +- `pub const current_version: i32 = 1;` +- `pub const minimum_supported_version: i32 = 1;` — files stamped below this refuse. An EXPLICIT BREAK in a future release is expressed here: bump `current_version`, set `minimum_supported_version = current_version`, ship no step. The chain then cannot reach the new version from below the minimum and `open` refuses the old file — a break is always versioned, always refused at runtime, never silent. +- `pub const legacy_fingerprint: i32 = 1975011655;` — the literal `user_version` stamp the 0.0.12/0.0.13 binaries wrote (CRC32 of their DDL text). FROZEN literal, derived from nothing; comment cites v0.0.12. +- `pub const version_floor_guard: i32 = 1_000_000;` +- Comptime asserts: `minimum_supported_version >= 1`; `minimum_supported_version <= current_version`; `current_version <= version_floor_guard`; `legacy_fingerprint` outside `[0, version_floor_guard]`; `step_sql.len == current_version - minimum_supported_version`. +- `pub const step_sql: []const [:0]const u8 = &.{};` — step i migrates version `minimum_supported_version + i` to `+ i + 1`; each entry is `@embedFile("migrations/v.sql")`. EMPTY this milestone. +- **Steps are SQL-only. There are no migration hooks.** A rebuild that m36-style projections would need is expressible as plain SQL (the recompute statements are SQL); a future change that truly cannot be SQL must amend this design explicitly in its own spec. This keeps every shipped migration byte-comparable (B.2 Gate 2) with no mutable code path. +- Shipped step files `src/storage/migrations/v.sql` and fixtures (B.1) are immutable once released; the cut gate byte-compares them against the previous tag. + +### A.2 Runner module and the rebuild rule + +New file `src/storage/querylog_migrations.zig` (SQLite side): the runner and the equivalence oracle. + +- `pub fn migrateSteps(database: *db.Db, sql: []const [:0]const u8, from: i32, target: i32) (db.Error || error{TransactionViolation})!void` — runs the steps and the final `PRAGMA user_version = target` stamp inside the caller's already-open transaction. SLICING CONTRACT: `sql` is exactly the `[from, target)` suffix — `sql[0]` migrates `from -> from + 1`; asserted: `sql.len == @intCast(target - from)`. Production callers slice `step_sql[from - minimum_supported_version ..]`. While steps execute, the runner installs SQLite's authorizer (`sqlite3_set_authorizer`; expose a scoped install/clear pair on the db wrapper) denying `SQLITE_TRANSACTION` — a step cannot BEGIN/COMMIT/ROLLBACK at all, which is the only reliable guard (a step containing `COMMIT; BEGIN IMMEDIATE;` would pass a post-step autocommit check while breaking atomicity; that exact bypass is a required negative test, and the test must also assert the authorizer is cleared after the rejection: the rollback succeeds and the SAME connection can then execute transaction statements normally — a leaked authorizer would block cleanup and strand the connection inside the migration transaction). The authorizer is cleared on every exit path. Belt: the post-step `sqlite3_get_autocommit(db) == 0` check stays. `migrateSteps`'s error set is `(db.Error || error{TransactionViolation})`; `runMigration` maps `TransactionViolation` to `error.MigrationFailed`. The no-transaction-statements rule is also in the step-authoring doc comment. +- `pub fn runMigration(io: std.Io, dir: std.Io.Dir, path: [:0]const u8, database: *db.Db, sql: []const [:0]const u8, from: i32, target: i32) Error!void` — the full orchestration seam: backup (A.4 step 1), transaction + `migrateSteps` + commit (step 2), failure handling (step 3), retention (step 4). `open` calls it with production metadata; synthetic tests call it directly with test chains, so the REAL backup/collision/retention/error paths are what the tests prove. +- **Rebuild rule** (doc comment on `step_sql`): a step that changes a table's shape must produce a table whose stored CREATE text is byte-identical to the fresh DDL's. The RUNNER brackets every migration with: `PRAGMA foreign_keys = OFF` and `PRAGMA legacy_alter_table = ON` BEFORE `BEGIN IMMEDIATE` (with `foreign_keys` on — which `db.applyPragmas` enables — a rename of a referenced parent rewrites child tables' FK text to `_old`, corrupting them the moment the old table drops; `legacy_alter_table` alone does not prevent that), and restores both pragmas on EVERY exit path, success or failure (they are connection-global and non-transactional). Before COMMIT the runner runs `PRAGMA foreign_key_check` and fails the migration on any row. Step sequence: `ALTER TABLE RENAME TO _old`, `CREATE TABLE ...` pasted VERBATIM from the target `querylog_schema.ddl`, `INSERT INTO SELECT ...` mapping, `DROP TABLE _old`, recreate EVERY dependent object of `` verbatim from the target DDL — indexes AND triggers (both dropped with `_old`). Views are NOT dropped by the rename or the drop (with `legacy_alter_table` on they keep naming ``), so a step DROPs each view over `` FIRST and recreates it verbatim LAST — recreating without the drop fails with "view already exists". `ALTER TABLE ... ADD/RENAME COLUMN` on a kept table is forbidden — SQLite rewrites stored CREATE text under it and the oracle's text layer would rightly fail. + +### A.3 The open path (rework `querylog_schema.open`) + +`open` owns the file exclusively: nxdns opens querylog.db once at startup before serving, and no other process shares a data dir (existing deployment contract; restate in `open`'s doc comment — the backup-then-lock sequence relies on it). + +The version-handling half of `open` is factored as `openVersioned(io, dir, path, handle, plan) Error!void` where `handle: *?db.Db` is an optional SLOT: `openVersioned` closes and nulls it on every error path, so the caller's `errdefer` no-ops and single-close is structural rather than a convention (as built 2026-08-28; the post-commit test asserts `handle == null`). `plan: Plan = .{ .minimum: i32, .current: i32, .legacy_fingerprint: i32, .step_sql: []const [:0]const u8 }`. Production `open` passes the constant plan from `querylog_versions`; tests inject synthetic plans, which is what makes classification, migration, the post-commit mapping, and the sole-close ownership all testable through the REAL open path even while the production chain is empty. Classification itself stays a pure function of `(stamped, plan)`. + +Classify a healthy existing file: read `PRAGMA user_version` as `stamped`, map to a logical version FIRST, mutate NOTHING during classification: + +| condition | logical version | action | +| --- | --- | --- | +| `stamped == legacy_fingerprint` | 1 | classify version 1 by the rows below; if it lands on "current" or "supported older", first restamp to 1 (one transaction, A.5 error mapping), then proceed | +| `stamped == current_version` | stamped | open as today | +| `minimum_supported_version <= v < current_version` | v | migrate via `runMigration` | +| `current_version < v <= version_floor_guard` | v | REFUSE: `error.SchemaTooNew` | +| anything else (0, negatives, other fingerprints, below minimum) | — | REFUSE: `error.SchemaUnsupported` | + +The order matters: after a future explicit break raises the minimum above 1, a legacy-fingerprint file maps to version 1, classifies as below-minimum, and refuses WITHOUT the restamp — an unsupported file is never modified. + +REFUSE: the canonical file stays in place, logically untouched (schema, rows, watermark, stamp unchanged — WAL/SHM sidecar bytes may change from the probe; not a violation), nothing set aside, no new file, `open` errors, the server does not start. The log line names the path, the stamped value, the supported range, and `docs/how-to/troubleshoot.md` ("The server refuses to start over querylog.db"). + +Recreate lanes `missing`, `not_a_database`, `corrupt`, `quick_check_failed` unchanged. `RecreateReason.fingerprint_mismatch` and the `schema-changed` aside tag are DELETED. + +Fresh files: after executing `ddl`, stamp `PRAGMA user_version = current_version` (the stamp is already a separate statement; the DDL text does not change this milestone, so `querylog_schema.fingerprint` does not move). + +Backup retention has two passes with different authority. A migration's step 4 KNOWS the newest backup — this run's exact filename — and deletes every other `querylog.db.pre-migrate-*`; it is the primary mechanism. A plain successful open at current version runs a CONSERVATIVE retry for cleanups that once failed: parse `` and the optional `-N` collision suffix from each name, delete only files whose epoch is STRICTLY below the maximum, keep every file tied at the maximum epoch, and never delete a name that does not parse. This pass EXPLICITLY assumes forward-moving wall clock between migrations (record the assumption in its doc comment): under a clock rollback an older high-epoch name could outrank a genuinely newer backup, which is why the authoritative exact-name pass in step 4 is the primary mechanism and this pass is only the retry for its failures. + +### A.4 Running a migration (`runMigration`) + +1. **Backup.** `VACUUM INTO` on the live connection (no open transaction) to `querylog.db.pre-migrate-` in the database's directory. Destination must not pre-exist: on collision retry `--2`, `-3`, … The path enters the statement through an SQL string-literal quoting helper (double every `'`), never raw interpolation. On failure: delete the partial destination just created (only that file; an older valid backup survives), REFUSE with `error.MigrationBackupFailed`. +2. **One transaction.** `BEGIN IMMEDIATE`; re-read `user_version` under the lock. If it no longer equals `from`: ROLLBACK, delete this run's backup, REFUSE with `error.MigrationFailed` (exclusive ownership makes this outside interference). Otherwise `migrateSteps(db, sql, from, target)` — every step and the stamp in this one transaction — then COMMIT once. +3. **On PRE-COMMIT failure:** ROLLBACK, delete this run's backup, REFUSE with `error.MigrationFailed`, log the failing step index. Canonical file keeps its logical state. Never fall through to recreate. +3b. **On POST-COMMIT failure** (pragma restore or anything after a successful COMMIT): the file IS at `target` and that is said plainly in the log; the backup is KEPT (never deleted on this path). `runMigration` does NOT close the borrowed connection — it returns the distinct internal error `error.MigrationCommittedButUnclean`, and `querylog_schema.open`, which owns the handle and already has the sole error-path close, performs that one close and surfaces `error.MigrationFailed` to its caller. The next start takes the current-version lane cleanly. No post-commit path may claim the file unchanged or delete the backup. +4. **On success:** best-effort delete of every OTHER `querylog.db.pre-migrate-*` (keep this run's). Deletion errors warn and do not fail startup; A.3's every-open retention retries later. Log one line naming `from -> target` and the kept backup. + +### A.5 Legacy restamp error mapping + +The fingerprint→1 restamp is this milestone's only real mutation of operator data. It runs in one transaction; any failure (statement or commit) maps to `error.MigrationFailed`, rolls back, and leaves the legacy stamp and every row intact — REFUSE semantics, never recreate. Session A adds a test-only fault-injection seam to the db wrapper (`src/storage/db.zig`, following its existing `ReadTx.commit` injection style): one SQL-substring-matched one-shot seam on `Db.exec` covers statement and commit alike (both restamp statements pass through `Db.exec`), and the same seam drives the post-commit pragma-restore failure. Refusal paths log at `err`, which the test runner treats as failure, so `querylog_migrations.expected_failures` (begin/end/capturing, modelled on `db.read_tx_faults`) captures EXPECTED refusal logs per test; an unexpected refusal elsewhere still fails its test (as built 2026-08-28). Acceptance tests: the restamp forced to fail at (a) the statement and (b) the commit each leave `user_version == legacy_fingerprint` and the rows readable by a subsequent successful open. + +### A.6 Schema equivalence oracle + +`pub fn schemaEquivalent(gpa: std.mem.Allocator, a: *db.Db, b: *db.Db) (db.Error || std.mem.Allocator.Error)!bool` in `querylog_migrations.zig`. Two layers, both must agree: + +1. **Textual, exact:** for every non-`sqlite_` object in `sqlite_schema` (tables, indexes, views, triggers), compare `(type, name, tbl_name, sql)` with `sql` compared byte-for-byte. No normalization: the A.2 rebuild rule guarantees a migrated table carries the verbatim fresh CREATE text, and a fresh file trivially does. This layer sees CHECK constraints, foreign keys, WITHOUT ROWID, partial-index predicates, trigger/view bodies. +2. **Structural belt:** per table, `PRAGMA table_xinfo` rows and `pragma_table_list` `wr`/`strict` flags; per table, `PRAGMA foreign_key_list`; per index, `PRAGMA index_xinfo` plus `index_list` `unique`/`origin`/`partial` flags. + +Sort object and row lists before comparison. Negative tests: dropped `CHECK (rcode BETWEEN 0 AND 4095)`; dropped `REFERENCES domains(id)`; dropped `WITHOUT ROWID`; added column; and a table rebuilt via `ALTER TABLE ... RENAME` WITHOUT the verbatim-text rule compares UNEQUAL (proves the text layer catches SQLite's rename rewrite). + +### A.7 Acceptance criteria + +- [ ] Fresh file stamps `user_version = 1`, opens as current. +- [ ] A file stamped `1975011655` opens, restamps to 1, keeps every row; second open takes the current lane. +- [ ] `SchemaTooNew` and `SchemaUnsupported` refuse: schema dump, row count, watermark, stamp unchanged after refusal; no aside, no new file. One byte-hash variant on a checkpointed, sidecar-free fixture. +- [ ] Legacy-below-minimum ordering: with a test-local metadata view where minimum > 1 (drive the classification helper directly with injected constants — classification must be a pure function of `(stamped, minimum, current)` for exactly this reason), a legacy-fingerprint stamp classifies as REFUSE and no restamp happens. +- [ ] A.5 restamp-failure test. +- [ ] Synthetic chain through `runMigration` (1→3, two SQL steps, the second using the full A.2 rebuild sequence on a real table): backup exists, is a valid db, contains pre-migration rows; `user_version` lands on 3; rows survived; the rebuilt table's CREATE text equals the injected target text. +- [ ] Referenced-parent rebuild: a synthetic step rebuilds `domains` (referenced by `query_log`); after the migration, `query_log`'s stored FK text still says `REFERENCES domains(id)` (not `domains_old`), `PRAGMA foreign_key_check` is empty, and both pragmas read their defaults (`foreign_keys` per `applyPragmas`, `legacy_alter_table` off) after success AND after a forced failure. +- [ ] Mid-chain failure (step 2's SQL errors): canonical file logically unchanged (still version 1, rows intact), this run's backup deleted, an older backup preserved, `error.MigrationFailed`. +- [ ] `legacy_alter_table` pragma is OFF after both success and failure paths. +- [ ] Post-commit failure branch, driven through `openVersioned` with an injected synthetic plan (not by calling `runMigration` directly): force the pragma restore to fail after a successful COMMIT (fault seam) and assert: the file is at the target version with the migrated schema, the backup remains, the connection is closed exactly once (by the open path), that startup refuses with `error.MigrationFailed`, and the NEXT `openVersioned` under the same plan succeeds through the current-version lane. +- [ ] Backup retention: two successful synthetic migrations leave exactly one `pre-migrate-*`, the newer (step-4 authority, exact name). A directory seeded with an older epoch, a newest epoch, and a `-2` suffix tied at the newest epoch has a plain successful open delete only the older epoch — both max-epoch ties survive; an unparseable `pre-migrate-*` name survives untouched. +- [ ] Backup consistency: a row committed but not checkpointed (WAL-only) is present in the backup. +- [ ] `PRAGMA user_version` transactionality: set inside a transaction, ROLLBACK, original value observed. +- [ ] Oracle: fresh==fresh true; every A.6 negative test false; a `runMigration`-migrated file vs a fresh file at the target schema true. +- [ ] Grep scoped to `src/` and `tools/`: the `fingerprint_mismatch` identifier and the `schema-changed` aside-tag string are gone from active code (docs, specs, and changelog legitimately keep the words — the downgrade recovery text names the aside). Both suites green. + +--- + +## Session B: cut gate inversion + fixture proof + +### B.1 Fixtures + +- `src/storage/testdata/querylog-v1-schema.sql` — the version-1 DDL frozen verbatim (today's `querylog_schema.ddl` text; the stamp is NOT part of it — the loader applies `PRAGMA user_version = 1`). +- `src/storage/testdata/querylog-v1-data.sql` — representative COHERENT content: query_log rows covering every `route_kind` and the NULL variants (qtype, cache_hit, response_time_us, upstream, forward_zone), matching `domains` rows, a non-default `available_since`, and `bucket_*` projection rows consistent with the raw rows. A fixture-validity test loads it and runs the projection-coherence oracle BEFORE any migration, so an incoherent fixture fails on its own. +- Immutable once shipped (header comment). From here on, every supported logical version in `[minimum_supported_version, current_version]` has a fixture pair — the current version's pair is the next migration's starting fixture, and an explicit break ships the new baseline pair. + +The **fixture proof tests** (appended to `querylog_migrations.zig` by Session B, sequenced after A): + +1. For EVERY starting version in `[minimum_supported_version, current_version)`: load that version's fixture pair, stamp it, run the real production chain, assert `schemaEquivalent` against a fresh-`ddl` db, every row survived, `available_since` preserved, projection coherence holds. Empty today; load-bearing without edits the day the chain grows. +2. The CURRENT version's fixture pair, stamped `current_version`, opens on the current lane, is `schemaEquivalent` to a fresh-`ddl` db, and passes projection coherence — the pair whose existence Gate 2 requires is thereby proven coherent, since the `[minimum, current)` loop never exercises it. +3. The legacy-stamp variant: a v1-fixture file stamped `1975011655` — while `minimum_supported_version == 1` it opens, restamps, and passes the same assertions as (2); the test is written against the classification helper's injected constants so that when a future break raises the minimum above 1, its companion assertion (legacy stamp + minimum > 1 REFUSES with `error.SchemaUnsupported`, file untouched) is already in the suite. + +### B.2 The gate in tools/cut.zig + +`cut` imports `querylog_versions` (pure, no SQLite — the link contract is why A.1 is separate). Two INDEPENDENT gates replace the disclose-a-reset gate. Let `prev_version` be the previous tag's `current_version` (parse `git show :src/storage/querylog_versions.zig` with the existing simple-extraction style; a tag predating the module means 1). + +**Gate 1 — schema text.** Fingerprint the previous tag's DDL text vs the tree's. If changed, require ONE of: +- **Migration lane:** `current_version > prev_version` AND `prev_version >= minimum_supported_version` (the previous release's files are actually reachable — an explicit break can never wear this lane) AND the chain covers `[prev_version, current_version)` (with contiguous per-step files, that is `step_sql.len == current_version - minimum_supported_version` plus the fixture/file checks of Gate 2). +- **Explicit-break lane:** `current_version > prev_version` AND `minimum_supported_version == current_version` AND the changelog section contains BOTH "resets your query history" AND a `### Restoring your query history` heading with a non-empty body. +- Neither: FAIL. + +**Gate 2 — migration metadata.** Runs INDEPENDENTLY of Gate 1 (catches data-only migrations and prefix edits when the DDL is unchanged): +- Every `src/storage/migrations/v.sql` present at the previous tag: byte-identical in the tree; missing: FAIL. +- Every `src/storage/testdata/querylog-v*-{schema,data}.sql` present at the previous tag: byte-identical; missing: FAIL. +- A fixture pair exists for every version in `[minimum_supported_version, current_version]`: else FAIL. +- `current_version < prev_version`: FAIL (never regresses). +- `current_version > prev_version` with neither a new step file nor a break (`minimum == current`): FAIL. +- `current_version > prev_version` via new step(s) — REGARDLESS of whether the DDL fingerprint moved (data-only migrations included): the changelog section must contain "migrates your query log in place"; else FAIL. +- Let `prev_minimum` be the previous tag's `minimum_supported_version` (module absent at tag: 1). `minimum_supported_version < prev_minimum`: FAIL. `minimum_supported_version > prev_minimum` is ONLY acceptable as the full explicit break — `minimum == current` AND `current_version > prev_version` AND the break-lane changelog requirements — REGARDLESS of the DDL fingerprint; any other raise: FAIL (a release must never silently drop supported schemas). +- The tree's `legacy_fingerprint` is not the literal `1975011655`: FAIL (the legacy anchor is frozen forever; editing it strands unupgraded 0.0.12/0.0.13 files). + +### B.3 Acceptance criteria + +- [ ] Gate unit tests (pure functions over injected inputs, house style): unchanged schema + unchanged metadata passes; migration lane passes; explicit-break lane passes; changed schema with neither FAILS; break metadata (`minimum == current`) presented with the migration phrase FAILS Gate 1's migration lane; version bump with short chain FAILS; edited shipped step FAILS despite a version append; edited fixture FAILS; deleted step file FAILS; missing target-version fixture pair FAILS; version regression FAILS; version bump with no step and no break FAILS; data-only step (unchanged DDL) without the migration phrase FAILS; minimum regression FAILS; minimum raised without the full break FAILS (unchanged DDL variant included); edited `legacy_fingerprint` FAILS; previous tag without `querylog_versions.zig` maps to `prev_version == 1` and `prev_minimum == 1`. +- [ ] Fixture-validity test and fixture proof loop pass in the plain suite. +- [ ] `zig build cut` compiles; both suites green. + +--- + +## Session C: docs (after A) + +- `docs/how-to/troubleshoot.md`: new section "The server refuses to start over querylog.db" — `SchemaTooNew` (downgraded binary: return to the newer release, or restore the matching `pre-migrate` backup), `SchemaUnsupported` (file predates 0.0.12 or is foreign: not migratable; how to set it aside by hand if starting empty is acceptable), `MigrationFailed`/`MigrationBackupFailed` (the server never starts empty on its own; before the migration committed the file is untouched, and in the rare committed-but-unclean case the log says the migration DID complete, the backup is kept, and the next start simply proceeds). +- `docs/reference/` page on the query-log lifecycle: version stamp, in-place migration, one kept backup, the honest downgrade contract (downgrading to 0.0.13 or older RESETS the log — those binaries predate this contract; migration-aware binaries refuse cleanly), corruption as the only automatic recreate, the explicit-break contract (versioned, refused at startup, changelog carries restore instructions). +- Update the documents that still state the old contract: `PLAN.md`, `docs/explanation/architecture.md`, `specs/release-cut.md` — surgical edits to the stale sentences only. + +Acceptance: prose accurate against A/B behavior, unwrapped lines, both suites still green. + +--- + +## Module Layout + +- `src/storage/querylog_versions.zig` — NEW: pure version/step metadata (cut-importable, no hooks by design). +- `src/storage/querylog_migrations.zig` — NEW: `migrateSteps`, `runMigration`, `schemaEquivalent`, fixture proof tests. +- `src/storage/migrations/` — one immutable SQL file per shipped step. NOT created this milestone (empty chain; git carries no empty directory) — the first real step creates it. +- `src/storage/querylog_schema.zig` — open-path rework, stamp change, lane deletions, every-open retention. +- `src/storage/testdata/querylog-v1-schema.sql`, `querylog-v1-data.sql` — NEW frozen fixtures. +- `src/storage/querylog_fixtures.zig` — NEW (Session B, as built): fixture loading and the survival oracle — full-content comparison against a pristine copy, each value encoded type-tag + byte-length + bytes so the comparison is injective (review round 2026-08-28). +- `tools/cut.zig` — two-gate rework. +- Session C's doc files. + +## File Ownership + +A: both new storage modules, `migrations/` dir, `querylog_schema.zig`, callers touched by lane deletion. B (after A): `tools/cut.zig`, `testdata/`, appends tests to `querylog_migrations.zig`, and makes the projection-coherence checker in `queries_repo.zig` `pub` (export-only edit — the checker is currently private to that file, which no session otherwise owns; B's fixture tests need it). C (after A): docs, `PLAN.md`, `specs/release-cut.md`. Orchestrator: CHANGELOG.md, spec sync. + +A also owns the fault-injection seam addition in `src/storage/db.zig` (A.5). + +## Changelog requirement (orchestrator) + +This milestone's own changelog entry must disclose the one hazard neither gate can see: opening querylog.db under this release restamps it from the legacy fingerprint to version 1, so a LATER downgrade to 0.0.13 or older treats the numeric stamp as a fingerprint mismatch, renames the file to a `.schema-changed-` aside, and starts an empty log. The restamp itself creates NO backup, so the accurate recovery is: return to a migration-aware release; stop the server; move the empty downgrade-created `querylog.db` out of the way AND delete its `querylog.db-wal`/`querylog.db-shm` sidecars (replaying the empty file's sidecars into the restored history would corrupt it — the recreate code documents this); move the downgrade-created `.schema-changed-` aside back to `querylog.db`; start. The entry states the hazard and exactly that procedure. + +## Acceptance Criteria (Milestone Complete) + +- [ ] No code path recreates or sets aside a healthy querylog.db (grep proves the lane gone). +- [ ] A 0.0.13-created file (v1 schema + `1975011655` stamp) opens under the new binary with every row intact. +- [ ] Refusals and pre-commit migration failures leave the file logically untouched; a post-commit `MigrationFailed` leaves it successfully migrated to `target` (backup kept) and only refuses that one startup; the restamp is this milestone's only real mutation and its failure refuses without loss. +- [ ] The cut gate refuses: a schema change with neither lane, any edit to shipped steps or fixtures, a data-only migration without disclosure, and an explicit break without versioning + restore instructions. +- [ ] Both suites green, fmt clean. + +## Anti-Requirements + +- NO migration steps for pre-0.0.12 schemas (refusal with instructions is the contract). +- NO real chain step this milestone; synthetic chains live in tests only. +- NO migration hooks — steps are SQL files, period; a future need amends the design in its own spec. +- NO generic column-intersection salvage. +- NO `ALTER TABLE ADD/RENAME COLUMN` on kept tables in future steps (rebuild rule; recorded in doc comments, machine-enforced only via the oracle's exact-text layer). +- NO admin UI/API surface for migrations; startup log lines are the interface. +- NO config knob for backup retention. +- NO change to config.db handling. diff --git a/specs/release-cut.md b/specs/release-cut.md index b69ab6a..9ae1370 100644 --- a/specs/release-cut.md +++ b/specs/release-cut.md @@ -57,7 +57,9 @@ Pure functions unit-tested: semver validation (accept/reject table incl. leading ## Addendum: the schema gate (post-0.0.9) -0.0.9 changed the `query_log` DDL and its announcement said nothing about it. `querylog.db` is never migrated: the server stamps `PRAGMA user_version` with a CRC32 of the DDL text, and on a mismatch it renames the file aside and creates an empty one, so the first start after such a release destroys the operator's query history. Nothing in the cut noticed, because nothing in the cut had ever read the schema. +> Superseded by milestone 38. The addendum below records the gate as it was first built, when `querylog.db` was never migrated. The server now versions and migrates that file in place (`docs/reference/query-log-lifecycle.md`), and the single disclose-a-reset check described here was replaced by the two independent gates of `specs/milestone-38.md` §B.2. + +0.0.9 changed the `query_log` DDL and its announcement said nothing about it. At the time `querylog.db` was never migrated: the server stamped `PRAGMA user_version` with a CRC32 of the DDL text, and on a mismatch it renamed the file aside and created an empty one, so the first start after such a release destroyed the operator's query history. Nothing in the cut noticed, because nothing in the cut had ever read the schema. `schema-gate` is a read-only preflight check beside the others. It compares releases, not commits: diff --git a/src/app.zig b/src/app.zig index 0f892c8..fa5db52 100644 --- a/src/app.zig +++ b/src/app.zig @@ -58,6 +58,7 @@ const model = @import("config/model.zig"); const pause = @import("server/pause.zig"); const queries_repo = @import("storage/repositories/queries_repo.zig"); const query_sink = @import("server/query_sink.zig"); +const querylog_migrations = @import("storage/querylog_migrations.zig"); const querylog_schema = @import("storage/querylog_schema.zig"); const rate_limiter = @import("server/rate_limiter.zig"); const reconcile = @import("config/reconcile.zig"); @@ -1316,9 +1317,9 @@ test "the recreated detail names the aside and the new coverage start" { var buf: [events.Store.max_detail_len]u8 = undefined; try std.testing.expectEqualStrings( - "previous file kept as 'querylog.db.schema-changed-1700000000'; " ++ + "previous file kept as 'querylog.db.quick-check-failed-1700000000'; " ++ "query history is available from 1700000001", - recreatedDetail(&buf, "querylog.db.schema-changed-1700000000", 1700000001), + recreatedDetail(&buf, "querylog.db.quick-check-failed-1700000000", 1700000001), ); // A fresh file that will not answer is a separate failure; the line still @@ -1437,7 +1438,7 @@ const m29_ddl: [:0]const u8 = \\VALUES (1, unixepoch(), unixepoch() + 1); ; -test "an m29 query log is set aside and recreated without the upstream-history tables" { +test "an m29 query log refuses the startup and is left exactly as it is" { var threaded: std.Io.Threaded = .init(testing.allocator, .{}); defer threaded.deinit(); const io = threaded.io(); @@ -1448,10 +1449,6 @@ test "an m29 query log is set aside and recreated without the upstream-history t var path_buf: [256]u8 = undefined; const path = try std.fmt.bufPrintZ(&path_buf, ".zig-cache/tmp/{s}/querylog.db", .{tmp.sub_path}); - var fx: events_fixture.Fixture = .{}; - try fx.init(io, 1000); - defer fx.deinit(); - // The fixture is only worth anything while it is still a *different* // schema from this build's, and one that carries the deleted tables. try testing.expect(!std.mem.eql(u8, m29_ddl, querylog_schema.ddl)); @@ -1461,9 +1458,8 @@ test "an m29 query log is set aside and recreated without the upstream-history t // edit to the literal cannot satisfy by changing what it is compared to. try testing.expectEqual(m29_fingerprint, @as(i32, @bitCast(std.hash.Crc32.hash(m29_ddl)))); - // A healthy m29 file, stamped with the fingerprint m29's own DDL produced - // and backdated so its coverage promise is visibly the older one. - const m29_coverage = blk: { + // A healthy m29 file, stamped with the fingerprint m29's own DDL produced. + { var m29 = try db.Db.open(path, .{ .mode = .read_write_create }); defer m29.close(); try db.applyPragmas(&m29, .{}); @@ -1481,43 +1477,47 @@ test "an m29 query log is set aside and recreated without the upstream-history t "PRAGMA user_version = {d};", .{m29_fingerprint}, )); - break :blk try m29.queryInt("SELECT available_since FROM querylog_meta"); - }; - try testing.expectEqual(m29_available_since, m29_coverage); - - var opened = try querylog_schema.open(io, std.Io.Dir.cwd(), path); - defer opened.database.close(); - - // Set aside under the name that says the file was healthy and this build - // moved, and still on disk for an operator who wants it. - try testing.expectEqual(querylog_schema.RecreateReason.fingerprint_mismatch, opened.recreated.?); - try testing.expect(std.mem.indexOf(u8, opened.aside(), ".schema-changed-") != null); - try tmp.dir.access(io, std.fs.path.basename(opened.aside()), .{}); - - // The two tables are gone from the file this process will write to. - for ([_][]const u8{ "upstream_targets", "upstream_minute", "idx_upstream_minute_ts" }) |name| { - var stmt = try opened.database.prepare("SELECT count(*) FROM sqlite_schema WHERE name = ?1"); - defer stmt.deinit(); - try stmt.bindText(1, name); - try testing.expect(try stmt.step()); - try testing.expectEqual(@as(i64, 0), stmt.columnInt(0)); } - // Coverage restarts: the new file does not inherit the replaced one's - // promise about what it can answer. Strictly newer, not merely not-older — - // a recreation that copied the watermark across would pass the weaker test. - const coverage = try queries_repo.availableSince(&opened.database); - try testing.expect(coverage > m29_coverage); + // m29 predates the version stamp entirely: its `user_version` is a CRC of a + // schema no migration chain starts from, so the only honest answer is to + // refuse and say so. The pre-0.0.12 contract — set it aside and start empty + // — is gone. + querylog_migrations.expected_failures.begin(); + defer querylog_migrations.expected_failures.end(); + try testing.expectError( + error.SchemaUnsupported, + querylog_schema.open(io, std.Io.Dir.cwd(), path), + ); - reportQuerylogRecreated(&fx.store, io, 2000, &opened, &opened.database); - try testing.expectEqualStrings("query_log.recreated", try fx.text("SELECT code FROM operational_events")); - try testing.expectEqualStrings( - "fingerprint_mismatch", - try fx.text("SELECT subject_key FROM operational_events"), + // Nothing was renamed, nothing was created, and the file still answers for + // itself: the operator can downgrade and keep the history. + var entries: usize = 0; + var it = tmp.dir.iterate(); + while (try it.next(io)) |entry| { + try testing.expect(std.mem.startsWith(u8, entry.name, "querylog.db")); + try testing.expect(std.mem.indexOfScalar(u8, entry.name[10..], '.') == null); + entries += 1; + } + try testing.expect(entries >= 1); + + var reopened = try db.Db.open(path, .{ .mode = .read_write_existing }); + defer reopened.close(); + try testing.expectEqual( + @as(i64, m29_fingerprint), + try reopened.queryInt("PRAGMA user_version"), + ); + try testing.expectEqual( + m29_available_since, + try reopened.queryInt("SELECT available_since FROM querylog_meta"), + ); + try testing.expectEqual( + @as(i64, 1), + try reopened.queryInt("SELECT count(*) FROM upstream_targets"), ); } -test "a fingerprint recreate files a resolved event naming the real aside and watermark" { +test "a recreate files a resolved event naming the real aside and watermark" { var threaded: std.Io.Threaded = .init(testing.allocator, .{}); defer threaded.deinit(); const io = threaded.io(); @@ -1539,27 +1539,17 @@ test "a fingerprint recreate files a resolved event naming the real aside and wa created.database.close(); try testing.expectEqual(@as(i64, 0), try fx.count("SELECT count(*) FROM operational_events")); - // A healthy file this build's DDL no longer matches, which is what an - // upgrade that edits the schema produces. - { - var stamped = try db.Db.open(path, .{ .mode = .read_write_existing }); - defer stamped.close(); - var sql_buf: [64]u8 = undefined; - try stamped.exec(try std.fmt.bufPrintZ( - &sql_buf, - "PRAGMA user_version = {d};", - .{querylog_schema.fingerprint +% 1}, - )); - } + // Real damage: corruption is the only thing that recreates now. + try tmp.dir.writeFile(io, .{ .sub_path = "querylog.db", .data = "not a database at all" }); var recreated = try querylog_schema.open(io, std.Io.Dir.cwd(), path); defer recreated.database.close(); - try testing.expectEqual(querylog_schema.RecreateReason.fingerprint_mismatch, recreated.recreated.?); + try testing.expectEqual(querylog_schema.RecreateReason.not_a_database, recreated.recreated.?); reportQuerylogRecreated(&fx.store, io, 2000, &recreated, &recreated.database); try testing.expectEqualStrings("query_log.recreated", try fx.text("SELECT code FROM operational_events")); - try testing.expectEqualStrings("fingerprint_mismatch", try fx.text("SELECT subject_key FROM operational_events")); + try testing.expectEqualStrings("not_a_database", try fx.text("SELECT subject_key FROM operational_events")); try testing.expectEqualStrings("warning", try fx.text("SELECT severity FROM operational_events")); // One-shot: already over when it is filed, so it never becomes an open // episode `/api/health` counts. @@ -1612,18 +1602,17 @@ test "a recreate under a long data directory keeps the watermark and a usable na { var created = try querylog_schema.open(io, std.Io.Dir.cwd(), path); - defer created.database.close(); - var sql_buf: [64]u8 = undefined; - try created.database.exec(try std.fmt.bufPrintZ( - &sql_buf, - "PRAGMA user_version = {d};", - .{querylog_schema.fingerprint +% 1}, - )); + created.database.close(); + } + { + var deep_dir = try tmp.dir.openDir(io, nested, .{}); + defer deep_dir.close(io); + try deep_dir.writeFile(io, .{ .sub_path = "querylog.db", .data = "not a database at all" }); } var recreated = try querylog_schema.open(io, std.Io.Dir.cwd(), path); defer recreated.database.close(); - try testing.expectEqual(querylog_schema.RecreateReason.fingerprint_mismatch, recreated.recreated.?); + try testing.expectEqual(querylog_schema.RecreateReason.not_a_database, recreated.recreated.?); const line_overhead = "previous file kept as ''; query history is available from ".len; try testing.expect(recreated.aside().len + line_overhead > events.Store.max_detail_len); @@ -1682,6 +1671,13 @@ test "run maps a rejected configuration to exit 2 and everything else to exit 1" try std.testing.expectEqual(cli.exit_check, failureExitCode(error.ParseZon)); try std.testing.expectEqual(cli.exit_check, failureExitCode(error.NoUsableUpstreams)); try std.testing.expectEqual(cli.exit_check, failureExitCode(error.BadCertificate)); + // The query-log schema refusals: `run` is the only command that reaches + // them, and exit 1 would put a deliberate refusal under the unit's + // `Restart=on-failure`. + try std.testing.expectEqual(cli.exit_check, failureExitCode(error.SchemaTooNew)); + try std.testing.expectEqual(cli.exit_check, failureExitCode(error.SchemaUnsupported)); + try std.testing.expectEqual(cli.exit_check, failureExitCode(error.MigrationFailed)); + try std.testing.expectEqual(cli.exit_check, failureExitCode(error.MigrationBackupFailed)); try std.testing.expectEqual(cli.exit_runtime, failureExitCode(error.AccessDenied)); try std.testing.expectEqual(cli.exit_runtime, failureExitCode(error.OutOfMemory)); } diff --git a/src/config/faults.zig b/src/config/faults.zig index bd6efb4..872884b 100644 --- a/src/config/faults.zig +++ b/src/config/faults.zig @@ -13,7 +13,7 @@ const validate = @import("validate.zig"); /// `ValidateError` enters as a whole set rather than variant by variant, so a /// variant added to the validator cannot silently fall through to exit 1. The -/// five extras are the configuration faults raised outside the validator: the +/// first five extras are the configuration faults raised outside the validator: the /// ZON reader (`ParseZon`), the file size limit (`ConfigTooLarge`), the managed /// file the operator named and this process cannot open /// (`ManagedConfigUnreadable`, milestone-20 ruling 2), the composition root's @@ -26,6 +26,16 @@ const validate = @import("validate.zig"); /// missing file anywhere else stays a runtime failure. `config/loader.zig` owns /// the conversion and the closed set of open errors that qualify. /// +/// The four querylog schema refusals are here for the exit code, not because a +/// `.zon` file is wrong: `SchemaTooNew` and `SchemaUnsupported` are a deliberate +/// refusal to touch a `querylog.db` this binary does not understand, and +/// `MigrationFailed` and `MigrationBackupFailed` are a deliberate refusal to run +/// on a database whose migration or pre-migration backup did not complete. All +/// four need an operator, and none of them will resolve on a retry — exit 1 puts +/// them under systemd's `Restart=on-failure` and restart-loops a server that is +/// refusing on purpose. The unit's `RestartPreventExitStatus=2 64` is what exit +/// 2 buys them. +/// /// Not here on purpose: `error.DestructiveImport`, which reports what an import /// would do to the database rather than the content of a file, and is the one /// config-shaped exit 2 `cli` decides for itself. @@ -35,6 +45,10 @@ const ConfigFault = validate.ValidateError || error{ ManagedConfigUnreadable, NoUsableUpstreams, BadCertificate, + SchemaTooNew, + SchemaUnsupported, + MigrationFailed, + MigrationBackupFailed, }; const faults: []const anyerror = blk: { @@ -119,6 +133,16 @@ test "the seed-file errors that used to exit 1 from run are configuration faults try testing.expect(isConfigFault(error.NoUpstreams)); } +test "the querylog schema refusals exit 2 so systemd does not restart-loop them" { + try testing.expect(isConfigFault(error.SchemaTooNew)); + try testing.expect(isConfigFault(error.SchemaUnsupported)); + try testing.expect(isConfigFault(error.MigrationFailed)); + try testing.expect(isConfigFault(error.MigrationBackupFailed)); + // The refusals are a closed set. A neighbouring schema error is a corrupt + // database, not a refusal, and stays a runtime failure. + try testing.expect(!isConfigFault(error.SchemaCorrupt)); +} + test "a runtime failure is not a configuration fault" { try testing.expect(!isConfigFault(error.OutOfMemory)); try testing.expect(!isConfigFault(error.AccessDenied)); diff --git a/src/storage/db.zig b/src/storage/db.zig index e500462..a319a24 100644 --- a/src/storage/db.zig +++ b/src/storage/db.zig @@ -54,9 +54,24 @@ pub const c = struct { pub extern fn sqlite3_column_count(stmt: *c.Stmt) c_int; pub extern fn sqlite3_column_type(stmt: *c.Stmt, col: c_int) c_int; pub extern fn sqlite3_column_int64(stmt: *c.Stmt, col: c_int) i64; + pub extern fn sqlite3_column_double(stmt: *c.Stmt, col: c_int) f64; pub extern fn sqlite3_column_text(stmt: *c.Stmt, col: c_int) ?[*]const u8; pub extern fn sqlite3_column_bytes(stmt: *c.Stmt, col: c_int) c_int; + pub extern fn sqlite3_set_authorizer(db: *Sqlite3, xAuth: ?*const AuthCallback, user_data: ?*anyopaque) c_int; + pub extern fn sqlite3_get_autocommit(db: *Sqlite3) c_int; pub extern fn sqlite3_last_insert_rowid(db: *Sqlite3) i64; + + /// `int (*)(void*, int, const char*, const char*, const char*, const char*)`. + /// The four name arguments are NULL for actions that do not use them, which + /// `SQLITE_TRANSACTION` is except for its operation name. + pub const AuthCallback = fn ( + user_data: ?*anyopaque, + action: c_int, + arg1: ?[*:0]const u8, + arg2: ?[*:0]const u8, + database: ?[*:0]const u8, + trigger_or_view: ?[*:0]const u8, + ) callconv(.c) c_int; pub extern fn sqlite3_changes(db: *Sqlite3) c_int; pub extern fn sqlite3_total_changes(db: *Sqlite3) c_int; }; @@ -105,6 +120,20 @@ pub const open_flag = struct { pub const exrescode: c_int = 0x2000000; }; +/// The authorizer verdicts and the one action code nxdns denies, from the +/// vendored `sqlite3.h` (3.53.4). `SQLITE_DENY` fails the *prepare* with +/// `SQLITE_AUTH`, which is what makes it a real guard: the statement never +/// runs at all. +pub const auth = struct { + pub const deny: c_int = 1; + pub const transaction: c_int = 22; + /// `SAVEPOINT`, `RELEASE` and `ROLLBACK TO` report under this code, not + /// under `transaction`. A savepoint at the outermost level opens a real + /// transaction and `RELEASE` commits it, so a step using one splits the + /// migration exactly as a bare `COMMIT` would. + pub const savepoint: c_int = 32; +}; + /// Column type codes returned by `sqlite3_column_type`. pub const column_type = struct { pub const integer: c_int = 1; @@ -367,6 +396,7 @@ pub const Db = struct { /// message is read back through `sqlite3_errmsg`, so there is no /// `sqlite3_free` obligation. pub fn exec(self: *Db, sql: [:0]const u8) Error!void { + if (execFaultTripped(sql)) return error.Internal; return check(c.sqlite3_exec(self.handle, sql.ptr, null, null, null)); } @@ -424,6 +454,38 @@ pub const Db = struct { return value; } + /// False while a transaction is open on this connection. The migration + /// runner's belt check: a step that somehow ended the runner's transaction + /// must not be allowed to look like a success. + pub fn inTransaction(self: *Db) bool { + return c.sqlite3_get_autocommit(self.handle) == 0; + } + + /// Denies every transaction statement — `BEGIN`, `COMMIT`, `ROLLBACK`, + /// `SAVEPOINT`, `RELEASE` — until `clearAuthorizer` runs. + /// + /// `guard` must outlive the installed window: SQLite keeps the pointer. + /// Install and clear are a scoped pair; the migration runner clears on + /// every exit path, because a leaked authorizer would go on denying the + /// ROLLBACK that cleans up after the very statement it rejected. + pub fn denyTransactions(self: *Db, guard: *TransactionGuard) Error!void { + guard.* = .{}; + return check(c.sqlite3_set_authorizer(self.handle, transactionAuthorizer, guard)); + } + + /// Never fails in a way the caller can act on: passing a null callback only + /// clears state SQLite already holds. A failure is logged and swallowed so + /// this stays usable in `defer`. + pub fn clearAuthorizer(self: *Db) void { + const rc = c.sqlite3_set_authorizer(self.handle, null, null); + if (rc != result.ok) { + log.err("sqlite3_set_authorizer(null) returned {s} (code {d})", .{ + std.mem.span(c.sqlite3_errstr(rc)), + rc, + }); + } + } + pub fn lastInsertRowid(self: *Db) i64 { return c.sqlite3_last_insert_rowid(self.handle); } @@ -442,6 +504,34 @@ pub const Db = struct { } }; +/// Records whether the authorizer installed by `Db.denyTransactions` actually +/// rejected anything. The rejection reaches the caller as `error.Auth`, which +/// is indistinguishable from any other authorization failure; this flag is what +/// lets the migration runner name the real cause. +pub const TransactionGuard = struct { + denied: bool = false, +}; + +fn transactionAuthorizer( + user_data: ?*anyopaque, + action: c_int, + arg1: ?[*:0]const u8, + arg2: ?[*:0]const u8, + database: ?[*:0]const u8, + trigger_or_view: ?[*:0]const u8, +) callconv(.c) c_int { + _ = arg2; + _ = database; + _ = trigger_or_view; + if (action != auth.transaction and action != auth.savepoint) return result.ok; + const guard: *TransactionGuard = @ptrCast(@alignCast(user_data.?)); + guard.denied = true; + log.warn("migration step attempted a transaction statement: {s}", .{ + if (arg1) |op| std.mem.span(op) else "(unnamed)", + }); + return auth.deny; +} + fn openHandle(filename: [:0]const u8, flags: c_int) Error!*c.Sqlite3 { var handle: ?*c.Sqlite3 = null; const rc = c.sqlite3_open_v2(filename.ptr, &handle, flags, null); @@ -963,6 +1053,45 @@ pub const read_tx_faults = if (builtin.is_test) struct { } } else struct {}; +/// Fails one chosen `Db.exec` so a test can drive a failure SQLite itself will +/// not produce on demand. Test builds only, same shape as `read_tx_seam`. +/// +/// The migration paths this exists for — the legacy restamp and the +/// post-commit pragma restore — run statements that always succeed against a +/// healthy file, and their recovery behaviour is the whole point of the +/// milestone. Matching on the SQL text rather than counting calls keeps a test +/// naming the statement it means. +const exec_seam = if (builtin.is_test) struct { + var fail_matching: ?[]const u8 = null; +} else struct {}; + +fn execFaultTripped(sql: []const u8) bool { + if (!builtin.is_test) return false; + const needle = exec_seam.fail_matching orelse return false; + if (std.mem.indexOf(u8, sql, needle) == null) return false; + exec_seam.fail_matching = null; + return true; +} + +/// The seam's controls, for tests in this file and in the storage layer. +pub const exec_faults = if (builtin.is_test) struct { + /// Arms the next `Db.exec` whose SQL contains `needle` to fail with + /// `error.Internal` before the statement reaches SQLite. One shot: it + /// disarms itself when it trips. Pair it with `defer disarm()` so a test + /// that never trips the fault cannot leak it into the next one. + pub fn failNextMatching(needle: []const u8) void { + exec_seam.fail_matching = needle; + } + + pub fn disarm() void { + exec_seam.fail_matching = null; + } + + pub fn armed() bool { + return exec_seam.fail_matching != null; + } +} else struct {}; + const testing = std.testing; fn openMemory() Error!Db { @@ -1479,3 +1608,57 @@ test "a duplicate insert into a UNIQUE column returns error.Constraint" { try stmt.bindText(1, "only"); try testing.expectError(error.Constraint, stmt.step()); } + +test "the transaction authorizer denies transaction statements and clears cleanly" { + var db = try openMemory(); + defer db.close(); + try db.exec("CREATE TABLE t (id INTEGER PRIMARY KEY);"); + + try db.exec("BEGIN IMMEDIATE;"); + try testing.expect(db.inTransaction()); + + var guard: TransactionGuard = .{}; + try db.denyTransactions(&guard); + // Ordinary work still runs: only the transaction statements are refused. + try db.exec("INSERT INTO t (id) VALUES (1);"); + try testing.expect(!guard.denied); + try testing.expectError(error.Auth, db.exec("COMMIT;")); + try testing.expect(guard.denied); + // The deny happens at prepare, so the transaction is still open. + try testing.expect(db.inTransaction()); + + // SAVEPOINT and its RELEASE report under a different action code, and they + // are the same bypass: at the outermost level they are a transaction under + // another name, and inside one they can still discard the migration's work. + guard.denied = false; + try testing.expectError(error.Auth, db.exec("SAVEPOINT half_a_migration;")); + try testing.expect(guard.denied); + guard.denied = false; + try testing.expectError(error.Auth, db.exec("RELEASE half_a_migration;")); + try testing.expect(guard.denied); + try testing.expect(db.inTransaction()); + + db.clearAuthorizer(); + // The same connection is usable again — a leaked authorizer would strand it + // inside the transaction by denying this too. + try db.exec("ROLLBACK;"); + try testing.expect(!db.inTransaction()); + try testing.expectEqual(@as(i64, 0), try db.queryInt("SELECT count(*) FROM t")); +} + +test "the exec fault seam fires once, on the statement it names" { + var db = try openMemory(); + defer db.close(); + try db.exec("CREATE TABLE t (id INTEGER PRIMARY KEY);"); + + exec_faults.failNextMatching("COMMIT"); + defer exec_faults.disarm(); + + try db.exec("BEGIN IMMEDIATE;"); + try db.exec("INSERT INTO t (id) VALUES (1);"); + try testing.expectError(error.Internal, db.exec("COMMIT;")); + try testing.expect(!exec_faults.armed()); + // Disarmed: the retry is a real COMMIT. + try db.exec("COMMIT;"); + try testing.expectEqual(@as(i64, 1), try db.queryInt("SELECT count(*) FROM t")); +} diff --git a/src/storage/phase6_integration_test.zig b/src/storage/phase6_integration_test.zig index 651950b..364244f 100644 --- a/src/storage/phase6_integration_test.zig +++ b/src/storage/phase6_integration_test.zig @@ -27,6 +27,7 @@ const disk_monitor = @import("disk_monitor.zig"); const logger = @import("logger.zig"); const queries_repo = @import("repositories/queries_repo.zig"); const querylog_schema = @import("querylog_schema.zig"); +const querylog_versions = @import("querylog_versions.zig"); const retention = @import("retention.zig"); const testing = std.testing; @@ -233,7 +234,7 @@ test "S8 case 1: the logger writes a real querylog.db end to end" { try testing.expectEqual(@as(i64, 250), try queries_repo.countRows(log_db.database())); try testing.expectEqual(@as(i64, 10), try queries_repo.countDomains(log_db.database())); try testing.expectEqual( - @as(i64, querylog_schema.fingerprint), + @as(i64, querylog_versions.current_version), try log_db.database().queryInt("PRAGMA user_version"), ); } diff --git a/src/storage/querylog_fixtures.zig b/src/storage/querylog_fixtures.zig new file mode 100644 index 0000000..9e4fb2d --- /dev/null +++ b/src/storage/querylog_fixtures.zig @@ -0,0 +1,535 @@ +//! The shipped `querylog.db` fixtures, and the proof that every supported +//! schema version reaches the current one with the operator's rows intact. +//! +//! A file of its own, not a section of `querylog_migrations.zig`, because of the +//! link contract that split `querylog_versions.zig` out in the first place. The +//! assertions here need `repositories/queries_repo.zig`'s projection-coherence +//! oracle, and that file reaches across `src/` for the config and filter types; +//! importing it from `querylog_migrations.zig` would pull all of it into the +//! module `tools/cut.zig` builds `querylog_schema.zig` as, where those paths lie +//! outside the module root and do not compile. +//! +//! The fixtures themselves — `testdata/querylog-v-{schema,data}.sql` — are +//! immutable once released. Every version in `[minimum_supported_version, +//! current_version]` has a pair: the current version's pair is the next +//! migration's starting point, and an explicit break ships the new baseline. + +const std = @import("std"); + +const db = @import("db.zig"); +const migrations = @import("querylog_migrations.zig"); +const queries_repo = @import("repositories/queries_repo.zig"); +const querylog_schema = @import("querylog_schema.zig"); +const versions = @import("querylog_versions.zig"); + +const testing = std.testing; + +/// A temporary directory and the `querylog.db` path inside it. Deliberately not +/// `querylog_migrations.zig`'s test harness: that one builds synthetic schemas, +/// while everything here starts from the shipped fixture files. +const Harness = struct { + threaded: std.Io.Threaded, + tmp: std.testing.TmpDir, + buf: [256]u8 = undefined, + + fn init() Harness { + return .{ + .threaded = .init(testing.allocator, .{}), + .tmp = testing.tmpDir(.{ .iterate = true }), + }; + } + + fn deinit(self: *Harness) void { + self.tmp.cleanup(); + self.threaded.deinit(); + } + + fn io(self: *Harness) std.Io { + return self.threaded.io(); + } + + fn path(self: *Harness) [:0]const u8 { + return std.fmt.bufPrintZ(&self.buf, ".zig-cache/tmp/{s}/querylog.db", .{self.tmp.sub_path}) catch + unreachable; + } + + fn openLive(self: *Harness) !db.Db { + var database = try db.Db.open(self.path(), .{ .mode = .read_write_existing }); + errdefer database.close(); + try db.applyPragmas(&database, .{}); + return database; + } +}; + +/// One shipped schema version's frozen pair. Both halves are immutable once +/// released — the release cut byte-compares them against the previous tag — and +/// every version in `[minimum_supported_version, current_version]` must have a +/// pair, which the cut gate also enforces. +const Fixture = struct { + version: i32, + schema: [:0]const u8, + data: [:0]const u8, +}; + +const fixtures = [_]Fixture{ + .{ + .version = 1, + .schema = @embedFile("testdata/querylog-v1-schema.sql"), + .data = @embedFile("testdata/querylog-v1-data.sql"), + }, +}; + +fn fixtureFor(version: i32) ?Fixture { + for (fixtures) |fixture| { + if (fixture.version == version) return fixture; + } + return null; +} + +/// Writes a fixture pair to `path` and stamps it. `stamp` is a parameter rather +/// than `fixture.version` because the legacy-fingerprint file is the same +/// version-1 bytes under a different stamp. +fn writeFixture(path: [:0]const u8, fixture: Fixture, stamp: i32) !void { + var database = try db.Db.open(path, .{ .mode = .read_write_create }); + defer database.close(); + try db.applyPragmas(&database, .{}); + try database.exec(fixture.schema); + try database.exec(fixture.data); + + var buf: [64]u8 = undefined; + const sql = std.fmt.bufPrintZ(&buf, "PRAGMA user_version = {d};", .{stamp}) catch unreachable; + try database.exec(sql); +} + +/// A second path in the harness's directory, for the reference databases the +/// assertions below compare against. +fn sidePath(h: *Harness, buf: []u8, name: []const u8) [:0]const u8 { + return std.fmt.bufPrintZ(buf, ".zig-cache/tmp/{s}/{s}", .{ h.tmp.sub_path, name }) catch unreachable; +} + +/// A database holding nothing but the current `ddl`, which is what a file +/// created by this build is. +fn openFreshCurrent(h: *Harness, buf: []u8) !db.Db { + var database = try db.Db.open(sidePath(h, buf, "fresh.db"), .{ .mode = .read_write_create }); + errdefer database.close(); + try db.applyPragmas(&database, .{}); + try database.exec(querylog_schema.ddl); + return database; +} + +/// An untouched load of `fixture`, to compare a migrated or opened file against +/// rather than restating the fixture's contents in the assertions. +fn openPristine(h: *Harness, buf: []u8, fixture: Fixture) !db.Db { + const path = sidePath(h, buf, "pristine.db"); + try writeFixture(path, fixture, fixture.version); + var database = try db.Db.open(path, .{ .mode = .read_write_existing }); + errdefer database.close(); + try db.applyPragmas(&database, .{}); + return database; +} + +/// Every row of every table, as one canonical text. Exact equality is only the +/// right question for a file no migration has reshaped; a migrated file is +/// checked by the counts and the watermark instead. +fn dumpContent(gpa: std.mem.Allocator, database: *db.Db, out: *std.ArrayList(u8)) !void { + var tables: std.ArrayList([]u8) = .empty; + defer freeOwned(gpa, &tables); + { + var stmt = try database.prepare( + \\SELECT name FROM sqlite_schema + \\WHERE type = 'table' AND name NOT LIKE 'sqlite\_%' ESCAPE '\' + \\ORDER BY name + ); + defer stmt.deinit(); + // Copied out before the per-table statements step: a borrowed + // `columnText` would not survive them. + while (try stmt.step()) try tables.append(gpa, try stmt.columnTextAlloc(gpa, 0)); + } + + // The lines are sorted here rather than by the query: the four `bucket_*` + // tables are WITHOUT ROWID, so `ORDER BY rowid` is not available to all of + // them and no single column list is. + var lines: std.ArrayList([]u8) = .empty; + defer freeOwned(gpa, &lines); + + for (tables.items) |table| { + var sql_buf: [256]u8 = undefined; + const sql = std.fmt.bufPrint(&sql_buf, "SELECT * FROM \"{s}\"", .{table}) catch unreachable; + var stmt = try database.prepare(sql); + defer stmt.deinit(); + while (try stmt.step()) { + var line: std.ArrayList(u8) = .empty; + errdefer line.deinit(gpa); + try line.print(gpa, "R|{s}", .{table}); + var col: c_int = 0; + const columns: c_int = db.c.sqlite3_column_count(stmt.handle); + while (col < columns) : (col += 1) { + try line.print(gpa, "|{s}", .{stmt.columnTextOrNull(col) orelse ""}); + } + try lines.append(gpa, try line.toOwnedSlice(gpa)); + } + } + + std.mem.sortUnstable([]u8, lines.items, {}, struct { + fn lessThan(_: void, a: []u8, b: []u8) bool { + return std.mem.lessThan(u8, a, b); + } + }.lessThan); + for (lines.items) |line| { + try out.appendSlice(gpa, line); + try out.append(gpa, '\n'); + } +} + +/// One value, encoded so that no two different values can produce the same +/// bytes: a type tag, the byte length, then the bytes themselves. +/// +/// Nothing here is a sentinel and nothing is escaped, which is the point. A +/// serialization that wrote NULL as `` cannot tell a NULL apart from the +/// six-character string of the same name, and one that separated values with +/// `|` cannot tell `a|b` in one column from `a` and `b` in two — so a migration +/// that turned a NULL `upstream` into text, or shifted a value from one column +/// into its neighbour, would compare EQUAL to the original. The length prefix +/// makes the stream uniquely decodable, so equal encodings mean equal rows. +/// +/// A float travels as its bit pattern rather than as printed digits: the +/// question here is whether the value survived, not whether it rounds the same. +fn writeValue(gpa: std.mem.Allocator, stmt: *db.Stmt, col: c_int, out: *std.ArrayList(u8)) !void { + switch (db.c.sqlite3_column_type(stmt.handle, col)) { + db.column_type.null_value => try out.print(gpa, "n0:", .{}), + db.column_type.integer => { + var buf: [24]u8 = undefined; + const text = std.fmt.bufPrint(&buf, "{d}", .{stmt.columnInt(col)}) catch unreachable; + try out.print(gpa, "i{d}:{s}", .{ text.len, text }); + }, + db.column_type.float => { + const bits: u64 = @bitCast(db.c.sqlite3_column_double(stmt.handle, col)); + var buf: [24]u8 = undefined; + const text = std.fmt.bufPrint(&buf, "{d}", .{bits}) catch unreachable; + try out.print(gpa, "f{d}:{s}", .{ text.len, text }); + }, + db.column_type.blob => { + // `columnText` on a blob hands back the same bytes SQLite stores, + // which is what this compares; it is not read as text. + const bytes = stmt.columnText(col); + try out.print(gpa, "b{d}:{s}", .{ bytes.len, bytes }); + }, + else => { + const bytes = stmt.columnText(col); + try out.print(gpa, "t{d}:{s}", .{ bytes.len, bytes }); + }, + } +} + +/// One query's rows, in the order the query returns them, tagged with `label` so +/// that a difference names the relation it came from. The column count is part +/// of each row for the same reason the lengths are part of each value. +fn dumpQuery( + gpa: std.mem.Allocator, + database: *db.Db, + label: []const u8, + sql: [:0]const u8, + out: *std.ArrayList(u8), +) !void { + var stmt = try database.prepare(sql); + defer stmt.deinit(); + while (try stmt.step()) { + const columns: c_int = db.c.sqlite3_column_count(stmt.handle); + try out.print(gpa, "{s}:{d}:", .{ label, columns }); + var col: c_int = 0; + while (col < columns) : (col += 1) { + try writeValue(gpa, &stmt, col, out); + } + try out.append(gpa, '\n'); + } +} + +/// The operator's data as the application means it, in an order a migration +/// cannot permute. +/// +/// `q.*` rather than a column list on purpose: a migration that adds a column +/// must show that column here, and a hand-written list would quietly stop +/// covering the table the day it grows. `d.domain` rides along so that the text +/// a row names is compared, not only the id it happens to hold. +const logical_relations = [_]struct { label: []const u8, sql: [:0]const u8 }{ + .{ + .label = "query_log", + .sql = + \\SELECT q.*, d.domain FROM query_log q + \\JOIN domains d ON d.id = q.domain_id + \\ORDER BY q.id + , + }, + .{ .label = "domains", .sql = "SELECT * FROM domains ORDER BY id" }, + .{ .label = "querylog_meta", .sql = "SELECT * FROM querylog_meta ORDER BY id" }, +}; + +fn dumpLogical(gpa: std.mem.Allocator, database: *db.Db, out: *std.ArrayList(u8)) !void { + for (logical_relations) |relation| { + try dumpQuery(gpa, database, relation.label, relation.sql, out); + } +} + +fn freeOwned(gpa: std.mem.Allocator, list: *std.ArrayList([]u8)) void { + for (list.items) |item| gpa.free(item); + list.deinit(gpa); +} + +fn expectSameContent(a: *db.Db, b: *db.Db) !void { + var text_a: std.ArrayList(u8) = .empty; + defer text_a.deinit(testing.allocator); + var text_b: std.ArrayList(u8) = .empty; + defer text_b.deinit(testing.allocator); + try dumpContent(testing.allocator, a, &text_a); + try dumpContent(testing.allocator, b, &text_b); + try testing.expectEqualStrings(text_a.items, text_b.items); +} + +fn rowCount(database: *db.Db, table: []const u8) !i64 { + var buf: [128]u8 = undefined; + return database.queryInt(std.fmt.bufPrint(&buf, "SELECT count(*) FROM \"{s}\"", .{table}) catch unreachable); +} + +/// What must hold after a fixture has come through `openVersioned`, whatever +/// path it took: every row the operator had is still there with the same +/// CONTENT, and the projections still agree with the raw rows. +/// +/// Counting rows and checking one watermark is what this used to do, and a +/// migration that shifted a timestamp, dropped a `qtype` or crossed two rows' +/// `client_ip` values passed it. The comparison is therefore the full logical +/// content of the three operator tables, `available_since` included as one +/// column of `querylog_meta` among the rest. The counts stay because a count +/// difference is the failure worth naming plainly. +/// +/// The relations are named rather than derived because they are the ones holding +/// operator data; the `bucket_*` projections are derived from them, which +/// `expectProjectionsMatchRecompute` is the right check for. A future migration +/// that renames a table amends this alongside the step that does it. So does one +/// that renumbers `id` values: this asserts they survive, which every rebuild +/// written to the A.2 rule does. +fn expectFixtureSurvived(opened: *db.Db, pristine: *db.Db) !void { + for ([_][]const u8{ "query_log", "domains", "querylog_meta" }) |table| { + try testing.expectEqual(try rowCount(pristine, table), try rowCount(opened, table)); + } + + var opened_text: std.ArrayList(u8) = .empty; + defer opened_text.deinit(testing.allocator); + var pristine_text: std.ArrayList(u8) = .empty; + defer pristine_text.deinit(testing.allocator); + try dumpLogical(testing.allocator, opened, &opened_text); + try dumpLogical(testing.allocator, pristine, &pristine_text); + try testing.expectEqualStrings(pristine_text.items, opened_text.items); + + try queries_repo.expectProjectionsMatchRecompute(opened); +} + +test "the survival comparison tells a NULL apart from text that looks like one" { + // The mutation a count-and-watermark check misses entirely, and a + // sentinel-string serialization misses just as completely: one column of one + // row stops being NULL and becomes the very text the sentinel used. Every + // count, the watermark and the projections all still agree. + const fixture = fixtureFor(1) orelse return error.MissingFixture; + + var h: Harness = .init(); + defer h.deinit(); + try writeFixture(h.path(), fixture, fixture.version); + + var side: [256]u8 = undefined; + var pristine = try openPristine(&h, &side, fixture); + defer pristine.close(); + + var corrupted = try h.openLive(); + defer corrupted.close(); + try testing.expectEqual(@as(i64, 0), corrupted.changes()); + try corrupted.exec( + \\UPDATE query_log SET upstream = '' + \\WHERE id = (SELECT min(id) FROM query_log WHERE upstream IS NULL) + ); + // The fixture has to actually carry a NULL `upstream` for this to be a test + // of anything. + try testing.expectEqual(@as(i64, 1), corrupted.changes()); + + // Everything the old check looked at still agrees, which is why it passed. + for ([_][]const u8{ "query_log", "domains", "querylog_meta" }) |table| { + try testing.expectEqual(try rowCount(&pristine, table), try rowCount(&corrupted, table)); + } + try testing.expectEqual( + try pristine.queryInt("SELECT available_since FROM querylog_meta WHERE id = 1"), + try corrupted.queryInt("SELECT available_since FROM querylog_meta WHERE id = 1"), + ); + try queries_repo.expectProjectionsMatchRecompute(&corrupted); + + // The dumps are compared here rather than through `expectFixtureSurvived` + // so that a PASSING run stays silent: `expectEqualStrings` prints the whole + // diff before it returns its error, and this is the comparison that function + // makes. + var corrupted_text: std.ArrayList(u8) = .empty; + defer corrupted_text.deinit(testing.allocator); + var pristine_text: std.ArrayList(u8) = .empty; + defer pristine_text.deinit(testing.allocator); + try dumpLogical(testing.allocator, &corrupted, &corrupted_text); + try dumpLogical(testing.allocator, &pristine, &pristine_text); + try testing.expect(!std.mem.eql(u8, corrupted_text.items, pristine_text.items)); +} + +test "every supported version ships a fixture pair" { + // The cut gate enforces this against a release; the suite enforces it + // against a commit, so a version bump that forgot its fixtures fails here + // long before anyone reaches for `zig build cut`. + var version = versions.minimum_supported_version; + while (version <= versions.current_version) : (version += 1) { + try testing.expect(fixtureFor(version) != null); + } +} + +test "each shipped fixture pair is coherent before any migration touches it" { + for (fixtures) |fixture| { + var h: Harness = .init(); + defer h.deinit(); + try writeFixture(h.path(), fixture, fixture.version); + + var database = try h.openLive(); + defer database.close(); + try testing.expectEqual( + @as(i64, fixture.version), + try database.queryInt("PRAGMA user_version"), + ); + // An incoherent fixture has to fail as a fixture, not later as a + // migration that appears to have corrupted the projections. + try queries_repo.expectProjectionsMatchRecompute(&database); + try testing.expectEqual(@as(i64, 0), try database.queryInt("SELECT count(*) FROM pragma_foreign_key_check")); + } +} + +test "every fixture below the current version migrates to it through the shipped chain" { + // Empty while the chain is: `minimum_supported_version == current_version` + // today. It is written as the loop so that the day a step ships, the + // fixture it starts from is proved through the REAL production plan with no + // edit to this test. + var version = versions.minimum_supported_version; + while (version < versions.current_version) : (version += 1) { + const fixture = fixtureFor(version) orelse return error.MissingFixture; + + var h: Harness = .init(); + defer h.deinit(); + try writeFixture(h.path(), fixture, version); + + var side: [256]u8 = undefined; + var pristine = try openPristine(&h, &side, fixture); + defer pristine.close(); + + var result = try querylog_schema.open(h.io(), std.Io.Dir.cwd(), h.path()); + defer result.database.close(); + try testing.expectEqual(@as(?querylog_schema.RecreateReason, null), result.recreated); + try testing.expectEqual( + @as(i64, versions.current_version), + try result.database.queryInt("PRAGMA user_version"), + ); + + var fresh_buf: [256]u8 = undefined; + var fresh = try openFreshCurrent(&h, &fresh_buf); + defer fresh.close(); + try testing.expect(try migrations.schemaEquivalent(testing.allocator, &result.database, &fresh)); + try expectFixtureSurvived(&result.database, &pristine); + } +} + +test "the current version's fixture pair opens on the current lane unchanged" { + // The loop above never reaches this pair, and Gate 2 of the cut requires it + // to exist. This is what proves it is a real, coherent file rather than one + // shipped to satisfy a gate. + const fixture = fixtureFor(versions.current_version) orelse return error.MissingFixture; + + var h: Harness = .init(); + defer h.deinit(); + try writeFixture(h.path(), fixture, versions.current_version); + + var side: [256]u8 = undefined; + var pristine = try openPristine(&h, &side, fixture); + defer pristine.close(); + + var result = try querylog_schema.open(h.io(), std.Io.Dir.cwd(), h.path()); + defer result.database.close(); + try testing.expectEqual(@as(?querylog_schema.RecreateReason, null), result.recreated); + try testing.expectEqual( + @as(i64, versions.current_version), + try result.database.queryInt("PRAGMA user_version"), + ); + + var fresh_buf: [256]u8 = undefined; + var fresh = try openFreshCurrent(&h, &fresh_buf); + defer fresh.close(); + try testing.expect(try migrations.schemaEquivalent(testing.allocator, &result.database, &fresh)); + try expectFixtureSurvived(&result.database, &pristine); + // No migration ran, so nothing reshaped anything: byte-for-byte the rows + // that were loaded. + try expectSameContent(&result.database, &pristine); +} + +test "a fixture carrying the 0.0.12 fingerprint restamps, and refuses once the minimum rises" { + const fixture = fixtureFor(1) orelse return error.MissingFixture; + + // Version 1 is the legacy fingerprint's logical version, so the pair only + // has anything to say while 1 is still supported. + if (versions.minimum_supported_version <= 1 and versions.current_version >= 1) { + var h: Harness = .init(); + defer h.deinit(); + try writeFixture(h.path(), fixture, versions.legacy_fingerprint); + + var side: [256]u8 = undefined; + var pristine = try openPristine(&h, &side, fixture); + defer pristine.close(); + + var result = try querylog_schema.open(h.io(), std.Io.Dir.cwd(), h.path()); + defer result.database.close(); + try testing.expectEqual(@as(?querylog_schema.RecreateReason, null), result.recreated); + try testing.expectEqual( + @as(i64, versions.current_version), + try result.database.queryInt("PRAGMA user_version"), + ); + + var fresh_buf: [256]u8 = undefined; + var fresh = try openFreshCurrent(&h, &fresh_buf); + defer fresh.close(); + try testing.expect(try migrations.schemaEquivalent(testing.allocator, &result.database, &fresh)); + try expectFixtureSurvived(&result.database, &pristine); + } + + // The companion, already in the suite for the day a break raises the + // minimum above 1: the same bytes under the same stamp are then a file this + // build cannot reach, and it is refused without being touched. + var h: Harness = .init(); + defer h.deinit(); + try writeFixture(h.path(), fixture, versions.legacy_fingerprint); + + const after_break: querylog_schema.Plan = .{ + .minimum = 2, + .current = 2, + .legacy_fingerprint = versions.legacy_fingerprint, + .step_sql = &.{}, + }; + try testing.expect(querylog_schema.classify(versions.legacy_fingerprint, after_break).action == + .refuse_unsupported); + try testing.expect(!querylog_schema.classify(versions.legacy_fingerprint, after_break).restamp); + + var handle: ?db.Db = try h.openLive(); + defer if (handle) |*open_db| open_db.close(); + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + try testing.expectError( + error.SchemaUnsupported, + querylog_schema.openVersioned(h.io(), std.Io.Dir.cwd(), h.path(), &handle, after_break), + ); + + var check = try h.openLive(); + defer check.close(); + try testing.expectEqual( + @as(i64, versions.legacy_fingerprint), + try check.queryInt("PRAGMA user_version"), + ); + var pristine_buf: [256]u8 = undefined; + var pristine = try openPristine(&h, &pristine_buf, fixture); + defer pristine.close(); + try expectSameContent(&check, &pristine); +} diff --git a/src/storage/querylog_migrations.zig b/src/storage/querylog_migrations.zig new file mode 100644 index 0000000..be4f3ba --- /dev/null +++ b/src/storage/querylog_migrations.zig @@ -0,0 +1,1316 @@ +//! Running a `querylog.db` schema migration, and proving one landed. +//! +//! `querylog_versions.zig` says which versions exist and which SQL gets from +//! one to the next. This file executes that chain against a real file, and +//! `schemaEquivalent` is the oracle a test uses to assert the migrated file is +//! indistinguishable from a freshly created one. +//! +//! The contract the whole design rests on: a migration either lands completely +//! or leaves the file exactly as it was. There is no third state and no +//! salvage path — a failure refuses the startup rather than recreating +//! anything, because the operator's query history is the thing being protected. + +const std = @import("std"); +const builtin = @import("builtin"); +const assert = std.debug.assert; + +const db = @import("db.zig"); + +const log = std.log.scoped(.querylog_migrations); + +pub const Error = error{ + /// The migration did not happen. The file is logically what it was before. + MigrationFailed, + /// The pre-migration backup could not be made, so the migration was never + /// attempted. The file is untouched. + MigrationBackupFailed, + /// Internal to this module and `querylog_schema.open`: the migration + /// COMMITTED, and something after the commit failed. The file is at the + /// target version and the backup is kept; the connection is no longer + /// trustworthy, so the open path closes it and refuses this one startup. + MigrationCommittedButUnclean, + /// Internal, the same route as the one above for a different reason: the + /// migration did NOT happen — the file is logically what it was and this + /// run's backup is gone — but restoring the connection's pragmas failed, so + /// it may still carry `foreign_keys = OFF` or `legacy_alter_table = ON`. + /// `MigrationFailed` promises a usable connection and this cannot keep that + /// promise, so it travels as its own error and the open path closes the + /// handle rather than handing a half-configured connection to the server. + MigrationFailedUnclean, + NameTooLong, +} || db.Error; + +/// Reports a condition that refuses a startup or abandons a migration. +/// +/// `err`, except while a test has said it is deliberately causing one. The test +/// runner fails a test that logs at `err`, and every refusal path in this +/// milestone is a path some test has to drive on purpose. Capture is opt-in, so +/// an unexpected refusal in an unrelated test still fails it — the same +/// reasoning as `db.read_tx_faults`. +fn fail(comptime fmt: []const u8, args: anytype) void { + if (expected_failures.capturing()) { + log.warn(fmt, args); + } else { + log.err(fmt, args); + } +} + +const failure_seam = if (builtin.is_test) struct { + var capturing: bool = false; +} else struct {}; + +/// Shared by `querylog_schema.zig`, whose refusal and restamp-failure lines are +/// the same class of deliberate fault. +pub const expected_failures = struct { + pub fn capturing() bool { + if (!builtin.is_test) return false; + return failure_seam.capturing; + } + + pub fn begin() void { + if (builtin.is_test) failure_seam.capturing = true; + } + + pub fn end() void { + if (builtin.is_test) failure_seam.capturing = false; + } +}; + +/// Long enough for any path this program will be handed, plus a backup suffix. +const path_buf_len = 4096 + 64; + +const backup_infix = ".pre-migrate-"; + +// --------------------------------------------------------------------------- +// the step runner +// --------------------------------------------------------------------------- + +/// Executes `sql` and stamps `target`, inside the transaction the CALLER has +/// already opened. It never opens or closes one. +/// +/// SLICING CONTRACT: `sql` is exactly the `[from, target)` suffix of the chain +/// — `sql[0]` migrates `from` to `from + 1`. Production callers slice +/// `step_sql[from - minimum_supported_version ..]`; tests pass synthetic +/// chains. The assert below is the whole guard against an off-by-one that would +/// stamp a version the file does not have. +/// +/// While the steps run, SQLite's authorizer denies every transaction statement. +/// A post-step `sqlite3_get_autocommit` check alone would not do: a step +/// containing `COMMIT; BEGIN IMMEDIATE;` leaves autocommit off and looks fine, +/// having silently split the migration into two transactions. Denying at +/// prepare time means such a step cannot run at all. The autocommit check stays +/// as a belt against anything that ends a transaction without a statement the +/// authorizer sees. +pub fn migrateSteps( + database: *db.Db, + sql: []const [:0]const u8, + from: i32, + target: i32, +) (db.Error || error{TransactionViolation})!void { + assert(target > from); + assert(sql.len == @as(usize, @intCast(target - from))); + assert(database.inTransaction()); + + var guard: db.TransactionGuard = .{}; + try database.denyTransactions(&guard); + defer database.clearAuthorizer(); + + for (sql, 0..) |step, index| { + database.exec(step) catch |e| { + if (guard.denied) { + fail("migration step {d} ({d} -> {d}) contains a transaction statement", .{ + index, + from + @as(i32, @intCast(index)), + from + @as(i32, @intCast(index)) + 1, + }); + return error.TransactionViolation; + } + var buf: [256]u8 = undefined; + fail("migration step {d} ({d} -> {d}) failed: {s}", .{ + index, + from + @as(i32, @intCast(index)), + from + @as(i32, @intCast(index)) + 1, + database.lastError(&buf), + }); + return e; + }; + if (!database.inTransaction()) { + fail("migration step {d} ended the migration transaction", .{index}); + return error.TransactionViolation; + } + } + + var stamp_buf: [64]u8 = undefined; + const stamp = std.fmt.bufPrintZ(&stamp_buf, "PRAGMA user_version = {d};", .{target}) catch + unreachable; // an i32 and a fixed prefix cannot overrun 64 bytes + try database.exec(stamp); +} + +// --------------------------------------------------------------------------- +// the orchestration +// --------------------------------------------------------------------------- + +/// Backs the file up, migrates it from `from` to `target` in one transaction, +/// and keeps exactly this run's backup. +/// +/// `database` is borrowed and this function never closes it: the caller owns the +/// handle and already has the one error-path close. What the caller gets back +/// decides what that handle is worth. +/// +/// - Success and `error.MigrationFailed`, `error.MigrationBackupFailed`: the +/// connection is USABLE. Its pragmas are what `applyPragmas` guarantees. +/// - `error.MigrationCommittedButUnclean` (the commit succeeded, the restore did +/// not) and `error.MigrationFailedUnclean` (the migration did not happen and +/// the restore did not either): the connection MUST NOT be used, because it +/// may still hold `foreign_keys = OFF` or `legacy_alter_table = ON`. Both are +/// internal to this module and `querylog_schema.openVersioned`, which closes +/// the handle on each and surfaces `error.MigrationFailed`. +/// +/// A failed restore is therefore never swallowed. Reporting "the migration +/// failed, carry on with this connection" while foreign keys are off would let +/// the server write rows no constraint ever checked. +/// +/// `dir` and `path` follow `querylog_schema.open`'s resolution rule: either an +/// absolute path with `dir` on its parent, or `std.Io.Dir.cwd()` with a +/// cwd-relative path. The backup is written beside `path`. +pub fn runMigration( + io: std.Io, + dir: std.Io.Dir, + path: [:0]const u8, + database: *db.Db, + sql: []const [:0]const u8, + from: i32, + target: i32, +) Error!void { + assert(target > from); + assert(sql.len == @as(usize, @intCast(target - from))); + + var backup_buf: [path_buf_len]u8 = undefined; + const backup = try makeBackup(io, dir, path, database, &backup_buf); + + // Both pragmas are connection-global and non-transactional, so they are set + // before BEGIN and restored on every way out. + // + // `foreign_keys = OFF` is not a convenience. `applyPragmas` turns foreign + // keys ON, and with them on, renaming a table that another table REFERENCES + // rewrites the child's stored FK text to point at `_old` — which then + // dangles the moment `_old` drops. `legacy_alter_table = ON` alone does + // not prevent that; it governs views and the renamed table's own text. + enterMigrationMode(database) catch { + var buf: [256]u8 = undefined; + fail("cannot prepare the connection for migration {d} -> {d}: {s}", .{ + from, target, database.lastError(&buf), + }); + deleteQuietly(io, dir, backup); + return restoreAfterFailure(database, from, target); + }; + + migrateInTransaction(database, sql, from, target) catch |e| { + var buf: [256]u8 = undefined; + fail("querylog migration {d} -> {d} failed before commit ({t}): {s}; " ++ + "the database is unchanged", .{ from, target, e, database.lastError(&buf) }); + deleteQuietly(io, dir, backup); + return restoreAfterFailure(database, from, target); + }; + + leaveMigrationMode(database) catch { + var buf: [256]u8 = undefined; + fail("querylog migration {d} -> {d} COMMITTED and the database IS at version {d}, " ++ + "but the connection could not be restored: {s}; the backup '{s}' is kept and " ++ + "the next start will open the migrated file normally", .{ + from, target, target, database.lastError(&buf), backup, + }); + return error.MigrationCommittedButUnclean; + }; + + // Authoritative retention: this run knows the exact name to keep, so every + // other backup goes now. `querylog_schema`'s every-open pass only retries + // what fails here. + pruneBackupsExcept(io, dir, path, std.fs.path.basename(backup)); + log.info("migrated querylog database '{s}' from version {d} to {d}; " ++ + "pre-migration backup kept as '{s}'", .{ path, from, target, backup }); +} + +/// BEGIN IMMEDIATE, re-check the stamp under the lock, run the chain, verify +/// referential integrity, COMMIT. Any error leaves the transaction rolled back. +fn migrateInTransaction( + database: *db.Db, + sql: []const [:0]const u8, + from: i32, + target: i32, +) !void { + var tx = try db.Tx.begin(database); + errdefer tx.rollback(); + + // nxdns owns this file exclusively, so a stamp that moved between the + // classification and the lock means another process is writing to it. + // Migrating on top of that would be guessing. + const stamped = try database.queryInt("PRAGMA user_version"); + if (stamped != from) { + fail("querylog user_version changed from {d} to {d} while the migration was starting; " ++ + "another process is using the file", .{ from, stamped }); + return error.MigrationFailed; + } + + try migrateSteps(database, sql, from, target); + + var check = try database.prepare("PRAGMA foreign_key_check"); + defer check.deinit(); + if (try check.step()) { + fail("querylog migration {d} -> {d} left a foreign key violation in table '{s}'", .{ + from, target, check.columnText(0), + }); + return error.MigrationFailed; + } + + try tx.commit(); +} + +/// Restores the connection after a migration that did not happen, and says which +/// error the caller is looking at. +/// +/// The caller has already logged WHY the migration failed and deleted this run's +/// backup; the file is logically untouched either way. The only question left is +/// whether the connection came back, and the answer changes the error rather +/// than being dropped. +fn restoreAfterFailure(database: *db.Db, from: i32, target: i32) Error { + leaveMigrationMode(database) catch { + var buf: [256]u8 = undefined; + fail("querylog migration {d} -> {d} did NOT happen and the database is unchanged, " ++ + "but the connection's pragmas could not be restored: {s}; the connection is " ++ + "closed and the server does not start", .{ from, target, database.lastError(&buf) }); + return error.MigrationFailedUnclean; + }; + return error.MigrationFailed; +} + +fn enterMigrationMode(database: *db.Db) db.Error!void { + try database.exec("PRAGMA foreign_keys = OFF;"); + try database.exec("PRAGMA legacy_alter_table = ON;"); +} + +/// Restores what `applyPragmas` guarantees every connection to this file has. +fn leaveMigrationMode(database: *db.Db) db.Error!void { + try database.exec("PRAGMA legacy_alter_table = OFF;"); + try database.exec("PRAGMA foreign_keys = ON;"); +} + +// --------------------------------------------------------------------------- +// the backup +// --------------------------------------------------------------------------- + +/// `VACUUM INTO` on the live connection with no transaction open: it produces a +/// complete, consistent copy of the committed database including whatever is +/// still only in the write-ahead log, which a byte copy of the main file would +/// miss. +/// +/// Returns the backup path, borrowed from `buf`. +fn makeBackup( + io: std.Io, + dir: std.Io.Dir, + path: [:0]const u8, + database: *db.Db, + buf: *[path_buf_len]u8, +) Error![]const u8 { + assert(!database.inTransaction()); + const seconds = std.Io.Clock.real.now(io).toSeconds(); + + var attempt: u32 = 1; + while (attempt < 100) : (attempt += 1) { + const backup = (if (attempt == 1) + std.fmt.bufPrint(buf, "{s}{s}{d}", .{ path, backup_infix, seconds }) + else + std.fmt.bufPrint(buf, "{s}{s}{d}-{d}", .{ path, backup_infix, seconds, attempt })) catch + return error.NameTooLong; + + // `VACUUM INTO` refuses an existing destination itself; probing first is + // what turns that refusal into a retry under a different name instead of + // a failed migration. + dir.access(io, backup, .{}) catch |e| switch (e) { + error.FileNotFound => { + try vacuumInto(database, backup, io, dir); + return backup; + }, + else => { + fail("cannot probe backup destination '{s}': {t}", .{ backup, e }); + return error.MigrationBackupFailed; + }, + }; + } + fail("cannot find an unused backup name beside '{s}'", .{path}); + return error.MigrationBackupFailed; +} + +fn vacuumInto(database: *db.Db, backup: []const u8, io: std.Io, dir: std.Io.Dir) Error!void { + var sql_buf: [path_buf_len * 2 + 32]u8 = undefined; + var quoted_buf: [path_buf_len * 2]u8 = undefined; + const quoted = quoteLiteral("ed_buf, backup) catch return error.NameTooLong; + const sql = std.fmt.bufPrintZ(&sql_buf, "VACUUM INTO '{s}';", .{quoted}) catch + return error.NameTooLong; + + database.exec(sql) catch { + var buf: [256]u8 = undefined; + fail("cannot back up the querylog database to '{s}': {s}", .{ backup, database.lastError(&buf) }); + // Only the destination this call just named: an older, valid backup + // beside it is the operator's last copy and must survive a failure here. + deleteQuietly(io, dir, backup); + return error.MigrationBackupFailed; + }; + + // SQLite creates the destination at `0644 & ~umask`, as it does the live + // file, and the backup holds the same query history. A backup this process + // cannot lock down may be world-readable, so it is a failed backup: drop it + // and refuse, exactly as a failed `VACUUM INTO` does. + dir.setFilePermissions(io, backup, .fromMode(0o600), .{}) catch |e| { + fail("cannot restrict the querylog backup '{s}' to 0600: {t}", .{ backup, e }); + deleteQuietly(io, dir, backup); + return error.MigrationBackupFailed; + }; +} + +/// Doubles every `'` so `text` can sit inside a single-quoted SQL literal. +/// `VACUUM INTO` takes an expression, not a bindable parameter, so the path has +/// to enter the statement as text — and a data directory is operator input. +fn quoteLiteral(buf: []u8, text: []const u8) error{NoSpaceLeft}![]const u8 { + var out: usize = 0; + for (text) |ch| { + const width: usize = if (ch == '\'') 2 else 1; + if (out + width > buf.len) return error.NoSpaceLeft; + if (ch == '\'') { + buf[out] = '\''; + out += 1; + } + buf[out] = ch; + out += 1; + } + return buf[0..out]; +} + +fn deleteQuietly(io: std.Io, dir: std.Io.Dir, target: []const u8) void { + dir.deleteFile(io, target) catch |e| switch (e) { + error.FileNotFound => {}, + else => log.warn("cannot delete '{s}': {t}", .{ target, e }), + }; +} + +// --------------------------------------------------------------------------- +// backup retention +// --------------------------------------------------------------------------- + +/// The authoritative pass: this run made `keep`, so everything else beside it +/// is a previous migration's backup and the contract keeps only the newest. +/// +/// Best effort by design. A backup that will not delete is disk to reclaim, not +/// a reason to refuse a startup whose migration already committed. +pub fn pruneBackupsExcept(io: std.Io, dir: std.Io.Dir, path: []const u8, keep: []const u8) void { + sweep(io, dir, path, .{ .keep_name = keep }); +} + +/// The retry pass, run on every plain successful open: deletes backups from +/// migrations whose own step-4 cleanup failed. +/// +/// Deliberately weaker than the pass above, because it does not know which name +/// this run made. It deletes only names whose epoch is STRICTLY below the +/// maximum epoch present, so every file tied at the newest epoch survives — +/// including collision suffixes — and a name that does not parse is never +/// touched. +/// +/// **This assumes a forward-moving wall clock between migrations.** Under a +/// clock rollback an older backup could carry the higher epoch and outrank a +/// genuinely newer one. That is why the exact-name pass above is the primary +/// mechanism and this is only its retry. +pub fn pruneBackupsConservative(io: std.Io, dir: std.Io.Dir, path: []const u8) void { + var prefix_buf: [path_buf_len]u8 = undefined; + const prefix = backupPrefix(&prefix_buf, path) orelse return; + + var newest: ?i64 = null; + { + var parent = openParent(io, dir, path) orelse return; + defer parent.close(io); + var it = parent.iterate(); + while (it.next(io) catch |e| { + log.warn("cannot scan for querylog backups: {t}", .{e}); + return; + }) |entry| { + if (entry.kind != .file) continue; + const epoch = backupEpoch(entry.name, prefix) orelse continue; + if (newest == null or epoch > newest.?) newest = epoch; + } + } + const max = newest orelse return; + sweep(io, dir, path, .{ .below_epoch = max }); +} + +const SweepRule = union(enum) { + /// Delete every backup whose name is not this one. + keep_name: []const u8, + /// Delete every backup whose epoch is strictly below this. + below_epoch: i64, +}; + +/// One pass over the directory, collecting victims before deleting any of them: +/// removing entries from under a live directory iterator is not something the +/// `std.Io.Dir` iterator promises to survive. +fn sweep(io: std.Io, dir: std.Io.Dir, path: []const u8, rule: SweepRule) void { + var prefix_buf: [path_buf_len]u8 = undefined; + const prefix = backupPrefix(&prefix_buf, path) orelse return; + + var rounds: u32 = 0; + while (rounds < 64) : (rounds += 1) { + var victims: Victims = .{}; + var parent = openParent(io, dir, path) orelse return; + defer parent.close(io); + { + var it = parent.iterate(); + while (it.next(io) catch |e| { + log.warn("cannot scan for querylog backups: {t}", .{e}); + return; + }) |entry| { + if (entry.kind != .file) continue; + const epoch = backupEpoch(entry.name, prefix) orelse continue; + const doomed = switch (rule) { + .keep_name => |keep| !std.mem.eql(u8, entry.name, keep), + .below_epoch => |max| epoch < max, + }; + if (!doomed) continue; + if (!victims.push(entry.name)) break; + } + } + for (0..victims.count) |i| { + deleteQuietly(io, parent, victims.names[i][0..victims.lengths[i]]); + } + if (!victims.overflowed) return; + } + log.warn("gave up sweeping querylog backups beside '{s}'", .{path}); +} + +/// A bounded batch of names to delete. The list is short in every real +/// deployment — one migration leaves one backup — so overflow means "sweep +/// again" rather than "allocate". +const Victims = struct { + const capacity = 16; + const name_len = 256; + + names: [capacity][name_len]u8 = undefined, + lengths: [capacity]usize = undefined, + count: usize = 0, + overflowed: bool = false, + + fn push(self: *Victims, name: []const u8) bool { + if (name.len > name_len) { + log.warn("querylog backup name is too long to delete: '{s}'", .{name}); + return true; + } + if (self.count == capacity) { + self.overflowed = true; + return false; + } + @memcpy(self.names[self.count][0..name.len], name); + self.lengths[self.count] = name.len; + self.count += 1; + return true; + } +}; + +fn openParent(io: std.Io, dir: std.Io.Dir, path: []const u8) ?std.Io.Dir { + const parent_path = std.fs.path.dirname(path) orelse "."; + return dir.openDir(io, parent_path, .{ .iterate = true }) catch |e| { + log.warn("cannot open '{s}' to manage querylog backups: {t}", .{ parent_path, e }); + return null; + }; +} + +/// ``: what every backup beside `path` starts with. +fn backupPrefix(buf: []u8, path: []const u8) ?[]const u8 { + return std.fmt.bufPrint(buf, "{s}{s}", .{ std.fs.path.basename(path), backup_infix }) catch { + log.warn("querylog path is too long to manage backups for: '{s}'", .{path}); + return null; + }; +} + +/// The epoch in `.pre-migrate-` or +/// `.pre-migrate--`, or null when `name` is not a backup +/// this code wrote. Anything that does not parse is somebody else's file. +fn backupEpoch(name: []const u8, prefix: []const u8) ?i64 { + if (!std.mem.startsWith(u8, name, prefix)) return null; + const rest = name[prefix.len..]; + if (rest.len == 0) return null; + const digits = if (std.mem.indexOfScalar(u8, rest, '-')) |dash| blk: { + const suffix = rest[dash + 1 ..]; + if (suffix.len == 0) return null; + _ = std.fmt.parseInt(u32, suffix, 10) catch return null; + break :blk rest[0..dash]; + } else rest; + if (digits.len == 0) return null; + return std.fmt.parseInt(i64, digits, 10) catch null; +} + +// --------------------------------------------------------------------------- +// the equivalence oracle +// --------------------------------------------------------------------------- + +/// True when `a` and `b` carry the same schema, in the sense that no SQL +/// statement could tell them apart. +/// +/// Two layers, both of which must agree: +/// +/// 1. **Exact text.** Every non-`sqlite_` object's `(type, name, tbl_name, +/// sql)`, with `sql` compared byte for byte. No normalization is needed or +/// wanted: the rebuild rule in `querylog_versions.step_sql` requires a +/// migrated table to carry the verbatim fresh CREATE text, and a fresh file +/// trivially does. This layer is what sees CHECK constraints, foreign key +/// clauses, WITHOUT ROWID, partial-index predicates and trigger bodies. +/// 2. **Structural belt.** Per table, `table_xinfo` plus the `table_list` +/// `wr`/`strict` flags and `foreign_key_list`; per index, `index_list` +/// flags plus `index_xinfo`. This is what stops a change SQLite records +/// somewhere other than the CREATE text from passing. +/// +/// Both layers land in one canonical text per database, which is then compared +/// whole — the sort orders come from the queries themselves. +pub fn schemaEquivalent( + gpa: std.mem.Allocator, + a: *db.Db, + b: *db.Db, +) (db.Error || std.mem.Allocator.Error)!bool { + var text_a: std.ArrayList(u8) = .empty; + defer text_a.deinit(gpa); + var text_b: std.ArrayList(u8) = .empty; + defer text_b.deinit(gpa); + + try describeSchema(gpa, a, &text_a); + try describeSchema(gpa, b, &text_b); + return std.mem.eql(u8, text_a.items, text_b.items); +} + +const DescribeError = db.Error || std.mem.Allocator.Error; + +fn describeSchema(gpa: std.mem.Allocator, database: *db.Db, out: *std.ArrayList(u8)) DescribeError!void { + { + var stmt = try database.prepare( + \\SELECT type, name, tbl_name, coalesce(sql, '') + \\FROM sqlite_schema WHERE name NOT LIKE 'sqlite\_%' ESCAPE '\' + \\ORDER BY type, name, tbl_name + ); + defer stmt.deinit(); + while (try stmt.step()) { + try out.print(gpa, "O|{s}|{s}|{s}|{s}\n", .{ + stmt.columnText(0), stmt.columnText(1), stmt.columnText(2), stmt.columnText(3), + }); + } + } + + const table_names_sql = + \\SELECT name FROM sqlite_schema + \\WHERE type = 'table' AND name NOT LIKE 'sqlite\_%' ESCAPE '\' + \\ORDER BY name + ; + var tables = try collectNames(gpa, database, table_names_sql, null); + defer freeNames(gpa, &tables); + + for (tables.items) |table| { + try describeRows(gpa, database, out, "T", table, + \\SELECT wr, strict FROM pragma_table_list + \\WHERE schema = 'main' AND name = ?1 + ); + try describeRows(gpa, database, out, "X", table, + \\SELECT cid, name, type, "notnull", coalesce(dflt_value, ''), pk, hidden + \\FROM pragma_table_xinfo(?1) ORDER BY cid + ); + try describeRows(gpa, database, out, "F", table, + \\SELECT id, seq, "table", "from", coalesce("to", ''), on_update, on_delete, match + \\FROM pragma_foreign_key_list(?1) ORDER BY id, seq + ); + // Autoindexes are included on purpose: a UNIQUE constraint's index is + // part of what the file enforces, and both sides list them the same way. + try describeRows(gpa, database, out, "L", table, + \\SELECT name, "unique", origin, partial + \\FROM pragma_index_list(?1) ORDER BY name + ); + + var indexes = try collectNames( + gpa, + database, + "SELECT name FROM pragma_index_list(?1) ORDER BY name", + table, + ); + defer freeNames(gpa, &indexes); + for (indexes.items) |index| { + try describeRows(gpa, database, out, "I", index, + \\SELECT seqno, cid, coalesce(name, ''), desc, coll, key + \\FROM pragma_index_xinfo(?1) ORDER BY seqno + ); + } + } +} + +/// One line per row, tagged and prefixed with the object the pragma was asked +/// about, so a difference names what differs rather than just where. +fn describeRows( + gpa: std.mem.Allocator, + database: *db.Db, + out: *std.ArrayList(u8), + comptime tag: []const u8, + subject: []const u8, + sql: []const u8, +) DescribeError!void { + var stmt = try database.prepare(sql); + defer stmt.deinit(); + try stmt.bindText(1, subject); + while (try stmt.step()) { + try out.print(gpa, tag ++ "|{s}", .{subject}); + var col: c_int = 0; + const columns = std.math.cast(c_int, columnCount(&stmt)) orelse unreachable; + while (col < columns) : (col += 1) { + try out.print(gpa, "|{s}", .{stmt.columnTextOrNull(col) orelse ""}); + } + try out.append(gpa, '\n'); + } +} + +fn columnCount(stmt: *db.Stmt) usize { + return @intCast(db.c.sqlite3_column_count(stmt.handle)); +} + +/// Names from a query with one text column, owned by `gpa`. `subject`, when +/// present, is bound as `?1` — the table a pragma function is asked about. +/// +/// The names are copied out before anything else runs against the connection: +/// the pragma statements below step while this list is alive, and a borrowed +/// `columnText` would not survive that. +fn collectNames( + gpa: std.mem.Allocator, + database: *db.Db, + sql: []const u8, + subject: ?[]const u8, +) DescribeError!std.ArrayList([]u8) { + var names: std.ArrayList([]u8) = .empty; + errdefer freeNames(gpa, &names); + + var stmt = try database.prepare(sql); + defer stmt.deinit(); + if (subject) |t| try stmt.bindText(1, t); + while (try stmt.step()) { + try names.append(gpa, try stmt.columnTextAlloc(gpa, 0)); + } + return names; +} + +fn freeNames(gpa: std.mem.Allocator, names: *std.ArrayList([]u8)) void { + for (names.items) |name| gpa.free(name); + names.deinit(gpa); +} + +// --------------------------------------------------------------------------- +// tests +// --------------------------------------------------------------------------- + +const testing = std.testing; + +/// A synthetic schema, not `querylog_schema.ddl`. The chains below are test +/// fixtures for the runner, and pinning them to the real schema would make +/// every future schema edit rewrite tests that are not about the schema. The +/// two table names are the real ones only because the parent/child rebuild +/// hazard is what several of these tests are about. +const syn_domains = + \\CREATE TABLE domains ( + \\ id INTEGER PRIMARY KEY, + \\ domain TEXT NOT NULL UNIQUE + \\) +; + +const syn_query_log_v1 = + \\CREATE TABLE query_log ( + \\ id INTEGER PRIMARY KEY, + \\ domain_id INTEGER NOT NULL REFERENCES domains(id), + \\ client_ip TEXT NOT NULL + \\) +; + +/// The verbatim text a v3 file must end up storing, whether it was created +/// fresh or rebuilt by step 2. That equality is the rebuild rule. +const syn_query_log_v3 = + \\CREATE TABLE query_log ( + \\ id INTEGER PRIMARY KEY, + \\ domain_id INTEGER NOT NULL REFERENCES domains(id), + \\ client_ip TEXT NOT NULL, + \\ qtype INTEGER, + \\ CHECK (qtype IS NULL OR qtype BETWEEN 0 AND 65535) + \\) +; + +const syn_index = "CREATE INDEX idx_query_log_domain ON query_log(domain_id)"; +const syn_notes = "CREATE TABLE notes (id INTEGER PRIMARY KEY, note TEXT NOT NULL)"; + +const syn_ddl_v1: [:0]const u8 = syn_domains ++ ";\n" ++ syn_query_log_v1 ++ ";\n" ++ syn_index ++ ";\n"; +const syn_ddl_v3: [:0]const u8 = + syn_domains ++ ";\n" ++ syn_query_log_v3 ++ ";\n" ++ syn_index ++ ";\n" ++ syn_notes ++ ";\n"; + +const syn_step_1_to_2: [:0]const u8 = syn_notes ++ ";\nINSERT INTO notes (id, note) VALUES (1, 'from step 1');\n"; + +/// The full A.2 rebuild sequence: rename, create verbatim, copy, drop, recreate +/// the dependent index. +const syn_step_2_to_3: [:0]const u8 = + \\ALTER TABLE query_log RENAME TO query_log_old; +++ "\n" ++ syn_query_log_v3 ++ ";\n" ++ + \\INSERT INTO query_log (id, domain_id, client_ip) + \\SELECT id, domain_id, client_ip FROM query_log_old; + \\DROP TABLE query_log_old; +++ "\n" ++ syn_index ++ ";\n"; + +const syn_chain: []const [:0]const u8 = &.{ syn_step_1_to_2, syn_step_2_to_3 }; + +/// A step that rebuilds the table the other one REFERENCES. With +/// `foreign_keys` left on, the rename would rewrite `query_log`'s stored FK +/// text to `domains_old`, and dropping `domains_old` would leave it dangling. +const syn_step_rebuild_parent: [:0]const u8 = + \\ALTER TABLE domains RENAME TO domains_old; +++ "\n" ++ syn_domains ++ ";\n" ++ + \\INSERT INTO domains (id, domain) SELECT id, domain FROM domains_old; + \\DROP TABLE domains_old; + \\ +; + +const Harness = struct { + threaded: std.Io.Threaded, + tmp: std.testing.TmpDir, + buf: [256]u8 = undefined, + + fn init() Harness { + return .{ + .threaded = .init(testing.allocator, .{}), + .tmp = testing.tmpDir(.{ .iterate = true }), + }; + } + + fn deinit(self: *Harness) void { + self.tmp.cleanup(); + self.threaded.deinit(); + } + + fn io(self: *Harness) std.Io { + return self.threaded.io(); + } + + fn path(self: *Harness) [:0]const u8 { + return std.fmt.bufPrintZ(&self.buf, ".zig-cache/tmp/{s}/querylog.db", .{self.tmp.sub_path}) catch + unreachable; + } + + /// A v1 file with two domains and two log rows, stamped and closed. + fn createV1(self: *Harness) !void { + var database = try db.Db.open(self.path(), .{ .mode = .read_write_create }); + defer database.close(); + try db.applyPragmas(&database, .{}); + try database.exec(syn_ddl_v1); + try database.exec( + \\INSERT INTO domains (id, domain) VALUES (1, 'a.example'), (2, 'b.example'); + \\INSERT INTO query_log (id, domain_id, client_ip) + \\VALUES (1, 1, '10.0.0.1'), (2, 2, '10.0.0.2'); + ); + try database.exec("PRAGMA user_version = 1;"); + } + + fn openLive(self: *Harness) !db.Db { + var database = try db.Db.open(self.path(), .{ .mode = .read_write_existing }); + errdefer database.close(); + try db.applyPragmas(&database, .{}); + return database; + } + + fn backupNames(self: *Harness, out: *std.ArrayList([]u8)) !void { + var it = self.tmp.dir.iterate(); + while (try it.next(self.io())) |entry| { + if (!std.mem.startsWith(u8, entry.name, "querylog.db" ++ backup_infix)) continue; + try out.append(testing.allocator, try testing.allocator.dupe(u8, entry.name)); + } + } +}; + +fn freeOwned(list: *std.ArrayList([]u8)) void { + for (list.items) |item| testing.allocator.free(item); + list.deinit(testing.allocator); +} + +fn schemaTextOf(database: *db.Db, name: []const u8) ![]const u8 { + var stmt = try database.prepare("SELECT sql FROM sqlite_schema WHERE name = ?1"); + defer stmt.deinit(); + try stmt.bindText(1, name); + if (!try stmt.step()) return error.NoSuchObject; + return testing.allocator.dupe(u8, stmt.columnText(0)); +} + +test "a synthetic chain migrates 1 to 3, keeps every row, and rebuilds verbatim" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + + try runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, syn_chain, 1, 3); + + try testing.expectEqual(@as(i64, 3), try database.queryInt("PRAGMA user_version")); + try testing.expectEqual(@as(i64, 2), try database.queryInt("SELECT count(*) FROM query_log")); + try testing.expectEqual(@as(i64, 2), try database.queryInt("SELECT count(*) FROM domains")); + try testing.expectEqual(@as(i64, 1), try database.queryInt("SELECT count(*) FROM notes")); + + // The rebuild rule, which is the whole reason the oracle's text layer can + // be exact: the migrated table stores the target's bytes, not SQLite's + // rewrite of the old ones. + const rebuilt = try schemaTextOf(&database, "query_log"); + defer testing.allocator.free(rebuilt); + try testing.expectEqualStrings(syn_query_log_v3, rebuilt); + + // Exactly one backup, and it is a real database holding the pre-migration + // state — still at version 1, and without the table step 1 added. + var names: std.ArrayList([]u8) = .empty; + defer freeOwned(&names); + try h.backupNames(&names); + try testing.expectEqual(@as(usize, 1), names.items.len); + + var backup_path_buf: [512]u8 = undefined; + const backup_path = try std.fmt.bufPrintZ( + &backup_path_buf, + ".zig-cache/tmp/{s}/{s}", + .{ h.tmp.sub_path, names.items[0] }, + ); + var backup = try db.Db.open(backup_path, .{ .mode = .read_only }); + defer backup.close(); + try testing.expectEqual(@as(i64, 1), try backup.queryInt("PRAGMA user_version")); + try testing.expectEqual(@as(i64, 2), try backup.queryInt("SELECT count(*) FROM query_log")); + try testing.expectEqual( + @as(i64, 0), + try backup.queryInt("SELECT count(*) FROM sqlite_schema WHERE name = 'notes'"), + ); + + // And it is equivalent to a file freshly created at the target schema. + var fresh = try db.Db.open(":memory:", .{ .mode = .memory }); + defer fresh.close(); + try fresh.exec(syn_ddl_v3); + try testing.expect(try schemaEquivalent(testing.allocator, &database, &fresh)); +} + +test "rebuilding a referenced parent leaves the child's foreign key intact" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + + const chain: []const [:0]const u8 = &.{syn_step_rebuild_parent}; + try runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, chain, 1, 2); + + // The hazard this exists for: with `foreign_keys` on, the rename would have + // rewritten this text to `REFERENCES "domains_old"(id)`, and the DROP two + // statements later would have left it pointing at nothing. + const child = try schemaTextOf(&database, "query_log"); + defer testing.allocator.free(child); + try testing.expectEqualStrings(syn_query_log_v1, child); + + var check = try database.prepare("PRAGMA foreign_key_check"); + defer check.deinit(); + try testing.expect(!try check.step()); + + try testing.expectEqual(@as(i64, 1), try database.queryInt("PRAGMA foreign_keys")); + try testing.expectEqual(@as(i64, 0), try database.queryInt("PRAGMA legacy_alter_table")); + try testing.expectEqual(@as(i64, 2), try database.queryInt("PRAGMA user_version")); +} + +test "a failed parent rebuild restores both pragmas and leaves the parent whole" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + + // The rebuild, then a statement that cannot run: the failure lands with + // `domains_old` already dropped inside the transaction, which is the worst + // moment for the rollback and the pragma restore to be wrong. + const failing: [:0]const u8 = syn_step_rebuild_parent ++ "UPDATE no_such_table SET x = 1;\n"; + + expected_failures.begin(); + defer expected_failures.end(); + try testing.expectError( + error.MigrationFailed, + runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, &.{failing}, 1, 2), + ); + + // Connection-global and non-transactional: the rollback cannot restore + // these, so the runner has to, on this path as much as on the happy one. + try testing.expectEqual(@as(i64, 1), try database.queryInt("PRAGMA foreign_keys")); + try testing.expectEqual(@as(i64, 0), try database.queryInt("PRAGMA legacy_alter_table")); + + try testing.expectEqual(@as(i64, 1), try database.queryInt("PRAGMA user_version")); + try testing.expectEqual(@as(i64, 2), try database.queryInt("SELECT count(*) FROM domains")); + const child = try schemaTextOf(&database, "query_log"); + defer testing.allocator.free(child); + try testing.expectEqualStrings(syn_query_log_v1, child); +} + +test "a mid-chain failure leaves the file untouched and deletes only this run's backup" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + // A backup from an earlier, successful migration. It is the operator's last + // copy and a failed run must not take it. + try h.tmp.dir.writeFile(h.io(), .{ + .sub_path = "querylog.db" ++ backup_infix ++ "1600000000", + .data = "an older backup", + }); + + var database = try h.openLive(); + defer database.close(); + + const broken: [:0]const u8 = "UPDATE no_such_table SET x = 1;"; + const chain: []const [:0]const u8 = &.{ syn_step_1_to_2, broken }; + + expected_failures.begin(); + defer expected_failures.end(); + try testing.expectError( + error.MigrationFailed, + runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, chain, 1, 3), + ); + + // Logically unchanged: step 1's table is gone with the rolled-back + // transaction, the stamp never moved, and the rows are all there. + try testing.expectEqual(@as(i64, 1), try database.queryInt("PRAGMA user_version")); + try testing.expectEqual(@as(i64, 2), try database.queryInt("SELECT count(*) FROM query_log")); + try testing.expectEqual( + @as(i64, 0), + try database.queryInt("SELECT count(*) FROM sqlite_schema WHERE name = 'notes'"), + ); + // The pragmas are back to what every connection to this file must have. + try testing.expectEqual(@as(i64, 1), try database.queryInt("PRAGMA foreign_keys")); + try testing.expectEqual(@as(i64, 0), try database.queryInt("PRAGMA legacy_alter_table")); + + var names: std.ArrayList([]u8) = .empty; + defer freeOwned(&names); + try h.backupNames(&names); + try testing.expectEqual(@as(usize, 1), names.items.len); + try testing.expectEqualStrings("querylog.db" ++ backup_infix ++ "1600000000", names.items[0]); +} + +test "a step cannot smuggle a transaction statement past the runner" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + + // The exact bypass a post-step autocommit check would miss: after this the + // connection IS inside a transaction again, so autocommit reads 0 and the + // step looks clean — while the work before it has been committed + // irreversibly. + const smuggler: [:0]const u8 = + \\CREATE TABLE half_done (id INTEGER PRIMARY KEY); + \\COMMIT; + \\BEGIN IMMEDIATE; + ; + + expected_failures.begin(); + defer expected_failures.end(); + + var tx = try db.Tx.begin(&database); + try testing.expectError( + error.TransactionViolation, + migrateSteps(&database, &.{smuggler}, 1, 2), + ); + + // The authorizer is gone, so the cleanup this failure requires can run at + // all. A leaked one would deny this ROLLBACK and strand the connection + // inside the migration transaction forever. + tx.rollback(); + try testing.expect(!database.inTransaction()); + try database.exec("BEGIN IMMEDIATE;"); + try database.exec("COMMIT;"); + + try testing.expectEqual(@as(i64, 1), try database.queryInt("PRAGMA user_version")); + try testing.expectEqual( + @as(i64, 0), + try database.queryInt("SELECT count(*) FROM sqlite_schema WHERE name = 'half_done'"), + ); +} + +test "migrateSteps stamps the target and refuses a chain that does not span the range" { + var database = try db.Db.open(":memory:", .{ .mode = .memory }); + defer database.close(); + try database.exec(syn_ddl_v1); + try database.exec("PRAGMA user_version = 1;"); + + var tx = try db.Tx.begin(&database); + try migrateSteps(&database, syn_chain, 1, 3); + try tx.commit(); + + try testing.expectEqual(@as(i64, 3), try database.queryInt("PRAGMA user_version")); +} + +test "PRAGMA user_version is transactional" { + var database = try db.Db.open(":memory:", .{ .mode = .memory }); + defer database.close(); + try database.exec("PRAGMA user_version = 41;"); + + var tx = try db.Tx.begin(&database); + try database.exec("PRAGMA user_version = 42;"); + try testing.expectEqual(@as(i64, 42), try database.queryInt("PRAGMA user_version")); + tx.rollback(); + + // The stamp rides the transaction, which is what lets the runner put the + // steps and the stamp in one unit and get all-or-nothing from SQLite. + try testing.expectEqual(@as(i64, 41), try database.queryInt("PRAGMA user_version")); +} + +test "the backup carries rows that are still only in the write-ahead log" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + + // Committed, but deliberately not checkpointed: a byte copy of the main + // file would miss this row entirely. + try database.exec("INSERT INTO domains (id, domain) VALUES (3, 'wal-only.example');"); + const wal = try h.tmp.dir.statFile(h.io(), "querylog.db-wal", .{}); + try testing.expect(wal.size > 0); + + const chain: []const [:0]const u8 = &.{syn_step_1_to_2}; + try runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, chain, 1, 2); + + var names: std.ArrayList([]u8) = .empty; + defer freeOwned(&names); + try h.backupNames(&names); + try testing.expectEqual(@as(usize, 1), names.items.len); + + var backup_path_buf: [512]u8 = undefined; + const backup_path = try std.fmt.bufPrintZ( + &backup_path_buf, + ".zig-cache/tmp/{s}/{s}", + .{ h.tmp.sub_path, names.items[0] }, + ); + var backup = try db.Db.open(backup_path, .{ .mode = .read_only }); + defer backup.close(); + try testing.expectEqual(@as(i64, 3), try backup.queryInt("SELECT count(*) FROM domains")); + var stmt = try backup.prepare("SELECT domain FROM domains WHERE id = 3"); + defer stmt.deinit(); + try testing.expect(try stmt.step()); + try testing.expectEqualStrings("wal-only.example", stmt.columnText(0)); +} + +test "two migrations leave exactly one backup, the newer" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + + try runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, &.{syn_step_1_to_2}, 1, 2); + var first: std.ArrayList([]u8) = .empty; + defer freeOwned(&first); + try h.backupNames(&first); + try testing.expectEqual(@as(usize, 1), first.items.len); + + try runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, &.{syn_step_2_to_3}, 2, 3); + var second: std.ArrayList([]u8) = .empty; + defer freeOwned(&second); + try h.backupNames(&second); + + // Step 4 knows the exact name it just made, so retention needs no clock + // comparison at all: everything else goes, even a same-second collision + // name from the previous run. + try testing.expectEqual(@as(usize, 1), second.items.len); + var older = try db.Db.open(blk: { + var b: [512]u8 = undefined; + break :blk try std.fmt.bufPrintZ(&b, ".zig-cache/tmp/{s}/{s}", .{ h.tmp.sub_path, second.items[0] }); + }, .{ .mode = .read_only }); + defer older.close(); + // The kept backup is the one from the SECOND migration: version 2, not 1. + try testing.expectEqual(@as(i64, 2), try older.queryInt("PRAGMA user_version")); +} + +test "the conservative pass keeps every backup tied at the newest epoch" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + const seeded = [_][]const u8{ + "querylog.db" ++ backup_infix ++ "1600000000", + "querylog.db" ++ backup_infix ++ "1700000000", + "querylog.db" ++ backup_infix ++ "1700000000-2", + // Not a name this code writes, so it is not this code's to delete. + "querylog.db" ++ backup_infix ++ "not-an-epoch", + }; + for (seeded) |name| { + try h.tmp.dir.writeFile(h.io(), .{ .sub_path = name, .data = "x" }); + } + + pruneBackupsConservative(h.io(), std.Io.Dir.cwd(), h.path()); + + var left: std.ArrayList([]u8) = .empty; + defer freeOwned(&left); + try h.backupNames(&left); + try testing.expectEqual(@as(usize, 3), left.items.len); + for (left.items) |name| { + try testing.expect(!std.mem.eql(u8, name, seeded[0])); + } +} + +test "backupEpoch parses only the names this code writes" { + const prefix = "querylog.db" ++ backup_infix; + try testing.expectEqual(@as(i64, 1700000000), backupEpoch(prefix ++ "1700000000", prefix).?); + try testing.expectEqual(@as(i64, 1700000000), backupEpoch(prefix ++ "1700000000-2", prefix).?); + try testing.expect(backupEpoch(prefix ++ "1700000000-", prefix) == null); + try testing.expect(backupEpoch(prefix ++ "1700000000-x", prefix) == null); + try testing.expect(backupEpoch(prefix ++ "not-an-epoch", prefix) == null); + try testing.expect(backupEpoch(prefix, prefix) == null); + try testing.expect(backupEpoch("querylog.db", prefix) == null); + try testing.expect(backupEpoch("querylog.db.corrupt-1700000000", prefix) == null); +} + +test "quoteLiteral doubles every quote so a path cannot end the literal" { + var buf: [64]u8 = undefined; + try testing.expectEqualStrings("/var/lib/nxdns", try quoteLiteral(&buf, "/var/lib/nxdns")); + try testing.expectEqualStrings("/o''brien/db", try quoteLiteral(&buf, "/o'brien/db")); + try testing.expectEqualStrings("''''", try quoteLiteral(&buf, "''")); + var tight: [3]u8 = undefined; + try testing.expectError(error.NoSpaceLeft, quoteLiteral(&tight, "a'b")); +} + +// --------------------------------------------------------------------------- +// the oracle +// --------------------------------------------------------------------------- + +fn openWith(sql: [:0]const u8) !db.Db { + var database = try db.Db.open(":memory:", .{ .mode = .memory }); + errdefer database.close(); + try database.exec(sql); + return database; +} + +fn expectDifference(a_sql: [:0]const u8, b_sql: [:0]const u8) !void { + var a = try openWith(a_sql); + defer a.close(); + var b = try openWith(b_sql); + defer b.close(); + try testing.expect(!try schemaEquivalent(testing.allocator, &a, &b)); +} + +const oracle_base: [:0]const u8 = + \\CREATE TABLE parent (id INTEGER PRIMARY KEY, name TEXT NOT NULL UNIQUE); + \\CREATE TABLE child ( + \\ id INTEGER PRIMARY KEY, + \\ parent_id INTEGER NOT NULL REFERENCES parent(id), + \\ rcode INTEGER NOT NULL, + \\ CHECK (rcode BETWEEN 0 AND 4095) + \\); + \\CREATE INDEX idx_child_parent ON child(parent_id) WHERE parent_id > 0; + \\CREATE VIEW child_view AS SELECT id, rcode FROM child; + \\CREATE TABLE bucket (a INTEGER NOT NULL, b TEXT NOT NULL, PRIMARY KEY (a, b)) WITHOUT ROWID; +; + +test "the oracle calls two fresh databases equal" { + var a = try openWith(oracle_base); + defer a.close(); + var b = try openWith(oracle_base); + defer b.close(); + try testing.expect(try schemaEquivalent(testing.allocator, &a, &b)); +} + +test "the oracle catches a dropped CHECK constraint" { + const without_check: [:0]const u8 = + \\CREATE TABLE parent (id INTEGER PRIMARY KEY, name TEXT NOT NULL UNIQUE); + \\CREATE TABLE child ( + \\ id INTEGER PRIMARY KEY, + \\ parent_id INTEGER NOT NULL REFERENCES parent(id), + \\ rcode INTEGER NOT NULL + \\); + \\CREATE INDEX idx_child_parent ON child(parent_id) WHERE parent_id > 0; + \\CREATE VIEW child_view AS SELECT id, rcode FROM child; + \\CREATE TABLE bucket (a INTEGER NOT NULL, b TEXT NOT NULL, PRIMARY KEY (a, b)) WITHOUT ROWID; + ; + try expectDifference(oracle_base, without_check); +} + +test "the oracle catches a dropped foreign key" { + const without_fk: [:0]const u8 = + \\CREATE TABLE parent (id INTEGER PRIMARY KEY, name TEXT NOT NULL UNIQUE); + \\CREATE TABLE child ( + \\ id INTEGER PRIMARY KEY, + \\ parent_id INTEGER NOT NULL, + \\ rcode INTEGER NOT NULL, + \\ CHECK (rcode BETWEEN 0 AND 4095) + \\); + \\CREATE INDEX idx_child_parent ON child(parent_id) WHERE parent_id > 0; + \\CREATE VIEW child_view AS SELECT id, rcode FROM child; + \\CREATE TABLE bucket (a INTEGER NOT NULL, b TEXT NOT NULL, PRIMARY KEY (a, b)) WITHOUT ROWID; + ; + try expectDifference(oracle_base, without_fk); +} + +test "the oracle catches a dropped WITHOUT ROWID" { + const with_rowid: [:0]const u8 = + \\CREATE TABLE parent (id INTEGER PRIMARY KEY, name TEXT NOT NULL UNIQUE); + \\CREATE TABLE child ( + \\ id INTEGER PRIMARY KEY, + \\ parent_id INTEGER NOT NULL REFERENCES parent(id), + \\ rcode INTEGER NOT NULL, + \\ CHECK (rcode BETWEEN 0 AND 4095) + \\); + \\CREATE INDEX idx_child_parent ON child(parent_id) WHERE parent_id > 0; + \\CREATE VIEW child_view AS SELECT id, rcode FROM child; + \\CREATE TABLE bucket (a INTEGER NOT NULL, b TEXT NOT NULL, PRIMARY KEY (a, b)); + ; + try expectDifference(oracle_base, with_rowid); +} + +test "the oracle catches an added column" { + const extra: [:0]const u8 = oracle_base ++ "ALTER TABLE bucket ADD COLUMN c INTEGER;"; + try expectDifference(oracle_base, extra); +} + +test "the oracle catches a rename that skipped the verbatim-text rule" { + // The rebuild rule exists because SQLite rewrites the stored CREATE text + // under a rename. A round trip back to the original NAME does not restore + // the original BYTES, and the exact-text layer is what sees that. + const renamed: [:0]const u8 = oracle_base ++ + \\ALTER TABLE child RENAME TO child_tmp; + \\ALTER TABLE child_tmp RENAME TO child; + ; + + var a = try openWith(oracle_base); + defer a.close(); + var b = try openWith(renamed); + defer b.close(); + + const original = try schemaTextOf(&a, "child"); + defer testing.allocator.free(original); + const rewritten = try schemaTextOf(&b, "child"); + defer testing.allocator.free(rewritten); + + // The premise, proved rather than assumed. + try testing.expect(!std.mem.eql(u8, original, rewritten)); + try testing.expect(!try schemaEquivalent(testing.allocator, &a, &b)); +} + +test "the oracle calls a migrated file equal to a fresh one at the same version" { + var h: Harness = .init(); + defer h.deinit(); + try h.createV1(); + + var database = try h.openLive(); + defer database.close(); + try runMigration(h.io(), std.Io.Dir.cwd(), h.path(), &database, syn_chain, 1, 3); + + var fresh = try openWith(syn_ddl_v3); + defer fresh.close(); + try testing.expect(try schemaEquivalent(testing.allocator, &database, &fresh)); + + // And it is not simply calling everything equal. + var stale = try openWith(syn_ddl_v1); + defer stale.close(); + try testing.expect(!try schemaEquivalent(testing.allocator, &database, &stale)); +} diff --git a/src/storage/querylog_schema.zig b/src/storage/querylog_schema.zig index 0bddcf7..e4723eb 100644 --- a/src/storage/querylog_schema.zig +++ b/src/storage/querylog_schema.zig @@ -1,14 +1,15 @@ -//! The `querylog.db` schema and its open-or-recreate policy. +//! The `querylog.db` schema and its open policy. //! -//! `querylog.db` is never migrated (PLAN §3.7). It holds expendable log rows, -//! so a schema change replaces the file instead of upgrading it. The -//! replacement trigger is a fingerprint derived from the DDL text itself, so -//! editing the schema below automatically invalidates every existing file — the -//! policy cannot drift out of sync with the SQL. +//! `querylog.db` carries a logical schema version in `PRAGMA user_version` +//! (`querylog_versions.zig`), and a file stamped below the current version is +//! MIGRATED in place. A healthy file is never replaced and never set aside: a +//! version this build cannot reach refuses the startup with instructions +//! instead, because the operator's query history is not this program's to +//! discard. //! //! **Recreating is destructive, so the predicate is a positive whitelist.** Only -//! a missing file, `error.Corrupt`, `error.NotADb`, a failed `PRAGMA -//! quick_check` and a fingerprint mismatch recreate. Every other error +//! a missing file, `error.Corrupt`, `error.NotADb` and a failed `PRAGMA +//! quick_check` recreate — genuine corruption, nothing else. Every other error //! propagates and the file on disk is not touched. `error.Busy` / `error.Locked` //! mean another process holds the write lock — waiting is right, deleting is //! catastrophic. `error.OutOfMemory` is this process's problem. `error.CantOpen` @@ -19,9 +20,21 @@ const std = @import("std"); const db = @import("db.zig"); +const migrations = @import("querylog_migrations.zig"); +const versions = @import("querylog_versions.zig"); const log = std.log.scoped(.querylog_schema); +/// See `querylog_migrations.fail`: `err`, unless a test has said it is causing +/// this refusal on purpose. +fn fail(comptime fmt: []const u8, args: anytype) void { + if (migrations.expected_failures.capturing()) { + log.warn(fmt, args); + } else { + log.err(fmt, args); + } +} + /// PLAN §11.3, plus the coverage watermark of milestone 28. Multi-statement /// text — it goes through `db.Db.exec`, never through `prepare`. /// @@ -139,14 +152,21 @@ pub const fingerprint: i32 = blk: { break :blk fingerprintOf(ddl); }; -const set_user_version = std.fmt.comptimePrint("PRAGMA user_version = {d};", .{fingerprint}); +/// What a fresh file is stamped with. The logical version, not the fingerprint: +/// from this release on, `user_version` is a version number. +const set_user_version = std.fmt.comptimePrint( + "PRAGMA user_version = {d};", + .{versions.current_version}, +); /// Long enough for any path this program will be handed, plus the aside suffix. /// A longer path is `error.NameTooLong`, which is what the filesystem calls /// would have returned anyway. const path_buf_len = 4096 + 64; -pub const RecreateReason = enum { missing, corrupt, not_a_database, quick_check_failed, fingerprint_mismatch }; +/// Corruption, and nothing else. A healthy file with a version this build does +/// not handle refuses the startup; it is never recreated and never set aside. +pub const RecreateReason = enum { missing, corrupt, not_a_database, quick_check_failed }; pub const OpenResult = struct { database: db.Db, @@ -166,10 +186,75 @@ pub const OpenResult = struct { } }; -pub const Error = db.Error || error{AsideNameCollision} || - std.Io.Dir.RenamePreserveError || std.Io.Dir.DeleteFileError || std.Io.Dir.AccessError; +pub const Error = db.Error || error{ + AsideNameCollision, + /// The file's version is above this build's. A downgrade, almost always. + SchemaTooNew, + /// The file's version is one this build cannot migrate from: older than + /// `minimum_supported_version`, or not a stamp nxdns ever wrote. + SchemaUnsupported, + MigrationFailed, + MigrationBackupFailed, + NameTooLong, +} || std.Io.Dir.RenamePreserveError || std.Io.Dir.DeleteFileError || std.Io.Dir.AccessError; -/// Opens `path`, recreating it if and only if it is genuinely unusable. +/// The version metadata `openVersioned` works against. Production passes +/// `production_plan`; tests inject synthetic chains, which is what makes +/// migration, the post-commit branch and the ownership of the handle testable +/// through the real open path while the shipped chain is still empty. +pub const Plan = struct { + minimum: i32, + current: i32, + legacy_fingerprint: i32, + step_sql: []const [:0]const u8, +}; + +pub const production_plan: Plan = .{ + .minimum = versions.minimum_supported_version, + .current = versions.current_version, + .legacy_fingerprint = versions.legacy_fingerprint, + .step_sql = versions.step_sql, +}; + +/// What a stamped `user_version` means. A pure function of the stamp and the +/// plan's three numbers — no file, no clock, no mutation. +pub const Action = enum { open_current, migrate, refuse_too_new, refuse_unsupported }; + +pub const Classification = struct { + /// The stamp mapped onto the version line. Equal to the stamp except for + /// the legacy fingerprint, which IS version 1. + logical: i32, + action: Action, + /// The legacy fingerprint must be replaced by its logical number before the + /// file is used — but only on a lane that accepts the file. An unsupported + /// file is never modified. + restamp: bool, +}; + +/// The classification table. The order is load-bearing: the legacy fingerprint +/// becomes version 1 FIRST, and only then is version 1 judged against the +/// plan's range. After a future explicit break raises the minimum above 1, a +/// legacy-stamped file therefore classifies as below-minimum and refuses +/// without ever being restamped. +pub fn classify(stamped: i32, plan: Plan) Classification { + const legacy = stamped == plan.legacy_fingerprint; + const logical: i32 = if (legacy) 1 else stamped; + + const action: Action = if (logical == plan.current) + .open_current + else if (logical >= plan.minimum and logical < plan.current) + .migrate + else if (logical > plan.current and logical <= versions.version_floor_guard) + .refuse_too_new + else + .refuse_unsupported; + + const accepted = action == .open_current or action == .migrate; + return .{ .logical = logical, .action = action, .restamp = legacy and accepted }; +} + +/// Opens `path`, recreating it if and only if it is genuinely unusable, and +/// migrating it if and only if it carries an older supported version. /// /// `path` is resolved twice by two different mechanisms: `dir`-relative for the /// filesystem calls, and process-cwd-relative by SQLite's VFS, which knows @@ -180,7 +265,7 @@ pub fn open(io: std.Io, dir: std.Io.Dir, path: [:0]const u8) Error!OpenResult { var handle: ?db.Db = null; errdefer if (handle) |*h| h.close(); - const reason: ?RecreateReason = probe: { + const cause: RecreateReason = probe: { dir.access(io, path, .{}) catch |e| switch (e) { error.FileNotFound => break :probe .missing, else => |other| return other, @@ -197,15 +282,10 @@ pub fn open(io: std.Io, dir: std.Io.Dir, path: [:0]const u8) Error!OpenResult { break :probe recreatable(e) orelse return e; if (!healthy) break :probe .quick_check_failed; - const stamped = opened.queryInt("PRAGMA user_version") catch |e| - break :probe recreatable(e) orelse return e; - if (stamped != fingerprint) break :probe .fingerprint_mismatch; - - break :probe null; + try openVersioned(io, dir, path, &handle, production_plan); + return .{ .database = handle.?, .recreated = null }; }; - const cause = reason orelse return .{ .database = handle.?, .recreated = null }; - // Close first, so SQLite checkpoints and drops `-wal`/`-shm` where it can. if (handle) |*h| h.close(); handle = null; @@ -240,6 +320,131 @@ pub fn open(io: std.Io, dir: std.Io.Dir, path: [:0]const u8) Error!OpenResult { return result; } +/// The version half of `open`, against an injectable `plan`. +/// +/// `handle` is the slot holding the healthy, pragma-applied connection to +/// `path`. On success the connection stays in it, at `plan.current`. On every +/// failure this function closes the connection and sets the slot to null, so +/// the caller's own error-path close cannot double-close it — including the +/// committed-but-unclean path, which is the one place the handle must be +/// dropped even though the file on disk is fine. +/// +/// **nxdns owns `path` exclusively.** It opens `querylog.db` once at startup, +/// before it serves anything, and no second process shares a data directory — +/// the standing deployment contract. The backup-then-lock sequence in +/// `querylog_migrations.runMigration` relies on it: between the `VACUUM INTO` +/// and the `BEGIN IMMEDIATE` there is no lock, and the re-read of +/// `user_version` under the lock is what turns a violation of that contract +/// into a refusal instead of a corrupted migration. +pub fn openVersioned( + io: std.Io, + dir: std.Io.Dir, + path: [:0]const u8, + handle: *?db.Db, + plan: Plan, +) Error!void { + const database = &(handle.*.?); + errdefer closeSlot(handle); + + const stamped64 = try database.queryInt("PRAGMA user_version"); + const stamped = std.math.cast(i32, stamped64) orelse { + // `user_version` is a signed 32-bit field, so this cannot come from + // SQLite. Refusing is the same answer any other foreign stamp gets. + refusalLog(path, stamped64, plan, "SchemaUnsupported"); + return error.SchemaUnsupported; + }; + + const verdict = classify(stamped, plan); + switch (verdict.action) { + .refuse_too_new => { + refusalLog(path, stamped64, plan, "SchemaTooNew"); + return error.SchemaTooNew; + }, + .refuse_unsupported => { + refusalLog(path, stamped64, plan, "SchemaUnsupported"); + return error.SchemaUnsupported; + }, + .open_current, .migrate => {}, + } + + if (verdict.restamp) try restampLegacy(database, path, verdict.logical); + + switch (verdict.action) { + .migrate => { + const first = @as(usize, @intCast(verdict.logical - plan.minimum)); + migrations.runMigration( + io, + dir, + path, + database, + plan.step_sql[first..], + verdict.logical, + plan.current, + ) catch |e| switch (e) { + // Two different states of the FILE — migrated and kept with its + // backup, or logically untouched — and one shared state of the + // CONNECTION: its pragmas are not what `applyPragmas` + // guarantees, so it must not serve. Closing it here is the + // single close either path gets. After a commit the next start + // opens the migrated file on the current-version lane; after a + // failure it retries the migration from the top. + error.MigrationCommittedButUnclean, error.MigrationFailedUnclean => { + closeSlot(handle); + return error.MigrationFailed; + }, + error.MigrationBackupFailed => return error.MigrationBackupFailed, + error.MigrationFailed, error.NameTooLong => return error.MigrationFailed, + else => |other| return other, + }; + }, + .open_current => migrations.pruneBackupsConservative(io, dir, path), + else => unreachable, + } +} + +/// Replaces the 0.0.12/0.0.13 fingerprint stamp with the logical version it +/// stands for. This is the milestone's only real mutation of operator data, so +/// it runs in its own transaction and any failure leaves the legacy stamp and +/// every row exactly as they were — a refusal, never a recreate. +fn restampLegacy(database: *db.Db, path: []const u8, logical: i32) Error!void { + var stamp_buf: [64]u8 = undefined; + const stamp = std.fmt.bufPrintZ(&stamp_buf, "PRAGMA user_version = {d};", .{logical}) catch + unreachable; // an i32 and a fixed prefix cannot overrun 64 bytes + + restamp: { + var tx = db.Tx.begin(database) catch break :restamp; + database.exec(stamp) catch { + tx.rollback(); + break :restamp; + }; + tx.commit() catch { + tx.rollback(); + break :restamp; + }; + log.info("querylog database '{s}' carried the 0.0.12 schema fingerprint; " ++ + "restamped as schema version {d}", .{ path, logical }); + return; + } + + var buf: [256]u8 = undefined; + fail("cannot restamp querylog database '{s}' as schema version {d}: {s}; " ++ + "the file is unchanged", .{ path, logical, database.lastError(&buf) }); + return error.MigrationFailed; +} + +fn closeSlot(handle: *?db.Db) void { + if (handle.*) |*h| h.close(); + handle.* = null; +} + +fn refusalLog(path: []const u8, stamped: i64, plan: Plan, name: []const u8) void { + fail("refusing to open querylog database '{s}': it is stamped {d}, and this build " ++ + "supports schema versions {d} to {d} ({s}). The file is left exactly as it is; " ++ + "see docs/how-to/troubleshoot.md, \"The server refuses to start over querylog.db\"", .{ + path, stamped, plan.minimum, plan.current, name, + }); +} + /// An additional connection to a `querylog.db` that `open` has already /// established, with the pragmas every connection to the file needs. /// @@ -285,16 +490,15 @@ fn quickCheck(database: *db.Db) db.Error!bool { /// What the aside file's name calls the reason it was set aside. /// /// The name is the only account of the reason an operator gets: the log line -/// naming it scrolls away, the file stays for months. `fingerprint_mismatch` is -/// a database with nothing wrong with it — this build's DDL moved — so calling -/// its file "corrupt" invites the operator to delete evidence of a healthy file. +/// naming it scrolls away, the file stays for months. Every tag here names real +/// damage, which is the whole set of reasons left — a healthy file whose +/// version this build cannot handle refuses the startup and is not renamed. fn asideTag(reason: RecreateReason) []const u8 { return switch (reason) { .missing => unreachable, // there is no file to rename .corrupt => "corrupt", .not_a_database => "not-a-database", .quick_check_failed => "quick-check-failed", - .fingerprint_mismatch => "schema-changed", }; } @@ -455,11 +659,14 @@ test "querylog_meta is seeded with one row the schema will not let a second join try testing.expectEqual(@as(i64, 1), try database.queryInt("SELECT count(*) FROM querylog_meta")); } -test "the user_version statement stamps the fingerprint" { +test "the user_version statement stamps the current schema version" { var database = try db.Db.open(":memory:", .{ .mode = .memory }); defer database.close(); try database.exec(set_user_version); - try testing.expectEqual(@as(i64, fingerprint), try database.queryInt("PRAGMA user_version")); + try testing.expectEqual( + @as(i64, versions.current_version), + try database.queryInt("PRAGMA user_version"), + ); } // The behaviour these two cases describe — a resource error leaves the file on @@ -488,7 +695,6 @@ test "the aside name says why, and a healthy file is never called corrupt" { try testing.expectEqualStrings("corrupt", asideTag(.corrupt)); try testing.expectEqualStrings("not-a-database", asideTag(.not_a_database)); try testing.expectEqualStrings("quick-check-failed", asideTag(.quick_check_failed)); - try testing.expectEqualStrings("schema-changed", asideTag(.fingerprint_mismatch)); } test "recreatable selects exactly two of db.Error's members" { @@ -532,7 +738,7 @@ test "a recreate returns the aside name by value and a fresh create returns none try tmp.dir.access(io, kept, .{}); } -test "a recreate resets coverage to the new file and keeps the old one aside" { +test "a corrupt file is recreated, coverage restarts, and the old one is kept aside" { var threaded: std.Io.Threaded = .init(testing.allocator, .{}); defer threaded.deinit(); const io = threaded.io(); @@ -549,21 +755,14 @@ test "a recreate resets coverage to the new file and keeps the old one aside" { try created.database.exec("INSERT INTO domains (domain) VALUES ('old.example');"); created.database.close(); - // A healthy file this build's DDL no longer matches — the case milestone - // 28's own schema edit produces on every upgrade. - { - var stamped = try db.Db.open(path, .{ .mode = .read_write_existing }); - defer stamped.close(); - var sql_buf: [64]u8 = undefined; - try stamped.exec(try std.fmt.bufPrintZ(&sql_buf, "PRAGMA user_version = {d};", .{fingerprint +% 1})); - } + // Real damage, which is now the only thing that recreates. + try tmp.dir.writeFile(io, .{ .sub_path = "querylog.db", .data = "not a database at all" }); var recreated = try open(io, std.Io.Dir.cwd(), path); defer recreated.database.close(); - try testing.expectEqual(RecreateReason.fingerprint_mismatch, recreated.recreated.?); - // The name says the file was healthy and this build moved, not that it rotted. - try testing.expect(std.mem.indexOf(u8, recreated.aside(), ".schema-changed-") != null); + try testing.expectEqual(RecreateReason.not_a_database, recreated.recreated.?); + try testing.expect(std.mem.indexOf(u8, recreated.aside(), ".not-a-database-") != null); try tmp.dir.access(io, std.fs.path.basename(recreated.aside()), .{}); // Exactly one meta row, and coverage starts at the recreate rather than @@ -602,3 +801,486 @@ test "a clean reopen reports no recreate and no aside" { try testing.expectEqual(@as(?RecreateReason, null), second.recreated); try testing.expectEqualStrings("", second.aside()); } + +test "classification is a pure function of the stamp and the plan's three numbers" { + const plan: Plan = .{ + .minimum = 1, + .current = 3, + .legacy_fingerprint = versions.legacy_fingerprint, + .step_sql = &.{}, + }; + + try testing.expectEqual(Action.open_current, classify(3, plan).action); + try testing.expectEqual(Action.migrate, classify(1, plan).action); + try testing.expectEqual(Action.migrate, classify(2, plan).action); + try testing.expectEqual(Action.refuse_too_new, classify(4, plan).action); + try testing.expectEqual(Action.refuse_too_new, classify(versions.version_floor_guard, plan).action); + // Above the floor guard is not a version this project ever wrote. + try testing.expectEqual( + Action.refuse_unsupported, + classify(versions.version_floor_guard + 1, plan).action, + ); + for ([_]i32{ 0, -1, -1_000_000, 603440875 }) |foreign| { + try testing.expectEqual(Action.refuse_unsupported, classify(foreign, plan).action); + try testing.expect(!classify(foreign, plan).restamp); + } + + // The legacy fingerprint IS version 1, and being version 1 is what decides + // its lane. + const legacy = classify(versions.legacy_fingerprint, plan); + try testing.expectEqual(@as(i32, 1), legacy.logical); + try testing.expectEqual(Action.migrate, legacy.action); + try testing.expect(legacy.restamp); +} + +test "a legacy stamp below a raised minimum refuses without a restamp" { + // What a future explicit break looks like from this side: the minimum has + // moved past 1, so the 0.0.12 file is no longer reachable. The ORDER is the + // point — mapping to 1 first and judging second is what stops the restamp + // from mutating a file this build will refuse anyway. + const after_break: Plan = .{ + .minimum = 3, + .current = 3, + .legacy_fingerprint = versions.legacy_fingerprint, + .step_sql = &.{}, + }; + + const legacy = classify(versions.legacy_fingerprint, after_break); + try testing.expectEqual(@as(i32, 1), legacy.logical); + try testing.expectEqual(Action.refuse_unsupported, legacy.action); + try testing.expect(!legacy.restamp); + + // And versions 1 and 2, which the break dropped, refuse the same way. + try testing.expectEqual(Action.refuse_unsupported, classify(1, after_break).action); + try testing.expectEqual(Action.refuse_unsupported, classify(2, after_break).action); + try testing.expectEqual(Action.open_current, classify(3, after_break).action); +} + +test "the production plan classifies a fresh stamp as current" { + try testing.expectEqual( + Action.open_current, + classify(versions.current_version, production_plan).action, + ); + try testing.expect(classify(versions.legacy_fingerprint, production_plan).restamp); +} + +/// The five lines every file-backed test below opens with. +const Fixture = struct { + threaded: std.Io.Threaded, + tmp: std.testing.TmpDir, + buf: [256]u8 = undefined, + + fn init() Fixture { + return .{ + .threaded = .init(testing.allocator, .{}), + .tmp = testing.tmpDir(.{ .iterate = true }), + }; + } + + fn deinit(self: *Fixture) void { + self.tmp.cleanup(); + self.threaded.deinit(); + } + + fn io(self: *Fixture) std.Io { + return self.threaded.io(); + } + + fn path(self: *Fixture) [:0]const u8 { + return std.fmt.bufPrintZ(&self.buf, ".zig-cache/tmp/{s}/querylog.db", .{self.tmp.sub_path}) catch + unreachable; + } + + fn stamp(self: *Fixture, value: i32) !void { + var database = try db.Db.open(self.path(), .{ .mode = .read_write_existing }); + defer database.close(); + var sql: [64]u8 = undefined; + try database.exec(try std.fmt.bufPrintZ(&sql, "PRAGMA user_version = {d};", .{value})); + } + + fn liveHandle(self: *Fixture) !?db.Db { + var database = try db.Db.open(self.path(), .{ .mode = .read_write_existing }); + errdefer database.close(); + try db.applyPragmas(&database, .{}); + return database; + } + + fn countMatching(self: *Fixture, prefix: []const u8) !usize { + var found: usize = 0; + var it = self.tmp.dir.iterate(); + while (try it.next(self.io())) |entry| { + if (std.mem.startsWith(u8, entry.name, prefix)) found += 1; + } + return found; + } + + fn expectModeOfOnlyMatch(self: *Fixture, prefix: []const u8, expected: std.posix.mode_t) !void { + var it = self.tmp.dir.iterate(); + while (try it.next(self.io())) |entry| { + if (!std.mem.startsWith(u8, entry.name, prefix)) continue; + const stat = try self.tmp.dir.statFile(self.io(), entry.name, .{}); + const mode = stat.permissions.toMode() & 0o777; + if (mode != expected) { + std.debug.print("mode of '{s}' is {o}, expected {o}\n", .{ entry.name, mode, expected }); + return error.TestUnexpectedResult; + } + return; + } + std.debug.print("no file starting with '{s}'\n", .{prefix}); + return error.TestUnexpectedResult; + } +}; + +test "a fresh file is stamped with the current schema version" { + var f: Fixture = .init(); + defer f.deinit(); + + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer created.database.close(); + try testing.expectEqual(RecreateReason.missing, created.recreated.?); + try testing.expectEqual( + @as(i64, versions.current_version), + try created.database.queryInt("PRAGMA user_version"), + ); +} + +test "a 0.0.13 file is restamped as version 1 and keeps every row" { + var f: Fixture = .init(); + defer f.deinit(); + + { + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer created.database.close(); + try created.database.exec("INSERT INTO domains (domain) VALUES ('kept.example');"); + } + // Exactly what 0.0.12 and 0.0.13 wrote: the CRC of their DDL, which is this + // build's DDL unchanged. + try f.stamp(versions.legacy_fingerprint); + try testing.expectEqual(versions.legacy_fingerprint, fingerprint); + + { + var upgraded = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer upgraded.database.close(); + try testing.expectEqual(@as(?RecreateReason, null), upgraded.recreated); + try testing.expectEqual(@as(i64, 1), try upgraded.database.queryInt("PRAGMA user_version")); + try testing.expectEqual( + @as(i64, 1), + try upgraded.database.queryInt("SELECT count(*) FROM domains WHERE domain = 'kept.example'"), + ); + } + + // Nothing was set aside on the way, and the second start is an ordinary + // current-version open. + try testing.expectEqual(@as(usize, 0), try f.countMatching("querylog.db.schema")); + var again = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer again.database.close(); + try testing.expectEqual(@as(?RecreateReason, null), again.recreated); + try testing.expectEqual(@as(i64, 1), try again.database.queryInt("PRAGMA user_version")); +} + +test "a version this build cannot handle refuses and leaves the file alone" { + var f: Fixture = .init(); + defer f.deinit(); + + { + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer created.database.close(); + try created.database.exec("INSERT INTO domains (domain) VALUES ('kept.example');"); + } + const watermark = blk: { + var probe = try db.Db.open(f.path(), .{ .mode = .read_write_existing }); + defer probe.close(); + break :blk try probe.queryInt("SELECT available_since FROM querylog_meta"); + }; + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + + const lanes = [_]struct { stamp: i32, expected: anyerror }{ + .{ .stamp = versions.current_version + 1, .expected = error.SchemaTooNew }, + .{ .stamp = versions.version_floor_guard, .expected = error.SchemaTooNew }, + .{ .stamp = 0, .expected = error.SchemaUnsupported }, + .{ .stamp = -3, .expected = error.SchemaUnsupported }, + .{ .stamp = 603440875, .expected = error.SchemaUnsupported }, + }; + + for (lanes) |lane| { + try f.stamp(lane.stamp); + try testing.expectError(lane.expected, open(f.io(), std.Io.Dir.cwd(), f.path())); + + // Schema, rows, watermark and stamp all as they were, and nothing new + // beside the file. + var probe = try db.Db.open(f.path(), .{ .mode = .read_write_existing }); + defer probe.close(); + try testing.expectEqual(@as(i64, lane.stamp), try probe.queryInt("PRAGMA user_version")); + try testing.expectEqual( + @as(i64, 1), + try probe.queryInt("SELECT count(*) FROM domains WHERE domain = 'kept.example'"), + ); + try testing.expectEqual( + watermark, + try probe.queryInt("SELECT available_since FROM querylog_meta"), + ); + try testing.expectEqual(@as(usize, 0), try f.countMatching("querylog.db.")); + } +} + +test "a restamp that fails at the statement or at the commit refuses without loss" { + for ([_][]const u8{ "PRAGMA user_version = 1;", "COMMIT;" }) |failing| { + var f: Fixture = .init(); + defer f.deinit(); + + { + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer created.database.close(); + try created.database.exec("INSERT INTO domains (domain) VALUES ('kept.example');"); + } + try f.stamp(versions.legacy_fingerprint); + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + db.exec_faults.failNextMatching(failing); + defer db.exec_faults.disarm(); + + try testing.expectError( + error.MigrationFailed, + open(f.io(), std.Io.Dir.cwd(), f.path()), + ); + try testing.expect(!db.exec_faults.armed()); + + // The legacy stamp and every row are exactly as they were: this is a + // refusal, and a refusal never costs the operator anything. + { + var probe = try db.Db.open(f.path(), .{ .mode = .read_write_existing }); + defer probe.close(); + try testing.expectEqual( + @as(i64, versions.legacy_fingerprint), + try probe.queryInt("PRAGMA user_version"), + ); + } + + // And the next start, with nothing injected, does the restamp properly. + var recovered = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer recovered.database.close(); + try testing.expectEqual(@as(i64, 1), try recovered.database.queryInt("PRAGMA user_version")); + try testing.expectEqual( + @as(i64, 1), + try recovered.database.queryInt("SELECT count(*) FROM domains WHERE domain = 'kept.example'"), + ); + } +} + +/// A one-step chain from the real current version to one above it. Nothing in +/// the shipped chain can exercise migration while `step_sql` is empty, so the +/// open path's migration lanes are driven through this instead. +const synthetic_plan: Plan = .{ + .minimum = versions.current_version, + .current = versions.current_version + 1, + .legacy_fingerprint = versions.legacy_fingerprint, + .step_sql = &.{"CREATE TABLE migration_marker (id INTEGER PRIMARY KEY);"}, +}; + +test "a post-commit failure keeps the migration, keeps the backup, and refuses once" { + var f: Fixture = .init(); + defer f.deinit(); + + { + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer created.database.close(); + try created.database.exec("INSERT INTO domains (domain) VALUES ('kept.example');"); + } + + { + var handle = try f.liveHandle(); + errdefer if (handle) |*h| h.close(); + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + // The first statement of the pragma restore, which runs only after a + // successful COMMIT. + db.exec_faults.failNextMatching("legacy_alter_table = OFF"); + defer db.exec_faults.disarm(); + + try testing.expectError( + error.MigrationFailed, + openVersioned(f.io(), std.Io.Dir.cwd(), f.path(), &handle, synthetic_plan), + ); + // Closed exactly once, by the open path: the slot it was handed is + // empty, so no caller can close it again. + try testing.expect(handle == null); + } + + // The file IS migrated. The log said so, and this is what it meant. + { + var probe = try db.Db.open(f.path(), .{ .mode = .read_write_existing }); + defer probe.close(); + try testing.expectEqual( + @as(i64, synthetic_plan.current), + try probe.queryInt("PRAGMA user_version"), + ); + try testing.expectEqual( + @as(i64, 1), + try probe.queryInt("SELECT count(*) FROM sqlite_schema WHERE name = 'migration_marker'"), + ); + try testing.expectEqual( + @as(i64, 1), + try probe.queryInt("SELECT count(*) FROM domains WHERE domain = 'kept.example'"), + ); + } + try testing.expectEqual(@as(usize, 1), try f.countMatching("querylog.db.pre-migrate-")); + // The backup is the whole query history at the moment of the migration, so + // it carries the live file's mode and not SQLite's `0644 & ~umask`. + try f.expectModeOfOnlyMatch("querylog.db.pre-migrate-", 0o600); + + // The next start is ordinary: the current-version lane, no second + // migration, and the one backup still there for the operator. + var handle = try f.liveHandle(); + defer if (handle) |*h| h.close(); + try openVersioned(f.io(), std.Io.Dir.cwd(), f.path(), &handle, synthetic_plan); + try testing.expect(handle != null); + try testing.expectEqual( + @as(i64, synthetic_plan.current), + try handle.?.queryInt("PRAGMA user_version"), + ); + try testing.expectEqual(@as(usize, 1), try f.countMatching("querylog.db.pre-migrate-")); +} + +/// The same shape as `synthetic_plan`, with a step SQLite refuses to prepare. It +/// drives the pre-commit failure path without any fault seam, leaving the seam +/// free for the pragma restore. +const failing_plan: Plan = .{ + .minimum = versions.current_version, + .current = versions.current_version + 1, + .legacy_fingerprint = versions.legacy_fingerprint, + .step_sql = &.{"CREATE TABLE migration_marker (id INTEGER PRIMARY KEY) NOT A STATEMENT;"}, +}; + +test "a pre-commit failure whose pragma restore also fails closes the connection" { + var f: Fixture = .init(); + defer f.deinit(); + + { + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer created.database.close(); + try created.database.exec("INSERT INTO domains (domain) VALUES ('kept.example');"); + } + + // The runner's own answer first: a failed restore is a DIFFERENT error from + // a failed migration, because the two leave the connection in different + // states even though they leave the file in the same one. + { + var handle = try f.liveHandle(); + defer if (handle) |*h| h.close(); + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + db.exec_faults.failNextMatching("legacy_alter_table = OFF"); + defer db.exec_faults.disarm(); + + try testing.expectError(error.MigrationFailedUnclean, migrations.runMigration( + f.io(), + std.Io.Dir.cwd(), + f.path(), + &handle.?, + failing_plan.step_sql, + failing_plan.minimum, + failing_plan.current, + )); + try testing.expect(!db.exec_faults.armed()); + } + + // And the open path's answer: the handle is closed, exactly as it is after a + // post-commit restore failure. A connection that may still hold + // `foreign_keys = OFF` never reaches the server. + { + var handle = try f.liveHandle(); + errdefer if (handle) |*h| h.close(); + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + db.exec_faults.failNextMatching("legacy_alter_table = OFF"); + defer db.exec_faults.disarm(); + + try testing.expectError( + error.MigrationFailed, + openVersioned(f.io(), std.Io.Dir.cwd(), f.path(), &handle, failing_plan), + ); + try testing.expect(handle == null); + } + + // The contrast that makes the rule visible, back at the runner, where the + // connection survives to be inspected: the same failing step with the + // restore working is a plain `MigrationFailed`, and that error promises the + // pragmas `applyPragmas` guarantees. + { + var handle = try f.liveHandle(); + defer if (handle) |*h| h.close(); + + migrations.expected_failures.begin(); + defer migrations.expected_failures.end(); + try testing.expectError(error.MigrationFailed, migrations.runMigration( + f.io(), + std.Io.Dir.cwd(), + f.path(), + &handle.?, + failing_plan.step_sql, + failing_plan.minimum, + failing_plan.current, + )); + try testing.expectEqual(@as(i64, 1), try handle.?.queryInt("PRAGMA foreign_keys")); + try testing.expectEqual(@as(i64, 0), try handle.?.queryInt("PRAGMA legacy_alter_table")); + } + + // No path committed anything, and every run deleted its own backup. + var probe = try db.Db.open(f.path(), .{ .mode = .read_write_existing }); + defer probe.close(); + try testing.expectEqual( + @as(i64, failing_plan.minimum), + try probe.queryInt("PRAGMA user_version"), + ); + try testing.expectEqual( + @as(i64, 1), + try probe.queryInt("SELECT count(*) FROM domains WHERE domain = 'kept.example'"), + ); + try testing.expectEqual(@as(usize, 0), try f.countMatching("querylog.db.pre-migrate-")); +} + +test "a plain open retries a retention cleanup that once failed" { + var f: Fixture = .init(); + defer f.deinit(); + + { + var created = try open(f.io(), std.Io.Dir.cwd(), f.path()); + created.database.close(); + } + + // What a migration whose step-4 cleanup failed leaves behind: an older + // epoch, the newest epoch, and a same-second collision name tied with it. + const seeded = [_][]const u8{ + "querylog.db.pre-migrate-1600000000", + "querylog.db.pre-migrate-1700000000", + "querylog.db.pre-migrate-1700000000-2", + "querylog.db.pre-migrate-handwritten", + }; + for (seeded) |name| { + try f.tmp.dir.writeFile(f.io(), .{ .sub_path = name, .data = "x" }); + } + + var opened = try open(f.io(), std.Io.Dir.cwd(), f.path()); + defer opened.database.close(); + try testing.expectEqual(@as(?RecreateReason, null), opened.recreated); + + // Only the strictly older epoch goes: both files tied at the newest epoch + // survive, because this pass cannot tell which of them a migration made, + // and a name it did not write is never its to delete. + try testing.expect(!try exists(&f, seeded[0])); + for (seeded[1..]) |name| try testing.expect(try exists(&f, name)); +} + +fn exists(f: *Fixture, name: []const u8) !bool { + f.tmp.dir.access(f.io(), name, .{}) catch |e| switch (e) { + error.FileNotFound => return false, + else => return e, + }; + return true; +} diff --git a/src/storage/querylog_versions.zig b/src/storage/querylog_versions.zig new file mode 100644 index 0000000..8a0bde8 --- /dev/null +++ b/src/storage/querylog_versions.zig @@ -0,0 +1,89 @@ +//! The `querylog.db` schema version chain: comptime metadata and nothing else. +//! +//! Separate from `querylog_migrations.zig` so `tools/cut.zig` can import it +//! without linking SQLite. Nothing in this file may reach for `db.zig`, for a +//! C symbol, or for an allocator — the release gate reads these constants at +//! build time, and a dependency here would drag the whole storage layer into +//! the cut tool. +//! +//! **There are no migration hooks.** A step is a SQL file, period. Every +//! shipped step is therefore byte-comparable against the previous tag, which is +//! what lets the cut gate prove a released migration was never edited. A future +//! change that genuinely cannot be expressed in SQL must amend this design in +//! its own spec rather than adding a code path here. + +const std = @import("std"); + +/// The version a file created by this build carries in `PRAGMA user_version`. +pub const current_version: i32 = 1; + +/// The oldest stamped version this build can reach `current_version` from. +/// A file stamped below this refuses to open. +/// +/// An EXPLICIT BREAK in a future release is expressed here and only here: bump +/// `current_version`, set `minimum_supported_version = current_version`, and +/// ship no step. The chain then cannot reach the new version from below the +/// minimum, so `open` refuses the old file by the ordinary rules. A break is +/// always versioned, always refused at runtime, and never silent. +pub const minimum_supported_version: i32 = 1; + +/// The literal `user_version` the 0.0.12 and 0.0.13 binaries stamped: the CRC32 +/// of their DDL text, under the pre-migration policy where a stamp mismatch +/// meant "replace the file". +/// +/// FROZEN. It is derived from nothing at build time on purpose — recomputing it +/// from today's DDL would silently stop recognising the files it exists to +/// recognise the moment the schema moves. Editing it strands every 0.0.12 and +/// 0.0.13 file that has not yet been opened by a migration-aware build, which +/// is why the cut gate fails on any change to this line. +pub const legacy_fingerprint: i32 = 1975011655; + +/// Logical versions live far below any plausible CRC32 stamp. A value above +/// this is not a version this project ever wrote, so it classifies as +/// unsupported rather than as a from-the-future schema. +pub const version_floor_guard: i32 = 1_000_000; + +/// One entry per shipped step: `step_sql[i]` migrates version +/// `minimum_supported_version + i` to `minimum_supported_version + i + 1`. +/// Each entry is `@embedFile("migrations/v.sql")`, and each such file is +/// immutable once released. +/// +/// **Step-authoring rules** (the runner enforces the first, the equivalence +/// oracle catches violations of the rest): +/// +/// - A step contains no transaction statement. No `BEGIN`, no `COMMIT`, no +/// `ROLLBACK`, no `SAVEPOINT`: the runner wraps the whole chain in one +/// transaction and installs an authorizer that denies them outright. +/// - A step that changes a table's shape must REBUILD it, so that the CREATE +/// text SQLite stores ends up byte-identical to the fresh DDL's: +/// `DROP` every view over `` first; `ALTER TABLE RENAME TO _old`; +/// `CREATE TABLE ...` pasted verbatim from `querylog_schema.ddl`; +/// `INSERT INTO SELECT ... FROM _old`; `DROP TABLE _old`; recreate +/// every index and trigger of `` verbatim; recreate the dropped views +/// verbatim last. +/// - `ALTER TABLE ... ADD COLUMN` and `ALTER TABLE ... RENAME COLUMN` on a kept +/// table are forbidden. SQLite rewrites the stored CREATE text under them, +/// and the oracle's exact-text layer would rightly call the result unequal. +pub const step_sql: []const [:0]const u8 = &.{}; + +comptime { + std.debug.assert(minimum_supported_version >= 1); + std.debug.assert(minimum_supported_version <= current_version); + std.debug.assert(current_version <= version_floor_guard); + std.debug.assert(legacy_fingerprint < 0 or legacy_fingerprint > version_floor_guard); + std.debug.assert(step_sql.len == @as(usize, @intCast(current_version - minimum_supported_version))); +} + +test "the chain covers exactly the supported range" { + try std.testing.expectEqual( + @as(usize, @intCast(current_version - minimum_supported_version)), + step_sql.len, + ); +} + +test "the legacy anchor is the literal 0.0.12 stamp" { + // Not `fingerprintOf(ddl)`. The number is a historical fact about released + // binaries, so a test that recomputed it would move with the schema and + // prove nothing. + try std.testing.expectEqual(@as(i32, 1975011655), legacy_fingerprint); +} diff --git a/src/storage/repositories/queries_repo.zig b/src/storage/repositories/queries_repo.zig index 8e3fe8c..1d359e5 100644 --- a/src/storage/repositories/queries_repo.zig +++ b/src/storage/repositories/queries_repo.zig @@ -2518,7 +2518,10 @@ const recompute_checks = [_]struct { projection: []const u8, recompute: []const }, }; -fn expectProjectionsMatchRecompute(database: *db.Db) !void { +/// Exported for `querylog_migrations.zig`'s fixture and migration tests: the +/// authority on projection coherence is this file, and a second copy of the +/// recompute SQL there would be free to drift from the writer it checks. +pub fn expectProjectionsMatchRecompute(database: *db.Db) !void { for (recompute_checks) |check| { var buf: [4096]u8 = undefined; const sql = try std.fmt.bufPrint( diff --git a/src/storage/storage_integration_test.zig b/src/storage/storage_integration_test.zig index e0e4e62..e766553 100644 --- a/src/storage/storage_integration_test.zig +++ b/src/storage/storage_integration_test.zig @@ -28,7 +28,9 @@ const model = @import("../config/model.zig"); const validate = @import("../config/validate.zig"); const db = @import("db.zig"); const migrations = @import("migrations.zig"); +const querylog_migrations = @import("querylog_migrations.zig"); const querylog_schema = @import("querylog_schema.zig"); +const querylog_versions = @import("querylog_versions.zig"); const testing = std.testing; @@ -354,7 +356,7 @@ test "S7 case 1: querylog open on a fresh directory creates the schema" { try testing.expectEqual(querylog_schema.RecreateReason.missing, result.recreated.?); try testing.expectEqual( - @as(i64, querylog_schema.fingerprint), + @as(i64, querylog_versions.current_version), try result.database.queryInt("PRAGMA user_version"), ); try testing.expectEqual( @@ -387,7 +389,7 @@ test "S7 case 2: reopening a healthy querylog recreates nothing" { try testing.expectEqual(@as(usize, 0), asides.items.items.len); } -test "S7 case 3: a wrong user_version recreates and keeps the old file aside" { +test "S7 case 3: a version this build cannot handle refuses and touches nothing" { if (!build_options.integration) return error.SkipZigTest; var f: Fixture = .init(); @@ -396,29 +398,42 @@ test "S7 case 3: a wrong user_version recreates and keeps the old file aside" { var buf: [path_buf_len]u8 = undefined; const path = try f.pathZ(&buf, "querylog.db"); - try stampUserVersion(path, querylog_schema.fingerprint +% 1); - const original = try f.read("querylog.db"); - defer testing.allocator.free(original); + querylog_migrations.expected_failures.begin(); + defer querylog_migrations.expected_failures.end(); - var result = try querylog_schema.open(io, std.Io.Dir.cwd(), path); - defer result.database.close(); - try testing.expectEqual( - querylog_schema.RecreateReason.fingerprint_mismatch, - result.recreated.?, - ); + // Both refusal lanes, against a checkpointed file with no sidecars beside + // it: a stamp above this build's version (a downgrade) and a stamp that is + // not a version at all (a pre-0.0.12 fingerprint, or a foreign file). + const refusals = [_]struct { stamp: i32, expected: anyerror }{ + .{ .stamp = querylog_versions.current_version + 1, .expected = error.SchemaTooNew }, + .{ .stamp = 0, .expected = error.SchemaUnsupported }, + .{ .stamp = -7, .expected = error.SchemaUnsupported }, + .{ .stamp = 603440875, .expected = error.SchemaUnsupported }, + }; - var asides = try collectAsides(&f); - defer asides.deinit(); - try testing.expectEqual(@as(usize, 1), asides.items.items.len); + for (refusals) |lane| { + try stampUserVersion(path, lane.stamp); + try testing.expect(!try f.exists("querylog.db-wal")); - // The file was healthy: this build's schema moved, the database did not rot. - // An operator who reads "corrupt" here deletes a file that was never broken. - try testing.expect(std.mem.startsWith(u8, asides.items.items[0], "querylog.db.schema-changed-")); + const original = try f.read("querylog.db"); + defer testing.allocator.free(original); - const kept = try f.read(asides.items.items[0]); - defer testing.allocator.free(kept); - try testing.expectEqualSlices(u8, original, kept); + try testing.expectError( + lane.expected, + querylog_schema.open(io, std.Io.Dir.cwd(), path), + ); + + // Byte-identical, not merely "still readable": nothing was rewritten, + // no aside was made, and no fresh database was created beside it. + const after = try f.read("querylog.db"); + defer testing.allocator.free(after); + try testing.expectEqualSlices(u8, original, after); + + var asides = try collectAsides(&f); + defer asides.deinit(); + try testing.expectEqual(@as(usize, 0), asides.items.items.len); + } } test "S7 case 4: a garbage file recreates and the garbage is preserved" { @@ -468,11 +483,11 @@ test "S7 case 5: two recreates in the same second produce two distinct aside fil var round: usize = 0; while (round < 2) : (round += 1) { - try stampUserVersion(path, querylog_schema.fingerprint +% 1); + try f.write("querylog.db", "not a database at all"); var result = try querylog_schema.open(io, std.Io.Dir.cwd(), path); defer result.database.close(); try testing.expectEqual( - querylog_schema.RecreateReason.fingerprint_mismatch, + querylog_schema.RecreateReason.not_a_database, result.recreated.?, ); } @@ -492,7 +507,7 @@ test "S7 case 6: a stale write-ahead log is removed before the fresh database is var buf: [path_buf_len]u8 = undefined; const path = try f.pathZ(&buf, "querylog.db"); - try stampUserVersion(path, querylog_schema.fingerprint +% 1); + try f.write("querylog.db", "not a database at all"); // Existence alone proves nothing: the fresh database turns WAL on again and // writes its own `-wal`. The marker is what distinguishes the stale file @@ -503,7 +518,7 @@ test "S7 case 6: a stale write-ahead log is removed before the fresh database is var result = try querylog_schema.open(io, std.Io.Dir.cwd(), path); defer result.database.close(); try testing.expectEqual( - querylog_schema.RecreateReason.fingerprint_mismatch, + querylog_schema.RecreateReason.not_a_database, result.recreated.?, ); @@ -591,7 +606,7 @@ test "S7 case 23: a locked querylog propagates Busy and is never destroyed" { defer reopened.database.close(); try testing.expectEqual(@as(?querylog_schema.RecreateReason, null), reopened.recreated); try testing.expectEqual( - @as(i64, querylog_schema.fingerprint), + @as(i64, querylog_versions.current_version), try reopened.database.queryInt("PRAGMA user_version"), ); diff --git a/src/storage/testdata/querylog-v1-data.sql b/src/storage/testdata/querylog-v1-data.sql new file mode 100644 index 0000000..b61f43c --- /dev/null +++ b/src/storage/testdata/querylog-v1-data.sql @@ -0,0 +1,88 @@ +-- FROZEN FIXTURE. Representative content for a querylog.db at schema version 1, +-- loaded on top of `querylog-v1-schema.sql`. +-- +-- IMMUTABLE once released, for the same reason as its schema half: the release +-- gate byte-compares it against the previous tag. A future schema version ships +-- a NEW pair rather than editing this one. +-- +-- The rows are chosen to be hard on a migration rather than realistic: every +-- `route_kind`, the NULL variants of `qtype`, `cache_hit`, `response_time_us`, +-- `upstream` and `forward_zone`, a non-IN qclass with a non-zero rcode, the +-- group and source id/name pairs, `cname_target`/`safe_search_target`, an +-- `available_since` that has been advanced away from its DDL default, and +-- timestamps that straddle two 30-minute projection buckets. +-- +-- The `bucket_*` rows below are the recomputation of the raw rows, transcribed +-- from the same SQL `queries_repo`'s coherence oracle recomputes with. A +-- fixture-validity test runs that oracle over this file BEFORE any migration, +-- so an incoherent transcription fails on its own rather than as a migration +-- bug. + +UPDATE querylog_meta SET created_at = 1699998000, available_since = 1699998600 WHERE id = 1; + +INSERT INTO domains (id, domain) VALUES + (1, 'ads.example'), + (2, 'news.example'), + (3, 'chat.example'), + (4, 'printer.lan'), + (5, 'nas.lan'), + (6, 'cdn.example'); + +INSERT INTO query_log ( + id, timestamp, domain_id, client_ip, qtype, blocked, response_time_us, + cache_hit, upstream, qclass, rcode, group_id, group_name, policy_action, + policy_reason, matched, source_id, source_name, cname_target, + safe_search_target, route_kind, forward_zone +) VALUES + (1, 1699999260, 1, '10.0.0.1', 1, 1, NULL, 0, NULL, 1, 0, 7, 'kids', + 'block', 'blocklist_domain', 'ads.example', 3, 'stevenblack', NULL, NULL, + 'blocked', NULL), + (2, 1699999320, 2, '10.0.0.1', 28, 0, 1500, 0, '9.9.9.9:853', 1, 0, NULL, + NULL, 'allow', 'no_match', NULL, NULL, NULL, NULL, NULL, 'upstream', NULL), + (3, 1699999380, 2, '10.0.0.2', 1, 0, 90, 1, NULL, 1, 0, NULL, NULL, + 'allow', 'no_match', NULL, NULL, NULL, NULL, NULL, 'cache', NULL), + (4, 1699999440, 3, '10.0.0.2', NULL, 0, NULL, NULL, NULL, 3, 4, NULL, NULL, + 'not_evaluated', 'non_in_class', NULL, NULL, NULL, NULL, NULL, 'rejected', + NULL), + (5, 1699999500, 4, '10.0.0.3', 1, 0, 200, 0, NULL, 1, 0, NULL, NULL, + 'not_evaluated', 'local_record', NULL, NULL, NULL, NULL, NULL, 'local', + NULL), + (6, 1700000700, 5, '10.0.0.3', 15, 0, 3400, 0, NULL, 1, 0, NULL, NULL, + 'not_evaluated', 'forward_zone', NULL, NULL, NULL, NULL, NULL, + 'forward_zone', 'lan.example'), + (7, 1700001060, 1, '10.0.0.1', 1, 1, NULL, 0, NULL, 1, 0, 7, 'kids', + 'block', 'blocklist_wildcard', '*.ads.example', 3, 'stevenblack', NULL, + NULL, 'blocked', NULL), + (8, 1700001120, 6, '10.0.0.4', 65, 0, 2500, 1, '1.1.1.1:853', 1, 0, NULL, + NULL, 'allow', 'no_match', NULL, NULL, NULL, 'edge.cdn.example', + 'forcesafesearch.example', 'upstream', NULL); + +INSERT INTO bucket_totals (bucket, queries, blocked, cached, rt_sum, rt_count) VALUES + (1699999200, 6, 1, 1, 5190, 4), + (1700001000, 2, 1, 1, 2500, 1); + +INSERT INTO bucket_clients (bucket, client_ip, queries) VALUES + (1699999200, '10.0.0.1', 2), + (1699999200, '10.0.0.2', 2), + (1699999200, '10.0.0.3', 2), + (1700001000, '10.0.0.1', 1), + (1700001000, '10.0.0.4', 1); + +-- qtype -1 is the lossless encoding of the NULL qtype on row 4. +INSERT INTO bucket_types (bucket, qtype, count) VALUES + (1699999200, -1, 1), + (1699999200, 1, 3), + (1699999200, 15, 1), + (1699999200, 28, 1), + (1700001000, 1, 1), + (1700001000, 65, 1); + +INSERT INTO bucket_routes (bucket, route_kind, source_present, source_text, count) VALUES + (1699999200, 'blocked', 0, '', 1), + (1699999200, 'cache', 0, '', 1), + (1699999200, 'forward_zone', 1, 'lan.example', 1), + (1699999200, 'local', 0, '', 1), + (1699999200, 'rejected', 0, '', 1), + (1699999200, 'upstream', 1, '9.9.9.9:853', 1), + (1700001000, 'blocked', 0, '', 1), + (1700001000, 'upstream', 1, '1.1.1.1:853', 1); diff --git a/src/storage/testdata/querylog-v1-schema.sql b/src/storage/testdata/querylog-v1-schema.sql new file mode 100644 index 0000000..dcb8018 --- /dev/null +++ b/src/storage/testdata/querylog-v1-schema.sql @@ -0,0 +1,88 @@ +-- FROZEN FIXTURE. querylog.db schema version 1: byte-for-byte the `ddl` text of +-- `src/storage/querylog_schema.zig`, which is the schema the 0.0.12 and 0.0.13 +-- binaries created and the one version 1 names. +-- +-- IMMUTABLE once released. The release gate byte-compares this file against the +-- previous tag and fails the cut on any edit, because it is the starting point +-- every future migration is proved against: editing it would prove a migration +-- against a file no operator ever had. A new schema version ships a NEW pair. +-- +-- The `PRAGMA user_version` stamp is deliberately NOT part of this file. The +-- loader applies it, which is what lets one fixture serve both the version-1 +-- stamp and the 0.0.12/0.0.13 legacy fingerprint. + +CREATE TABLE domains ( + id INTEGER PRIMARY KEY, + domain TEXT NOT NULL UNIQUE +); + +CREATE TABLE query_log ( + id INTEGER PRIMARY KEY, + timestamp INTEGER NOT NULL, + domain_id INTEGER NOT NULL REFERENCES domains(id), + client_ip TEXT NOT NULL, -- text, not a FK: log rows are immutable facts + qtype INTEGER, + blocked INTEGER NOT NULL, + response_time_us INTEGER, + cache_hit INTEGER, + upstream TEXT, + qclass INTEGER NOT NULL, + rcode INTEGER NOT NULL, + group_id INTEGER, -- text/id pairs, not FKs: a renamed + group_name TEXT, -- group must not rewrite history + policy_action TEXT NOT NULL, + policy_reason TEXT NOT NULL, + matched TEXT, + source_id INTEGER, + source_name TEXT, + cname_target TEXT, + safe_search_target TEXT, + route_kind TEXT NOT NULL, + forward_zone TEXT, + CHECK (rcode BETWEEN 0 AND 4095) -- twelve bits (RFC 6891 6.1.3) +); +CREATE INDEX idx_query_log_ts ON query_log(timestamp); +CREATE INDEX idx_query_log_client ON query_log(client_ip); +CREATE INDEX idx_query_log_domain ON query_log(domain_id); + +CREATE TABLE querylog_meta ( + id INTEGER PRIMARY KEY CHECK (id = 1), -- one row, enforced by the schema + created_at INTEGER NOT NULL, + available_since INTEGER NOT NULL +); +INSERT INTO querylog_meta (id, created_at, available_since) +VALUES (1, unixepoch(), unixepoch() + 1); + +CREATE TABLE bucket_totals ( + bucket INTEGER PRIMARY KEY, + queries INTEGER NOT NULL, + blocked INTEGER NOT NULL, + cached INTEGER NOT NULL, + rt_sum INTEGER NOT NULL, -- sum(response_time_us) over timed rows + rt_count INTEGER NOT NULL -- count(response_time_us) +) WITHOUT ROWID; + +CREATE TABLE bucket_clients ( + bucket INTEGER NOT NULL, + client_ip TEXT NOT NULL, + queries INTEGER NOT NULL, + PRIMARY KEY (bucket, client_ip) +) WITHOUT ROWID; + +CREATE TABLE bucket_types ( + bucket INTEGER NOT NULL, + qtype INTEGER NOT NULL, -- -1 encodes a NULL qtype, losslessly + count INTEGER NOT NULL, + PRIMARY KEY (bucket, qtype) +) WITHOUT ROWID; + +CREATE TABLE bucket_routes ( + bucket INTEGER NOT NULL, + route_kind TEXT NOT NULL, + source_present INTEGER NOT NULL, -- 0: source NULL; 1: source = source_text + source_text TEXT NOT NULL, -- '' when source_present = 0 + count INTEGER NOT NULL, + PRIMARY KEY (bucket, route_kind, source_present, source_text), + CHECK (source_present IN (0, 1)), + CHECK (source_present = 1 OR source_text = '') +) WITHOUT ROWID; diff --git a/src/tests.zig b/src/tests.zig index cfac243..b96d3fb 100644 --- a/src/tests.zig +++ b/src/tests.zig @@ -40,7 +40,10 @@ comptime { _ = @import("config/faults.zig"); _ = @import("storage/config_schema.zig"); _ = @import("storage/migrations.zig"); + _ = @import("storage/querylog_migrations.zig"); _ = @import("storage/querylog_schema.zig"); + _ = @import("storage/querylog_versions.zig"); + _ = @import("storage/querylog_fixtures.zig"); _ = @import("storage/provenance.zig"); _ = @import("storage/repositories/context.zig"); _ = @import("storage/repositories/crud.zig"); diff --git a/tools/cut.zig b/tools/cut.zig index ac9d0a1..46bb77a 100644 --- a/tools/cut.zig +++ b/tools/cut.zig @@ -57,6 +57,13 @@ const http = std.http; /// the expression that computes it. const querylog_schema = @import("querylog_schema"); +/// The migration metadata the gates below judge: the supported version range +/// and the step chain, as `querylog_versions.zig` declares it and +/// `querylog_schema.open` runs it. Reached through `production_plan` rather +/// than as a second module because `querylog_schema.zig` already imports that +/// file, and one source file cannot belong to two modules. +const querylog_versions = querylog_schema.production_plan; + const max_input_bytes = 1 << 30; /// The only repository this program can ever act on. There is no flag for it: @@ -561,9 +568,283 @@ fn disclosesHistoryReset(section: []const u8) bool { return std.mem.indexOf(u8, section, history_reset_phrase) != null; } -/// The file whose DDL decides whether `querylog.db` survives an upgrade. +/// The phrase a changelog section must carry to release a MIGRATION. It is the +/// other operator-facing consequence: the history survives, and the first start +/// after the upgrade rewrites the file to get there. +const migration_phrase = "migrates your query log in place"; + +fn disclosesMigration(section: []const u8) bool { + return std.mem.indexOf(u8, section, migration_phrase) != null; +} + +/// The heading under which an explicit break tells the operator how to get +/// their history back. A break is allowed; a break with nowhere to turn is not. +const restore_heading = "### Restoring your query history"; + +/// Whether the section carries `restore_heading` AND something under it. An +/// empty section under the heading is the failure mode this exists to catch: +/// the heading alone would satisfy a substring check while telling the operator +/// nothing at all. +fn disclosesRestoreInstructions(section: []const u8) bool { + var lines = std.mem.splitScalar(u8, section, '\n'); + var under_heading = false; + while (lines.next()) |raw| { + const line = std.mem.trim(u8, std.mem.trimEnd(u8, raw, "\r"), " \t"); + if (under_heading) { + if (std.mem.startsWith(u8, line, "#")) return false; + if (!isBlank(line)) return true; + continue; + } + if (std.mem.eql(u8, line, restore_heading)) under_heading = true; + } + return false; +} + +// --------------------------------------------------------------------------- +// the two migration gates (specs/milestone-38.md B.2) +// --------------------------------------------------------------------------- + +/// A file that is immutable once released, and what became of it in this tree. +/// +/// The gate never sees the bytes. Reading two revisions of a file is the +/// driver's job; deciding what a difference means is a pure function of these +/// three states, which is what makes every rule below a unit test. +const ShippedFile = struct { + kind: enum { step, fixture }, + path: []const u8, + status: enum { identical, differs, missing }, +}; + +/// One link of the chain this build ships: the bytes `querylog_versions.step_sql` +/// carries for it, and the bytes of the tree file it is supposed to be an +/// `@embedFile` of. +/// +/// The pair is what makes "a step is a SQL file, period" checkable. Counting +/// steps proves only that the chain is the right LENGTH; comparing these two +/// byte strings proves each link is the frozen file the previous release can be +/// diffed against, so inline SQL, a reordered chain and an edited file all fail. +const ChainStep = struct { + embedded: []const u8, + /// The tree's `src/storage/migrations/v.sql`, or null when that file + /// does not exist. + on_disk: ?[]const u8, +}; + +/// Everything the gates judge: the tree's migration metadata, the previous +/// release's, whether the schema text moved, what became of the files the +/// previous release froze, and the changelog section for this version. +const GateInput = struct { + ddl_changed: bool, + current_version: i32, + minimum_version: i32, + legacy_fingerprint: i32, + /// The chain in `step_sql` order: `chain[i]` migrates + /// `minimum_version + i` to `+ i + 1`. + chain: []const ChainStep, + prev_version: i32, + prev_minimum: i32, + /// One entry per step file and fixture file the PREVIOUS tag shipped. + shipped: []const ShippedFile, + /// The versions in the tree that have BOTH halves of a fixture pair. + fixture_versions: []const i32, + /// The `## []` section, or empty when CHANGELOG.md could not be + /// read — which fails every rule that needs a disclosure, on purpose. + changelog_section: []const u8, +}; + +/// Which lane, if any, a schema text change is released under. +const Gate1 = enum { + /// The DDL is byte-identical to the previous release's, so this gate has + /// nothing to say. Gate 2 still runs. + unchanged, + migration_lane, + break_lane, + /// The schema moved under neither lane. This is the v0.0.9 failure. + no_lane, +}; + +/// The metadata a release can only have by being an explicit break: a new +/// version, no way back from the previous one, and a changelog that says so and +/// says how to recover. +fn isExplicitBreak(in: GateInput) bool { + return in.current_version > in.prev_version and + in.minimum_version == in.current_version and + disclosesHistoryReset(in.changelog_section) and + disclosesRestoreInstructions(in.changelog_section); +} + +fn gate1(in: GateInput) Gate1 { + if (!in.ddl_changed) return .unchanged; + + // `prev_minimum <= prev_version` is what makes the previous release's files + // reachable. An explicit break sets `minimum == current > prev_version`, so + // it fails this test and can never wear the migration lane. + const chain_spans_range = in.chain.len == stepsBetween(in.minimum_version, in.current_version); + if (in.current_version > in.prev_version and + in.prev_version >= in.minimum_version and + chain_spans_range) return .migration_lane; + + if (isExplicitBreak(in)) return .break_lane; + return .no_lane; +} + +/// How many steps a contiguous chain from `from` to `to` has. Zero when the +/// range is empty or inverted, so a regressed version cannot produce a negative +/// count that would wrap. +fn stepsBetween(from: i32, to: i32) usize { + if (to <= from) return 0; + return @intCast(to - from); +} + +/// Everything Gate 2 refuses. It runs whether or not the DDL moved: a +/// data-only migration and an edit to a released step file both leave the +/// schema text alone. +const Gate2Reason = enum { + step_edited, + step_missing, + step_has_no_file, + step_not_its_file, + fixture_edited, + fixture_missing, + fixture_pair_absent, + legacy_fingerprint_edited, + version_regressed, + minimum_regressed, + minimum_raised_without_break, + bump_without_step_or_break, + migration_undisclosed, +}; + +const Gate2Problem = struct { + reason: Gate2Reason, + /// The file or version the reason is about, for the message. Empty when the + /// reason is about the metadata as a whole. + subject: []const u8 = "", +}; + +/// The literal `querylog_versions.legacy_fingerprint` is frozen forever: +/// editing it strands every 0.0.12/0.0.13 file that has not yet been opened by +/// a migration-aware build. The gate holds the same number the module does. +const frozen_legacy_fingerprint: i32 = 1975011655; + +fn gate2(arena: Allocator, in: GateInput) ?Gate2Problem { + if (in.legacy_fingerprint != frozen_legacy_fingerprint) { + return .{ .reason = .legacy_fingerprint_edited }; + } + + for (in.shipped) |file| { + const reason: ?Gate2Reason = switch (file.status) { + .identical => null, + .differs => switch (file.kind) { + .step => .step_edited, + .fixture => .fixture_edited, + }, + .missing => switch (file.kind) { + .step => .step_missing, + .fixture => .fixture_missing, + }, + }; + if (reason) |r| return .{ .reason = r, .subject = file.path }; + } + + // Every link of the chain is the frozen file at its own index. The path is + // computed here rather than taken from the input, so a step can only clear + // this rule by being the `@embedFile` of the one file the next release will + // byte-compare against its predecessor. + for (in.chain, 0..) |step, index| { + const from = in.minimum_version + @as(i32, @intCast(index)); + const path = std.fmt.allocPrint(arena, "{s}/v{d}.sql", .{ migrations_dir, from }) catch @panic("OOM"); + const on_disk = step.on_disk orelse return .{ .reason = .step_has_no_file, .subject = path }; + if (!std.mem.eql(u8, on_disk, step.embedded)) { + return .{ .reason = .step_not_its_file, .subject = path }; + } + } + + var version = in.minimum_version; + while (version <= in.current_version) : (version += 1) { + if (std.mem.indexOfScalar(i32, in.fixture_versions, version) == null) { + return .{ + .reason = .fixture_pair_absent, + .subject = std.fmt.allocPrint(arena, "{d}", .{version}) catch @panic("OOM"), + }; + } + } + + if (in.current_version < in.prev_version) return .{ .reason = .version_regressed }; + if (in.minimum_version < in.prev_minimum) return .{ .reason = .minimum_regressed }; + // Raising the minimum drops support for schemas the previous release + // carried. That is allowed exactly once per break and never quietly, and + // the DDL fingerprint has no say in it — a break can leave the text alone. + if (in.minimum_version > in.prev_minimum and !isExplicitBreak(in)) { + return .{ .reason = .minimum_raised_without_break }; + } + + if (in.current_version > in.prev_version) { + const new_steps = in.chain.len > stepsBetween(in.prev_minimum, in.prev_version); + const a_break = in.minimum_version == in.current_version; + if (!new_steps and !a_break) return .{ .reason = .bump_without_step_or_break }; + // A break discloses under Gate 1's break lane instead: its history does + // not migrate, it is thrown away. + if (new_steps and !a_break and !disclosesMigration(in.changelog_section)) { + return .{ .reason = .migration_undisclosed }; + } + } + + return null; +} + +/// The file whose DDL decides what shape `querylog.db` has. const querylog_schema_path = "src/storage/querylog_schema.zig"; +/// The file whose constants decide whether an existing `querylog.db` survives +/// the upgrade, and how. +const querylog_versions_path = "src/storage/querylog_versions.zig"; + +const migrations_dir = "src/storage/migrations"; +const fixtures_dir = "src/storage/testdata"; + +/// A `pub const : i32 = ;` out of any revision of +/// `querylog_versions.zig`, read as text for the same reason `extractDdl` reads +/// the DDL as text: the previous release's copy only exists as `git show` +/// output. Null when the declaration is absent or is not a plain literal, which +/// is a refusal rather than a default — guessing a version would let a gate +/// pass a release it never measured. +fn extractVersionConst(file_text: []const u8, name: []const u8) ?i32 { + var lines = std.mem.splitScalar(u8, file_text, '\n'); + while (lines.next()) |raw| { + const line = std.mem.trim(u8, std.mem.trimEnd(u8, raw, "\r"), " \t"); + var prefix_buf: [64]u8 = undefined; + const prefix = std.fmt.bufPrint(&prefix_buf, "pub const {s}: i32 = ", .{name}) catch return null; + if (!std.mem.startsWith(u8, line, prefix)) continue; + + const rest = line[prefix.len..]; + const end = std.mem.indexOfScalar(u8, rest, ';') orelse return null; + var digits: [32]u8 = undefined; + var len: usize = 0; + for (std.mem.trim(u8, rest[0..end], " \t")) |ch| { + if (ch == '_') continue; + if (len == digits.len) return null; + digits[len] = ch; + len += 1; + } + return std.fmt.parseInt(i32, digits[0..len], 10) catch null; + } + return null; +} + +/// The version a fixture path names, for either half of a pair. Null for any +/// name that is not one, so an unrelated file in `testdata/` is ignored rather +/// than parsed into a version that does not exist. +fn fixtureVersionOf(name: []const u8) ?i32 { + const prefix = "querylog-v"; + if (!std.mem.startsWith(u8, name, prefix)) return null; + const rest = name[prefix.len..]; + const dash = std.mem.indexOfScalar(u8, rest, '-') orelse return null; + const suffix = rest[dash..]; + if (!std.mem.eql(u8, suffix, "-schema.sql") and !std.mem.eql(u8, suffix, "-data.sql")) return null; + return std.fmt.parseInt(i32, rest[0..dash], 10) catch null; +} + /// The declaration line the DDL follows, matched whole so no other `ddl` in the /// file can be mistaken for it. const ddl_declaration = "pub const ddl: [:0]const u8 ="; @@ -1561,22 +1842,26 @@ fn preflight(ctx: *Ctx, version: []const u8, bump_needed: bool, plan: Plan) !Pre return result; } -/// Refuses a release that changes the querylog schema without saying so. +/// Refuses a release whose querylog schema or migration metadata moved without +/// the release saying what that costs the operator. /// -/// `querylog.db` is never migrated: the server compares the file's stamped -/// fingerprint against this build's and, on a mismatch, sets the file aside and -/// creates an empty one. Every query the operator ever logged is gone on the -/// first start after the upgrade. v0.0.9 shipped exactly that while its -/// announcement claimed no such change, which is what this check exists to stop. +/// TWO INDEPENDENT GATES, both measured against the previous release TAG rather +/// than the last commit, because the tag is what an operator upgrades from. /// -/// The comparison is between the DDL of the previous release tag and this -/// tree's, so it measures the release, not the last commit. Every step that can -/// fail — listing the tags, reading the old file, parsing it — is a refusal -/// naming the step: a gate that cannot tell whether the schema moved must not -/// report that it did not. +/// Gate 1 is about the schema TEXT. A changed DDL has to be released under one +/// of exactly two lanes: a migration that carries the file forward, or an +/// explicit break that throws the history away and says how to get it back. +/// v0.0.9 shipped a silent break while its announcement claimed no such change, +/// which is what this gate exists to stop. +/// +/// Gate 2 is about the migration METADATA, and it runs whether or not the text +/// moved: a data-only migration, an edit to a step that has already shipped, an +/// edited fixture and a quietly raised minimum all leave the DDL alone. +/// +/// Every step that can fail — listing the tags, reading the old files, parsing +/// them — is a refusal naming the step. A gate that cannot tell whether +/// something moved must not report that it did not. fn schemaGate(ctx: *Ctx, version: []const u8, target: Semver, changelog: ?[]const u8) !void { - const current = querylog_schema.fingerprint; - const tags = try gitCapture(ctx, &.{ "git", "ls-remote", "--tags", "origin" }, git_network_timeout_s); if (!tags.ok()) { ctx.soft("schema-gate", "`git ls-remote --tags origin` exited {d}: {s}", .{ @@ -1611,33 +1896,184 @@ fn schemaGate(ctx: *Ctx, version: []const u8, target: Semver, changelog: ?[]cons }); return; }; - const old = querylog_schema.fingerprintOf(old_ddl); + const old_fingerprint = querylog_schema.fingerprintOf(old_ddl); + const current_fingerprint = querylog_schema.fingerprint; - if (old == current) { - ctx.pass("schema-gate", "the querylog schema is unchanged since {s} (fingerprint {d})", .{ previous_tag, current }); - return; + // The previous release's metadata. `querylog_versions.zig` did not exist + // before milestone 38, and every file such a release created is a version-1 + // file — that is what the legacy fingerprint stands for — so an absent + // module is 1 and 1 rather than a refusal. The object itself is known good + // by now: the DDL above came out of it. + var prev_version: i32 = 1; + var prev_minimum: i32 = 1; + const old_versions = try gitCapture(ctx, &.{ + "git", "show", ctx.fmt("{s}:{s}", .{ previous.object, querylog_versions_path }), + }, git_local_timeout_s); + if (old_versions.ok()) { + prev_version = extractVersionConst(old_versions.stdout, "current_version") orelse { + ctx.soft("schema-gate", "cannot read `current_version` out of {s}:{s} ({s})", .{ + previous.object, querylog_versions_path, previous_tag, + }); + return; + }; + prev_minimum = extractVersionConst(old_versions.stdout, "minimum_supported_version") orelse { + ctx.soft("schema-gate", "cannot read `minimum_supported_version` out of {s}:{s} ({s})", .{ + previous.object, querylog_versions_path, previous_tag, + }); + return; + }; + } else { + ctx.note("schema-gate: {s} predates {s}, so it is read as schema version 1", .{ + previous_tag, querylog_versions_path, + }); } - const source = changelog orelse { + const shipped = frozenFiles(ctx, previous.object, previous_tag) catch |err| switch (err) { + error.CheckFailed => return, + else => return err, + }; + + const in: GateInput = .{ + .ddl_changed = old_fingerprint != current_fingerprint, + .current_version = querylog_versions.current, + .minimum_version = querylog_versions.minimum, + .legacy_fingerprint = querylog_versions.legacy_fingerprint, + .chain = treeChain(ctx), + .prev_version = prev_version, + .prev_minimum = prev_minimum, + .shipped = shipped, + .fixture_versions = treeFixtureVersions(ctx), + .changelog_section = if (changelog) |source| changelogSection(source, version) orelse "" else "", + }; + + if (changelog == null) { // The changelog check already reported why it could not be read; this - // reports what that costs, because the gate has no way to clear itself. - ctx.soft("schema-gate", "the querylog schema changed since {s} ({d} to {d}) and CHANGELOG.md could not be read to check the disclosure", .{ - previous_tag, old, current, + // reports what that costs, because neither gate can clear itself + // without the disclosure it is looking for. + ctx.soft("schema-gate", "CHANGELOG.md could not be read, so no disclosure can be checked", .{}); + } + + switch (gate1(in)) { + .unchanged => ctx.pass("schema-gate", "the querylog schema is unchanged since {s} (fingerprint {d})", .{ + previous_tag, current_fingerprint, + }), + .migration_lane => ctx.pass("schema-gate", "the querylog schema changed since {s} ({d} to {d}) and schema version {d} migrates to {d} in place", .{ + previous_tag, old_fingerprint, current_fingerprint, prev_version, in.current_version, + }), + .break_lane => ctx.pass("schema-gate", "the querylog schema changed since {s} ({d} to {d}) as an explicit break to schema version {d}, and the `## [{s}]` section says so and says how to recover", .{ + previous_tag, old_fingerprint, current_fingerprint, in.current_version, version, + }), + .no_lane => ctx.soft( + "schema-gate", + "the querylog schema changed since {s} ({d} to {d}) under neither lane. Either ship a migration (raise `current_version` above {d}, keeping `minimum_supported_version` at or below it, with a step per version) or declare an explicit break (`minimum_supported_version == current_version`) and give the `## [{s}]` section both the phrase '{s}' and a `{s}` section with recovery steps", + .{ previous_tag, old_fingerprint, current_fingerprint, prev_version, version, history_reset_phrase, restore_heading }, + ), + } + + const problem = gate2(ctx.arena, in) orelse { + ctx.pass("schema-gate-metadata", "the migration metadata is consistent with {s}: schema versions {d}..{d}, {d} step(s), every released step and fixture untouched", .{ + previous_tag, in.minimum_version, in.current_version, in.chain.len, }); return; }; - const section = changelogSection(source, version) orelse ""; - if (!disclosesHistoryReset(section)) { - ctx.soft( - "schema-gate", - "the querylog schema changed since {s} ({d} to {d}), so the first start after this release sets querylog.db aside and creates an empty one; say so in the `## [{s}]` section, which must contain the phrase '{s}'", - .{ previous_tag, old, current, version, history_reset_phrase }, - ); - return; + switch (problem.reason) { + .step_edited => ctx.soft("schema-gate-metadata", "`{s}` shipped in {s} and this tree changes it; a released migration step is immutable, so add a new step instead", .{ problem.subject, previous_tag }), + .step_missing => ctx.soft("schema-gate-metadata", "`{s}` shipped in {s} and is gone from this tree; a released migration step is immutable and every operator still below its target needs it", .{ problem.subject, previous_tag }), + .step_has_no_file => ctx.soft("schema-gate-metadata", "step {s} of the chain has no `{s}`; a step is a SQL file and nothing else, so inline SQL leaves the next release nothing to byte-compare and no operator a way to audit what ran", .{ problem.subject, problem.subject }), + .step_not_its_file => ctx.soft("schema-gate-metadata", "the chain's bytes for `{s}` are not that file's bytes; every step is the `@embedFile` of its own `v.sql`, so rebuild the chain from the files rather than editing one side of the pair", .{problem.subject}), + .fixture_edited => ctx.soft("schema-gate-metadata", "`{s}` shipped in {s} and this tree changes it; a released fixture is the file the next migration is proved against, so a new schema version ships a NEW pair", .{ problem.subject, previous_tag }), + .fixture_missing => ctx.soft("schema-gate-metadata", "`{s}` shipped in {s} and is gone from this tree; a released fixture is immutable", .{ problem.subject, previous_tag }), + .fixture_pair_absent => ctx.soft("schema-gate-metadata", "schema version {s} is supported but has no `{s}/querylog-v{s}-schema.sql` and `-data.sql` pair; every version in {d}..{d} needs one", .{ problem.subject, fixtures_dir, problem.subject, in.minimum_version, in.current_version }), + .legacy_fingerprint_edited => ctx.soft("schema-gate-metadata", "`legacy_fingerprint` is {d}, not the frozen {d}; it is the literal stamp the 0.0.12 and 0.0.13 binaries wrote, and changing it strands every such file that no migration-aware build has opened yet", .{ in.legacy_fingerprint, frozen_legacy_fingerprint }), + .version_regressed => ctx.soft("schema-gate-metadata", "`current_version` is {d} and {s} shipped {d}; the schema version never regresses", .{ in.current_version, previous_tag, prev_version }), + .minimum_regressed => ctx.soft("schema-gate-metadata", "`minimum_supported_version` is {d} and {s} shipped {d}; this build claims to migrate files the previous one could not, with no step to do it", .{ in.minimum_version, previous_tag, prev_minimum }), + .minimum_raised_without_break => ctx.soft("schema-gate-metadata", "`minimum_supported_version` rises from {d} to {d}, which drops support for schemas {s} could open. That is only releasable as the full explicit break: `minimum_supported_version == current_version`, a `current_version` above {d}, and a `## [{s}]` section carrying both '{s}' and a `{s}` section", .{ prev_minimum, in.minimum_version, previous_tag, prev_version, version, history_reset_phrase, restore_heading }), + .bump_without_step_or_break => ctx.soft("schema-gate-metadata", "`current_version` rises from {d} to {d} with no new step file and no explicit break; a version an operator's file cannot reach and is not refused for is a silent reset", .{ prev_version, in.current_version }), + .migration_undisclosed => ctx.soft("schema-gate-metadata", "this release migrates querylog.db from schema version {d} to {d}, so the `## [{s}]` section must contain the phrase '{s}'", .{ prev_version, in.current_version, version, migration_phrase }), } - ctx.pass("schema-gate", "the querylog schema changed since {s} ({d} to {d}) and the `## [{s}]` section discloses it", .{ - previous_tag, old, current, version, - }); +} + +/// The step and fixture files the previous tag froze, each paired with what +/// this tree did to it. +/// +/// `git ls-tree` lists the tag's side; the tree's side is read off disk, +/// because a fixture added in this working copy is not in any index yet. +fn frozenFiles(ctx: *Ctx, object: []const u8, previous_tag: []const u8) ![]const ShippedFile { + const listing = try gitCapture(ctx, &.{ + "git", "ls-tree", "-r", "--name-only", object, "--", migrations_dir, fixtures_dir, + }, git_local_timeout_s); + if (!listing.ok()) { + ctx.soft("schema-gate-metadata", "`git ls-tree {s}` for {s} exited {d}: {s}", .{ + object, previous_tag, listing.code, std.mem.trimEnd(u8, listing.combined(ctx.arena), "\n"), + }); + return CheckFailed; + } + + var files: std.ArrayList(ShippedFile) = .empty; + var lines = std.mem.splitScalar(u8, listing.stdout, '\n'); + while (lines.next()) |raw| { + const path = std.mem.trim(u8, raw, " \t\r"); + if (path.len == 0) continue; + + const kind: @FieldType(ShippedFile, "kind") = if (std.mem.startsWith(u8, path, migrations_dir ++ "/")) + .step + else if (fixtureVersionOf(std.fs.path.basename(path)) != null) + .fixture + else + // Anything else under `testdata/` belongs to some other test and + // carries no immutability promise. + continue; + + const released = try gitCapture(ctx, &.{ + "git", "show", ctx.fmt("{s}:{s}", .{ object, path }), + }, git_local_timeout_s); + if (!released.ok()) { + ctx.soft("schema-gate-metadata", "`git show {s}:{s}` exited {d}: {s}", .{ + object, path, released.code, std.mem.trimEnd(u8, released.combined(ctx.arena), "\n"), + }); + return CheckFailed; + } + + const status: @FieldType(ShippedFile, "status") = blk: { + const current = Io.Dir.cwd().readFileAlloc(ctx.io, path, ctx.arena, .limited(max_input_bytes)) catch + break :blk .missing; + break :blk if (std.mem.eql(u8, current, released.stdout)) .identical else .differs; + }; + files.append(ctx.arena, .{ .kind = kind, .path = path, .status = status }) catch @panic("OOM"); + } + return files.items; +} + +/// The chain this build embedded, each step paired with the tree file it claims +/// to be. Reading the file is all this does; whether the two agree is Gate 2's +/// rule, and an unreadable file reads as absent so that the gate names the step +/// rather than the syscall. +fn treeChain(ctx: *Ctx) []const ChainStep { + var chain: std.ArrayList(ChainStep) = .empty; + for (querylog_versions.step_sql, 0..) |embedded, index| { + const from = querylog_versions.minimum + @as(i32, @intCast(index)); + const path = ctx.fmt("{s}/v{d}.sql", .{ migrations_dir, from }); + const on_disk = Io.Dir.cwd().readFileAlloc(ctx.io, path, ctx.arena, .limited(max_input_bytes)) catch null; + chain.append(ctx.arena, .{ .embedded = embedded, .on_disk = on_disk }) catch @panic("OOM"); + } + return chain.items; +} + +/// The versions this tree has BOTH halves of a fixture pair for, over the range +/// the metadata claims to support. Probing the range beats listing the +/// directory: the range is what the rule is about, and a stray `querylog-v9-` +/// file for some unsupported version proves nothing either way. +fn treeFixtureVersions(ctx: *Ctx) []const i32 { + var found: std.ArrayList(i32) = .empty; + var version = querylog_versions.minimum; + while (version <= querylog_versions.current) : (version += 1) { + const schema = ctx.fmt("{s}/querylog-v{d}-schema.sql", .{ fixtures_dir, version }); + const data = ctx.fmt("{s}/querylog-v{d}-data.sql", .{ fixtures_dir, version }); + _ = Io.Dir.cwd().readFileAlloc(ctx.io, schema, ctx.arena, .limited(max_input_bytes)) catch continue; + _ = Io.Dir.cwd().readFileAlloc(ctx.io, data, ctx.arena, .limited(max_input_bytes)) catch continue; + found.append(ctx.arena, version) catch @panic("OOM"); + } + return found.items; } /// What to do about a `v` tag that exists locally. @@ -2673,3 +3109,360 @@ test "a published release is only reported from a payload that carries one" { // A missing tag_name yields the empty string, which never equals a tag. try testing.expectEqualStrings("", jsonString(no_assets.object, "id")); } + +// --------------------------------------------------------------------------- +// the two migration gates +// --------------------------------------------------------------------------- + +/// A release with nothing to declare: the schema is unchanged, the metadata is +/// the previous release's, and every frozen file is where it was. Each test +/// below changes exactly the fields its rule is about, so what it is testing is +/// what it names. +fn baseGateInput() GateInput { + return .{ + .ddl_changed = false, + .current_version = 1, + .minimum_version = 1, + .legacy_fingerprint = frozen_legacy_fingerprint, + .chain = &.{}, + .prev_version = 1, + .prev_minimum = 1, + .shipped = &.{}, + .fixture_versions = &.{1}, + .changelog_section = "", + }; +} + +/// Two steps of plausible SQL, and the chains a correctly authored release +/// carries them in: the embedded bytes ARE the file's bytes. +const step_v1_sql = "ALTER TABLE domains RENAME TO domains_old;\n"; +const step_v2_sql = "DROP VIEW recent_queries;\n"; +const one_frozen_step: []const ChainStep = &.{ + .{ .embedded = step_v1_sql, .on_disk = step_v1_sql }, +}; +const two_frozen_steps: []const ChainStep = &.{ + .{ .embedded = step_v1_sql, .on_disk = step_v1_sql }, + .{ .embedded = step_v2_sql, .on_disk = step_v2_sql }, +}; + +const migration_section = "This release " ++ migration_phrase ++ ", so nothing is lost.\n"; +const break_section = "This release " ++ history_reset_phrase ++ ".\n\n" ++ + restore_heading ++ "\n\nStop the server and move the aside file back.\n"; + +/// A release that migrates schema version 1 to 2: one new step, one new fixture +/// pair, and the changelog phrase that discloses it. +fn migratingGateInput() GateInput { + var in = baseGateInput(); + in.ddl_changed = true; + in.current_version = 2; + in.chain = one_frozen_step; + in.fixture_versions = &.{ 1, 2 }; + in.changelog_section = migration_section; + return in; +} + +/// A release that abandons schema version 1 instead of migrating it. +fn breakingGateInput() GateInput { + var in = baseGateInput(); + in.ddl_changed = true; + in.current_version = 2; + in.minimum_version = 2; + in.chain = &.{}; + in.fixture_versions = &.{2}; + in.changelog_section = break_section; + return in; +} + +fn expectGate2(in: GateInput, expected: ?Gate2Reason) !void { + var arena_state = std.heap.ArenaAllocator.init(testing.allocator); + defer arena_state.deinit(); + const problem = gate2(arena_state.allocator(), in); + if (expected) |reason| { + try testing.expectEqual(reason, (problem orelse return error.GatePassed).reason); + } else { + if (problem) |actual| { + std.debug.print("unexpected gate 2 failure: {t} ({s})\n", .{ actual.reason, actual.subject }); + return error.GateFailed; + } + } +} + +test "a release that touches neither the schema nor the metadata passes both gates" { + const in = baseGateInput(); + try testing.expectEqual(Gate1.unchanged, gate1(in)); + try expectGate2(in, null); +} + +test "a migration is released under the migration lane" { + const in = migratingGateInput(); + try testing.expectEqual(Gate1.migration_lane, gate1(in)); + try expectGate2(in, null); +} + +test "an explicit break is released under the break lane" { + const in = breakingGateInput(); + try testing.expectEqual(Gate1.break_lane, gate1(in)); + try expectGate2(in, null); +} + +test "a schema change under neither lane is refused" { + var in = baseGateInput(); + in.ddl_changed = true; + // The v0.0.9 shape exactly: the DDL moved and nothing else did. + try testing.expectEqual(Gate1.no_lane, gate1(in)); +} + +test "break metadata cannot be released as a migration" { + var in = breakingGateInput(); + // `minimum == current` means the previous release's files cannot reach the + // new version at all. Saying they migrate does not make them. + in.changelog_section = migration_section; + try testing.expectEqual(Gate1.no_lane, gate1(in)); +} + +test "a version bump whose chain does not span the supported range is refused" { + var in = migratingGateInput(); + in.current_version = 3; + in.fixture_versions = &.{ 1, 2, 3 }; + // One step cannot carry a file from 1 to 3. + try testing.expectEqual(Gate1.no_lane, gate1(in)); +} + +test "an edited or deleted released step is refused however the version moved" { + const path = migrations_dir ++ "/v1.sql"; + for ([_]@FieldType(ShippedFile, "status"){ .differs, .missing }) |status| { + var in = migratingGateInput(); + // A perfectly well-formed version append, which is exactly the case + // that must not launder an edit to a step already in operators' hands. + in.current_version = 3; + in.chain = two_frozen_steps; + in.fixture_versions = &.{ 1, 2, 3 }; + in.shipped = &.{.{ .kind = .step, .path = path, .status = status }}; + + try testing.expectEqual(Gate1.migration_lane, gate1(in)); + try expectGate2(in, if (status == .differs) .step_edited else .step_missing); + } +} + +test "every step of the chain must be the frozen file at its own index" { + // Matching bytes are the whole rule, so start by proving they pass. + const frozen = migratingGateInput(); + try expectGate2(frozen, null); + + // A step written inline, with no `v1.sql` for the next release to compare + // against. Counting steps calls this chain complete; the byte comparison + // does not. + var inline_only = migratingGateInput(); + inline_only.chain = &.{.{ .embedded = step_v1_sql, .on_disk = null }}; + try testing.expectEqual(Gate1.migration_lane, gate1(inline_only)); + try expectGate2(inline_only, .step_has_no_file); + + // The file edited after the fact, so the binary runs SQL the audited file no + // longer contains. + var edited = migratingGateInput(); + edited.chain = &.{.{ .embedded = step_v1_sql, .on_disk = step_v1_sql ++ "DROP TABLE domains;\n" }}; + try expectGate2(edited, .step_not_its_file); + + // And a chain listing its files out of order: index 0 must be `v1.sql`. + var reordered = migratingGateInput(); + reordered.current_version = 3; + reordered.fixture_versions = &.{ 1, 2, 3 }; + reordered.chain = &.{ + .{ .embedded = step_v2_sql, .on_disk = step_v1_sql }, + .{ .embedded = step_v1_sql, .on_disk = step_v2_sql }, + }; + try expectGate2(reordered, .step_not_its_file); +} + +test "an edited or deleted released fixture is refused" { + const path = fixtures_dir ++ "/querylog-v1-data.sql"; + for ([_]@FieldType(ShippedFile, "status"){ .differs, .missing }) |status| { + var in = migratingGateInput(); + in.shipped = &.{.{ .kind = .fixture, .path = path, .status = status }}; + try expectGate2(in, if (status == .differs) .fixture_edited else .fixture_missing); + } +} + +test "a supported version with no fixture pair is refused" { + var in = migratingGateInput(); + // The starting fixture is there; the version being released has none, so + // the migration it ships was never proved to land anywhere. + in.fixture_versions = &.{1}; + try expectGate2(in, .fixture_pair_absent); +} + +test "the schema version never regresses" { + var in = baseGateInput(); + in.prev_version = 3; + in.prev_minimum = 1; + in.fixture_versions = &.{1}; + try expectGate2(in, .version_regressed); +} + +test "a version bump with neither a step nor a break is refused" { + var in = baseGateInput(); + in.current_version = 2; + in.fixture_versions = &.{ 1, 2 }; + in.changelog_section = migration_section; + try expectGate2(in, .bump_without_step_or_break); +} + +test "a data-only migration must disclose itself even though the schema text held still" { + var in = migratingGateInput(); + in.ddl_changed = false; + in.changelog_section = ""; + // Gate 1 has nothing to say, which is the whole reason Gate 2 runs + // independently of it. + try testing.expectEqual(Gate1.unchanged, gate1(in)); + try expectGate2(in, .migration_undisclosed); + + in.changelog_section = migration_section; + try expectGate2(in, null); +} + +test "the supported minimum never regresses" { + var in = baseGateInput(); + in.prev_minimum = 2; + in.minimum_version = 1; + in.current_version = 2; + in.prev_version = 2; + in.fixture_versions = &.{ 1, 2 }; + try expectGate2(in, .minimum_regressed); +} + +test "raising the minimum is only releasable as the full explicit break" { + // Dropping support for a schema is the one change that silently discards an + // operator's history, so every half-measure below is refused — including the + // one where the schema text did not move at all. + var partial = breakingGateInput(); + partial.changelog_section = migration_section; + try expectGate2(partial, .minimum_raised_without_break); + + var no_heading = breakingGateInput(); + no_heading.changelog_section = "This release " ++ history_reset_phrase ++ ".\n"; + try expectGate2(no_heading, .minimum_raised_without_break); + + var empty_heading = breakingGateInput(); + empty_heading.changelog_section = "This release " ++ history_reset_phrase ++ ".\n\n" ++ + restore_heading ++ "\n\n## [0.0.1] - 2020-01-01\n"; + try expectGate2(empty_heading, .minimum_raised_without_break); + + var same_version = breakingGateInput(); + same_version.current_version = 1; + same_version.minimum_version = 1; + same_version.prev_minimum = 0; + same_version.fixture_versions = &.{1}; + try expectGate2(same_version, .minimum_raised_without_break); + + var unchanged_ddl = breakingGateInput(); + unchanged_ddl.ddl_changed = false; + unchanged_ddl.changelog_section = migration_section; + try expectGate2(unchanged_ddl, .minimum_raised_without_break); +} + +test "the legacy fingerprint is frozen" { + // Recomputing the anchor from a later DDL is the plausible way it gets + // edited, so the substitute is any other CRC-shaped number. + var in = baseGateInput(); + in.legacy_fingerprint = 603440875; + try expectGate2(in, .legacy_fingerprint_edited); + + // And the tree's own constant is the frozen one, which is what makes the + // rule above a check on this repository rather than on its own literal. + try testing.expectEqual(frozen_legacy_fingerprint, querylog_versions.legacy_fingerprint); + + // The DDL has not moved since 0.0.12, so today the anchor and the schema + // fingerprint are the same number. They are not the same THING: the anchor + // is frozen at that value forever, and the fingerprint follows the schema. + try testing.expectEqual(frozen_legacy_fingerprint, querylog_schema.fingerprint); +} + +test "this tree passes both gates against itself" { + // The state every cut starts from: nothing moved since the previous + // release. A tree that cannot pass this has a metadata bug, not a + // disclosure one. + var in = baseGateInput(); + in.current_version = querylog_versions.current; + in.minimum_version = querylog_versions.minimum; + in.legacy_fingerprint = querylog_versions.legacy_fingerprint; + in.prev_version = querylog_versions.current; + in.prev_minimum = querylog_versions.minimum; + + // The tree's own chain, each step paired with itself: reading the file off + // disk is `treeChain`'s job and needs an `Io` this test has no business + // holding. What this covers is the metadata — the chain's LENGTH against the + // supported range — which is the part a self-test can judge. + var chain: std.ArrayList(ChainStep) = .empty; + defer chain.deinit(testing.allocator); + for (querylog_versions.step_sql) |sql| { + try chain.append(testing.allocator, .{ .embedded = sql, .on_disk = sql }); + } + in.chain = chain.items; + + var versions: std.ArrayList(i32) = .empty; + defer versions.deinit(testing.allocator); + var version = querylog_versions.minimum; + while (version <= querylog_versions.current) : (version += 1) { + try versions.append(testing.allocator, version); + } + in.fixture_versions = versions.items; + + try testing.expectEqual(Gate1.unchanged, gate1(in)); + try expectGate2(in, null); +} + +test "a previous tag without the versions module reads as schema version 1" { + // What `git show :src/storage/querylog_versions.zig` hands back is + // nothing at all, and the driver answers 1 and 1 — every file such a release + // created is a version-1 file, which is what the legacy fingerprint stands + // for. This proves the extractor does not invent a number from a file that + // has no such declaration. + try testing.expect(extractVersionConst("pub const ddl = \"\";\n", "current_version") == null); + try testing.expect(extractVersionConst("", "minimum_supported_version") == null); + + const in = baseGateInput(); + try testing.expectEqual(@as(i32, 1), in.prev_version); + try testing.expectEqual(@as(i32, 1), in.prev_minimum); + try testing.expectEqual(Gate1.unchanged, gate1(in)); + try expectGate2(in, null); +} + +test "the version constants of the file on disk are the ones the gate compiled" { + // The same round trip the DDL extractor gets: `git show` will hand this + // text to `extractVersionConst`, so the parse has to agree with the + // compiler on the file it can check. + var threaded: std.Io.Threaded = .init(testing.allocator, .{}); + defer threaded.deinit(); + + var arena_state = std.heap.ArenaAllocator.init(testing.allocator); + defer arena_state.deinit(); + const source = try Io.Dir.cwd().readFileAlloc( + threaded.io(), + querylog_versions_path, + arena_state.allocator(), + .limited(max_input_bytes), + ); + try testing.expectEqual(querylog_versions.current, extractVersionConst(source, "current_version").?); + try testing.expectEqual(querylog_versions.minimum, extractVersionConst(source, "minimum_supported_version").?); + try testing.expectEqual( + querylog_versions.legacy_fingerprint, + extractVersionConst(source, "legacy_fingerprint").?, + ); +} + +test "a fixture name yields its version, and nothing else does" { + try testing.expectEqual(@as(i32, 1), fixtureVersionOf("querylog-v1-schema.sql").?); + try testing.expectEqual(@as(i32, 12), fixtureVersionOf("querylog-v12-data.sql").?); + try testing.expect(fixtureVersionOf("querylog-v1-notes.sql") == null); + try testing.expect(fixtureVersionOf("querylog-schema.sql") == null); + try testing.expect(fixtureVersionOf("config-v1-schema.sql") == null); + try testing.expect(fixtureVersionOf("querylog-vx-data.sql") == null); +} + +test "restore instructions need a heading and something under it" { + try testing.expect(disclosesRestoreInstructions(break_section)); + try testing.expect(!disclosesRestoreInstructions(restore_heading ++ "\n\n")); + try testing.expect(!disclosesRestoreInstructions(restore_heading ++ "\n\n### Something else\nbody\n")); + try testing.expect(!disclosesRestoreInstructions("### Restoring\nbody\n")); + try testing.expect(disclosesRestoreInstructions("intro\n" ++ restore_heading ++ "\n- move it back\n")); +}