diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 2c39b7e..f268a14 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -3220,6 +3220,40 @@ the DEPLOY CLIENT immediately before the deploy**, and this project has already lost three bootstrap attempts to under-carved controller networking. Not covered by the four-empty- subnet approval, and **NOT deploy-blocking**: an extra ULA subnet in MAAS harms nothing. + **>>> SESSION CLOSE 2026-08-02 (part 2), GA-R4 bookend. DURABILITY TRIAD MEASURED. <<<** + vcloud **0 uncommitted, 0 unpushed**; **voffice1 SYNCED** (was 1 behind, cleared at close -- + it is the D-128 Plane-2 host where preflight runs); dc0 rack `~/repo-stage` **all 13 tracked + files MATCH** the repo, so the D-138 deploy client is in sync (this is the check that caught + the stale VIP overlay at the previous close). + **GATES AT CLOSE, quoted:** repo-lint **0 fail, 1 warn** (legacy D-001..018 ASCII carve-out); + `run-tests-all` **GAUNTLET: ALL GREEN (98 harnesses)**, 97 -> 98 for the one harness added + and its manifest recorded deliberately; `ledger-scan` 3 open decisions, **28 open SEC** (none + opened this session), next-free **D-141 / DOCFIX-208 / BUNDLEFIX-053**. **RECONCILED:** DOCFIX + moved 207 -> 208 matching the single number assigned; SEC unchanged at 28 matching zero rows + opened; harness manifest 97 -> 98 matching the one harness added. + **CLOSE SWEEP: `docs/audit/queued-findings-20260802-step6-queued-items.txt` -- SIX FIRST + SURFACE**, each grep-proven absent from every repo surface first. Highest-consequence: a broad + `Bash(ssh vr1-dc0-maas *)` allow rule added to GITIGNORED state, and -- measured by + consequence -- **four destructive MAAS subnet deletes that matched NO ask rule**, because the + committed rules pin the `maas admin` profile and the double-quoted form while this session used + `maas vr1-dc0-region` in single quotes. The deletes were separately operator-approved, so + nothing bypassed consent, but the GATE did not fire; this is the same rule-fails-to-MATCH class + as the 2026-07-30 DHCP cutover and is now RECURRING, not one-off. Then: `--diff=none` does not + mean a full re-download (the mirror is incremental, ~1.3 MB steady state against a 952 GiB + tree); LWP synthesises HTTP 500 for CLIENT-side failures; dc1's GUA carve is INCOMPLETE and + that is an unregistered dc1 work item; and the step-6 end state is NOT re-derivable from the + carve tool's report. + **GA-R7 MEMORY REVIEW: CLEAN** -- zero entries claiming operator policy, priority or posture. + One update: the instrument-currency memory gains instance **thirteen**, a NEW SHAPE (a + comparison that could never match, returning a clean ZERO) plus the meta-lesson that writing + the rule down does not inoculate against it. + **LEDGER ROTATED (GA-R4 rule 3):** 294 lines would have breached the 300 cap; the TWO oldest + closed summaries moved VERBATIM to `docs/archive/session-ledger-rotated-20260802b.md`; ledger + now 287. + **No orphaned session** -- every prior session carries a close bookend, so rule 7 is a no-op. + **NO STAGE OPENED OR CLOSED**; Stage 5 remains OPEN and this is a session bookend. + **NEXT SESSION'S FIRST ITEM: the preflight-P2 / phase4 machines-overlay asymmetry** -- P2 + validates a merged input the deploy never passes -- then the bundle deploy. - Project: Omega Cloud, VR1 DC-DC rehearsal -- a two-DC + Office1-headend virtual rehearsal on KVM (vcloud host), rehearsing the future bare-metal Roosevelt deployment (D-100, `docs/design-decisions.md:1946`). diff --git a/docs/archive/session-ledger-rotated-20260802b.md b/docs/archive/session-ledger-rotated-20260802b.md new file mode 100644 index 0000000..d0df8e6 --- /dev/null +++ b/docs/archive/session-ledger-rotated-20260802b.md @@ -0,0 +1,42 @@ +# Session-ledger rotation 2026-08-02 (b) + +Rotated VERBATIM out of `docs/session-ledger.md` at the second 2026-08-02 close +(GA-R4 rule 3 / F1). The live ledger stood at 294 lines and this close's bounded +summary would have breached the 300-line cap. + +--- + +## SESSION CLOSE 2026-07-30 (part 3) -- STAGE 5 OPENED; three bootstraps; D-138 + D-132 ruled; per-DC MAAS region LIVE at dc0 (bounded, GA-R4) + +- Branch `dc-dc-stage5-preconditions`, **12 commits** pushed (`726127d..`). **STAGE 5 OPENED**, not closed. Scan: 3 decisions, **SEC 23** (SEC-026, -027 opened), **D 139** / DOCFIX 206 / BUNDLEFIX 053. +- **5 RULINGS (GA-R5, quoted):** P5 accepted; bootstrap flags **"Use both flags"**; **D-138** cloud-facing client moves INTO the DC; **D-132 q1** per-DC MAAS region; its placement **VM at utility `.6`** (extends the D-134 octet map: .4 artifact / .5 juju / .6 region). +- **THREE BOOTSTRAP ATTEMPTS, each failing one layer deeper -- every failure a real defect no record would have shown.** (1) no juju-client->node-plane path anywhere (SEC-010 + isolated libvirt nets; never existed). (2) D-138 fixed the client, then the CONTROLLER VM had NO DEFAULT ROUTE -- under-carved, added after both Stage-4 carves. (3) route fixed, agent fetched on attempt 1, then `jujud` could not reach the MAAS region API. **SEC-010 is INCOMPATIBLE with D-104-in-DC-controller + Office1 region**; per-DC regions REMOVE the requirement rather than excepting it, so **SEC-010 stays UNAMENDED**. +- **BUILT:** D-134 `.5`+v6 on BOTH controllers (was ruled-but-not-built); region VM applied dc0 (`2 add/0/0`, MACs pinned, converged ZERO DIFF); **dc0 region LIVE** -- noble + MAAS 3.7.2 + PostgreSQL 16.14 matching Office1, API **200** on `10.12.8.6:5240`, `dhcpd` INACTIVE (no DHCP conflict), jammy+noble **Synced**. +- **G17 dc0 CAPTURED, both assertions PASS** -- the one-shot window opened during a FAILED bootstrap and was taken. Found G17's own text names `chronyc`, absent on the MAAS jammy image, so the gate as written could only REFUSE. +- **CREDENTIALS (operator-directed):** 3 region secrets minted on-VM, never printed, consolidated to `~/vr1-dc0-creds/` (sha256-verified identical), **SEC-027** opened; matrix now EXPECTS them at BOTH DCs (5 rows/DC, new `region` host-role enum, 4 `vm-secret-locations` rows). **dc1's correctly FAIL S2** -- the forward register making absence detectable. Harness 65/65. +- **MEASURED, DO NOT CONFLATE:** MAAS boot images come from `images.maas.io`, NOT the DC mirror (404s on every simplestreams path -- it mirrors apt only). Three artifact classes, ONE local. **The D-135/D-107 narrowing would BREAK image sync AND bootstrap; nothing sequences it after them** (queued F10 -- highest-consequence unruled item). +- **OWNED:** ran the DC checkers on the wrong host; then re-probed egress FROM THE RACK and called the window open when the NODE does the fetching -- same wrong-host class, made right after writing that note; scoped D-138 without enumerating that `jujud` is itself a MAAS client, which is why a 4th blocker appeared after the 3rd was fixed. All corrected on-surface. +- **NOTHING IS HALF-APPLIED.** Office1 still owns all 9 dc0 nodes, correctly carved and `Ready`. Pre-migration carve captured (`dc0-maas-carve-premigration-20260730.txt`, 179 lines) -- MAAS cannot move machines between regions, so the migration must recreate it. +- **AS-EXECUTED LOG IS PARTIAL** -- the classifier refused several `script -aqe` wrapped forms; index row says so (F6). A log that looks complete is worse than one declaring its gap. +- **NEXT:** recreate topology on the new region -> DHCP handover (two servers on one segment is the hazard) -> re-enrol/re-commission 9 nodes -> re-apply statics/br-ex/tags -> re-point Juju at `10.12.8.6:5240`. dc1's region VM authored, NOT applied. Start with fresh context: the destructive half is nine nodes. +- Sweep: `docs/audit/queued-findings-20260730-stage5.txt` (F1-F12). Bodies: `docs/changelog-20260730-stage5-open.md`. Status ONLY in CURRENT-STATE.md. + + + +--- + +## SESSION CLOSE 2026-07-30 (part 4) -- dc0 region topology BUILT 40/0; cutover BLOCKED on a permission wall (bounded, GA-R4) + +- Branch `dc-dc-stage5-preconditions`, **8 commits** pushed (`0ec9c97..b1afa42`). NO stage opened/closed. Scan unchanged: 3 decisions, SEC 23, D 139 / DOCFIX 206 / BUNDLEFIX 053. +- **STAGE 5 REMAINS BLOCKED** -- now on the DHCP cutover, refused by the harness classifier and NOT retried in an altered shape. Exact 4-command sequence, every parameter re-resolved live, is in `changelog-20260730-dc0-region-migration.md` item 10. **NOTHING half-applied** (Office1 verified read-only still `dhcp_on=True primary_rack=7chphy`). +- **NEW dc0 REGION TOPOLOGY LIVE: `dc-region-topology.sh check` = 40 assertions / 0 failed, EXIT 0.** 5 named plane fabrics, 6 spaces, 6 v4 subnets + gateways, 6 VLAN->space bindings, site tag, metal-admin dns/dynamic range. Power key installed + proven (9/9, real virsh enumerating 12 domains -- the snap ssh dir did not exist at all, so commissioning could not have powered a node on). +- **THREE TOOLS THE REPO NEVER HAD:** `maas-profile-assert.sh` (region identity by RACK IDENTITY -- a machine count is not proof), `maas-region-power-key.sh`, `dc-region-topology.sh`. A survey found the named fabrics, v4 subnets and site tag were built AD-HOC in the Stage-4 window, logged only to an as-executed file NOT in the repo -- there was nothing to re-run. dc1 now rebuilds from tools. +- **I INTRODUCED A DEFECT THAT PASSED MY OWN GATE 39/39.** The first apply CREATED the provider-public fabric and MOVED the subnet; MAAS does not bring interface links along, leaving the region VM's `enp2s0` on a subnet-less VLAN -- the same under-carve class that cost three bootstrap attempts. Live connectivity was unaffected so nothing flagged it. Fixed live + in the tool (RENAME, never move) + in the gate (stranded-link assertion). +- **A GATE WAS ALREADY RED AT HEAD:** `tests/node-vm` T8/T9 (11/66 vs assertions reading 10/60) -- commits `086c827`+`447315f` added the region VMs and the gauntlet was never re-run. PROVEN pre-existing by stashing this session's work. Re-pointed to the new invariant. +- **NEARLY RECORDED A FALSE FINDING:** a `+time=3 +tries=1` probe said the new region's BIND could not resolve externally, which would have forced keeping the cross-fiber DNS dependency. Cold-cache timeout; at `+time=5 +tries=3` it resolves everything. Node DNS now points at the DC-local `10.12.8.6`. +- **MEASURED, load-bearing:** VLAN row id 5005 is dc0 metal-admin in Office1 and `vr1-dc0-data-tenant` in the new region; `hot-kid` is `c3aqh8` there and `tw7ptw` in Office1. Ids do not cross regions. NIC->plane order recorded in `lib-hosts.sh` and is NOT `PLANE_CIDRS` order (walking that positionally strands PXE on provider-public). +- **`dc-node-v6-carve.py` FIXED, not just noted** -- it was env-blind (`export MAAS_PROFILE` silently hit OFFICE1) and tracebacked on an absent CLI. Harness 14/14. +- **DELIBERATELY NOT WRITTEN: `dc-node-carve.sh`** (60 NIC re-homes + 9 `br-ex` + 54 statics). Both defects found today were caught by LIVE runs, not fixtures, and it cannot be exercised until the nodes are re-enrolled. Its two hard inputs are now settled and recorded. +- **OWED:** wire `maas-profile-assert.sh` into `dc-plane-ipam.sh` / `maas-role-tags.sh` / `maas-node-power.sh` (all still default to `admin`, where dc1's nine nodes live); rack DNS forwarder upstream; remove the `vr1-dc0-region` profile from voffice1 (SEC-026) and Office1's dc0 power-key copy. +- Gauntlet **ALL GREEN (92)**; repo-lint 0 fail. AS-EXECUTED LOG NOT USED -- `run-logged.sh` needs an interactive shell; captures went to `docs/audit/*` and this is declared. Body: `docs/changelog-20260730-dc0-region-migration.md`. Status ONLY in CURRENT-STATE.md. + diff --git a/docs/audit/queued-findings-20260802-step6-queued-items.txt b/docs/audit/queued-findings-20260802-step6-queued-items.txt new file mode 100644 index 0000000..d96d69c --- /dev/null +++ b/docs/audit/queued-findings-20260802-step6-queued-items.txt @@ -0,0 +1,181 @@ +queued-findings-20260802-step6-queued-items.txt +================================================ +Close sweep for the SECOND 2026-08-02 session (post-/clear): the queued-findings +backlog (sweep F1-F6), the mirror root-cause, and D-139 execution step 6. + +Method (ruled 2026-07-31): read back over the whole session, enumerate every +finding/decision/measurement/mistake, then GREP each candidate against repo +surfaces. A hit = ALREADY ON SURFACE, and where. No hit = FIRST SURFACE and would +have been lost on a context clear. + +Session body: docs/changelog-20260802-queued-items.md (items 1-12). +Status claims live in docs/CURRENT-STATE.md ONLY. + +NUMBERING NOTE, because two registers collide: items below cite the SWEEP register +(F1-F6, docs/audit/queued-findings-20260802-stage5-edge-fold.txt). The runbook fold +register has its OWN F1-F12 with different meanings and is untouched this session. + +-------------------------------------------------------------------------------------- +FIRST SURFACE -- existed ONLY in the transcript. Listed first, by consequence. +-------------------------------------------------------------------------------------- + +G1. >>> A BROAD `Bash(ssh vr1-dc0-maas *)` ALLOW RULE WAS ADDED TO GITIGNORED STATE. <<< + grep "ssh vr1-dc0-maas" in docs/: 0 hits. + `.claude/settings.local.json` was MODIFIED this session (mtime 17:33) and now + holds allow=310 / ask=11 / deny=0. The 2026-08-02 (first session) sweep recorded + the counts as UNCHANGED from the 2026-07-30 verbatim record, so the growth is + this session's. + THE ITEM WORTH ATTENTION: `Bash(ssh vr1-dc0-maas *)` is a WILDCARD allow on the + MAAS REGION VM -- the same shape as the broad `Bash(ssh vr1-dc0-rack *)` that + the 2026-07-30 sweep flagged as "worth review". It was auto-added by an approval + during the pg_dump work. It permits ANY command on the host holding the region + database. Recommend narrowing to the read-only forms actually needed. + The file is GITIGNORED, so this text is the only recovery copy. The 11 ASK rules + verbatim (these are the gating ones and matter most): + Bash(script -aqe ~/as-executed/2026-07-23-stage4-carve.log -c 'ssh voffice1 "maas admin machine *) + Bash(ssh -i ~/vr1-dc0-creds/vr1-dc0_svc_ed25519 -J voffice1 jessea123@172.31.0.2 "sudo *) + Bash(ssh voffice1 "maas admin interface *) + Bash(ssh voffice1 "maas admin subnet *) + Bash(ssh voffice1 "maas admin tags *) + Bash(ssh voffice1 "maas admin tag *) + Bash(ssh voffice1 "maas admin machine *) + Bash(ssh *'maas * update*) + Bash(ssh *"maas * update*) + Bash(ssh *'maas * ipranges *) + Bash(ssh *"maas * ipranges *) + NOTE A GAP IN THOSE ASK RULES, measured by consequence this session: they pin the + `maas admin` profile and the `ssh voffice1 "..."` double-quoted form. This + session's MAAS SUBNET DELETES used `MAAS_PROFILE=vr1-dc0-region ... maas + vr1-dc0-region subnet delete` in SINGLE quotes and therefore did NOT match + `Bash(ssh voffice1 "maas admin subnet *)`. Four destructive deletes ran without + hitting an ask rule. This is the SAME rule-fails-to-MATCH class recorded on + 2026-07-30 (the DHCP cutover), and it is now recorded as recurring rather than + one-off. The deletes were separately operator-approved, so nothing bypassed + consent -- but the GATE did not fire, and that is the finding. + +G2. `--diff=none` DOES NOT MEAN "DOWNLOAD EVERYTHING", AND THE MIRROR IS INCREMENTAL. + grep "diff=none": 3 hits, all either the flag itself (dc-mirror.sh:187,193) or a + passing mention; NO hit explains the semantics. + MEASURED this session from the sync journal: + routine daily run ubuntu leg 1277 kiB + the failing run ubuntu leg 2867 kiB (313s, of which 300s was one timeout) + UCA leg, every run 15 kiB + on-disk ubuntu 952 G, cloud-archive 342 M + the 1048 MiB outlier (2026-07-31) = the day jammy-backports was ADDED to the + suite list by operator ruling; adding a suite pulls its content ONCE. + `--diff=none` disables PDIFFS -- the incremental-patch mechanism for INDEX files + -- so debmirror fetches each Release/Packages/dep11 index in full every run. The + POOL sync is still differential throughout. So a re-trigger costs seconds and + single-digit MB, not a 950 G pull. + WHY IT MATTERS: someone reading the flag name could refuse to re-trigger a sync + believing it means a full re-download, or could budget hours for it. + +G3. LWP SYNTHESISES HTTP 500 FOR CLIENT-SIDE FAILURES -- IT IS NOT A SERVER 500. + grep "synthesises 500": 1 hit, in this session's own root-cause capture only; + grep "synthesizes 500": 0 hits. Not on any DURABLE surface (platform-traps). + debmirror's `500 read timeout` is LWP reporting ITS OWN timeout + (/usr/bin/debmirror:629, `our $timeout=300;`), not archive.ubuntu.com returning + an error. Reading it as a server-side 500 sends the investigation upstream for + the wrong reason, which is exactly what happened here before it was corrected. + Belongs in platform-traps' verbatim-error index. LOGGED, NOT FIXED (hard rule 1). + +G4. dc1's GUA CARVE IS INCOMPLETE, AND IT IS A dc1 BLOCKER NOT JUST A TOOL FIXTURE. + grep "f03:20" / "dc1.*GUA.*incomplete": hits only inside this session's own + changelog item 10 and CURRENT-STATE, i.e. recorded as the REASON for a tool + guard rather than as a dc1 WORK ITEM. + MEASURED: only FOUR rows exist under `2602:f3e2:f03::/48`, all provider-public. + No `:f03:20::/64`, no `:f03:21::/64`. So for dc1, D-139 steps 1 and 2 have NOT + been run. Anyone reaching dc1's Stage 5 must run `d139-gua-carve.py --dc vr1-dc1 + --commit` (and the MAAS half) BEFORE step 6 -- the step-6 tool now refuses, so + the failure is loud, but the WORK is unregistered. + +G5. THE STEP-6 END STATE IS NOT RE-DERIVABLE FROM THE CARVE TOOL'S REPORT. + grep "counts existence, not status": 2 hits, both this session's own records. + `d139-gua-carve.py --dc vr1-dc0` STILL reports `RETIRE-REPORT 9` and `dependent + objects ... 26 ip-address(es)` AFTER step 6 completed successfully. That is + correct -- the rows still EXIST, deprecated not deleted, per the ruling -- but a + future session running the carve tool to check progress will read it as "step 6 + never ran". The carve tool has no status awareness and was not extended (hard + rule 1). A `--commit`-less step-6 run is the correct instrument: it reports + `CREATE 0 | ALREADY 26 | DEPRECATE-ADDR 0 | DEPRECATE-PFX 0`. + +G6. NO AS-EXECUTED LOG COVERS THIS WINDOW. + `run-logged.sh` was NOT used -- it needs an interactive shell. The newest file in + ~/as-executed/ is 2026-07-30-stage5-dc0-deploy.log. This session executed LIVE + MUTATIONS (26 apex creates, 35 apex deprecations, 4 MAAS subnet deletes, one + mirror sync trigger, one rack file copy) with NO as-executed coverage. All + evidence is in docs/audit/* captures and the commit messages. DECLARED here + rather than left to be discovered -- a log that looks complete and is not is + worse than one declaring its gap. + +-------------------------------------------------------------------------------------- +ALREADY ON SURFACE -- verified by grep, recorded for completeness +-------------------------------------------------------------------------------------- + +A1. DOCFIX-207, preflight P6 plan count 50/97 -> 56/108 -- scripts/preflight.sh + + changelog item 1 + CURRENT-STATE. +A2. sweep F3 (`systemctl show` is not an existence check) -- platform-traps 5c. +A3. sweep F4 (assert the harness CASE COUNT moved) + sweep F5 (two scripts probing + one endpoint must share the probe definition) -- script-authoring. +A4. sweep F1 CLOSED, rack deploy input re-staged, full 14-file enumeration -- + docs/audit/repo-stage-drift-dc0-20260802.txt + changelog items 4-5. +A5. sweep F6 CLOSED PASS, maasdb proven uncorrupted, and F6's own diagnosis being + wrong (role + transport, not confinement) -- + docs/audit/maasdb-pgdump-integrity-dc0-20260802.txt sections 7-11. +A6. sweep F2 diagnosed R1, and the sweep's premise being wrong (it quoted the + PASSING UCA leg) -- docs/audit/mirror-exitcode-diagnosis-dc0-20260802.txt. +A7. The mirror root cause: archive.ubuntu.com backend 91.189.92.23 hangs on one + object while serving its siblings; resolver rotates; 11/12 succeed -- + docs/audit/mirror-500-timeout-rootcause-20260802.txt. +A8. Four GA-R5/operational rulings with exact utterances -- "Re-trigger the sync + first, decide after"; "Root-cause the curl/debmirror anomaly first"; "Full step + 6 first, then deploy"; "Deprecate both, delete nothing" -- design-decisions.md + + CURRENT-STATE. Plus the two approvals ("Queue up the MAAS half ... Go ahead + with both now"; "Proceed with 1 and 2") in changelog items 11-12. +A9. The step-6 tool, its adversarial review and all four DEFects -- + docs/audit/d139-step6-tool-review-20260802.txt + changelog item 10. +A10. Step 6 EXECUTED, 4 MAAS subnets deleted, 1 HELD on the juju-controller / + region-VM hazard -- changelog item 11 + CURRENT-STATE. +A11. `retire-v6-ula` and the broken by-hand link check it caught -- changelog item + 12 + CURRENT-STATE + tests/dc-plane-ipam R1-R7. +A12. The preflight-P2 / phase4 machines-overlay asymmetry -- changelog item 5. +A13. SEC-029(3) recording ~/repo-stage as nine files when it is fourteen -- + changelog item 5. +A14. The netbox env FILENAME INVERSION (vr1-netbox-sandbox.env is the LIVE apex; + vr1-netbox.env is the v1 reference) -- CURRENT-STATE:4638-4641, pre-existing. +A15. The DOCFIX decoy-token counter inflation, twice, and its correction -- + CURRENT-STATE + commit 93973b6. + +-------------------------------------------------------------------------------------- +THE FIVE STRUCTURAL SWEEPS +-------------------------------------------------------------------------------------- + +S1. GITIGNORED STATE. `.claude/settings.local.json` MODIFIED this session: + allow=310 / ask=11 / deny=0. Verbatim ask rules and the broad-allow finding are + in G1 above. This sweep file is the only recovery copy. +S2. DANGLING REFERENCES. Every docs/audit path cited by this session's commits + resolves; the four new captures are committed + (repo-stage-drift, maasdb-pgdump-integrity, mirror-exitcode-diagnosis, + mirror-500-timeout-rootcause, d139-step6-tool-review). +S3. RULING FIDELITY. Four rulings + two approvals this session, each with its EXACT + utterance quoted, dated, and -- for the two GA-R5 step-6 rulings -- committed + and pushed BEFORE the dependent work began (934a1f0, 53aae78). None paraphrased. + Both step-6 rulings were correctly classed OPS (GA-R3) rather than new D-numbers. +S4. AS-EXECUTED LOG. NOT USED. See G6 -- this window is uncovered and it is declared. +S5. CONTRADICTION DETECTOR. + (a) The 2026-08-01 ordering ruling's stated premise ("the per-DC VIP overlays + carry v6 VIPs in the ULA range") is FALSE as of 2026-08-02; the overlay was + re-rendered onto GUA. Corrected in the 2026-08-02 ordering ruling itself + rather than left to contradict silently. + (b) The step-6 amendment first said the GUA records would be `status=active`; + MEASURED, the originals are `reserved`. Corrected in both design-decisions + and CURRENT-STATE under GA-R1 C2. + (c) "The gate is FLAKY, not stuck" (written mid-session about the mirror) does + NOT extend to the dep11 failure, which is persistent-per-backend. Withdrawn + and corrected in commit 4f65c09 rather than left standing. + (d) "Every action this tool takes is reversible" (step-6 docstring) was wrong for + the 26 CREATES. Corrected in the docstring and in the presentation. + (e) "Proven safe twice" (the four MAAS deletes) was HALF FALSE -- the link half + of that proof was an inert checker. Corrected in changelog item 12 and + CURRENT-STATE, with the safe OUTCOME recorded separately from the unsound + VERIFICATION. diff --git a/docs/session-ledger.md b/docs/session-ledger.md index f0c5515..88fac4d 100644 --- a/docs/session-ledger.md +++ b/docs/session-ledger.md @@ -172,35 +172,12 @@ owed, because the next session is directed straight at the juju deployment. Sessions from the 2026-07-27 Phase-0 close onward remain live below. -## SESSION CLOSE 2026-07-30 (part 3) -- STAGE 5 OPENED; three bootstraps; D-138 + D-132 ruled; per-DC MAAS region LIVE at dc0 (bounded, GA-R4) +## ROTATED 2026-08-02 (b) (GA-R4 rule 3 / F1 -- cap restored at this close) -- Branch `dc-dc-stage5-preconditions`, **12 commits** pushed (`726127d..`). **STAGE 5 OPENED**, not closed. Scan: 3 decisions, **SEC 23** (SEC-026, -027 opened), **D 139** / DOCFIX 206 / BUNDLEFIX 053. -- **5 RULINGS (GA-R5, quoted):** P5 accepted; bootstrap flags **"Use both flags"**; **D-138** cloud-facing client moves INTO the DC; **D-132 q1** per-DC MAAS region; its placement **VM at utility `.6`** (extends the D-134 octet map: .4 artifact / .5 juju / .6 region). -- **THREE BOOTSTRAP ATTEMPTS, each failing one layer deeper -- every failure a real defect no record would have shown.** (1) no juju-client->node-plane path anywhere (SEC-010 + isolated libvirt nets; never existed). (2) D-138 fixed the client, then the CONTROLLER VM had NO DEFAULT ROUTE -- under-carved, added after both Stage-4 carves. (3) route fixed, agent fetched on attempt 1, then `jujud` could not reach the MAAS region API. **SEC-010 is INCOMPATIBLE with D-104-in-DC-controller + Office1 region**; per-DC regions REMOVE the requirement rather than excepting it, so **SEC-010 stays UNAMENDED**. -- **BUILT:** D-134 `.5`+v6 on BOTH controllers (was ruled-but-not-built); region VM applied dc0 (`2 add/0/0`, MACs pinned, converged ZERO DIFF); **dc0 region LIVE** -- noble + MAAS 3.7.2 + PostgreSQL 16.14 matching Office1, API **200** on `10.12.8.6:5240`, `dhcpd` INACTIVE (no DHCP conflict), jammy+noble **Synced**. -- **G17 dc0 CAPTURED, both assertions PASS** -- the one-shot window opened during a FAILED bootstrap and was taken. Found G17's own text names `chronyc`, absent on the MAAS jammy image, so the gate as written could only REFUSE. -- **CREDENTIALS (operator-directed):** 3 region secrets minted on-VM, never printed, consolidated to `~/vr1-dc0-creds/` (sha256-verified identical), **SEC-027** opened; matrix now EXPECTS them at BOTH DCs (5 rows/DC, new `region` host-role enum, 4 `vm-secret-locations` rows). **dc1's correctly FAIL S2** -- the forward register making absence detectable. Harness 65/65. -- **MEASURED, DO NOT CONFLATE:** MAAS boot images come from `images.maas.io`, NOT the DC mirror (404s on every simplestreams path -- it mirrors apt only). Three artifact classes, ONE local. **The D-135/D-107 narrowing would BREAK image sync AND bootstrap; nothing sequences it after them** (queued F10 -- highest-consequence unruled item). -- **OWNED:** ran the DC checkers on the wrong host; then re-probed egress FROM THE RACK and called the window open when the NODE does the fetching -- same wrong-host class, made right after writing that note; scoped D-138 without enumerating that `jujud` is itself a MAAS client, which is why a 4th blocker appeared after the 3rd was fixed. All corrected on-surface. -- **NOTHING IS HALF-APPLIED.** Office1 still owns all 9 dc0 nodes, correctly carved and `Ready`. Pre-migration carve captured (`dc0-maas-carve-premigration-20260730.txt`, 179 lines) -- MAAS cannot move machines between regions, so the migration must recreate it. -- **AS-EXECUTED LOG IS PARTIAL** -- the classifier refused several `script -aqe` wrapped forms; index row says so (F6). A log that looks complete is worse than one declaring its gap. -- **NEXT:** recreate topology on the new region -> DHCP handover (two servers on one segment is the hazard) -> re-enrol/re-commission 9 nodes -> re-apply statics/br-ex/tags -> re-point Juju at `10.12.8.6:5240`. dc1's region VM authored, NOT applied. Start with fresh context: the destructive half is nine nodes. -- Sweep: `docs/audit/queued-findings-20260730-stage5.txt` (F1-F12). Bodies: `docs/changelog-20260730-stage5-open.md`. Status ONLY in CURRENT-STATE.md. - -## SESSION CLOSE 2026-07-30 (part 4) -- dc0 region topology BUILT 40/0; cutover BLOCKED on a permission wall (bounded, GA-R4) - -- Branch `dc-dc-stage5-preconditions`, **8 commits** pushed (`0ec9c97..b1afa42`). NO stage opened/closed. Scan unchanged: 3 decisions, SEC 23, D 139 / DOCFIX 206 / BUNDLEFIX 053. -- **STAGE 5 REMAINS BLOCKED** -- now on the DHCP cutover, refused by the harness classifier and NOT retried in an altered shape. Exact 4-command sequence, every parameter re-resolved live, is in `changelog-20260730-dc0-region-migration.md` item 10. **NOTHING half-applied** (Office1 verified read-only still `dhcp_on=True primary_rack=7chphy`). -- **NEW dc0 REGION TOPOLOGY LIVE: `dc-region-topology.sh check` = 40 assertions / 0 failed, EXIT 0.** 5 named plane fabrics, 6 spaces, 6 v4 subnets + gateways, 6 VLAN->space bindings, site tag, metal-admin dns/dynamic range. Power key installed + proven (9/9, real virsh enumerating 12 domains -- the snap ssh dir did not exist at all, so commissioning could not have powered a node on). -- **THREE TOOLS THE REPO NEVER HAD:** `maas-profile-assert.sh` (region identity by RACK IDENTITY -- a machine count is not proof), `maas-region-power-key.sh`, `dc-region-topology.sh`. A survey found the named fabrics, v4 subnets and site tag were built AD-HOC in the Stage-4 window, logged only to an as-executed file NOT in the repo -- there was nothing to re-run. dc1 now rebuilds from tools. -- **I INTRODUCED A DEFECT THAT PASSED MY OWN GATE 39/39.** The first apply CREATED the provider-public fabric and MOVED the subnet; MAAS does not bring interface links along, leaving the region VM's `enp2s0` on a subnet-less VLAN -- the same under-carve class that cost three bootstrap attempts. Live connectivity was unaffected so nothing flagged it. Fixed live + in the tool (RENAME, never move) + in the gate (stranded-link assertion). -- **A GATE WAS ALREADY RED AT HEAD:** `tests/node-vm` T8/T9 (11/66 vs assertions reading 10/60) -- commits `086c827`+`447315f` added the region VMs and the gauntlet was never re-run. PROVEN pre-existing by stashing this session's work. Re-pointed to the new invariant. -- **NEARLY RECORDED A FALSE FINDING:** a `+time=3 +tries=1` probe said the new region's BIND could not resolve externally, which would have forced keeping the cross-fiber DNS dependency. Cold-cache timeout; at `+time=5 +tries=3` it resolves everything. Node DNS now points at the DC-local `10.12.8.6`. -- **MEASURED, load-bearing:** VLAN row id 5005 is dc0 metal-admin in Office1 and `vr1-dc0-data-tenant` in the new region; `hot-kid` is `c3aqh8` there and `tw7ptw` in Office1. Ids do not cross regions. NIC->plane order recorded in `lib-hosts.sh` and is NOT `PLANE_CIDRS` order (walking that positionally strands PXE on provider-public). -- **`dc-node-v6-carve.py` FIXED, not just noted** -- it was env-blind (`export MAAS_PROFILE` silently hit OFFICE1) and tracebacked on an absent CLI. Harness 14/14. -- **DELIBERATELY NOT WRITTEN: `dc-node-carve.sh`** (60 NIC re-homes + 9 `br-ex` + 54 statics). Both defects found today were caught by LIVE runs, not fixtures, and it cannot be exercised until the nodes are re-enrolled. Its two hard inputs are now settled and recorded. -- **OWED:** wire `maas-profile-assert.sh` into `dc-plane-ipam.sh` / `maas-role-tags.sh` / `maas-node-power.sh` (all still default to `admin`, where dc1's nine nodes live); rack DNS forwarder upstream; remove the `vr1-dc0-region` profile from voffice1 (SEC-026) and Office1's dc0 power-key copy. -- Gauntlet **ALL GREEN (92)**; repo-lint 0 fail. AS-EXECUTED LOG NOT USED -- `run-logged.sh` needs an interactive shell; captures went to `docs/audit/*` and this is declared. Body: `docs/changelog-20260730-dc0-region-migration.md`. Status ONLY in CURRENT-STATE.md. +The TWO oldest live summaries (2026-07-30 part 3 -- Stage 5 opened, three bootstraps, +D-138 + D-132 ruled; and part 4 -- dc0 region topology built, cutover blocked on a +permission wall) moved VERBATIM to `docs/archive/session-ledger-rotated-20260802b.md`. +The live ledger stood at 294 lines and this close's summary would have breached the cap. ## SESSION CLOSE 2026-07-30 (part 5) -- dc0 MIGRATED to its own MAAS region: 10/10 Ready, 4 named gates green (bounded, GA-R4) @@ -292,3 +269,19 @@ - **OWNED:** I diagnosed the edge break as a partial update from the symptom's SHAPE and was wrong; my agent brief carried a **wrong base-image path** where the apply destroys the volume first and no rollback exists; I guessed `/srv/mirror/ubuntu` and a systemd unit name the repo already defines; two harness cases I wrote never ran while the suite said ALL PASS; one assertion passed on its own comment; and I pushed a red lint once by masking the exit code. - Gauntlet **ALL GREEN (97)**, repo-lint 0 fail / 1 warn, ledger-scan reconciled. **voffice1 1 commit behind** at close (not a loss). - **NEXT:** re-stage the rack's VIP overlay (sweep F1 -- it is STALE and is the deploy input), settle the mirror's exit-1-with-"All done" (F2), `pg_dump maasdb` (F6), then fold F2-F11 and the dc1 Phase-2 exercise. Sweep: `docs/audit/queued-findings-20260802-stage5-edge-fold.txt` (**6 FIRST SURFACE**). Status ONLY in CURRENT-STATE.md. + +## SESSION CLOSE 2026-08-02 (part 2) -- queued backlog cleared; mirror ROOT-CAUSED; D-139 step 6 EXECUTED (bounded, GA-R4) + +- Branch `dc-dc-stage5-preconditions`, **9 commits** pushed (`1cdd607..56b37f8`). NO stage opened or closed. Scan: 3 open decisions, SEC **28** (none opened this session), D 141 / **DOCFIX 208** / BUNDLEFIX 053 -- DOCFIX moved 207->208, reconciling with the one number assigned. +- **Sweep F1 and F6 CLOSED; F2 diagnosed then ROOT-CAUSED; F3/F4/F5 graduated to platform-traps + script-authoring; DOCFIX-207** corrected preflight P6's plan count (50/97, stale since 2026-07-10, against a measured 56/108). +- **F6 PASS -- the dc0 region DB is proven uncorrupted:** pg_dump read every page of `maasdb` (23,878,796 bytes / 37,199 lines / exit 0 / completion marker). **F6's own stated blocker was WRONG** -- the discriminators are ROLE and TRANSPORT, not snap confinement; over the unix socket the `maas` role needs no credential at all. +- **F2 ROOT CAUSE IS UPSTREAM:** one of NINE `archive.ubuntu.com` backends (`91.189.92.23`) hangs on ONE dep11 object while serving its directory siblings in 0.5s; the resolver rotates, and 11 of 12 fetches succeed. **"Not transient" WITHDRAWN.** apt is unaffected -- it fetches the `.xz`, which is present; `apt-get update` against the mirror returns rc=0. +- **4 rulings, exact utterances:** *"Re-trigger the sync first, decide after"*; *"Root-cause the curl/debmirror anomaly first"*; **"Full step 6 first, then deploy"**; **"Deprecate both, delete nothing"**. Both step-6 rulings were pushed BEFORE the dependent work (`934a1f0`, `53aae78`) and correctly classed OPS, not new D-numbers. +- **>>> D-139 STEP 6 EXECUTED -- the deploy's last stated blocker. <<<** Apex: 26 GUA VIP addresses created, 26 ULA addresses + 9 ULA prefixes deprecated, nothing deleted; idempotent on re-run. The 26 CREATE targets diff EXACTLY against the deploy overlay's 26 GUA VIP legs. +- **MAAS half: 4 of 5 ULA subnets deleted, 1 HELD.** `fd50:840e:74e2:220::/64` carries the juju controller (`::5`) and the MAAS region VM (`::6`), neither with a GUA counterpart -- deleting it would strip the deploy client's only recorded v6. +- **Two tools shipped:** `netbox/d139-step6-vip-rehome.py` (harness 20 cases) and `dc-plane-ipam.sh retire-v6-ula` (harness 25->32). An adversarial review returned **FIX FIRST on four defects, two CRITICAL** (a dc1 orphan-create; a dropped apex-identity guard) -- all fixed and verified live. +- **OWNED -- THREE of my checkers COULD NOT FAIL**, every one written AFTER I landed that exact rule into script-authoring this session: an assertion satisfied by a traceback; a grep covering one file while the tool inherited the other; and a `sid="'$id'"` comparison that returned a clean ZERO, on which four deletes proceeded. **None was caught by re-reading my own work** -- two by an adversarial reviewer, one by the live run. +- **Also owned:** called the four deletes "proven safe twice" when half that proof was inert (the OUTCOME was safe -- measured afterwards, nodes read v4=6 v6=6); wrote `status=active` into a ruling by inference (measured: `reserved`); and inflated the DOCFIX counter with a decoy token TWICE, the second time inside the sentence correcting the first. +- **Durability:** vcloud 0 uncommitted / 0 unpushed; **voffice1 synced** (was 1 behind); dc0 rack `~/repo-stage` all 13 tracked files MATCH the repo. Gates: gauntlet **ALL GREEN (98)**, repo-lint 0 fail / 1 legacy warn. +- **NEXT:** the preflight-P2 / phase4 machines-overlay asymmetry -- P2 validates a merged input the deploy never passes -- then the bundle deploy. The held subnet needs the controller's v6 re-homed to GUA first and is NOT deploy-blocking. +- Sweep: `docs/audit/queued-findings-20260802-step6-queued-items.txt` (**6 FIRST SURFACE**, incl. a broad `Bash(ssh vr1-dc0-maas *)` allow rule, and four destructive MAAS deletes that matched NO ask rule -- the rule-fails-to-MATCH class, now recurring). Body: `docs/changelog-20260802-queued-items.md`. Status ONLY in CURRENT-STATE.md.