diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 84f8d78..ae134e5 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -69,14 +69,22 @@ > redeploy (specs in days); this teardown/redeploy is the test of the module-deployment project, > and it compresses the Roosevelt-review Group-4 horizon from far-future to weeks-away (re-run that > assessment against real hardware sheets when they land). -> **CONTAINER-ELIM PASS SCOPED (2026-08-09), execution DEFERRED to a NEW session.** A multi-agent -> READ-ONLY planning pass (phases 0-4; standard 3-5 sonnet workers + 1 fable administrator/phase; -> final fable advisor) is fully scoped as a cold-start handoff: -> `docs/audit/container-elim-pass/SCOPE-AND-EXECUTION-PLAN.md` (+ `phase-prompts.md`). Locked: -> phases 0-4, layered module system (IaC + procedure), Phase-0 proposes the target topology -> -> operator CONFIRMS before Phase 1 (hard gate), operator gets ONE report at the end. Goal: plan the -> dc0/dc1 container-layer elimination for the 10.13 redeploy AND turn the redeploy steps into a -> repeatable layered module workflow feeding the pre-Roosevelt bare-metal test. NOT YET RUN. +> **CONTAINER-ELIM PASS RAN (2026-08-09) -- READ-ONLY planning complete; the [ARCH] ruling is now OWED.** +> The multi-agent read-only pass (phases 0-4, 4 sonnet workers + 1 administrator per phase + a final +> advisor; agents recorded as "the administrator"/"the advisor", no model name asserted per operator) +> executed this session. NOTHING was built or executed against the cloud -- it produced a PLAN. +> **Phase-0 operator gate:** operator confirmed **Option 1** (flat node VMs on vcloud libvirt + a small +> per-DC `vr1-dcN-client` VM for the D-138 client role) + cross-DC handling **(a)** (new vcloud +> host-level isolation control) -- a DIRECTIONAL planning confirmation, NOT the formal GA-R5 [ARCH] +> ruling. Deliverables in `docs/audit/container-elim-pass/`: **`operator-report.md`** (the single +> report), **`FINAL-PLAN.md`** (104-row change-set + L0-L5 module design + Part-A/B sequencing + +> decision package), **`FINAL-advisor-review.md`** (verdict SOUND + 2 follow-ups), + 20 worker/admin +> docs. **OWED (operator, GA-R5): the Tier-1 decision package** -- DEC-01 = adopt the container-elim as +> a new **D-144** (recommended; supersedes D-123 Model B; D-143 is the precedent) + amendments D-128/ +> D-134/D-131 + DEC-15 (power-key mitigation incl. reach) + DEC-14/16 ((a) control) + DEC-11 (root +> topology (B)) + DEC-08 (rack retirement, live-re-measure-gated) + DEC-24 (dc0<->dc1 mesh/Ceph path). +> Model B remains the LIVE shape until D-144 is ruled and the redeploy executes. Scope handoff (history): +> `docs/audit/container-elim-pass/SCOPE-AND-EXECUTION-PLAN.md`. > > **dc0 checkpoint scope (operator 2026-08-08): "activate + smoke-test"** -- networks + Octavia (1 test > LB) + Designate (1 test zone) + wrap gates (cloud-assert BOM, controller backup, verify-live diff --git a/docs/audit/container-elim-pass/FINAL-PLAN.md b/docs/audit/container-elim-pass/FINAL-PLAN.md new file mode 100644 index 0000000..0fce588 --- /dev/null +++ b/docs/audit/container-elim-pass/FINAL-PLAN.md @@ -0,0 +1,471 @@ +# FINAL PLAN -- container-layer elimination + layered module workflow (Phase-4 consolidated plan) + +**Author:** the Phase-4 administrator (multi-agent container-elim pass, +`SCOPE-AND-EXECUTION-PLAN.md` Section 4; no model name asserted, operator instruction). +**Date:** 2026-08-09. **Inputs read in full this session:** `SCOPE-AND-EXECUTION-PLAN.md`; +the four prior administrator reports (`pass0-admin-report.md` .. `pass3-admin-report.md`); +the four Phase-4 worker docs (`pass4-w1-master-change-inventory.md`, +`pass4-w2-module-workflow-design.md`, `pass4-w3-execution-sequencing.md`, +`pass4-w4-decision-framing.md`). **READ-ONLY planning synthesis** -- no mutation performed; +every owed artifact below is LOGGED, not built (hard rule 1). All recommendations are +graded "recommend," never "ruled" (GA-R5). This document is the input to the final advisor +review (`FINAL-advisor-review.md`) and the operator report. + +--- + +## 1. The confirmed target + as-is -> to-be + +**Target (operator-confirmed at the Phase-0 gate, 2026-08-09 -- a DIRECTIONAL PLANNING +CONFIRMATION, not yet the GA-R5 [ARCH] ruling; that ruling is Section 5's package):** + +- **Option 1:** flat node VMs on vcloud libvirt + one small non-hypervisor + `vr1-dcN-client` VM per DC (metal-admin + transit legs) carrying the D-138 client role + and that DC's credential residencies. +- **Cross-DC adjacency handling (a):** accept co-residency + a NEW vcloud-level host + isolation control (SEC-010's nftables pattern one layer up) with a mechanical `--check` + gate and its own SEC row. +- **MAAS region stays on `vr1-dcN-maas-01`** (.6, D-132 addendum -- no change). + +**As-is (Model B, D-122/D-123):** `vcloud (outer libvirt) -> vvr1-dcN (containment VM = +inner libvirt) -> node VMs`. Two tofu roots per DC + a bash bootstrap gate between them; a +qemu+ssh provider dial from voffice1 into the containment VM (D-126 keys, D-128 Plane 2); +D-125 bridge-in WAN plumbing (`modules/wan-bridge`, IP-less uplink NIC, `br-vr1-dcN-wan`); +SEC-010 FORWARD-drop on the containment VM's transit leg; nesting depth 4. + +**To-be (Option 1, flat 10.13):** `vcloud (libvirt) -> node VMs` directly -- depth 2 +(VR0-proven). Per DC: six planes + DC edge + 12 node VMs (9 D-121 role + juju-01/.5 + +maas-01/.6 + tailscale-01/.7) + the new `vr1-dcN-client` VM (recommended `.8`, ~4 vCPU / +8192 MiB / 80 GiB, non-hypervisor), all siblings in ONE flat apply per DC. Eliminated +outright: the containment VMs + inner libvirtd, the two inner roots AS roots, the +bootstrap gate's `--host-nodes` duty (~134 lines), the qemu+ssh dial + D-126 per-env keys +(no successor), `modules/wan-bridge` + netplan bridge (edge WAN attaches directly to the +per-DC `site-wan` NAT). Preserved unchanged: the mesh triangle + netem link, plane +CIDRs/families/MTU (D-139/D-143 own the values -- the removal changes NO byte budget), +Office1/voffice1 wholesale (D-114 -- a DIFFERENT, KEPT containment pattern), Stages 6-7. + +**What the flattening costs, carried honestly:** D-122's one-command site-down +(`virsh destroy vvr1-dcN`) is lost -- re-earned via the root-scoped teardown primitive +(owed #1) + emergency lever (#6). Three NEW isolation exposures are created and each gets +its own control (Section 5 / the three concerns): (i) cross-DC plane-bridge co-residency +on vcloud's one kernel, (ii) the SEC-010 transit-drop successor on the new endpoints, +(iii) the MAAS power-key blast radius (each DC's region key would open virsh control over +EVERY vcloud domain -- SEC-012/SEC-016's per-DC separation becomes vacuous without the +#11 mitigation). + +**Both changes ride one redeploy, two attributable axes:** `[D-143]` (10.12->10.13 value +substitution, ruled) vs `[CE]` (container-elim shape change, ruled by Section 5's package); +four confirmed dual-cause items carry `[both]` (G17; D-124 transit bearer; R7 revocation; +the juju execution-host/overlay step), per pass1 check 5. + +--- + +## 2. The master change-set (W4.1 -- 104 rows; summarized here, full table in `pass4-w1-master-change-inventory.md`) + +| Category | Rows | Character | +|---|---|---| +| DEC (open rulings) | 23 | Not changes -- the decisions that gate the change rows (Section 7 tiers them) | +| TF (tofu roots/modules) | 13 | 3 retire (containment modules+vars, inner roots, wan-bridge), 5 re-home unchanged bodies, 3 change (site-wan rewire, opnsense input, D-124 transit re-point), 2 NEW (`modules/dc-site`, per-DC flat roots). Zero module bodies rewritten | +| LB (lib-hosts/lib-net) | 3 | 2 lib-hosts edits (power-address re-derivation **BLOCKED on DEC-15**; comment currency); lib-net = ZERO container-elim edits (D-143 axis only, grep-verified) | +| SC (scripts/procedures) | 19 | 8 change, 3 retire/split, 8 NEW (the 13 owed artifacts' script halves). Carve/power/tag script BODIES: no code change (pure MAAS-API, ``-parameterized) | +| DC (workflow-doc prose) | 6 | Stage 3 = THE restructured stage; Stage-5 literals; gap-register updates; the D-114-vs-D-123 two-patterns distinction note | +| RB (runbooks) | 5 | teardown-rollback rewrite; phase2 heaviest rewrite; phase3 low-delta; phase4 RUN-LOCATION 3rd correction; phase6 pre-existing stale-D-138 ride-along fix | +| GT (gates) | 10 | NEW Stage-1 (a)-gate + A11a/A11b + P10; P4/P5/P8/P9 extensions; G9/G10 single successor gate; G17 dual-cause edit; G14 count flag | +| HN (harnesses) | 21 | 9 existing-change + 5 existing-retire + 7 new-build (matches pass3's decomposition exactly); 3 rides (#6, #8, #10) build nothing | +| SEC (ledger rows) | 4 | 2-or-3 new rows ((a) control; #11 -- critical path; concern-(ii) disposition may be a SEC-010 amendment) + register-row re-points | +| **Total** | **104** | (arithmetic shown in Appendix A.7) | + +**The three critical-path chains (W4.1 Section 2, verbatim import -- the sequence in +Section 4 must and does honor all three):** + +- **Chain A (power-key, owed #11 -- the single largest blocker):** DEC-15 mechanism choice + -> SEC row -> SC-11 artifact -> HN-C4 harness (itself a PRECONDITION, not just coverage) + -> lib-hosts re-derivation + `maas-region-power-key` shape -> call-site literals -> + A4/A5 harness edits (TOGETHER, same session -- hazard H1) -> A9/A11b -> P5 row. **Six + test/tool edits are frozen until DEC-15 rules**; the interim RED on A4/A5 is the desired + fail-loud state, never something to "fix" early. +- **Chain B (the (a)-control invariant, owed #2):** DEC-14 -> SEC row -> SC-10 -> HN-C2 -> + Stage-1 gate installed + `--check`-verified -> **MUST PRECEDE the first flat apply of + EITHER per-DC root** (fork-robust under any DEC-11 outcome) -> A11a re-verify at each + apply's close -> re-verified at Stage-5 live traffic. +- **Chain C (R7 + MAAS-release before destroy):** SC-12 R7 revocation checklist + SC-13 + MAAS record-release/rack-decommission run BEFORE any substrate destroy (revoking after + the hosts are gone degrades to "assume it's moot"); then inner-root destroys, then outer. + Governs the CURRENT 10.12 teardown; largely independent of DEC-11. +- **Cross-cutting:** the DEC-11 root-topology fork gates the largest single cluster of + "ready once ratified" rows -- ratify it EARLY (with the decision package, Section 7). + +--- + +## 3. The layered module-workflow design (W4.2 -- full design in `pass4-w2-module-workflow-design.md`) + +The backbone (adopted from pass1 W1.4, populated by pass2/pass3): + +``` +L0 Host & inter-site substrate (IaC) -- Stage 1 [mesh triangle, pools, office1-net, base image] +L1 Site/edge nodes (IaC) -- Stage 2/3 [voffice1 (D-114, untouched), DC edges, THE CLIENT VM] +L2 DC substrate: planes + node VMs (IaC) -- Stage 3 [modules/dc-site composing pool -> planes -> edge -> node-vm x12 -> client VM; ONE flat root/state per DC] +L3 Enlist/commission (procedure) -- Stage 4 [MAAS carve/power/tags -- bodies unchanged, values re-derive] +L4 Juju/OpenStack deploy (procedure) -- Stage 5-7 [D-140 PINS as procedure for this redeploy] +L5 Verify/gate (cross-cutting, re-invoked at every layer boundary) +``` + +**The load-bearing rule:** each layer's input is the layer below's OUTPUT only; the +IaC<->procedure boundary is IDENTITY (a MAC, an IP, a hostname), never orchestration (a +state-file read, a cross-host provider dial). The container layer was the ONE place this +rule was violated (the inner root's qemu+ssh dial into the outer root's own output) -- +Option 1 removes the violation structurally: L2 becomes IaC end to end, and the first live +dial into anything L2 produced is L3's MAAS commissioning, exactly where the boundary +belongs. + +**Placement of the pass's key objects** (W4.2 Section 6): client VM = L1 instance +(same `cloudinit-vm` module type as voffice1/edges), apply-grouped with its DC's flat root; +the (a) control = L0-scoped L5 gate (procedure, NOT a tofu module); SEC-010 successor = +L1-scoped procedure on both transit endpoints (client VM + voffice1, one role-agnostic +installer); power-key mitigation = L3-scoped credential control with an L3/L4-split gate +(P4 dependency + A11b standing); teardown primitive = L2/L5 boundary (procedure wrapping a +root-scoped `tofu destroy`, verified by an L5-style completeness check). + +**Design principles** (each grounded in an existing repo pattern): site-token +parameterization, never hardcoded DC identity; every module ships its tested harness from +the same commit (per-module harness contract, pass3 Section 6); idempotence at every layer; +no layer reaches past the one directly below; findings logged at their true layer; D-140 +is a distinct future axis (a later `-juju` root consumes `dc-site` outputs without +reshaping L0-L3). + +**Roosevelt transfer, judged per layer** (the operator's "module deployment project" +lens): L0 does NOT transfer (mesh/netem are virtualization shims); L1's client-VM PATTERN +is the direct pre-Roosevelt D-138 bastion deliverable; L2 transfers as CONTRACT +("booted object with correct MAC-per-NIC identity"), not as libvirt mechanism; **L3 and L4 +transfer verbatim -- they ARE the module deployment project's payload**; L5 mostly +transfers, EXCEPT the (a) control and the power-key mitigation, both artifacts of vcloud's +single-hypervisor co-residency with no bare-metal analog in the same shape (flagged so no +future session assumes parity). + +--- + +## 4. Execution sequencing (W4.3 -- full step tables in `pass4-w3-execution-sequencing.md`) + +**Part A -- teardown of the current 10.12 Model-B checkpoint** (today's tooling tears down +today's shape; unchanged by the target): A.1 state backups both roots/both hosts -> A.2 +MAAS census (two lenses) -> **A.3 R7 credential revocation (Chain C -- BEFORE any +destroy)** -> A.4 MAAS record release/delete + rack-controller decommission -> A.5/A.6 +plan destroys (inner first, then outer, `-target`ed -- mesh/netem EXCLUDED) -> A.7 apply +destroys (each individually operator-approved, never batched -- hard rule 3) -> A.8 drift +gate -> A.9 NetBox decommission -> A.10 repeat/batch per DC. + +**Part B -- redeploy on flat 10.13 (Option 1):** B.1 prerequisites (**B.1.1 FIT/capacity +HARD GATE** -- extended calculator + fresh vcloud measurement, closes before any apply; +B.1.2-B.1.4 NetBox apex re-carve, lib-net 10.13 literals, D-124 transit re-point) -> B.2 +vcloud/Office1 prep -> **B.3 the (a) control installed + `--check`-verified (Chain B HARD +GATE, before ANY flat apply)** -> B.4 the flat apply per DC (`dc-site` composition; client +VM; **B.4.3 MAC re-measurement before anything trusts a MAC**) -> B.5 surviving duties +re-targeted (rack-retirement recommend + owed live re-measure; D-131 retire-with-evidence; +artifact-service sizing via B.1.1's numbers; SEC-010 successor both ends) -> B.6 MAAS +enlist/commission/carve (**B.6.1 power-key mitigation closes, harness green, BEFORE the +B.6.2 lib-hosts edit -- Chain A**) -> B.7 juju/bundle from the client VM + verify-live +(Ceph-v6, geneve assert, the (a) `--check` RE-RUN under real traffic) -> B.8 close-out +(NetBox registration; **B.8.2 the [ARCH] decision record -- the redeploy is NOT "done" +while it is owed**). + +**Critical-path invariant verification (W4.3 Section 0, re-checked this synthesis -- +Appendix A.2):** all three chains are honored -- (a)-control B.3 before B.4 (Chain B); +#11/B.6.1 before B.6.2 (Chain A); R7/A.3-A.4 before A.5-A.7 (Chain C); every step carries +its axis tag (invariant 4). **One harmonization this synthesis adds** (advisor-reviewed): +W4.3 places the DEC-15 mechanism CHOICE at B.6.1 (its last responsible moment); W4.2's +composition needs the ruled key shape at its Stage 3.5.3. Both honor Chain A; the plan's +recommendation is that **DEC-15 (with DEC-14 and DEC-11/12) be RULED up front with the +decision package** (Section 7), so no operator ruling is discovered mid-Part-B inside a +live teardown window -- B.6.1 then marks where the built artifact + C4 harness must be +green, not where the ruling happens. This matches W4.1's own "ratify DEC-11 early" note. + +**Pre-Roosevelt bare-metal plug-in points:** hardware specs (owed "within days") plug into +B.1.1 (FIT re-derivation), B.4.1/B.4.2 (sizing + provider target -- same module +composition), B.4.3/B.6.4 (real NIC MACs via enlistment -- the re-measurement discipline +transfers, the injection mechanism does not), B.5.3 (artifact-service disk). B.1.2/B.1.3, +B.6.5, and B.7 are hardware-agnostic; **the module INVOCATION ORDER itself (B.1->B.8) is +the reusable deliverable**. B.5's placement should be RE-DECIDED for bare metal, not +carried blindly. The (a) control and #11 mitigation likely do NOT transfer as-is (flagged +explicitly). + +--- + +## 5. The decision package (W4.4 -- framed for the operator, GA-R5; nothing here is ruled) + +### 5.1 The core recommendation: mint a NEW D-number (next-free **D-144**, re-grep at mint) + +The GA-R3 admission test fires on all three prongs independently (each verified across the +passes): (1) architectural consequence beyond the stage -- Stage 3 restructured wholesale, +D-128's own definition amended, a D-124 clause re-caused, D-125 terminated, D-131's +standing-pattern status reopened, D-132-addendum premise mooted, a D-134 map addition, +D-138's concrete host changed; (2) Roosevelt-delta (A1 test) -- operator-stated: the +layered module workflow is the pre-Roosevelt deliverable a future build session would grep +before touching substrate shape; (3) supersession -- D-123's core Model-B ruling is +directly reversed. **The precedent is D-143 itself** (header verified verbatim this +session: "AMENDS D-115 premise; TERMINATES D-101 v4-inherit clause" -- yet minted as its +OWN number): a new entry states its verb against each affected decision rather than the +affected decision being edited to absorb the reversal. A third D-123 amendment reversing +D-123's own central ruling would have the entry amend itself out of existence; append-only +discipline is better served by a fresh entry naming D-123 SUPERSEDED. **Alternative +presented for completeness (not recommended):** amend D-123 in place -- one canonical +entry, but it fights the D-143 precedent and requires rewriting D-123's body to point at +the new shape anyway. + +D-144's body would need, at minimum (W4.4 Section 4): (a) the D-123 supersession; (b) the +D-125 termination + D-124 sizing re-cause as named consequences; (c) **the accepted +power-key blast-radius tradeoff named explicitly in the body** (flattening knowingly voids +SEC-012/016's per-DC scoping, mitigated by the #11 SEC-row control -- burying this only in +a SEC row nobody greps before touching architecture would repeat the failure class A1 +exists to prevent); (d) pointers to the amendments filed alongside. + +### 5.2 The seven ride-alongs (W4.4 Section 2 -- classifications, all recommendation-grade) + +| # | Item | Grade | Recommended form | +|---|---|---|---| +| 1 | D-128 Plane-2 shrink (substrate build becomes wholly Plane 1) | [ARCH] | Own dated D-128 AMENDMENT, filed with the package | +| 2 | D-125 bridge-in retirement | [ARCH] substance, no independent existence | TERMINATION recorded inside D-144's reconciliation ledger (the D-143/D-101 pattern), not its own entry | +| 3 | Client-VM `.8` octet into the D-134 map | [ARCH]-adjacent, narrow | Own dated D-134 AMENDMENT (identical shape to the `.5`/`.6`/`.7` precedents) | +| 4 | Root topology (B) + root naming | [OPS] | Ratified in the delivery change-set (changelog + workflow-doc text), no D-number | +| 5 | The three isolation controls' SEC rows | (i)/(ii) [OPS] SEC rows; (iii) [ARCH]-adjacent FINDING, [OPS] mechanism | (i)/(ii) SEC-ledger rows (next-free SEC-034, verified); (iii) finding folded into D-144's BODY, mechanism as its own SEC row | +| 6 | Rack-controller retirement + D-131 retire-with-evidence | Split: retirement [OPS]; D-131 [ARCH]-touching, independently triggered | Retirement rides owed #5; D-131 gets its own dated AMENDMENT, **gated on the owed live re-measure landing first** | +| 7 | Artifact-service placement/sizing | [OPS] | Resolved WITH the #7 FIT numbers at delivery; changelog only | + +**Net: 0 new D-numbers beyond D-144; 3 amendments filed with the package (D-128, D-134, +D-131) + 1 deferred amendment (D-127's client-VM autostart row, filed once DEC-22 rules +the value -- a reconciled note this synthesis adds: W4.4's "3 amendments" count and its own +D-127 ledger row are consistent only when the D-127 amendment is stated as DEFERRED, see +Appendix A.5); 2 findings folded into D-144's body (D-125 termination, power-key +tradeoff); the rest is SEC/changelog/runbook work.** + +### 5.3 The reconciliation ledger (D-144 -> every touched decision; verb vocabulary = D-143's) + +| Decision | Verb | One-line basis | +|---|---|---| +| D-123 (Model B) | **SUPERSEDED** | Core ruling reversed; history stays intact, append-only | +| D-125 (bridge-in) | **TERMINATES** | OBS-3's precondition disappears with the nesting; no successor | +| D-122 (site shape) | **AMENDS** | Intent preserved; the one-command site-down LITERAL regresses to a root-scoped destroy -- a real, honestly-stated capability loss | +| D-124 (transit) | **AMENDS** | Scheme-A addressing survives on the client VM; the Model-B sizing-void clause is RE-CAUSED (exact figure owed) | +| D-127 (autostart) | **AMENDS** | Containment row loses its object; client-VM row OWED (DEC-22), amendment deferred until ruled | +| D-128 (two-plane) | **AMENDS** | Plane 2's definition shrinks to MAAS/NetBox; the model itself stands | +| D-131 (node DNS) | **AMENDS** | Independently triggered by D-132's per-DC regions; gated on live re-measure; dc1 asymmetry is real | +| D-134 (octet map) | **AMENDS** | Adds `.8` = client VM; bands/CIDRs unchanged | +| D-126 (SSH convention) | **PRESERVES** | Pattern reusable; only the qemu+ssh consumer's key retires, no successor | +| D-132-addendum (region VM) | **PRESERVES** | Region stays on maas-01; hypervisor-fate rationale gets a premise-currency note only | +| D-138 (client in DC) | **PRESERVES** | The principle IS what the client VM realizes; only the concrete host changes | +| D-114, D-133, D-139, D-140, D-143 | **untouched** | Named to prevent scope creep (D-114 especially: a DIFFERENT, KEPT containment pattern) | + +(One divergence between this ledger and the Phase-4 tasking prompt's shorthand is logged +at Appendix A.6 -- the worker document, grounded in the pass evidence, governs.) + +--- + +## 6. OWED artifacts (13) + OWED live measurements + +### 6.1 The 13 owed artifacts (pass2 Section 5 spec; harness dispositions per pass3 -- all LOGGED, none built) + +| # | Artifact | Harness | Blocker | +|---|---|---|---| +| 1 | Teardown primitive (root-scoped gated `tofu destroy`) | C3 (new) | DEC-11 | +| 2 | The (a) cross-DC host isolation control | C2 (new) | DEC-14 | +| 3 | SEC-010 transit-leg successor (one installer, both ends) | A3 (extend) | DEC-16 (shape) | +| 4 | R7 credential-revocation checklist | C6 (new) | ready | +| 5 | MAAS record release/delete + rack-controller decommission | C7 (new) | ready (decommission half: DEC-08) | +| 6 | Emergency site-down lever | rides C3's fixture library | DEC-11 | +| 7 | FIT-calculator extension (+ artifact-service sizing) + fresh capacity measure | A8 (extend) | ready | +| 8 | MAC re-measurement pass post-apply | rides C1's MAC invariant | TF-13 | +| 9 | NetBox DCIM migration | A6 (extend) | DEC-21 | +| 10 | Post-build live asserts (geneve/jumbo; gap-#20 re-verify) | rides `geneve-encap-assert` verbatim, new invocation only | TF-13 | +| 11 | **Power-key blast-radius mitigation -- CRITICAL PATH** | C4 (new; itself a precondition for A4/A5/A9) | DEC-15 | +| 12 | `modules/dc-site` + per-DC flat roots | C1 (new) | DEC-11/12/13 | +| 13 | D-131 retirement-evidence checker | C5 (new; dc1 = standing-RED until live retirement) | DEC-09 | + +Every net-new artifact ships its `tests//run-tests.sh` FROM THE SAME COMMIT +(per-module harness contract), with changelog + revert, repo-lint clean. + +### 6.2 OWED live measurements (read-only where pre-build; NONE performed by this pass) + +**Before the recommendation-grade rulings are ratified:** (1) current-day +`primary_rack`/rackd state, BOTH DCs (gates DEC-08/DEC-09 -- the pass2 cites are +2-10 days old, instrument-currency #20/#25); (2) vcloud's LIVE polkit/libvirt access +config (input to DEC-15's mechanism design). + +**Before/at build:** (3) fresh vcloud host-capacity measurement + (4) exact FIT for the +flat 12+1-VM/DC roster (until then "~176 GiB freed" stays DIRECTIONAL only -- B.1.1 hard +gate); (5) post-build geneve/jumbo live assert on the vcloud-level planes (analytically +unchanged, live proof owed); (6) MAC re-measurement post-apply (B.4.3); (7) the client +VM's transit NIC name (dc0's live was `enp1s0`, not `mgmt` -- before any SEC-010-successor +rule is written); (8) D-131 dig-evidence per fresh region (owed #13); (9) gap-#20 verdict +re-verify post-build (its own expiry clause triggers). + +--- + +## 7. >>> OPEN DECISIONS FOR THE OPERATOR <<< + +Two tiers. Nothing below is ruled by this pass (GA-R5). The 23 DEC rows of the master +inventory partition exactly into these tiers + already-covered ledger rows (Appendix A.3). + +### Tier 1 -- the [ARCH] decision package (rule together, BEFORE Part B's Stage 3 ever runs) + +1. **DEC-01 -- THE ruling: adopt the container-elim as new D-144 (recommended) or as a + D-123 amendment.** **This is the single most important decision in the plan**: it is + the root every ride-along hangs off, W4.2's composition names it a PRECONDITION for the + first flat Stage-3 apply, and per SCOPE Section 7 the redeploy is NOT done while it is + owed. +2. **DEC-02 / DEC-13 / DEC-09** -- the three amendments filed with it: D-128 (Plane-2 + shrink), D-134 (`.8` client VM), D-131 (retire-with-evidence -- **only after owed live + measurement (1)**). D-127's amendment (DEC-22, autostart value) is deferred until ruled. +3. **DEC-15 -- the power-key mitigation mechanism** (restricted SSH key / per-DC virsh + wrapper / polkit ACL) + its SEC row. **The sharpest EXECUTION blocker: six test/tool + edits are frozen until it rules** (Chain A). Recommended: rule it WITH this package, + not mid-sequence (Section 4's harmonization). **AMENDED by the advisor review + (Section 8, follow-up 1): the mechanism must decide REACH as well as authorization** -- + how an in-plane `vr1-dcN-maas-01` reaches vcloud's own libvirtd at L3 without putting a + host address on a plane bridge (which the (a) control forbids). Reach + auth ruled together. +4. **DEC-14 / DEC-16** -- the (a) control's concrete nftables mechanism + SEC row; the + SEC-010-successor row disposition (new row vs SEC-010 amendment) + endpoint + ratification (client VM + voffice1, recommended). +5. **DEC-11 / DEC-12 / DEC-23** -- ratify root topology (B) shared-outer + per-DC-flat + (recommended; state-isolation only -- it mitigates NONE of the three concerns), root + naming (`vr1-dcN-flat` vs reserving `-substrate`), and the state-blast-radius weighing + that rides it. Unblocks the largest cluster of "ready once ratified" rows. +6. **DEC-08 -- rack-controller retirement ratification** (recommended; gated on owed live + measurement (1)). +7. **DEC-24 -- the dc0<->dc1 mesh-leg consumer + cross-DC Ceph replication path** (NEW, from + the advisor review, Section 8 follow-up 2). Who holds the dc0<->dc1 mesh leg (`virbr5`, + netem target) endpoints post-flatten, and how the flat-topology replication plane routes + cross-DC Ceph traffic (D-100/D-108) THROUGH netem without violating the (a) control. Pairs + with DEC-14. Partly pre-existing (cross-DC replication never built -- dc1 HELD), so the flat + design must DEFINE this path; it is not a regression into a working path. OWED: a design + + its gate; not a blocker for the D-144 package. + +### Tier 2 -- delivery-grade rulings (needed before their specific rows, not before the package) + +- **DEC-10** artifact-service (`.4`) placement + sizing -- ruled WITH the #7 FIT numbers. +- **DEC-22** D-127 client-VM autostart value (unblocks harness A1's new case + the + deferred D-127 amendment). +- **DEC-21** NetBox-migration design (rename-in-place vs concept retirement, shapes A6). +- **DEC-20** A11's home (fold into `cloud-assert.sh` vs a dedicated `isolation-assert.sh`). +- **DEC-17** `wan-bridge` module directory: delete vs leave-unreferenced (append-only bias). +- **DEC-18** SEC-013 `maas-vm-host` retire-or-keep (flagged to its owner, not this pass). +- **DEC-19** `maas-fabric-prune.sh`/`maas_fabric_classify.py` harness gap: build vs + accept-as-named-exception (pre-existing, container-elim-ADJACENT only). +- DEC-03..DEC-07 are covered by the package's reconciliation ledger (they are the named + consequence-notes inside/alongside D-144, not separate operator forks). + +### Also owed from the operator (SCOPE Section 8, unchanged) + +Pre-Roosevelt hardware specs ("within a few days" -- plug-in points mapped in Section 4); +any external "module deployment project" artifacts (none identified in-repo at any phase). + +--- + +## 8. Advisor-review follow-ups (added post-synthesis; full text in `FINAL-advisor-review.md`) + +The advisor reviewed the aggregate and returned **verdict SOUND** -- Option 1, root (B), and the +D-144 framing all stand. It raised TWO follow-up gaps of the same shape (an attachment object is +deleted; a survivor glossed "unchanged"), both VERIFIED against the full pass docs this session: + +- **Follow-up 1 (folded into DEC-15):** the MAAS power-dial REACH path (in-plane `maas-01` -> + vcloud's own libvirtd) is assumed, not designed, and collides with the (a) control's + no-host-address-on-a-plane-bridge rule. The DEC-15 mechanism ruling must decide reach + auth + together. Sharpens DEC-15; does not block the package. +- **Follow-up 2 (new DEC-24):** the dc0<->dc1 mesh leg (`virbr5`, netem target) is the cross-DC + Ceph replication carrier (D-100/D-108); the pass reassigned the Office1-transit legs cleanly to + the client VM but left the dc0<->dc1 leg's post-flatten consumer + netem routing undefined. + Partly pre-existing (dc1 HELD -> never built), so the flat design must DEFINE it. Own open item. + +Neither changes the plan's direction; each is now a named open item (Section 7 DEC-15 / DEC-24). + +--- + +## Appendix A -- cross-consistency check results (this synthesis's adversarial pass) + +**Verdict: CONSISTENT -- zero contradictions across the four Phase-4 dimensions; six +reconciled notes, enumerated below. No manufactured contradiction survives into this plan.** + +**A.1 W4.1 rows <-> W4.3 sequence steps.** Every Part-A/Part-B step that invokes a +change-artifact resolves to a W4.1 inventory row (spot-mapped: A.3=SC-12, A.4=SC-13, +A.9/B.8.1=SC-18, B.1.1=SC-16, B.3.1=SC-10/GT-01/SEC-01/HN-C2, B.4.1=TF-12/TF-13, +B.4.3=SC-17, B.5.1=DEC-08/SC-05, B.5.2=SC-15/HN-C5, B.5.4=SC-04/GT-03, B.6.1=DEC-15/SC-11/ +HN-C4/SEC-02, B.6.2=LB-01, B.6.3=SC-01, B.7.3=SEC-04/GT-06, B.7.5=SC-19/GT-02, +B.8.2=DEC-01). *Reconciled note 1:* W4.3's B.1.2/B.1.3a (NetBox apex re-carve; 10.13 +naming-collision DOCFIX) are NOT W4.1 rows -- correctly so: they are D-143's OWN +owed-execution items (D-143 axis), and W4.1's LB-03 explicitly quarantines that axis. +By-design separation, not a gap. + +**A.2 The three critical-path chains vs the sequence.** All honored: Chain B at B.3-before- +B.4 (fork-robust "before ANY flat apply", not merely "before the second DC's"); +Chain A at B.6.1-before-B.6.2 (with A4/A5 edited together, same session); Chain C at +A.3/A.4-before-A.5-A.7. *Reconciled note 2 (elevated into Section 4):* W4.2's Stage 3.5.3 +needs the #11-ruled key shape earlier in its composition than W4.3's B.6.1 choice-point -- +no invariant is violated (both keep the ruling before B.6.2/any literal), but the plan +recommends DEC-15 be ruled with the up-front package so the choice never lands mid-window. + +**A.3 W4.4's decision list <-> W4.1's 23 DEC rows.** Not identical, and correctly so: +W4.4 frames the [ARCH]-relevant subset -- DEC-01 (core), DEC-02..09/13 (ride-alongs + +ledger rows), DEC-10..12, DEC-14..16, DEC-22 (via the D-127 ledger row), DEC-23 (rides +DEC-11) = 18 of 23. The remaining five (DEC-17, 18, 19, 20, 21) are OPS/delivery-grade +decisions with clean provenance in pass2 Section 6 / pass3 Section 7 open lists -- no +invented rows, no dropped [ARCH] item. *Reconciled note 3:* the FINAL-PLAN unions them as +Tier 2 (Section 7) so the operator sees all 23. + +**A.4 W4.2's modules <-> W4.1's 13 TF rows.** Full coverage both directions: every W4.2 +layer-table IaC artifact maps to a TF row or a confirmed-unchanged note (mesh-link x3, +netem-link, office1-network -- deliberately not itemized; maas-vm-host dead/orthogonal -> +DEC-18); every TF row appears in W4.2's design. *Reconciled note 4 (cosmetic):* W4.1's TF +IDs skip TF-11 (TF-01..10, 12, 13, 14) -- the count of 13 rows is CORRECT; the gap is a +numbering artifact only. Flagged so no future reader "finds" a missing row. + +**A.5 The D-144 package's internal soundness.** GA-R3 three-prong argument checked against +the pass evidence -- each prong independently grounded (Section 5.1); the D-143 precedent +verified VERBATIM against `docs/design-decisions.md:8083` this session; next-free D-144 +re-verified this session by direct grep (highest = D-143), next-free SEC-034 re-verified +(highest = SEC-033) -- both re-grepped again at mint time per numbering discipline. +*Reconciled note 5:* W4.4's summary line "3 existing-decision amendments (D-128, D-134, +D-131) filed separately" is consistent with its own D-127 AMENDS ledger row only when the +D-127 amendment is stated as DEFERRED on DEC-22's value ruling -- Section 5.2 states it +that way ("3 filed + 1 deferred"). + +**A.6 The reconciliation ledger vs decision statuses.** No contradiction found between +W4.4's ledger and any decision's verified status (D-128/D-138/D-131/D-134/D-123 texts were +direct-read by prior admins at cited lines; D-122/D-123/D-125/D-131/D-140/D-143 headers +re-confirmed present this session). *Reconciled note 6:* the Phase-4 tasking prompt's +shorthand ("D-122 ... preserved"; D-127 omitted) DIVERGES from W4.4's evidence-grounded +ledger (D-122 = AMENDS -- the site-down literal regresses; D-127 = AMENDS-deferred). The +worker document governs; this plan carries W4.4's verbs. Named explicitly so a later +reader cannot manufacture a contradiction from the prompt text. + +**A.7 Count arithmetic (lesson #25: show the addition, don't assert it).** +23 (DEC) + 13 (TF) + 3 (LB) + 19 (SC) + 6 (DC) + 5 (RB) + 10 (GT) + 21 (HN) + 4 (SEC) += **104**. Harness decomposition 9 + 5 + 7 = 21 matches pass3 exactly; 13 distinct +owed-artifact tags all appear; the three no-double-count folds (rack decommission -> #5; +artifact-service sizing -> #7; SEC-010-writer extraction = #3's shape) carried intact. + +**A.8 Two-axis separation, end to end.** Every W4.1 row and every W4.3 step carries +`[CE]`/`[D-143]`/`[both]`; lib-net is quarantined to D-143 (grep-verified at pass2); +the four dual-cause items are dual-labeled, never folded. HOLDS. + +**A.9 Read-only / logged-not-built.** All eight input documents state READ-ONLY with no +mutation; this synthesis performed only reads + greps of repo files and writes only this +planning document. All 13 owed artifacts remain LOGGED, none built; all mutations in +Section 4 are PLANNED steps for later gated execution (hard rules 1/3 respected). + +**A.10 Remaining inferred/uncited-claim sweep.** None found beyond items already marked +OWED by their sources. Recommendation-grade items still needing a live re-measure before +ratification: rack retirement (DEC-08/09) and FIT/capacity (Section 6.2 items 1-4). +Standing doc-currency nits carried, not fixed here: CLAUDE.md "98 harnesses" vs 103; the +workflow doc's stale "~416 GiB"; `maas-node-power` cosmetic URI fixture. + +--- + +## Verification note + +Author = "the administrator" (no model name asserted, operator instruction). Direct +measurements THIS session: `grep -nE '^## D-1(22|23|25|31|40|43)'` + +`grep -oE 'D-1[0-9]{2}' | sort -u | tail` on `docs/design-decisions.md` (D-143 header +verbatim incl. the AMENDS/TERMINATES precedent; highest D = D-143 -> next-free D-144) and +`grep -oE 'SEC-0[0-9]{2}'` on `docs/security-ledger.md` (highest = SEC-033 -> next-free +SEC-034). Everything else rests on the four Phase-4 worker docs and the four prior +administrator reports, read IN FULL this session, whose own verification notes record the +underlying direct reads/greps -- their citations were relied upon as verified there, not +re-derived wholesale here; no claim was upgraded beyond its source's stated grade. +READ-ONLY; findings LOGGED only; nothing executed against the cloud. diff --git a/docs/audit/container-elim-pass/FINAL-advisor-review.md b/docs/audit/container-elim-pass/FINAL-advisor-review.md new file mode 100644 index 0000000..4cd364e --- /dev/null +++ b/docs/audit/container-elim-pass/FINAL-advisor-review.md @@ -0,0 +1,64 @@ +# FINAL advisor review -- container-layer-elimination pass (aggregate) + +**Reviewer:** the advisor (`advisor()` tool; recorded as "the advisor" -- no model name asserted, +per operator instruction). **Date:** 2026-08-09. **Input:** the full pass transcript (all worker +bounded summaries, the five administrator reports, `FINAL-PLAN.md`). + +## Verdict: SOUND + +Nothing in the aggregate overturns the pass's three load-bearing conclusions: +- **Option 1** (flat node VMs on vcloud libvirt + a small per-DC `vr1-dcN-client` VM) as the target; +- **Root topology (B)** (shared-outer + per-DC-flat roots) as the recommended tofu shape; +- **New D-144** (superseding D-123 Model B) as the decision form, with the D-143 precedent. + +The plan is honest about its open items and its recommendation-grade items pending live re-measure. +The advisor raised TWO follow-up gaps -- both the SAME failure shape (an attachment object is +deleted; a survivor is glossed "unchanged" when its attachment point no longer exists). Both were +VERIFIED against the full pass docs by the orchestrator this session. Neither changes the plan's +direction; each becomes a NAMED open item. + +## Follow-up 1 -- MAAS power-dial REACH (distinct from authorization scope). CONFIRMED, fold into DEC-15. + +The pass covered the power-key AUTHORIZATION concern thoroughly (both DCs' region VMs dial one vcloud +`qemu:///system`; the #11 mitigation / DEC-15 scopes it). It did NOT resolve REACHABILITY: how an +in-plane `vr1-dcN-maas-01` gets an L3 path to vcloud's own libvirtd at all. Model B dialed the +containment VM at an in-plane metal-admin address (`10.12.x.2`); flat, the target is vcloud's own +libvirtd, and vcloud by design holds no address on any DC plane. The naive fix -- give vcloud an +address on the metal-admin bridge -- is EXACTLY "a host address on a plane bridge", which W3.3's (a)-control +fixture treats as a failure-to-catch. Evidence: `pass2-w3-scripts.md:184-201` and `pass4-w3` B.6.1/B.6.2 +assume "one endpoint reachable from BOTH DCs' region VMs" without spelling out the reach path. +**Disposition:** DEC-15's framing is AMENDED -- the power-key mechanism ruling must decide REACH and +AUTHORIZATION together (the reach path and the (a) control's no-host-on-plane-bridge rule must be +reconciled, not left to a delivery-time improvisation). Does not block the D-144 package; sharpens DEC-15. + +## Follow-up 2 -- the dc0<->dc1 mesh leg's post-flatten consumer + cross-DC Ceph path. CONFIRMED, NEW open item DEC-24. + +Verified precisely: the containment VMs attach ONLY to the two Office1-TRANSIT mesh legs +(`opentofu/main.tf:192-193` region-end, `:430`/`:554-555` DC-end) -- which the pass reassigns cleanly to +the `vr1-dcN-client` VM's transit NIC (`pass2-w1-tofu-modules.md` row `mesh-link`; `pass4-w3` B.1.4). But +the **dc0<->dc1 leg** is a DIFFERENT segment: `bridge virbr5` (`main.tf:356`), the netem target +(`main.tf:334-356`), designed as the cross-DC **Ceph replication** carrier (D-100 mesh-not-star rationale +`design-decisions.md:2235`; D-108 rbd-mirror/radosgw-multisite rides the IPv6 replication plane +`:3051`; `dc0<->dc1 stays replication-only` `:4993`). `pass2-w1` graded it "untouched either way" -- true +for the BRIDGE's existence, but it does NOT define what attaches the (now vcloud-level) replication plane +to virbr5 and routes cross-DC traffic THROUGH netem once the containment layer is gone. +**Disposition:** NEW open design item **DEC-24** (Tier 1-adjacent, pairs with DEC-14 the (a)-control): +"who holds the dc0<->dc1 mesh-leg endpoints and how does the flat-topology replication plane route +cross-DC Ceph traffic through netem (WAN fidelity) without violating the (a) control." Nuance carried +honestly: this is PARTLY pre-existing -- cross-DC replication was never built (dc1 HELD, rbd-mirror blocked), +so the flat design must DEFINE this path, it is not a regression the flattening introduces into a working path. +Roosevelt-delta: on bare metal the dark-fiber dc0<->dc1 leg is real (`design-decisions.md:2235`), so this +routing must be defined regardless -- the flat rehearsal is the right place to design it. + +## Report mechanics (advisor-directed, executed by the orchestrator) + +1. This file (`FINAL-advisor-review.md`) captures the review. [done] +2. `operator-report.md` leads with the Tier-1 decisions (DEC-01 first; DEC-15 as the sharpest execution + blocker, now carrying the reach dimension; DEC-24 added), carries W4.4's evidence-grounded ledger verbs + (D-122 AMENDS, D-127 AMENDS-deferred -- NOT the tasking shorthand), and marks the recommendation-grade + items (DEC-08/09 rack retirement, FIT/capacity) as pending live re-measure. It situates Part A: the + teardown presumes the dc0 checkpoint wrap-gates (cloud-assert BOM, controller backup, verify-live Ceph) + close first -- still owed per the ledger. +3. Durability: commit + push the pass folder; repo-lint first (24 worker-authored docs risk non-ASCII); + update CURRENT-STATE's "CONTAINER-ELIM PASS SCOPED ... NOT YET RUN" line to point at FINAL-PLAN.md + + operator-report.md IN THE SAME COMMIT (GA-R1 C1). The full session bookend (savegame) stays the operator's call. diff --git a/docs/audit/container-elim-pass/operator-report.md b/docs/audit/container-elim-pass/operator-report.md new file mode 100644 index 0000000..1a76ac6 --- /dev/null +++ b/docs/audit/container-elim-pass/operator-report.md @@ -0,0 +1,127 @@ +# Container-layer elimination -- OPERATOR REPORT + +**Date:** 2026-08-09. **Pass:** multi-agent, READ-ONLY planning (phases 0-4 + advisor review). +**Nothing was executed against the cloud; every owed artifact is LOGGED, not built.** Full plan: +`FINAL-PLAN.md`. Decision framing: `pass4-w4-decision-framing.md`. Advisor review: +`FINAL-advisor-review.md`. All 24 pass documents are in this folder. + +--- + +## 1. What you confirmed, and what's still yours to rule + +At the Phase-0 gate you confirmed (directional -- NOT the formal ruling): **Option 1** (flat node +VMs on vcloud libvirt + one small per-DC `vr1-dcN-client` VM for the D-138 client role), and +handling **(a)** for the cross-DC gap (a new vcloud host-level isolation control). The pass then +planned the whole change against that target. **The formal [ARCH] ruling is still owed -- it is the +single most important decision below.** + +## 2. The one decision everything hangs off: DEC-01 + +**Mint a new `D-144` superseding D-123 (Model B), OR record it as a D-123 amendment.** +Recommendation: **new D-144.** The GA-R3 test fires on all three prongs (architectural consequence +beyond the stage, the operator-stated Roosevelt-delta, direct supersession of D-123's core ruling), +and the on-point precedent is **D-143 itself** -- it AMENDED D-115 and TERMINATED a D-101 clause yet +was minted as its own number rather than editing either. A third D-123 amendment reversing D-123's +own central ruling would amend the entry out of existence; append-only discipline prefers a fresh +entry naming D-123 SUPERSEDED. Per SCOPE, **the redeploy is not "done" while DEC-01 is owed** -- and +W4.2 names the ruling a PRECONDITION for the first flat Stage-3 apply. + +D-144's body should carry, explicitly: the D-123 supersession; the D-125 termination + the D-124 +sizing re-cause; and -- important -- **the accepted power-key blast-radius tradeoff named in the body, +not buried in a SEC row** (below). + +## 3. The Tier-1 package to rule WITH DEC-01 (before Part B's Stage 3 runs) + +- **DEC-02 / DEC-13 / DEC-09** -- three amendments filed alongside: **D-128** (Plane-2 shrinks -- the + substrate build becomes wholly Plane 1), **D-134** (adds `.8` = the client-VM octet), **D-131** + (retire-with-evidence -- *gated on a fresh live re-measure first*). D-127's autostart amendment is + DEFERRED until DEC-22 rules the value. +- **DEC-15 -- the power-key mitigation mechanism** (restricted SSH key / per-DC virsh wrapper / polkit + ACL) + its SEC row. **The sharpest execution blocker:** six test/tool edits are frozen until it + rules, and the interim RED on `dc-selector`/`maas-region-power-key` once `lib-hosts` changes is the + *wanted* fail-loud state -- do not green it early. **The advisor sharpened this: the mechanism must + also decide REACH** (how an in-plane `maas-01` reaches vcloud's own libvirtd without a host address + on a plane bridge -- which the (a) control forbids). Reach + authorization ruled together. +- **DEC-14 / DEC-16** -- the (a) control's concrete nftables mechanism + SEC row; the SEC-010-successor + disposition (new row vs SEC-010 amendment) + endpoints (client VM + voffice1, recommended). +- **DEC-11 / DEC-12 / DEC-23** -- ratify **root topology (B)** (shared-outer + per-DC-flat roots; + recommended -- but note it's tofu **state** isolation only, it mitigates NONE of the three + concerns), root naming, and the state-blast-radius weighing. Unblocks the largest cluster of rows. +- **DEC-08** -- rack-controller retirement (recommended; gated on the same fresh live re-measure). +- **DEC-24 (NEW, from the advisor review)** -- the **dc0<->dc1 mesh-leg consumer + cross-DC Ceph + replication path**. `virbr5` (the netem'd leg) is the D-100/D-108 replication carrier; the pass + reassigned the Office1-transit legs to the client VM cleanly but left this leg's post-flatten + attachment undefined. Partly pre-existing (dc1 HELD -> never built), so the flat design must DEFINE + it. Pairs with DEC-14; owed a design + gate, not a package blocker. + +Tier-2 delivery-grade rulings (needed before their specific rows, not the package): artifact-service +sizing (DEC-10, with the FIT numbers), D-127 autostart value (DEC-22), NetBox migration (DEC-21), +the A11 gate home (DEC-20), `wan-bridge` dir delete-vs-leave (DEC-17), and two pre-existing items +adjacent to the pass (DEC-18 `maas-vm-host`, DEC-19 `maas-fabric-prune` harness gap). + +## 4. What the flattening buys -- and what it costs, stated honestly + +**Buys:** two tofu roots + a bash bootstrap gate + a cross-host `qemu+ssh` provider dial collapse to +**one apply per DC, one state axis, no gate**; nesting depth **4 -> 2** (the VR0-proven shape). The +deepest structural win (found independently by two workers): the container layer is the **one +violation** of the repo's IaC->procedure boundary rule -- the inner root's `qemu+ssh` provider is +IaC reaching across a live-dial boundary -- and Option 1 removes it structurally. Module impact is +small: **6 IaC modules re-home with zero body changes**, `wan-bridge` collapses, one new `dc-site` +module composes the per-DC stack. + +**Costs (each carried, none papered over):** +- D-122's **one-command site-down** (`virsh destroy vvr1-dcN`) is lost -- re-earned via a root-scoped + teardown primitive + an emergency virsh-loop lever (both owed, don't exist yet). +- **Three NEW isolation exposures**, each with its own owed control: + 1. **cross-DC plane-bridge co-residency** -- both DCs' bridges on vcloud's one kernel -> the (a) + host-level control (Stage-1 gate + cloud-assert A11a; must REFUSE on partial resolution); + 2. the **SEC-010 transit-drop successor** on the new endpoints (client VM + voffice1; preflight P10); + 3. the **MAAS power-key blast radius** -- and this is the pass's most consequential finding: + re-deriving node power to vcloud's own libvirtd would give **each DC's region key virsh control + over BOTH DCs' fleets + voffice1 + vcloud** (verified: voffice1 is a vcloud-libvirtd domain, + `main.tf:175`; no libvirt access-scoping exists in-repo). **Flattening silently makes SEC-012/016's + per-DC key separation vacuous** unless the DEC-15 mitigation ships. Its gate is a *negative* test + (the key CANNOT reach the other DC), not existence-only. + +## 5. The layered module workflow (the pass's headline deliverable) + +The redeploy becomes an L0-L5 layered module system -- IaC modules (OpenTofu) + procedure modules +(runbooks/scripts) -- built from the EXISTING 12 modules + 8 `$DC`-parameterized stage runbooks, not +invented. Composition = a `(once)` prefix (mesh/pools + install-and-verify the (a) control) then a +per-`$SITE` loop (`dc-site` apply -> installs -> MAAS enlist -> juju/bundle -> verify). **Roosevelt +transfer** (what the operator wants tested): the **L1 client-VM + L3 MAAS + L4 juju/bundle procedures +transfer to bare metal unchanged** -- that is the "module deployment project" payload; L0/L2's +virtualization IaC and the two co-residency-specific controls (the (a) control, the power-key +mitigation) are the parts that do NOT carry over (bare metal uses separate hosts + IPMI/Redfish). The +pre-Roosevelt bare-metal hardware specs plug in at named points (FIT/roster/MAC/sizing) -- mapped, so +when the specs land the plug-in points are known. + +## 6. Scale, owed work, and the two-axis discipline + +- **Master change-set: 104 rows** (23 decisions, 13 tofu-module, 3 lib, 19 script, 6 doc, 5 runbook, + 10 gate, 21 harness, 4 SEC). Full table: `pass4-w1-master-change-inventory.md`. +- **13 owed artifacts** (all logged, none built), each shipping its own failable harness. **7 new-build + harnesses / 9 change / 5 retire** (10 of 103 existing harnesses touched; ~90 grep-confirmed clean). +- **9 owed live measurements** -- 2 needed BEFORE ratifying the recommendation-grade items + (current-day `primary_rack`/rackd state both DCs; vcloud's live polkit/libvirt config), 7 at build + time (fresh capacity + exact FIT -- until then "~176 GiB freed" is DIRECTIONAL only; geneve/jumbo + live assert; MAC re-measure; client-VM transit NIC name; D-131 dig-evidence; gap-#20 re-verify). +- **Two-axis separation holds end to end:** every change is attributable to `[D-143]` (the 10.13 value + substitution, already ruled) or `[CE]` (this shape change); four dual-cause items carry `[both]`. + `lib-net.sh` carries ZERO container-elim edits (D-143 axis only, grep-verified). + +## 7. Sequencing note (do not miss) + +The teardown (Part A) presumes the **dc0 checkpoint wrap-gates close first** -- cloud-assert BOM, +controller backup, verify-live Ceph -- which are **still owed per the session ledger**. The redeploy +(Part B) order honors three hard invariants: the (a) control before ANY flat apply; the power-key +mitigation before the `lib-hosts` power-address re-derivation; R7 revocation + MAAS record-release +before any substrate destroy. Full ordered step tables: `pass4-w3-execution-sequencing.md`. + +## 8. Recommended next step + +Rule the **Tier-1 package** (DEC-01 D-144 + its amendments + DEC-15/14/16/11/08/24) in GA-R5 +exchanges -- ONE decision per exchange, each with its exact utterance. That ratification is what turns +this plan into buildable, gated work and unblocks the redeploy design. Everything below the package is +delivery-grade and can follow. The pre-Roosevelt hardware specs, when they land, slot into the mapped +plug-in points without reopening the topology. diff --git a/docs/audit/container-elim-pass/pass0-admin-report.md b/docs/audit/container-elim-pass/pass0-admin-report.md new file mode 100644 index 0000000..f58de6d --- /dev/null +++ b/docs/audit/container-elim-pass/pass0-admin-report.md @@ -0,0 +1,312 @@ +# Pass 0 -- ADMINISTRATOR REPORT: baseline & target topology (container-layer elimination) + +**Author:** the Phase-0 administrator (multi-agent pass, `SCOPE-AND-EXECUTION-PLAN.md` Section 4). +**Date:** 2026-08-09. **Inputs:** `pass0-w1-substrate-map.md`, `pass0-w2-network-map.md`, +`pass0-w3-consumers.md`, `pass0-w4-targets.md` -- read in full, adversarially cross-checked +against repo ground truth (checks logged in Section 3). READ-ONLY synthesis; no mutation. + +**This report ends in the Phase-0 OPERATOR GATE (Section 7). Do not start Phase 1 before the +operator rules on it.** + +--- + +## 1. As-is map (W0.1 + W0.2 reconciled into one picture) + +### 1.1 The substrate (Model B, D-122/D-123) + +``` +vcloud (outer libvirt, qemu:///system) <- OUTER ROOT: opentofu/main.tf +|-- Office1 arm: office1_network / office1-wan / office1_opnsense / voffice1 (NOT container-layer) +|-- ubuntu_noble_base (shared base image) (persists) +|-- D-100 mesh TRIANGLE (mesh-link x3, MTU 9000) + netem on dc0<->dc1 (persists) +| office1<->dc0 leg = D-124 transit: voffice1 NIC2 172.31.0.1/30 <-> vvr1-dc0 NIC1 172.31.0.2/30 +| office1<->dc1 leg = transit: voffice1 NIC3 172.31.0.5/30 <-> vvr1-dc1 NIC1 172.31.0.6/30 +|-- vr1_dc0_uplink / vr1_dc1_uplink (site-wan NAT 172.30.2.0/24 / 172.30.3.0/24, MTU 1500, D-125) +|-- vr1_dc0_storage / vr1_dc1_storage (outer pool -- holds ONLY the containment VM's own disk) +| +`-- vvr1-dc0 / vvr1-dc1 ==== CONTAINMENT VM (D-123 Model B) ==== + 108 vCPU / 480 GiB / ~3000 GiB (opentofu/variables.tf:137-156; RAM raised 416->480 2026-08-01) + NIC1 = transit (SEC-010 FORWARD-drop keys here); NIC2 = IP-less, enslaved into the + br-vr1-dcN-wan netplan bridge (D-125 bridge-in); expose_nested_virt = true + | + |== BOOTSTRAP GATE (scripts/site-headend-install.sh --role rack --host-nodes): + | nested libvirtd + inner pool dir + kvm nested=1 + OPNsense image staging + + | SEC-010 nftables writer + D-125 bridge verify. Bash, NOT tofu -- sits BETWEEN the roots. + | + `== INNER ROOT: opentofu/vr1-dcN-substrate/ (qemu+ssh into the containment VM, run + from voffice1 per D-128; D-126 per-env key auth) + |-- inner_storage pool (inside the containment VM's own disk image) + |-- the SIX planes (dc-planes, isolated L2, MTU 9000) -- RELOCATED here by D-123 + |-- vr1_dcN_wan (wan-bridge onto br-vr1-dcN-wan) + vr1_dcN_opnsense (DC edge) + `-- vr1_dcN_node x 12: 9 D-121 role nodes (3 control / 2 compute / 4 storage) + + vr1-dcN-juju-01 (.5, D-104) + vr1-dcN-maas-01 (.6, D-132 addendum) + + vr1-dcN-tailscale-01 (.7, D-129(iii)); all MAC-pinned +``` + +### 1.2 The two-stage apply ordering (structurally forced, not preferred) + +OUTER apply (vcloud, Plane 1) -> BOOTSTRAP GATE (bash, over SSH) -> INNER apply (voffice1, +Plane 2, qemu+ssh). Forced because "a libvirt provider cannot be configured from a resource +created in the same apply" (`opentofu/main.tf:28-29`). Eliminating the containment VM removes +the forcing function entirely: one root, one state axis, no gate script, no cross-host provider +dial. This is the single largest structural simplification the elimination buys (W0.1 Sec 3). + +### 1.3 What the "container layer" is, precisely + +The `vvr1-dcN` VMs + their inner libvirtd + the inner roots + the bootstrap gate's node-host +mode + the D-125 bridge-in plumbing (wan-bridge module, IP-less uplink NIC, `br-vr1-dcN-wan`) ++ the qemu+ssh provider dial and its D-126 keys. It is NOT: the six planes, the node VMs, the +DC edge, the mesh triangle, the uplink NATs, or the LXD API-charm containers on deployed nodes +(all of which persist; the first three re-home, W0.1 Sec 5 / W0.2 Sec 4). + +### 1.4 Network facts the elimination must preserve (W0.2) + +- MTU/geneve budget: 9000 jumbo underlay -> tenant MTU 1500 with ~56-byte v6-geneve overhead. + The containment hop was a same-MTU bridge with no extra encapsulation -- **removing it changes + no byte budget** (W0.2 Sec 3). Do not let later phases imply an MTU benefit. +- The 2026-08-08/09 geneve-over-v6 defects (encap-family split, bracketed encap-ip) were + OVN/OVS-layer, NOT containment-layer -- container-elim neither caused nor fixes them. +- SEC-010 (CLOSED) is an interface-scoped FORWARD-drop on the transit leg of vvr1-dcN and + voffice1. Its ledger rationale is nesting-specific: vvr1-dcN "bridges ALL 6 inner planes + + the transit" -- SEC-010 is doing the isolation a separate host would do for free. + +--- + +## 2. Consolidated blast-radius inventory (W0.3 + W0.1/W0.2, deduplicated) + +Census baseline: 424 `vvr1-dc` occurrences across 78 canonical files; ~26 in scripts, 17 tofu +files; the bulk is decision-record prose (append-only -- superseding notes only, never edits). + +| # | Consumer (file:line) | Keys to the containment layer via | On flattening | Severity | +|---|---|---|---|---| +| 1 | `opentofu/main.tf:410-519,537-623` (`module vvr1_dc0/_dc1`) + sizing vars `variables.tf:137-156,175-194` + D-124 rack-addressing vars `:196-244` + D-126 pubkey vars | IS the containment VM + its sizing/addressing/keys | DELETED. Sizing/rack-addressing/pubkey vars go with it; `dc-dc-whole-host-budget.py` containment-overhead flags re-derived or dropped | HIGH | +| 2 | `opentofu/vr1-dc0-substrate/`, `vr1-dc1-substrate/` (whole inner roots + states) | qemu+ssh provider into vvr1-dcN | RETIRED as roots; their module CALLS (planes, edge, 12 node VMs, storage) re-home into the outer/flat root -- bodies unchanged, provider changes | HIGH | +| 3 | `modules/wan-bridge` + `vr1_dcN_wan` calls + the IP-less uplink NIC + `br-vr1-dcN-wan` netplan | D-125 bridge-in (exists only to fix OBS-3 nesting egress) | DELETED; DC edge WAN attaches directly to the outer `vr1_dcN_uplink` NAT (Model A item 8 -- same /24, no re-address) | MEDIUM | +| 4 | `scripts/lib-hosts.sh:52,157-168,212-214,246-251` (`VIRSH_POWER_ADDRESS*`) | qemu+ssh power URIs dial the containment VM's libvirtd | Re-derived to vcloud's own virsh; the FROM_OFFICE1/FROM_DCREGION split may collapse. Wrong values = "MAAS unreachable" masquerading as a network fault | HIGH | +| 5 | `scripts/maas-node-power.sh` | power address as ARGUMENT (topology-agnostic) | No code change; every invocation site/runbook example needs the new address | MEDIUM | +| 6 | `scripts/site-headend-install.sh` `--host-nodes` (~140 lines) + SEC-010 writer | THE containment-VM bootstrap + SEC-010 enforcement point | node-host mode = dead code for VR1; SEC-010's implementation must be REBUILT for the new boundary (Section 5), not moved. Rack-role remainder needs a re-homing decision (Section 6) | HIGH | +| 7 | `scripts/dc-rack-net.sh` (rack bridge-leg IPs `.2`/`.3` + D-131 node DNS forwarder) | runs ON vvr1-dcN; addresses are the containment VM's identity on its own inner bridges | Needs a NEW HOME (the Option-1 client VM is the natural candidate) or a ruled retirement -- D-131 is a DNS-availability control; silently losing it reintroduces the SERVFAIL bug | HIGH | +| 8 | SEC-028 juju service key + SEC-029 Octavia PKI overlay, resident on "the rack"; `vm-secret-locations` `rack` rows; SEC-026 isolation control | vvr1-dcN is a credential-bearing host (first `rack` rows in the register) | Residencies MIGRATE to the replacement client host; matrix/register rows re-point; SEC-026 "only that DC's credential" must be re-asserted; rotation triggers ("if the rack is rebuilt") FIRE on this change | HIGH | +| 9 | `runbooks/dc-dc-teardown-rollback.md` (Paths A/B/C/M) | `vvr1-dcN` as the teardown unit; D-122's "site-down = one `virsh destroy`" | Re-authored around a scripted group-destroy of the `vr1-dcN-*` domain set. The one-command site-down is a REAL D-123 regression to carry honestly; module design should re-earn it (tofu-module-scoped destroy) | MEDIUM-HIGH | +| 10 | D-128/D-138 execution-host surface (`runbooks/dc-dc-phase3:424,430`, `phase4:165-167`; `ssh -J voffice1` to the rack transit IP) | vvr1-dcN is the concrete D-138 client host | THE load-bearing decision -- answered by proposal in Section 4 (Option 1's client VM), ruled at the gate | CRITICAL (gate item) | +| 11 | `netbox/dc-rack-mgmt-import.py` + its harness | imports vvr1-dcN as a NetBox DCIM device | NetBox-side decommission/repoint (system-of-record data migration, not just code) | MEDIUM | +| 12 | `scripts/site-baseleg.sh:40-48` DC rows | deferred BECAUSE DCs nest behind qemu+ssh (deferral answered by D-138) | premise disappears; new leg only if the target needs one -- currently a no-op | LOW | +| 13 | `scripts/dc-mirror.sh:6`, `maas-profile-assert.sh:39`, `maas-role-tags.sh:37`, `dc-egress-check.sh:63,72` | doc-comments / usage examples naming vvr1-dcN or "the rack host" | doc-currency only; logic is host-agnostic | LOW | +| 14 | `lib-hosts.sh:69-72` `NIC_PLANE_ORDER`/`BREX_PARENT_NIC`; `:29-34` `CARVE_AUX_HOSTS` | conventions realized by the inner root's MAC pinning | Conventions carry forward; the FILE that encodes the pinning changes (row 2). MAC re-capture owed post-rebuild (moot-ish: 10.13 rebuilds the fleet anyway) | LOW-MEDIUM | +| 15 | **Non-consumers, named to prevent double-counting:** `scripts/lib-net.sh` (0 hits -- IPAM facts change under D-143 only, a separate ruled axis); plane CIDRs/families (D-139); the 66+17 design-decision/CURRENT-STATE hits (append-only record) | -- | -- | NONE | + +--- + +## 3. Adversarial-check results (worker claims corrected or confirmed) + +All checks run against repo ground truth this session; a plausible finding was not accepted +as a verified one. + +1. **W0.3's "three ruled roles on vvr1-dcN" CONFLATES region and rack -- corrected.** The + D-132 AMENDMENT ADDENDUM (`docs/design-decisions.md:7216-7243`, RULED 2026-07-30, exact + utterance "Dedicated region VM per DC at utility .6 (Recommended)") puts each DC's MAAS + REGION (regiond + PostgreSQL) in its OWN VM -- `vr1-dcN-maas-01`, already one of the 12 + inner-root node VMs. **The region survives flattening as a flat sibling with no redesign.** + What vvr1-dcN actually carries: the MAAS **rack** controller (`--role rack`), the D-131 + node-DNS forwarder, and the D-138 client + SEC-028/SEC-029 credential residencies. This + RE-GRADES W0.3's CRITICAL: the execution-host question is real but narrower than "three + roles orphaned," and it is now ANSWERED-BY-PROPOSAL (Option 1's client VM), pending the gate. +2. **"Model A fallback is a validated artifact" -- TRUE but NOT TURNKEY; do not overstate.** + Verified: `docs/archive/model-a-fallback-plan.md` exists; tag `model-a-fallback` resolves to + `114d392`, dated **2026-07-16**. That PREDATES the D-132 amendment + addendum and D-138 + (both 2026-07-30), D-139's IPv6 family rulings (07-31), D-143 (10.13), the three utility + node VMs, and the lb-mgmt plane. Its own MAAS line -- "region on Office1 + rack" -- is + **contradicted by D-132 as now ruled**. Its Section-4 revert procedure (`git checkout + model-a-fallback -- opentofu/`) must NOT be run: it would resurrect a 10.12-era, 9-node, + Office1-region substrate. **The real build path is re-homing the CURRENT inner-root module + calls (W0.1's map) into a flat root, with Model A as shape precedent** -- "Model A PLUS the + rulings Model A never had," freshly validated. One thing the archive DOES verify in Option + 1's favor: Model A itself kept a small non-hypervisor `vvr1-dc0` headend (4/8192/80, + `expose_nested_virt=false`, metal-admin + transit legs) -- Option 1's client VM is exactly + that shape, repurposed. +3. **D-138 citation CONFIRMED.** `docs/design-decisions.md:7118-7124`: "the per-DC management + bastion holds only its own DC's cloud credential" / "Each DC has a management entry point + inside it; the NOC reaches that entry point." The RULING is "the cloud-facing client lives + IN the DC" -- a principle, with vvr1-dcN only as the then-concrete host. Option 1's client + VM (inside the DC's planes, holding one DC's credential) is **consistent with D-138's + principle**; the concrete-host change still needs recording (amendment vs. new D -- Phase 4 + frames it, Section 6). +4. **Cross-DC adjacency gap: W0.4 does NOT address it -- OPEN.** Grep-confirmed: the only + "cross-DC" text in `pass0-w4-targets.md` is the credential row of the tradeoff table. Both + options put both DCs' plane bridges and node VMs on vcloud's single libvirtd. See Section 5. +5. **Worker contradictions logged:** (a) W0.2's diagram says "416 GiB" for the containment VM + -- STALE; `variables.tf:143-150` shows 416->480 GiB raised 2026-08-01 (W0.1/W0.4 are + correct). (b) MTU: W0.2 rules the budget analytically unaffected; W0.4 lists it OWED -- + reconciled as: **analytically unchanged** (same `underlay_mtu=9000` var; vcloud-level + mesh/office1 networks already run 9000), with a live jumbo/geneve assert still owed + post-build (Section 8). (c) W0.3 row 4's "MAAS region is SHARED" echo (from SEC-026's + pre-amendment text) is superseded by D-132's per-DC regions -- SEC-026's regional-blast + -radius note is historical context, not current topology. +6. **Option-2 counter-evidence found during verification:** + `runbooks/dc-dc-phase4-juju-bundle-per-dc.md:167` -- tool-placement table row: "anything at + all | **never the vcloud jumphost** ... It picks the transport only" (measured 2026-07-29: + no juju/maas binaries on vcloud). Runbook doctrine, not a D-ruling -- but Option 2 would + reverse it, on top of inverting SEC-026's isolation control. +7. **No inferred-value violations found** in the four worker docs beyond the items above -- + claims spot-checked (ssh -J runbook lines, lib-hosts power URIs, SEC ledger rows, D-132 + addendum, sizing vars) all resolved to the cited lines. + +--- + +## 4. Target-topology options + recommendation + +Both options share the flat core: 12 node VMs + DC edge as vcloud-libvirt siblings; the six +planes re-homed to vcloud level (same CIDRs/families/MTU -- IPAM identity untouched); D-125 +bridge-in deleted (edge WAN -> direct NAT); mesh triangle unchanged; nesting depth 4 -> 2 +(VR0-proven); one tofu root/state axis, no bootstrap gate, no qemu+ssh dial. + +### Option 1 (RECOMMENDED) -- flat nodes + a small per-DC client VM (Model A shape, D-132/D-138-updated) + +A dedicated ~4/8192/80 non-hypervisor VM per DC (`vr1-dcN-client`), legs = metal-admin + +transit, carrying: the D-138 client role, the SEC-028/SEC-029 credential residencies, and +(pending Section 6) the D-131 forwarder / rack-controller remainder. It is a flat utility +sibling -- no nested libvirt, no `expose_nested_virt` -- so the containment PATTERN is gone +even though a small VM remains. + +- **For:** lowest-delta path (Model A precedent + current inner-root module bodies re-homed); + preserves SEC-026 per-DC credential isolation exactly as today; is LITERALLY the D-138 + Roosevelt bastion analog, rehearsed early (transfers to the pre-Roosevelt bare-metal test); + keeps the phase-4 "never the vcloud jumphost" doctrine intact. +- **Against:** SEC-010 must be re-authored + re-measured for the new VM (new NIC-naming trap); + D-123's one-command site-down is lost (re-earn via module-scoped destroy); the cross-DC gap + (Section 5) still needs a ruling; ~2 small VMs of overhead retained (~8 GiB each). +- **Capacity:** directionally frees ~176 GiB host RAM (2 x 96 GiB containment overhead removed, + minus 2 x 8 GiB client VMs); vCPU a wash. QUALITATIVE -- exact FIT owed (Section 8). + +### Option 2 (NOT recommended) -- fully flat: client + credentials on vcloud itself + +Same flat core; no client VM; juju/openstack + both DCs' MAAS-admin-scoped keys land on the +shared jumphost via a new host-side leg into metal-admin. + +- **For:** ~8 GiB/DC less overhead; one fewer VM class. +- **Against:** inverts SEC-026's load-bearing isolation control (both DCs' credentials on the + widest-blast-radius host); contradicts the phase-4 runbook's "never the vcloud jumphost" row; + no Roosevelt analog -- throwaway work undone at the bare-metal test; net-new unprecedented + surface (host-side veth/bridge into a MAAS plane on the live jumphost). + +**Recommendation: Option 1.** Every verified axis -- risk delta, credential isolation, +Roosevelt fidelity, runbook doctrine -- favors it; Option 2's only advantage is a marginal +~8 GiB/DC. The cross-DC gap (Section 5) applies EQUALLY to both options, so it is a gate +condition on the flattening itself, not a tiebreaker. Framing per W0.4: "eliminate the +container layer" is fully satisfied by Option 1 -- no VM is a hypervisor for another VM; the +retained client VM is the same class as juju-01/maas-01. If the operator's intent is literally +zero additional VMs, that is Option 2 with the above accepted knowingly. + +--- + +## 5. THE CROSS-DC ADJACENCY GAP -- unresolved by W0.4, stated plainly + +**W0.2's finding (Sec 4, "the single largest wiring risk"):** today dc0's and dc1's six plane +bridges live on SEPARATE kernels (each containment VM's own libvirtd). Flattening puts BOTH +DCs' plane bridges + node VMs co-resident on vcloud's ONE libvirtd/kernel for the first time +in any deployed shape. SEC-010 -- interface-scoped to transit legs that no longer carry this +function -- does NOT cover host-level cross-DC leakage (an accidental host address on a plane +bridge, a forward rule, br_netfilter interactions). **Neither W0.4 option addresses this; +verified absent from the doc.** Nuances carried honestly: the adjacency existed in Model A's +committed design, but Model A was never deployed and SEC-010 was priced AFTER it against +Model-B's separate-kernel shape -- so no ruled control covers the flattened adjacency. This is +DISTINCT from rebuilding SEC-010's transit drop on the client VM (row 6) -- two controls. + +Handling options for the gate (presented, not picked -- GA-R5): +- **(a) Accept co-residency + a new host-level control:** a vcloud-level nftables/isolation + artifact asserting no inter-plane/inter-DC forwarding, with a mechanical `--check` gate and + its own SEC-NNN row (SEC-010's proven pattern, one layer up). Designed in Phase 1/2. +- **(b) A per-DC isolation mechanism in the target design itself** (e.g., per-DC network + namespaces or equivalent separation on vcloud) -- stronger boundary, more engineering, + partially re-introduces the complexity being eliminated. +- **(c) Reject single-host flattening** for both DCs (not what the operator's directive + implies, listed for completeness). + +--- + +## 6. UNRESOLVED at Phase 0 / needs operator or later-phase decision + +1. **Target topology itself** -- the gate (Section 7). +2. **Cross-DC isolation mechanism** -- Section 5's (a)/(b)/(c); at minimum a Phase-1 design + item + a new SEC row if flattening proceeds. +3. **Rack-controller remainder + D-131 forwarder + mirror home:** does MAAS rackd co-locate + onto `vr1-dcN-maas-01` (region VM), onto the Option-1 client VM, or elsewhere? Ditto the + D-131 node-DNS forwarder and the `.4` artifact-service placement (`dc-mirror.sh` says "runs + on the rack host" today). W0.3 flagged; no worker resolved it. +4. **The client VM's octet + naming:** inherit the rack's `.2` metal-admin identity or take a + new utility-band octet (D-134 standing map is a cross-DC STANDARD -- needs a ruled octet); + name should NOT be `vvr1-dcN` (avoid conflation with the eliminated containment class). +5. **Transit leg's surviving purpose + SEC-010 re-pin points:** qemu+ssh dial disappears; + operator `ssh -J` access and any Office1-originated flows remain -- which ends get the + re-authored FORWARD-drop (client VM + voffice1?), and does the mesh triangle survive + unchanged (W0.4 implicitly kept it; W0.2 flagged it as a decision)? +6. **Credential-residency migration plan** (SEC-026/-028/-029 + `vm-secret-locations` `rack` + rows + rotation obligations that FIRE on rack rebuild) -- Phase 1/2 work, listed so it is + not lost. +7. **Teardown-primitive replacement:** the module-scoped group-destroy that re-earns D-122's + one-command site-down (Phase 4 module design). +8. **Decision-record form:** D-123 amendment vs. new D-number for the container-elim (+ the + D-138 concrete-host change and the D-125 bridge-in retirement ride along). Phase 4 FRAMES + it; the operator rules it. NOT ruled at this gate. +9. **NetBox DCIM migration** for the vvr1-dcN device records (row 11). + +--- + +## 7. >>> OPERATOR GATE (Phase-0 hard stop -- SCOPE plan Section 4) <<< + +**Decision put to the operator:** + +1. **Target topology -- confirm one:** + - **Option 1 (recommended):** flat node VMs on vcloud libvirt + one small non-hypervisor + per-DC client VM (metal-admin + transit legs) carrying the D-138 client role and the + per-DC credential residencies. + - **Option 2:** fully flat; client + both DCs' credentials on vcloud itself (weakens + SEC-026 isolation; contradicts the phase-4 "never the vcloud jumphost" doctrine row; no + Roosevelt analog). +2. **Cross-DC isolation gap (applies to EITHER option):** choose handling -- (a) accept + co-residency + a new host-level isolation control with a mechanical gate + SEC row + (designed Phase 1); (b) require a per-DC isolation mechanism in the target design; or + (c) reject single-host flattening. +3. **Role re-homing:** confirm the MAAS region stays on `vr1-dcN-maas-01` (per D-132 addendum + -- no change needed), and rule where the rack-controller remainder + D-131 forwarder land + (client VM / maas-01 / retire-with-evidence). + +Items 4-9 of Section 6 are flagged as downstream (Phase 1-4) work, not gate blockers. + +--- + +## 7a. PHASE-0 GATE OUTCOME (operator, 2026-08-09) + +Recorded as a DIRECTIONAL PLANNING CONFIRMATION at the Phase-0 gate -- NOT a GA-R5 [ARCH] +ruling and NOT a D-number (the formal container-elim ruling, D-123 amendment vs. new D, is +framed in Phase 4 and ruled then, per SCOPE Section 7). The operator was shown the +layer+separation diagram (artifact, Option-1 target) before confirming. + +- **Target topology: OPTION 1 CONFIRMED** -- flat node VMs on vcloud libvirt + one small + non-hypervisor per-DC client VM (metal-admin + transit legs) carrying the D-138 client role + and that DC's credential residencies. Phases 1-4 plan against this. +- **Cross-DC adjacency gap (Section 5): handling (a) CONFIRMED** -- accept co-residency and + design a new vcloud-level host isolation control (SEC-010's nftables pattern one layer up) + with a mechanical `--check` gate and its own SEC-NNN row. This is now a Phase-1 DESIGN ITEM, + not an open gate fork. +- **Role re-homing:** the MAAS region stays on `vr1-dcN-maas-01` (.6, D-132 addendum -- no + change). The rack-controller remainder + D-131 forwarder + artifact-service placement (client + VM vs. maas-01 vs. retire-with-evidence) is carried into Phase 1 as an open placement decision + (Section 6 item 3), not resolved at this gate. + +Feeds forward to Phase 1: plan against Option 1; treat (a) as a required design deliverable; +keep the container-elim change-set DISTINGUISHABLE from D-143's 10.12->10.13 address change +(two rulable axes riding one redeploy). + +## 8. OWED live measurements (read-only; do not infer -- W0.4 Sec 5 + reconciliation) + +1. **vcloud host-capacity currency:** `dc-dc-whole-host-budget.py:66-67`'s committed "MEASURED + host budget" (256 vCPU / 1024 GiB / 10240 GiB) needs a fresh read-only measurement before + any FIT verdict. +2. **Exact FIT for the flat 12-VM/DC roster:** the calculator lacks flags for the 3 utility + node classes (juju/maas/tailscale, added after authoring) -- extend or hand-total before + any freed-capacity number enters the Phase-4 change-set. Until then "~176 GiB freed" is + directional only. +3. **MTU/jumbo/geneve assert on vcloud-level planes post-build:** analytically unchanged + (Section 3.5), but the live gate (`geneve-encap-assert.sh` + a jumbo-path check) is owed + once the planes exist at vcloud level. diff --git a/docs/audit/container-elim-pass/pass0-w1-substrate-map.md b/docs/audit/container-elim-pass/pass0-w1-substrate-map.md new file mode 100644 index 0000000..d2b8052 --- /dev/null +++ b/docs/audit/container-elim-pass/pass0-w1-substrate-map.md @@ -0,0 +1,374 @@ +# Pass 0 / W0.1 -- Substrate/tofu map of the container layer + +READ-ONLY planning artifact. Container-layer-elimination pass, Phase 0, Worker 1. +Repo HEAD at write time: branch `dc-dc-stage5-preconditions`. All refs are path:line +against files as read this session; re-check line numbers if the file has moved since. + +--- + +## 1. As-is substrate diagram (text) + +``` +vcloud (outer libvirt, qemu:///system) <- OUTER ROOT: opentofu/main.tf +| +|-- module.office1_storage / office1_network / office1_opnsense (Office1 -- NOT container-layer) +|-- module.voffice1 (Office1 headend VM -- NOT container-layer) +|-- module.ubuntu_noble_base (shared base image, multi-consumer) +|-- module.mesh_vr1_dc0_vr1_dc1 / _office1 / vr1_dc1_office1 (D-100 mesh triangle -- outer, persists) +|-- module.netem_vr1_dc0_vr1_dc1 (tc netem on a mesh bridge -- outer, persists) +|-- module.vr1_dc0_storage / vr1_dc1_storage (outer per-DC pool -- holds ONLY the containment VM's own disk) +|-- module.vr1_dc0_uplink / vr1_dc1_uplink (D-125 vcloud-level ISP NAT /24 -- the ONE real NAT egress) +| +|-- module.vvr1_dc0 <==== CONTAINMENT VM (D-122/D-123 Model B) ==== +| cloudinit-vm: 108 vCPU / 480 GiB / ~3000 GiB disk (opentofu/variables.tf:137-156) +| NIC1 -> mesh_vr1_dc0_office1 (transit, region-facing, SEC-010 keys here) +| NIC2 -> vr1_dc0_uplink (IP-less bridge port -> br-vr1-dc0-wan) +| expose_nested_virt = true (LOAD-BEARING: svm passthrough for inner KVM) +| | +| |==== BOOTSTRAP GATE (scripts/site-headend-install.sh, node-host mode) ==== +| | installs: nested libvirtd, inner storage-pool dir + AppArmor grant, +| | kvm nested=1, stages OPNsense base image, verifies bridge + SEC-010 +| | FORWARD-drop scoped to the transit NIC. NOT an OpenTofu artifact -- +| | a bash script gate between the two tofu roots. +| | +| +-- INNER ROOT: opentofu/vr1-dc0-substrate/ (provider = qemu+ssh -> vvr1-dc0, R-5) +| |-- module.inner_storage (dc-storage-pool, INSIDE vvr1-dc0) +| |-- module.vr1_dc0_planes (dc-planes x6 -- the RELOCATED planes, D-123 Model B) +| |-- module.vr1_dc0_wan (wan-bridge -> br-vr1-dc0-wan, D-125 bridge-in) +| |-- module.vr1_dc0_opnsense (opnsense-edge, LAN=provider-public, WAN=vr1_dc0_wan) +| +-- module.vr1_dc0_node[9] (node-vm x9: 3 control+2 compute+4 storage, +| PLUS juju-01, maas-01, tailscale-01 = +| 12 for_each entries total, D-104/D-132/D-129(iii)) +| ++-- module.vvr1_dc1 <==== CONTAINMENT VM (mirror of vvr1_dc0) ==== + cloudinit-vm: 108 vCPU / 480 GiB / ~3000 GiB disk (opentofu/variables.tf:175-194) + NIC1 -> mesh_vr1_dc1_office1, NIC2 -> vr1_dc1_uplink + | + |==== BOOTSTRAP GATE (same script, dc1 target) ==== + | + +-- INNER ROOT: opentofu/vr1-dc1-substrate/ (provider = qemu+ssh -> vvr1-dc1) + mirrors dc0's inner root exactly: inner_storage, vr1_dc1_planes (x6), + vr1_dc1_wan, vr1_dc1_opnsense, vr1_dc1_node[N] (schematic MAC scheme + 52:54:01:d1:NN:PP, pre-pinned at authoring -- opentofu/vr1-dc1-substrate/main.tf:84-119) +``` + +Target (post-elimination, NOT yet confirmed -- Phase-0 W0.4's job, referenced here only +for contrast): `vcloud (libvirt) -> node VMs` directly, planes/wan/edge modules re-homed +onto the outer root or vcloud itself, containment VM + bootstrap gate + inner root deleted. + +--- + +## 2. Every file/module involved, with path:line + +### 2.1 Outer root (`opentofu/`) + +- `opentofu/main.tf` -- the whole outer root. Load-bearing container-layer blocks: + - `module "vvr1_dc0"` at `main.tf:410-519` -- the containment VM itself (cloudinit-vm). + - `module "vvr1_dc1"` at `main.tf:537-623` -- mirror for dc1. + - `module "vr1_dc0_uplink"` at `main.tf:379-384`, `module "vr1_dc1_uplink"` at + `main.tf:391-396` -- the D-125 vcloud-level ISP NAT that the containment VM's NIC2 + bridges into (site-wan module; this is the SOLE real internet egress for each DC). + - `module "vr1_dc0_storage"` (`main.tf:35-39`) / `vr1_dc1_storage` (`main.tf:49-53`) -- + outer per-DC pool. Post-Model-B this pool holds ONLY the containment VM's own boot + disk/seed, NOT node disks (those moved to `inner_storage` in the inner root). + - `module "mesh_vr1_dc0_vr1_dc1"` / `_office1` / `mesh_vr1_dc1_office1` (`main.tf:125-141`) + -- the D-100 mesh triangle. dc0/dc1<->office1 legs carry the D-124 transit the + containment VM's NIC1 rides; NOT container-layer-specific themselves (mesh survives + elimination, only what rides it changes -- see W0.2 network map). + - `module "netem_vr1_dc0_vr1_dc1"` (`main.tf:353-358`) -- tc netem on mesh bridge + virbr5; outer, independent of the containment layer. + - Comment block `main.tf:22-33` -- explicitly documents the D-123 Model B move: the + six vr1-dc0 planes moved to the inner root because "a libvirt provider cannot be + configured from a resource created in the same apply." + - Comment block `main.tf:300-331` -- names the apply order and the bootstrap-gate + script explicitly (quoted in section 3 below). + - `moved` blocks `main.tf:275-298` -- state-address rewrites from the pre-D-119/pre- + Model-B naming; historical, not currently container-layer-live, but relevant if + Phase 1 needs precedent for a future `moved{}` migration when the layer collapses. + +- `opentofu/variables.tf` -- sizing/addressing vars for the containment VMs: + - `vvr1_dc0_vcpu`/`_memory_mib`/`_disk_bytes` (`variables.tf:137-156`, defaults + 108 / 491520 MiB / 3,221,225,472,000 bytes) and the dc1 mirror + (`variables.tf:175-194`) -- derived via `scripts/dc-dc-whole-host-budget.py` + (comment cites the derivation, not invented). + - `vr1_dc0_rack_metal_admin_ip` / `_transit_ip` / `_transit_prefix` / `_transit_peer_ip` + (`variables.tf:226-244`) and the dc1 mirror (`variables.tf:201-219`) -- NO defaults, + HELD values, populated from `opentofu/d124-rack.auto.tfvars` (gitignored; confirmed + content read this session -- dc0 `10.12.8.2` / `172.31.0.2/30` peer `172.31.0.1`; + dc1 `10.12.68.2` / `172.31.0.6/30` peer `172.31.0.5`). + - `vr1_dc0_ssh_pubkey_path` / `vr1_dc1_ssh_pubkey_path` (`variables.tf:105-114,158-168`) + -- D-126 per-env dedicated SSH key each containment VM's cloud-init authorizes; the + inner root's qemu+ssh provider authenticates with the matching private half. + - `vr1_dc0_planes` / `vr1_dc1_planes` (`variables.tf:39-86`) -- the six-plane CIDR maps. + Declared in the OUTER root's variables.tf but consumed by the INNER root's + `dc-planes` module call, not the outer root itself (outer no longer instantiates + `dc-planes` for either DC) -- the outer var is the values-of-record copy the inner + tfvars mirrors (comment `variables.tf:70-73`). + +- `opentofu/d124-rack.auto.tfvars` -- gitignored; the four rack-addressing values per DC + (metal-admin IP, transit IP/prefix/peer) that key the containment VM's netplan AND the + inner root's qemu+ssh connection variable (`vvr1_dc0_transit_ip` in the inner root is a + SEPARATE var populated from the SAME measured IP, not auto-derived from this file -- + confirmed by reading `vr1-dc0-substrate/variables.tf:4-7`, which is a plain string var + with no default). + +### 2.2 Inner roots + +- `opentofu/vr1-dc0-substrate/main.tf` (266 lines, read in full): + - `provider "libvirt"` block (`main.tf:12-22`) -- the qemu+ssh dial into vvr1-dc0, + keyfile+sshauth REQUIRED (measured trap, no default identity works). + - `module "inner_storage"` (`main.tf:26-30`) -- `dc-storage-pool` INSIDE vvr1-dc0. + - `module "vr1_dc0_planes"` (`main.tf:34-40`) -- the six RELOCATED planes (`dc-planes` + module reused verbatim from `../modules/`). + - `module "vr1_dc0_wan"` (`main.tf:51-56`) -- `wan-bridge` onto `br-vr1-dc0-wan`. + - `module "vr1_dc0_opnsense"` (`main.tf:60-73`) -- the DC edge, `opnsense-edge` module. + - `locals.vr1_dc0_node_nics` (`main.tf:77-85`) + `locals.vr1_dc0_nodes` map + (`main.tf:96-247`) -- 12 entries: 9 D-121 Option-C role nodes (3 control/2 compute/4 + storage) + `vr1-dc0-juju-01` (D-104 amendment) + `vr1-dc0-maas-01` (D-132(iii)) + + `vr1-dc0-tailscale-01` (D-129(iii)) -- all MAC-pinned from live measurement. + - `module "vr1_dc0_node"` (`main.tf:250-265`) -- `for_each` over the 12-entry map, + instantiating `../modules/node-vm` once per node. +- `opentofu/vr1-dc0-substrate/variables.tf` (48 lines) -- `vvr1_dc0_transit_ip` / + `vvr1_dc0_ssh_user` / `vvr1_dc0_ssh_keyfile` (connection vars, MEASURED post-outer- + apply, no defaults on the IP/keyfile), `inner_pool_path` (default + `/var/lib/libvirt/vr1-dc0-inner`), `opnsense_base_path` (default staged path, must be + on the EXECUTING host's filesystem -- measured trap, see `variables.tf:26-33`), + `domain_suffix`, `underlay_mtu`, `vr1_dc0_planes` (mirrors outer var). +- `opentofu/vr1-dc0-substrate/versions.tf` -- separate `terraform{}` block, same + provider pin (`dmacvicar/libvirt` 0.9.8) as the outer root's `versions.tf`. +- `opentofu/vr1-dc1-substrate/{main,variables,versions}.tf` -- structural MIRROR of the + dc0 inner root (confirmed by reading both `main.tf` files side by side): same module + set (`inner_storage`, `vr1_dc1_planes`, `vr1_dc1_wan`, `vr1_dc1_opnsense`, + `vr1_dc1_node[for_each]`), only the addressing/MAC values differ (dc1 uses a + DETERMINISTIC pre-assigned MAC scheme `52:54:01:d1:NN:PP`, `main.tf:84-98`, vs dc0's + measured-and-pinned `52:54:00:*` values -- a genuine divergence in HOW the MACs were + assigned, not a container-layer-elimination-relevant difference). + +### 2.3 Modules touched by the container layer + +Reused VERBATIM by both the outer root (for the containment VM itself) and the inner +roots (for what runs inside it) -- same module source, different provider/root: + +| Module | Used by outer root for... | Used by inner root(s) for... | +|---|---|---| +| `modules/cloudinit-vm` | the containment VM (vvr1-dc0/vvr1-dc1) itself, AND voffice1/office1 edge (non-container uses) | not used inside | +| `modules/dc-storage-pool` | outer per-DC pool (now holds only the containment VM's own disk) | `inner_storage` (holds node + edge disks) | +| `modules/dc-planes` | NOT instantiated (retired outer call, per `main.tf:44-47` comment) | `vr1_dc0_planes` / `vr1_dc1_planes` -- the six planes, now INSIDE the containment VM | +| `modules/wan-bridge` | not used | `vr1_dc0_wan` / `vr1_dc1_wan` -- D-125 bridge-in onto the containment VM's own `br-vr1-dc0-wan` bridge | +| `modules/opnsense-edge` | office1_opnsense (non-container) | `vr1_dc0_opnsense` / `vr1_dc1_opnsense` -- the DC edge, now INSIDE the containment VM | +| `modules/node-vm` | not used (node VMs never lived on vcloud) | `vr1_dc0_node[*]` / `vr1_dc1_node[*]` -- all 12+12 node VMs | +| `modules/site-wan` | `vr1_dc0_uplink`/`vr1_dc1_uplink` -- the OUTER vcloud-level ISP NAT (persists post-elimination; not container-specific) | office1-wan is also this module (non-container) | +| `modules/mesh-link` | the D-100 mesh triangle (persists; carries the transit the containment VM currently uses) | not used | +| `modules/netem-link` | dc0<->dc1 mesh netem (persists) | not used | +| `modules/base-image` | shared Ubuntu noble base (persists; containment VM boots from it via cloudinit-vm) | not used | +| `modules/maas-vm-host` | NOT instantiated anywhere (DEFERRED/REFUTED, see its own header: snap MAAS cannot open the local libvirt socket, qemu+ssh pod fails `domblkinfo`) | same -- refuted for DC use by measurement 2026-07-20 | +| `modules/office1-network` | office1-local L2 (non-container, Office1-only) | not used | + +None of `modules/{cloudinit-vm,dc-storage-pool,dc-planes,wan-bridge,opnsense-edge, +node-vm,site-wan,mesh-link,netem-link,base-image}` contain any hardcoded reference to +"vvr1-dc0"/"containment" in their own HCL -- they are generic and reused verbatim. The +container-layer-SPECIFIC facts live entirely in the two ROOTS (which provider/state each +module instance is created under) and in the containment VM's own module call +(`vvr1_dc0`/`vvr1_dc1` in the outer root). This is exactly what `main.tf:326-330`'s +comment states: "Their HCL is UNCHANGED -- the same node-vm / site-wan / opnsense-edge / +dc-planes modules are reused verbatim; only the provider they run against moved from +vcloud to vvr1-dc0." + +--- + +## 3. Outer<->inner apply ordering (as documented, `main.tf:326-330`) + +Verbatim from the outer root's own comment: + +> "Apply order: OUTER (this root -- boots + sizes vvr1-dc0) -> BOOTSTRAP GATE +> (site-headend-install.sh: install nested libvirtd + inner pool + kvm nested=1 + +> stage the opnsense base image) -> INNER root (planes/wan/edge/nodes)." + +Concretely, three distinct steps with two different tools: + +1. **OUTER apply** (`opentofu/` root, `qemu:///system` on vcloud, run per D-128 Plane 1 + from vcloud itself, confirmed `docs/CURRENT-STATE.md:159,178`) -- creates + `vvr1_dc0`/`vvr1_dc1` (the containment VM), the D-125 uplink NAT, the mesh legs, the + outer per-DC storage pool. +2. **BOOTSTRAP GATE** (`scripts/site-headend-install.sh`, "node-host mode" / `--role rack`) + -- a bash script, NOT OpenTofu, run against the now-booted containment VM over SSH. + Installs nested libvirtd, creates the inner pool directory + AppArmor grant + (`site-headend-install.sh:264-272`), sets `kvm nested=1`, stages the OPNsense base + image onto the EXECUTING host's filesystem (voffice1, per the inner root's + `opnsense_base_path` var comment), and verifies the transit/uplink bridges + the + SEC-010 FORWARD-drop scoped to the transit interface. +3. **INNER apply** (`opentofu/vr1-dc0-substrate/` or `vr1-dc1-substrate/`, `qemu+ssh` + dialed from the EXECUTING host -- per D-128 the inner tofu roots run from voffice1, + confirmed `docs/CURRENT-STATE.md:3678` "D-128 has it run the INNER tofu roots and + `tofu init` leaves a provider cache") -- creates the inner storage pool, the six + relocated planes, the WAN bridge, the DC edge, and the 12 node VMs per DC. + +This ordering is STRUCTURALLY FORCED, not a preference: `main.tf:28-29`'s comment states +the reason plainly -- "A libvirt provider cannot be configured from a resource created +in the same apply, so the inner substrate is a separate root/state applied AFTER +vvr1-dc0 is up (the bootstrap gate)." Eliminating the containment VM removes this +structural forcing function entirely -- a flat topology would let ONE root (or a set of +modules under one root/state) create the planes/edge/nodes directly on vcloud's own +`qemu:///system` provider, with no bootstrap-gate script and no second `tofu init`/apply +cycle. This is very likely the single largest simplification the elimination buys, +structurally speaking (Phase 1/2 should confirm the operational-workflow-step count this +removes). + +--- + +## 4. Disk/pool/cloud-init/seed-volume dependency chain + +Two INDEPENDENT chains today, one per root: + +**Outer chain** (containment VM's own boot disk): +`module.ubuntu_noble_base` (`base-image`, downloaded once, `main.tf:165-173`) --backing_store--> +`module.vvr1_dc0`/`vvr1_dc1` (`cloudinit-vm`) which creates, per instance: + - `libvirt_volume.disk` (COW backing off the shared base image, + `modules/cloudinit-vm/main.tf:24-47`) + - `libvirt_cloudinit_disk.seed` + `libvirt_volume.seed` (the NoCloud ISO, + `modules/cloudinit-vm/main.tf:49-76`) -- `lifecycle { ignore_changes = [create] }` + guards against the D-130 staging-path-remint-on-reboot trap. + - both volumes live in `module.vr1_dc0_storage`/`vr1_dc1_storage` (the OUTER per-DC + pool, `dc-storage-pool`, `main.tf:35-53`). + +**Inner chain** (everything the containment VM hosts): + - `module.inner_storage` (`dc-storage-pool`, `vr1-dc0-substrate/main.tf:26-30`) is a + FRESH pool created at `var.inner_pool_path` (default + `/var/lib/libvirt/vr1-dc0-inner`) -- a directory the BOOTSTRAP GATE (not OpenTofu) + provisions on the containment VM's own filesystem before this root can apply. + - `module.vr1_dc0_node[*]` (`node-vm`) creates a BLANK `libvirt_volume.disk` per node + (PXE-boot pattern, no backing image -- `modules/node-vm/main.tf:44-62`) plus an + opt-in OSD volume for storage-role nodes, all in `inner_storage`. + - `module.vr1_dc0_opnsense` (`opnsense-edge`) creates its disk as a DIRECT COPY (not + COW) of `var.opnsense_base_path` (`modules/opnsense-edge/main.tf:55-85`) -- a file + that must exist on the EXECUTING host's filesystem (the qemu+ssh provider uploads + from its own client-side filesystem, NOT from a path on vvr1-dc0 -- measured trap, + `vr1-dc0-substrate/variables.tf:26-33`), also placed in `inner_storage`. + +**Key coupling for elimination planning:** the inner chain's `inner_storage` pool is +NOT the same pool as the outer `vr1_dc0_storage`/`vr1_dc1_storage` pool -- they are two +separate `dc-storage-pool` module instances under two separate roots/providers, with the +inner one's target directory living INSIDE the containment VM's own disk image (itself a +volume in the outer pool). A flat topology collapses this to ONE pool per DC on vcloud's +own filesystem holding node/edge/pool volumes directly -- no nested-disk-inside-a-disk +indirection, and no bootstrap-gate-provisioned directory dependency. + +--- + +## 5. Bulleted list -- exactly what is containment-layer-specific + +**Would be REMOVED entirely if the containment VM is eliminated:** +- `module "vvr1_dc0"` / `module "vvr1_dc1"` (`opentofu/main.tf:410-519`, `537-623`) -- + the containment VM domain + its disk + seed volumes. +- The ENTIRE inner roots: `opentofu/vr1-dc0-substrate/{main,variables,versions}.tf` and + `opentofu/vr1-dc1-substrate/{main,variables,versions}.tf` as separate roots/states + (their MODULE CALLS would likely be re-homed into the outer root or a new flat root, + not deleted -- see below). + - `module.inner_storage` in each inner root -- re-homed (folds into, or replaces, the + outer per-DC storage pool). + - The BOOTSTRAP GATE step of `scripts/site-headend-install.sh` "node-host mode" -- + nested-libvirtd / inner-pool / kvm-nested=1 install -- has NOTHING to install once + there is no nested hypervisor. +- The two `qemu+ssh` provider connections (`vr1-dc0-substrate/main.tf:11-22`, + `vr1-dc1-substrate/main.tf:11-19`) and their connection variables + (`vvr1_dc0_transit_ip`/`_ssh_user`/`_ssh_keyfile` and dc1 mirror) -- a flat topology + needs no remote-libvirt dial at all; the outer root's `qemu:///system` covers + everything. +- `module "vr1_dc0_wan"` / `vr1_dc1_wan` (`wan-bridge` module calls) -- these exist ONLY + to bridge the inner OPNsense edge onto the containment VM's own `br-vr1-dc0-wan` + netplan bridge; with no containment VM there is no such bridge to attach to. The DC + edge's WAN would need to reconnect DIRECTLY to `vr1_dc0_uplink`/`vr1_dc1_uplink` (the + outer D-125 NAT network) instead -- likely reverting toward the pre-Model-B ("Model A") + shape `docs/model-a-fallback-plan.md` already describes as the fallback. +- The containment VM's own two-NIC netplan (`main.tf:495-518`, `599-621`): the + `br-vr1-dc0-wan`/`br-vr1-dc1-wan` bridge declarations, the transit-leg static + addressing consumed only by the inner qemu+ssh dial, and the D-124 rack-addressing + vars that exist ONLY to address the containment VM + (`vr1_dc0_rack_transit_ip`/`_prefix`/`_peer_ip`, `vr1_dc0_rack_metal_admin_ip`, and + dc1 mirrors -- `variables.tf:196-244`). +- `vvr1_dc0_vcpu`/`_memory_mib`/`_disk_bytes` and dc1 mirror (`variables.tf:137-156, + 175-194`) -- sizing vars that exist only because the containment VM must hold the + ENTIRE node fleet (104 node vCPU + 4-96 GiB overhead derivation). A flat topology + sizes each node VM directly against vcloud's own budget instead (no "containment + overhead" line item at all -- `scripts/dc-dc-whole-host-budget.py`'s + `--containment-overhead-*` flags would need re-deriving or dropping). +- `vr1_dc0_ssh_pubkey_path`/`vr1_dc1_ssh_pubkey_path` (D-126 per-env keys) and the + `~/vr1-dc0-creds/`, `~/vr1-dc1-creds/` key material they authorize -- these exist + specifically to authenticate the inner qemu+ssh provider to the containment VM; a flat + topology has no SSH-to-a-nested-hypervisor step to key. + +**Would be RE-HOMED (kept, but moved to the outer root / vcloud's own provider), not +deleted:** +- `module "vr1_dc0_planes"` / `vr1_dc1_planes` (`dc-planes` x6 each) -- the six planes + themselves are NOT container-layer artifacts; only WHERE they are created (inside vs. + outside the containment VM) is. This is precisely the D-123 Model B move already + documented at `main.tf:22-33` and `CURRENT-STATE.md:7676-7682` (2.3-iii) -- eliminating + the container layer is largely UNDOING that specific move: planes go back to being + created directly by the outer root (or a flat per-DC root) on vcloud's own libvirt, + the way they stood before D-123 Model B (Model A). +- `module "vr1_dc0_opnsense"` / `vr1_dc1_opnsense` (`opnsense-edge`) -- the DC edge VM + itself; only its WAN attachment (currently `vr1_dc0_wan`/wan-bridge) needs to change + to attach directly to the outer `vr1_dc0_uplink`/`vr1_dc1_uplink` NAT network instead. +- `module "vr1_dc0_node"` / `vr1_dc1_node` (`for_each` node-vm, 12 entries each) -- the + node VMs themselves; only their provider (containment-VM qemu+ssh -> vcloud + qemu:///system) and their storage-pool parent (inner_storage -> a per-DC pool on + vcloud) change. +- `module.inner_storage` -- effectively MERGES with (or is replaced by) the existing + outer `vr1_dc0_storage`/`vr1_dc1_storage` pool, since there is no longer an "inner" + vs "outer" distinction once everything is on vcloud. + +**Would be UNCHANGED / NOT container-layer-specific (persist regardless):** +- `module "voffice1"`, `office1_opnsense`, `office1_network`, `office1_storage`, + `ubuntu_noble_base` -- Office1 headend, untouched. +- `module "mesh_vr1_dc0_vr1_dc1"`, `mesh_vr1_dc0_office1`, `mesh_vr1_dc1_office1` + (mesh-link) and `netem_vr1_dc0_vr1_dc1` -- the D-100 mesh triangle. NOTE for W0.2: what + RIDES the office1<->dc0/dc1 mesh legs changes (today: the transit to the containment + VM's NIC1; post-elimination: whatever replaces the region<->rack path, if anything is + still needed once there is no separate rack/region split) -- flagged as a W0.2 item, + not resolved here. +- `module "vr1_dc0_uplink"` / `vr1_dc1_uplink` (site-wan, the D-125 vcloud-level ISP + NAT) -- persists as the DC's egress; only what attaches to it changes (today: the + containment VM's bridge port; post-elimination: directly the DC edge's WAN NIC, likely + reverting to a `site-wan`-module-style NAT attachment matching Office1's own pattern). +- `modules/maas-vm-host` -- already NOT instantiated for either DC (refuted by + measurement, unrelated to this pass). + +--- + +## 6. Top risks / unknowns for eliminating this layer (structural, from this dimension only) + +1. **The D-125 bridge-in mechanism (`wan-bridge` module) has no proven Model-A + equivalent in this codebase's CURRENT config** -- Model A's original NAT-based WAN + shape was RETIRED when Model B replaced it (`main.tf:44-47`'s comment: "the old + pre-Model-B `module vr1_dc1_planes` skeleton... is retired -- do not re-add"). Reverting + needs either resurrecting that retired shape or re-deriving it; `docs/model-a-fallback- + plan.md` exists and should be W0.4's/Phase-1's first read, not re-invented. +2. **Node VM count triples per DC's libvirt host under a flat topology**: today ONE + containment-VM domain exists per DC at the outer-root level; post-elimination the + outer root (or its replacement) would directly own 12 node-vm domains + 1 edge domain + + 6 planes PER DC = ~19 domains x 2 DCs = ~38 domains, all on vcloud's single outer + libvirt/state. This is a state-management and blast-radius change beyond "flatten + the topology" -- Phase 4's module design should weigh one-state-per-DC vs. one-state- + total for this reason. +3. **MAC-pinning discipline must survive the move.** Every node/edge MAC is currently + pinned from LIVE MEASUREMENT taken against the containment VM's virsh + (`virsh domiflist` on vvr1-dc0/vvr1-dc1). Re-homing these `for_each` maps to a new + provider will very likely register as a `-/+` (force-replace, since `name` is + ForceNew per the existing `moved{}` block precedent at `main.tf:256-271`) or at minimum + demands re-measuring and re-pinning 12x2=24 nodes' MACs plus 2 edges' MACs after the + move -- MAAS re-enlistment risk is real and large, not cosmetic (the exact "MAC + regeneration trap" incident class this repo already suffered once, per multiple + comments cross-referenced above). +4. **Whether the outer root's `qemu:///system` (vcloud) provider is even the RIGHT single + target is a W0.4/operator question, not resolved here** -- this doc only maps what + exists; it takes no position on the target topology. +5. **The bootstrap-gate script (`site-headend-install.sh`) has a SECOND role** (region+ + rack MAAS enrollment, SEC-010 nftables) beyond nested-libvirt bring-up -- eliminating + the containment layer removes the nested-libvirt PART of its job but the script likely + still needs SOME of its other duties (rack MAAS role, SEC-010) reassigned somewhere; + W0.3 (consumer inventory) owns confirming exactly which parts survive. +6. **Not verified this session (out of W0.1's read-only substrate scope, flagged for + W0.2/W0.3):** whether `underlay_mtu` (jumbo, 9000) assumptions inside the currently- + nested plane networks change once they run on vcloud's own bridges directly (host NIC + MTU vs. a nested-guest's virtio MTU) -- geneve-over-v6 budget is explicitly a W0.2 + deliverable per `phase-prompts.md`; this doc does not check it. diff --git a/docs/audit/container-elim-pass/pass0-w2-network-map.md b/docs/audit/container-elim-pass/pass0-w2-network-map.md new file mode 100644 index 0000000..f5e12a3 --- /dev/null +++ b/docs/audit/container-elim-pass/pass0-w2-network-map.md @@ -0,0 +1,217 @@ +# Pass 0 / W0.2 -- Network & wiring map of the container layer + +**Agent:** W0.2 (sonnet worker), container-elim pass, Phase 0. Read-only. All values below are +cited to path:line or to a ruled D-NNN/SEC-NNN entry; nothing is inferred. Dated 2026-08-09. + +--- + +## 1. As-is topology -- text diagram (dc0 arm; dc1 is the structural mirror) + +``` +vcloud (outer libvirt, qemu:///system) -- opentofu/main.tf +| +|-- office1_network (module, MTU=underlay_mtu=9000) -- main.tf:76-80 +| `-- voffice1 (LXD/MAAS-region/NetBox host) -- main.tf:175 (NIC1 enp1s0, DHCP/Kea) +|-- office1-wan (site-wan NAT, MTU 1500, not jumbo) +|-- office1_opnsense edge -- main.tf:98-114 +| +|-- D-100 dark-fiber mesh TRIANGLE (mesh-link module, MTU=underlay_mtu=9000) -- main.tf:125-141 +| |-- mesh_vr1_dc0_vr1_dc1 (dc0<->dc1) bridge virbr5 (MEASURED, `virsh net-info`) +| | netem: delay 3ms jitter 1ms loss 0.01% (PLACEHOLDER, D-100 gap #11 unruled) +| | -- main.tf:352-357 netem_vr1_dc0_vr1_dc1 +| |-- mesh_vr1_dc0_office1 (dc0<->office1, region<->rack TRANSIT) +| | region end: voffice1 NIC2 static 172.31.0.1/30 -- main.tf:191-192 +| | rack end: vvr1-dc0 NIC1 (enp1s0) static 172.31.0.2/30 -- main.tf:498-505, D-124 +| `-- mesh_vr1_dc1_office1 (dc1<->office1 transit) +| region end: voffice1 NIC3 static 172.31.0.5/30 -- main.tf:193 +| rack end: vvr1-dc1 NIC1 static 172.31.0.6/30 -- main.tf:602-609, D-124 amdt +| (transit supernet 172.31.0.0/24, D-124 AMENDMENT 2026-07-16, design-decisions.md:5031) +| +|-- vr1_dc0_uplink (site-wan NAT, 172.30.2.0/24, MTU 1500) -- main.tf:379-385, D-125/D-115 +|-- vr1_dc1_uplink (site-wan NAT, 172.30.3.0/24, MTU 1500) -- main.tf:388-395 +| +`-- vvr1-dc0 CONTAINMENT VM (416 GiB/108 vCPU, expose_nested_virt=true) -- main.tf:410-522, D-123/D-124 + NIC1 enp1s0 = transit (region-facing; SEC-010 --transit-if KEYS HERE) + NIC2 enp2s0 = uplink, IP-LESS, enslaved into netplan bridge br-vr1-dc0-wan (declared + IN vvr1-dc0's own cloud-init netplan, D-125 bridge-in) -- main.tf:507-522 + | + `== INNER root: opentofu/vr1-dc0-substrate/ (qemu+ssh to vvr1-dc0 over the transit, + R-5, run from Office1 per D-128) == + |-- inner_storage pool -- vr1-dc0-substrate/main.tf:26-30 + |-- vr1_dc0_planes (dc-planes module, isolated L2, no forward/DHCP, + | MTU = var.underlay_mtu = 9000) -- vr1-dc0-substrate/main.tf:32-40 + | SIX PLANES (D-052/D-100 template, D-139-amended families): + | provider-public 10.12.4.0/22 dual-stack v4 + GUA 2602:f3e2:f02:10::/64 + | metal-admin 10.12.8.0/22 dual-stack v4 + GUA 2602:f3e2:f02:20::/64 + | metal-internal 10.12.12.0/22 IPv6-ONLY GUA 2602:f3e2:f02:21::/64 (v4 REMOVED, D-139 Ruling A) + | data-tenant 10.12.16.0/22 IPv6-ONLY GUA 2602:f3e2:f02:30::/64 (geneve underlay) + | storage 10.12.32.0/22 IPv6-ONLY GUA 2602:f3e2:f02:40::/64 (Ceph public) + | replication 10.12.36.0/22 IPv6-ONLY GUA 2602:f3e2:f02:50::/64 (Ceph cluster, cross-DC leg) + | lb-mgmt (NEW) n/a IPv6-ONLY GUA 2602:f3e2:f02:80::/64 (RESERVED, no charm consumer -- G18) + | lib-net.sh mirrors dc0's v4 CIDRs: PLANE_CIDRS (lib-net.sh:22), PLANE_NAME (:23-30) + |-- vr1_dc0_wan (wan-bridge module: forward={mode=bridge} onto br-vr1-dc0-wan; + | NO mtu block -- provider REJECTS on bridge-mode; inherits host + | bridge MTU=1500 from the outer netplan) -- vr1-dc0-substrate/main.tf:51-56 + |-- vr1_dc0_opnsense DC edge (LAN=provider-public, WAN=vr1_dc0_wan bridge) + | -- vr1-dc0-substrate/main.tf:60-73 + `-- vr1_dc0_node x 12 (9 role + juju-01 + maas-01 + tailscale-01) + NIC_PLANE_ORDER (lib-hosts.sh:69): metal-admin, provider-public, + metal-internal, data-tenant, storage, replication + role nodes (9): all 6 planes + OVS br-ex parented on enp2s0 + (provider-public leg; BREX_PARENT_NIC, lib-hosts.sh:72) -- enp2s0 + itself carries NO L3, br-ex carries the static (Pattern A, D-100/D-060) + juju-01 / maas-01 / tailscale-01 (utility VMs): 2 planes only + (metal-admin + provider-public, raw NIC WITH gateway, NO br-ex) + -- lib-hosts.sh:74-93 +``` + +**dc1 arm** is structurally identical: `vvr1-dc1` (main.tf:524-618), its own inner root +`opentofu/vr1-dc1-substrate/`, planes 10.12.64/68/72/76/80/84.0/22 (lib-net.sh:171-179, +D-124 AMENDMENT 2026-07-21), transit 172.31.0.4/30, uplink 172.30.3.0/24. + +--- + +## 2. SEC-010 -- transit FORWARD-drop (forwarding choke point) + +- **Scope:** interface-scoped FORWARD-drop on the TRANSIT leg only (`enp1s0` on vvr1-dc0), + never a global `ip_forward=0` and never globalized on the bridge -- `br_netfilter` makes + bridged WAN frames traverse the L3 FORWARD chain, so a global drop would silently kill + `br-vr1-dc0-wan` (design-decisions.md D-125, "br_netfilter CONSTRAINT"). +- **Artifact:** `scripts/site-headend-install.sh --host-nodes` writes + `/etc/nftables-sec010.nft` + boot-persistent `sec010-fw.service`; `--host-nodes --check` + is the mechanical gate (fails if the rule is absent OR the keyed transit interface does + not exist -- an nftables oifname on an absent iface loads clean but matches nothing, + fail-open class). CLOSED 2026-07-20, both ends (security-ledger.md SEC-010 row). +- **Why it sits where it does (nesting-specific rationale, verbatim from the ledger):** + *"after the D-123 Model B reshape ... vvr1-dc0 now bridges ALL 6 inner planes + the + transit, and belongs on the inner libvirt host, not just a 2-leg rack."* SEC-010 is + therefore the control that keeps vvr1-dc0's six bridged inner planes from being reachable + FROM the region side of the transit -- it is doing the isolation job a separate host would + otherwise do "for free." +- **The same pin follows onto voffice1's own transit leg** (its `enp2s0`), independently + verified (SEC-010 row, security-ledger.md). + +--- + +## 3. MTU / jumbo / geneve-over-v6 budget + +- `underlay_mtu = 9000` (jumbo) is the ONE knob applied to every plane network AND every + mesh-link network: `opentofu/dc-dc-phase0.auto.tfvars:7` -- *"Step 3 ruling: jumbo internal + fabric (tenant MTU stays 1500)"*. Consumed by `vr1-dc0-substrate/main.tf:38` (`mtu = + var.underlay_mtu`) and `main.tf:127/133/139` (mesh links) and `main.tf:78` + (office1_network). +- **The WAN path is deliberately NOT jumbo:** `modules/wan-bridge` (bridge-mode, no `mtu` + block -- the provider REJECTS it in bridge mode, measured 2026-07-20) inherits the host + bridge's 1500 from vvr1-dc0's outer netplan; `modules/site-wan` (office1-wan, the two + uplink NATs) defaults to 1500 -- "the ISP-uplink domain; NOT the jumbo planes/mesh" + (`vr1-dc0-substrate/main.tf:55`). +- **D-101 tenant-MTU sub-policy** (design-decisions.md ~line 2280): geneve-over-v6 overhead + is roughly 56 bytes (IPv6 40 + UDP 8 + Geneve base 8) before nested-virt/options. With the + jumbo (9000) underlay, tenant MTU stays 1500 and amphora/geneve fit; if the underlay were + pinned at 1500 instead, tenant MTU would drop to ~1444 (v6-geneve) and must be set + consistently across ovn geneve, tenant-network MTU, and amphora. **The jumbo underlay is + what makes tenant MTU 1500 possible at all** -- this is the plane-level fact the + container-elim must preserve. +- **D-139 (2026-07-31):** data-tenant (the geneve underlay plane) moved from IPv4 to + IPv6-only GUA (`2602:f3e2:f02:30::/64` dc0) -- geneve now rides GUA, not the historically + planned ULA (Ruling B, "full GUA on every plane"). +- **The 2026-08-08/09 geneve-over-v6 root cause (memory + `scripts/geneve-encap-assert.sh` + header) was NOT an MTU/byte-budget defect.** Two distinct causes, both at the OVN/OVS + layer, neither is the vvr1-dcN containment hop itself: + 1. Encap-family SPLIT between the containerized control plane (LXD, v4) and metal compute + chassis (v6) -- cross-family tunnels never form. Gate: `geneve-encap-assert.sh` C1. + 2. ovn-chassis 24.03 emitted a BRACKETED v6 `ovn-encap-ip` (`[2602:...]`) that OVS's + geneve implementation rejects ("bad geneve 'remote_ip'"), leaving v6 tunnels at + ofport -1. Gate: `geneve-encap-assert.sh` C2. + Fixed live: unbracket + carve v6 on the LXD/containerized chassis + `overlay_ip_version=6`; + VM->VM 8/8 0% loss confirmed (memory: dc0-checkpoint-then-reip-redeploy.md). **NAMING TRAP + reminder (per this pass's own framing): the "containerized control plane" here is the LXD + API-charm containers on nodes, NOT the `vvr1-dcN` containment VM this pass eliminates -- + do not credit or blame container-elim for this fix.** +- **Net conclusion:** the plane-level MTU/geneve byte budget (9000 underlay / 1500 tenant / + ~56-byte v6-geneve overhead) is **unaffected by container-elim** -- the vvr1-dcN hop for + plane traffic was already a same-MTU (9000) isolated libvirt bridge with no extra + encapsulation, so removing it removes a bridging HOP, not a budget constraint. + +--- + +## 4. What "simplify the wiring" precisely touches if vvr1-dcN is eliminated + +**COLLAPSES entirely (exists only because of the nesting):** +- vvr1-dc0's own two outer NICs (transit `enp1s0`, uplink `enp2s0`) and the containment VM + itself -- their sole purpose is to give the inner `qemu+ssh` provider (R-5) something to + dial and to give the inner WAN bridge an egress path. +- `br-vr1-dc0-wan` (the netplan bridge INSIDE vvr1-dc0) and the `modules/wan-bridge` + network realizing D-125's bridge-in -- this exists purely to fix OBS-3 (nested WAN NAT + losing egress under Model B). Flat topology reverts to Model A's direct `site-wan` NAT; + the repo already names this exact revert: "Fallback: `docs/archive/model-a-fallback-plan.md` + section 3 (revert removes the uplink NIC/network/bridge + `wan-bridge` and restores the + OPNsense WAN addr)" (D-125 entry). +- The two-stage OpenTofu apply ordering (outer boots+sizes vvr1-dc0 -> bootstrap gate + installs inner libvirtd -> inner root applies) and the `qemu+ssh` inner-provider URI + itself (`vvr1_dc0_transit_ip` dial) -- one fewer apply stage, one fewer credential/URI to + carry (SEC-026 already flags the MAAS-admin-key residency problem this ordering created). + +**RE-HOMES onto vcloud libvirt directly (same module, same shape, different provider target):** +- The SIX plane bridges (`modules/dc-planes`) -- same CIDRs, same MTU var (9000), same + isolated/no-forward/no-DHCP posture (NetBox/OpenStack still own L3). D-139's family matrix + is an OpenStack/NetBox-layer concern, not a libvirt-nesting one -- unaffected. +- The 12 node-vm module calls per DC (9 role + juju-01/maas-01/tailscale-01) -- same + NIC_PLANE_ORDER, same per-role carve shape (6 planes + br-ex for role nodes; 2 planes, no + br-ex, for utility VMs), same MAC-pin-after-first-apply discipline. +- The DC OPNsense edge -- LAN stays provider-public; WAN reverts to a direct `site-wan` NAT + (Model A shape) instead of routing through a bridge inside a containment VM. + +**REQUIRES A TARGET-TOPOLOGY DECISION (flagged for W0.4, not resolved here):** +- **SEC-010's isolation boundary.** Today SEC-010 protects ONE DC's forwarding surface + because vvr1-dc0 and vvr1-dc1 are SEPARATE nested libvirt hosts -- each bridges only its + own six planes. If both DCs' planes flatten onto vcloud's SAME libvirtd, dc0's and dc1's + plane bridges become co-resident on ONE host for the first time, a cross-DC adjacency + that did not exist before and that SEC-010 (interface-scoped to a transit leg that no + longer carries this function) does not cover. **This is the single largest wiring risk of + elimination** and needs an explicit new host-level control (or a proposed topology that + avoids co-residency) before Phase 1. +- **The transit leg's (D-124) surviving purpose.** Its `qemu+ssh`-dial role disappears with + the inner root, but MAAS PXE/commissioning reachability from Office1 to a DC-local + region (D-132) and jujud's MAAS-API dial may still need SOME region<->rack leg depending + on whether the target topology keeps "DC-local" as a routing/administrative concept at + all. W0.3/W0.4's call. +- **The D-100 mesh triangle's fate.** The three mesh-link segments represent inter-SITE + links (dark-fiber stand-ins), a concept orthogonal to intra-site containment -- they likely + survive container-elim UNLESS the target topology also collapses "DC" as a distinct site + into one shared vcloud fabric with no site-to-site fiction needed. Also a W0.4 call. + +--- + +## 5. Files / lines this map is built from (for W0.4 and later phases) + +- `opentofu/main.tf`: 76-148 (office1 net/edge, mesh triangle), 175-233 (voffice1 3-NIC + transit host), 341-397 (netem, uplink NATs), 399-522 (vvr1-dc0 containment VM), 524-618 + (vvr1-dc1 mirror). +- `opentofu/vr1-dc0-substrate/main.tf`: 1-73 (inner provider, planes, wan-bridge, edge), + 75-265 (node fleet + utility VMs, NIC order, carved planes per role). +- `opentofu/modules/dc-planes/main.tf`, `opentofu/modules/wan-bridge/main.tf`, + `opentofu/modules/mesh-link/main.tf` -- the three bridge-shape module bodies. +- `scripts/lib-net.sh`: 22-30 (PLANE_CIDRS/PLANE_NAME dc0), 171-179 (dc1). +- `scripts/lib-hosts.sh`: 53-100 (NIC_PLANE_ORDER, BREX_PARENT_NIC, per-role carve shape). +- `scripts/geneve-encap-assert.sh`: 1-30 (the C1/C2 gate header, root-cause citation). +- `docs/design-decisions.md`: D-100 (~2223), D-101 (~2246-2330), D-124 AMENDMENTs (~5022-5083), + D-125 (~5084-5161), D-139 (~7253-7370). +- `docs/security-ledger.md`: SEC-010 row (line 21), SEC-026 row (line 79). +- `docs/CURRENT-STATE.md`: 7606-7710 (Sec 2.3-iii containment relocation note, Sec 3 authored-not- + applied inventory) -- NOTE this section is dated 2026-07-18/20 and predates the later + amphora/geneve/D-139 work; used here only for the nesting mechanism, not current apply state. + +--- + +## 6. Top risks for the administrator / W0.4 + +1. **SEC-010 co-residency gap** (Section 4) -- the single biggest wiring risk; no existing + artifact addresses cross-DC plane-bridge adjacency on one flattened host. +2. **D-125's whole bridge-in mechanism is Model-B-only debt** -- eliminating the container + layer should DELETE it (revert to Model A direct NAT), not carry it forward unused. +3. **MTU/geneve budget is NOT the source of container-elim benefit** -- do not let Phase 4's + change-set imply an MTU fix; the real payoff is removing OBS-3-class nesting fragility + and one apply-ordering stage, not a byte-budget change. +4. **Transit leg's surviving purpose is undetermined** pending the target topology (W0.4) -- + this doc intentionally stops short of ruling it out. diff --git a/docs/audit/container-elim-pass/pass0-w3-consumers.md b/docs/audit/container-elim-pass/pass0-w3-consumers.md new file mode 100644 index 0000000..17c32e1 --- /dev/null +++ b/docs/audit/container-elim-pass/pass0-w3-consumers.md @@ -0,0 +1,159 @@ +# Pass 0 / W0.3 -- Consumer/dependency inventory for the container-layer-elimination pass + +READ-ONLY planning finding. Repo HEAD at time of read: branch `dc-dc-stage5-preconditions`, +2026-08-09. This is worker W0.3 of Phase 0 (`docs/audit/container-elim-pass/SCOPE-AND- +EXECUTION-PLAN.md`, `phase-prompts.md`). No command was executed; every row cites path:line. + +## 0. Framing recap (do not re-derive) + +Model B (D-122/D-123): `vcloud (outer libvirt) -> vvr1-dcN (containment VM = inner libvirt, +ALSO the DC rack: MAAS rack/region + Juju/openstack client host) -> node VMs`. Eliminating the +containment VM collapses this to `vcloud (libvirt) -> node VMs` directly. `vvr1-dcN` (the +containment VM) is NOT the LXD API-charm containers that run on deployed OpenStack nodes -- +those are untouched by this pass. + +**The load-bearing fact this dimension surfaces:** `vvr1-dcN` is not only a nesting boundary, +it is also the NAMED EXECUTION HOST for three separate ruled roles: (1) the D-123 nested-libvirt +node host, (2) the D-132-amendment per-DC MAAS region+rack, and (3) the D-138 cloud-facing +client host (`juju`/`openstack` run FROM `vvr1-dcN`, reached `ssh -J voffice1 jessea123@` -- `runbooks/dc-dc-phase3-maas-enlist-deploy.md:424,430`, +`runbooks/dc-dc-phase4-juju-bundle-per-dc.md:165,357,1108`). Flattening removes the host that +carries all three roles simultaneously -- each must be re-homed, and they need not all land on +the same place. + +## 1. Repo-wide grep census -- `vvr1-dc` + +Canonical tree (excludes the stale `.claude/worktrees/skill-repackage-20260727/` git-worktree +snapshot, which mirrors ~384 additional hits from an unrelated 2026-07-27 skill-repackage task +and is not part of the live repo surface): + +``` +grep -rn "vvr1-dc" --include="*.sh" --include="*.py" --include="*.md" --include="*.yaml" \ + --include="*.tf" . (from repo root, worktree copy excluded) +``` +**424 occurrences across 78 files.** + +Heaviest files (occurrence count): +| File | Count | Nature | +|---|---|---| +| `docs/design-decisions.md` | 66 | D-122/123/124/125/126/128/138 rulings, amendments | +| `opentofu/main.tf` | 34 | outer root: creates the `vvr1_dc0`/`vvr1_dc1` containment VM modules | +| `scripts/site-headend-install.sh` | 18 | `--role rack` = installs on `vvr1-dcN` | +| `docs/CURRENT-STATE.md` | 17 | living status references | +| `opentofu/variables.tf` | 16 | `vvr1_dc0`/`vvr1_dc1` module var blocks ("Model B: holds the inner libvirt pool") | +| `opentofu/vr1-dc0-substrate/main.tf` | 13 | inner root: `qemu+ssh` provider targets `vvr1-dc0` | +| `runbooks/dc-dc-phase2-tofu-dc-substrate.md` | 13 | outer/inner apply runbook | +| `docs/dc0-deploy-readiness.md` | 12 | readiness record | +| `opentofu/vr1-dc1-substrate/main.tf` | 9 | inner root, dc1 | +| `docs/dc-dc-deployment-workflow.md` | 9 | workflow doc, plane-split doctrine | +| `runbooks/dc-dc-teardown-rollback.md` | 8 | Paths A/B/C/M name `vvr1-dcN` as the teardown target | +| `opentofu/vr1-dc0-substrate/variables.tf` | 6 | | +| `docs/audit/record-inventory.md` | 6 | | +| `netbox/dc-rack-mgmt-import.py` | 5 | imports `vvr1-dcN` as a NetBox rack-mgmt device | +| `tests/dc-rack-mgmt-import/test_logic.py` | 5 | harness for the above | +| ... | | remaining 64 files carry 1-4 hits each (changelogs, security-ledger, archive) | + +`scripts/*.sh` + `scripts/*.py` combined: **26 occurrences** across +`site-headend-install.sh`, `dc-rack-net.sh`, `site-baseleg.sh`, `dc-mirror.sh`, +`dc-egress-check.sh`, `maas-profile-assert.sh`, `dc-dc-whole-host-budget.py`. +`opentofu/`: **17 files**. Most of the remainder is `docs/` (design-decisions, CURRENT-STATE, +changelogs, security-ledger, archive) and `runbooks/` -- i.e. DECISION RECORD, not live +tooling; those do not need code changes on flattening, only a superseding note (GA-R1/R2 +append-only discipline -- do not edit history). + +## 2. Consumer table + +| Consumer (file:line) | What it keys to the containment layer | What breaks / changes on flattening | Severity | +|---|---|---|---| +| `scripts/lib-hosts.sh:52,157-168,212-214,246-251` `VIRSH_POWER_ADDRESS`, `VIRSH_POWER_ADDRESS_FROM_OFFICE1`, `VIRSH_POWER_ADDRESS_FROM_DCREGION` | Both forms are `qemu+ssh://` URIs that dial the **containment VM's libvirtd** (`172.31.0.2`/`10.12.8.2` for dc0, `172.31.0.6`/`10.12.68.2` for dc1) -- this is how MAAS powers on/off node VMs, because the node VMs live INSIDE `vvr1-dcN`'s libvirt, not vcloud's. | If node VMs move to vcloud's own libvirt, the power address becomes **vcloud's own libvirt** (`qemu:///system` locally, or a `qemu+ssh` back to vcloud from wherever MAAS's region runs). The FROM_OFFICE1/FROM_DCREGION split (which exists because "the reachable rack address differs by region", `lib-hosts.sh:159-167`) may collapse entirely if there is no separate inner libvirt host to be reachable FROM. Every value in both DC arms (`vr1-dc0`/`vr1-dc1` cases, lines 212-214, 246-251) needs re-derivation, not a rename. | **HIGH** -- this is the mechanism MAAS uses to power every node; wrong here = "MAAS unreachable" masquerading as a network fault (a failure mode this repo has already hit once, `maas-role-tags.sh:48-51`). | +| `scripts/lib-hosts.sh:69-72` `NIC_PLANE_ORDER`, `BREX_PARENT_NIC` | Describes the NIC-to-plane mapping the INNER tofu root pins per node (`main.tf macs[0..5]` order) -- a property of how `opentofu/vr1-dcN-substrate/main.tf` (the inner root, applied via `qemu+ssh` to `vvr1-dcN`) attaches MACs to node VM NICs. | If nodes are created directly by the (now single) tofu root against vcloud libvirt, the NIC/MAC-to-plane pinning logic must be re-authored in that root, but the ORDER CONVENTION itself (metal-admin first, not `PLANE_CIDRS` order) is a fact about how MAAS/libvirt XML enumerates interfaces, independent of nesting -- it likely CARRIES FORWARD unchanged. Confirm at Phase 2 (tooling), not a Phase-0 blocker. | **LOW-MEDIUM** -- convention likely survives; the FILE that encodes it (inner substrate root) is what changes. | +| `scripts/lib-hosts.sh:29-34` `CARVE_AUX_HOSTS` | Declares the aux (non-Juju) `-tailscale-01`/`-maas-01` VMs as `--host`-only carve targets; these ALSO live inside `vvr1-dcN`'s inner libvirt today. | No conceptual change on flattening -- these become ordinary flat-libvirt VMs alongside the 9 role nodes + juju controller. Their power address inherits the same `VIRSH_POWER_ADDRESS` fix above. | **LOW** -- mechanical, rides the same fix as the row above. | +| `scripts/lib-net.sh` (whole file) | **No `vvr1-dc` string appears here** (verified: `grep vvr1-dc scripts/lib-net.sh` = 0 hits). `PLANE_CIDRS`/`PLANE_GW`/VIP prefixes are IP-plan facts, not containment-topology facts. | UNCHANGED by container elimination as such. It WILL change under the concurrent D-143 re-IP (10.12->10.13), but that is a SEPARATE, already-ruled axis this pass must keep distinguishable (SCOPE-AND-EXECUTION-PLAN.md Section 7 says exactly this). | **NONE** for this dimension -- flag as a non-consumer so Phase 4's change-set does not double-count it. | +| `scripts/maas-node-power.sh:2-46` (usage, `POWER_ADDRESS` arg, `VIRSH_URI`) | Takes a `qemu+ssh://` power address as an **argument** (not hardcoded) and virsh-lists domains AT THAT URI to MAC-match them to MAAS machines. Every worked example in the header (`qemu+ssh://jessea123@172.31.0.2/system vr1-dc0`) targets the containment VM. | The script itself is topology-agnostic (it takes the URI as input) -- it needs NO code change. But every CALLER that supplies `172.31.0.2`/`10.12.8.2`-class addresses must supply the new (flat) address instead, which is exactly the `lib-hosts.sh` fix above. Also note: the header's D-103 rationale for per-machine (not pod) power was ALREADY about avoiding MAAS-virsh-pod storage incompatibilities (`domblkinfo` on volume-ref disks) -- that reasoning is unrelated to nesting and stays valid post-flatten. | **MEDIUM** -- no code change, but every invocation site needs the new address; a stale hardcoded address in a runbook/changelog would silently power the WRONG (or a nonexistent) host. | +| `scripts/site-headend-install.sh:6-20,82-129,204-343,371-409` `--role rack`, `--host-nodes`, `node_host_check()`, `node_host_setup()` | This is **THE script that turns `vvr1-dcN` into a nested-libvirt node host**: installs qemu-kvm/libvirtd, turns on nested KVM, creates the inner pool dir, the SEC-010 FORWARD-drop keyed on `$TRANSIT_IF` ("mgmt"), and verifies the D-125 WAN bridge/uplink. `--role rack` (no `--host-nodes`) is the narrower "just a MAAS rack enrolling to Office1's region" case, ALSO run on `vvr1-dcN` today. | If nodes run directly on vcloud, `--host-nodes` (node_host_setup/node_host_check, ~140 lines) becomes **DEAD CODE for VR1** -- there is no separate "rack that also hosts nested KVM" anymore; vcloud itself already IS the outer libvirt host and does not need this bootstrap. The **rack role itself** (MAAS enrollment) may still be needed if D-132's "MAAS region controller per DC" is KEPT as a normal flat VM (see D-132 row below) -- but it would no longer carry `--host-nodes`. The SEC-010 FORWARD-drop (`TRANSIT_IF="mgmt"`) is keyed to a transit LEG that exists BECAUSE of the containment VM's two-leg (transit + inner-plane) design; a flat topology may not have a "transit" interface to drop-FORWARD on at all, or the DC-LOCAL boundary the drop enforces may need to move to a different device (a vcloud-host-level nftables rule, or a per-VM firewall) -- this is a design question for Phase 0's target-topology worker (W0.4), flagged here as a dependency. | **HIGH** -- largest single file by consumer surface (18 `vvr1-dc` hits, ~140 lines of `--host-nodes` logic); also the SEC-010 enforcement point (a security control, not just wiring) lives here. | +| `scripts/dc-rack-net.sh:2-14,55-85` (whole script; site table for `dc0`/`dc1`) | RUNS ON THE DC RACK HOST (i.e. `vvr1-dcN`) per its own header; persists rack bridge-leg IPs (`vr1-dc0-metal-admin=10.12.8.2/22` etc.) and the D-131 node-facing DNS forwarder, both keyed to addresses the CONTAINMENT VM itself holds on its own inner bridges. | If there is no containment VM, "the rack's own inner bridges" do not exist in the same shape -- these addresses (`.2`, `.3` on metal-admin) were the containment VM's identity on its OWN nested networks. Under a flat topology this functionality (rack-leg persistence + node DNS forwarder) needs a new HOME: either the per-DC MAAS region VM (if D-132's region-per-DC survives as a flat VM) or is retired if the forwarder's reason (rack-only resolver SERVFAILs, `dc-rack-net.sh:27-30`) no longer applies once "the rack" is gone. Cannot be resolved without W0.4's target topology. | **HIGH** -- a security/DNS-availability control (D-131), not cosmetic; silently losing it would reintroduce the SERVFAIL bug this script was built to fix. | +| `scripts/site-baseleg.sh:40-48` `LEGS["office1"]`, the commented-out `[vr1-dc0]` row | The `office1` row is UNRELATED to the containment VM (it's the vcloud<->voffice1 base leg). The **DC rows are explicitly DEFERRED** with the comment "the DCs nest inside vvr1-dc0 and are reached by qemu+ssh ... add a row ONLY if/when a DC needs one" -- and D-138 (`design-decisions.md:7126-7131`) already closed this deferral: "the DC rows stay DEFERRED and their MEASURE-first note is now ANSWERED: ... no host-side leg is wanted." | If nodes are flat on vcloud, the "DC rows deferred because reached by qemu+ssh through vvr1-dc0" premise disappears -- vcloud already has an L3 view of its own libvirt guests without an SSH hop. Whether a NEW base-leg row is needed depends entirely on whether the flat node VMs get bridged IPs vcloud itself must route to, which again is W0.4's call. | **LOW** -- currently an explicit no-op; only becomes live work if the target topology needs a new leg. | +| `scripts/maas-role-tags.sh:37,48-51` | The comment "Run this where maas lives (the D-128 Plane-2 headend)" and the `REFUSE` message both assume the operator KNOWS which host that is -- today, resolvable to `voffice1` (region) OR `vvr1-dcN` (rack), a distinction this pass changes. No `vvr1-dc` LITERAL string appears in this script's logic (it is genuinely site/host-agnostic, driven by `lib-hosts.sh`/`MAAS_PROFILE`). | No code change required. Purely a DOC/comment currency issue once the execution-host question (Section 3 below) is answered -- the "headend" language should be corrected to match whatever host actually runs `maas` post-flatten. | **LOW** -- comment-only. | +| `opentofu/main.tf:25-33,187,234,309-422` (outer root, `module "vvr1_dc0"`/`"vvr1_dc1"`) | This IS the resource that creates the containment VM (`vm_name = "vvr1-dc0"`, line 412) -- the single biggest and most literal consumer. Line 405: "Site-down = a single `virsh destroy vvr1-dc0` (the D-122 intent)" -- i.e. the outer root's entire per-DC unit of work today IS this one VM. | The ENTIRE outer-root module for a DC either (a) is deleted and replaced by direct node-VM resources targeting vcloud libvirt (the flattened design), or (b) is repurposed to create something smaller (e.g. just a MAAS-region VM, if that role is kept). This is squarely Phase-0-W0.1's (substrate/tofu map) territory, cross-referenced here because it is also the top literal-reference file in the census. | **HIGH** (cross-referenced to W0.1, not re-owned here). | +| `opentofu/vr1-dc0-substrate/main.tf`, `vr1-dc1-substrate/main.tf` (inner roots) | Provider block targets `qemu+ssh` INTO `vvr1-dc0`/`vvr1-dc1` (per `main.tf:25-33` comment on the outer root, and the inner root's own `versions.tf`/provider config) -- this is THE apply that creates the 9+1(+2 aux) node VMs, currently nested. | If flattened, there is no longer an "inner root run via qemu+ssh into the containment VM" -- either this root's provider target becomes `qemu:///system` on vcloud directly (merging outer+inner into one tofu apply), or it becomes `qemu+ssh` FROM wherever the new execution host is, straight to vcloud. Either way this is a provider-config + apply-ordering change, not a pure rename. Cross-referenced to W0.1. | **HIGH** (cross-referenced to W0.1). | +| `netbox/dc-rack-mgmt-import.py`, `tests/dc-rack-mgmt-import/*` | Imports `vvr1-dcN` into NetBox as a **rack-management device record** (a physical/virtual asset in the DCIM model representing the containment VM as "the rack"). | If there is no containment VM, this NetBox device either needs to be retired (with a NetBox-side decommission, not silent deletion -- NetBox is the IPAM/DCIM system of record) or repointed to represent something else (e.g., vcloud itself, or nothing, if "the rack" concept is retired outright). This is a DATA-migration consumer, distinct from a code consumer -- flag for Phase 2/W2.4 (which tools become which modules) since it touches an external system of record, not just repo files. | **MEDIUM** -- external-system consumer, easy to miss because it is not a runbook or script that gets "read" in the normal grep sweep for tooling. | +| `scripts/dc-egress-check.sh:63,72` | Comments citing MEASURED `ip route` output taken directly ON `vvr1-dc0`/`vvr1-dc1` (the default-route-via-provider-public fact). Historical/measurement citations, not live logic keyed to the containment layer. | No functional change -- the DEFAULT ROUTE FACT itself (default via `.4.1`/`.64.1`) is a plane-gateway fact (`lib-net.sh PLANE_GW`), not a containment fact, and would still hold once nodes are flat as long as the same plane/gateway design is kept. Re-verify empirically post-flatten rather than assume; the citation just needs a currency note. | **LOW**. | +| `scripts/dc-mirror.sh:6` | Doc-comment: "RUNS ON THE DC RACK HOST (e.g. vvr1-dc0), not on vcloud/voffice1." | Same class as `dc-rack-net.sh` above but for the apt mirror service -- if "the rack host" no longer exists as a distinct entity, this script's stated execution host must be re-pointed to wherever the mirror service actually runs post-flatten (likely still a dedicated per-DC utility VM at the `.4` octet, per `lib-hosts.sh:96-100` `REGION_HOST_SUFFIX`/utility-band convention -- that utility VM is unaffected by containment removal, it just stops being INSIDE a nested host). | **LOW-MEDIUM** -- doc-currency + confirm the utility VM's own libvirt parent changes from "inner" to "vcloud direct", nothing else. | +| `scripts/maas-profile-assert.sh:39` | Usage example lists `voffice1,vvr1-dc0,vvr1-dc1` as the three MAAS-profile-bearing hosts whose region identity this script proves ("a machine count is NOT proof"). | If the per-DC MAAS region (D-132 amendment) no longer lives ON the containment VM (because there is no containment VM) but instead on a flat per-DC utility VM, the usage example needs updating but the SCRIPT'S JOB (prove which region a profile resolves to, by rack identity) is unchanged and arguably MORE necessary post-flatten (more hosts look alike without the containment boundary as a visual/structural cue). | **LOW** -- doc-currency; the script's function is robust to the topology change. | +| **D-128/D-138 execution-host assumption** (see Section 3, its own severity call-out below) | `vvr1-dcN` is the D-138-RULED host for the Juju client + `openstack` CLI, and (per D-132 amendment, `design-decisions.md:7161-7168`) the per-DC MAAS region controller. Both rulings say "in the DC" / "on the rack" and CITE `vvr1-dc0` transit-IP SSH as the concrete mechanism (`runbooks/dc-dc-phase3-maas-enlist-deploy.md:424,430`; `phase4-juju-bundle-per-dc.md:165`). | **THE OWED QUESTION THIS PASS MUST ANSWER, NOT ASSUME:** with no containment VM, where do `juju`/`openstack`/the MAAS region run from? See Section 3. | **CRITICAL / BLOCKING** -- this is not a mechanical rename; it is an unresolved architecture question that D-128/D-132/D-138 do not answer on their own once their named host disappears. | +| `runbooks/dc-dc-teardown-rollback.md` (8 hits; Paths A/B/C/M) | Teardown paths name `vvr1-dcN` as the target of `tofu destroy` / `virsh destroy` / juju-controller-down recovery. | Once the containment VM does not exist, these paths must be re-authored around whatever the new substrate shape is (destroying flat node VMs individually or via a module, not "destroy one containment VM = destroy the site"). The D-122 intent quoted in `opentofu/main.tf:405` ("site-down = a single `virsh destroy vvr1-dc0`") is a SIMPLICITY PROPERTY that flattening explicitly TRADES AWAY -- teardown of N+1 individual VMs is inherently more steps than one. Flag as a design tradeoff for W0.4/W4.2 (module design should re-earn this simplicity via a tofu-module-scoped destroy, not lose it). | **MEDIUM-HIGH** -- teardown is a frequently-exercised, high-consequence path (see D-061 register / `docs/tool-index.md`'s own origin story) and losing "one command tears down a DC" without a replacement is a real regression. | + +## 3. The D-128/D-138 execution-host consequence -- flagged explicitly, per instructions + +**Today:** `vvr1-dcN` simultaneously serves three ruled roles: +1. **D-123 Model B node host** -- nested libvirt for the 9+1(+2) node VMs. +2. **D-132-amendment MAAS region+rack for its own DC** -- installed "in the DC" so `jujud`'s + continuous MAAS-API dependency never crosses the Office1 fiber (SEC-010 preserved + unamended by removing the cross-fiber requirement rather than exempting it). +3. **D-138 cloud-facing CLIENT host** -- `juju`/`openstack` run FROM `vvr1-dcN`, reached via + `ssh -J voffice1 jessea123@`, because SEC-010 + D-052 make the node + planes unreachable at L3 from `voffice1`, and D-138's own reasoning is explicit: "a MAAS + rack proxies at the application layer and needs no kernel forwarding... that is TRUE of + every Plane-2 tool then in use and FALSE of Juju, which dials the MACHINE at L3." + +**If `vvr1-dcN` is eliminated, all three roles lose their named host simultaneously,** and +they do NOT have to be re-homed together -- each has a different constraint: +- Role 1 (node host) simply DISAPPEARS as a distinct thing -- vcloud's own libvirt takes over, + no replacement host is needed. +- Role 2 (MAAS region) needs SOME host inside "the DC's" network segment (wherever that + segment's boundary now is) so `jujud`'s L3 MAAS dependency stays local -- this could be a + flat per-DC utility VM (the existing `.6`/`REGION_HOST_SUFFIX` convention already assumes a + VM distinct from the role nodes, so it may not even need a new design, just a new libvirt + parent). +- Role 3 (Juju/openstack CLIENT) needs a host that can reach the flat node VMs at L3 for the + full port set D-138 enumerated (22, 17070, 5000, 9292, 8774, 9696, 8776, 8778, 9876, 9311, + 9511, 443+ -- `design-decisions.md:7107`). If node VMs are flat on vcloud and reachable from + vcloud's own network without SEC-010's transit-drop boundary in the way, **the client could + plausibly run FROM VCLOUD ITSELF** (collapsing D-128's Plane-1/Plane-2 split for this + specific tool-class) -- but this is only true if the flattened topology does NOT recreate an + equivalent transit/forwarding boundary between vcloud and the flat node VMs. If it DOES + (e.g., to keep simulating the Roosevelt fiber-boundary realism the whole VR1 exercise exists + to rehearse), then a replacement "DC-local client host" is still needed, and D-138's + reasoning (why the client cannot live at the far side of that boundary) still applies + verbatim to whatever plays the transit-boundary role post-flatten. + +**This is the single largest open question this dimension surfaces, and W0.3 cannot close it +alone** -- it is jointly a W0.2 (network/wiring: does a transit/SEC-010-equivalent boundary +still exist post-flatten?) and W0.4 (target-topology: what replaces `vvr1-dcN` as a named +host) question. Recorded here as the explicit hand-off the SCOPE doc's Section 4 instructs: +**"if there is no containment VM, where do juju/openstack run from?" has NO answer in the +current repo** -- D-128/D-132/D-138 all name `vvr1-dcN` as the concrete host, and none of them +anticipates its removal. Whatever Phase 0 / W0.4 proposes as the target topology MUST include +an explicit answer to this, stated as plainly as D-138's own ruling was stated, because it +carries the same SEC-010/D-052 security-boundary weight D-138 was ruled to resolve -- getting +it wrong reopens exactly the "narrow SEC-010 exception vs. re-plan the control-plane topology" +choice D-132's amendment closed. + +## 4. Non-consumers worth naming (to prevent Phase 4 double-counting) + +- `scripts/lib-net.sh` -- no containment-layer coupling; only D-143 (re-IP) touches it. Keep + the two axes distinguishable per SCOPE-AND-EXECUTION-PLAN.md Section 7. +- `scripts/maas-role-tags.sh`, `scripts/dc-egress-check.sh`, `scripts/maas-profile-assert.sh` + logic (not comments) -- host-agnostic by construction (driven by `lib-hosts.sh`/env), so they + need NO code change, only doc-comment currency once Section 3's question is answered. +- The 66 `design-decisions.md` hits and ~17 `CURRENT-STATE.md` hits are the DECISION RECORD -- + per GA-R1/append-only discipline these are not edited; the container-elim [ARCH] decision + (owed, per SCOPE doc Section 7) will AMEND or SUPERSEDE the relevant entries (D-122/123 most + directly), not rewrite them in place. Not a "consumer" in the code sense; flagged so Phase 4 + does not propose editing history. + +## 5. Summary counts for the bounded return + +- **Distinct functional consumers identified (code/config, excluding decision-record docs and + the stale worktree copy): 15** (table rows in Section 2, excluding the two doc-only rows for + teardown-runbook and netbox-import which are procedural/data consumers, counted separately). +- **Highest severity: CRITICAL/BLOCKING** -- the D-128/D-138 execution-host question + (Section 3): no repo surface currently answers "where do juju/openstack run from with no + containment VM", and D-132/D-138's own security reasoning (SEC-010, the app-layer-proxy vs. + L3-dial distinction) applies to whatever replaces `vvr1-dcN` in that role. +- **HIGH severity (concrete, non-blocking but large):** `scripts/lib-hosts.sh` power-address + maps (every DC arm), `scripts/site-headend-install.sh` `--host-nodes` machinery (~140 lines + becomes dead code for VR1), `scripts/dc-rack-net.sh` (rack-leg + D-131 DNS forwarder needs a + new home), the outer/inner OpenTofu roots (cross-referenced to W0.1), and + `runbooks/dc-dc-teardown-rollback.md` (loses the "one `virsh destroy`" simplicity D-122 + intended, needs a re-earned equivalent). diff --git a/docs/audit/container-elim-pass/pass0-w4-targets.md b/docs/audit/container-elim-pass/pass0-w4-targets.md new file mode 100644 index 0000000..9f72eb5 --- /dev/null +++ b/docs/audit/container-elim-pass/pass0-w4-targets.md @@ -0,0 +1,282 @@ +# Pass 0 / W0.4 -- Target-topology options (container-layer elimination) + +**Worker:** W0.4 (target-topology options). **Feeds:** the Phase-0 OPERATOR GATE +(`SCOPE-AND-EXECUTION-PLAN.md` Section 4). **Scope discipline:** READ-ONLY; every claim below +cites `path:line` or a named durable record; anything not resolvable from the repo is marked +**OWED** rather than invented (hard rule 2). + +--- + +## 0. Framing recap (verified against the repo, not assumed) + +- **Current nesting (Model B, D-123):** `vcloud (outer libvirt, qemu:///system) -> vvr1-dcN + (containment VM, inner libvirtd) -> node VMs`. The outer root creates `vvr1-dc0`/`vvr1-dc1` + (`opentofu/main.tf:410-519,537-623`); the INNER root (`opentofu/vr1-dc0-substrate/main.tf:1-266`) + creates the 6 planes + WAN bridge + edge + **12** node VMs (9 D-121 role nodes + the D-104 + juju-controller `vr1-dc0-juju-01` + the D-132 MAAS-region `vr1-dc0-maas-01` + the D-129(iii) + Tailscale router `vr1-dc0-tailscale-01`, `opentofu/vr1-dc0-substrate/main.tf:96-247`) via a + `qemu+ssh` provider dialled FROM Office1 (`opentofu/vr1-dc0-substrate/main.tf:12-22`). +- **The eliminate-and-pull-up-one target, stated in the operator's own words** (quoted in + `SCOPE-AND-EXECUTION-PLAN.md:34-36`): collapse to `vcloud (libvirt) -> node VMs` directly. +- **This is not a green-field question.** Before Model B was ruled, the exact flat shape existed, + was committed, and validated: `docs/archive/model-a-fallback-plan.md` ("Model A", git tag + `model-a-fallback` at `114d392`, R-3-compliant, `tofu validate` 11/11 modules). Both options + below are read against that precedent rather than invented from scratch. +- **Two rulings POST-DATE Model A and change what a flat topology must additionally provide** -- + Model A's own spec did not have to solve them: + - **D-138** (`docs/design-decisions.md:7064-7126`, RULED 2026-07-30): the cloud-facing Juju + client + `openstack` CLI (phase-03..phase-06) MUST run from a host with L3 reach to the DC's + node planes (`voffice1` cannot reach them; SEC-010/D-052/D-125 forbid routing there). CURRENT + STATE: that host is `vvr1-dc0` itself -- "the dc0 rack" (`docs/CURRENT-STATE.md:1046,3014-3016`; + `lib-hosts.sh:212-213` shows `vvr1-dc0` holds a live metal-admin address, + `qemu+ssh://jessea123@10.12.8.2/system`, the D-134 rack-utility `.2` -- distinct from + `.5` juju-01, `.6` maas-01, `.7` tailscale-01). + - **D-132 amendment** (`docs/design-decisions.md:7135-7204`, RULED 2026-07-30): each DC gets its + OWN MAAS region (now a node VM, `vr1-dc0-maas-01`), not a migrated/shared Office1 region. + - **SEC-026/SEC-028** (`docs/security-ledger.md:79,81`): "the rack" is now a **credential-bearing + host** -- it holds the DC-scoped MAAS admin key (SEC-018) and the per-DC Juju service key + (SEC-028), OPEN rotation obligations, and an explicit per-DC isolation requirement ("each DC's + client host receives ONLY THAT DC's credential"). + - **Consequence for this pass:** eliminating the container layer must still answer "where does + the D-138 client + its SEC-026/SEC-028 credential live", because that requirement did not exist + when Model A was archived. This is the crux item (d) below. + +--- + +## 1. OPTION 1 (RECOMMENDED) -- "Model A, D-132/D-138-updated": flat nodes + a small per-DC client VM + +``` +vcloud (host, L0, qemu:///system) + |-- vr1-dc0-client small VM (4 vCPU / 8 GiB / 80 GiB, D-138 role only -- see (d)) + | legs: metal-admin + office1<->dc0 transit; NO nested libvirt, + | NOT a hypervisor for anything -- a flat sibling VM like any other + |-- vr1-dc0-control-01..03 node VMs (16/65536/150) \ + |-- vr1-dc0-compute-01..02 node VMs (12/49152/100) \ vcloud-level libvirt siblings, + |-- vr1-dc0-storage-01..04 node VMs (8/24576/550) > attached DIRECTLY to the 6 + |-- vr1-dc0-juju-01 node VM (4/8192/100) / vcloud-level plane networks + |-- vr1-dc0-maas-01 node VM (4/8192/150) / + |-- vr1-dc0-tailscale-01 node VM (2/2048/25) / + |-- vr1-dc0-opnsense edge (2/2048, 2-NIC: provider-public LAN + vr1-dc0-wan WAN) + |-- vr1-dc0-{provider-public,metal-admin,metal-internal,data-tenant,storage,replication} + | 6 isolated-L2 libvirt networks (dc-planes), AT VCLOUD LEVEL + |-- vr1-dc0-wan NAT /24 simulated ISP uplink, AT VCLOUD LEVEL (site-wan, direct) + `-- mesh-vr1-dc0-office1 transit leg (office1 <-> dc0), unchanged + +nesting depth = 2 (vcloud -> node VM -> nova KVM guest) <- VR0-PROVEN, same as Model A +site-down = scripted group-destroy of the vr1-dc0-* domain set (Model A's mechanism) +``` + +### (a) Where node VMs live +Directly on vcloud's outer libvirt provider (`qemu:///system`), as siblings -- identical to +`docs/archive/model-a-fallback-plan.md:24-39`'s Model A shape, but with the 3 utility node +classes (juju-controller, MAAS region, Tailscale router) added as additional flat siblings using +the SAME `modules/node-vm` bodies already authored in +`opentofu/vr1-dc0-substrate/main.tf:250-265` (only the provider target changes, per that file's +own header note at lines 1-10: "All module bodies are UNCHANGED from the outer root -- only the +provider they run against moved"). + +### (b) Where the six planes land +Back at vcloud level as `modules/dc-planes` outputs, reversing the D-123 move documented at +`opentofu/main.tf:22-33` ("the 6 vr1-dc0 planes MOVED to the INNER root ... under Model B"). +`lib-net.sh`'s `PLANE_CIDRS`/`PLANE_NAME` (unchanged, `scripts/lib-net.sh:21-30`) and the D-101/ +D-134 CIDR values stay IDENTICAL -- only the libvirt network objects' host moves, not their IPAM +identity. metal-admin remains the PXE/boot plane, `NIC_PLANE_ORDER` unchanged (`lib-hosts.sh:69`). + +### (c) Where transit/uplink/mesh/br-ex land +- Mesh legs (office1<->dc0, office1<->dc1, dc0<->dc1): UNCHANGED, already vcloud-level + (`opentofu/main.tf:125-141`). +- Uplink/WAN: the D-125 bridge-in plumbing (`module "vr1_dc0_uplink"` NAT + `wan-bridge` module + + the IP-less 2nd NIC + `br-vr1-dc0-wan` netplan bridge, `opentofu/main.tf:360-397,505-518`) + is REMOVED. The DC edge's WAN attaches to a vcloud-level NAT directly, exactly like Office1's + edge does today (`opentofu/main.tf:99-117`) -- this is Model A's item 8 + (`docs/archive/model-a-fallback-plan.md:67-75`): "same `172.30.2.0/24` ... no uplink NIC, no + bridge, no `wan-bridge` module ... NO re-address needed". +- br-ex: unchanged in kind -- still an OVS bridge parented on each ROLE node's provider-public + NIC (`lib-hosts.sh:70-72`), a per-node fact independent of where the node's hypervisor sits. + +### (d) Where the juju/openstack execution host goes (THE CRUX) +A **dedicated, small, per-DC client VM** on vcloud libvirt, occupying the ROLE Model A's original +rack headend already had by construction: two legs (metal-admin + office1 transit, +`docs/archive/model-a-fallback-plan.md:25-26`), 4 vCPU / 8 GiB / 80 GiB. Model A's own headend was +authored BEFORE D-138/D-132 existed and was scoped as "MAAS rack headend" only -- this option +REPURPOSES that same artifact shape to also satisfy D-138 (cloud-facing client execution) and +SEC-026/SEC-028 (the DC-scoped credential residency), because it already has the one property +those rulings require: an L3 leg on metal-admin. It is explicitly **not a hypervisor for anything** +-- no `expose_nested_virt`, no inner libvirtd, no nodes attached to it -- so it does not +reintroduce a container layer; it is a flat utility sibling, same class as `vr1-dc0-juju-01`. +Per-DC credential isolation (SEC-026 "(1) ISOLATION IS THE LOAD-BEARING CONTROL") is preserved +exactly as it is today, since this VM plays the identical role the current rack plays. +maas-vm-host registration (Step 9, deferred DOCFIX-179) targets **vcloud's own virsh**, not this +VM's (Model A table row 4, `docs/archive/model-a-fallback-plan.md:48`) -- MAAS discovers the +vcloud-level node domains directly. + +### (e) Capacity/FIT impact on vcloud +Directionally FREES host RAM, does not need it. Today's containment VM is sized to hold BOTH the +node fleet AND its own containment overhead (`opentofu/variables.tf:143-150`: RAM raised +416->480 GiB per DC, "384 GiB node fleet + 96 GiB overhead"; the 96 GiB is inner-libvirtd/OS/page- +cache overhead that exists ONLY because of the nesting). Flat placement removes that per-DC 96 GiB +overhead layer and replaces it with the small client VM's 8 GiB -- a swing of roughly (96-8) x 2 +DCs ~ 176 GiB of host RAM, all else equal. vCPU impact is smaller: Model B's overhead is 4 vCPU/DC +(`opentofu/variables.tf:137-141`, unchanged since the RAM-only raise), and Model A's headend is +also 4 vCPU, so vCPU is roughly a wash. **This is a qualitative, repo-grounded direction, not a +verified number** -- see Section 4 OWED items: the sizing calculator +(`scripts/dc-dc-whole-host-budget.py`) does not yet have flags for the 3 utility node classes +(juju/maas/tailscale) added after its authoring, so an exact FIT verdict needs either extending +it or hand-totaling before the operator relies on a number. + +### (f) What breaks +- The inner OpenTofu root (`opentofu/vr1-dc0-substrate/`, its own state file, its `qemu+ssh` + provider dialled from Office1) is retired; its module bodies fold back into the outer root. +- The "OUTER -> BOOTSTRAP GATE -> INNER" apply-ordering contract (`opentofu/main.tf:301-331`) + goes away -- no bootstrap gate, no cross-root apply sequencing. +- `site-headend-install.sh`'s "node-host mode" (nested libvirtd, `kvm nested=1`, inner pool, + the SEC-010 nftables writer scoped to the inner host, D-125 `--uplink-if`/`--wan-bridge` + verification) is retired; only whatever the small client VM still needs (metal-admin address + assignment, the D-138 client role) remains, in a much smaller mode. +- **SEC-010's CLOSED implementation must be REBUILT, not moved.** Today's mechanism + (`docs/security-ledger.md:21`) is `nftables-sec010.nft` + `sec010-fw.service` written onto + `vvr1-dc0`'s MEASURED interface names (`enp1s0` transit / `br-vr1-dc0-wan`/`enp2s0` uplink, + captured live 2026-07-20). The new client VM is a DIFFERENT artifact with its own NIC-naming + trap (`opentofu/main.tf:473-481` documents this exact trap for `vvr1-dc0`'s q35 shape) -- + its transit-scoped FORWARD-drop has to be re-authored and re-measured, not copy-pasted. +- `VIRSH_POWER_ADDRESS_FROM_OFFICE1`/`_FROM_DCREGION` (`lib-hosts.sh:212-213,246-250`) currently + point at the rack's transit/metal-admin address for `qemu+ssh` NODE POWER CONTROL through the + inner libvirt. With nodes flat on vcloud, power control becomes vcloud's own local virsh + (`qemu:///system` or a local-equivalent address) -- these constants and every consumer that + reads them (`maas-node-power.sh`, the teardown runbook's virsh probes) need retargeting. +- **The D-123 Model B ruling's headline benefit is lost:** "site-down = a single `virsh destroy + vvr1-dc0`" (`opentofu/main.tf:405`) reverts to Model A's scripted group-destroy of the + `vr1-dc0-*` domain set (`docs/archive/model-a-fallback-plan.md:36,64`). This is a REAL + regression against what D-123 was ruled for, not a free simplification -- flagged here so + Phase 4's [ARCH] decision framing carries it honestly, per + `SCOPE-AND-EXECUTION-PLAN.md:197-200`. +- MAC-pin discipline for the 12 node VMs (`opentofu/vr1-dc0-substrate/main.tf:87-245`) needs a + fresh capture pass once nodes are re-created flat -- moot for the 10.13 redeploy itself since + the whole fleet is being torn down and rebuilt regardless. + +### (g) What simplifies +- Nesting depth 4 -> 2 (VR0-proven), eliminating the entire "unproven depth-4 nested-virt" risk + class that Model A's own fallback-trigger list names (`docs/archive/model-a-fallback-plan.md:97, + 100`: "nova-compute guests fail to boot or are unusably slow at 3x-nested KVM"; "expose_nested_ + virt=true ... destabilises the headend"). +- No inner-root/outer-root apply-ordering gate, no separate inner state file, no cross-host + `qemu+ssh`-from-Office1 provider dial for the substrate apply. +- No D-125 bridge-in plumbing at all (`wan-bridge` module, the IP-less uplink NIC, the + `br-vr1-dc0-wan` bridge, the "unprovable pre-apply" deploy-time gate at + `opentofu/main.tf:375-378`) -- the DC edge's WAN becomes an ordinary direct NAT, one fewer + bespoke module and one fewer "unprovable until you try it" gate. +- Fewer NIC-naming traps to carry (today TWO artifacts -- `vvr1-dc0` and `voffice1` -- each + independently measured `enp1s0`/`enp2s0` on first boot; flat placement needs this measured + once, for the single small client VM, not for a containment VM AND its inner substrate). +- `lib-net.sh` plane CIDRs/names are untouched (still 6 planes/DC, same D-101/D-134 values) -- + only the low-level libvirt/carve/power scripts (`lib-hosts.sh`'s power-address constants, + `maas-node-power.sh`, `dc-node-carve`) need retargeting, not the IPAM layer. + +### (h) Roosevelt-delta +Strongly POSITIVE. Bare metal has no hypervisor containment layer at all -- Roosevelt's node +"VMs" are literal servers on the DC's physical planes, so flat placement is the closer bare-metal +analog and nesting was always a simulation-only artifact (Model A's fallback doc never claimed +Model B carried Roosevelt fidelity; its one stated advantage was the single-object site-down +convenience, which is a simulation-only DR primitive, not a transferable property). The retained +small client VM in (d) DOES transfer: D-138 names its own Roosevelt analog explicitly -- +"each DC has a management entry point inside it; the NOC reaches that entry point ... the per-DC +management bastion holds only its own DC's cloud credential" (`docs/design-decisions.md:7121- +7124`). So this option's client VM is not throwaway scaffolding; it is that bastion, rehearsed +early, minimizing delta-to-Roosevelt exactly as `SCOPE-AND-EXECUTION-PLAN.md:203-204` asks. + +--- + +## 2. OPTION 2 (NOT RECOMMENDED, given for contrast) -- fully flat: client runs on vcloud itself + +``` +vcloud (host, L0) <-- juju/openstack CLI + the D-138/SEC-026/SEC-028 credential live HERE, + | on the jumphost OS itself, via a new host-side NIC/veth onto the + | (now vcloud-level) metal-admin plane bridge + |-- vr1-dc0-control-01..03 / compute-01..02 / storage-01..04 / juju-01 / maas-01 / tailscale-01 + | (same flat node-VM siblings as Option 1 -- (a)/(b)/(c) IDENTICAL to Option 1) + `-- ... (planes/uplink/mesh identical to Option 1; NO small client VM) +``` + +Everything in (a)/(b)/(c)/(f)/(g) is IDENTICAL to Option 1 -- the only difference is (d): no +per-DC client VM at all; the D-138 execution host collapses one layer further, onto vcloud's own +OS. + +### (d) Execution host -- the crux, and why this is worse +The juju/openstack CLI, and the SEC-026/SEC-028 per-DC credentials, would live directly on the +shared jumphost that ALSO runs Plane-1 substrate tooling for BOTH DCs and Office1 +(`docs/design-decisions.md:5354-5358`, D-128 Plane 1). This inverts SEC-026's own stated control: +"(1) ISOLATION IS THE LOAD-BEARING CONTROL: each DC's client host receives ONLY THAT DC's +credential" (`docs/security-ledger.md:79`) -- on vcloud, BOTH DCs' credentials would sit on the +one host that already has the widest blast radius in the whole project (it is the CLAUDE.md +"live operations clone", governed by the PreToolUse guard and every permission `ask` rule). A +credential-scoping boundary that today survives a "rebuild the rack VM" event would no longer +exist to survive anything short of rebuilding vcloud itself. + +### (e) Capacity/FIT +Same node-fleet math as Option 1 (nodes are identical either way); this option additionally saves +the ~8 GiB the small client VM would have cost. Negligible vs. Option 1's ~176 GiB swing. + +### (h) Roosevelt-delta -- negative +No bare-metal analog exists for "run the DC client from the shared jumphost" -- Roosevelt's own +management model is per-DC bastions (D-138's own analog, quoted above), which is what Option 1's +small client VM already rehearses. Option 2 would need to be UNDONE and re-built as something +like Option 1's shape at the pre-Roosevelt bare-metal test, making it throwaway work rather than +transferable. + +--- + +## 3. Tradeoff table + +| Dimension | Option 1 (small per-DC client VM) | Option 2 (client on vcloud) | +|---|---|---| +| Nesting depth | 2 (same as Option 1's node placement) | 2 | +| D-138 execution host | dedicated, isolated, per-DC | shared jumphost, cross-DC | +| SEC-026/SEC-028 isolation | preserved (matches today's model) | WEAKENED (both DCs' admin-scoped MAAS keys land on one shared host) | +| Host RAM freed vs. today | ~176 GiB (2 DCs, qualitative -- see OWED) | ~184 GiB (marginally more) | +| Roosevelt-delta | POSITIVE -- the client VM IS the D-138-named per-DC bastion | NEGATIVE -- no bare-metal analog, throwaway | +| Git/spec precedent | Model A archive (`docs/archive/model-a-fallback-plan.md`), already validated 11/11 | none -- net-new design | +| D-123 site-down benefit | lost (reverts to Model A's group-destroy) | lost, identically | +| New engineering needed | SEC-010 re-author for the new small VM; power-address retarget; MAC re-capture | same, PLUS a new host-side bridge/veth into a MAAS plane on vcloud itself (unprecedented surface) | + +--- + +## 4. Recommendation + +**Option 1** (flat nodes + a small, single-purpose, per-DC client VM). Rationale, in order of +weight: +1. It is the lower-risk delta -- Model A is a previously-validated, git-tagged artifact, not a + new design; only the D-138/D-132 layer on top of it is new work, and that work is small (repurpose + an existing headend shape, not invent a host class). +2. It preserves the credential-isolation property SEC-026/SEC-028 were opened to protect, where + Option 2 actively erodes it. +3. It has a NAMED Roosevelt analog in D-138's own text, so it is not scaffolding to be re-done + at the bare-metal test -- it satisfies `SCOPE-AND-EXECUTION-PLAN.md`'s minimize-delta-to- + Roosevelt framing directly. +4. The RAM difference between the two options (~176 vs ~184 GiB freed) is marginal next to the + isolation and Roosevelt-fidelity costs of Option 2. + +**Framing note for the operator gate:** "eliminate the container layer" is fully satisfied by +Option 1 under the reading that matters -- no VM in this topology is a HYPERVISOR for another VM +(no nested libvirt, no `expose_nested_virt`, no inner OpenTofu root). The retained small VM is a +flat utility sibling, the same class of thing as the Juju-controller or MAAS-region node VMs +already in the fleet, not a resurrection of the containment pattern. If the operator's intent is +literally zero additional VMs of any kind, that is Option 2, with the tradeoffs above accepted +knowingly. + +--- + +## 5. OWED live measurements (not resolvable from the repo; do not infer) + +1. **Current vcloud host capacity.** `scripts/dc-dc-whole-host-budget.py:66-67` carries a + committed default it labels "MEASURED host budget" (256 vCPU / 1024 GiB / 10240 GiB), but its + CURRENCY for the 10.13 redeploy (same hardware? any change since it was set?) needs a fresh + read-only measurement on vcloud before any FIT verdict is finalized. +2. **An exact FIT number for Option 1 vs. today's Model B sizing.** The calculator's `--control`/ + `--compute`/`--storage` flags do not yet cover the 3 utility node classes added after it was + authored (`vr1-dc0-juju-01`, `vr1-dc0-maas-01`, `vr1-dc0-tailscale-01` -- + `opentofu/vr1-dc0-substrate/main.tf:133-245`). Extending the script (or hand-totaling) and + re-running `--model A` vs. `--model B` with the CURRENT 12-VM/DC roster is owed before citing + a precise freed-capacity number in the Phase-4 change-set. +3. **MTU/jumbo budget for vcloud-level plane bridges** once the 6 planes move up a layer -- + this is W0.2's (network/wiring map) dimension, not measured here; flag as a dependency for + the Phase-0 administrator synthesis. diff --git a/docs/audit/container-elim-pass/pass1-admin-report.md b/docs/audit/container-elim-pass/pass1-admin-report.md new file mode 100644 index 0000000..9685bb7 --- /dev/null +++ b/docs/audit/container-elim-pass/pass1-admin-report.md @@ -0,0 +1,373 @@ +# Pass 1 -- ADMINISTRATOR REPORT: planning review (container-layer elimination) + +**Author:** the Phase-1 administrator (multi-agent pass, `SCOPE-AND-EXECUTION-PLAN.md` Section 4). +**Date:** 2026-08-09. **Inputs:** `pass1-w1-redeploy-teardown.md`, `pass1-w2-workflow-gates.md`, +`pass1-w3-sequencing.md`, `pass1-w4-module-planning.md` -- read in full, adversarially +cross-checked against repo ground truth (checks logged in Section 1). Baseline consumed: +`pass0-admin-report.md` incl. Section 7a gate outcome -- **Option 1 CONFIRMED** (flat node VMs +on vcloud libvirt + small non-hypervisor `vr1-dcN-client` VM per DC), **cross-DC handling (a) +CONFIRMED** (new vcloud-level host isolation control + `--check` gate + SEC-NNN row), **MAAS +region stays on `vr1-dcN-maas-01`**, rack-controller-remainder placement OPEN, all riding +D-143. READ-ONLY synthesis; no mutation; findings are logged, not executed. + +--- + +## 1. Adversarial-check results (run against repo ground truth this session) + +1. **W1.2's Stage-2-vs-Stage-3 correction: CONFIRMED, load-bearing.** Read directly: + `docs/dc-dc-deployment-workflow.md:60-88` -- Stage 2 is the Office1 SITE standup, its Owns + line is **D-114** (site containment VM `voffice1` + MAAS-composed LXD VMs), not D-123. + `docs/dc-dc-deployment-workflow.md:149-158` -- Stage 3 is the per-DC substrate, Build line + verbatim: *"D-123 **Model B** (TWO OpenTofu roots + a bootstrap gate between)"* (`:154`). + **Stage 3 is the stage container-elim restructures; Stage 2 is untouched.** D-114's + containment pattern (voffice1) is a SEPARATE, KEPT decision -- the project carries two + different "containment VM" patterns going forward (D-123's retired, D-114's kept) and the + Phase-4 change-set must name that explicitly so name-similarity does not sweep Stage 2 in. + Any earlier prompt text saying "Stage 2 restructures" is wrong; every table below keys off + Stage 3. + +2. **Root-topology fork: NOT a contradiction, but W1.3's headline over-asserts -- reconciled.** + W1.1 Sec 5 item 1 leaves "one merged root vs shared-outer + per-DC roots" OPEN (its destroy + commands are marked illustrative for exactly this reason). W1.3's B.4 step 8 headline says + "ONE FLAT `tofu` root, ONE apply cycle" -- but W1.3's own risk #5 explicitly leaves + one-state-per-DC vs one-state-total open ("this document leaves that split as a Phase-4 + decision"). Ruling of this synthesis: **the sequence holds under EITHER root shape.** The + invariant B.4 actually establishes is: one apply CYCLE per DC, run on vcloud against local + `qemu:///system`, no bootstrap-gate seam, no inner/outer ordering, no qemu+ssh dial. What + the fork changes is only (i) how destroy SCOPING is achieved (`-target` set vs "pick the + root"), (ii) per-state blast radius (~38 domains x 2 DCs under one state vs split), and + (iii) the exact wording of the teardown primitive. Read W1.3's "ONE FLAT root" as "one flat + apply scope," not a resolved root design. **Pointer discrepancy also reconciled:** W1.1 says + Phase 2/W2.1 decides; W1.3 risk 5 says Phase 4 -- resolution: **W2.1 (tofu module design) + DESIGNS the root split; Phase 4 RATIFIES it in the module-workflow synthesis.** Carried to + Phase 2 as explicitly open (Section 7). Interaction with the (a) control: under per-DC + roots the second DC's apply is a discrete event the control can precede; under one merged + root the FIRST full apply may create both DCs' planes at once -- so the only fork-robust + ordering is "control installed + verified before ANY flat substrate apply" (Section 3). + +3. **Two SEC-010 successors: CONFIRMED distinct, BOTH OPEN, neither dropped.** All three + sources agree (pass0 Section 5 "two controls"; W1.1 phase-2 table step-B row "related but + distinct"; W1.3 B.3 vs B.5 step 17 "do not read B.3 as closing SEC-010's full scope"): + - **(i) The NEW vcloud-level cross-DC host-isolation control** (Phase-0 handling (a), + gate-confirmed): asserts no inter-plane/inter-DC forwarding on vcloud's own kernel. + Mechanism undesigned; SEC-NNN unminted; stage home resolved by recommendation in + Section 3. Container-elim-only content. + - **(ii) The re-authored SEC-010 transit-leg FORWARD-drop** on the new boundary: today's + SEC-010 is interface-scoped to `vvr1-dcN`'s transit NIC (+ voffice1 peer); the qemu+ssh + purpose dissolves but operator `ssh -J` and Office1-originated flows (e.g. MAAS rack + enrollment traffic) still ride the transit -- WHICH ends get the drop (client VM + + voffice1?) is OPEN (pass0 Section 6 item 5). Sequenced at W1.3 B.5 step 17, deliberately + separate from B.3. + Neither exists yet; both are owed artifacts (Section 6). + +4. **The (a) control consolidated into ONE design requirement** -- Section 3. + +5. **D-143 orthogonality: holds as "every diff attributable," NOT as "no intersections."** + W1.1's framing (D-143 = value/data substitution; container-elim = shape change) is correct + at the structural level and the attribution discipline is sound. But four line-items are + genuinely `[both]` and must carry dual labels (W1.3 Part C + W1.2 Stage-4 row): + - **G17** (per-DC artifact source): new ADDRESS family (D-143) + new HOST (container-elim + placement) in one gate-literal edit. + - **B.1.4 / D-143 item 4** (transit routes + DC-side statics): octet-preserving math is + D-143; the BEARER host changes rack -> client VM (container-elim). + - **R7 / D-143 item 5** (credential revocation): owed BY D-143; the hosts being revoked + are container-elim-eliminated classes -- W1.1: "the SAME step wearing two decision-labels." + - **B.7** (juju/bundle deploy): overlay literals carry 10.13 (D-143); execution-host + identity changes rack -> client VM (container-elim). + Plus B.1.1 (capacity re-measurement) FEEDS both axes. Everything else in W1.3's Part C + table separates cleanly. The change-set reviewer's test stands: every diff hunk names the + axis (or, for these four, both) it belongs to. + +6. **D-128 amendment: REAL, not a doc-currency nit -- verified.** `docs/design-decisions.md: + 5354-5364` read directly: Plane 2's DEFINITION includes *"the INNER `tofu` root + (`opentofu/vr1-dc0-substrate/`, `qemu+ssh` FROM Office1 into `vvr1-dc0`, R-5)"* executing + on voffice1. Under Option 1 that object ceases to exist -- the entire substrate build + becomes Plane 1 (vcloud-local), and Plane 2 shrinks to MAAS/NetBox (with juju/openstack + already moved to the DC client by D-138). That is a substantive scope change to a RULED + decision's own definition, not stale prose. **Flagged for the Phase-4 [ARCH] decision + framing: the D-128 amendment rides alongside the D-123 amendment/new-D, the D-125 + retirement, and the D-138 concrete-host change.** + +7. **Worker-claim verification + contradictions:** + - **W1.1's phase6 "ALREADY WRONG" claim: CONFIRMED by direct read.** `runbooks/ + dc-dc-phase6-designate-cos-magnum.md:437-444` says "Run them from the Plane-2 host + (voffice1) per D-128" / "CHECK ... from the Office1 headend (voffice1)" before a + `juju status` command. Its U9 measurement (2026-07-27) predates D-138 (2026-07-30); + phase4's own RUN-LOCATION table records the correction. A pre-existing stale-D-138 + defect the rewrite sweep must also fix -- not a new container-elim delta. + - **W1.3 step 18 wording vs the refuted vm-host mechanism -- reconciled.** W1.3 B.6 + step 18 says nodes are discovered "via the vcloud-registered virsh `vm-host`"; the + workflow doc's Stage-3 State note (`:164-166`) records that `maas-vm-host` registration + was REFUTED for DCs and replaced by per-machine virsh power (D-103/D-123 amendments + 2026-07-20; W1.1's Step-D row + DOCFIX-179 agree). **Correct reading: the mechanism is + per-machine `power_type=virsh`; only the power-address VALUE re-derives** from "dial the + containment VM's libvirtd" to "dial vcloud's own libvirtd" (`lib-hosts.sh` + `VIRSH_POWER_ADDRESS*`, pass0 rows 4-5). The sequence position is right; the noun is not. + - **"Migrate" vs "re-mint" credentials -- compatible, stated so it can't read as a + contradiction.** Pass0 row 8 says residencies "MIGRATE"; W1.3 B.7.22 says credentials are + "(re-)MINTED here, not migrated." Both true at different layers: the REGISTER rows + (`vm-secret-locations`, SEC-026/-028/-029 pointers) re-point to the client-VM host class; + the credential MATERIAL is revoked with the old host (Part A step 3) and freshly minted + on the new one. No secret material moves between hosts. + - **W1.2 vs W1.4 on the (a) control's artifact kind -- taxonomy tension, resolved in + Section 3** ([3] in W1.2's IaC-layer sketch vs L5 in W1.4's model; W1.2's placement was + about INVOCATION ORDER, not artifact kind -- do not let Phase 2 build it twice). + - **W1.2 vs W1.4 on the client VM's grouping -- not a contradiction.** W1.2's sketch + bundles the client VM into the per-DC `dc-substrate-flat` apply; W1.4 classifies it L1 + (a `cloudinit-vm` instance, same module type as voffice1/edges). Layer CLASSIFICATION vs + apply GROUPING -- both can hold; Phase 2's module design picks the call site. + - **Nits carried (doc-currency, ride the rewrites):** W1.3 Part E risk-2 has off-by-one + step references ("17 (MAAS enlist target)" is step 18; "20 (juju execution host)" is + step 21). The workflow doc's Stage-3 Build line still says "~416 GiB" -- stale vs 480 + (`variables.tf:143-150`, already flagged at pass0 check 5a). W1.2's Gap-#20 verdict + ("probably unchanged") is properly hedged as needing re-measurement post-build -- carried + as owed, not as fact. + - **No inferred-value violations found** beyond the items above; spot-checked citations + (workflow doc stage rows, D-128, D-138, phase6, W1.4's module listings vs + `opentofu/main.tf`) resolved to the cited lines. + +--- + +## 2. The planning change-set (per-runbook, per-stage, per-gate) + +### 2.1 Runbooks + +| Runbook | Verdict | Change (all against confirmed Option 1) | +|---|---|---| +| `dc-dc-teardown-rollback.md` | **REWRITE** (same weight as its own 2026-07-16 "MODEL B RESHAPE" banner) | Two-clones/two-hosts preamble DELETED (one clone, vcloud); Path A root-choice becomes shared-outer vs per-DC-flat (per root-fork resolution); Step-4 verify collapses to ONE vcloud-local `virsh` block (the "vcloud grep passes against an intact DC" trap goes moot); Path B "everything" = N+1 roots on ONE host, qemu+ssh guard branch dropped; rollback-tree Question 0 redrawn (fewer roots); **item 5's `virsh destroy vvr1-dcN` lever needs a REPLACEMENT artifact** (Section 6 #2); mesh-leg caveat REWORDED (no qemu+ssh path; still carries client-VM reach + `ssh -J`); Paths **M and C are UNCHANGED as procedures** (above the substrate); footer needs a fresh re-measurement pass against the new module tree at delivery | +| `dc-dc-phase0-vcloud-prep.md` | NO CHANGE | Zero containment hits (W1.1-verified) | +| `dc-dc-phase1-office1-standup.md` | NO CHANGE | Its containment VM is `voffice1` under **D-114** -- out of scope (check 1) | +| `dc-dc-phase2-tofu-dc-substrate.md` | **HEAVIEST REWRITE** | Step B (bootstrap gate) ELIMINATED WHOLESALE; Step C (inner apply) MERGES into Step A (one apply, one root-scope, one host, one state); the DC-substrate USAGE of `expose_nested_virt` drops (`vvr1-dcN`'s true-setting; the module VARIABLE stays -- `voffice1` consumes it per the Stage-2 Build line, and the client VM sets it false); D-125 bridge-in rows DELETED (`modules/wan-bridge` dead for VR1; edge WAN -> direct `vr1_dcN_uplink` NAT, same /24); Step D loses the "inner virsh" framing (DOCFIX-179 deferral stands); netem caveat reworded only; step list re-grouped fresh-linear (no A-E lettering needed); definition-of-done re-homed off `--host-nodes`; Step-13 backup = ONE state file on vcloud | +| `dc-dc-phase3-maas-enlist-deploy.md` | LOW DELTA | Two SSH-jump-target lines (`:424,430`): far end becomes the ruled placement host (client VM OR maas-01 -- **OPEN, do not pre-pick**); mechanics otherwise topology-agnostic (per-machine virsh power; only the power-address VALUE changes) | +| `dc-dc-phase4-juju-bundle-per-dc.md` | MODERATE | RUN-LOCATION table gets its THIRD correction: row 1 juju/openstack CLI -> `vr1-dcN-client` ("the rack" is retired vocabulary); row 2 (voffice1: maas/NetBox/tofu*) and row 3 ("never vcloud") unchanged -- *tofu's Plane-2 half shrinks per the D-128 amendment; staged-scripts caveat re-points its noun (client VM still has no repo clone); below-juju-destroy caution re-points at the rewritten teardown runbook | +| `dc-dc-phase5-dr-failover-drill.md` | NO DIRECT CHANGE | Inherits phase4's table | +| `dc-dc-phase6-designate-cos-magnum.md` | RIDE-ALONG FIX | `:437-444` pre-existing stale-D-138 defect (verified, check 7) -- fix in the same sweep as phase4's table; not a container-elim delta | +| `dc-dc-office1-service-reip.md` | OUT OF SCOPE | D-114 territory | + +### 2.2 Workflow doc (`docs/dc-dc-deployment-workflow.md`) + +| Stage | Change | +|---|---| +| Stage 1 | Gate content UNCHANGED (nested-KVM still needed for Stage 2's voffice1). Relationship note: the six per-DC planes RETURN to vcloud level -- Stage 3 reconverges onto Stage 1's own execution shape. **Recommended new owner of the (a) control** (Section 3) | +| Stage 2 | **NONE** (D-114; check 1). Phase-4 change-set must state the two-containment-patterns distinction explicitly | +| Stage 3 | **THE restructured stage.** Build line: Model-B two-root text -> single flat apply (planes + edge + 12 node VMs + client VM); also fix stale "~416 GiB" en route to deleting the sizing. Gate line: depth-4 boot gate -> depth-2; bridge-in isolation test -> direct-NAT egress assertion; add the (a) `--check`. Owns line: D-123/D-125 phrasing VOID pending the Phase-4 ruling; D-124 survives only for the client VM's transit leg. Reuse-vs-new: "NEW, no precedent" should be revisited -- the flat shape is MORE reusable, closer to Stage 1's | +| Stage 4 | Placement, not mechanism: gate-line prose + G17 literal need the ruled artifact-source host + 10.13 address (dual-cause, check 5) | +| Stage 5 | No gate-content change; every literal naming the containment VM's transit IP re-points to the client VM (e.g. `docs/CURRENT-STATE.md:7829` "openstackclient ... ON THE dc0 RACK (172.31.0.2)"). D-140 keeps Stage 5 a PROCEDURE module -- settled constraint, not open | +| Stages 6-7 | No container-layer dependency found (verified by W1.2, not assumed) | +| Gap register | #2 reshapes (first Stage-3 exercise becomes a single flat apply; wan-bridge deleted; inner roots retire); #17 closing-mechanism note becomes HISTORICAL (doc-currency addendum owed at ruling time); #19b unaffected; #20 verdict RE-VERIFY post-build (its own expiry clause triggers); #21, #22 unaffected (checked, excluded); **NEW register entry owed for the (a) control** (no entry exists today) | + +### 2.3 G-series gates (`docs/CURRENT-STATE.md` section 6) + +| Gate | Status | +|---|---| +| G9/G10 | CLOSED/historical; their SUCCESSOR for the 10.13 rebuild is a SINGLE apply-and-verify gate (no outer/inner pair): substrate apply + (a) `--check` + depth-2 boot proof + direct-NAT egress test. The retired `--host-nodes --check` sub-item is replaced by the (a) control's check -- which is host-scoped, hence the Stage-1 home (Section 3) | +| G12 | CLOSED/historical -- read as "the shape being replaced," never a template | +| G17 | OPEN; the one dual-cause edit (D-143 address + container-elim placement) -- record as two line-items landing in one edit | +| G14 | Indirect: residency re-points + >=1 new SEC row are COUNT-affecting; flag for the next `ledger-scan.sh` reader (instrument-currency lesson #25) | +| G18, G1-G8, G11, G13, G15, G16 | No container-layer dependency (verified per-gate by W1.2) | + +--- + +## 3. The (a) cross-DC isolation control -- ONE consolidated design requirement + +Consolidating Phase-0 7a (confirmed deliverable: `--check` gate + SEC-NNN row), W1.2 (no +stage home, no register entry, risk of ad-hoc under-gated build), and W1.3 (must precede +co-residency; E.1): + +**Requirement.** A vcloud-level host isolation artifact (nftables, SEC-010's proven pattern +one layer up) asserting no inter-plane / inter-DC forwarding on vcloud's own kernel, shipping +with: a mechanical `--check` gate, its own `tests//run-tests.sh` harness, a new SEC-NNN +ledger row, a NEW workflow-doc gap-register entry, and a gate row in the rebuild's G-series +successor. Parameterized by site token (W1.4 principle 1), never DC-hardcoded. + +**Ordering invariant (fork-robust).** The control must be installed and `--check`-verified +**BEFORE the first flat substrate apply that can make any two DCs' planes co-resident** -- +i.e. before ANY per-DC flat apply, since under a merged single root the first apply may create +both DCs' planes at once (check 2). "Before the second DC's apply" (W1.3's phrasing) is the +minimum; "before any flat apply" is the only ordering that survives the open root fork. + +**Stage placement recommendation: Stage 1 (vcloud host prep), as a host-level control**, +re-verified (i) at each per-DC substrate apply's close and (ii) at Stage-5 verify-live +(W1.3 B.7 step 24 -- the first point real traffic tests the claim). Rationale: it is +host-scoped, not per-DC-apply-scoped (W1.2's G10 analysis); Stage 1 gives it a named stage +owner so it cannot be orphaned; and the placement is indifferent to the root-topology fork. + +**Artifact kind (resolves W1.2-vs-W1.4):** SEC-010's actual pattern -- a script-installed +nftables control + `--check` + harness, i.e. a **procedure/L5 verify artifact**, NOT an +OpenTofu module. W1.2's sketch slot "[3]" expressed invocation ORDER (before any DC apply), +not artifact kind. Phase 2 builds ONE artifact. + +**Distinct from** the re-authored transit-leg FORWARD-drop (check 3) -- two controls, two SEC +rows, both owed. + +--- + +## 4. Ordered teardown -> redeploy sequence (consolidated from W1.3, corrected per Section 1) + +Two axes tagged throughout: `[D-143]` (value substitution) / `[CE]` (container-elim, shape) / +`[both]` (dual-labeled per check 5). + +### Part A -- teardown of the CURRENT 10.12 Model-B checkpoint (today's tooling; the runbook +already describes this shape correctly and needs no rewrite to tear it down) + +1. `[D-143]` Back up state -- BOTH roots, BOTH hosts (outer on vcloud; each inner on voffice1). +2. `[D-143]` MAAS machine census FIRST, from voffice1 (Step 2, both lenses). +3. `[both]` **R7 credential-revocation checklist (OWED BUILD -- Section 6 #5), run BEFORE any + substrate destroy** (revoking after the hosts are gone degrades to "assume it's moot"): + SEC-026 client credential, SEC-028 service credential, SEC-029 rack-local PKI copy (shred; + headend canonical unaffected), D-126 per-env qemu+ssh keys (clean RETIREMENT -- no + successor exists), rack enroll-secret residue. Enumerate from EVERY `vm-secret-locations` + row keyed to the rack host class, not just the named examples; mark rows RETIRED + (append-only). Each revocation individually confirmed with captured output. +4. `[D-143]` **MAAS machine-record release/delete (OWED sequenced step -- Section 6 #6)**: + guaranteed non-zero LENS-2 on a live checkpoint; release/delete per record + (operator-gated), re-run both lenses to zero; ALSO clean the region-side residue -- the + `vvr1-dcN` rack-controller's own enrollment record + the region's `primary_rack`/DHCP + reference. +5. `[D-143]` Plan destroy, INNER first (from voffice1), capturing the pre-destroy virsh + baseline while the containment VM is still reachable. +6. `[D-143]` Plan destroy, OUTER second (vcloud): `vvr1_dcN` + uplink + storage modules only. + Do NOT target mesh/netem legs (working assumption: mesh triangle SURVIVES the pivot -- + confirm at Phase 2, do not destroy speculatively). +7. `[D-143]` Apply destroys, INNER then OUTER; verify each half from the host that can see it. +8. `[D-143]` Untargeted `tofu plan` drift gate: zero unexpected drift. +9. `[D-143]` NetBox DCIM decommission of the `vvr1-dcN` device records. +10. Repeat per DC (Path A) or batched (Path B). + +### Part B -- redeploy on flat 10.13 (Option 1) + +- **B.1 prerequisites (once, before any apply):** `[both]` vcloud capacity re-measurement + + FIT-calculator extension (the "~176 GiB freed" stays directional until then); `[D-143]` + NetBox B2 apex re-carve (10.13.0.0/16, octet-preserving); `[D-143]` `lib-net.sh` 10.13 + literal blocks; `[both]` D-124 transit re-point -- octet math is D-143, the bearer host + changes rack -> client VM (CE). +- **B.2:** `[unchanged]` vcloud host prep; Office1 headend (if not already up). voffice1 + remains MAAS-region + Plane-2 host for what still needs it. +- **B.3:** `[CE]` **Install + verify the (a) isolation control -- BEFORE ANY flat substrate + apply** (Section 3's fork-robust invariant; supersedes W1.3's "before the second DC's + apply" as the minimum reading). +- **B.4:** `[CE]` **One flat apply CYCLE per DC** (root shape per the OPEN fork, Section 7), + vcloud-local `qemu:///system`: mesh legs + uplink NAT `[unchanged]`; per-DC storage pool + collapses to one; six planes re-homed (same CIDRs/families/MTU -- D-139/D-143 own the + values); edge re-homed, WAN -> direct NAT; 12 node VMs re-homed (**MAC re-pinning is a + likely force-replace -- re-measure every MAC after apply, before B.6 trusts one**); `[new]` + the `vr1-dcN-client` VM. VANISHES: `vvr1_dcN` modules + sizing/addressing/pubkey vars, the + two inner roots AS ROOTS, the qemu+ssh dial + D-126 keys (no successor), `wan-bridge` + + netplan bridge, the bootstrap gate's `--host-nodes` duty. +- **B.5 (surviving duties, re-targeted -- placement OPEN, a real sequencing dependency, not a + footnote):** rack enrollment (`site-headend-install.sh --role rack`, WITHOUT `--host-nodes`) + onto the ruled host; D-131 forwarder (`dc-rack-net.sh`) -- must land SOMEWHERE or the + SERVFAIL bug returns; artifact service (`.4`); **the re-authored transit-leg FORWARD-drop + (control (ii), Section 3 -- distinct from B.3, ends still open)**. +- **B.6:** `[unchanged*]` MAAS discovery/commission/carve -- mechanism is per-machine + `power_type=virsh` (check 7 wording correction); *only the power-address VALUE re-derives + to vcloud's own libvirtd (`lib-hosts.sh` re-derivation + every invocation-site example). +- **B.7:** `[unchanged]` Juju bootstrap + bundle from the `vr1-dcN-client` VM ("never the + vcloud jumphost" doctrine intact); SEC-026/028/029 freshly MINTED here (register rows + re-point; no material migrates -- check 7 reconciliation); verify-live gates: Ceph-over-v6 + (new literals per D-143), `geneve-encap-assert.sh` (MTU budget analytically unchanged -- + live assert still OWED), **plus the (a) control's `--check` re-run now that both DCs' + planes are actually co-resident**. +- **B.8 close-out:** `[CE]` NetBox DCIM registration of the client VM + flat roster; + `[deferred]` the container-elim [ARCH] ruling (Phase 4 frames, operator rules) -- the + redeploy is NOT "done" while that ruling is owed. + +--- + +## 5. The layer model (adopted from W1.4 -- the Phase-4 module-workflow backbone) + +L0 host/inter-site substrate (IaC: mesh, pools, office1-network -- Stage 1) -> L1 site/edge +nodes (IaC: voffice1, DC edges, **+ the client VM as a new instance of the SAME +`cloudinit-vm` module type**) -> L2 DC substrate (IaC: planes + node VMs; **the two-root +split collapses INTO L2 -- the one place the layer-boundary rule was violated** (inner root +dialing a cross-host provider) **and Option 1 removes that violation structurally**) -> L3 +enlist/commission (procedure; rack-remainder placement OPEN) -> L4 Juju/OpenStack deploy +(procedure; **D-140 PINS it as procedure for this redeploy -- settled, do not fold into +IaC**) -> L5 verify/gate (cross-cutting; the (a) control is a NEW L5 artifact per Section 3). + +Design principles carried to Phase 4: site-token parameterization (never DC-hardcoded); +every module ships its harness; idempotence at every layer; no layer reaches past the one +below; findings logged at their true layer; D-140 is a distinct axis (its later adoption +would make L4 IaC WITHOUT reshaping L0-L3); Roosevelt-transfer judged per layer (L0's +node-VM shim does NOT transfer; L1's client-VM pattern + L3/L4 procedures ARE the +pre-Roosevelt deliverable). D-143 is a PARAMETER change threaded through every layer, not a +layer -- the model keeps the two axes structurally distinguishable. + +--- + +## 6. OWED ARTIFACTS surfaced this phase (feed Phase 2 tools / Phase 3 tests) + +The five core artifacts: + +1. **Teardown-primitive** -- module/root-scoped `tofu` group-destroy procedure (+ harness) + re-earning D-122's one-command site-down for the flat shape (final form depends on the + root fork; Phase 2 designs, Phase 4 ratifies). +2. **The (a) cross-DC host isolation control** -- nftables artifact + `--check` gate + + harness + new SEC-NNN row + new gap-register entry + gate row (full spec Section 3). +3. **Re-authored SEC-010 transit-leg FORWARD-drop** -- control (ii); endpoint set (client + VM + voffice1?) still open; own check + SEC-row disposition. +4. **R7 credential-revocation checklist** -- built by enumerating `vm-secret-locations` + rows for the rack host class; gates Part A step 3. +5. **MAAS machine-record release/delete step** -- promoted from runbook contingency to a + sequenced, operator-gated Part-A step, incl. region-side rack-controller record + + `primary_rack`/DHCP-reference cleanup. + +Additional owed items surfaced (kept separate so the core count stays legible): +6. **Emergency site-down lever** -- scripted `virsh destroy` loop over the DC's domain set, + roster-derived from `lib-hosts.sh` (rollback-tree item-5 replacement; W1.1 Sec 1(b)) -- + distinct from #1 (emergency vs gated path). +7. **FIT-calculator extension** (3 utility-node classes) + fresh vcloud capacity measurement. +8. **MAC re-measurement pass** post-B.4, before B.6 (likely force-replace). +9. **NetBox DCIM migration** -- decommission `vvr1-dcN` records; register client VM + roster. +10. **Post-build live asserts** -- geneve/jumbo (`geneve-encap-assert.sh`) on the vcloud-level + planes; gap-#20 `site-baseleg` verdict re-verification (its own expiry clause triggered). + +--- + +## 7. Settled vs OPEN + +**SETTLED (do not re-open):** Option-1 target; handling (a) as a required deliverable; MAAS +region stays on `vr1-dcN-maas-01`; Stage 2/D-114 out of scope; Paths M + C unchanged; +D-140 pins Stage 5/L4 as a procedure module for this redeploy; mesh triangle persists as the +working assumption (Phase-2 confirmation, not speculation-destroy); the two-axis attribution +discipline (with the four named `[both]` items); the (a)-control ordering invariant +(before ANY flat apply) and its recommended Stage-1 home; the teardown ORDER of the current +checkpoint (Part A -- unchanged by container-elim). + +**OPEN -- carried to Phase 2 (design):** +1. **Root topology** -- merged single root vs shared-outer + per-DC roots (W2.1 designs; + Phase 4 ratifies; the sequence is invariant to it, the teardown primitive's wording and + state blast radius are not). +2. **Rack-controller remainder + D-131 forwarder + artifact-service (`.4`) placement** + (client VM vs maas-01 vs retire-with-evidence) -- blocks B.5, the phase3 SSH-target + edits, and G17's host half. THE highest-leverage open item: three sequence steps and two + runbook edits key off it. +3. **The (a) control's concrete mechanism** (nftables rule set, check shape, SEC number). +4. **Transit-drop endpoints** for control (ii). +5. **Client VM octet + name** (D-134 standing map needs a ruled octet; name must NOT be + `vvr1-dcN`). +6. **Mesh-triangle survival through the pivot** -- confirm module shape at Phase 2. +7. **`site-headend-install.sh` refactor scope** (rack-role remainder vs dead `--host-nodes`). + +**OPEN -- carried to Phase 4 (decision framing; operator rules, GA-R5):** +8. The container-elim [ARCH] ruling itself -- D-123 amendment vs new D-number. +9. **The D-128 amendment** (verified real, check 6): Plane 2 shrinks to MAAS/NetBox; the + substrate build becomes wholly Plane 1. +10. Ride-alongs on the same ruling: D-125 bridge-in retirement, D-138 concrete-host change, + D-122 site-down re-earn, D-124 sizing-void re-cause, Stage-3 Owns/Reuse-vs-new rewrite. +11. State-blast-radius weighing (rides open item 1). + +**OPEN -- operator inputs (SCOPE Section 8):** pre-Roosevelt hardware specs (plug into B.1.1 +capacity + B.4 sizing only -- and B.5's placement should be RE-DECIDED for bare metal, not +carried blindly); any external "module deployment project" artifacts. + +--- + +## 8. Verification note + +Author = "the administrator" (no model name asserted, operator instruction). Direct reads +this session: `docs/dc-dc-deployment-workflow.md:36-200` (Stage 1/2/3 identity), `docs/ +design-decisions.md:5344-5373` (D-128), `runbooks/dc-dc-phase6-designate-cos-magnum.md: +430-450` (the "ALREADY WRONG" claim), plus the six pass documents in full. Worker citations +were spot-checked, not re-derived wholesale; every correction in Section 1 names its source +lines. Findings are LOGGED only; nothing here was executed. diff --git a/docs/audit/container-elim-pass/pass1-w1-redeploy-teardown.md b/docs/audit/container-elim-pass/pass1-w1-redeploy-teardown.md new file mode 100644 index 0000000..7aa7372 --- /dev/null +++ b/docs/audit/container-elim-pass/pass1-w1-redeploy-teardown.md @@ -0,0 +1,300 @@ +# Pass 1 -- Worker W1.1: redeploy/teardown PLANNING under the confirmed Option-1 target + +**Author:** W1.1 (Phase 1, container-elimination pass). **Date:** 2026-08-09. **Scope:** +what the CONFIRMED Option-1 target topology (`pass0-admin-report.md` Section 7a) changes in +the teardown and redeploy PLANNING ARTIFACTS -- `runbooks/dc-dc-teardown-rollback.md`, the +`runbooks/dc-dc-phase0..6*.md` runbooks, and how D-143's re-IP interleaves with the +container-elim change-set in the redeploy sequence. **READ-ONLY.** No mutation, no live +commands. Every claim below is anchored to a `path:line` read this session; where a value is +not yet knowable it is marked UNKNOWN with what resolves it. + +**Reminder of the confirmed target (do not re-derive):** flat node VMs directly on vcloud +libvirt; `vvr1-dcN` ELIMINATED; a small non-hypervisor `vr1-dcN-client` VM (metal-admin + +transit legs) inherits the D-138 client role + that DC's credential residencies; MAAS region +stays on `vr1-dcN-maas-01` unchanged; rack-controller remainder / D-131 forwarder / artifact +service placement is OPEN (client VM vs maas-01, carried into this phase per Section 6 item 3 +of `pass0-admin-report.md`); cross-DC isolation (a) is a REQUIRED Phase-1/2 design deliverable +(a new vcloud-level host isolation control); this rides D-143 but must stay DISTINGUISHABLE +from it. + +--- + +## 1. The teardown-primitive answer + +**Today:** `runbooks/dc-dc-teardown-rollback.md`'s whole structure is built on the D-123 +Model-B TWO-ROOT/TWO-HOST shape -- an OUTER root on vcloud that creates `vvr1-dcN`, and an +INNER root run FROM THE OFFICE1 HEADEND (`voffice1`) via `qemu+ssh` into the containment VM +(lines 26-65, the "MODEL B RESHAPE" banner; lines 136-152, "TWO CLONES, TWO HOSTS"). D-122's +site-down lever is explicitly one object: `virsh destroy vvr1-dcN` (line 48). Every Path +(A/B/C/M) is written against this shape. + +**Under Option 1, the two-root/two-host shape COLLAPSES to one root on vcloud.** There is no +containment VM to hold an inner libvirt, so there is no `qemu+ssh` provider dial, no inner +state file on `voffice1`, and no D-122 single-object site-down (that lever is explicitly +named a "REAL D-123 regression to carry honestly" in `pass0-admin-report.md` row 9). The +teardown primitive becomes: + +> **A single-root, module-scoped `tofu destroy` per DC, run entirely on vcloud** -- +> `tofu destroy -target=module.vr1_dcN_planes -target=module.vr1_dcN_node -target=module.vr1_dcN_wan -target=module.vr1_dcN_opnsense -target=module.vr1_dcN_client -target=module.vr1_dcN_storage` (module +> names illustrative -- Phase 2/W2.1 owns the real post-flattening module names), OR (preferred, since +> a flat root holding exactly one DC's resources needs no `-target` at all, same logic the current +> runbook already uses for the per-DC INNER root, lines 751-758) **one dedicated root per DC** so a +> plain `tofu destroy` there is inherently scoped -- re-earning the "scope by choosing the root" +> discipline (current lines 155-168) without a second host. + +This is the **group-destroy of the `vr1-dcN-*` domain set**, not a single-object destroy -- +D-122's convenience is genuinely lost, and Section 5's cross-DC isolation gap (Section 5) is +one more reason NOT to try to re-earn it via "add a synthetic wrapper VM back" -- that would +just re-import the eliminated pattern. The honest replacement discipline: (a) a *tofu*-level +group-destroy scoped by root/module (safe, plan-reviewable, the mechanism above), and (b) for +a "stop this DC's compute NOW" emergency lever (the actual use case D-122's single-`virsh` +convenience served), a SCRIPTED loop over `virsh destroy` for the named domain set derived +from `lib-hosts.sh`'s roster for that DC -- not a manual enumeration. **Both are NEW artifacts +this pass must hand to Phase 2/4 (tooling + module design), not something already built.** + +### 1.1 Per-Path change list + +| Path | Today's assumption (path:line) | Option-1 change | +|---|---|---| +| **"TWO CLONES, TWO HOSTS" preamble** (`:136-152`) | Every command runs on `$REPO` (vcloud, outer) OR `$O1_REPO` (voffice1, inner substrate roots) | COLLAPSES to ONE clone/one host: `$REPO` on vcloud only. The whole two-host framing is DELETED; `$O1_REPO` no longer holds a DC substrate root (voffice1 keeps its OWN unrelated Office1 role, D-114 -- untouched, see Section 2) | +| **Path A -- scoped teardown** (`:154-173`, Steps 3-5) | "Scope by choosing the ROOT first" between OUTER (vcloud, all 3 sites) and INNER (per-DC, voffice1, `qemu+ssh`); order = INNER then OUTER, backups from TWO hosts (Step 1, `:600-631`) | Root choice becomes: OUTER (shared: Office1 + mesh + both DC-storage-pool-parent, if any survives flat) vs. a PER-DC flat root holding that DC's planes/nodes/edge/client VM. Order collapses to ONE apply per DC (no inner/outer split); Step 1's backup is ONE state file per DC root, taken on vcloud only. Step 3's module table (`:765-775`) is rewritten with post-flattening module names (Phase 2/W2.1 -- names not yet minted, UNKNOWN pending that design) | +| **Path A -- Step 2 (MAAS census)** (`:635-747`) | "WHERE: from the OFFICE1 HEADEND ... the region MAAS and its CLI profile live there" (`:655-658`); LENS 2 sources `lib-hosts.sh` boot-MAC roster via `lib_hosts_select_dc` | MAAS region is UNCHANGED (stays `vr1-dcN-maas-01`, confirmed at the Phase-0 gate) -- **this step's WHERE clause survives if maas CLI still runs from voffice1** (D-128 Plane 2), but the machine roster it's checking against no longer includes the containment VM itself (there is none) and now includes `vr1-dcN-client`. Mechanically the SAME two-lens structure, refreshed roster | +| **Path A -- Step 4 verify** (`:838-891`) | "The DC's node, edge and plane objects live inside the containment VM, so a local `virsh` on vcloud can never see them" (`:858-860`); verify INNER half via `qemu+ssh` from voffice1, OUTER half via local `virsh` on vcloud (two verify blocks) | COLLAPSES to ONE verify block, local `virsh` on vcloud only -- planes/nodes/edge/client VM are all now vcloud-libvirt-visible directly. The "vcloud-local grep 'passes' against a fully intact DC" caveat (`:861-864`) becomes MOOT (there is nothing hidden behind a second libvirt any more) | +| **Path B -- full VR1 teardown** (`:914-957`) | "everything" = THREE roots on TWO hosts (`:920-932`) | "everything" = N+1 roots (one shared/outer + one per DC, or a single unified root -- module design decides, Phase 2) on ONE host. Verify block (`:941-953`) drops its `qemu+ssh` refusal-guard branch entirely | +| **Path M -- juju MODEL teardown** (`:188-392`) | Untouched by containment: it is one layer ABOVE substrate (juju model on top of already-provisioned MAAS machines). No `qemu+ssh` or `vvr1-dcN` reference in the body | **UNCHANGED as a procedure.** Only its execution-HOST assumption (implicit: wherever `juju` runs, i.e. the DC rack today) moves with D-138's client-VM re-home -- see Section 3 below (phase4's table, not this runbook) | +| **Path C -- juju CONTROLLER teardown/rebuild** (`:394-596`) | Untouched by containment for the SAME reason as Path M -- it operates on MAAS machine records and the juju controller, one layer above substrate. C.1's MAAS census and C.6's tag check are host-agnostic (maas CLI location unaffected: still MAAS-region-side, i.e. voffice1 per D-128 Plane 2, unchanged) | **UNCHANGED as a procedure.** No edit needed in this runbook; the D-138 client-VM re-home is the phase4 change, not this one | +| **Mesh-link teardown** (`:960-995`) | Three OUTER-root legs (`mesh_vr1_dc0_vr1_dc1`, `mesh_vr1_dc0_office1`, `mesh_vr1_dc1_office1`) + netem, sharing the D-125 caveat that `mesh_vr1_dc0_office1` "carries the live rack<->region transit ... the inner root's own qemu+ssh path" (`:981-983`) | The mesh legs and netem module PERSIST (confirmed non-eliminated, `pass0-admin-report.md` Sec 1.3). The `mesh_vr1_dc0_office1`/`mesh_vr1_dc1_office1` caveat's WORDING changes: it no longer carries "the inner root's qemu+ssh path" (that's gone) but DOES still carry the `vr1-dcN-client`'s reach to voffice1/MAAS-region and the operator `ssh -J` path (Section 6 item 5 of `pass0-admin-report.md`, still OPEN) -- reword, don't delete the caution | +| **"Relationship to D-061"** (`:86-134`) | Frames the coordination principle as OpenTofu's resource view vs. MAAS's per-machine view (unaffected by which root creates the resources) | **UNCHANGED in principle.** No containment-specific text here to edit | +| **Rollback decision tree** (`:998-1069`) | "Question 0 -- WHICH ROOT failed? There are three" (`:1006-1010`); item 5's containment-VM-destroy caution (`:1059-1067`) is written entirely around `virsh destroy vvr1-dcN` as the reversible site-down lever | Question 0 becomes "which DC's flat root" (fewer roots, no inner/outer split). **Item 5 needs a REPLACEMENT lever entirely** -- there is no single reversible power-off object any more; the nearest analog is a scripted `virsh destroy` loop over that DC's node-VM domain set (irreversible-feeling but mechanically the same primitive, just N domains instead of 1) -- flag as a NEW tested artifact this pass owes (Section 1 above), not a drop-in rename | +| **Verification footer** (`:1071-1091`) | Notes all module/network/pool names were "RE-MEASURED against the tree and the live hosts on 2026-07-29" for the two-root shape | The whole footer's provenance is for a shape being eliminated -- a rewritten runbook needs its OWN re-measurement pass against the new flat module tree (Phase 2/4 delivery item, not this planning pass) | + +**Net verdict on this runbook:** it needs a REWRITE comparable in scope to the 2026-07-16 +"MODEL B RESHAPE" banner that was bolted on for the *previous* topology change -- likely the +same pattern (a prominent "OPTION-1 FLATTEN" banner up top plus the Steps re-grouped), not a +line-by-line patch. Paths M and C are the two genuinely LOW-delta paths (they sit above the +substrate layer entirely); Paths A/B and the rollback tree carry the real rewrite weight. + +--- + +## 2. Per-phase-runbook change list + +### 2.0 What does NOT change + +- **`runbooks/dc-dc-phase0-vcloud-prep.md`** -- grep for `vvr1-dc`/`containment`/`qemu+ssh`/ + `inner root`/`outer root` returns ZERO hits in the body (only a generic `qemu+ssh://...` example + at `:113,123` describing the libvirt-URI syntax itself, and `libvirt_uri = "qemu:///system" # or + qemu+ssh://... -- Step 1` at `:364`, which is the OUTER root's own provider block, unaffected). + This runbook covers vcloud host prep, the D-100 mesh triangle, storage pools, the shared base + image -- none of it is containment-specific. **No change required.** +- **`runbooks/dc-dc-phase1-office1-standup.md`** -- its "containment VM" is `voffice1` under + **D-114**, a DIFFERENT decision than D-123's Model-B DC containment (`:9,13,44-47,313-337` -- + D-114 is cited by name throughout, never D-123). `voffice1` simulates the Office1 FACILITY + and hosts MAAS-region/LXD for MAAS-composed service VMs (NetBox, tailscale) -- it is not the + object the container-elim pass targets, and Phase-0's scope explicitly named the container + layer as `vvr1-dcN` only (`SCOPE-AND-EXECUTION-PLAN.md` Sec 2, `pass0-admin-report.md` Sec 1.3 "It is + NOT: ... the DC edge ..." -- Office1 wasn't even in scope). **No change required for THIS pass.** + Flag for the record: D-114's own containment pattern is architecturally the same shape D-123 + used, and if the operator's "eliminate the container layer" intent is read as broader than the + Phase-0 gate scoped it, that is a SEPARATE decision outside this pass's mandate -- not assumed + here. + +### 2.1 `runbooks/dc-dc-phase2-tofu-dc-substrate.md` -- the largest single change + +This runbook IS the two-root apply sequence, and its "D-123 MODEL B RESHAPE" callout +(`:144-204`) is the direct predecessor of the change this pass makes. Per-step: + +| Step (path:line) | Today's assumption | Option-1 change | +|---|---|---| +| **A. OUTER apply** (`:151-152`) -- creates + sizes `vvr1-dc0` (~480 GiB, `expose_nested_virt=true`) + transit | Sizes and provisions the CONTAINMENT VM itself | Sizes and provisions the FLAT node-VM set + edge + `vr1-dcN-client` directly. `expose_nested_virt=true` DROPS (no nested libvirt needed anywhere in this DC -- confirm no OTHER consumer needs nesting before deleting the flag; none found in Phase-0's inventory). Overhead sizing (~480 GiB minus node/edge sizes) is FREED -- ties to the owed FIT recompute (`pass0-admin-report.md` Sec 8.2) | +| **B. BOOTSTRAP GATE** -- `site-headend-install.sh --role rack --host-nodes ...` (`:153-157`) | A DISTINCT stage between the two applies: enrolls the rack to the Office1 region AND makes `vvr1-dc0` a nested libvirt host (kvm nested=1, inner pool dir, AppArmor, SEC-010 transit FORWARD-drop). GATE: `--check` must pass | **THIS STAGE IS ELIMINATED WHOLESALE.** There is no "make a VM a nested libvirt host" step because there is no nested libvirt host. What SURVIVES from its payload, re-homed: (i) rack enrollment to the Office1 region -- lands wherever the rack-controller remainder is placed (client VM / maas-01 / retire -- Section 6 item 3 of `pass0-admin-report.md`, still OPEN, UNKNOWN pending that placement decision); (ii) the SEC-010 transit FORWARD-drop -- REBUILT for the new boundary, not moved (`pass0-admin-report.md` row 6, "must be REBUILT ... not moved"); this is the SAME control the Phase-0 gate's cross-DC isolation deliverable (handling (a)) needs a new design for -- the two are related but distinct (SEC-010 is transit-leg-scoped; (a) is host-level cross-DC). `site-headend-install.sh --host-nodes` mode becomes DEAD CODE for VR1 (confirmed `pass0-admin-report.md` row 6) | +| **C. INNER apply** (`:158-161`) -- provider = `qemu+ssh` to `vvr1-dc0`, run FROM voffice1; the 6 planes + wan + edge + inner pool + 9 node VMs | A SEPARATE apply, separate state, separate host, gated behind B | **MERGES INTO STEP A.** One apply, one root, one host (vcloud), one state file. The 6 planes / wan / edge / node VMs (now +1 for the client VM) are ALL created by the single OUTER apply -- there is no "inner" any more. Steps 5-6 (wiring `modules/opnsense-edge` and `modules/node-vm` calls into `main.tf`, `:129-141`) collapse from "authored in the inner root" to "authored directly in the (now singular) root's `main.tf`" | +| **D. maas-vm-host** (Step 9, `:162-163,588-640`) -- register `vvr1-dc0`'s OWN local `qemu:///system` virsh to the Office1 region | Registers the CONTAINMENT VM's inner libvirt as a MAAS vm-host | Registers vcloud's OWN libvirt (already the case for the OUTER root's objects) OR is retired if the flat topology never needed per-machine `power_type=virsh` registration to change shape -- **this step was already DEFERRED (DOCFIX-179) before this pass**; Option 1 does not resurrect it, it just removes the "vvr1-dc0's OWN inner virsh" framing since there is no inner virsh | +| **E. netem** (Step 11, `:164-168`) -- runs on vcloud-level mesh bridges, OUTER, already correctly targets the dc0<->dc1 leg, NOT the transit leg that "carries ... the inner root's own qemu+ssh" | Already outer-scoped, mostly unaffected | The caveat text needs its "the inner root's qemu+ssh" clause reworded (that path no longer exists) but the netem TARGET and mechanism are unaffected -- low-delta edit | +| **D-125 bridge-in** (`:193-203`) -- OUTER creates the vcloud ISP NAT + a 2nd IP-less uplink NIC + `br-vr1-dc0-wan` bridge on `vvr1-dc0`; BOOTSTRAP `--check` verifies the bridge; INNER's `vr1-dc0-wan` is a BRIDGE (`modules/wan-bridge`) onto it | Exists SOLELY to fix OBS-3 nesting egress -- an artifact of the containment VM being a two-NIC transit-only host with no egress of its own | **DELETED per `pass0-admin-report.md` row 3.** The DC edge WAN attaches DIRECTLY to the outer `vr1_dc0_uplink` NAT (same /24, no re-address) -- there is no intermediate host needing a bridge-in fix. `modules/wan-bridge` becomes dead code for VR1; the BOOTSTRAP `--check`'s bridge-verify clause is deleted with the gate itself | +| **"Sequence" step list** (`:114-142`) | Numbered 1-12 against the single-root FRAMING, then regrouped A-E for Model B | Needs a THIRD regrouping (or, cleaner, a fresh linear list -- Option 1's flat shape doesn't need the A-E lettering since there's no bootstrap-gate seam to letter around). This is the step-list equivalent of the runbook-wide rewrite flagged in Section 1 | +| **DC standup definition-of-done** (`:822-887`) | References "the containment/service net" MEASUREMENT and "containment ssh shape" (`:830-833`) and objects "created by `site-headend-install.sh --host-nodes`" (`:842`) | Both references need re-homing to "how vcloud reaches the DC's node-VM planes directly" and drop the `--host-nodes` citation (dead per row above) | +| **Step 13 backup** (`:741-822`) | Backs up "the inner tfstate" via `voffice1` (`:764-804`), noting a `qemu+ssh`-only provider dependency is "strong[er]" evidence of correctness (`:764`) | Backs up ONE state file (the flat per-DC root), taken and stored on vcloud -- the voffice1-hop and its "no key in state" argument are MOOT (no `qemu+ssh` provider block exists to make that argument about) | + +**Net verdict:** Phase 2 needs the SAME weight of rewrite as its own prior D-123 reshape -- +arguably heavier, since this pass ELIMINATES a whole apply stage (B) rather than adding one. + +### 2.2 `runbooks/dc-dc-phase3-maas-enlist-deploy.md` + +- **`:424,430`** -- `ssh -i ~/vr1-dcN-creds/... -J voffice1 jessea123@ 'sudo bash -s' -- check dcN < scripts/dc-mirror.sh` (dc0) / `dc-cache-proxy.sh` (dc1). **Today's assumption:** the jump target is the rack transit IP, i.e. the CONTAINMENT VM (`vvr1-dcN`'s own transit leg, per `lib-hosts.sh`'s `VIRSH_POWER_ADDRESS`). **Option-1 change:** the jump target becomes `vr1-dcN-client`'s transit IP -- the artifact-service check (dc-mirror/dc-cache-proxy) still runs via SSH-through-voffice1 into the DC, but the far end of that hop is a different (non-hypervisor) VM. The artifact-service PLACEMENT itself (client VM vs. `vr1-dcN-maas-01` vs. retire) is the OPEN item from `pass0-admin-report.md` Section 6 item 3 -- if it lands on `maas-01` instead, this whole SSH target changes to that VM's address, not the client VM's. **UNKNOWN pending that placement ruling; do not pick one here.** +- Everything else in Phase 3 (MAAS enlist/commission/deploy of the node fleet, Steps 1-7 generally) is about MAAS machine-level operations against a fleet that is already flat from MAAS's point of view TODAY (per-machine `power_type=virsh`, no pod) -- `pass0-admin-report.md` confirms this is NOT a two-root/containment-keyed concern (row 4/5: only the POWER ADDRESS changes, the mechanism is topology-agnostic). **Low delta beyond the two SSH-target lines above.** + +### 2.3 `runbooks/dc-dc-phase4-juju-bundle-per-dc.md` + +- **`:158-185`, the "RUN LOCATION" table** -- this is the runbook's own record of getting D-138 + wrong once already (superseded text quoted at `:169-178`: "Every juju and maas command runs + on voffice1" was WRONG; D-138 moved juju/openstack CLI onto "THIS DC's RACK"). **Under + Option 1, "THIS DC's RACK" is retired vocabulary --** the concrete host D-138's principle + points at becomes `vr1-dcN-client` (per `pass0-admin-report.md` Sec 3.3, Option 1's client VM is + "consistent with D-138's principle"; the concrete-host CHANGE still needs recording, Phase 4's + [ARCH] decision). This table needs a THIRD correction pass: row 1 ("juju/openstack CLI") -> + runs on `vr1-dcN-client`, not "the rack" (the rack as a distinct host no longer exists); row 2 + (maas/NetBox/tofu) stays on voffice1, UNCHANGED (D-128 Plane 2 survives -- MAAS region and + NetBox are not containment-layer objects); row 3 (never vcloud) UNCHANGED. +- **`:180-185`, the staged-scripts caveat** -- "the rack has NO repo clone" and needs + `~/repo-stage/scripts/` staged there. This caveat TRANSFERS verbatim to `vr1-dcN-client` + (still a non-repo-clone host by design, same staging discipline needed) -- no logic change, + just a re-pointed noun. +- **`:840-848`, the "considering destroying anything BELOW the juju layer" caution** -- names + "the containment VM, the DC's libvirt resources, MAAS records" and points at the (currently + Model-B-shaped) teardown runbook. **Needs updating to point at the REWRITTEN teardown + runbook (Section 1 above) and drop the containment-VM noun** -- the caution's PRINCIPLE (don't + reach for a substrate destroy while nodes are MAAS-enrolled) is unchanged. +- Path M and Path C content itself (the juju model/controller procedures phase4 delegates to + the teardown runbook) needs NO change here -- confirmed low-delta in Section 1. + +### 2.4 `runbooks/dc-dc-phase5-dr-failover-drill.md` + +- No `vvr1-dc`/`containment`/`qemu+ssh` hits (grep confirmed). Its only execution-host + reference worth flagging is INDIRECT: it delegates the per-DC deploy command to Phase 4 + (`:430-433`) and otherwise operates at the juju/openstack layer, which is D-138-scoped + (client VM under Option 1) rather than containment-scoped. **No direct edit required by + container-elim**; it inherits whatever Phase 4's table says. + +### 2.5 `runbooks/dc-dc-phase6-designate-cos-magnum.md` + +- **`:437-444`** -- "`juju` is ABSENT on vcloud ... Run them from the Plane-2 host (voffice1) + per D-128" and "**CHECK (read-only) -- from the Office1 headend (voffice1)**" immediately + precedes a `juju status -m ...` command. **This is ALREADY WRONG independent of the + container-elim** -- it predates phase4's own D-138 correction (juju runs on the DC's + execution host, not voffice1; `pass0-admin-report.md` and phase4's `:174-178` both confirm + `juju controllers` on voffice1 returns "No controllers registered"). Flagging here because + the SAME fix phase4 needed (rack -> client VM) applies to this line too, and both should be + corrected in the SAME pass since they share the root cause (a stale D-138 pointer). Not + itself a NEW container-elim delta, but a pre-existing defect this pass's rewrite sweep should + not leave behind. + +### 2.6 `runbooks/dc-dc-office1-service-reip.md` + +- Entirely Office1-scoped (`voffice1`'s own MAAS-composed LXD VMs -- NetBox, tailscale). No + `vvr1-dcN` reference; **out of this pass's scope** (D-114, not D-123). + +--- + +## 3. D-143 (re-IP) vs. container-elim: keeping the two axes distinguishable in the redeploy sequence + +D-143's OWED EXECUTION list (`docs/design-decisions.md:8163-8174`) is FIVE items, all +address-only: + +1. NetBox apex -- new B2 role `Cloud -- VR1 rebuild` owning `10.13.0.0/16`; re-carve VR1 + prefixes/VIPs/ranges under it, octet-preserving. +2. `scripts/lib-net.sh` -- flat defaults STAY 10.12; `vr1-dc0`/`vr1-dc1` arms get full 10.13 + literal blocks; fix the stale ":124-134" comment. +3. `10.13` naming-collision DOCFIX (`netbox/README.md:49`, a test literal). +4. D-124 transit routes re-point + rack statics shift (`10.12.8.2 -> 10.13.8.2`, etc.). +5. Teardown must REVOKE the dc0-substrate credentials before redeploy. + +**None of these five items touch topology** -- they are a value substitution (12->13) plus +one NetBox role-model addition. The container-elim change-set is orthogonal: it changes WHICH +objects exist and WHERE they run, not what numbers they carry. Concretely, keeping them +distinguishable in the redeploy plan means: + +- **Sequence within one redeploy, not two redeploys.** Since both changes ride the SAME + teardown+rebuild, the interleave is: (i) teardown the CURRENT 10.12 Model-B checkpoint + (Section 4 below) -- an ADDRESS-AGNOSTIC, TOPOLOGY-AGNOSTIC action (you're destroying + objects, not re-carving them); (ii) apply D-143 item 1 (NetBox apex) and item 2 (`lib-net.sh` + literal blocks) as the FIRST rebuild inputs -- these are pure data, no topology dependency, + so they can be done before or independent of the module rewrite; (iii) rebuild against the + Option-1 flat module design (this pass's Phase 2/4 deliverable), consuming the NEW 10.13 + literals from step (ii) -- the flat root's tfvars/module calls reference `lib-net.sh`'s + 10.13 arm the same way the old two-root shape referenced its 10.12 arm, so the flattening + and the re-IP compose cleanly (one substitutes VALUES, the other substitutes SHAPE) rather + than colliding. +- **Attribution discipline for the change-set (Phase 4's job, flagged here so it isn't lost):** + every diff in the rebuilt `opentofu/` tree should be traceable to EITHER "this line's value + changed because of D-143" OR "this line/module/step exists or doesn't because of + container-elim" -- not both folded into one undifferentiated rewrite. A rewritten Phase-2 + runbook that silently also renumbers octets (or vice versa) would make a future revert-one- + axis-not-the-other impossible to reason about. Concretely: the module RENAME/RESTRUCTURE + (e.g., `module.vvr1_dc0` disappearing, `module.vr1_dc0_client` appearing) is container-elim; + the VALUES inside surviving modules changing from `10.12.x` to `10.13.x` is D-143. A resource + that is BOTH new AND carries a 10.13 literal (e.g., the client VM's transit IP) is fine to + create once, but its commit/changelog framing should name both governing changes, not blur + them into "the redeploy." +- **D-143 item 5 (revoke dc0-substrate credentials before redeploy) interacts directly with + the container-elim's credential-residency migration** (`pass0-admin-report.md` row 8 / + Section 6 item 6): the OLD credentials being revoked are resident on the (about to be + eliminated) `vvr1-dc0`; the NEW credentials being minted for the rebuild are resident on + `vr1-dc0-client`. This is naturally sequenced (revoke old -> destroy old host -> build new + host -> mint new credential) but it means D-143 item 5 and the container-elim's SEC-028/ + SEC-029 re-pointing are the SAME step wearing two decision-labels -- call it out as one + action justified by both D-NNNs when it's executed, not two separate ones. + +--- + +## 4. Teardown of the CURRENT (Model-B, 10.12) live checkpoint -- order + +This is the FIRST action of the redeploy (dc0 is already at the confirmed FULL-deployment +checkpoint per the operator's pivot memory; dc1 is held). Per the teardown runbook's own +discipline (Section 1 above, still valid for tearing down the CURRENT Model-B shape -- the +runbook doesn't need to be rewritten to tear down what it already describes correctly): + +1. **Path M first, per DC with a live model** -- tear down the juju MODEL(s) before touching + substrate (`dc-dc-teardown-rollback.md` Path M, `:188-392`). This is the D-143 item-5 + credential-revoke's natural anchor point too (revoke, then destroy). +2. **D-143 item 5** -- revoke the dc0-substrate credentials (folded into or immediately after + step 1, per Section 3 above). +3. **Path C, if a controller needs to come down** with the model (only if the controller itself + is being retired, not just the model) -- `:394-596`. +4. **Path B substrate teardown, ordered INNER then OUTER, per DC** -- `:914-957`: from + voffice1, `tofu destroy` each `opentofu/vr1-dcN-substrate/` root; THEN, on vcloud, `tofu + destroy` the outer root's `vvr1_dc0`/`vvr1_dc1`/`vr1_dcN_uplink`/`vr1_dcN_storage` modules. + Step 2's MAAS census (`:635-747`) runs FIRST, per DC, before any destroy -- this is + unaffected by container-elim (still the load-bearing "attribute every `power_type=virsh` + record" gate). +5. **Mesh legs -- destroy ONLY if the whole VR1 layer (both DCs) is going away together** + (`:960-995`); a checkpoint-then-redeploy that keeps the SAME two DCs (new IPs, new shape, + same DCs) should almost certainly PRESERVE the mesh triangle and simply let the flat rebuild + re-target it -- destroying and recreating the mesh legs is extra churn with no benefit unless + the mesh module ITSELF needs to change shape for the flat topology (UNKNOWN pending Phase + 2/W2.1's tofu-module design -- the mesh legs are NOT named as changing in `pass0-admin- + report.md`'s "persists" list, Sec 1.3, so the working assumption is they survive; confirm at + Phase 2, don't destroy speculatively). +6. **Only after 1-5 converge (Step 5's untargeted `tofu plan` gate, `:894-911`, showing zero + unexpected drift)** does the rebuild begin -- against the Option-1 flat module design, + consuming D-143's new NetBox/lib-net.sh literals per Section 3's sequencing. + +**This order is UNCHANGED by container-elim** -- it is the teardown runbook's existing Path +M -> C -> B ordering, applied to what is CURRENTLY built (Model-B, 10.12). Container-elim only +changes what gets BUILT next (Section 2), not the order in which the current thing comes down. + +--- + +## 5. Open items this pass surfaces (not resolved here -- feed to Phase 2/4) + +1. **Root topology: one merged root per DC, or a shared outer + per-DC root split retained + minus the containment seam?** Section 1/2.1 above illustrate BOTH as plausible; Phase + 2/W2.1 (OpenTofu module design) is where this gets decided, not this planning pass. This + document's `tofu destroy` illustrations are DELIBERATELY marked illustrative for that + reason. +2. **The new teardown-primitive script (module-scoped group-destroy / `virsh destroy` loop) + does not exist yet** -- it is a NEW tested artifact this pass identifies as owed, to be + built and harnessed per repo discipline (CLAUDE.md "Delivery") once the module design lands. +3. **D-128's Plane-1/Plane-2 split is written entirely around the two-root shape** + (`docs/design-decisions.md:5354-5364`). Under Option 1 there is no more "Plane 2 = the + INNER tofu root run from voffice1" -- Plane 2 shrinks to MAAS/NetBox only (tofu is now + entirely Plane 1, vcloud). **This is a D-128 amendment this pass's Phase 4 [ARCH] framing + should fold in alongside the D-123 container-elim ruling itself** -- flagged here because it + was found while reading phase2/phase4's execution-host tables, not because W1.1 is ruling on + it. +4. **Artifact-service / rack-controller-remainder / D-131-forwarder placement remains OPEN** + (`pass0-admin-report.md` Section 6 item 3) and DIRECTLY determines the SSH-jump-target edit + in Section 2.2 above and the maas-vm-host re-homing in Section 2.1's Step-D row. This + pass's findings are written generically ("client VM OR maas-01") precisely because that + placement is not yet ruled -- do not let a later pass silently pick one without updating both + citations here. + +--- + +## 6. Verification note + +Every `path:line` cite above was read directly this session from `runbooks/dc-dc-teardown- +rollback.md`, `runbooks/dc-dc-phase0-vcloud-prep.md` through `phase6-designate-cos-magnum.md`, +`runbooks/dc-dc-office1-service-reip.md`, `opentofu/main.tf` (module list), and +`docs/design-decisions.md` (D-128, D-143). No value in this document is inferred from a prior +session's summary or from memory -- where a placement or a module name is not yet knowable it +is marked UNKNOWN with what resolves it, per hard rule 2. diff --git a/docs/audit/container-elim-pass/pass1-w2-workflow-gates.md b/docs/audit/container-elim-pass/pass1-w2-workflow-gates.md new file mode 100644 index 0000000..93165cb --- /dev/null +++ b/docs/audit/container-elim-pass/pass1-w2-workflow-gates.md @@ -0,0 +1,251 @@ +# Pass 1 -- W1.2: deployment-workflow doc + stage/gate structure vs the layered-module framing + +**Worker:** W1.2 (Phase 1, container-layer-elimination pass). **Date:** 2026-08-09. +**Scope:** `docs/dc-dc-deployment-workflow.md` (stage identity, gates, tooling gap register) and +the G-series gate register in `docs/CURRENT-STATE.md` section 6, read against the CONFIRMED +Phase-0 outcome (`pass0-admin-report.md`): **Option 1** (flat node VMs on vcloud libvirt + +per-DC non-hypervisor `vr1-dcN-client` VM), cross-DC isolation handling **(a)** (new vcloud-level +host-isolation control, new SEC-NNN row, mechanical `--check` gate), MAAS region unchanged +(`vr1-dcN-maas-01`). READ-ONLY. No inferred values -- every claim below cites path:line from +`docs/dc-dc-deployment-workflow.md`, `docs/CURRENT-STATE.md`, `docs/design-decisions.md`, or +`pass0-admin-report.md`. + +--- + +## 0. A framing correction before the analysis (measured, not inferred) + +The task brief's example -- "does the Stage-2 substrate / bootstrap-gate stage collapse" -- +names the wrong stage number. **Measured against the workflow doc:** + +- **Stage 2** (`docs/dc-dc-deployment-workflow.md:60-146`) is the **Office1 SITE standup** + (D-114): it builds `voffice1`, a *different* containment VM (MAAS-region + LXD 5.21-track + + MAAS-composed service VMs). Pass0 confirmed `voffice1` is explicitly **NOT container-layer** + (`pass0-admin-report.md:19,60`, "Office1 arm ... NOT container-layer"). Container-elim does + **not** touch Stage 2's shape. This distinction matters for module design later in this doc: + the project will carry TWO different "containment VM" patterns going forward -- one retired + (Model B `vvr1-dcN`), one kept (`voffice1`'s D-114 LXD-composition model) -- and a future + session must not conflate them. +- **Stage 3** (`docs/dc-dc-deployment-workflow.md:149-197`) is the actual **bootstrap-gate / + two-root Model-B DC-substrate stage** -- its own Build line says it verbatim: *"D-123 **Model + B** (TWO OpenTofu roots + a bootstrap gate between)"* (`:154`). **This is the stage + container-elim restructures.** All findings below key off Stage 3, not Stage 2. + +--- + +## 1. Stage-by-stage change table + +| Stage / Gate | Container-layer dependency (measured) | Option-1 change | Maps to (module layer) | +|---|---|---|---| +| **Stage 0** (decision ratification) | None | None | N/A (history) | +| **Stage 1** (`:36-57`, vcloud host prep) | None directly, but its 2026-07-10 as-built note (`:47-56`) shows the **six DC planes were originally built HERE, at vcloud level**, before D-123 later relocated them into the per-DC inner root (`opentofu/main.tf:22-33`, quoted at `pass0-w4-targets.md:79-80`: *"the 6 vr1-dc0 planes MOVED to the INNER root ... under Model B"*). Nested-KVM enablement here is a **shared prerequisite for two different purposes**: (i) `voffice1`'s LXD nesting (Stage 2, KEPT) and (ii) the now-eliminated per-DC containment VM's inner libvirtd (Stage 3, GONE). | Gate content is UNCHANGED (nested-KVM still needed for Stage 2's voffice1; pool/MTU/mesh gates untouched, D-100/D-101 own them regardless). **What changes is Stage 1's RELATIONSHIP to Stage 3**: under Option 1 the six per-DC planes land back at vcloud level via `modules/dc-planes`, the SAME level and SAME module family Stage 1 already establishes for the mesh/pools -- Stage 3 RECONVERGES onto Stage 1's own execution shape (single vcloud-local `qemu:///system` root) rather than diverging into it via D-123's later two-root split. | IaC module -- `vcloud-host-prep` (host-wide, invoked ONCE) | +| **Stage 2** (`:60-146`, Office1/voffice1) | **NOT a consumer** (pass0-confirmed; Section 0 above) | **NONE.** D-114's containment-VM pattern is architecturally distinct from D-123's Model B and is out of this pass's scope. Flag this explicitly in the Phase-4 change-set so it is not swept in by name-similarity ("containment VM" appears in both stages' prose). | Unaffected -- stays its own IaC module (`office1-site`) + procedure module (MAAS/LXD compose steps) | +| **Stage 3** (`:149-197`, per-DC OpenTofu substrate) -- **THE STAGE MOST RESTRUCTURED** | **Total.** Build line = Model B outer+bootstrap+inner (`:154`); Gate line = `vvr1-dc0` sizing/nested-KVM + "inner node VMs boot at depth-4 (D-114/D-123 boot gate)" (`:155`); Owns D-103-as-amended-by-D-123, D-122, D-123, D-124, D-125 (`:156`); Reuse-vs-new explicitly calls it NEW with "no OpenTofu/multi-rack precedent" BECAUSE of the two-root/bootstrap-gate mechanics (`:157`) | **Collapses to ONE tofu root, ONE apply, ONE execution host.** Per `pass0-w4-targets.md:69-76,97-111`: node VMs, planes, edge, and the new client VM all target vcloud's own `qemu:///system` DIRECTLY -- no `qemu+ssh`-from-Office1 dial for the substrate build at all. This means Stage 3's substrate creation shifts from a **two-plane execution stage** (outer root = Plane 1/vcloud; inner root + bootstrap gate = Plane 2/voffice1 via qemu+ssh, D-128 `docs/design-decisions.md:5354-5358`) to a **Plane-1-ONLY stage** -- the entire DC substrate is built from vcloud with local virsh/tofu, matching Stage 1's own execution shape. The "OUTER -> BOOTSTRAP GATE -> INNER" apply-ordering contract (`opentofu/main.tf:301-331`, cited `pass0-w4-targets.md:130-131`) disappears outright: no bootstrap gate script (`site-headend-install.sh --host-nodes` node-host mode retires, `pass0-admin-report.md:87`), no cross-root state, no D-126 per-env qemu+ssh key auth for the substrate dial. **Nesting depth 4 -> 2** (`pass0-admin-report.md` Section 4, `pass0-w4-targets.md:158`), so the D-114/D-123 "depth-4 boot gate" concept becomes a depth-2 boot gate -- a materially lower-risk check, closer to Stage 2's already-proven depth-3 (vcloud->voffice1->LXD) shape than to anything Stage 3 has attempted before. Stage 3's own "Reuse vs new" framing (`:157`, "NEW, no precedent") should be revisited in Phase 4: under Option 1 the stage is now MORE reusable (closer to Stage-1's flat-vcloud-libvirt shape) not less. **Owns line changes**: D-123 as currently phrased ("Model B site-down + two-root") is VOID under Option 1 and needs its Phase-4-framed amendment/new-D (per `SCOPE-AND-EXECUTION-PLAN.md:197-200`); D-125's "bridge-in" realization is retired (edge WAN -> direct NAT, `pass0-admin-report.md:84,169`), so D-125's Owns-line text needs updating alongside whatever D-number the container-elim ruling becomes. D-124's rack-transit addressing survives only for the CLIENT VM's transit leg, not for a rack-sizing purpose (D-124's own Stage-3 entry already flags its sizing clause VOID under Model B, `:180`; under Option 1 it stays void, now for a different reason -- no rack to size at all, only the small client VM). | IaC module -- `dc-substrate-flat` (planes + edge + 12 node VMs + client VM, ONE apply, invoked per DC) -- see Section 3 | +| **Stage 4** (`:200-212`, MAAS enlist/commission) | Indirect, via the rack-hosted services. `dc-rack-net.sh` (D-131 forwarder) and the per-DC artifact mirror/proxy (`.4`) run **on the containment VM today** (`pass0-admin-report.md` row 6-7, table). G17's dc0 check literal (`curl -fsS http://10.12.8.4/...`) targets that same host. | **Placement, not mechanism, changes -- and it is UNRESOLVED at Phase 0** (`pass0-admin-report.md` Section 6 item 3: "client VM / maas-01 / retire-with-evidence", carried into Phase 1 as an open item). Whichever placement is ruled, Stage 4's Gate-line prose (`:206`, "per-DC artifact source answering on its own address") and G17's check literal both need the new address substituted -- **this is a genuine intersection with D-143's 10.12->10.13 re-IP**: G17 will need BOTH edits (new address family per D-143, new HOST per container-elim) in the same pass, but they are two distinguishable causes per `SCOPE-AND-EXECUTION-PLAN.md:194-196`'s own instruction, and this doc should record them as two line-items even though they land in one edit. Stage 4's own commission/deploy mechanics (`:206-209`) are otherwise UNCHANGED -- node VMs still PXE/commission the same way regardless of hypervisor nesting depth. | Procedure module -- `maas-enlist-commission` (per-DC); its config INPUT (mirror/forwarder host) changes, its STEPS do not | +| **Stage 5** (`:215-230`, Juju + bundle) | **Structural, via D-138.** D-138 (`docs/design-decisions.md:7118-7180`, RULED 2026-07-30) already moved the cloud-facing client (Juju CLI, `openstack` CLI) INTO the DC because `voffice1` has no L3 path to any DC node plane and SEC-010/D-052/D-100 forbid opening one. The concrete host D-138 names is `vvr1-dc0` (the containment VM) -- it is currently the D-138 execution host by NECESSITY, not by a separate design choice. | **The D-138 PRINCIPLE is preserved, its concrete host changes.** Option 1's `vr1-dcN-client` VM is explicitly built to BE the D-138 host (`pass0-admin-report.md:163-169`, "carrying: the D-138 client role"; `pass0-w4-targets.md:97-111` confirms the client VM occupies exactly D-138's required shape -- metal-admin + transit legs, an L3 presence). **No Stage-5 gate content changes** -- `preflight.sh` PASS, post-deploy `cloud-assert.sh --capture`, controller backup, geneve-over-v6/Ceph-over-v6 verification (`:221`) are all unaffected by WHICH VM the Juju client runs from. What DOES change: every literal address/host-alias in Stage-5 runbooks and CURRENT-STATE's version-pins table that names the containment VM's transit IP (e.g. `docs/CURRENT-STATE.md:7829`, "openstackclient ... INSTALLED ON THE dc0 RACK (172.31.0.2)" -- that IP is `vvr1-dc0`'s transit address) must be re-pointed to the client VM's new address once it exists. **D-140 constrains this stage's module mapping directly** (Section 4 below) -- the Juju layer itself stays a PROCEDURE module (scripted, bundle-based) for this redeploy by an ALREADY-RULED pin (`docs/design-decisions.md:7781-7787`), not a design choice this pass is free to make. | Procedure module -- `juju-bundle-deploy` (per-DC); execution-host param changes (vvr1-dcN client VM), D-140 keeps it out of IaC scope for now | +| **Stage 6** (`:233-244`, DR/failover) | **None found.** Ceph replication, radosgw multisite, rbd-mirror all run on deployed OpenStack units, not on the containment layer. | None. | Procedure module -- `dr-failover-drill` (unaffected) | +| **Stage 7** (`:248-259`, Designate/COS/Magnum) | **None found.** | None. | Procedure module -- `designate-cos-magnum` (unaffected) | +| **Cross-cutting: branch-per-stage/merge-at-close discipline** (`:424-436`) | Indirect -- Stage 3's branch would carry the container-elim change-set as part of its definition-of-done | Stage 3's branch (whatever it is named for the 10.13 rebuild) is where the module collapse actually lands; its close-out sweep must also reconcile `.claude/skills/openstack-cloud-ops/` for the new invariant (per `:430-433`'s standing rule) | N/A -- process discipline, not a module | +| **Teardown runbook** (`dc-dc-teardown-rollback.md`, gap #19a, `:958-975`) | **Direct** (pass0 row 9: "`vvr1-dcN` as the teardown unit; D-122's 'site-down = one `virsh destroy`'") | Re-authored around a scripted group-destroy of the flat `vr1-dcN-*` domain set OR, better, a **module-scoped `tofu destroy -target=module.vr1_dc0`** now that the whole DC is ONE module call under the single flat root -- this would PARTIALLY re-earn the lost D-123 one-command convenience via tofu itself rather than virsh scripting (flagged for Phase 4, pass0 Section 6 item 7 already names this as owed design work) | IaC-module-scoped teardown, one call per DC module | + +--- + +## 2. G-series gate register: which are keyed to the container layer + +Read against `docs/CURRENT-STATE.md` section 6 (`:7786-7813`) in full. + +- **G9** (DC0 outer apply, `:7803`) -- keyed to the two-stage apply by its own SEC dependency + note: *"SEC-010's transit FORWARD-drop is applied+verified at deploy step B via + `site-headend-install.sh --host-nodes --check` on vvr1-dc0 (gate G10)"*. This gate is CLOSED + (historical, dc0's checkpoint build) -- it does not reopen, but it is the **template** the + next DC-substrate apply gate (whatever supersedes it for the 10.13 rebuild) must NOT copy + verbatim: under Option 1 there is no outer/inner split for it to key off, so a redrawn gate + is a SINGLE apply-and-verify, not a two-step G9+G10 pair. +- **G10** (`:7804`, deploy steps B-E) -- the sub-items most exposed: + - *"SEC-010 `--host-nodes --check` on vvr1-dc0"* -- this exact check RETIRES with + `site-headend-install.sh`'s node-host mode (`pass0-admin-report.md:87`). Its REPLACEMENT is + the new cross-DC host-isolation control's own `--check` (the confirmed Phase-0 gate handling + (a), `pass0-admin-report.md:288-291`) -- but that control is scoped to the WHOLE vcloud host + (both DCs' co-residency), not to a single DC's apply step, so it likely does NOT slot into a + per-DC G10-style gate at all; it is closer in shape to a Stage-1-level, apply-once gate (see + Section 4). This is a genuine **stage-home question for Phase 1/2 design**, not resolved by + this worker's dimension alone -- flagged forward. + - *"depth-4 nested boot"* -- becomes depth-2 (Section 1, Stage 3 row). + - *"D-125 foreign-MAC egress test"* -- the MECHANISM changes (no bridge-in to test the + isolation of), but an analogous direct-NAT egress-isolation assertion is still owed; do not + read this as "the egress gate goes away", only that what it tests changes shape. + - *"MAAS reachability + `TF_VAR_maas_api_key`"* and *"netem placeholder"* -- both UNAFFECTED + (Plane-2/mesh concerns, not containment-layer concerns). +- **G12** (`:7806`, vr1-dc1 build) -- CLOSED/historical, but it is the **fullest worked example** + of what a per-DC Stage-3 gate sequence looks like today (outer apply -> Step B SEC-010 check + -> inner apply -> rack standup -> commissioning). Its structure is the thing that collapses; + future sessions writing the reshaped gate should read G12's own text as "the shape being + replaced," not as a template to repeat. +- **G17** (`:7811`, per-DC artifact source reachable from a node) -- **OPEN today**, and its + check literals directly name the containment-VM-hosted mirror address (`10.12.8.4` for dc0). + As covered in Section 1's Stage-4 row: this gate's eventual close needs BOTH the D-143 + address-family edit and the container-elim placement edit, kept distinguishable per the + SCOPE plan's instruction. +- **G18** (`:7813`, lb-mgmt IPAM apex) -- **no container-layer dependency found**; Octavia's + charm-owned lb-mgmt network is unrelated to the substrate hypervisor topology. Named here only + to confirm it was checked and correctly excluded (a "manufactured contradiction" risk this + repo's own instrument-currency record warns about -- do not let G18 get swept into the + container-elim change-set by proximity in the same table). +- **G1-G8, G11, G13, G15, G16** -- read; **no container-layer dependency found** in any of them + (audit/ratification gates, D-068/D-071 rulings, office1 edge state-surgery gates). Confirmed + by direct read of `:7795-7802,7805,7807,7809-7810`, not assumed from gate NAME. +- **G14** (`:7808`, open SEC rows) -- indirect only: container-elim MIGRATES several credential + residencies (SEC-026/-028/-029, `vm-secret-locations` `rack` rows, per pass0 row 8) and ADDS + at least one new SEC-NNN row for the cross-DC isolation control -- both are G14 COUNT-affecting + events when they land, not gate-SHAPE changes. Flag for whoever next runs `ledger-scan.sh` + after this pass's changes land, so the count is not read as unexplained drift (the project's + own instrument-currency lesson #25 applies directly here). + +**Net finding:** exactly **one G-series gate concept is directly restructured (G9/G10's DC-apply +sequence)**; **one is a placement-pending intersection with D-143 (G17)**; the rest are either +unaffected or affected only as a downstream credential-count/ledger consequence (G14). No gate +needs to be DELETED outright -- G9/G10's successor for the 10.13 rebuild is a simplified +redraw, not a removal, since a substrate apply + SEC verification + boot proof + egress test are +all still real checks, just fewer steps and one execution host instead of two. + +--- + +## 3. Tooling gap register -- what closes, reshapes, or is added + +Read against `docs/dc-dc-deployment-workflow.md:447-1354` in full. + +- **Gap #2 (OpenTofu, `:499-578`)** -- RESHAPES. The still-open line *"`tofu plan`/`apply` has + NOT been exercised for the Stage 3 DC substrate"* (`:537`) changes SCOPE under Option 1: the + first real exercise of Stage-3 apply becomes a single flat-root apply, not an outer+bootstrap + +inner sequence. The module inventory itself shrinks: `modules/wan-bridge` (D-125 bridge-in) + is DELETED (pass0 row 3); the inner-root files (`opentofu/vr1-dc0-substrate/`, + `vr1-dc1-substrate/`) RETIRE as separate roots, their module CALLS folding into the outer/flat + root (pass0 row 2). `modules/maas-vm-host`'s target changes from `vvr1-dc0`'s inner virsh to + vcloud's own virsh directly (`pass0-w4-targets.md:109-111`) -- this closes the DOCFIX-179 + deferral note this gap still carries (`:553-554`, "deliberately NOT `maas_vm_host_machine`") + differently than originally anticipated, since there is no longer an inner virsh to register. +- **Gap #17 (per-site ISP uplink, `:857-933`)** -- STAYS CLOSED, but its "how it was actually + closed" mechanism description (`:894-904`, the bridge-in shape: outer NAT + inner bridge onto + `br-vr1-dcN-wan`) becomes HISTORICAL under Option 1's direct-NAT attachment + (`pass0-w4-targets.md:88-93`, "exactly like Office1's edge does today ... Model A's item 8"). + Flag so a future reader does not rebuild the retired bridge-in mechanism by reading this gap's + closing note as current practice; a doc-currency addendum is owed here at the point the + container-elim ruling lands, not a reopening of the gap itself. +- **Gap #19b (`$DC` region-qualified namespace, `:976-1038`)** -- UNAFFECTED. D-119's + region-qualified selectors (`vr1-dc0`/`vr1-dc1`) are an IPAM/naming-layer fix, orthogonal to + hypervisor nesting depth. +- **Gap #20 (host->site reachability / `site-baseleg.sh`, `:1039-1109`)** -- **verdict likely + still holds, but the reasoning chain it rests on shifts and should be RE-VERIFIED, not + carried forward silently.** Today's "no leg required on vcloud" verdict (`:1075-1093`) rests on + D-128: *"no vcloud-originated DC operation exists ... Plane 2 ... EXECUTES on `voffice1`"*. + Under Option 1 the SUBSTRATE build itself now originates from vcloud directly via a local + `qemu:///system` provider (Section 1, Stage 3 row) -- but that is Plane 1 (host-VM creation), + the SAME category D-128 already assigns to vcloud today for the outer root, not a new + vcloud-originated L3 dial into a DC's own MAAS/Juju/openstack layer (which stays Plane 2, via + `voffice1`, or D-138's client VM). The verdict's PREMISE is therefore probably unchanged, but + this gap's own text explicitly warns *"If a future change makes vcloud originate to a DC + directly, this verdict expires and branch 1 applies"* (`:1092`) -- container-elim is exactly + such a change to the SUBSTRATE layer, even if not to the Plane-2 boundary, so a fresh + re-measurement (not an inference from today's text) is owed once the flat root is built. Named + here as a Phase-2/3 tooling-review item, not resolved by this worker. +- **Gap #21 (per-DC Tailscale, `:1189-1310`)** -- **UNAFFECTED in shape.** The tailscale router + VM (`vr1-dcN-tailscale-01`, utility `.7`) is already one of the 12 flat node-VM siblings inside + today's inner root (pass0 table row 2's "12 node VMs") and folds into the flat root exactly + like every other node VM -- same re-homing mechanism, no gap-content change. Confirmed + explicitly so it is not mis-swept into the container-elim change-set by proximity (it is a + D-129(iii) item, not a D-123/Model-B item). +- **Gap #22 (ceph-rbd-mirror HA, `:1310-1354`)** -- **NOT a consumer.** Confirmed by direct read; + this is a D-108 application-layer HA question, unrelated to substrate nesting. Named to record + it was checked, per this repo's "a clean negative still needs the check shown" discipline. +- **New gap to REGISTER (not yet in the doc):** the cross-DC adjacency control itself + (`pass0-admin-report.md` Section 5, handling (a)) has **no register entry today** because it + is a Phase-0-confirmed-but-not-yet-designed deliverable. It should land as a NEW numbered item + in this register (next-free per the register's own numbering, distinct from the D-NNN/SEC-NNN + numbering it also needs) once Phase 1/2 design work produces its concrete shape, so it is + trackable the same way every other gap here is. + +--- + +## 4. The current stage/gate flow re-expressed as layered modules (first sketch) + +Per `SCOPE-AND-EXECUTION-PLAN.md:40-43`, "module" means BOTH IaC (OpenTofu) modules for the +substrate AND procedure/runbook modules for the orchestration layered on top. This sketch maps +today's Stage 1-7 flow onto that split, incorporating Option 1 and the (a) isolation control. + +``` +IaC-MODULE LAYER (OpenTofu, vcloud-executed, Plane 1 -- D-128) + [1] vcloud-host-prep (Stage 1) -- ONCE. pools, mesh triangle, MTU, nested-KVM base. + [2] office1-site (Stage 2) -- ONCE. voffice1 containment VM (D-114, UNCHANGED + by this pass -- kept as its own IaC module, its + LXD-compose half stays a procedure step inside it). + [3] cross-dc-isolation (NEW) -- ONCE, vcloud-level. The Phase-0-confirmed (a) + deliverable: nftables/isolation artifact + + mechanical --check + new SEC-NNN row. STAGE-HOME + OPEN (Section 2): candidate slots are (i) folded + into [1] as a host-prep-time control, applied once + and re-verified whenever a second DC lands, or + (ii) its own micro-module invoked once, before the + first per-DC substrate apply. Phase-1/2 design item, + not resolved here. + [4] dc-substrate-flat (Stage 3) -- PER DC (dc0, dc1, futureN). The collapsed module: + 6 planes + edge + 12 node VMs + client VM, ONE + apply. Directly supersedes today's + outer+bootstrap-gate+inner three-part sequence. + Gate: single apply-and-verify (Section 2) -- + SEC check via [3], depth-2 boot, direct-NAT egress + isolation test. + +PROCEDURE-MODULE LAYER (scripted/runbook, mixed Plane 1/2/DC-local execution) + [5] maas-enlist-commission (Stage 4) -- PER DC. Unaffected mechanically; artifact-mirror/ + rack-controller/D-131-forwarder PLACEMENT is an + input this module now takes as a parameter (client + VM vs maas-01 vs retire -- pass0 Section 6 item 3), + not a hardcoded containment-VM address. + [6] juju-bundle-deploy (Stage 5) -- PER DC. D-138's client role executes here, now on + the [4]-created client VM instead of the retired + containment VM. STAYS a procedure module, not an + IaC module, by the ALREADY-RULED D-140 pin (Section + 1, Stage 5 row) -- this pass's module design must + not silently fold Juju into [4]. + [7] dr-failover-drill (Stage 6) -- PER DC (or cross-DC pair). Unaffected. + [8] designate-cos-magnum (Stage 7) -- PER DC. Unaffected. + +TEARDOWN (mirrors the IaC layer, module-scoped) + [T] tofu destroy -target=module. -- PER DC, one call per [4] module instance -- the + candidate re-earning of D-123's lost one-command + site-down (Section 1, teardown row). +``` + +**Sequencing note for W4.3 (not this worker's dimension, flagged forward):** `[3]` cross-DC +isolation logically needs to exist and be verified BEFORE the second DC's `[4]` apply lands +(otherwise dc0-alone co-residency has nothing yet to leak into, but dc0+dc1 co-residency does) -- +this is a real ordering constraint on the module-invocation sequence Phase 4 designs, not merely +a stage-numbering question. + +--- + +## 5. Top risks for this dimension + +1. **The (a) isolation control has no stage home yet** -- it is a confirmed Phase-0 DESIGN + REQUIREMENT with no gate row, no stage slot, and no register entry (Section 3, "New gap to + register"). If Phase 1/2 do not give it one explicitly, it risks being built ad hoc and + under-gated relative to the SEC-010 pattern it is explicitly modeled on. +2. **G17's dual-cause edit (D-143 address + container-elim placement) is easy to under-scope** + if a future session only remembers the address-family half; this doc names both causes + explicitly so neither is dropped. +3. **D-140 is an easy module-design trap** -- the layered-module goal's own IaC-modules-for- + substrate framing could tempt folding Stage 5 (Juju) into the IaC layer "while we're at it"; + D-140 already RULES that OUT for this redeploy (pinned to end-of-deployment review). Phase 4 + must treat Stage 5 as a procedure module, full stop, for the 10.13 rebuild. +4. **Stage 3's "Owns" D-number list is stale the moment Option 1 is ruled formally** (D-123, + D-125 both need their Phase-4-framed amendment text); this doc's Section 1 table can serve as + the input list for that edit but does not itself perform it (out of scope, GA-R5 territory). + +--- + +## 6. Sources consulted (path:line, for the adversarial-check pass) + +`docs/dc-dc-deployment-workflow.md` (full read, 1354 lines); `docs/CURRENT-STATE.md:7786-7835` +(section 6 gate register + section 7 version pins); `docs/design-decisions.md:7118-7180` (D-138 +full text), `:7781-7787` (D-140 full text), `:8083-8143` (D-143 excerpt); `pass0-admin-report.md` +(full read); `pass0-w4-targets.md:1-220` (full read, targeted at execution-host/plane detail not +fully covered in the admin report's synthesis). diff --git a/docs/audit/container-elim-pass/pass1-w3-sequencing.md b/docs/audit/container-elim-pass/pass1-w3-sequencing.md new file mode 100644 index 0000000..32368a6 --- /dev/null +++ b/docs/audit/container-elim-pass/pass1-w3-sequencing.md @@ -0,0 +1,462 @@ +# Pass 1 / W1.3 -- Teardown -> Redeploy SEQUENCING for the flat (Option 1) topology + +READ-ONLY planning artifact. Container-layer-elimination pass, Phase 1, Worker 3. +Confirmed Phase-0 outcome consumed as ground truth (`pass0-admin-report.md` Section 7a): +**Option 1** (flat node VMs on vcloud libvirt + a small per-DC `vr1-dcN-client` VM, +D-138 shape); cross-DC handling **(a)** (accept co-residency + a new vcloud-level host +isolation control, DESIGN ITEM, not yet built); MAAS region stays on `vr1-dcN-maas-01`; +rack-controller remainder / D-131 forwarder / artifact-service placement is **OPEN**, +carried here as an unresolved slot in the sequence, not invented. + +No mutation performed or proposed as executable here -- this is the ORDERED PLAN only. +Tags per item: `[vanishes]` `[reorders]` `[new]` `[unchanged]`, and `[D-143]` / +`[container-elim]` / `[both]` for the two riding axes (SCOPE-AND-EXECUTION-PLAN.md +Section 7's "keep the two changes distinguishable" instruction). + +--- + +## 0. Two teardown questions, kept separate + +The worker prompt asks two different things that must not be conflated: + +- **0.A -- Teardown of the CURRENT 10.12 Model-B checkpoint.** This tears down what is + ACTUALLY DEPLOYED today: the two-root, containment-VM, qemu+ssh shape. It uses TODAY's + tooling (`runbooks/dc-dc-teardown-rollback.md`) because that is what matches the live + state -- the flat-topology design does not change what commands tear down the OLD + shape. This is D-143-tagged work (the checkpoint-then-redeploy pivot), not a + container-elim mechanism in itself. +- **0.B -- Where D-122's one-command site-down PRIMITIVE gets replaced going forward.** + This is a container-elim design question: once the flat topology is BUILT, what does + ITS OWN future teardown/DR primitive look like (Phase 4's module-scoped destroy). It is + answered here as a target-state design note, not as an action taken during 0.A. + +--- + +## PART A -- TEARDOWN of the current 10.12 checkpoint (Section 0.A) + +Baseline: `runbooks/dc-dc-teardown-rollback.md` Steps 1-5 (Path A, per-DC) / Path B (both +DCs), read in full this session. This IS today's two-root shape; most steps below are +`[unchanged]` mechanically -- the D-143/container-elim axes contribute exactly TWO new +steps the runbook does NOT yet contain: A.3 (the R7 credential-revocation checklist, +confirmed absent by the admin report Section 6 item 6 and +`open-items-review-20260809.md` R7/R15) and A.4 (the MAAS machine-record release/delete +that Step 2's own gate implies but the runbook only states as a contingency, not a +sequenced step -- advisor-flagged, see below). + +1. **[unchanged][D-143]** Back up state, BOTH roots, BOTH hosts (teardown-rollback.md + Step 1): outer `opentofu/terraform.tfstate` on vcloud; each DC's INNER + `opentofu/vr1-dcN-substrate/terraform.tfstate`, which lives on the **Office1 headend** + (D-128 Plane 2) per the runbook's own callout (`:608-615`) -- a backup that only + touched the vcloud copy has not backed up the substrate about to be destroyed. +2. **[unchanged][D-143]** MAAS-side machine census FIRST, from the Office1 headend + (Step 2, `:635-694`): LENS 1 (enumerate every `power_type=virsh` record and attribute + it) + LENS 2 (corroborate against `lib-hosts.sh`'s pinned boot-MAC roster for the + site). GATE: both lenses succeed and LENS 2 returns zero unattributed records. +3. **[new][container-elim, D-143]** **R7 credential-revocation checklist (owed build, + `open-items-review-20260809.md:80,398,431`; `design-decisions.md:8174` D-143 owed-item + 5).** Run BEFORE the substrate destroy (the credential-bearing hosts must still be up + and reachable to revoke cleanly; revoking after they are gone is a "trust it expired" + guess, not a verified revocation). Per the admin report's credential-residency + inventory (Section 2 row 8) and the security-ledger rows read this session: + - **SEC-026** (`security-ledger.md:79`): the Juju client credential + (`juju-vr1-dc0-cred`/`juju-vr1-dc1-cred`, SEC-018/019) resident on the DC client + host (today: `vvr1-dcN` itself, since D-138 put the client inside the containment + VM). Confirm `juju credentials --client` on that host, then REVOKE/rotate the + underlying MAAS admin-scoped key per SEC-026's own rotation-trigger note ("if the DC + client host is rebuilt or shared" -- this teardown IS that trigger). + - **SEC-028** (`security-ledger.md:81`): the per-DC `juju-vr1-dcN` SERVICE credential + minted ON the region VM and distributed to the rack (`~/vr1-dcN-creds/`, 0600). + Revoke/rotate per its own OPEN rotation obligation -- "at v1 close, or immediately if + the rack or region VM is rebuilt" (this teardown rebuilds both). + - **SEC-029** (`security-ledger.md:82`): the Octavia PKI overlay COPY on the rack + (`~/repo-stage/overlays/vr1-dcN-octavia-pki.yaml`). Not a mint (SEC-004 is the + authority) -- shred the rack-local copy; no upstream rotation needed for the copy + itself, but confirm the headend's canonical copy is unaffected. + - **D-126 per-env SSH keypair** (`~/vr1-dc0-creds/`, `~/vr1-dc1-creds/` private halves; + `opentofu/variables.tf:105-114,158-168`): authenticated the now-retired qemu+ssh + inner-provider dial. Revoke/shred -- this credential class has NO successor under + container-elim (Part B has no qemu+ssh dial at all), so this is a clean retirement, + not a rotation. + - **MAAS rack enrollment secret** consumed at the rack's original enrollment (one-shot, + already spent) -- confirm no residual copy sits readable on the rack filesystem. + - Register updates riding the same pass (not new mints, bookkeeping): the checklist's + BUILD should enumerate from EVERY `vm-secret-locations` row keyed to the + `rack`/`vr1-dcN-client` host class, not only the three SEC rows named above by + example -- the register-can't-see-a-row-it-doesn't-have discipline (this repo's own + rule, `creds-folder-convention` memory pointer) applies here too: mark each such row + RETIRED, not deleted (append-only discipline). + GATE: every row above individually confirmed revoked/shredded/rotated, with the + confirming command's output captured (not asserted from memory). +4. **[new][D-143]** **MAAS machine-record release/delete [MUTATION, operator-gated per + record class].** Step 2's LENS 2 will return NON-ZERO by construction on a live + commissioned checkpoint -- the runbook's own text treats this as a contingency + ("if records DO exist and you intend to lose them... release or delete them via MAAS's + own documented flow FIRST, then re-run both lenses," `:686-690`), but for THIS teardown + it is a guaranteed sequenced step, not an edge case. Release/delete every attributed + record for the site (individually operator-approved, MAAS's own current documented + flow, never invented here), then RE-RUN both lenses and confirm LENS 2 now returns + zero -- only then does the GATE in Step 2 actually close. Also clean the REGION-SIDE + residue that goes stale once the rack host is destroyed and is otherwise easy to miss: + the `vvr1-dcN` rack-controller's own enrollment record on the Office1 region (it + enrolled via the one-shot enroll-secret, `site-headend-install.sh:25-29`), and the + region's `primary_rack`/DHCP-on-metal-admin reference to it. Flag both for confirmation + as part of this step, not left implicit. +5. **[unchanged][D-143]** Pick the root, INNER first (Step 3, `:751-836`): plan + `-destroy` in `opentofu/vr1-dcN-substrate/` from the Office1 headend; capture the + pre-destroy `virsh -c "$VIRSH_POWER_ADDRESS" list/net-list/pool-list --all` baseline + while the containment VM is still reachable. +6. **[unchanged][D-143]** Pick the root, OUTER second (Step 3 cont'd, `:838-882`): plan + `-destroy -target=module.vvr1_dcN -target=module.vr1_dcN_uplink + -target=module.vr1_dcN_storage` on vcloud. Do NOT target the mesh-link/netem modules + (shared with the surviving/adjacent site legs) -- carried forward unchanged into the + redeploy since the mesh triangle persists per Part B step B.4. +7. **[unchanged][D-143]** Apply, INNER then OUTER (Step 4, `:838-935`): `tofu destroy + teardown-vr1-dcN-inner.tfplan` from the Office1 headend, THEN `tofu destroy + teardown-vr1-dcN-outer.tfplan` on vcloud. Verify each half from the host that can + actually see it (inner via the containment VM's virsh BEFORE it is gone; outer via + vcloud's own virsh after) -- the runbook's own trap: a vcloud-local `virsh list` never + saw the DC's inner objects even when they were fully intact. +8. **[unchanged][D-143]** Confirm no drift outside the torn-down scope (Step 5, untargeted + `tofu plan` on the outer root; equivalent check on the surviving DC's inner root if one + exists). GATE: zero unexpected drift. +9. **[new][D-143]** NetBox DCIM decommission of the `vvr1-dcN` device record(s) (admin + report Section 2 row 11, `netbox/dc-rack-mgmt-import.py`) -- a system-of-record data + migration, not a code change; run against the LIVE NetBox after the substrate is + confirmed gone. +10. **[unchanged]** Repeat 1-9 per DC (Path A) or run the Path B batched form + (`teardown-rollback.md:914-958`) for both DCs at once if tearing down together; either + way the mesh-link teardown note (`:960-998`) applies only if BOTH ends of a leg are + going away -- confirm against Part B before deciding whether the mesh triangle itself + is torn down or kept live through the pivot (Part B step B.4 assumes it is KEPT). + +**A.10 -- where does D-122's one-command site-down get replaced, for THIS teardown.** It +doesn't, in Part A -- Part A tears down the shape that HAS the one-command primitive (a +single `virsh destroy vvr1-dcN` is still literally available here, per D-122 AMENDMENT +2026-07-16, `design-decisions.md:4838-4842`), even though this runbook does not use it +raw (module-scoped `tofu destroy` is the reviewed, gated path; a raw `virsh destroy` +bypasses state and is the exact class of error CLAUDE.md hard rule 4 was added to +prevent, 2026-08-03 incident). **The primitive's REPLACEMENT is a target-state property of +the NEW flat topology, not of this teardown** -- see Part B.0 below. + +--- + +## PART B -- REDEPLOY on flat 10.13 (Option 1) + +Numbered against TODAY's chain (worker-prompt's own framing, confirmed by +`pass0-w1-substrate-map.md` Section 3): **outer-apply -> bootstrap-gate +(`site-headend-install.sh --host-nodes`) -> inner-apply (qemu+ssh, from voffice1) -> +MAAS enlist/commission -> juju deploy.** Runbook baseline for the surviving mechanics: +`runbooks/dc-dc-phase0..phase6`, `docs/tool-index.md`. + +### B.0 -- the site-down primitive's replacement (design note, answers Part A.10) + +Under Option 1 there is no single VM whose destruction equals "the DC is gone" -- the DC +is now N flat sibling domains (planes x6, edge, 12 node VMs, 1 client VM) on vcloud's own +libvirt/state, the same shape `docs/design-decisions.md:4903-4914`'s Model A already +named: **"Site-down DR = destroy the `vr1-dcN-*` domain group... as a scripted op (owned +by `dc-dc-teardown-rollback.md`), NOT a single `virsh destroy`."** Concretely this becomes +**ONE module-scoped `tofu destroy -target=...` (or a full-root destroy once the DC is the +whole root/state) against the flat root**, listing every `vr1_dcN_*` module in one +plan/apply pair -- functionally Part A's steps 5-6 OUTER half alone (the runbook's own +Step 3's OUTER sub-plan), because there is no separate inner half left to sequence around +it. This re-earns the "one scripted command" +property (not literally "one `virsh destroy`") and is exactly what admin report Section 6 +item 7 flags as owed to Phase 4's module design -- **noted here as the target shape, not +built here.** + +### B.1 -- capacity + IPAM prerequisites (run once, ahead of any apply) + +1. **[unchanged][both]** vcloud host-capacity re-measurement (admin report Section 8 item + 1) -- `dc-dc-whole-host-budget.py`'s committed 256 vCPU / 1024 GiB / 10240 GiB needs a + fresh read before any FIT verdict for the flat 12+1-VM/DC roster (item 2, calculator + currently lacks the 3 utility-node flags -- extend or hand-total first). Feeds BOTH + axes: D-143 needs it for the rebuild sizing, container-elim needs it for the + "~176 GiB freed" claim to stop being directional-only. +2. **[unchanged][D-143]** NetBox apex re-carve: mint the B2 role `Cloud -- VR1 rebuild` + owning `10.13.0.0/16`; re-carve VR1 prefixes/VIPs/ranges octet-preserving + (`design-decisions.md:8165-8166`, D-143 owed-execution item 1). Pure IPAM, topology- + agnostic -- runs whether the topology is flat or nested. +3. **[unchanged][D-143]** `scripts/lib-net.sh`: keep flat defaults at 10.12 (vr0-dc0 no- + op preserved); give `vr1-dc0`/`vr1-dc1` full 10.13 literal blocks; update the now-false + `:124-134` inherits-VR0 comment (owed-execution item 2, F13). +4. **[both]** D-124 transit routes re-point + the DC-side static IP shifts 12->13 + (owed-execution item 4). Under container-elim the ADDRESS this step touches changes + HOST from `vvr1-dcN`'s rack-transit IP to the new `vr1-dcN-client` VM's transit leg -- + same octet-preserving math (D-143), different bearer (container-elim). Tag `[both]` + because neither axis is separable at this one line item; call it out explicitly in the + change-set so a reviewer does not assume it is pure D-143. + +### B.2 -- Office1 + vcloud prep (persists, off the container-layer entirely) + +5. **[unchanged]** vcloud host prep (Phase 0 runbook) -- OpenTofu reaches vcloud libvirt, + host-level prerequisites. Not container-layer; runs regardless of target topology. +6. **[unchanged]** Office1 headend standup (Phase 1 runbook), IF not already up from the + dc0 checkpoint -- voffice1 remains the MAAS region host (D-132 addendum unaffected) and + remains the D-128 Plane-2 execution host for whatever STILL needs a remote dial + (mesh-link tofu, if any; NOT the DC substrate apply any more -- see B.4). + +### B.3 -- the new cross-DC isolation control (container-elim design item, gate-confirmed) + +7. **[new][container-elim]** Build + install the vcloud-level host isolation control + (admin report Section 5 handling (a), CONFIRMED at the Phase-0 gate, Section 7a): an + nftables/isolation artifact asserting no inter-plane/inter-DC forwarding on vcloud + itself, with a mechanical `--check` gate and its own new SEC-NNN ledger row (SEC-010's + pattern, one layer up). **MUST run BEFORE step B.4** -- the moment two DCs' plane + bridges share vcloud's one kernel (the instant the flat apply creates the second DC's + planes) is the moment the gap becomes live; the control has to already exist, not be + retrofitted after the fact. This is a genuinely NEW artifact with no direct precursor + in today's chain (SEC-010 was interface-scoped to the transit leg, a different boundary + -- see admin report Section 1.4/Section 5 for why it does not already cover this). + **This is ONE of TWO SEC-010-successor controls, not the whole story** -- the + re-authored transit-leg FORWARD-drop itself is a SEPARATE, still-OPEN item, sequenced + at B.5 step 17; do not read this step as closing SEC-010's full scope. + +### B.4 -- the (formerly) outer + bootstrap-gate + inner sequence, COLLAPSED + +8. **[reorders][container-elim]** **ONE FLAT `tofu` root, ONE apply cycle**, run FROM + VCLOUD directly on `qemu:///system` (no qemu+ssh dial exists any more -- the + `opentofu/vr1-dcN-substrate/` roots are RETIRED as separate roots; their module bodies + -- planes x6, WAN, edge, 12 node VMs -- are UNCHANGED HCL, re-homed into the flat + root/provider per `pass0-w1-substrate-map.md` Section 2.3's own observation that none + of those modules hardcode "vvr1-dcN"). Per DC, this single apply creates: + - `[unchanged]` mesh triangle legs, uplink NAT `/24` (D-125 addressing/IPAM identity + untouched, `design-decisions.md:4830-4844`) -- these already exist at vcloud level + and are UNCHANGED by the flattening. + - `[reorders]` the per-DC storage pool -- was two separate pools (outer holding only + the containment VM's disk, inner holding everything else); COLLAPSES to one pool per + DC on vcloud's own filesystem, no nested-disk-inside-a-disk indirection + (`pass0-w1-substrate-map.md` Section 4). + - `[reorders]` the six planes (`dc-planes` x6) -- re-homed from the inner root's + provider to the flat root's `qemu:///system`; same module, same CIDRs/MTU/family + (D-139 unaffected). + - `[reorders]` the DC edge (`opnsense-edge`) -- re-homed; its WAN attachment changes + from `vr1_dcN_wan` (wan-bridge onto the containment VM's own bridge) to attaching + DIRECTLY to the outer `vr1_dcN_uplink` NAT (Model A shape, D-122 AMENDMENT + `:4844-4846`'s intent preserved, mechanism reverted). + - `[reorders]` the 12 role/utility node VMs (`node-vm x9` + juju-01 + maas-01 + + tailscale-01) -- re-homed to the flat root's provider; MAC-pinning must be + RE-MEASURED post-move (admin report/W0.1 risk #3 -- treat every MAC as unmeasured + until re-pinned, hard rule 2). + - `[new]` the per-DC `vr1-dcN-client` VM (Option 1's small non-hypervisor VM, + ~4/8192/80, metal-admin + transit legs, Model-A-headend shape per admin report + Section 3 item 2) -- genuinely new module instance, no D-143 content, pure + container-elim. +9. **[vanishes][container-elim]** The containment VM itself (`module "vvr1_dc0"` / + `"vvr1_dc1"`, `opentofu/main.tf:410-519,537-623`) and its sizing/rack-addressing/pubkey + vars (`variables.tf:137-156,175-194,196-244,105-114,158-168`) -- deleted from the + module set entirely, not re-homed. +10. **[vanishes][container-elim]** The two `opentofu/vr1-dcN-substrate/` roots AS + SEPARATE ROOTS/STATES -- their bodies survive (step 8), the ROOT BOUNDARY does not. +11. **[vanishes][container-elim]** The qemu+ssh provider dial + connection vars + (`vvr1_dcN_transit_ip`/`_ssh_user`/`_ssh_keyfile`) and D-126 per-env SSH keys -- + already revoked in Part A step 3; here confirmed as having NO successor to re-mint. +12. **[vanishes][container-elim]** `modules/wan-bridge` calls (`vr1_dcN_wan`) and the + `br-vr1-dcN-wan` netplan bridge -- no containment-VM bridge to attach to; folded into + step 8's direct-NAT edge attachment. +13. **[vanishes][container-elim]** The BOOTSTRAP GATE's node-host duty -- + `site-headend-install.sh --host-nodes` (nested libvirtd, inner pool dir + AppArmor + grant, `kvm nested=1`, OPNsense base-image staging) -- confirmed dead code for VR1 + once step 8 removes the nested hypervisor it exists to prepare + (`scripts/site-headend-install.sh:82-85,122-123` gates `--host-nodes` on `--role rack` + explicitly; with no rack-as-node-host, this flag is simply never passed again). + +### B.5 -- the bootstrap gate's SURVIVING duty, re-targeted (OPEN placement) + +14. **[reorders][container-elim]** `site-headend-install.sh --role rack` **WITHOUT** + `--host-nodes` -- MAAS rack enrollment against the Office1 region + (`--region-url .../MAAS --enroll-secret-file ...`) -- still real work + (`site-headend-install.sh:14-21`), just stripped of the node-host half. **Runs AFTER + step 8 creates its target host, not before an inner apply that no longer exists.** + **Placement is OPEN** (admin report Section 6 item 3, gate outcome Section 7a "not + resolved at this gate"): the leading candidate is the new `vr1-dcN-client` VM (same + host that already carries D-138's client role and the SEC-026/028/029 + residencies -- co-locating the rack role there is the lowest-delta reading of Option + 1), but `vr1-dcN-maas-01` (the region-adjacent utility VM) and a ruled retirement are + both still live options per the gate outcome. **This sequencing document does not + pick one** -- it places a slot here and flags it for Phase 2 (tools review, which owns + `site-headend-install.sh` changes) to close. +15. **[reorders][container-elim]** D-131 node-DNS forwarder + rack-legs persistence + (`scripts/dc-rack-net.sh install `, `dc-rack-net.sh:1-19`) -- same OPEN + placement question as step 14 (today it "RUNS ON THE DC RACK HOST", admin report + Section 2 row 7); sequenced immediately after step 14 since both target the same host + class. Losing this silently reintroduces the SERVFAIL bug the forwarder exists to + prevent (admin report Section 2 row 7) -- it must land SOMEWHERE, not be dropped by + omission. +16. **[reorders][container-elim]** Artifact-service placement (`.4`, `dc-mirror.sh` + comment "runs on the rack host") -- same OPEN question, same slot; dc0's debmirror / + dc1's apt-caching-proxy content is unaffected, only WHICH flat VM serves it. +17. **[new][container-elim] -- OPEN, distinct from B.3.** The re-authored SEC-010 + TRANSIT-LEG FORWARD-drop. B.3 builds the NEW cross-DC vcloud-level control; this is a + SEPARATE control -- the admin report is explicit that the two are not the same thing + ("DISTINCT from rebuilding SEC-010's transit drop on the client VM (row 6) -- two + controls," Section 5) and that WHICH ends receive the re-authored drop is itself + unresolved (Section 6 item 5: "which ends get the re-authored FORWARD-drop -- client + VM + voffice1?"). Today SEC-010 is scoped to the transit NIC on `vvr1-dcN` (and its + voffice1-side peer) precisely because the qemu+ssh dial and the inner-planes-bridging + made that interface the isolation boundary; B.4 step 11 removes the qemu+ssh dial + entirely, so the boundary's PURPOSE partly dissolves, but any surviving operator + `ssh -J` access / Office1-originated management flow (B.5 step 14's MAAS enrollment + traffic, for one) still rides SOME leg that needs a FORWARD-drop decision. Sequenced + here, alongside steps 14-16, because it shares their "which flat host" open question -- + NOT resolved by this document; a reader must not conclude B.3 already covers it. + +### B.6 -- MAAS enlist / commission / carve (Phase 3, largely unchanged mechanics) + +18. **[unchanged]*** MAAS discovers the node VMs via the vcloud-registered virsh + `vm-host` -> Office1 region (the per-machine virsh power mechanism, D-123 AMENDMENT + 2026-07-20, `design-decisions.md` -- the pod mechanism was already refuted and + replaced by per-machine virsh BEFORE this pass; that mechanism itself does not change + under flattening). **`*` = the VALUE changes, the STEP does not:** + `scripts/lib-hosts.sh`'s `VIRSH_POWER_ADDRESS*` re-derives from "dial the containment + VM's libvirtd" to "dial vcloud's own libvirtd directly" (admin report Section 2 row + 4) -- a measured-value change, not a new step or a reordering. `maas-node-power.sh` + itself needs NO code change (power address is an argument, topology-agnostic, row 5) + -- every invocation SITE (runbook examples, scripts) needs the new address. +19. **[unchanged]** Commission each discovered node (Phase 3 Step 3); tag each node, + nodes stay Ready (Step 4); Pattern-A interface carve via `scripts/dc-node-carve.sh` + (Step 5) -- mechanically identical against flat-topology node VMs; simpler in that + there is only ONE virsh hop to reason about (no "which host can even see this domain" + confusion the containment layer introduced, teardown-rollback.md `:868-878`). +20. **[unchanged]** PXE/boot-fabric verify (Step 6); per-DC artifact-source + time- + authority verify (Step 7, target host is whatever step 16 lands on); topology + consistency check (Step 8). + +### B.7 -- Juju + OpenStack bundle deploy (Phase 4, mechanically unchanged, host renamed) + +21. **[unchanged]** Juju controller bootstrap on a tagged machine, `preflight.sh` gate, + model/spaces setup, DC egress gate, `juju deploy bundle.yaml` + overlays, dry-run + first, mid-deploy watch (Phase 4 Steps 1-4b) -- runs from the `vr1-dcN-client` VM + (D-138's execution host, "never the vcloud jumphost" doctrine intact per admin report + Section 3 item 6) instead of `vvr1-dcN`. Same commands, renamed host. +22. **[unchanged]** SEC-026/028/029 credentials are (re-)MINTED here, not migrated -- Part + A already revoked the old copies, so this is a clean mint onto the new client VM, the + same "OPEN, rotate at v1 close or on rebuild" posture the ledger already carries. +23. **[unchanged]** vault bring-up (Step 5), IPv6 family-matrix overlay (Step 6), + phase-03/04/05 (Steps 7-9), `cloud-assert.sh --capture` (Step 10), controller backup + (Step 11). +24. **[reorders][both]** VERIFY-LIVE gates (Step 12): Ceph-over-v6 bind check (unchanged + mechanically, D-143 gives it new literal addresses); geneve-over-v6 encap-family + assert (`scripts/geneve-encap-assert.sh`, unaffected by container-elim per admin + report Section 1.4 -- same MTU budget, no extra hop removed or added). **Add here** + the container-elim's OWN new live check: the B.3 cross-DC isolation control's + `--check` gate, re-run now that BOTH DCs' planes are actually co-resident on vcloud -- + this is the first point in the sequence where the control's claim becomes testable + against real traffic. + +### B.8 -- close-out (both axes) + +25. **[new][container-elim]** NetBox DCIM: register the new `vr1-dcN-client` VM + the + flat node roster device records (mirrors Part A step 9's decommission, other + direction). +26. **[new, deferred]** Enter the container-elim [ARCH] decision record (D-123 amendment + vs. new D-number, per SCOPE Section 7 -- Phase 4 frames it, the operator rules it, + GA-R5). Not an executable step; flagged here so the sequence does not read as "done" + without it. The redeploy should not be called closed while this ruling is still owed. + +--- + +## PART C -- D-143 vs container-elim, pulled apart explicitly + +| Redeploy step # | D-143 content (10.12->10.13 address shift) | container-elim content (flatten the topology) | +|---|---|---| +| B.1.1 (capacity) | sizing feeds the rebuild | sizing also proves/disproves the "~176 GiB freed" claim | +| B.1.2 (NetBox apex) | ENTIRE content -- new B2 role, re-carve | none | +| B.1.3 (lib-net.sh) | ENTIRE content -- 10.13 literal blocks | none | +| B.1.4 (transit/statics) | the octet-preserving math | the BEARER host changes (rack -> client VM) | +| B.3 (isolation control) | none | ENTIRE content -- new artifact, new SEC row | +| B.4 (flat apply) | node/plane addresses land in 10.13 (via lib-net.sh) | ENTIRE structural content -- root collapse, module re-homing, client VM | +| B.5 (rack-role slot) | none | ENTIRE content -- placement question, MAAS enrollment retarget | +| B.6 (MAAS enlist) | none (addresses are downstream of lib-net.sh, not re-derived here) | the power-address VALUE changes (row 17) | +| B.7 (juju/bundle) | overlay literals carry 10.13 addresses | execution HOST identity changes (rack -> client VM) | +| Part A.3 (R7 revocation) | owed BY the re-IP ruling (D-143 item 5) | the credentials being revoked are container-elim-eliminated host classes | + +**Reading:** the two axes are NOT cleanly separable at every line (B.1.4, B.6, B.7, +A.3 are genuinely `[both]`), but the STRUCTURAL steps (root collapse, module re-homing, +client VM, isolation control, rack-role placement) are container-elim-only, and the +ADDRESSING steps (NetBox apex, lib-net.sh, overlay literals) are D-143-only. A change-set +reviewer can use this table to confirm a given diff hunk belongs to the axis it claims. + +--- + +## PART D -- where the pre-Roosevelt bare-metal hardware specs plug in + +Per the operator's own framing (SCOPE Section 1, "a good test of the module deployment +project... during the teardown and redeploy") and D-138's own history (the `vr1-dcN-client` +VM is explicitly "the D-138 Roosevelt bastion analog, rehearsed early," admin report +Section 4): the hardware specs, once provided, feed **B.1.1's capacity math** and **B.4's +module bodies**, not a new sequence. Concretely: + +- **B.4's flat-apply module set** (planes x6, edge, node-vm roster, client VM) is the + SAME module composition intended to transfer -- the hardware specs let Phase 4's module + design re-derive sizing for bare metal instead of vcloud-hosted VMs, WITHOUT changing + which modules exist or their call order (only their provider target and concrete + sizing numbers change, same class of change as B.4 itself already is relative to + today's inner root). +- **Nesting-depth honesty, not overclaimed:** flat-Option-1 is `vcloud (libvirt) -> node + VM -> nova KVM guest` = depth 2 (VR0-proven, admin report Section 4). Roosevelt bare + metal removes the vcloud hypervisor layer entirely -- `bare host -> KVM guest` = depth + 1. The flattening is a rehearsal of the MODULE SHAPE and the single-apply/no-bootstrap- + gate WORKFLOW, not of the exact nesting depth; do not let a future session assume + parity it does not have. +- **B.5's OPEN rack-role placement** is the one item that should be RE-DECIDED, not + carried forward blindly, once real hardware specs exist -- a bare-metal Roosevelt build + may have a materially different answer for where MAAS-rack/D-131/artifact-service duties + land than a VM-hosted `vr1-dcN-client` does. +- Not required to unblock THIS sequencing document (SCOPE Section 8: "hardware-agnostic"; + hardware specs are noted here per the worker prompt's explicit ask, not treated as a + blocking input). + +--- + +## PART E -- top sequencing risks (this dimension only) + +1. **B.3 (isolation control) must precede B.4's second-DC apply, not follow it** -- the + gap is live the instant both DCs' planes exist on one kernel; sequencing it after would + leave a real window of unmitigated cross-DC adjacency during the build itself, not just + in the final state. +2. **B.5's OPEN placement (rack-role/D-131/artifact-service) is a genuine sequencing + dependency, not a footnote** -- steps 14-16, 17 (MAAS enlist target), and 20 (juju + execution host) all key off "which flat VM plays this role," and it is currently + UNRESOLVED. A session that runs this sequence without closing it first will improvise + the answer live, which is exactly the class of error CLAUDE.md hard rule 4 exists to + prevent. +3. **MAC re-pinning (B.4 step 8) is understated if treated as cosmetic** -- 24 node MACs + + 2 edge MACs move provider; `pass0-w1-substrate-map.md` risk #3 flags this as a likely + force-replace, not a metadata update, with real MAAS re-enlistment cost. Sequence a + re-measurement pass explicitly after B.4, before B.6 trusts any MAC. +4. **A.3-4 (R7 revocation + MAAS record release) run too late is a real gap, not a + formality** -- if the substrate destroy (A.5-7) runs before revocation, the + credential-bearing hosts are gone and + "revoke" degrades to "assume it's moot," which is exactly the conditional-supersession + risk `open-items-review-20260809.md:396-399` names explicitly. +5. **The flat root's single-state blast radius (admin report/W0.1 risk #2, ~38 domains x + 2 DCs under one state/provider) changes the RISK PROFILE of every apply in B.4** even + though it removes complexity elsewhere -- Phase 4's module design should weigh + one-state-per-DC vs. one-state-total before B.4 is treated as a single monolithic step + in the executable runbook (this document leaves that split as a Phase-4 decision, not + pre-empting it). +6. **This document does not verify the MTU/geneve claim it inherits** ("analytically + unchanged," admin report Section 1.4/8 item 3) -- B.7 step 23 carries the OWED live + assert forward explicitly so it is not silently dropped. + +--- + +## Sources read this session (path:line, for a future session to re-verify) + +- `docs/audit/container-elim-pass/SCOPE-AND-EXECUTION-PLAN.md` (full) +- `docs/audit/container-elim-pass/pass0-admin-report.md` (full) +- `docs/audit/container-elim-pass/pass0-w1-substrate-map.md` (full) +- `runbooks/dc-dc-teardown-rollback.md:600-998` (Steps 1-5, Path A/B, mesh-link note) +- `docs/design-decisions.md:4830-4965` (D-122 amendments, D-123 body + both amendments) +- `docs/design-decisions.md:8083-8182` (D-143 full ruling + owed-execution list) +- `docs/security-ledger.md:79,81,82` (SEC-026, SEC-028, SEC-029) +- `docs/audit/open-items-review-20260809.md:60-100,380-435` (R1, R7, R15, R16) +- `docs/audit/roosevelt-held-decisions-review-20260809.md:70-95` (D-137/R7 pull-forward) +- `scripts/site-headend-install.sh:1-125` (`--role`, `--host-nodes` gating, full flag set) +- `scripts/dc-node-carve.sh:1-13`, `scripts/dc-rack-net.sh:1-19` (headers) +- `docs/tool-index.md:96-134` (tested-artifact lookups used for every script cited above) +- `runbooks/dc-dc-phase2-tofu-dc-substrate.md`, `phase3-maas-enlist-deploy.md`, + `phase4-juju-bundle-per-dc.md` (section headers, for today's step names/numbers) + +No inferred values used; every address, path, and status cited above resolves to the +line given. Items marked OPEN are stated as open, not resolved by inference. diff --git a/docs/audit/container-elim-pass/pass1-w4-module-planning.md b/docs/audit/container-elim-pass/pass1-w4-module-planning.md new file mode 100644 index 0000000..b81dedd --- /dev/null +++ b/docs/audit/container-elim-pass/pass1-w4-module-planning.md @@ -0,0 +1,194 @@ +# Pass 1 -- W1.4: how planning artifacts map onto the layered module model + +**Worker:** W1.4 (Phase 1, container-layer-elimination pass). **Date:** 2026-08-09. +**Scope:** READ-ONLY. This is the conceptual backbone for Phase 4's module-workflow design +(`SCOPE-AND-EXECUTION-PLAN.md` W4.2). Inputs read in full: `SCOPE-AND-EXECUTION-PLAN.md`, +`pass0-admin-report.md` (Option 1 CONFIRMED at the Phase-0 gate, Section 7a), plus live repo +survey below. No inferred values -- every claim below cites path:line or a directory listing +taken this session. + +--- + +## 1. What module structure ALREADY exists (survey, not proposal) + +### 1.1 IaC modules -- `opentofu/modules/` (12 modules, `opentofu/modules/*`) + +Listing (`find opentofu/modules -maxdepth 2 -type d`, this session): +`base-image`, `cloudinit-vm`, `dc-planes`, `dc-storage-pool`, `maas-vm-host`, `mesh-link`, +`netem-link`, `node-vm`, `office1-network`, `opnsense-edge`, `site-wan`, `wan-bridge`. + +Each module is already: (a) independently `tofu validate`-able (`scripts/opentofu-validate.sh`, +`opentofu/README.md:5-6,40-41` -- "validates EVERY module standalone"), (b) called from one or +both of the two CURRENT roots -- the outer root `opentofu/main.tf` (vcloud provider) and the +inner root `opentofu/vr1-dc0-substrate/main.tf` (qemu+ssh provider, D-123 Model B) -- and +(c) already parameterized (no hardcoded DC identity inside a module body; the calling root +supplies `dc`/CIDR/sizing values). This is a working IaC-module layer TODAY; the container-elim +does not need to invent module mechanics, only re-home module CALLS from two roots into one and +delete the modules that existed only to bridge the two-root split. + +Outer-root calls (`grep '^module "' opentofu/main.tf`): `vr1_dc0_storage`, `vr1_dc1_storage`, +`office1_storage`, `office1_network`, `office1_opnsense`, three `mesh_*` legs, `ubuntu_noble_base`, +`voffice1`, `netem_vr1_dc0_vr1_dc1`, `vr1_dc0_uplink`, `vr1_dc1_uplink`, `vvr1_dc0`, `vvr1_dc1` +(`opentofu/main.tf:35-537`). + +Inner-root calls (`grep '^module "' opentofu/vr1-dc0-substrate/main.tf`): `inner_storage`, +`vr1_dc0_planes`, `vr1_dc0_wan`, `vr1_dc0_opnsense`, `vr1_dc0_node` (x12, D-121/R-3) +(`opentofu/vr1-dc0-substrate/main.tf:26-250`). + +The `vvr1_dc0`/`vvr1_dc1` outer-root calls ARE the containment layer (`pass0-admin-report.md` +Section 1.1); everything in the inner-root list is a module CALL that Option 1 re-homes, not a +module BODY that changes shape (`pass0-admin-report.md` row 2: "module CALLS ... re-home into +the outer/flat root -- bodies unchanged, provider changes"). + +### 1.2 Procedure/runbook modularity -- `runbooks/dc-dc-phaseN-*.md` + `docs/dc-dc-deployment-workflow.md` + +`docs/dc-dc-deployment-workflow.md` already defines an 8-stage sequence (Stage 0 decision +ratification through Stage 7 Designate/COS/Magnum), each stage with a fixed schema -- **Goal / +Build / Gate / Owns (D-numbers) / Reuse-vs-new / Authoring status** -- and each stage's +authoring status POINTS AT one `runbooks/dc-dc-phaseN-*.md` file: + +| Stage | Runbook | IaC layer it drives | +|---|---|---| +| 1 | `dc-dc-phase0-vcloud-prep.md` | outer root: mesh/planes/pools/office1-network | +| 2 | `dc-dc-phase1-office1-standup.md` | `cloudinit-vm`+`base-image` (`voffice1`), MAAS-composed LXD | +| 3 | `dc-dc-phase2-tofu-dc-substrate.md` | outer `vvr1_dcN` + bootstrap gate + inner root (THE container layer) | +| 4 | `dc-dc-phase3-maas-enlist-deploy.md` | no tofu; MAAS commission/deploy of already-created node VMs | +| 5 | `dc-dc-phase4-juju-bundle-per-dc.md` | no tofu; Juju controller + `bundle.yaml` | +| 6 | `dc-dc-phase5-dr-failover-drill.md` | `netem-link` mechanism + Ceph replication | +| 7 | `dc-dc-phase6-designate-cos-magnum.md` | no tofu; Designate/COS/Magnum additive apps (D-106/D-105) | + +Each runbook is already parameterized by `$DC` (`lib_net_select_dc "$DC"` / +`lib_hosts_select_dc "$DC"`, DOCFIX-151, `docs/dc-dc-deployment-workflow.md:485-498`) and each +stage's close-out is gated by repo-lint + a harness sweep + a changelog + a branch-merge +(`docs/dc-dc-deployment-workflow.md:424-436`, "Cross-cutting discipline"). **This IS a procedure- +module system already** -- the container-elim pass does not need to invent the concept, only +(a) rewrite Stage 3's content for the flat topology and (b) formalize the pattern the operator +is asking for so it is named and reusable, not just an emergent convention. + +Underneath the stage runbooks, two shared libraries act as the procedure layer's own "IaC +modules": `scripts/lib-net.sh` and `scripts/lib-hosts.sh` -- every stage-4/5 script sources them +for CIDR/host facts, keyed by the `$DC` selector (`vr1-dc0`/`vr1-dc1`, D-119). This is the +existing parameterization mechanism the layer model below reuses rather than replaces. + +### 1.3 Test-harness convention (governs BOTH module kinds) + +- IaC: one shared gate, `scripts/opentofu-validate.sh` (`tests/opentofu-validate/`), validates + every module standalone plus both roots. +- Procedure: per-script harnesses, `tests//run-tests.sh` (tool-index: "65 scripts + with their own `tests//` harness"; `docs/tool-index.md:27`). + +No module of either kind ships without its harness (CLAUDE.md "Delivery" rule); this is already +the repo norm the layer model's design principles (Section 4) restate for the new modules +Option 1 introduces, not a new rule. + +--- + +## 2. Proposed layer model + +Six layers, named to match the EXISTING Stage numbers where a layer is Stage-owned, so the +layer model is a REFRAME of what's already built, not a parallel taxonomy. "IaC or procedure" +marks whether the layer's artifact is an OpenTofu module/root or a runbook+script procedure. +Verify layers (L5) are cross-cutting and re-invoked at every layer boundary, not a one-time +final step. + +| Layer | Name | Kind | Inputs | Outputs | Current artifacts | Stage owner | +|---|---|---|---|---|---|---| +| **L0** | Host & inter-site substrate | IaC module | vcloud libvirt connection; site tokens (Office1/dc0/dc1); MTU plan (D-101) | mesh triangle (dark-fiber legs), per-site storage pools, `office1-network` | `modules/mesh-link`, `modules/dc-storage-pool`, `modules/office1-network`, outer `main.tf:35-163` | Stage 1 | +| **L1** | Site/edge nodes | IaC module | L0 outputs (a network + a pool); per-VM cloud-init or edge-image spec | Office1 headend VM (`voffice1`); per-site OPNsense edge VMs; (Option 1 NEW) the per-DC client VM | `modules/cloudinit-vm`, `modules/base-image`, `modules/opnsense-edge`, `modules/site-wan` | Stage 2 (Office1); Stage 3 (DC edges, currently split by the container layer -- Option 1 flattens into this same layer) | +| **L2** | DC substrate: planes + node VMs | IaC module | L0 pool + L1 edge; `$DC` token; D-121/R-3 node counts/sizing; D-134 octet map | six per-DC plane networks; 9+3 node-VM libvirt domains (role nodes + utility nodes) per DC | `modules/dc-planes`, `modules/node-vm`; TODAY split outer(`vvr1_dcN`)/inner(`vr1_dc0_planes`,`vr1_dc0_node`x12) -- Option 1 COLLAPSES this to ONE root, one state | Stage 3 | +| **L3** | Enlist/commission procedure | Procedure module | L2 node VMs (MAC-pinned, discoverable); MAAS region reachable (L1's `vr1-dcN-maas-01`, D-132 addendum) | READY, carved, tagged MAAS machines; per-DC artifact-mirror answering | `runbooks/dc-dc-phase3-maas-enlist-deploy.md`, `scripts/dc-rack-net.sh`, `scripts/maas-node-power.sh`, `scripts/site-headend-install.sh` (rack-role remainder), `modules/maas-vm-host` | Stage 4 | +| **L4** | Juju/OpenStack deploy procedure | Procedure module | L3's READY machines; `$DC` selector; `bundle.yaml` + per-DC overlays; Vault root | a running independent OpenStack cloud per DC (controller + bundle + Vault) | `runbooks/dc-dc-phase4-juju-bundle-per-dc.md`, `runbooks/phase-01..08-*.md` (the VR0 template it runs twice), `scripts/preflight.sh` (entry gate) | Stage 5 (+ Stage 6 DR, Stage 7 Designate/COS/Magnum as ADDITIVE sub-procedures riding the same L4 mechanism, per `vr0-to-vr1-is-additive` -- not separate layers) | +| **L5** | Verify/gate | Procedure module (cross-cutting) | any layer's declared-done state | PASS/FAIL + a committed BOM at milestones | `scripts/opentofu-validate.sh` (L0-L2), `scripts/cloud-assert.sh`/`--capture` (L4), `scripts/preflight.sh` (L3->L4 boundary), `scripts/geneve-encap-assert.sh` (L2/L4 network correctness), `scripts/repo-lint.sh` (every layer's artifact hygiene) | Invoked at every stage close, not owned by one stage | + +**Layer-boundary rule (why this ordering, not invented):** each layer's INPUT is the layer +below's OUTPUT only -- L1 does not reach into L3, L3 does not dial L0 directly. The one place +this rule is currently VIOLATED is exactly the container layer: the inner root (part of L2) +dials OUT to a cross-host qemu+ssh provider (`opentofu/main.tf:28-29`, "a libvirt provider +cannot be configured from a resource created in the same apply") because L2's outer half +(`vvr1_dcN`) had to exist as infrastructure BEFORE L2's inner half (planes/nodes) could be +declared. Option 1 removes this violation structurally (Section 3) -- it is the single clearest +justification, in layer-model terms, for why flattening also SIMPLIFIES the module system, not +just the topology. + +--- + +## 3. How Option 1's changes map onto the layer model + +| Option-1 change (from `pass0-admin-report.md`) | Layer | What happens to the module/procedure | +|---|---|---| +| Flat single tofu root/state (no outer/inner split) | **L2** | The two-root split collapses INTO L2: `vvr1_dc0`/`vvr1_dc1` module calls (outer `main.tf:410-623`) are DELETED; the inner root's module calls (`vr1_dc0_planes`, `vr1_dc0_node` x12, `vr1_dc0_opnsense`) become direct calls in the single flat root, same module bodies, `qemu+ssh` provider replaced by the flat root's own `qemu:///system` (`pass0-admin-report.md` row 2). The BOOTSTRAP GATE script (`site-headend-install.sh --host-nodes`) that sat BETWEEN the two roots is eliminated as a step -- L2 becomes a single `tofu apply`, no cross-host provider dial. | +| `modules/wan-bridge` deletion; DC edge WAN attaches directly to the outer uplink NAT | **L1/L2 boundary** | `modules/wan-bridge` (existed only to fix the OBS-3 nesting-egress problem, D-125) is DELETED; the DC OPNsense edge (L1) attaches its WAN leg directly to `vr1_dcN_uplink` (already an L0/L1 artifact) -- no re-address, same `/24` (`pass0-admin-report.md` row 3). | +| Per-DC client VM standup (D-138 role, Model-A shape) | **L1 (new module instance)** | A NEW small `cloudinit-vm` instantiation per DC (~4/8192/80, non-hypervisor, `expose_nested_virt=false`) -- the SAME L1 module type Office1's `voffice1` and the DC edges already use, not a new module. Carries the D-138 client role + SEC-028/SEC-029 credential residencies (a procedure/config concern layered on top at L3/L4, not an IaC concern). | +| (a) cross-DC host-level isolation control | **L5 (new artifact)** | Confirmed at the Phase-0 gate (`pass0-admin-report.md` Section 7a) as a Phase-1 DESIGN ITEM: a NEW vcloud-level nftables isolation artifact + mechanical `--check` gate + its own SEC-NNN row -- structurally the SAME pattern as `scripts/geneve-encap-assert.sh`/`cloud-assert.sh` (an L5 verify module), one layer up from SEC-010's interface-scoped drop. This is a genuinely NEW L5 module, not a re-home of an existing one. | +| Rack-controller remainder (MAAS `--role rack`) + D-131 forwarder + artifact-service (`.4`) placement | **L3 (open placement, not yet a module change)** | Currently `site-headend-install.sh --host-nodes` bootstraps the rack role INSIDE `vvr1-dcN`; that call site becomes dead code for VR1 once the containment VM is gone. Where the rack role, `dc-rack-net.sh`'s forwarder, and the mirror land (the new L1 client VM vs. `vr1-dcN-maas-01` vs. retire-with-evidence) is UNRESOLVED -- `pass0-admin-report.md` Section 6 item 3, carried into Phase 1 explicitly as an open item, not resolved by this worker. | +| MAAS region stays on `vr1-dcN-maas-01` | **L2 (no change)** | Already a flat L2 sibling node VM under D-132's addendum; flattening does not touch it. | +| Credential residencies (SEC-028/029) migrate to the client VM | **L3/L4 (procedure, not IaC)** | The register rows (`vm-secret-locations`, SEC-026 isolation control) re-point to the new L1 client VM; this is a data/procedure change at the L3-L4 boundary, not a module-body change. | +| Teardown primitive (`virsh destroy vvr1-dcN` = site-down) | **L5/cross-layer** | D-123's one-command site-down is LOST by flattening (no single containment domain to destroy); Option-1's design must re-earn it as a MODULE-SCOPED group-destroy (a tofu-state-scoped `destroy -target` set, or a scripted domain-group teardown) -- an L2-layer verify/teardown primitive, owed to Phase 4's module-workflow design (`pass0-admin-report.md` Section 6 item 7). | + +**D-143 stays a separate axis, threaded through every layer, not a layer of its own.** The +10.12->10.13 re-IP touches L0's CIDR inputs, L2's node addressing, and L4's overlay files +identically whether or not the container layer exists -- `SCOPE-AND-EXECUTION-PLAN.md` Section 7 +requires the two changes stay distinguishable in the change-set, and the layer model above keeps +that distinguishability structural: D-143 is a PARAMETER change at every layer's input; the +container-elim is a LAYER-COUNT/BOUNDARY change at L1/L2/L5 specifically. + +--- + +## 4. Design principles for the module system (repeatable + Roosevelt-transferable) + +1. **Parameterize by site token, not by hardcoded identity.** The `$DC` selector convention + (D-119, DOCFIX-151, `lib_net_select_dc`/`lib_hosts_select_dc`) already does this for the + procedure layer; the IaC layer already does it structurally (no module body names a DC). + Option 1's new artifacts (the client VM, the L5 isolation control) MUST take the same token + rather than being written DC0-specific and copy-pasted for DC1 -- the exact anti-pattern + `dc-dc-deployment-workflow.md` gap-register item 1 was created to close. +2. **Every module ships its tested harness, no exception** (CLAUDE.md "Delivery"; Section 1.3 + above). A new L1 client-VM instantiation reuses `opentofu-validate.sh`'s existing coverage + (it's the same `cloudinit-vm` module, no new module body); the NEW L5 isolation control needs + its own `tests//run-tests.sh` from day one, matching `geneve-encap-assert.sh`'s and + SEC-010's own precedent -- a mechanical `--check`, not a prose claim. +3. **Idempotence at every layer.** L0-L2 already get this from `tofu apply`; L3/L4 procedure + modules must stay re-run-safe (the existing MAAS "READY not deployed" handoff and the + `preflight.sh` drift-detection gate already enforce this at the L3/L4 boundary -- D-140's own + "OWED BEFORE THE REVIEW" note on P8 drift-detection value is the same principle one layer up). +4. **Clear layer boundaries -- no layer reaches past the one directly below it.** Section 2's + layer-boundary rule; the container layer is the one place this was violated (L2's inner half + dialing a cross-host provider), and removing it is what makes the OTHER four principles easier + to hold, not a side effect. +5. **Findings/design stay LOGGED at their true layer, not folded upward.** The rack-controller/ + D-131/mirror placement question (Section 3) is explicitly NOT resolved here -- it is an L3 + procedure-placement decision that this worker's L1/L2 mapping cannot answer without inventing + a value, so it stays an open item for the phase administrator and Phase-2's tools workers. +6. **D-140 is a DISTINCT axis from this layer model -- do not conflate.** D-140 + (`docs/design-decisions.md:7781-7822`) is PINNED, not ruled, and governs whether the L4 Juju + layer ITSELF becomes OpenTofu-managed -- sequenced strictly AFTER "dc0 deploys and is + HARDENED ... that deployment method is TESTED." It does not currently govern how L0-L2's IaC + modules are structured (that structure already exists and predates D-140). The layer model + above is written so that IF D-140 is later adopted, L4 becomes an IaC-module layer too WITHOUT + changing L0-L3's shape -- but that is a future trigger, not something this pass proposes now. +7. **Roosevelt-transfer lens applied per layer, not per artifact.** L0's node-VM/netem-link shim + (no bare-metal analog, `dc-dc-deployment-workflow.md:11-14` Section-9 shim register) does NOT + transfer; L1's client-VM pattern, L3's MAAS enlist procedure, and L4's Juju/bundle procedure + ARE the D-138 Roosevelt bastion analog and the direct pre-Roosevelt bare-metal deliverable + (`pass0-admin-report.md` Option-1 "For" bullet: "rehearsed early ... transfers to the + pre-Roosevelt bare-metal test"). The layer model's job for Phase 4 is to keep these two classes + visibly separate so the module-workflow design does not present shim-layer work as reusable. + +--- + +## 5. Open questions for Phase 4 (not resolved by this worker) + +1. **L3 rack-role/D-131/mirror placement** (Section 3, `pass0-admin-report.md` Section 6 item 3) + -- unresolved; blocks writing the concrete L3 module change-set until ruled or proposed. +2. **The L5 cross-DC isolation control's exact mechanism** -- confirmed IN SCOPE (handling (a)) at + the Phase-0 gate, but its concrete design (nftables rule set, `--check` gate shape, SEC-NNN + number) is Phase-1/2 work, not yet started. +3. **The site-down teardown primitive's replacement shape** (module-scoped group-destroy) -- + named as owed (Section 3, L5/cross-layer row) but not designed here; it is the biggest single + piece of NEW tooling this layer model implies beyond re-homing existing module calls. +4. **Biggest open question about the layering itself:** whether L3 (MAAS enlist) and L4 (Juju + deploy) should be reframed as OpenTofu-orchestrated procedure invocations once D-140 is + eventually adopted (Section 4 item 6) -- this pass deliberately keeps them procedure modules + now, consistent with D-140's PINNED status, but Phase 4's module-workflow design should name + this as a FUTURE layer-model revision trigger rather than silently assuming procedure-only is + permanent. diff --git a/docs/audit/container-elim-pass/pass2-admin-report.md b/docs/audit/container-elim-pass/pass2-admin-report.md new file mode 100644 index 0000000..ac8df24 --- /dev/null +++ b/docs/audit/container-elim-pass/pass2-admin-report.md @@ -0,0 +1,397 @@ +# Pass 2 -- ADMINISTRATOR REPORT: tools review (container-layer elimination) + +**Author:** the Phase-2 administrator (multi-agent pass, `SCOPE-AND-EXECUTION-PLAN.md` +Section 4). **Date:** 2026-08-09. **Inputs:** `pass2-w1-tofu-modules.md`, +`pass2-w2-lib-hosts-net.md`, `pass2-w3-scripts.md`, `pass2-w4-module-decomposition.md` -- +read in full, adversarially cross-checked against repo ground truth (checks logged in +Section 1). Baseline consumed: `pass0-admin-report.md` incl. Section 7a (**Option 1 +CONFIRMED**; **cross-DC handling (a) CONFIRMED**; MAAS region stays on `vr1-dcN-maas-01`) +and `pass1-admin-report.md` (planning change-set; Section 7 open items 1-7 carried to this +phase). READ-ONLY synthesis; no mutation; findings are LOGGED, not executed. All +recommendations here are proposals for Phase 4 -> operator (GA-R5); "recommend," never +"ruled." + +--- + +## 1. Adversarial-check results (run against repo ground truth this session) + +1. **THREE isolation concerns confirmed distinct, each real, each open -- Section 2 is the + canonical separation.** (i) the cross-DC plane-bridge adjacency on vcloud's kernel + (Phase-0 Section 5, handling (a) operator-confirmed); (ii) the SEC-010 transit-leg + FORWARD-drop successor (pass1 check 3); (iii) **NEW this phase (W2.3 Sec 3): the MAAS + power-key blast radius** -- verified REAL this session, evidence in Section 2.3. They + are two network controls and one credential-scope control; no one artifact covers two + of them; root-shape (B) mitigates NONE of them (check 3 below). + +2. **W2.3's "rack controller already retired in practice" -- CONFIRMED, with one citation + correction.** Spot-checked all three cites by direct read: + - `docs/changelog-20260807-dc1-region-sequence.md:80-89`: dc1 DHCP cutover executed + 2026-08-07 -- `primary_rack=qtw8pm` in `vr1-dc1-region`; verified `ss :67` on the + region VM's `enp1s0` only; "vvr1-dc1 has NO `:67` on the node segment." CONFIRMED. + - `docs/changelog-20260730-dc0-region-migration.md` Item 15 (~:532-549): dc0 DHCP + handover table -- `pgrep -c dhcpd` on the rack = **0**, `ps -ef` on `hot-kid` = 2 + dhcpd on the region VM; `primary_rack=c3aqh8` (= hot-kid in its own region). + CONFIRMED, process-measured. + - **Citation correction:** W2.3 cites `changelog-20260807-dc0-tailscale-provisioning.md + :200,254` for "BOTH maas-01 VMs installed `region+rack`." Read directly: line 200 is + the PLANNED init instruction and line 254 the EXECUTED `maas init region+rack ... + --maas-url http://10.12.68.6:5240/MAAS --force` -- **both are dc1's `.6`**. No dc0 + init transcript was found in the changelogs (grep `region+rack` across changelog-*). + dc0's region+rack status rests on FUNCTIONAL evidence instead: `hot-kid` is a + registered rack in `vr1-dc0-region` (`changelog-20260730:49` "racks=hot-kid"), is + `primary_rack`, and measurably runs the segment's only dhcpd (rackd-managed in MAAS). + Evidence grade: dc1 = transcript; dc0 = functional. The claim carries at both grades; + the DIRECTION (retire the standalone rack) is unaffected. W2.3's own caveat stands: + **a current-day live re-measure (both DCs' `primary_rack`, `vvr1-dcN` rackd state) is + OWED before "settled"** -- the cites are 2-10 days old (instrument-currency #20/#25). + +3. **Root-topology (B): rationale confirmed, and it is about tofu STATE isolation, NOT L2 + network isolation.** W2.1 Sec 2.2's three winning axes are state blast radius, destroy + scoping, and apply-ordering simplicity -- all state/apply-plane facts. Under (B) both + DCs' plane bridges and node VMs STILL share vcloud's one libvirtd/kernel; W2.1 Sec 2.3 + itself keeps the (a) control's `--check` "before the first `tofu apply` of EITHER + per-DC-flat root." **Do not read per-DC roots as addressing the cross-DC network gap + (concern i) or the power-key blast radius (concern iii) -- (B) changes neither.** + +4. **dc0/dc1 D-131 asymmetry: CONFIRMED by direct read.** dc0's forwarder is NOT + load-bearing: `changelog-20260730-dc0-region-migration.md` Item 9 (~:336-352) -- + `dns_servers=10.12.8.6` (the region's own BIND), dig-proven (`archive.ubuntu.com`, + `flags: qr rd ra`, ANSWER: 9), "Pointing node DNS at the DC-LOCAL region is the correct + end state." dc1's forwarder IS load-bearing: `changelog-20260807-dc1-region-sequence.md + :80-89` -- cutover config "replicated verbatim," `dns_servers=10.12.68.3` (the D-131 + forwarder alias). A real rebuild cleanup item: the 10.13 build sets each fresh region's + own BIND from the start and collapses the asymmetry (Section 4.2). + +5. **Client-VM octet `.8` / name `vr1-dcN-client` (W2.2) vs W2.1's grouping: NO CONFLICT.** + W2.2 assigns identity (D-134 utility band, `.8` = next free slot after the `.7` + amendment; cross-DC standard so it binds every DC); W2.1 assigns the apply site (the + per-DC flat root, via `dc-site`'s `client_vm` input taking `metal_admin_ip`/`transit_ip` + as NetBox-assigned inputs). Identity vs grouping -- orthogonal, both hold. W2.2's hedge + is proper: `.8` reads on the metal-admin plane; the transit leg keeps its D-124 /30 + addressing (not a plane octet); exact leg addressing confirmed at build, not inferred. + +6. **Worker contradictions reconciled + inferred-claim sweep:** + - **W2.4 Sec 5 row 2 vs W2.3 rack retirement -- superseded conditional, not a + contradiction.** W2.4 has `site-headend-install.sh --role rack` landing on the client + VM "pending the OPEN placement ruling"; W2.3's measured finding recommends RETIRING + the standalone rack enrollment entirely. W2.4's row was explicitly conditional; with + retirement adopted, **nothing rack-shaped lands on the client VM**, whose duty roster + shrinks to: D-138 client role + SEC-028/SEC-029 credential residencies + concern-(ii) + DC-side endpoint + the transit leg. This simplifies W2.1's open `dc-site` + `client_vm.user_data` content question. + - **W2.4 "no tooling gap here" (Sec 5 row 4) vs W2.3's decommission GAP -- reconciled.** + W2.4's claim was about re-targeting existing script bodies (true); W2.3 correctly + identifies one missing artifact on the retirement path: a `maas rack-controller + delete` decommission step for `vvr1-dcN`'s Office1-registered rack object exists in + NO script. Resolution: it FOLDS INTO owed artifact #5 (the MAAS machine-record + release/delete step, which pass1 already scoped to include "region-side residue -- + the rack-controller's own enrollment record + `primary_rack`/DHCP reference") -- + amended to name the command class explicitly, plus W2.3's runbook note (new builds + install directly as `region+rack` on the region VM, never a separate `--role rack` + enrollment). Not double-counted (Section 5). + - **W2.4's "net new procedure-module bodies: exactly TWO" -- CORRECTED, do not carry.** + With W2.3's findings adopted, net-new bodies are at least FOUR: the (a) control, the + teardown primitive, the power-key mitigation (+ its SEC row), and the `dc-site` IaC + module (W2.4's count was procedure-only and predates the power-key finding). The + SEC-010-writer extraction is a refactor of an existing body; the rack decommission + step folds into #5. + - **W2.2 Sec 4 item 1 and W2.3 Sec 3 are the SAME object seen from two sides -- MERGED** + (Section 2.3): the undetermined power-address target (W2.2) and the blast radius of + re-deriving it (W2.3). Consequence wired into the change-set: the `lib-hosts.sh` + power-address re-derivation AND every call-site/runbook literal update (pass0 row 5) + are **BLOCKED on the mitigation design**, because the mitigation (restricted key / + wrapper / ACL) determines the URI and key shape. Not a literal swap. + - **No inferred-value violations found.** W2.2 and W2.3 both state unknowns as OWED + (power-address target, client-VM plane legs, dc1 cache sizing) rather than asserting. + Spot-checked load-bearing cites resolved: `maas-region-power-key.sh:1-40` (key + resident ON the region host; SEC-012/SEC-016 per-DC), `maas-node-power.sh:28-30` + ("MEASURED to be the REGION" dials power), `opentofu/vr1-dc0-substrate/main.tf: + 176-181` (maas-01's 150 GiB earmarked for PostgreSQL + boot images, "no spare"), + `opentofu/main.tf:175` (voffice1 on the outer/vcloud root), D-134 octet map + D-143 + C.3 as quoted by W2.2. + +--- + +## 2. THE THREE ISOLATION CONCERNS -- separated, not conflated + +Three distinct controls/mitigations are owed. Different layers, different attack surfaces, +different SEC rows. None substitutes for another; root-shape (B) resolves none of them. + +### 2.1 Concern (i): cross-DC plane-bridge NETWORK adjacency -> the (a) vcloud host control + +Both DCs' six plane bridges + node VMs become co-resident on vcloud's ONE libvirtd/kernel +for the first time in any deployed shape. **Status: operator-confirmed at the Phase-0 gate +(handling (a)); design requirement consolidated at pass1 Section 3** -- a vcloud-level +nftables artifact asserting no inter-plane/inter-DC forwarding, with a mechanical `--check` +gate, its own harness, a new SEC-NNN row, a gap-register entry, Stage-1 home, installed and +verified **before the first flat apply of EITHER per-DC root** (fork-robust invariant, +unchanged by (B)). Artifact kind: procedure/L5 gate, NOT a tofu module (W2.1 Sec 3.4 and W2.4 +Sec 5 row 1 both re-confirm; built ONCE). This is a network/forwarding control on vcloud's own +kernel. Owed artifact #2. + +### 2.2 Concern (ii): the SEC-010 transit-leg FORWARD-drop successor + +The interface-scoped FORWARD-drop on the Office1<->DC transit legs. The qemu+ssh purpose +dissolves, but operator `ssh -J` and Office1-originated flows still ride the transit, and +the protective claim (nothing routes from DC planes across the transit) is unchanged in +spirit. **Endpoints resolved by recommendation (W2.3 Sec 4, this phase's owned decision): +client VM (DC side) + voffice1 (Office1 side).** The client VM is structurally forced -- +it is the ONLY DC-side VM with a transit leg under Option 1, and it reproduces SEC-010's +original exposure shape ("straddles metal-admin + the transit," `security-ledger.md:21`) +verbatim. voffice1's end never was containment-bound and does not move. Implementation +shape: **extract the SEC-010 nftables writer out of `site-headend-install.sh`'s +`node_host_setup()` (`:273-320`) into a role-agnostic subcommand that installs BOTH ends** +-- closing today's hand-mirrored voffice1 install rather than carrying it forward. Carry +the NIC-naming trap (live dc0 interface was `enp1s0`, not `mgmt`; re-measure the client +VM's transit NIC before writing the rule). This is a network control on the transit +endpoints. Owed artifact #3 (spec amended). + +### 2.3 Concern (iii): the MAAS power-key BLAST RADIUS -- NEW, VERIFIED REAL + +**The most consequential risk Phase 2 surfaced (W2.3 Sec 3), verified this session on all +three legs:** + +1. **The region VM holds the power key.** `scripts/maas-region-power-key.sh` installs "the + per-DC MAAS->libvirt power key inside a MAAS region's snap. Runs ON the region host" + (SEC-012 dc0 / SEC-016 dc1, per-DC keys, never cross-DC reuse). `maas-node-power.sh: + 28-30`: "MAAS dials the power address from whichever controller it chooses -- MEASURED + to be the REGION." So each `vr1-dcN-maas-01` holds a live qemu+ssh virsh credential. +2. **voffice1 is a domain on vcloud's libvirtd.** `opentofu/main.tf:175` -- `voffice1` is + created by the OUTER root (`qemu:///system` on vcloud), alongside (post-flatten) both + DCs' entire node fleets, edges, and client VMs. +3. **One libvirtd connection scopes to all its domains.** libvirt's default `qemu:///system` + access model is all-or-nothing per connection -- no per-domain scoping. **No repo + artifact configures any libvirt access scoping** (grep of `scripts/` + `opentofu/` for + polkit/ACL: only an unrelated file-ACL comment in `modules/opnsense-edge`). A read of + vcloud's LIVE polkit/libvirt config is delivery work, not asserted here. + +**Consequence:** re-deriving the power address to vcloud's own libvirtd hands EACH DC's +region VM a virsh credential with power control (start/destroy/undefine/console) over +**every domain vcloud manages: both DCs' fleets + voffice1 + the vcloud substrate +(mesh/pool objects)** -- virsh control, not a host shell; do not overstate. This recreates, +through the power-credential door, exactly the cross-DC blast radius SEC-026/D-132 removed +for the MAAS/cloud credential. **SEC-012/SEC-016's per-DC key separation becomes vacuous +under Option 1: the keys stay distinct but both open the same libvirtd.** Neither concern +(i) nor (ii) covers this (both are network controls; this is credential scope at the +libvirt layer), and root-shape (B) does not touch it (state isolation, not connection +scoping). + +**Owed: a THIRD control -- its own mitigation artifact + its own SEC-NNN row** (owed +artifact #11). Candidate mechanisms (Phase-4 choice, not picked here): a +`command=`-restricted SSH key on vcloud (virsh RPC / per-DC domain-set allowlist), a +per-DC-scoped virsh wrapper as the forced command, or a libvirt polkit ACL scoping each +key's connection. **Dependency wired into the change-set: the `lib-hosts.sh` +power-address re-derivation (rows 2-3 of Section 3.2) and every `maas-node-power.sh` +call-site/runbook literal are BLOCKED on this mitigation's design** -- the mechanism +determines the URI and key shape. + +--- + +## 3. The tools change-set + +### 3.1 OpenTofu roots + modules (W2.1 -- full table in `pass2-w1-tofu-modules.md` Sec 1) + +**Zero module bodies need rewriting.** Of 12 module types: **1 collapses** (`wan-bridge` -- +D-125 bridge-in dead; edge WAN -> direct `site-wan` NAT), **1 dead/orthogonal** +(`maas-vm-host`, never instantiated; SEC-013 retire-or-keep flagged to its owner, not +acted on), **3 untouched** (`office1-network`, `mesh-link` -- resolving pass1 open item 6: +the mesh triangle survives, only the office1-leg consumer changes from `vvr1-dcN` NIC1 to +the client VM's transit NIC -- and `netem-link`), **1 rewired input** (`site-wan` output +feeds the DC edge directly), **6 re-home with unchanged bodies** (`cloudinit-vm` -- loses +the 2 containment calls, gains the client-VM call; `dc-planes`; `dc-storage-pool` -- +2-per-DC collapses to 1; `node-vm` x12/DC; `opnsense-edge` -- one input re-pointed; +`base-image`). + +**Root topology: recommend (B) -- shared-outer + per-DC-flat roots (3 total).** Rationale +(state blast radius incl. this repo's own CLAUDE.md-cited incident class; destroy scoping = +`cd && tofu destroy`, the closest re-earn of D-122's one-command site-down; +apply-ordering: DC1's planes structurally cannot be created from DC0's root, so the (a) +ordering invariant reduces to "check before the first apply of either root"). (A)'s only +edge (DRY) is absorbed by `dc-site`. **STATE isolation only -- see check 3.** Directional, +Phase 4 ratifies. Root naming (`vr1-dcN-flat` vs reserving `-substrate`) is Phase 4's call. + +**New IaC: `modules/dc-site`** (highest-leverage new artifact) -- composes pool + 6 planes ++ edge + 12 node VMs + client VM; replaces the ~230-266-line copy-pasted per-DC inner-root +bodies; interface sketched W2.1 Sec 3.1 (site-token parameterized; `client_vm.user_data` stays +a pass-through input -- now simpler, per the shrunken client-VM duty roster of check 6). A +`site-client-vm` wrapper is explicitly NOT built yet. The client VM apply-groups with its +DC's flat root (teardown symmetry, SEC-026 state-level isolation, D-138 fidelity). D-140 +stays a separate future L4 root consuming `dc-site` outputs -- noted, not folded in. + +### 3.2 `lib-hosts.sh` / `lib-net.sh` (W2.2 -- full table in `pass2-w2-lib-hosts-net.md` Sec 1) + +| Surface | Change | +|---|---| +| `VIRSH_POWER_ADDRESS` (VR0 default, `:52`) | UNCHANGED (VR0 out of scope) | +| `VIRSH_POWER_ADDRESS_FROM_OFFICE1` / `_FROM_DCREGION` (`:212-213,246-251`) | **THE two load-bearing edits, value UNKNOWN and BLOCKED on the concern-(iii) mitigation design** (Section 2.3). Whether the FROM_OFFICE1/FROM_DCREGION split survives (both DCs' dials may converge on one vcloud endpoint) is part of that same design. Wrong value = "MAAS unreachable" masquerading as a network fault (pass0 row 4) | +| `CARVE_AUX_HOSTS`, `NIC_PLANE_ORDER`, `BREX_PARENT_NIC`, `HOST_OCTET` maps, suffixes, `HOST_TAG`, resolver fns | UNCHANGED under container-elim (octet maps change under D-143 only). **The client VM does NOT join `CARVE_AUX_HOSTS`** -- it is L1 `cloudinit-vm`, not MAAS-carved | +| Client-VM lib-hosts identity | **Resolved by recommendation:** NO `lib-hosts.sh` row now -- the client VM is not MAAS/virsh-power-managed; its identity lives in tofu (`dc-site` inputs) + NetBox IPAM. Revisit only if a script consumer appears | +| `REGION_HOST_SUFFIX` comment (`:95-100`) | Comment-currency: the build-via-Office1-profile procedure needs re-verification for the flat path (rides the rewrites) | +| **`lib-net.sh` (whole file)** | **ZERO container-elim edits -- D-143 address-axis ONLY** (grep-verified zero `vvr1` hits; D-143 C.3 shape (i) governs). Keeps the two rulable axes cleanly separable for this file | + +### 3.3 Scripts (W2.3 -- full table in `pass2-w3-scripts.md` Sec 1) + +| Script | Change | +|---|---| +| `maas-node-power.sh` | **NO CODE CHANGE** (address is `$1`). Every invocation-site/runbook literal updates -- BLOCKED on the concern-(iii) mitigation (Section 2.3) | +| `dc-rack-net.sh` | **Legs half RETIRES with the containment layer** (no flat VM is a libvirt host with its own bridges; `LEGS`/`br_of()` has no home to move to; `DNS_UPSTREAM=10.10.0.20` confirmed STALE for dc0). DNS-forwarder half: per the D-131 resolution (Section 4.2) | +| `site-headend-install.sh` | `node_host_setup()`/`node_host_check()` (~134 lines): **DEAD, delete wholesale**. **EXTRACT the SEC-010 writer (`:273-320`) into a role-agnostic subcommand installing BOTH concern-(ii) ends** (Section 2.2). `--role rack`: **retired IF the rack-retirement recommendation is adopted** (Section 4.2) -- do not delete prematurely. D-125 WAN-bridge verify code: dead | +| `dc-node-carve.sh`, `dc-node-v6-carve.py`, `carve-host-interfaces.sh`, `maas-role-tags.sh` | **NO CODE CHANGE** (grep-verified zero containment hits; pure MAAS-API, ``-parameterized). Invocation-host currency only (D-128 amendment territory) | +| `site-baseleg.sh` | Stays a no-op; comment block re-cites D-138 + the (a) control instead of the retired qemu+ssh premise (doc-currency) | +| `dc-mirror.sh` / `dc-cache-proxy.sh` (`dc-snap-proxy.sh` rides) | **New host + EXPLICIT disk sizing, not a relabel.** Neither Option-1 VM fits dc0's several-hundred-GB mirror as authored (maas-01: 150 GiB earmarked, verified `vr1-dc0-substrate/main.tf:176-181`; client VM: ~80 GiB). Decision rides the FIT-calculator extension (Section 4.4). The `.4` alias becomes a guest-netplan/MAAS-static concern | +| `maas-region-power-key.sh` | Body unchanged; the KEY it installs is the concern-(iii) object -- its URI/key shape re-derives per that mitigation | + +### 3.4 Module decomposition (W2.4 -- full tables in `pass2-w4-module-decomposition.md`) + +51 deploy-path scripts classified: 4 libraries, 13 gates, 34 procedure modules (32/34 with +harnesses; the 2 gaps routed in Section 6). The procedure-module contract (7 points, all +grounded in existing repo patterns: single-purpose, `$SITE`-parameterized, own harness, +declared I/O + exit contract, idempotent, layer-respecting, changelog+revert) is ADOPTED as +the Phase-4 backbone. The IaC<->procedure boundary is **identity, not orchestration**: tofu +ends at "booted domain with correct MACs/planes"; procedures re-derive everything from live +identity (MAC/hostname/API) and never read tofu state -- and Option 1 structurally removes +the one existing violation (the inner root's qemu+ssh provider dial). **Net-new bodies: +corrected to at least FOUR** (check 6): the (a) control (L5 gate), the teardown primitive +(L2/L5 wrapper), the concern-(iii) mitigation, and the `dc-site` IaC module. Everything +else is re-instantiation or re-targeting via existing parameterization. + +--- + +## 4. Resolutions of the Phase-2-owned open decisions (all: recommendation grade, Phase 4 ratifies) + +### 4.1 Root topology (pass1 open item 1): **(B) shared-outer + per-DC-flat roots** +Per Section 3.1. Teardown primitive builds against (B): gated `tofu destroy` of one DC root ++ emergency `virsh destroy` loop over that root's state-listed domains. NOT an L2-isolation +answer (check 3). + +### 4.2 Rack-controller remainder (pass1 open item 2, "THE highest-leverage open item") -- +**largely DISSOLVES, three components, three answers (W2.3, verified check 2):** +- **(i) MAAS rack controller: RETIRE the standalone registration.** Both maas-01 VMs + already run `region+rack` (dc1 transcript-grade, dc0 functional-grade -- check 2); DHCP + authority measurably cut over on both DCs (2026-07-30 / 2026-08-07). `vvr1-dcN`'s rackd + is a vestigial Office1-region registration doing no work. Delta artifact = the + decommission step folded into owed #5 + the runbook note (new builds init `region+rack` + on the region VM directly; `--role rack` retires). **OWED before settled: current-day + live re-measure.** Phase-4 ride-along: this touches D-132-addendum PREMISES (nothing is + being co-located; the addendum's hypervisor-fate rationale is moot under Option 1) -- + name it in the [ARCH] package. +- **(ii) D-131 forwarder: RETIRE-WITH-EVIDENCE as the 10.13 end state for BOTH DCs.** + D-131's precondition (rack-only controller, remote region) is removed by D-132's per-DC + regions; dc0 already proves the end state (dig-verified region BIND); dc1's forwarder is + presently load-bearing (verbatim-replicated config) -- the asymmetry is real and must + not be assumed equal (check 4). The fresh 10.13 regions set their own BIND from the + start; the retirement EVIDENCE step (the dig test dc0's migration used) is owed artifact + #13. If retirement is rejected, the forwarder co-locates with component (i)'s host. +- **(iii) Artifact service (`.4`): a right-sized, explicitly-provisioned home -- a STORAGE + decision, not a MAAS-adjacency one.** Neither existing VM fits dc0's full mirror as + authored (verified sizing cites, Section 3.3). Options: dedicated volume on a chosen VM, + or a small dedicated utility VM. Decided WITH NUMBERS via the FIT-calculator extension + (owed #7, spec amended to include mirror-disk sizing; dc0 full-mirror vs dc1 cache per + D-135) -- not guessed here. + +**Net effect: the client VM's duty roster SHRINKS to** D-138 client role + SEC-028/029 +credential residencies + concern-(ii) DC-side endpoint + transit leg (check 6) -- no rackd, +no forwarder, and the artifact service only if the sizing decision puts it there. + +### 4.3 Client-VM octet + name (pass1 open item 5): **octet `.8`, name `vr1-dcN-client`** +(W2.2 Sec 2; D-134's ruled utility-band table enumerates through `.7`; cross-DC standard so +`.8` binds every DC; octet-preserving under D-143). Consistent with W2.1's flat-root +grouping (check 5). The D-134 map addition is a RULING -- Phase 4 package. + +### 4.4 SEC-010 successor endpoints (pass1 open item 4): **client VM + voffice1**, one +extracted installer for both ends (Section 2.2). + +### 4.5 Also resolved at this phase +- **Mesh-triangle survival (pass1 open item 6): CONFIRMED at the module level** -- all 3 + `mesh-link` legs + `netem-link` unchanged; only the office1-leg consumer re-points. +- **`site-headend-install.sh` refactor scope (pass1 open item 7):** delete `--host-nodes` + (~134 lines) + extract the SEC-010 writer + retire `--role rack` contingent on 4.2(i). +- **Client-VM lib-hosts identity:** no row (Section 3.2). +- **The (a) control's mechanism (pass1 open item 3): NOT resolved this phase** -- concrete + nftables rule set/check shape/SEC number remain Phase-3/4 design work; kind, home, and + ordering are settled (Section 2.1). + +--- + +## 5. OWED ARTIFACTS -- consolidated (Phase-1 list updated; count = 13) + +| # | Artifact | Status vs Phase 1 | +|---|---|---| +| 1 | Teardown primitive (gated, root-scoped `tofu destroy` of one DC root) | Spec sharpened by (B) | +| 2 | The (a) cross-DC host isolation control (concern i): nftables + `--check` + harness + SEC-NNN + gap-register + gate row | Unchanged spec (pass1 Sec 3) | +| 3 | SEC-010 transit-leg successor (concern ii): **endpoints resolved (client VM + voffice1); ONE extracted role-agnostic installer for both ends** (out of `node_host_setup()`); own SEC-row disposition | Spec AMENDED this phase | +| 4 | R7 credential-revocation checklist | Unchanged | +| 5 | MAAS machine-record release/delete step -- **AMENDED: + `maas rack-controller delete` decommission of `vvr1-dcN`'s Office1 registration + the region+rack runbook note** | Amended (absorbs W2.3's gap; check 6) | +| 6 | Emergency site-down lever (`virsh destroy` loop over the DC root's domain set) | Spec sharpened by (B) | +| 7 | FIT-calculator extension + fresh vcloud capacity measure -- **AMENDED: + artifact-service disk sizing (dc0 mirror vs dc1 cache)** | Amended (4.2 iii) | +| 8 | MAC re-measurement pass post-apply | Unchanged | +| 9 | NetBox DCIM migration (decommission `vvr1-dcN`; register client VM + roster) | Unchanged | +| 10 | Post-build live asserts (geneve/jumbo; gap-#20 re-verify) | Unchanged | +| 11 | **NEW: concern-(iii) power-key blast-radius mitigation** (restricted key / wrapper / libvirt ACL) **+ its own SEC-NNN row**; BLOCKS the lib-hosts power-address re-derivation + all call-site literals | NEW this phase | +| 12 | **NEW: `modules/dc-site` + the per-DC flat-root pattern** (the container-elim's IaC deliverable) | NEW this phase | +| 13 | **NEW: D-131 retirement-evidence step** (dig test against each fresh region's BIND; collapses the dc1 asymmetry) | NEW this phase | + +Not double-counted: the rack decommission (folds into #5), artifact-service sizing (rides +#7), the SEC-010-writer extraction (IS #3's implementation shape). + +--- + +## 6. OPEN -- carried forward + +**To Phase 3 (tests):** harness requirements for #1/#2/#3/#11/#12; the +`maas-fabric-prune.sh`/`maas_fabric_classify.py` harness gap (ROUTE as a named decision: +build vs accept-as-named-exception -- pre-existing, surfaced W2.4); `lib-identity.sh` +harness status (W2.4 flagged to W2.2, which did not cover it -- an unclosed worker handoff, +re-routed to Phase 3); which existing harnesses assume the container layer (Phase 3's own +charter). + +**To Phase 4 (decision framing; operator rules, GA-R5):** +1. The container-elim [ARCH] ruling (D-123 amendment vs new D) + ride-alongs: D-125 + retirement, D-138 concrete-host, D-128 amendment (Plane 2 shrinks), D-122 site-down + re-earn, D-124 sizing-void, **+ NEW: the D-132-addendum premises note (4.2 i)**. +2. Ratify root topology (B) + root NAMING (`-flat` vs reserving `-substrate`). +3. Ratify rack-retirement (i) / D-131 retire-with-evidence (ii) -- after the owed live + re-measure -- and rule the artifact-service home WITH the #7 numbers. +4. Rule the client-VM `.8` octet into the D-134 standing map + the name. +5. Choose the concern-(iii) mitigation mechanism + mint its SEC row; then unblock the + power-address re-derivation (URI + FROM_OFFICE1/FROM_DCREGION split question). +6. The (a) control's concrete mechanism + SEC number (with Phase 3's harness spec). +7. `wan-bridge` module directory: delete vs leave-unreferenced (repo append-only bias). +8. SEC-013 `maas-vm-host` retire-or-keep -- flagged to its owner, not this pass's ruling. +9. State-blast-radius weighing rides item 2 (per pass1). + +**Owed live measurements (delivery, not this pass):** current-day rack/`primary_rack` +state (4.2 i); vcloud's live polkit/libvirt access config (2.3); plus pass0 Sec 8's standing +three (capacity, FIT, geneve/jumbo). + +--- + +## 7. Settled vs open -- two grades, kept distinct + +**OPERATOR-CONFIRMED (do not re-open):** Option-1 target; handling (a) as a required +deliverable; MAAS region stays on `vr1-dcN-maas-01`; plus pass1's settled list (Stage-2/ +D-114 out of scope; D-140 pins L4 as procedure; two-axis attribution; (a) ordering +invariant + Stage-1 home). + +**RESOLVED BY RECOMMENDATION -- pending Phase-4 ratification (GA-R5):** root topology (B); +rack-controller retirement; D-131 retire-with-evidence; artifact-service = explicit sizing +decision via #7; client-VM `.8` + `vr1-dcN-client`; SEC-010 endpoints (client VM + +voffice1) + single-installer shape; client VM out of `lib-hosts.sh`/`CARVE_AUX_HOSTS`; +mesh triangle unchanged; `dc-site` as the IaC composition unit. + +**OPEN:** everything in Section 6; the concern-(iii) mechanism; the (a) mechanism; the +power-address value (blocked, by design, on #11). + +--- + +## 8. Verification note + +Author = "the administrator" (no model name asserted, operator instruction). Direct reads +this session: the four worker docs in full; `docs/changelog-20260807-dc1-region-sequence.md +:70-95`; `docs/changelog-20260730-dc0-region-migration.md:330-355,525-555` + grep for +region+rack/`hot-kid`; `docs/changelog-20260807-dc0-tailscale-provisioning.md:180-260` +(incl. the :200/:254 cite correction); `scripts/maas-region-power-key.sh:1-60`; +`scripts/maas-node-power.sh:25-35`; `opentofu/vr1-dc0-substrate/main.tf:172-185`; +`opentofu/main.tf` module greps; repo-wide polkit/ACL grep. Worker citations were +spot-checked at their load-bearing points, not re-derived wholesale; every correction in +Section 1 names its source lines. READ-ONLY; findings LOGGED only; nothing executed. diff --git a/docs/audit/container-elim-pass/pass2-w1-tofu-modules.md b/docs/audit/container-elim-pass/pass2-w1-tofu-modules.md new file mode 100644 index 0000000..342ce4a --- /dev/null +++ b/docs/audit/container-elim-pass/pass2-w1-tofu-modules.md @@ -0,0 +1,295 @@ +# Pass 2 / W2.1 -- OpenTofu roots/modules (container-layer elimination) + +**Author:** W2.1 (Phase-2 sonnet worker, `SCOPE-AND-EXECUTION-PLAN.md` Section 4). +**Date:** 2026-08-09. READ-ONLY planning; no mutation. Baseline consumed in full: +`SCOPE-AND-EXECUTION-PLAN.md`, `pass0-admin-report.md` (Option 1 CONFIRMED at the Phase-0 +gate; cross-DC handling (a) CONFIRMED), `pass1-admin-report.md` (planning change-set; root +topology left explicitly OPEN for this worker to design, Phase 4 to ratify -- Section 7 item 1 ++ Section 1 check 2). All citations below are from direct reads this session of +`opentofu/main.tf`, `opentofu/variables.tf`, `opentofu/vr1-dc0-substrate/main.tf`, +`opentofu/vr1-dc1-substrate/main.tf`, and every file under `opentofu/modules/*`. + +--- + +## 0. What exists today (ground truth, for the disposition table's "current use" column) + +Three roots, 16 module call sites, 12 distinct module types: + +- **Outer root** (`opentofu/main.tf`, 623 lines, `qemu:///system` local to vcloud): DC storage + pools (x3, incl. office1), office1-network, office1-opnsense, 3 mesh-triangle legs, netem, + base-image, voffice1 (cloudinit-vm), 2 site-wan uplink NATs, **`vvr1_dc0`/`vvr1_dc1`** + (cloudinit-vm -- the containment VMs, `:410-519`/`:537-623`). +- **Inner roots** (`opentofu/vr1-dc0-substrate/main.tf`, `vr1-dc1-substrate/main.tf`, + qemu+ssh into `vvr1-dcN`): inner storage pool, 6 planes (dc-planes), wan-bridge, DC edge + (opnsense-edge), 12 node VMs (node-vm `for_each`: 9 D-121 role nodes + juju-01 + maas-01 + + tailscale-01). +- **12 module types**, all in `opentofu/modules/*`: `base-image`, `cloudinit-vm`, `dc-planes`, + `dc-storage-pool`, `maas-vm-host`, `mesh-link`, `netem-link`, `node-vm`, `office1-network`, + `opnsense-edge`, `site-wan`, `wan-bridge`. + +--- + +## 1. Per-module disposition table + +| Module | Current call sites (path:line) | Option-1 disposition | Reusable-module shape / notes | +|---|---|---|---| +| `cloudinit-vm` | `voffice1` (`main.tf:175`), `vvr1_dc0` (`:410`), `vvr1_dc1` (`:537`) | **PARTIAL COLLAPSE + NEW USE.** The two containment-VM instantiations (`vvr1_dc0`/`vvr1_dc1`) DELETE with the containment layer. The module BODY is unchanged and gains a NEW instantiation: the `vr1-dcN-client` VM (pass0's Option-1 shape: ~4/8192/80, metal-admin + transit legs, `expose_nested_virt=false`) is "L1, same module type as voffice1/edges" (pass1 Section 5). `voffice1`'s call is untouched (D-114, out of scope, pass1 check 1). | Stays a leaf module; no interface change needed. The client VM's `network_names`/`network_config` need a metal-admin plane ref (new -- containment VMs never touched a plane directly) instead of an uplink NIC. | +| `dc-planes` | `vr1_dc0_planes` (inner `:34-40`), `vr1_dc1_planes` (inner `:32-38`) | **RE-HOME, body unchanged.** Moves from the inner (qemu+ssh) root to the flat per-DC root, still one call per DC, same `planes` map / `mtu` / `domain_suffix` inputs (D-139/D-143 own the CIDR *values*, not this module). | No shape change. Becomes a direct child of the new `dc-site` composing module (Section 3). | +| `dc-storage-pool` | `vr1_dc0_storage`/`vr1_dc1_storage` (outer `:35-39`/`:49-53`, held ONLY the containment VM's own disk), `office1_storage` (outer `:59-63`), `inner_storage` x2 (inner `:26-30`/`:23-27`, held the 12 node disks + edge disk) | **COLLAPSES 2-per-DC -> 1-per-DC.** The outer/inner split existed only because Model B had two disks per DC (containment VM's own + the pool inside it). With no containment VM there is one disk surface per DC: the current `inner_storage` call (renamed) becomes THE per-DC pool, sized to hold node disks + edge disk + the client VM's disk. `office1_storage` is untouched (D-114). | Module body unchanged (a `libvirt_pool type=dir`). Net effect: 4 DC-scoped pool calls today -> 2 (one per DC) + `office1_storage` unchanged = 3 total, down from 5. | +| `maas-vm-host` | **NONE** -- authored but never instantiated (REFUTED for DC use by measurement 2026-07-20; `main.tf:6-19` of the module itself; retained pending the SEC-013 retire-or-keep ruling) | **NO CHANGE -- orthogonal to container-elim.** MAAS discovery under Option 1 is still PXE-enlistment + per-machine `power_type=virsh` (pass1 check 7, D-103/D-123 amendments), never this module. Flattening does not resurrect it. | Flag only: this pass is a natural moment to close the SEC-013 retire-or-keep question, but that ruling belongs to whoever owns SEC-013, not this pass -- listed so it is not silently forgotten, not acted on here. | +| `mesh-link` | `mesh_vr1_dc0_vr1_dc1`, `mesh_vr1_dc0_office1`, `mesh_vr1_dc1_office1` (outer `:125-141`) | **STAYS, all 3 legs, module and call sites unchanged.** This resolves pass1's open item 6 at the module-survey level: the mesh triangle is not containment-layer at all (it wires DC<->DC and DC<->Office1 L2 segments, independent of who lives at each end). What changes is the CONSUMER on the two office1-legs: `vvr1_dc0`/`vvr1_dc1` NIC1 today -> the `vr1-dcN-client` VM's transit NIC under Option 1 (client VM inherits the D-138/D-124 transit-leg role). The dc0<->dc1 leg (netem target) is untouched either way. | No module change. The *purpose* of the surviving transit leg (ssh -J / Office1-originated flows / SEC-010-successor-(ii) endpoint) is still an open placement question (pass1 Section 7 item 4) -- a consumer-side decision, not a module-shape one. | +| `netem-link` | `netem_vr1_dc0_vr1_dc1` (outer `:353-358`) | **STAYS, untouched.** Targets the dc0<->dc1 mesh bridge directly (`virbr5`), runs local `tc` on vcloud (D-128) -- has no dependency on either DC's internal shape. Confirmed non-consumer, same class as `lib-net.sh` in pass0's row 15. | No change. | +| `node-vm` | `vr1_dc0_node`/`vr1_dc1_node` `for_each` (inner `:250-265`/`:214-229`) | **RE-HOME, body unchanged.** The single largest re-home by resource count (12 domains/DC): the `for_each` map + `locals` block move verbatim from the inner root into the flat per-DC root, attaching to the same `dc-planes` outputs (now created locally, not over qemu+ssh). MAC re-measurement is owed post-apply regardless (pass1 Section 6 item 8) -- a data concern, not a module-shape one. | No shape change. Becomes a direct child of `dc-site`. | +| `office1-network` | `office1_network` (outer `:76-80`) | **STAYS, untouched.** D-114/Stage-2 territory (pass1 check 1). | No change. | +| `opnsense-edge` | `office1_opnsense` (outer `:99-117`), `vr1_dc0_opnsense`/`vr1_dc1_opnsense` (inner `:60-73`/`:56-69`) | **RE-HOME (DC calls only), body unchanged, one input rewired.** The two DC-edge calls move to the flat per-DC root. `wan_network_name` currently takes `module.vr1_dc0_wan.network_name` (the now-dead `wan-bridge`); it becomes `module.vr1_dc0_uplink.network_name` (the outer root's `site-wan` NAT) directly -- literally the same string-typed input, no module-interface change, just a different upstream module feeding it (D-125 bridge-in removed, "direct-NAT" per pass0 Section 4). `office1_opnsense` untouched. | No shape change to the module itself. Precedent for cross-root string-valued network references already exists in this repo: `office1_opnsense`'s `wan_network_name = "office1-wan"` (`main.tf:112`) is a bare literal today, not even a module reference -- the DC edges' rewired input is the SAME pattern, just sourced from a real module output instead of a literal. | +| `site-wan` | `vr1_dc0_uplink`/`vr1_dc1_uplink` (outer `:379-384`/`:391-396`) | **STAYS, unchanged, gains a consumer.** Same 2 calls, same `172.30.2.0/24`/`172.30.3.0/24` (ruled literals, D-125 note at `variables.tf:246-251` -- IPAM identity untouched). Its output now feeds the DC edge's WAN *directly* (see `opnsense-edge` row) instead of via `wan-bridge`. | No change. This module absorbs the egress duty `wan-bridge` used to hand off. | +| `wan-bridge` | `vr1_dc0_wan`/`vr1_dc1_wan` (inner `:51-56`/`:45-50`) | **COLLAPSES / DELETED.** Its entire reason to exist (D-125 bridge-in) was that the containment VM's only routed leg was the SEC-010 FORWARD-dropped transit, so a NAT-mode WAN inside it would have had no egress (OBS-3). With no containment VM, the DC edge attaches directly to the outer `site-wan` NAT -- there is nothing left for this module to bridge. Confirmed dead per pass0 row 3 and pass1's B.4 "VANISHES" list. | Retire the module directory (or leave in place, unreferenced, per this repo's append-only bias -- Phase 4's call, not this pass's). Its host-bridge (`br-vr1-dcN-wan`) and the bootstrap's `--host-nodes` bridge-verify duty vanish with it. | +| `base-image` | `ubuntu_noble_base` (outer `:165-173`) | **STAYS, unchanged, gains consumers.** One call today, consumed by `voffice1` + both `vvr1_dcN`. Under Option 1 it is consumed by `voffice1` + both `vr1-dcN-client` VMs (same base image, same `backing_store` copy-on-write pattern) -- one fewer distinct consumer class (containment VMs gone) but the SAME call. | No change. | + +**Summary count:** of 12 module types, **1 collapses** (`wan-bridge`), **1 is dead/orthogonal** +(`maas-vm-host`), **3 stay untouched** (`office1-network`, `mesh-link`, `netem-link`), **1 +stays with a rewired input** (`site-wan`), and **6 re-home with unchanged bodies** +(`cloudinit-vm`, `dc-planes`, `dc-storage-pool`, `node-vm`, `opnsense-edge`, `base-image`). +**Zero modules need their HCL rewritten** -- every disposition is a re-homing / re-wiring / +deletion of call sites, never a module-body change. This is the direct payoff of the modules +already being written provider-agnostically (D-119's own design intent) plus the fact that +flattening does not change *what* gets built, only *where the provider dials*. + +--- + +## 2. THE ROOT-TOPOLOGY FORK -- analysis + recommendation + +**Carried open from Phase 1** (`pass1-admin-report.md` Section 1 check 2, Section 7 item 1): +merged single flat root vs. shared-outer + per-DC roots. Pass 1 established the sequence and +the (a)-control ordering hold under EITHER shape, and explicitly assigned this worker to +DESIGN it (Phase 4 ratifies). + +### 2.1 The two candidate shapes + +- **(A) One merged root.** Collapse outer + both inner roots into a single `opentofu/` + directory/state: mesh, uplinks, office1, voffice1, BOTH DCs' pools/planes/edges/12-node + fleets/client VMs, all in one state file, ideally driven by a single `for_each` over a + `var.sites` map (`{vr1-dc0 = {...}, vr1-dc1 = {...}}`). +- **(B) Shared-outer + per-DC-flat roots (RECOMMENDED).** Keep today's 3-root shape, but + delete the qemu+ssh indirection: the outer root keeps everything that is genuinely + cross-DC/shared (mesh, uplinks, office1, voffice1, base-image); each DC's current INNER root + becomes a **flat** root -- same directory, but its `provider "libvirt"` block changes from + `qemu+ssh://...@${transit_ip}/system?keyfile=...` (inner `:12-22`, dc1 `:11-19`) to the outer + root's own local `qemu:///system` (`variables.tf:1-4`) -- and it additionally gains that DC's + `vr1-dcN-client` VM (moved out of the outer root, see Section 2.4). + +### 2.2 Evaluation against the four named axes + +**State blast radius.** (A) puts every domain of both DCs -- the full node fleet, both edges, +both client VMs' credentials-bearing config -- in ONE state file. An untargeted mistake +(the exact incident class CLAUDE.md's hard rule 4 preamble names: *"an ad-hoc `juju +destroy-model` was run instead of the D-061 teardown scripts and took the dc0 controller +down"*) threatens both DCs at once. (B) makes DC0's and DC1's states structurally +un-reachable from each other -- a mistake in DC1's root cannot touch a DC0 resource because it +is not in DC1's state file, full stop, with no reliance on `-target` discipline being followed +correctly every time. **(B) wins decisively** -- this axis is where this repo has first-hand +incident history to weigh against. + +**Destroy scoping / the teardown primitive.** D-122's original intent was "site-down = one +`virsh destroy vvr1-dc0`" -- ONE OBJECT, ONE COMMAND. Under (B), the closest re-earn is +literally `cd opentofu/vr1-dc0-flat/ && tofu destroy` (or a thin wrapper script) -- one +directory, one command, matching the ORIGINAL intent almost exactly, and matching this repo's +EXISTING convention (the inner roots are already per-DC directories; `dc-dc-teardown- +rollback.md` Path A is already organized around "pick a root"). Under (A), site-down requires +a generated `-target=module.xxx` list (one entry per plane/node/edge/client-VM resource for +that DC) -- a new artifact to build AND keep in sync with the node roster, and a bare +`tofu destroy` run without it wipes BOTH DCs (again, the CLAUDE.md-named incident shape). +**(B) wins** -- lower delta from D-122's intent, reuses an existing convention, smaller new +surface to build and test. + +**Cross-DC (a)-isolation ordering.** Pass 1 Section 3's ordering invariant: the control must be +verified "before ANY flat substrate apply that can make any two DCs' planes co-resident... +since under a merged single root the FIRST apply may create both DCs' planes at once." Under +(A) this is a real, standing risk: the default (untargeted) `tofu apply` of a merged root is +ALREADY a both-DCs-co-resident event on its very first run, so safety depends on the operator +remembering to `-target` down to one DC even for that very first apply. Under (B) this is +structural, not procedural: DC1's planes literally cannot be created by applying DC0's root -- +there is no way to accidentally co-create them, so the ordering invariant reduces to "run the +(a) control's `--check` before the first per-DC-root apply of EITHER DC," a much simpler +operator contract. **(B) wins.** + +**`for_each`/module reuse.** Largely NEUTRAL. Regardless of root shape, every module body +listed in Section 1 is reused unchanged -- a module call does not care which root instantiates +it. Introducing the `dc-site` composing module (Section 3) gives BOTH shapes the same DRY +benefit: (A) would `for_each` `dc-site` over `var.sites`; (B) has each per-DC root make ONE +`dc-site` call with that DC's own tfvars file. (A) is marginally more DRY at the HCL-file level +(one file instead of two near-identical directories differing only by tfvars), but (B)'s +"two thin directories, one shared module" is not meaningfully more duplicative than (A)'s +"one file with a for_each" once `dc-site` exists -- this repo already carries the inner roots' +current near-duplication (dc0/dc1 substrate `main.tf` files today) and pass1/pass0 both +identify that duplication as tolerable precisely because the MODULE bodies (not the root files) +are the unit of reuse. **Mild edge to (A)**, not decisive against the three axes above. + +**Current outer/inner structure.** (B) is the direct, low-delta evolution of what exists today +-- the inner roots ALREADY are per-DC directories; the only structural change is deleting the +qemu+ssh provider indirection (one block per root) and merging the outer `vvr1_dcN`-disk pool +into the (renamed) inner pool. (A) requires collapsing 3 roots/states into 1, which means +re-deriving resource addresses for everything currently in the outer root's DC-scoped blocks +(the `vr1_dc0_storage`/`vr1_dc1_storage` pools, if kept renamed, need `moved{}` blocks -- this +repo has done exactly this kind of state-address migration once before, for a much smaller +rename, in `main.tf:275-298`'s D-119 `moved{}` blocks). **(B) wins** on migration risk and +delta size. + +### 2.3 Recommendation + +**Recommend (B): shared-outer root + per-DC-flat roots (three roots total).** Every +blast-radius, destroy-scoping, and isolation-ordering axis favors it, each for reasons +specific to THIS repo's own operating discipline and incident history (CLAUDE.md hard rule 4's +own worked example is exactly the failure mode (B) structurally prevents and (A) would +reproduce at a larger scale). (A)'s only edge -- marginally more DRY HCL -- is fully absorbed +by the `dc-site` module proposed in Section 3, which benefits both shapes equally. This +recommendation is **directional, for Phase 4 to ratify** (pass1 Section 7 item 1), not a +ruling. + +**What (B) implies for the teardown primitive:** the emergency/gated site-down lever (pass1 +Section 6 items 1 + 6) becomes "operate on that DC's root directory" -- `tofu destroy` (gated) +or a `virsh destroy` loop over that root's own state-listed domains (emergency), never touching +the other DC's or the outer root's state. This is the artifact Phase 2/W2.3 or W2.4 should +build against. + +**What (B) implies for the (a)-control ordering:** the control's `--check` gate (Stage 1 home, +per pass1 Section 3) precedes "the first `tofu apply` of EITHER per-DC-flat root" -- a single, +simple operator contract, not a per-root-shape-dependent one. No `-target` discipline is +load-bearing for safety under (B) (unlike (A)). + +### 2.4 The client-VM grouping call (resolves pass1's W1.2-vs-W1.4 tension, Section 1 check 7) + +**Recommend: the `vr1-dcN-client` VM's APPLY site is the per-DC-flat root** (grouped with that +DC's planes/nodes/edge), even though its MODULE TYPE is `cloudinit-vm` (L1, same type as +`voffice1` -- pass1 Section 5's layer model, which classifies by module type, not apply +grouping; both can hold per pass1's own reconciliation). Rationale, specific to (B): + +1. **Teardown symmetry.** If the client VM lived in the outer root, destroying "the DC" would + require touching TWO roots again (client VM in outer, everything else in the per-DC root) + -- reintroducing exactly the two-step INNER-then-OUTER choreography the flattening exists to + remove (pass0 Section 1.2's "single largest structural simplification"). +2. **Credential-isolation symmetry (SEC-026).** The client VM is a DC-scoped credential-bearing + host (SEC-028/029 residencies). Co-locating it in the shared outer root's state would put a + single-DC credential-bearing resource in the SAME state as the other DC's substrate -- + directly opposed to the per-DC isolation SEC-026 exists to guarantee. Per-DC-root placement + keeps that isolation at the STATE level, not just the network level. +3. **D-138 fidelity.** "The cloud-facing client lives IN the DC" (D-138's ruled principle, + `docs/design-decisions.md:7118-7124`) is honored more literally by the client VM sharing its + DC's own tofu root/state than by living in a shared outer one. + +--- + +## 3. Proposed NEW reusable IaC modules + +### 3.1 `modules/dc-site` (the composing module Option 1 wants -- HIGHEST LEVERAGE) + +Replaces the copy-pasted body of `vr1-dc0-substrate/main.tf` / `vr1-dc1-substrate/main.tf` +(currently ~230-266 lines each, identical shape, differing only in tfvars + dc1's schematic +MAC scheme) with ONE module both per-DC-flat roots call once. Composes, in order: +`dc-storage-pool` (1) -> `dc-planes` (6) -> `opnsense-edge` (1, direct-NAT wired) -> +`node-vm` `for_each` (12: 9 D-121 role nodes + juju-01 + maas-01 + tailscale-01) -> +`cloudinit-vm` (1, the client VM). + +**Inputs (sketch, no invented values -- every CIDR/MAC/sizing stays a real per-DC var, same +discipline `variables.tf` already enforces):** + +``` +site_token # e.g. "vr1-dc0" -- D-106 naming anchor, threaded into every child call +domain_suffix # anchor, D-106 +underlay_mtu # Phase-0 measured, no default +pool_path # host filesystem path for this DC's single storage pool +planes # map(plane_name -> {cidr}), same shape as var.vr1_dc0_planes today +opnsense_base_path # prepped nano image path (unchanged from today) +uplink_network_name # STRING -- the outer root's site-wan NAT name (e.g. "vr1-dc0-uplink"), + replaces module.vr1_dc0_wan.network_name; same cross-root string + pattern office1_opnsense already uses (main.tf:112) +nodes # map(vm_name -> {vcpu, mem, disk_gib, osd_gib?, macs}) -- the 12-entry + roster, identical shape to today's vr1_dc0_nodes/vr1_dc1_nodes locals +client_vm # object: vcpu/mem/disk, ssh_pubkey_path, macs, transit_network_name + (STRING -- outer root's mesh_vr1_dcN_office1 network name), + metal_admin_ip/transit_ip (NetBox-assigned, no defaults) +``` + +**Outputs:** plane `network_names` map (pass-through, for verify scripts / a future D-140 Juju +root to consume node/edge addressing), node `domain_names`/`ids` map, client-VM `domain_id`, +edge `domain_id` -- the same shape `dc-planes`/`node-vm`/`cloudinit-vm` already expose today, +just re-exported one level up. + +**What it deliberately does NOT own:** the (a) cross-DC isolation control (confirmed a +procedure/L5 artifact, not a tofu module, pass1 Section 3 -- do not build it twice here); the +rack-controller-remainder / D-131 forwarder / credential-bootstrap content of the client VM's +`user_data` (pass1 Section 7 OPEN item 2 -- still unruled; this module's `client_vm` input +should accept a `user_data`/`network_config` string exactly like `cloudinit-vm` does today, +NOT hardcode content that depends on an unruled placement decision). + +### 3.2 The per-DC-flat ROOT (not a module -- a root pattern) + +`opentofu/vr1-dc0-flat/`, `opentofu/vr1-dc1-flat/` (naming: avoid reusing `-substrate` unqualified +if that name is wanted for the pre-Roosevelt bare-metal target later -- Phase 4's naming call, +flagged not decided here). Each root: one local `provider "libvirt"` block (no keyfile/sshauth/ +known_hosts -- that whole trap class, `vr1-dc0-substrate/main.tf:12-22`'s header comment, +disappears with the qemu+ssh dial), one `module "site" { source = "../modules/dc-site" ... }` +call, one tfvars file. This is the direct, low-delta successor described in Section 2. + +### 3.3 NOT recommended yet: a dedicated `site-client-vm` wrapper module + +The client VM's `cloudinit-vm` call could someday warrant its own thin wrapper (to standardize +its user_data templating once the rack-controller-remainder + D-131 forwarder + credential +bootstrap logic is designed). **Not proposed now**: that content depends on pass1's still-OPEN +Section 7 item 2 (rack-controller-remainder placement), and building the wrapper before that +ruling would either invent the content (hard rule 2 violation) or ship an empty wrapper with no +real interface to validate. Flagged as a likely Phase-2/W2.4 or Phase-4 follow-on once that +placement is ruled, not built here. + +### 3.4 The (a) cross-DC host-isolation control -- explicitly OUT of `opentofu/modules/` + +Restated from this dimension's view so Phase 2's tools workers do not build it twice: pass1 +Section 3 already resolves its artifact kind as a **procedure/L5 verify artifact** (nftables + +`--check` + harness + SEC-NNN row), not an OpenTofu resource. Nothing in this survey changes +that. It sequences BEFORE the first `tofu apply` of either `vr1-dcN-flat` root (Section 2.3). + +--- + +## 4. D-140 alignment (pinned axis -- noted, not folded in) + +D-140 (`docs/design-decisions.md:7781-7819`, PINNED 2026-08-02 to the end-of-deployment +review) converts the **Juju layer** (applications/relations -- the `bundle.yaml`-equivalent +56-application/108-relation surface) to OpenTofu-managed resources, sequenced strictly AFTER a +hardened, tested dc0 deployment. It is an **L4 concern** (pass1 Section 5's layer model) and +does not touch any module in Section 1 or the `dc-site` module proposed here -- those are all +L0-L2 substrate. Where D-140 would eventually intersect this pass's deliverable, noted for +later, not designed now: + +- It would add a **separate, later root** (e.g. `opentofu/vr1-dcN-juju/`, Juju-provider-backed) + consuming this pass's `dc-site` OUTPUTS (node hostnames/addresses, once MAAS has assigned + them -- not available at `dc-site` apply time) as its inputs, not folding into `dc-site` + itself. +- The root-naming convention adopted in Section 3.2 should stay generic enough + (`-`) to admit a future `-juju` root per DC without renaming the + substrate roots this pass proposes. +- D-140's own review is still owed provider-capability research (`references/opentofu-provider- + docs.md`) before it can rule (a) model/config-layer-only, (b) full translation, or (c) + decline -- none of which this pass's substrate-module design depends on or should + pre-suppose. + +--- + +## 5. Summary for the Phase-2 administrator + +- **Root-topology recommendation: (B) shared-outer + per-DC-flat roots** (3 total), NOT a + merged single root -- every axis (state blast radius, destroy scoping, (a)-control ordering, + migration delta) favors it for reasons specific to this repo's own incident history; `for_each` + DRY-ness is the only edge for a merged root and is absorbed by the `dc-site` module either way. +- **Client VM apply-groups with its DC's flat root**, not the outer root (teardown symmetry + + SEC-026 state-level isolation + D-138 fidelity). +- **Zero module bodies need rewriting.** 6 modules re-home unchanged, 1 (`site-wan`) stays with + a rewired input, 3 stay fully untouched, 1 (`wan-bridge`) collapses, 1 (`maas-vm-host`) is + dead/orthogonal. +- **One new module wanted: `modules/dc-site`** (composes pool+planes+edge+12 nodes+client VM; + interface sketched Section 3.1) -- the single highest-leverage new artifact, replacing today's + copy-pasted per-DC inner-root bodies. +- **A `site-client-vm` wrapper is explicitly NOT proposed yet** -- blocked on pass1's open + rack-controller-remainder ruling; building it now would invent content. +- **The (a) isolation control stays OUT of `opentofu/modules/`** (procedure/L5, already ruled + by pass1) -- restated here so Phase 2's tools workers do not duplicate it. +- **D-140 noted, not folded in**: a later, separate L4 root consuming this design's outputs. diff --git a/docs/audit/container-elim-pass/pass2-w2-lib-hosts-net.md b/docs/audit/container-elim-pass/pass2-w2-lib-hosts-net.md new file mode 100644 index 0000000..01b958b --- /dev/null +++ b/docs/audit/container-elim-pass/pass2-w2-lib-hosts-net.md @@ -0,0 +1,187 @@ +# Pass 2 -- W2.2: `lib-hosts.sh` / `lib-net.sh` containment-keyed values (flat topology + 10.13) + +**Author:** W2.2 (Phase 2 -- Tools review), container-layer-elimination pass. READ-ONLY. No +mutation, no live commands run. Repo: `/home/jessea123/openstack-caracal-dc-dc`. +**Baseline consumed:** `SCOPE-AND-EXECUTION-PLAN.md`; `pass0-admin-report.md` (Option 1 +CONFIRMED: flat node VMs on vcloud libvirt + one small non-hypervisor per-DC `vr1-dcN-client` +VM carrying the D-138 client role + that DC's credential residencies; cross-DC handling (a) +CONFIRMED; MAAS region stays on `vr1-dcN-maas-01`); `pass1-admin-report.md` (client VM = +L1 `cloudinit-vm` module type, same class as voffice1/DC edges -- NOT an L2 MAAS-managed node; +rack-controller-remainder + D-131 forwarder placement OPEN, carried to Phase 2 as item #2; +client-VM octet+name carried to Phase 2 as item #5, owned here). Governing rulings read in +full: D-134 (`docs/design-decisions.md:5870-6007`, the standing octet map + its three +amendments), D-143 (`docs/design-decisions.md:8083-8182`, the 10.12->10.13 re-IP, C.3 +lib-net.sh shape). + +Sources read in full: `scripts/lib-hosts.sh` (269 lines), `scripts/lib-net.sh` (264 lines), +`scripts/maas-node-power.sh` (power-address call sites), `scripts/dc-rack-net.sh` (rack-leg / +D-131 forwarder addresses), `docs/tool-index.md` (operation lookups per Hard Rule 4). + +--- + +## 1. `scripts/lib-hosts.sh` -- per-value table + +| Value | Current containment binding | Option-1 flat value/shape | path:line | +|---|---|---|---| +| `VIRSH_POWER_ADDRESS` (flat/VR0 default) | **NOT containment-keyed.** `qemu+ssh://logxen@10.12.64.1/system` dials a real VR0 KVM host directly -- VR0 was never nested. This is the existing flat-topology PRECEDENT the VR1 arms are converging toward. | UNCHANGED (out of scope; VR0 is a separate live cloud, D-143 does not touch it, container-elim does not touch it) | `lib-hosts.sh:52` | +| `VIRSH_POWER_ADDRESS_FROM_OFFICE1` (per-DC) | **CONTAINMENT-KEYED.** `qemu+ssh://jessea123@172.31.0.2/system` (dc0) / `...172.31.0.6/system` (dc1) dials `vvr1-dcN`'s own libvirtd over its D-124 transit-leg NIC1 address, from voffice1 (Office1 MAAS region, historically the sole power-dial origin -- comment: "What the Office1 region has always used"). | **UNDETERMINED from this file alone -- flagged, not inferred.** The containment VM this address dials CEASES TO EXIST under Option 1; there is no libvirtd left at a `172.31.0.x` transit address to reach. Post-flatten, MAAS's power target must become vcloud's own libvirtd (whichever URI/host that resolves to for qemu+ssh from a MAAS region). Whether reaching it FROM Office1 still rides the 172.31.0.0/30 transit legs (now terminating at the `vr1-dcN-client` VM's transit NIC per pass1's "D-124 survives only for the client VM's transit leg") -- a non-hypervisor client VM has no libvirtd for the power dial to land on -- or whether the FROM_OFFICE1 form loses its meaning entirely, is a Phase-2 W2.1(tofu)/W2.3(scripts) DESIGN QUESTION, not resolvable from lib-hosts.sh's data alone. **See Section 4, biggest open item.** | `lib-hosts.sh:212,246`; called out by `lib-hosts.sh:162-172` comment block | +| `VIRSH_POWER_ADDRESS_FROM_DCREGION` (per-DC) | **CONTAINMENT-KEYED.** `qemu+ssh://jessea123@10.12.8.2/system` (dc0) / `...10.12.68.2/system` (dc1) dials `vvr1-dcN`'s libvirtd over its METAL-ADMIN bridge-leg address, from the DC-local MAAS region (`vr1-dcN-maas-01`, .6). Currently unused by default (comment: "Flipping the default ... is OWED once BOTH DCs have migrated"). | **Same undetermined status as FROM_OFFICE1, with an added structural wrinkle:** under Option 1, BOTH DCs' node VMs are co-resident on ONE vcloud libvirtd (pass0 Section 5, the cross-DC adjacency gap). If the region VMs dial vcloud directly, `vr1-dc0-maas-01` and `vr1-dc1-maas-01` may resolve to the SAME target address (one libvirtd for both DCs) rather than two distinct per-DC addresses as today -- itself a manifestation of the co-residency the cross-DC control (Section 3, pass1 admin report) is designed around, not a lib-hosts.sh-local fix. Flagging, not asserting a value. | `lib-hosts.sh:213,250`; comment `:159-172` | +| `CARVE_AUX_HOSTS` (per-DC populated: `vr1-dcN-tailscale-01`, `vr1-dcN-maas-01`) | Names inner-root-provisioned utility VMs, carved via `--host` (excluded from `HOSTS` so role-node consumers don't miscount them). Not itself a containment ADDRESS, but its members were built by the inner (containment-nested) root. | **UNCHANGED SHAPE.** `vr1-dcN-tailscale-01` and `vr1-dcN-maas-01` persist as flat vcloud-libvirt sibling VMs (pass0: the MAAS region "survives flattening as a flat sibling with no redesign"); they stay MAAS-carved and MAC-pinned, so they stay in `CARVE_AUX_HOSTS`, not `HOSTS`. **`vr1-dcN-client` does NOT join this array** -- per pass1's layer model it is an L1 `cloudinit-vm` (same module class as voffice1/DC edges), not an L3 MAAS-enrolled node, so it has no MAAS carve/power identity at all and is out of `lib-hosts.sh`'s namespace entirely (Section 2). | `lib-hosts.sh:34,219,257` | +| `NIC_PLANE_ORDER` | Realized by the INNER root's six pinned MACs per node (`main.tf macs[0..5]`), but the ordering CONVENTION (metal-admin first, matching PXE/boot requirements) is topology-agnostic. | **UNCHANGED.** pass0 row 14: "Conventions carry forward; the FILE that encodes the pinning changes" (the tofu module, not this constant). The re-homed flat module calls keep the same MAC-order-to-NIC-index contract. | `lib-hosts.sh:53-69` | +| `BREX_PARENT_NIC="enp2s0"` | Role-node carve convention (OVS br-ex parented on the provider-public NIC), independent of containment. | **UNCHANGED.** Same reasoning as `NIC_PLANE_ORDER`. | `lib-hosts.sh:72` | +| `HOST_OCTET` maps (`.100-.200` node bands, `.5`/`.6`/`.7` utility octets) | **NOT containment-keyed** -- these are D-121/D-134 node-identity addressing, orthogonal to WHERE the libvirtd lives. A non-consumer of the container-elim axis, same class as `lib-net.sh`'s plane CIDRs (Section 3). | **UNCHANGED under container-elim; changes ONLY under D-143** (second/third octet 12->13; last octet -- the D-134 band -- is offset-relative to the plane /22 and is untouched per D-143's own reconciliation, `design-decisions.md:5942`). This is where the client-VM octet question lives (Section 2). | `lib-hosts.sh:195-199,229-233` | +| `REGION_HOST_SUFFIX="maas-01"` | The constant itself is a naming suffix, containment-agnostic. Its COMMENT is containment-era-specific: "bootstrapped BY Office1 (deployed there, then MAAS installed on it) ... its carve runs against `--profile admin` with all Office1 racks in `--expect-rack`" -- describes a build procedure that assumed the Office1-reachable containment topology. | **Constant UNCHANGED.** The comment's PROCEDURE (build-via-Office1-profile) needs re-verification once the region VM is a flat vcloud sibling with no containment hop to traverse -- this is `site-headend-install.sh` / MAAS-profile-script territory (W2.1/W2.3), not a `lib-hosts.sh` value change. Flagged as a comment-currency item, not asserted as broken. | `lib-hosts.sh:95-100` | +| `JUJU_HOST_SUFFIX="juju-01"`, `TAILSCALE_HOST_SUFFIX="tailscale-01"` | Naming constants only. | UNCHANGED -- non-consumers. | `lib-hosts.sh:89,94` | +| `HOST_TAG` / per-DC `HOST_TAG="openstack-vr1-dcN"` | MAAS placement tag for bundle binding -- orthogonal to containment. | UNCHANGED. | `lib-hosts.sh:112,215,252` | +| `host_sysid()` / `host_sysid_by_bootmac()` | Resolution functions against the live MAAS API -- topology-agnostic (they resolve by hostname/MAC, never by containment address). | UNCHANGED. | `lib-hosts.sh:124-138` | +| `lib_hosts_select_dc()` case-arm structure (`vr1-dc0`/`vr1-dc1`) | Houses all of the above per-DC. | Structure UNCHANGED; the two power-address lines inside each arm are the load-bearing edits (rows 2-3 above). | `lib-hosts.sh:177-269` | + +**Transit IPs, precisely:** the only "transit IP" values IN `lib-hosts.sh` are the +`VIRSH_POWER_ADDRESS_FROM_OFFICE1` targets (`172.31.0.2`, `172.31.0.6` -- the D-124 transit +`/30` legs into `vvr1-dcN` NIC1). No other transit literal exists in this file. Rack-leg +`.2`/`.3` addresses (metal-admin MAAS/DHCP leg, D-131 DNS forwarder) live in +`scripts/dc-rack-net.sh:59-82`, NOT `lib-hosts.sh` -- confirmed by direct read; that script +carries its own re-homing question (pass0 row 7, still OPEN per pass1 item #2) and is W2.3's +dimension, not this one. + +**`maas-node-power.sh` (row 5, pass0) -- CONFIRMED topology-agnostic, verified this session.** +It takes the power address as a positional ARGUMENT (`scripts/maas-node-power.sh +qemu+ssh://jessea123@172.31.0.2/system vr1-dc0`, `:4-5,46`) and writes it via +`power_parameters_power_address` (`:109`). No code change is needed in this script under +Option 1 -- only every call-site/runbook example carrying the old containment address needs +updating once the new target is decided (Section 4's open item). + +--- + +## 2. Client-VM octet + name -- recommendation (the decision this worker owns) + +**Recommendation: octet `.8`, name `vr1-dcN-client`.** + +**Octet rationale, grounded in D-134 (`docs/design-decisions.md:5996-6007`, the most recent +amendment).** The STANDING cross-DC utility-band octet map, as ruled, is: + + .1 gateway (routed planes) + .2 rack (MAAS rack-controller leg) + .3 node-DNS forwarder (D-131) + .4 artifact service (mirror/proxy) + .5 Juju controller (D-134 amendment 2026-07-29) + .6 MAAS region (D-134 amendment 2026-07-29 / D-132 addendum) + .7 Tailscale subnet router (D-134 amendment 2026-08-07) + .8 + +`.8` is the next free slot in the `.4-.49` utility band (`design-decisions.md:5929`, "RESERVED: +future per-DC utility/infra nodes (46 slots)") -- no measurement is needed to establish +freeness; the ruled table enumerates every occupied utility octet through `.7` and stops +there. D-134's 2026-07-29 amendment is explicit that this is not a per-DC choice: **"the +octet map is a STANDING CROSS-DC STANDARD... Assigning it in one DC assigns it in all of +them... Divergence between DCs at the same octet is a DEFECT"** (`:5977-5987`). So `.8` must +be reserved for `vr1-dcN-client` at **every** DC, not decided per-standup -- consistent with +the task framing. Under D-143's octet-preserving shift the last-octet value is untouched +(D-134 endorses the 1:1 shift precisely because it is "offset-relative to the plane /22, not +tied to the second octet," `design-decisions.md:5142-5146` reconciliation), so `.8` reads as +`10.12.8.8`/`10.12.68.8` today and `10.13.8.8`/`10.13.68.8` post-re-IP, on whichever plane(s) +the client VM's legs attach to (metal-admin per the D-138/D-124 transit shape; confirm exact +plane(s) at build time -- not inferred here). + +**Naming rationale.** `vr1-dcN-client` (already the name used consistently across +`SCOPE-AND-EXECUTION-PLAN.md`, `pass0-admin-report.md`, `pass1-admin-report.md`) satisfies the +task's hard constraint -- it does NOT read as `vvr1-dcN` (no risk of conflation with the +eliminated containment class) and it follows the existing `--NN` family the utility +octets already use (`-juju-01`, `-maas-01`, `-tailscale-01`). It is also NOT MAAS-carved +(Section 1's `CARVE_AUX_HOSTS` finding: L1 `cloudinit-vm`, no boot MAC / power-type dance), so +the `-NN` suffix convention is cosmetic consistency, not a functional requirement the way it +is for MAAS-enrolled siblings. + +**One open placement question this recommendation does NOT resolve** (correctly deferred to +pass1's open item #2, not re-decided here): whether the rack-controller remainder + D-131 +forwarder + `.4` artifact service co-locate onto `.8`'s client VM or onto `.6`'s region VM. +That is a role-placement decision, orthogonal to the octet-map slot assignment above -- `.8` +is reserved for the client VM's OWN identity regardless of which additional duties later land +on it. + +**Register/register-adjacent note (not this worker's artifact to build, flagged for +Phase 2/4):** once ruled, `.8` needs a `HOST_OCTET` entry (or equivalent) added to both +`vr1-dc0`/`vr1-dc1` arms if the client VM is to be resolved by the same lib-hosts.sh +machinery as the utility nodes -- but per Section 1's `CARVE_AUX_HOSTS` finding, it is not +a MAAS/virsh-power object, so whether it belongs in `lib-hosts.sh` at all (vs. purely in the +tofu `cloudinit-vm` definition + IPAM record) is itself a small open design choice for W2.1. + +--- + +## 3. `scripts/lib-net.sh` -- axis-separation statement (verified, not assumed) + +**Confirmed by direct full read this session: `lib-net.sh` contains ZERO `vvr1-dc` / +containment references.** Grep-confirmed (`grep -n "vvr1" scripts/lib-net.sh` -> no hits; +matches Phase-0's finding, pass0-admin-report.md row 15). Every value in the file is IPAM +literal: `PLANE_CIDRS`, `PLANE_NAME`, `PLANE_GW`, `DATA_PLANE_CIDRS`, `METAL_INTERNAL_*`, +`VIP_PREFIX_*`, `VIP_OCTET_MIN/MAX`, `VIP_COUNT_EXPECT`, `FIP_POOL_START/END`, +`KEYSTONE_VIP_DEFAULT`, and the `lib_net_select_dc()` per-arm overrides -- none of these +encode a containment VM, a qemu+ssh dial, or a nested-libvirt fact. They are entirely D-052 / +D-119 / D-133 / D-134 / R9 / R11 address-plane facts. + +**D-143 is EXPLICIT and load-bearing here, not inferred.** The ruling itself specifies the +exact shape lib-net.sh takes (`design-decisions.md:8116-8123`, Exchange 2 C.3, operator exact +utterance **"(i) Keep flat defaults at 10.12; VR1 arms get full 10.13 blocks"**): the flat +(`vr0-dc0`, no-op) defaults stay at 10.12 (the live cloud's real values); the `vr1-dc0` and +`vr1-dc1` case arms each gain a complete explicit 10.13 literal block. Owed execution item 2 +of the ruling (`:8167-8168`) names this file directly: *"`scripts/lib-net.sh`: keep flat +defaults at 10.12; give `vr1-dc0`/`vr1-dc1` full 10.13 literal blocks; UPDATE the now-false +`:124-134` 'inherits VR0 unchanged' comment (F13)."* Every edit this ruling requires is a +value substitution (10.12.x.y -> 10.13.x.y across `PLANE_CIDRS`, `PLANE_GW`, `VIP_PREFIX_*`, +`FIP_POOL_*`, `KEYSTONE_VIP_DEFAULT`) plus one comment-currency fix (F13) -- no shape change, +no new containment-dependent value, nothing container-elim touches. + +**Verdict: `lib-net.sh`'s changes are D-143 ADDRESS-AXIS ONLY.** The container-elim change-set +should carry ZERO line-item edits to `lib-net.sh`. This keeps the two axes cleanly separable +for this file specifically (unlike the four `[both]`-tagged items pass1 found elsewhere -- +G17, B.1.4, R7, B.7 -- none of which are `lib-net.sh` edits). + +**Plane-CIDR / MTU values the flat topology touches -- checked, none found IN this file.** +`lib-net.sh` carries no MTU constant at all (grep-confirmed: no `MTU` / `mtu` token in the +file). The plane CIDR VALUES (`PLANE_CIDRS`) are untouched by flattening -- pass0/pass1 are +explicit that the six planes re-home to vcloud-level bridges with "SAME CIDRs/families/MTU -- +IPAM identity untouched" (pass0 Section 4) and that the containment hop was "a same-MTU bridge +with no extra encapsulation... removing it changes no byte budget" (pass0 Section 1.4). Where +those CIDRs are REALIZED (which host's bridges carry them) is a tofu/substrate fact +(`opentofu/vr1-dcN-substrate` -> the flat root, W2.1's dimension), not a `lib-net.sh` fact -- +this file only ever held the address-plane VALUES, never the hosting topology. + +--- + +## 4. Top risks / open items (this dimension) + +1. **BIGGEST FINDING: `VIRSH_POWER_ADDRESS_FROM_OFFICE1`/`_FROM_DCREGION`'s target is + UNDETERMINED, not just "re-derived."** Both today dial `vvr1-dcN`'s libvirtd -- an object + Option 1 deletes outright. There is no drop-in replacement host with a libvirtd at a + client-VM-shaped address (the client VM is explicitly non-hypervisor). The power dial must + land on vcloud's own libvirtd; whether the FROM_OFFICE1/FROM_DCREGION SPLIT still means + anything once both DCs' targets may collapse toward the same vcloud host is a design + question for W2.1 (tofu placement) and W2.3 (script/runbook), not something resolvable by + editing this file's literals alone. Getting this wrong reproduces exactly the CLAUDE.md-cited + incident class (wrong power address masquerading as a network fault, pass0 row 4). +2. **Cross-DC co-residency implication surfaces here too, not just in the network-wiring + dimension.** If both DCs' power dials converge on one vcloud libvirtd target, that is a + second, host-identity-shaped face of pass0 Section 5's cross-DC adjacency gap -- worth + flagging to whichever Phase-2 worker owns the (a) isolation-control design so it accounts + for the operational surface, not only the data-plane one. +3. **`REGION_HOST_SUFFIX`'s comment procedure (build-via-Office1-`--profile admin`) is + containment-era language** that needs re-verification against the flat build path + (W2.1/W2.3), even though the constant itself does not change. +4. **The client-VM octet's placement in `lib-hosts.sh`'s own data structures is undecided** + (Section 2, register note) -- contingent on W2.1's tofu-module classification of the + client VM (whether it gets any lib-hosts.sh-visible identity at all, given it is not + MAAS/virsh-power-managed). +5. **Every VIRSH_POWER_ADDRESS call-site and runbook example** (row 5's finding, echoed from + pass0) needs a literal update once item 1 above is resolved -- listed here as a delivery + dependency, not re-enumerated (pass0 row 5 already owns the site inventory). + +--- + +## 5. Method note + +No live commands were run. All findings above are grep/Read-verified against repo HEAD this +session (`scripts/lib-hosts.sh`, `scripts/lib-net.sh`, `scripts/maas-node-power.sh`, +`scripts/dc-rack-net.sh`, `docs/design-decisions.md` D-134 section incl. all three amendments, +D-143 in full, `docs/tool-index.md`). Where a value could not be determined from the repo as +it stands (Section 4 item 1), it is stated as UNKNOWN / owed, per the pass's no-inferred-value +rule, rather than asserted. diff --git a/docs/audit/container-elim-pass/pass2-w3-scripts.md b/docs/audit/container-elim-pass/pass2-w3-scripts.md new file mode 100644 index 0000000..10cd100 --- /dev/null +++ b/docs/audit/container-elim-pass/pass2-w3-scripts.md @@ -0,0 +1,266 @@ +# Pass 2 -- WORKER W2.3: carve/power/network scripts (container-layer elimination) + +**Author:** Phase-2 worker W2.3 (multi-agent pass, `SCOPE-AND-EXECUTION-PLAN.md` Section 4). +**Date:** 2026-08-09. **Scope:** `scripts/maas-node-power.sh`, `scripts/dc-rack-net.sh`, +`scripts/site-headend-install.sh`, the carve scripts (`dc-node-carve.sh`, +`dc-node-v6-carve.py`, `carve-host-interfaces.sh`, `maas-role-tags.sh`), `scripts/ +site-baseleg.sh`, `scripts/dc-mirror.sh`/`dc-cache-proxy.sh`. Owns the two +highest-leverage Phase-1 open decisions (rack-remainder placement; SEC-010 successor +endpoints). READ-ONLY; findings LOGGED only, nothing executed. Inputs read in full: +`SCOPE-AND-EXECUTION-PLAN.md`, `pass0-admin-report.md`, `pass1-admin-report.md`, plus the +scripts themselves and the design-decisions / security-ledger / changelog citations below. + +**Baseline consumed:** Option 1 CONFIRMED (flat node VMs on vcloud libvirt + one small +non-hypervisor `vr1-dcN-client` VM per DC carrying the D-138 client role + +SEC-028/SEC-029 credential residencies); cross-DC handling (a) CONFIRMED (new vcloud-level +host isolation control, Phase-1 design item, not this worker's dimension); MAAS region +stays on `vr1-dcN-maas-01` (no change); rack-controller-remainder placement OPEN +(pass1 Section 7 open item 2, "THE highest-leverage open item"). + +--- + +## 0. A load-bearing fact this worker surfaced, not present in pass0/pass1 + +**The "rack-controller remainder" is substantially ALREADY RETIRED IN PRACTICE for both +DCs' DHCP/enrollment duty -- this is measured, live-executed history, not a proposal.** + +- **dc0** (`docs/changelog-20260730-dc0-region-migration.md` items 9, 10, 15): the D-132 + region migration moved DHCP from Office1's rack (`primary_rack=7chphy`, i.e. `vvr1-dc0`) + to the region VM `vr1-dc0-maas-01` ("hot-kid", `primary_rack=c3aqh8` **in its own + region**). Executed 2026-07-30, read back by PROCESS (`pgrep`/`ps -ef`), not MAAS + self-report: **"2 dhcpd on the region VM; rack still 0"** (item 15 step 4). Node DNS was + ALSO measured and switched: item 9 -- `dns_servers=10.12.8.6` (the region VM's own BIND, + not the D-131 forwarder `.3`), proven live (`dig` against `10.12.8.6` answers + `archive.ubuntu.com` / `maas-internal` SOA correctly, `flags: qr rd ra`). Item 9's own + words: **"Pointing node DNS at the DC-LOCAL region is the correct end state... removes + the cross-fiber dependency that D-132 q1 exists to remove."** +- **dc1** (`docs/changelog-20260807-dc1-region-sequence.md` Item 2): the same DHCP handover + ran 2026-08-07 (`primary_rack=qtw8pm` in `vr1-dc1-region`; verified `ss :67` on `enp1s0` + only, `vvr1-dc1` has NO `:67`). **But DNS was NOT re-derived** -- the cutover explicitly + **"replicated verbatim"** the old config, so dc1's region still carries + `dns_servers=10.12.68.3` (the D-131 forwarder alias), unlike dc0's corrected + `10.12.8.6`. This is a real, present ASYMMETRY between the two DCs, not a documentation + gap: dc1 still depends on the forwarder today; dc0 does not. + +**Implication for Decision 1:** the MAAS-rack-as-DHCP-server function has already left +`vvr1-dcN` for both DCs. What remains genuinely resident on `vvr1-dcN` today is (a) its own +idle rack-controller *registration* (a rackd process enrolled to Office1's region, doing no +DHCP work), (b) the D-131 forwarder (load-bearing for dc1, provably NOT load-bearing for +dc0), and (c) the artifact service. These three components do not need one placement +answer -- they were never one thing operationally, even though they shared a host. + +--- + +## 1. Per-script disposition table + +| Script:line | Containment binding today | Option-1 change | +|---|---|---| +| `scripts/maas-node-power.sh` (whole file, 123 lines) | **NONE in code.** `POWER_ADDRESS` is `$1` (`:41`), validated only as a `qemu*://` URI shape (`:51-54`); `PREFIX` is `$2` (`:42`). No hostname, IP, or containment literal is baked in -- confirmed by full read. | **NO CODE CHANGE.** Every *invocation site* / runbook example that passes `qemu+ssh://@172.31.0.2/system` (the containment VM's libvirtd) must instead pass vcloud's own libvirt endpoint. The new value is **UNKNOWN -- an OWED live measurement**, not inferred (hard rule 2). See Section 4 for why this value change is not "just a literal swap." | +| `scripts/dc-rack-net.sh` (226 lines) | Whole-file: runs ON the rack host (`:8`, "RUNS ON THE DC RACK HOST (e.g. vvr1-dc0)"); `LEGS` table (`:62-64`, `:78-80`) is bare `ip addr` on the containment VM's **own inner libvirt bridges**, re-applied via `br_of()` -> `virsh -c qemu:///system net-dumpxml` (`:102-104`) -- a host-with-its-own-bridges pattern that only exists because `vvr1-dcN` IS the inner libvirtd. `DNS_UPSTREAM` (`:66`, `:81`) hardcodes `10.10.0.20` (Office1's BIND) for BOTH sites -- STALE for dc0 per Section 0 (measured 2026-07-30 that dc0 now points `dns_servers` at its own region, not this forwarder's upstream chain at all). | **RETIRE-WITH-EVIDENCE for the legs half; component-dependent for the DNS half (Section 4).** No per-DC VM under Option 1 is a libvirt host with its own bridges -- flat sibling VMs get their addresses from guest netplan / MAAS statics, and the (a) cross-DC control's whole posture is "vcloud keeps no L3 legs on DC planes" (pass1 Section 3). The `LEGS`/`br_of()` mechanism has no home to move to; it retires with the containment layer itself. The DNS-forwarder half is a SEPARATE decision, Section 4. | +| `scripts/site-headend-install.sh` (565 lines) | `--role rack` (`:14-20`, `:371-409`): enrolls a DC rack to Office1's region -- STILL A LIVE CODE PATH even though DHCP has moved off it (Section 0); `--host-nodes` / `node_host_setup()` (`:248-342`, ~95 lines) + `node_host_check()` (`:206-244`, ~39 lines) = **~134 lines, the D-123 Model-B node-hosting bootstrap** (nested KVM, inner pool dir, AppArmor grant, OPNsense base staging, D-125 WAN-bridge verify). Embedded in `node_host_setup()`: the SEC-010 nftables writer (`:273-320`, ~48 lines) -- the one piece of this block that is NOT dead. | **`node_host_setup()`/`node_host_check()` (~134 lines): DEAD, delete wholesale** -- no inner root, no nested libvirt, no OPNsense-inner-edge staging under Option 1. **EXTRACT the SEC-010 writer (`:273-320`) OUT of `node_host_setup()` into its own role-agnostic subcommand** (e.g. `--transit-drop --transit-if `), so it can run on the client VM (and, per Section 5, on voffice1 too, replacing today's hand-mirrored install) WITHOUT dragging in the dead nested-KVM setup. `--role rack` itself (region-enrollment only, no `--host-nodes`): its disposition is **CONTINGENT on Decision 1** -- retired outright if the rack-controller identity fully consolidates onto `vr1-dcN-maas-01` (Section 4 recommendation); kept, retargeted to the client VM, only if the operator rejects that consolidation. D-125 WAN-bridge verify code (`:90-94`, `:228-243`, `:322-335`): dead, D-125 bridge-in itself retires per pass0 (edge WAN -> direct NAT). | +| `scripts/dc-node-carve.sh` (485 lines), `scripts/dc-node-v6-carve.py` (301 lines), `scripts/carve-host-interfaces.sh` (301 lines), `scripts/maas-role-tags.sh` (210 lines) | **NONE found.** Grepped all four for `vvr1`, `containment`, `qemu+ssh`, `inner`, `172.31`, `POWER`, `voffice1` -- zero hits in every file. Confirmed by header read: these operate purely against **the MAAS API** (machine records by MAC/tag, `maas machines/interfaces/...`), run "where the maas CLI lives / the D-128 Plane-2 host" (`dc-node-v6-carve.py:109`, `maas-role-tags.sh:51`) -- a value/profile question, not a containment-binding one. | **NO CODE CHANGE.** These are already topology-agnostic; the only currency item is *where they are invoked from* (D-128's Plane-2 host shrinks per the D-128 amendment pass1 flagged, check 6) and *which `MAAS_PROFILE`* they target -- both are invocation-parameter concerns, already handled by the existing `MAAS_PROFILE`/`--profile` plumbing these scripts carry. | +| `scripts/site-baseleg.sh:40-48` | DC rows are commented-out placeholders (`# [vr1-dc0]="|..."`, `:47`), explicitly deferred: "the DCs nest inside vvr1-dc0 and are reached by qemu+ssh... NOT necessarily an L3 leg" (`:41-42`). D-138 (2026-07-30) already answered this for the CURRENT shape: "no host-side leg is wanted" (design-decisions.md:7127-7129, cross-referenced from this file's own deferred-row comment). | **Premise moot, not merely re-answered.** Under Option 1 there is still no vcloud-side L3 leg wanted onto a DC plane (the (a) control's entire point is the opposite -- no cross-plane forwarding on vcloud's kernel). Stays a no-op; the comment block should be updated to cite D-138 + the (a) control rather than the retired qemu+ssh premise, but this is a doc-currency edit, not a behavior change. LOW. | +| `scripts/dc-mirror.sh` (385 lines), `scripts/dc-cache-proxy.sh` (396 lines) | Whole-file: "RUNS ON THE DC RACK HOST (e.g. vvr1-dc0)" (`dc-mirror.sh:6`, `dc-cache-proxy.sh:13`), except `dc-cache-proxy.sh node` which runs on a DC node (`:18`, unaffected). `.4` listen alias added on a metal-admin libvirt bridge via the same rack-legs mechanism as `dc-rack-net.sh` (`dc-mirror.sh:16-23`). dc0's mirror pulls "several hundred GB" (jammy main/restricted/universe/multiverse + jammy-updates/security + UCA jammy-updates/caracal, `:29-34`); dc1's is a lighter apt-cacher-ng cache, not a full mirror (D-135 amendment, per-DC strategy split). | **New host + explicit disk sizing, NOT a doc-only relabel.** The rack-legs `.4`-alias mechanism retires with `dc-rack-net.sh`'s legs half (row above); the artifact SERVICE itself needs a fresh host binding. Neither Option-1 utility VM is sized for dc0's mirror footprint as currently authored: `vr1-dc0-maas-01` is `4 vCPU / 8 GiB / 150 GiB disk` and that 150 GiB is EXPLICITLY earmarked for "the region's PostgreSQL AND its boot-image set" (`opentofu/vr1-dc0-substrate/main.tf:179-180`), not spare; the client VM is sized `~4/8192/80` (pass0 Section 4). See Section 4 component 3. dc1's cache-proxy footprint is materially smaller and less likely to force a resizing decision, but should not be assumed to fit without the same sizing pass. | + +--- + +## 2. OPEN DECISION 1 -- rack-controller-remainder placement (per component) + +Three components, three separate recommendations -- they were never one placement problem +(Section 0). + +### Component (i): the MAAS rack controller + +**Recommendation: RETIRE the standalone rack registration; let `vr1-dcN-maas-01`'s own +`region+rack` install (already present) be the DC's sole rack.** + +- **This is smaller than "co-locate rack onto the region VM" -- it is already true today.** + `vr1-dc0-maas-01` and `vr1-dc1-maas-01` were BOTH installed with `maas init region+rack` + (`docs/changelog-20260807-dc0-tailscale-provisioning.md:200`, `:254`; the region+rack + form, not region-only), so each already runs its own local rackd. DHCP authority for both + DCs was measured and cut over to that local rackd 2026-07-30 (dc0) / 2026-08-07 (dc1), + verified by live process, not self-report (Section 0). `vvr1-dcN`'s own rackd is a + vestigial registration to Office1's region doing no work. +- **Do NOT cite `site-headend-install.sh --role region+rack` as the delta artifact.** That + role is the D-114 Office1/voffice1 build path and carries LXD install + LXD vm-host + registration + compose-network DHCP (traps 1-4 of that script) -- none of which + `vr1-dcN-maas-01` needs or has. The maas-01 VMs were already built correctly via a direct + `maas init region+rack --database-uri ... --maas-url ...` sequence + (`docs/changelog-20260807-dc0-tailscale-provisioning.md:196-203`), NOT via this script's + region+rack role. The delta artifact this recommendation needs is therefore a GAP, not a + reuse: a formal decommission step for `vvr1-dcN`'s Office1-registered rack object + (`maas admin rack-controller delete` or equivalent -- not yet in any script) and a runbook + note that new-build racks going forward install DIRECTLY as `region+rack` on the region VM + the way maas-01 already was, never as a separate `--role rack` enrollment. +- **`scripts/maas-node-power.sh --profile`/`MAAS_PROFILE` plumbing is unaffected** -- power + config already targets whichever profile is passed; this is a rack-identity question, not + a power-config one. +- **OWED before this is treated as settled, not before it is recommended:** re-measure + CURRENT live state (both DCs' `primary_rack` binding, `vvr1-dcN`'s rackd process state) as + of THIS session (2026-08-09) -- the cited changelogs are 2026-07-30/2026-08-07, and this + is a READ-ONLY planning pass, so the measurement is Phase-2/3 delivery work, not asserted + here as current fact. The DIRECTION is well-evidenced; the CURRENT-DAY confirmation is not + yet taken. +- **Ride-along note for Phase 4 [ARCH] framing (not ruled here, GA-R5):** eliminating the + separate rack does not touch the D-132 addendum's actual RULING ("region... NOT on the + rack host" -- design-decisions.md:7232-7234) in the way "co-locate rack onto region" + would have, because nothing is being co-located; the separate rack is being retired, and + the addendum's own stated rationale (region must not share fate with "the hypervisor + running every node it manages") is structurally moot under Option 1 regardless (no VM is a + hypervisor for another VM). Still name this explicitly in the Phase-4 package alongside + the D-128/D-125/D-138 ride-alongs pass1 already flagged (Section 7 item 10) -- it touches + the same ruled decision's premises even though it does not reverse its letter. + +### Component (ii): the D-131 node-DNS forwarder + +**Recommendation: RETIRE-WITH-EVIDENCE as the target end state for BOTH DCs under the +10.13 rebuild; carry the asymmetry honestly in the interim.** + +- D-131's own title scopes it to "rack-only controllers" (`design-decisions.md:5653`) -- + its SERVFAIL bug is specifically the rack-only agent resolver walking public root hints + with a remote region (`docs/audit/commissioning-diag-20260721.txt`, cited + `dc-rack-net.sh:29`). D-132's per-DC region (regiond + BIND, on metal-admin) removes the + precondition: a node can take `dns_servers=` directly, and the + workaround has nothing left to work around. +- **dc0 already proves this, measured, not hypothesized:** `dns_servers=10.12.8.6` (the + region VM's own BIND) answers `archive.ubuntu.com`/`maas-internal` correctly with `dig` + (`docs/changelog-20260730-dc0-region-migration.md:337-349`); the migration's own words: + "Pointing node DNS at the DC-LOCAL region is the correct end state... removes the + cross-fiber dependency that D-132 q1 exists to remove." + `dc-rack-net.sh:66,81`'s `DNS_UPSTREAM="10.10.0.20"` for both sites is confirmed STALE by + the same changelog entry (item 9's own follow-up note) -- an instrument-currency finding + in this script's own site table, not a live-state guess. + **dc1 is NOT yet at this end state** -- its 2026-08-07 cutover explicitly "replicated + verbatim" the old forwarder-pointed config (`dns_servers=10.12.68.3`, + `docs/changelog-20260807-dc1-region-sequence.md:85-87`) rather than re-deriving it, so + dc1's forwarder is presently load-bearing. +- **For the 10.13 rebuild specifically:** since BOTH DCs are being rebuilt from scratch + under Option 1, the forwarder need not be stood up at all -- set `dns_servers` at each + fresh region's own BIND address from the start (dc0's already-proven pattern), and + RE-VERIFY with the same live-answer test dc0's migration used (`dig` for + `archive.ubuntu.com` + the `maas-internal` SOA) against the new build before calling this + closed. That collapses the asymmetry rather than carrying it into 10.13. +- **If retirement is rejected** (e.g. a Roosevelt-transfer argument for keeping a + forwarder pattern rehearsed): co-locate it with whichever host carries the rack-controller + registration decision (component i) -- `dc-rack-net.sh`'s forwarder half is coupled to the + same host's metal-admin leg by construction (`gen_dns_unit`'s `Requires=${SITE}-rack-legs. + service`, `:155`), so it has no independent placement logic once the legs half is decided. + +### Component (iii): the artifact service (`.4` mirror/proxy) + +**Recommendation: a right-sized, explicitly-provisioned home, decided at Phase 2 (W2.1 IaC +module design) -- NOT silently inherited from whichever VM absorbs components (i)/(ii).** + +- This is a **storage** decision, not a MAAS-adjacency decision -- unlike the rack + controller and forwarder, the mirror/proxy has no dependency on being co-resident with + MAAS. Sizing rules it out of both existing Option-1 candidates as currently authored: + `vr1-dcN-maas-01` is `150 GiB` disk, already earmarked for "the region's PostgreSQL AND + its boot-image set" (`opentofu/vr1-dc0-substrate/main.tf:179-180`, explicit comment, no + spare capacity claimed); the client VM is sized `~4/8192/80` (pass0 Section 4) -- neither + has headroom for dc0's "several hundred GB" full debmirror (`dc-mirror.sh:29-34`: jammy + main/restricted/universe/multiverse + updates/security + UCA jammy-updates/caracal). +- dc1's `dc-cache-proxy.sh` footprint (an apt-cacher-ng cache, not a full mirror -- D-135 + amendment's deliberate per-DC strategy split) is materially smaller and a plausible fit on + either existing VM with a modest disk bump, but should not be assumed without the same + sizing pass -- an inferred disk-size claim here would be exactly the hard-rule-2 trap. +- **Concretely:** either (a) attach a dedicated volume to whichever VM ends up hosting it, + sized explicitly for D-135's known per-DC footprint (dc0 full-mirror vs dc1 cache), or (b) + keep it a distinct small utility VM if the operator wants mirror-storage growth isolated + from either control-plane VM's disk. Both are legitimate; the FIT-calculator extension + pass1 already flagged as owed (Section 6 item 7) is the right place to settle it with + numbers rather than here with a guess. +- `dc-rack-net.sh`'s legs mechanism that the mirror's `.4` alias rides today retires with + the rest of that script's legs half (Section 1); the new host's `.4` alias becomes a + guest-netplan/MAAS-static concern like every other flat-VM address, not a host-level + `ip addr replace` unit. + +--- + +## 3. New finding: the re-derived power address is a cross-DC credential blast-radius risk + +`scripts/maas-node-power.sh` needs no code change (Section 1), but the VALUE change is not +innocent. Today, per DC, the qemu+ssh power key dials ONLY that DC's own nested libvirtd +(the containment VM) -- `POWER_ADDRESS` per DC is scoped to that DC's own hardware. Under +Option 1, node VMs move to being **flat siblings directly on vcloud's own libvirtd**, so the +re-derived power address becomes vcloud's own `qemu:///system` (or a `qemu+ssh://` dial into +vcloud) -- **one endpoint reachable from BOTH DCs' region VMs** (each DC's `vr1-dcN-maas-01` +independently dials it for its own power control, per the "credential note" already in the +script: `:28-30`, "MAAS dials the power address from... the REGION"). + +A per-DC region VM holding a virsh key that reaches vcloud's libvirtd has **virsh power over +everything on vcloud** -- both DCs' node fleets, `voffice1`, the jumphost's own substrate -- +not just its own DC's hardware. This recreates, through a different door, exactly the +cross-DC blast radius SEC-026/D-132 worked to remove (D-132's whole point was DC-local MAAS +so a DC-local host does not have region-wide reach; this reintroduces region-wide REACH via +the power-control credential even though MAAS itself stays DC-local). **Neither the (a) +cross-DC network-isolation control nor the SEC-010 successor (Section 4) covers this** -- both +are network/forwarding controls; this is a credential-scope problem at the libvirt layer. + +Not in pass0 rows 4-5 or pass1's carried-forward items. **Flagged as an owed SEC-row + +mitigation design** for Phase 2/4: a `command=`-restricted SSH key (virsh RPC allowlist) or a +per-DC-scoped virsh wrapper/ACL on vcloud's libvirtd, so each region VM's power key can only +touch its own DC's domain set, matching the per-DC isolation SEC-026 already establishes for +the MAAS/cloud credential. + +--- + +## 4. OPEN DECISION 2 -- the SEC-010 transit-leg FORWARD-drop successor endpoints + +**Recommendation: client VM (DC side) + voffice1 (Office1 side), unchanged shape, distinct +from the (a) control.** + +- **The client VM is the structurally-forced DC-side successor**, not a choice among + several: per the confirmed Option-1 shape, the client VM is the only DC-side VM that + carries a transit leg at all (`vr1-dcN-client`, "legs = metal-admin + transit", pass0 + Section 4). It reproduces SEC-010's ORIGINAL exposure shape essentially verbatim -- the + ledger's own words for what SEC-010 was written against: the host "straddles metal-admin + (DC-local) + the office1<->dc0 transit (crosses fiber)" + (`docs/security-ledger.md:21`) -- so the protective claim carries without + reinterpretation: nothing should route FROM the DC's node planes THROUGH the client VM + ACROSS the transit leg; only the client VM's own originated/terminated traffic (operator + `ssh -J`, any client-VM-terminated calls) should cross it. +- **voffice1's end is unchanged** -- it was never containment-bound (it is not `vvr1-dcN`), + and its peer role on the SAME physical mesh leg does not move under Option 1. +- **Re-author, do not blind-copy, the rule content:** the qemu+ssh purpose that motivated + the original drop is gone, but the drop's actual protective claim (no forwarding across + the transit leg) is unchanged in spirit, so the SAME interface-scoped nftables idiom + applies (`oifname "$TRANSIT_IF" drop` / `iifname "$TRANSIT_IF" drop` in the forward hook, + `site-headend-install.sh:299-305`) -- only `$TRANSIT_IF` re-targets to the client VM's own + transit NIC name. +- **Naming-trap precedent, cite it explicitly at build time:** the LIVE dc0 interface was + `enp1s0`, not the script's default `mgmt` -- "netplan set-name dropped" + (`docs/security-ledger.md:21`, close note). Do not assume the client VM's transit NIC + keeps any prior name; re-measure it live before writing the rule, same discipline hard + rule 2 already requires. +- **Consolidate the install, do not re-hand-mirror it:** today voffice1's end was applied as + an "identical scoped artifact" by HAND, separately from the rack's scripted install + (`docs/security-ledger.md:21`, "voffice1... identical scoped artifact + `sec010-fw.service` + enabled"). The extraction recommended in Section 1 (pull the SEC-010 writer out of + `node_host_setup()` into a standalone, role-agnostic subcommand) should make ONE tested + artifact install BOTH ends, closing that hand-mirroring gap rather than carrying it + forward. +- **Distinct from control (a), the new vcloud-level cross-DC host-isolation control + (pass0 Section 5, pass1 Section 3) -- two controls, two SEC rows, do not merge them.** + Control (a) guards vcloud's OWN kernel against inter-plane/inter-DC forwarding now that + both DCs' planes are co-resident on one libvirtd; the SEC-010 successor guards the + Office1<->DC transit leg specifically. Different hosts, different attack surfaces, + different SEC-NNN rows (pass1 check 3, "CONFIRMED distinct, BOTH OPEN, neither dropped"). + +--- + +## 5. Top risks (bounded) + +1. **Power-address blast radius (Section 3)** -- the highest-value new finding this worker + surfaced; no existing control covers it. +2. **dc1's D-131 forwarder is presently load-bearing, dc0's is not** -- an asymmetry that + must not be silently assumed equal when authoring the retirement step for 10.13. +3. **Artifact-service sizing** -- neither Option-1 candidate VM has spare disk for dc0's + full mirror as currently authored; a placement chosen without the FIT-calculator pass + (pass1 Section 6 item 7) risks an under-sized build discovered mid-sync (hours-long job). +4. **`--role rack`'s disposition in `site-headend-install.sh` is contingent on Decision 1's + outcome** -- do not delete it prematurely; it is only fully dead if the rack-consolidation + recommendation (Section 2, component i) is adopted. +5. All Decision-1/Decision-2 recommendations are proposals for the Phase-2 administrator -> + Phase 4 -> operator (GA-R5); none is ruled here. + +**Durable doc:** `docs/audit/container-elim-pass/pass2-w3-scripts.md` (this file). diff --git a/docs/audit/container-elim-pass/pass2-w4-module-decomposition.md b/docs/audit/container-elim-pass/pass2-w4-module-decomposition.md new file mode 100644 index 0000000..6e160eb --- /dev/null +++ b/docs/audit/container-elim-pass/pass2-w4-module-decomposition.md @@ -0,0 +1,244 @@ +# Pass 2 -- W2.4: module decomposition (which tools become procedure-modules; the IaC/procedure layering) + +**Worker:** W2.4 (Phase 2, container-layer-elimination pass). **Date:** 2026-08-09. **Scope:** +READ-ONLY. Builds on the L0-L5 layer model adopted at Phase 1 (`pass1-w4-module-planning.md`, +consolidated in `pass1-admin-report.md` Section 5) and the Option-1-confirmed target +(`pass0-admin-report.md` Section 7a). Inputs read in full: `SCOPE-AND-EXECUTION-PLAN.md`, +`pass0-admin-report.md`, `pass1-admin-report.md`, `pass1-w4-module-planning.md`, +`docs/tool-index.md`, plus a live inventory of `scripts/`, `tests/*/`, and script headers taken +this session (cited by path throughout). D-140 (L4-as-IaC) is PINNED, not ruled -- treated as a +future trigger, not folded into this decomposition (per CLAUDE.md instruction and +`pass1-w4-module-planning.md` Section 4 item 6). + +--- + +## 1. Method + +Every script under `scripts/` matching the task's deploy-path filter (`preflight`, +`cloud-assert`, `phase-NN-*`, `lib-*`, `dc-*`, `maas-*`, `site-*`, `carve-*`, +`geneve-encap-assert`, `opentofu-validate`) was read (header + role), placed on the L0-L5 model, +classed as **IaC / procedure / library / gate**, and checked against `tests/` for a harness +(`ls tests/*/` cross-referenced by basename; a script with no matching `tests//` dir is +flagged NO). 51 scripts matched the filter; `opentofu/modules/*` (12 IaC modules) are W2.1's +domain and are cited here only at the handoff boundary (Section 4), not re-decomposed. + +--- + +## 2. Script -> module decomposition table + +Legend: **kind** = IaC (OpenTofu-owned) / Procedure (bash/python driving live MAAS/juju/OS state) +/ Library (sourced, no independent action) / Gate (procedure, but its sole job is verify, never +mutate by default). **Layer** per Section 2 of `pass1-w4-module-planning.md`. **Harness** = +`tests//run-tests.sh` exists (checked live this session). + +| Tool | Current role | Kind | Layer | Harness? | +|---|---|---|---|---| +| `lib-hosts.sh` | host/power-address/NIC-map constants, sourced | Library | L2/L3 (consumed) | Selector fn covered by `tests/dc-selector/` (DOCFIX-151); the constants table itself is exercised only transitively via every consumer's harness -- no dedicated full-library harness | +| `lib-net.sh` | CIDR/plane/space constants, sourced | Library | L0/L2 (consumed) | Same as above (`tests/dc-selector/`) | +| `lib-identity.sh` | identity constants, sourced | Library | L2/L3 (consumed) | Not independently verified this session -- flag for W2.2 (its dimension) | +| `lib-validate.sh` | shared exit-contract + emit() lib for `scripts/checks/*.sh` | Library | L5 (verify-layer support) | `tests/lib-validate/` YES | +| `opentofu-validate.sh` | validates every IaC module standalone + both roots, S1-S3 guards | **Gate** | L0-L2 | `tests/opentofu-validate/` YES | +| `preflight.sh` | THE single pre-deploy gate; sequences P1-P4 | **Gate** | L3->L4 boundary | `tests/preflight/` YES | +| `cloud-assert.sh` | behavioral cloud verifier, `--capture` BOM | **Gate** | L4 | `tests/cloud-assert/` YES | +| `geneve-encap-assert.sh` | OVN geneve family/tunnel-health gate | **Gate** | L2/L4 (network correctness) | `tests/geneve-encap-assert/` YES | +| `dc-egress-check.sh` | layered DC-egress probe (read-only) | **Gate** | L3 (rack-host) | `tests/dc-egress-check/` YES | +| `dc-node-v6-verify.sh` | gate G19: v6 statics + plane forwarding | **Gate** | L2/L3 | `tests/dc-node-v6-verify/` YES | +| `dc-dc-mtu-geneve-budget.sh` | MTU/geneve budget calculator | **Gate** (calculator) | L0 | `tests/dc-dc-mtu-geneve-budget/` YES | +| `dc-dc-ceph-disk-budget.sh` | Ceph disk-budget calculator | **Gate** (calculator) | L2 | `tests/dc-dc-ceph-disk-budget/` YES | +| `dc-dc-whole-host-budget.py` | whole-host RAM/vCPU/disk FIT calculator | **Gate** (calculator) | L0/L2 | `tests/dc-dc-whole-host-budget/` YES | +| `maas-profile-assert.sh` | proves which region a profile resolves to | **Gate** | L3 | `tests/maas-profile-assert/` YES | +| `site-headend-install.sh` | installs MAAS region+rack (or rack-only) + LXD on the host it runs on | Procedure | L1 (install) / L3 (rack-role output) | `tests/site-headend-install/` YES | +| `dc-rack-net.sh` | rack bridge legs + D-131 node-DNS forwarder, runs ON the rack host | Procedure | L3 | `tests/dc-rack-net/` YES | +| `dc-node-carve.sh` | v4 NIC/br-ex carve for a DC's nodes | Procedure | L3 | `tests/dc-node-carve/` YES | +| `dc-node-v6-carve.py` | v6 static assignment mirroring the v4 carve | Procedure | L3 | `tests/dc-node-v6-carve/` YES | +| `carve-host-interfaces.sh` | Pattern-A single-host interface carve (VR0) | Procedure | L3 | `tests/carve-host-interfaces/` YES | +| `maas-node-power.sh` | sets `power_type=virsh` per enlisted machine, MAC-matched | Procedure | L2->L3 handoff (Section 4) | `tests/maas-node-power/` YES | +| `maas-role-tags.sh` | creates + applies per-role MAAS tags the bundle constrains on | Procedure | L3 | `tests/maas-role-tags/` YES | +| `maas-region-power-key.sh` | installs/verifies the per-DC MAAS->libvirt power key | Procedure | L3 | `tests/maas-region-power-key/` YES | +| `maas-fabric-prune.sh` | deletes orphaned auto-fabrics (recurring maintenance) | Procedure | L3 | **NO** -- no `tests/` dir found this session (gap, logged not executed) | +| `maas_fabric_classify.py` | pure classifier backing the above (no mutation) | Library (pure fn) | L3 (support) | **NO** -- same gap | +| `dc-region-topology.sh` | builds/verifies a per-DC MAAS region's fabric/space/subnet/tag topology | Procedure | L3 | `tests/dc-region-topology/` YES | +| `dc-plane-ipam.sh` | site-keyed plane IPAM state incl. v6 carve, D-134's executable gate | Procedure (+ gate mode) | L2/L3 | `tests/dc-plane-ipam/` YES | +| `dc-mirror.sh` | per-DC apt+UCA artifact mirror, runs ON the rack host | Procedure | L3 | `tests/dc-mirror/` YES | +| `dc-cache-proxy.sh` | per-DC apt caching proxy (interim/DC1-strategy artifact path) | Procedure | L3 | `tests/dc-cache-proxy/` YES | +| `dc-snap-proxy.sh` | per-DC snap forward proxy | Procedure | L3 | `tests/dc-snap-proxy/` YES | +| `dc-node-etchosts.sh` | cloudinit-userdata generator for node-local `/etc/hosts` fix | Procedure (generator, feeds L1 cloud-init) | L1/L3 boundary | `tests/dc-node-etchosts/` YES | +| `site-baseleg.sh` | durable base L3 leg vcloud -> site-local libvirt net | Procedure | L0/L1 | `tests/site-baseleg/` YES | +| `site-forward.sh` | rootless systemd port-forward jumphost -> site VM | Procedure | L1 (access) | `tests/site-forward/` YES | +| `site-ssh-config.sh` | ssh_config Host-alias generator for site VMs | Procedure | L1 (access) | `tests/site-ssh-config/` YES | +| `site-tailscale.sh` | per-DC Tailscale subnet-router install/check | Procedure | L1 | `tests/site-tailscale/` YES | +| `render-dc-overlays.py` | deterministic per-DC bundle-overlay renderer (derive/render split) | Procedure | L4 | `tests/render-dc-overlays/` YES | +| `phase-00-maas-standup.sh` | MAAS fabric/VLAN/subnet/space stand-up (VR0 plane scheme) | Procedure | L3 (VR0 template L4 reuses per-DC per `pass1-w4` Section 1.2) | `tests/phase-00-maas-standup/` YES | +| `phase-00-teardown-destroy.sh` / `-release.sh` | juju-model teardown, VR0-scoped (D-061) | Procedure | L4/L5 (destroy) | `tests/phase-00-teardown-d061/` YES | +| `phase-02-vault-preflight.sh` | Vault preflight for the VR0 template | Procedure | L4 | `tests/phase-02/` YES | +| `phase-03-admin-openrc.sh`, `phase-03-core-verify.sh` | admin creds + core-service verify | Procedure | L4 | `tests/phase-03-adminrc/`, `tests/phase-03/` YES | +| `phase-04-network-create.sh`, `-verify.sh`, `-internal-cert-san-verify.sh` | network stand-up + verify + cert SAN check | Procedure | L4 | `tests/phase-04-create/`, `tests/phase-04/`, `tests/phase-04-internal-cert-san/` YES | +| `phase-05-amphora-pipeline.sh`, `-octavia-verify.sh` | Octavia amphora image pipeline + verify | Procedure | L4 | `tests/phase-05-amphora/`, `tests/phase-05/` YES | +| `phase-06-bootstrap.sh`, `-capi-stack.sh`, `-k8s-bootstrap.sh`, `-kubeconfig-gate.sh`, `-mgmt-vm.sh`, `-net-setup.sh` | Magnum/CAPI tenant-K8s stand-up chain | Procedure | L4 (Stage-7 additive, per `vr0-to-vr1-is-additive`) | Each has its own `tests/phase-06-*/` dir -- YES | +| `phase-07-conductor-graft.sh` | Magnum conductor graft step | Procedure | L4 | `tests/phase-07-conductor-graft/` YES | +| `dc-dc-rbd-mirror.sh`, `dc-dc-radosgw-multisite.sh`, `dc-dc-dr-drill.sh` | Ceph rbd-mirror / radosgw multisite bootstrap + failover-failback sequences (D-108) | Procedure | L4 (Stage 6, additive) | `tests/dc-dc-rbd-mirror/`, `tests/dc-dc-radosgw-multisite/`, `tests/dc-dc-dr-drill/` YES | + +**Shape of the decomposition (this filtered set of 51 scripts):** 4 library units (2 with a +dedicated selector-mechanism harness, 2 flagged for W2.2), 13 gates (verify-only, cross-cutting +L5 or layer-scoped calculators), 34 procedure modules (L1-L4), of which **32/34 ship a harness +today and 2 do not** (`maas-fabric-prune.sh` / `maas_fabric_classify.py` -- a pre-existing gap, +unrelated to container-elim, logged here because this pass's harness-discipline principle +(`pass1-w4-module-planning.md` Section 4 item 2) would otherwise silently wave it through). + +--- + +## 3. What a "procedure module" IS in this repo's terms -- the contract + +Grounded entirely in patterns that already exist (`pass1-w4-module-planning.md` Section 1.2-1.3; +`docs/tool-index.md`; CLAUDE.md "Delivery"), not invented for this pass. A procedure module is: + +1. **A named, single-purpose script** under `scripts/`, one file = one job (the repo's existing + granularity -- `dc-rack-net.sh` does rack-net only, `dc-mirror.sh` does the mirror only; no + script is asked to do two jobs at once). +2. **`$SITE`/`$DC`-parameterized, never DC-hardcoded** -- the DOCFIX-151 convention + (`lib_net_select_dc "$DC"` / `lib_hosts_select_dc "$DC"`), or a positional `` argument + read by the same underlying selector (every `dc-*.sh` / `maas-*.sh` script in Section 2 takes + `` this way). This is the mechanism that lets one module body run DC0 today and DC1 + tomorrow without a copy-paste fork -- named explicitly as the anti-pattern to avoid in + `pass1-w4-module-planning.md` Section 4 item 1. +3. **Independently testable with its own harness** -- `tests//run-tests.sh` + (`docs/tool-index.md:27`, "65 scripts with their own `tests//` harness"; CLAUDE.md + "Delivery": "every script change ships with its `tests//run-tests.sh` harness green"). + Section 2's table shows this held for 32/34 procedure modules in the filtered set. +4. **Declared inputs, declared outputs, a stated exit contract.** The pattern is explicit in the + library layer already (`lib-validate.sh`'s 0/1/2/3/4 PASS/FAIL/HOLD/PASS_PENDING_MANUAL/SKIPPED + vocabulary) and echoed in every gate script (`geneve-encap-assert.sh`: "Exit: 0 all pass | 1 + any FAIL | 2 usage/precondition"). A procedure module's *inputs* are its CLI args + whatever it + reads live (never a baked table -- hard rule 2); its *outputs* are the mutation it performs (or, + in `check` mode, the PASS/FAIL verdict) plus what it leaves behind for the layer above to + consume (e.g. `dc-node-carve.sh` leaves a carved `br-ex`; `dc-node-v6-carve.py` needs that + `br-ex` to already exist -- an explicit inter-module input/output chain, not implicit state). +5. **Idempotent, safe to re-run.** Universal across the table's procedure modules: `check`/`apply` + split with `apply` DRY by default and `--commit` required to mutate + (`dc-region-topology.sh`, `dc-plane-ipam.sh`, `phase-00-maas-standup.sh`); `install` actions + that no-op when already correct (`dc-rack-net.sh`, `dc-mirror.sh`, `site-tailscale.sh` + explicitly say "idempotent; safe to re-run"). +6. **Composes onto the layer strictly below it, never reaches past it.** The layer-boundary rule + from `pass1-w4-module-planning.md` Section 2: an L3 module's live inputs come from L2's + output (MAC-pinned node VMs) or L1's output (a reachable MAAS region), never by dialing L0 or + L2's *infrastructure* directly (Section 4 makes this concrete for the container-elim case). +7. **Ships with a changelog entry + a revert, and is `repo-lint` clean** (CLAUDE.md "Delivery" -- + applies identically to procedure and IaC modules; not a procedure-only rule, restated here so + the contract is complete). + +**Precedent this contract is built on, not invented against:** the `$DC`-parameterized Stage +runbooks (DOCFIX-151, `docs/dc-dc-deployment-workflow.md:485-498`) already run this exact +contract at the *stage* granularity (one runbook file, `$DC`-selected, gated by repo-lint + a +harness sweep + a changelog + a branch-merge per stage close-out, +`docs/dc-dc-deployment-workflow.md:424-436`). A procedure module is the SAME contract one level +down -- the runbook is the composition of several procedure-module invocations in a stated order; +the module is the individually-testable unit the runbook calls. + +--- + +## 4. The IaC <-> procedure BOUNDARY, stated precisely + +**The boundary is identity, not orchestration.** OpenTofu owns everything up to *"a booted +libvirt domain exists, with its network identity (MAC per NIC, and any statically-assigned IP) +correctly wired to the right plane bridges."* The procedure layer begins at the first +live dial into that object -- an SSH session, a MAAS API call, a libvirt power query -- and by +design **never reads OpenTofu state**. It re-derives everything it needs from LIVE, independently +observable identity: MAC address, hostname, or a fresh API/SSH probe. This is not a convenience; +it is a stated design rule with its own incident history: + +> `lib-hosts.sh:6-11` -- *"WHY hostname-keyed (NOT system_id-keyed): MAAS system_ids are minted +> fresh on every (re-)enrollment... The stable identities are the hostname and the libvirt domain +> name, so every map here keys on hostname and the live system_id is resolved AT RUNTIME."* + +> `dc-node-v6-carve.py` -- *"EVERYTHING IS DERIVED FROM LIVE STATE -- there is no plane table +> here (hard rule 2): site membership <- the MAAS tag openstack-."* + +> `maas-node-power.sh:8-9` -- *"Domains are matched to MAAS machines BY MAC ADDRESS -- never by +> name, because MAAS assigns its own random hostnames at enlistment."* + +**Concretely, at each L1/L2 -> L3 handoff point in Section 2's table:** + +- **Node VMs (L2, `modules/node-vm`) -> MAAS commissioning (L3):** the handoff artifact is the + **MAC address baked into the libvirt domain XML** (IaC output) that MAAS discovers at + commissioning and `maas-node-power.sh` later matches power config against (procedure input). + No file, no state read, no data pipe crosses the boundary -- only a MAC address that both sides + independently observe. +- **Client VM / edges / voffice1 (L1, `modules/cloudinit-vm`) -> install procedures (L1/L3):** + the handoff artifact is a **reachable IP + a cloud-init-seeded initial state** (IaC output); + `site-headend-install.sh` / `dc-rack-net.sh` / `site-tailscale.sh` then dial in over SSH and + configure from there, keyed by hostname/IP resolved live (`lib-hosts.sh`), never by reading + `opentofu/*.tfstate`. +- **DC edge WAN (L1) -> egress verification (L5):** `dc-egress-check.sh` probes the live path + end to end; it does not consult tofu at all. + +**Why this rule matters here specifically -- it is the diagnostic for the container-elim's core +defect.** `pass1-w4-module-planning.md` Section 2 already names the ONE place this boundary rule +is currently *violated*: the inner root (`opentofu/vr1-dc0-substrate/main.tf`) dials OUT to a +`qemu+ssh` provider INTO the outer root's own `vvr1_dc0` output, because "a libvirt provider +cannot be configured from a resource created in the same apply" (`opentofu/main.tf:28-29`). That +is IaC reaching *into* IaC across a live-dial boundary that should only ever be crossed by a +procedure module -- L2's inner half is doing L2-to-L2 what only L2-to-L3 is supposed to do (dial +a live object by observed identity, not by direct provider coupling). **Option 1 removes the +violation structurally**: with one flat root, L2 is IaC end to end (module bodies only, no +cross-host provider dial), and the FIRST live dial into anything L2 produced is L3's MAAS +commissioning -- exactly where the boundary rule says it should be. This is the single clearest +argument, in module-decomposition terms, for why flattening also simplifies the module system +and not merely the topology (echoing `pass1-w4-module-planning.md` Section 2's own framing). + +**Boundary summary table:** + +| From (IaC, L0-L2) | To (Procedure, L3+) | Crossing artifact | Never crosses | +|---|---|---|---| +| `modules/node-vm` (L2) | MAAS commissioning (L3) | MAC address (observed both sides) | tofu state, module output vars | +| `modules/cloudinit-vm` (L1) | `site-headend-install.sh` / access scripts (L1/L3) | IP + cloud-init seed | tofu state | +| `modules/dc-planes` (L2) | `dc-plane-ipam.sh` / `dc-node-v6-carve.py` (L3) | plane CIDR/VLAN (both re-derive from live MAAS, per D-134/D-139) | tofu state | +| Any IaC module | Any L5 gate | the live object itself, probed independently | tofu state (`opentofu-validate.sh` is the ONE gate that DOES read tofu directly -- because it IS the IaC-layer's own gate, not a procedure-layer consumer) | + +--- + +## 5. Option-1's new tools, mapped onto the decomposition + +| New tool (Option 1) | Module kind | Layer | Placement / what's actually new | +|---|---|---|---| +| **Cross-DC host-isolation control** (the "(a)" control, `pass0-admin-report.md` Section 7a / `pass1-admin-report.md` Section 3) | **Gate (procedure)** -- confirmed in `pass1-admin-report.md` Section 3: "SEC-010's actual pattern -- a script-installed nftables control + `--check` + harness ... NOT an OpenTofu module" | **L5**, new artifact | A genuinely NEW script + `tests//run-tests.sh` + SEC-NNN row, on the `geneve-encap-assert.sh`/SEC-010 pattern one layer up. Recommended stage home: Stage 1 (host-scoped, not per-DC-apply-scoped), re-verified at each per-DC apply's close and at Stage-5 live verify. **This is the one wholly new procedure-module BODY this pass's decomposition requires** -- everything else re-targets existing bodies. | +| **Client-VM standup** (`vr1-dcN-client`, D-138 role) | **Split: IaC instantiation + procedure configuration** | **L1 (IaC)** for the VM itself; **L1/L3 (procedure)** for what runs on it | The VM is a NEW *instance* of the EXISTING `modules/cloudinit-vm` module type (same body Office1's `voffice1` and the DC edges already use, per `pass1-w4-module-planning.md` Section 3) -- **no new IaC module**, so it rides `opentofu-validate.sh`'s existing coverage. What lands ON it once booted is EXISTING procedure modules RE-TARGETED, not new bodies: `site-headend-install.sh --role rack` (rack-controller portion, sans the now-dead `--host-nodes` bootstrap-gate duty), and -- pending the OPEN Section-6-item-3 placement ruling from Phase 0/1 -- `dc-rack-net.sh` (D-131 forwarder) and the artifact-mirror scripts (`dc-mirror.sh`/`dc-cache-proxy.sh`/`dc-snap-proxy.sh`). This is a call-site change (which host these scripts SSH into), not a module-body change -- consistent with the contract's "declared inputs" (the target host is an input, not baked in). | +| **Teardown-primitive** (module/root-scoped group-destroy, re-earning D-122's one-command site-down) | **Procedure (new)** | **L2/L5 boundary** -- a procedure module that WRAPS an IaC destroy (invokes `tofu destroy -target=...` against the flat root's DC-scoped module set, or drives a scripted `virsh destroy` loop over the roster `lib-hosts.sh` derives) | **The second wholly new procedure-module BODY.** Does not exist today (`pass1-admin-report.md` Section 6 item 1 / Section 6 item 6 names it as owed, distinct from the emergency `virsh destroy` loop). Its exact shape (root-scoped `-target` set vs. scripted domain-group destroy) depends on the OPEN root-topology fork (`pass1-admin-report.md` Section 7 item 1, W2.1's domain) -- this worker does not resolve that fork, only places the resulting artifact's layer. Must carry the same contract as Section 3: `$DC`-parameterized, its own harness, idempotent-on-already-torn-down state. | +| **Re-homed rack-controller / D-131 forwarder / artifact-service (`.4`)** | **No new module -- re-target of existing L3 procedure modules** | **L3**, placement OPEN | Confirmed by both `pass0-admin-report.md` Section 6 item 3 and `pass1-w4-module-planning.md` Section 3: this is a call-site/placement decision (client VM vs. `vr1-dcN-maas-01` vs. retire-with-evidence), not a module-body change. `dc-rack-net.sh`, `site-headend-install.sh --role rack`, and the mirror/proxy scripts already take `` as a parameter and run over SSH to whatever host is named -- the decomposition table in Section 2 shows every one of them already meets the procedure-module contract (Section 3) independent of which host wins. **This worker's contribution is confirming there is no tooling gap here** -- the gap is a ruling, not a build. | + +**Net new procedure-module bodies this pass's decomposition surfaces: exactly TWO** -- the (a) +cross-DC isolation gate and the teardown-primitive. Everything else Option 1 needs is either (i) +a new instance of an existing IaC module type (the client VM), or (ii) an existing procedure +module re-targeted at a new host via its existing ``/host-argument parameterization (no +body change). This matches `pass1-w4-module-planning.md` Section 1's framing: "the container-elim +does not need to invent module mechanics, only re-home." + +--- + +## 6. Open items (not resolved by this worker; feed the Phase-2 administrator / Phase 4) + +1. **`lib-identity.sh`'s harness status** -- not independently verified this session; flagged for + W2.2 (its assigned dimension, `lib-hosts`/`lib-net` containment-keyed values). +2. **`maas-fabric-prune.sh` / `maas_fabric_classify.py` have no harness** -- pre-existing gap, + unrelated to container-elim causally, but the harness-discipline principle this pass restates + (Section 3 item 3) would be inconsistent if silently excluded; logged for the Phase-2 + administrator to route (build-a-harness item vs. accept-as-a-named-exception). +3. **The teardown-primitive's exact shape** depends on W2.1's root-topology design (the OPEN fork + from `pass1-admin-report.md` Section 7 item 1) -- this worker places its LAYER, not its final + form. Feed forward to whichever worker/administrator resolves the fork. +4. **The rack-controller-remainder placement ruling** (Section 5 row 4) is the actual blocker for + writing the FINAL L3 change-set (which scripts get new call-site arguments); this worker + confirms readiness (no tooling gap) but cannot pick a value (hard rule 2). +5. **`opentofu/modules/*`'s own decomposition** (W2.1's domain) was consulted only at the + boundary (Section 4) -- this doc does not re-derive the IaC-module inventory or recommend + root-splitting; that is W2.1's report to write. + +--- + +## 7. Verification note + +Every script's role/kind/layer classification in Section 2 is grounded in a direct header read +this session (paths cited); harness presence was checked live via `ls tests/*/` cross-referenced +against each script's basename, not assumed from the tool-index's aggregate counts. The two +"NO harness" flags and the `lib-identity.sh` open item are the only claims in this doc not +independently confirmed against a second source. No mutation performed; no live cloud state +queried; findings are LOGGED only, per the pass's read-only-planning charter. diff --git a/docs/audit/container-elim-pass/pass3-admin-report.md b/docs/audit/container-elim-pass/pass3-admin-report.md new file mode 100644 index 0000000..1e78903 --- /dev/null +++ b/docs/audit/container-elim-pass/pass3-admin-report.md @@ -0,0 +1,323 @@ +# Pass 3 -- ADMINISTRATOR REPORT: tests review (container-layer elimination) + +**Author:** the Phase-3 administrator (multi-agent pass, `SCOPE-AND-EXECUTION-PLAN.md` +Section 4). **Date:** 2026-08-09. **Inputs:** `pass3-w1-harnesses-gauntlet.md`, +`pass3-w2-gates.md`, `pass3-w3-new-tests.md` -- read in full, adversarially cross-checked +against repo ground truth (checks logged in Section 1). Baseline consumed: +`pass0-admin-report.md` incl. Section 7a (**Option 1 CONFIRMED**; **cross-DC handling (a) +CONFIRMED**; MAAS region stays on `vr1-dcN-maas-01`), `pass1-admin-report.md` (planning +change-set; the (a)-control consolidated spec, Section 3), `pass2-admin-report.md` (tools +change-set; the THREE isolation concerns, Section 2; the 13 owed artifacts, Section 5). +READ-ONLY synthesis; no mutation; findings are LOGGED, not executed. All dispositions are +proposals feeding Phase 4 (GA-R5); contingent items are marked with their blocking ruling, +never pre-picked. + +Repo discipline governing every row: an assertion must be provably able to FAIL (GA-R6, +`SKILL.md:321-333`); assert on CONTENT not existence; REFUSE on unrecognised input, never +pass silently; a fix that makes an assertion stale gets REPLACED with the new invariant, +never deleted to go green. + +--- + +## 1. Reconciliation / adversarial-check results (run against repo ground truth this session) + +1. **The two harness universes UNIFY cleanly -- with two count corrections, no + contradictions.** + - **W3.1's "13 of 103 affected" needs granularity:** its own verdict line ("2 clean + RETIRE, 2 contingent RETIRE, 6 CHANGE, 1 contingent CHANGE") sums to 11 + verdict-blocks; the survey table names 14 harnesses (one, `netem-link`, a declared + grep false-positive). Reconciled: **13 harnesses carried a survey hit; 8 of them + carry CHANGE/RETIRE verdicts; 5 are graded STAY.** The unified table (Section 2) is + presented at verdict-block granularity so the counts fall out of it. + - **W3.3's summary line "6 new + 5 extend-existing + 2 ride" does not decompose + against its own rows -- CORRECTED.** By its own per-artifact dispositions: + **7 new-build** (#12 dc-site, #2 (a) control, #1 teardown primitive, #11 power-key + mitigation, #13 D-131 evidence, #4 R7 checklist, #5 MAAS record release/delete), + **3 extend-existing** (#3 SEC-010 successor -> `tests/site-headend-install` + (conditional: new harness if the extraction becomes its own script), #7 FIT + extension -> `tests/dc-dc-whole-host-budget`, #9 NetBox migration -> + `tests/dc-rack-mgmt-import`), **3 ride** (#6 rides #1's harness, #8 rides + `dc-site`'s MAC invariant, #10 rides `geneve-encap-assert` unchanged). 7+3+3 = 13; + all owed artifacts accounted for. (The phase prompt inherited the 6/5/2 figure from + W3.3's own summary line -- this is a worker-count correction.) + - **Overlap check (the load-bearing one):** of the 3 extend-existing, **#3 and #9 map + INTO W3.1's affected set** (`site-headend-install` Section-8 SEC-010 sub-case; + `dc-rack-mgmt-import` vvr1 pins) -- the same edit seen from two sides, MERGED in + Section 2, not double-counted. **#7 crosses the universe boundary knowingly:** + `dc-dc-whole-host-budget` sits OUTSIDE W3.1's 13 (its Sec 3.4 supplementary note + routes it to W3.2/W3.3 explicitly), and W3.3 #7 picks it up -- the one clean + boundary crossing, now inside the unified set (Section 2 row C8). No worker + contradicts another anywhere in the two universes. + +2. **The power-key (#11) critical-path dependency is CONSISTENT across all three + workers** -- W3.1 (edit-list item 4: `dc-selector` + `maas-region-power-key` edited + ONLY together with #11's ship, from the ruled URI/key shape), W3.2 (P4's power-address + assertion and P5's new register row BLOCKED on #11's mechanism; Sec 4 risk 1), W3.3 + (Sec 2.4: the harness + artifact are a stated PRECONDITION for the `lib-hosts.sh` + re-derivation and every call-site literal). Stated plainly as the sequencing + constraint in Section 4: **owed artifact #11 blocks a six-item cluster of test + edits**, and the interim RED those harnesses will show once `lib-hosts.sh` changes is + the DESIRED fail-loud state, not a defect to green out. + +3. **Gate homes are COHERENT and each assertion is failable** (Section 3 verified + row-by-row against W3.2 Sec 1-3 and W3.3 Sec 2): (i) Stage-1 gate + cloud-assert A11a -- + enumerates the live bridge set and REFUSES if fewer than the full plane count + resolves (SEC-010's fail-open lesson generalized), plus W3.3's + no-host-address-on-bridge and does-not-globalize failing fixtures; (ii) preflight + **P10** (P10 confirmed the next free slot; P6 is the reserved non-executing reminder + block) -- host-bound `--check` on the P7 model, REFUSE off-host, REFUSE on an absent + transit interface; (iii) P4-dependency + A11b -- the pass verdict is tied to a + NEGATIVE test (the key CANNOT reach the wrong domains) plus W3.3's positive-coverage + diff against the DC's own roster and a must-not-over-restrict control. None is an + existence-only check; none is a checker-that-cannot-fail. **One administrator spec + amendment (recommendation grade): P8's root enumeration must come from a DECLARED + root list, not a filesystem glob** -- a glob loop is a checker-that-cannot-fail for a + missing/renamed root (it silently drops out of coverage instead of FAIL/WARN). + Section 3 carries it. + +4. **Two coverage gaps the workers left dangling -- CLOSED BY MEASUREMENT this + session:** + - **`scripts/pre-flight-checks.sh` body sweep** (W3.2 Sec 4 risk 3 flagged it read by + neither worker): grep for W3.1's exact term set + (`vvr1|qemu+ssh|172.31|host-nodes|expose_nested|inner root|bootstrap gate|Model B| + node_host`) returned **zero hits (rc=1)**. Its container-layer exposure is entirely + INDIRECT via sourced `lib-hosts.sh` (`pre-flight-checks.sh:44-45`) -- the + dc-selector-covered surface. No direct edit owed to its body; the NEW post-#11 + content assertion (Section 3) is net-new, homed in `tests/pre-flight-checks` + (harness exists, verified `ls tests/`). + - **`lib-identity.sh` harness status** (pass2 Section 6's unclosed worker handoff): + ANSWERED by W3.1's no-hit list -- `tests/lib-validate` exercises the + `lib-hosts`/`lib-net`/`lib-identity` plumbing generically. Closed. + +5. **One pass2 Phase-3 handoff was addressed by NO worker -- carried OPEN, not + dropped:** the `maas-fabric-prune.sh`/`maas_fabric_classify.py` harness gap + (pre-existing; pass2 Section 6 routed "build vs accept-as-named-exception" to this + phase). Verified this session: `tests/maas-fabric-prune` does not exist. None of + W3.1/W3.2/W3.3 mentions it. Routed to Phase 4 as a named decision exactly as pass2 + framed it (Section 7 open item 8). It is container-elim-ADJACENT only (a + pre-existing gap), so parking it there loses nothing from this change-set. + +6. **Next-free SEC number: SEC-034 CONFIRMED** -- with the verification shape stated + honestly. `ledger-scan.sh` computes next-free for D/DOCFIX/BUNDLEFIX only (run this + session: D next-free=144, DOCFIX=214, BUNDLEFIX=059), NOT for SEC rows; the SEC + confirmation rests on the direct ledger grep: highest row in + `docs/security-ledger.md` = **SEC-033**, and the sole repo-wide SEC-034 mention is + W3.2's own doc. **2-or-3 new SEC rows are owed** ((a) control; #11 mitigation; and + concern (ii)'s row, which pass2 deliberately left as a "SEC-row disposition" -- it + may land as a SEC-010 amendment rather than a new number). Nothing assigned here; + delivery greps next-free at mint time per repo numbering discipline. + +7. **Denominator + doc-currency:** `tests/HARNESS-MANIFEST` lists 103 harnesses + (`tests/` holds 104 entries including the manifest itself) -- W3.1's denominator + confirmed. CLAUDE.md's "98 harnesses" line is stale vs 103 -- logged as a + doc-currency nit OUTSIDE this pass's change-set (rides the next CLAUDE.md touch; + no action here). + +8. **No inferred-value violations found** in the three worker docs. Spot-checked + load-bearing cites (`preflight.sh` P8 root literal `:284,298`; `cloud-assert.sh` + A0-A10 shape; `site-headend-install.sh` SEC-010 writer/check line ranges; + `node-vm/variables.tf:38-62` MAC validation; `security-ledger.md:21` fail-open + hardening note; the changelog cites for D-131 asymmetry) resolved to the cited + lines. Every unknown (root naming, #11 mechanism, (a) rule set, P10's body, D-127's + client-VM autostart value) is marked OWED in the source docs rather than guessed. + +--- + +## 2. THE UNIFIED TESTS CHANGE-SET (one table, three dispositions, no double-counting) + +Blocker legend: `ready` = editable/buildable at delivery with no ruling owed; +`#11` = blocked on the power-key mitigation design (Section 4); `root-naming`, +`rack-ret` (rack-controller retirement + owed live re-measure), `D-131` (retirement +ratification + owed live re-measure), `netbox` (NetBox migration design, owed #9), +`mechanism` (the artifact's own Phase-4 mechanism choice), `root-fork` (root-topology +ratification). + +### A. EXISTING -- CHANGE (9 verdict-blocks across 9 harnesses) + +| # | Harness (cases) | Edit | Blocker | Failing-direction fixture | +|---|---|---|---|---| +| A1 | `opentofu-validate` T14/T15 | Re-point autostart pins at the flat root's file; **ADD a new case for the `vr1-dcN-client` VM's autostart value** (D-127's table predates the client VM -- value is an OWED ruling, not inferred) | root-naming (+ the D-127 client-VM ruling for the new case) | Keep exact-value asserts (edge=true / node=false); a flipped bool or wrong path must redden | +| A2 | `node-vm` T8-T11 | Re-point hardcoded `INNER=` path at the flat root's file; count invariant likely unchanged (12 `macs=[` / 72 MACs -- the client VM is `cloudinit-vm`, not `node-vm`) -- confirm at build, do not assume | root-naming | Wrong count or missing MAC list must still redden (this harness has gone red twice before, `tests/node-vm/run-tests.sh:68-81`) | +| A3 | `site-headend-install` Section-8 SEC-010 sub-case **= owed #3's harness half (MERGED W3.1/W3.3 row)** | Re-target the extracted role-agnostic installer subcommand (both ends: client VM + voffice1); **MIGRATE every proven SEC-010 assertion** (transit-if existence check, br_netfilter never-global, idempotent declare-then-delete reload) -- a migration-completeness grep of the new location for every old trap-string is itself a required test | ready once #3's extraction shape is picked (subcommand -> extend this harness; own script -> new harness inheriting all of it) | Both-ends fixture (fake client role + fake voffice1 role, BOTH get the drop); wrong-role interface-name fixture must FAIL; NIC-naming trap carried (dc0 live was `enp1s0`, not `mgmt`) | +| A4 | `dc-selector` power-address rows (`:204-247`) | New values dial vcloud's own libvirtd -- URI shape + whether the FROM_OFFICE1/FROM_DCREGION split survives are #11 design outputs | **#11** | Edit ONLY together with A5, from the ruled URI shape; interim RED once `lib-hosts.sh` changes is correct fail-loud, not a defect | +| A5 | `maas-region-power-key` URI assertions (`:67,82,102,108`) | Same dependency, same ruled shape, same edit session as A4 (or they silently diverge) | **#11** | Same as A4 | +| A6 | `dc-rack-mgmt-import` vvr1 pins (`:77-78`) **= owed #9's harness half (MERGED W3.1/W3.3 row)** | Rename-in-place vs full-concept retirement is the NetBox-migration design's call -- do not pre-pick; when it lands, add the decommission-completeness assertion | netbox (+ rack-ret for the concept question) | A stale `vvr1-dcN` device record surviving import must FAIL; the D-124 transit-scheme pins (`:73,79` + `d124-transit-seed`) STAY -- mesh facts, not containment facts | +| A7 | `maas-profile-assert` office1-profile fixture (`:41,68,72,80`) | Drop the two `vvr1-dcN` rows from the simulated machine list; the client VM does NOT replace them (not MAAS-carved, pass2 3.2) | rack-ret | Post-retirement roster fixture; lowest-risk of the CHANGE set, independent of the #11 cluster | +| A8 | `dc-dc-whole-host-budget` **(W3.3 #7; the one universe-boundary crossing -- outside W3.1's 13 by its own routing)** | Extend with the 3 utility-node classes + the artifact-service disk-sizing branch + the flat model; the old Model-A/B comparison cases must not remain the only coverage (they stay green forever against a retired topology) | ready (spec = owed #7; capacity re-measure owed at delivery) | A roster total EXCEEDING the measured host budget must FAIL, not round/omit a class | +| A9 | `tests/pre-flight-checks` -- NEW case post-#11 | Assert the LIVE power-address in use matches the mitigation's issued form (restricted-key/wrapper shape), not a raw containment-shaped URI -- the regression detector for concern (iii) | **#11** | A containment-shaped (`qemu+ssh://...@172.31.0.2`-class) or unrestricted URI fixture must FAIL | + +### B. EXISTING -- RETIRE (5 verdict-blocks across 3 harnesses) + +| # | Harness (cases) | Disposition | Blocker | Failing-direction handling (retire rows still need one) | +|---|---|---|---|---| +| B1 | `opentofu-validate` T13 (D-127 containment autostart pin) | RETIRE -- object deleted, not renamed; D-127 boot-matrix comment block updated in the same edit | ready | **Recommend the negative-assertion case** ("no `vvr1` domain block exists in `opentofu/main.tf`") over W3.1's accept-residual-gap alternative -- a stray re-add then still reddens; grep-checkable comment-currency fixture rides along | +| B2 | `site-headend-install` `--host-nodes` block (~15 cases: `node_host_setup()`/`node_host_check()`, D-125 WAN-bridge asserts, hint-neutrality) | RETIRE wholesale (dead code, D-125 has no successor) -- EXCEPT the SEC-010 sub-case, which is A3 | ready (when the code deletion ships) | Re-derive every surviving arg-contract rc from the NEW parse contract (hazard H2) -- do not let a stale `-> 2` case pass on an unrecognized-flag coincidence | +| B3 | `site-headend-install` Section 6 (`--role rack`) | CONTINGENT retire; if ratified, a new case asserts `--role rack` is REFUSED/removed; if rejected, stays as-is | rack-ret | The REFUSE case IS the replacement failing-direction fixture -- hold unedited until the ruling lands | +| B4 | `dc-rack-net` LEGS cases (T3,T5,T13,T15,T17) | RETIRE (stronger-settled: structural consequence of Option 1 itself -- no flat VM is a libvirt host with own bridges) | ready at script retirement | Append-only bias: file stays in git history; removal via `--record-manifest` + changelog with revert | +| B5 | `dc-rack-net` DNS-forwarder cases (T4,T9,T16) + its 8 hygiene cases | CONTINGENT retire wholesale with `dc-rack-net.sh` if D-131 retirement ratifies; if D-131 is kept for a DC, the DNS half survives re-targeted | D-131 | Same manifest/changelog discipline as B4; note `DNS_UPSTREAM="10.10.0.20"` is a pre-existing stale defect either way | + +**STAY, named so they are not re-litigated:** `opentofu-validate` T8-T10 (S3 module-body +cases); `node-vm` T1-T7/T12-T15; `site-headend-install` Sections 1-5,7; `maas-node-power` +(URI fixture is an opaque pass-through arg -- cosmetic refresh optional); `preflight` +pending-change fixture; `geneve-encap-assert` (all cases -- its post-build live re-run is +a new INVOCATION, owed #10, not a harness edit); `site-baseleg` (guard already +forward-compatible; comment re-cite is doc-currency); `cloudinit-vm`; `d124-transit-seed`; +`netem-link` (declared grep false-positive); the 90 no-hit harnesses. + +### C. NEW-BUILD (7 harnesses, one per owed artifact; specs in `pass3-w3-new-tests.md` Sec 2-3) + +| # | Owed artifact | Harness (model) | Blocker | Signature failing-direction fixture | +|---|---|---|---|---| +| C1 | #12 `modules/dc-site` | `tests/dc-site/run-tests.sh`, static fixture `.tf` trees (`opentofu-validate` `--static-only` model) | module build (spec ready) | 5-plane tree FAILS count; 11/13-node roster FAILS; 2 client-VM calls FAILS exactly-one; empty `interface_macs` on post-apply tree FAILS; v4-only CIDR on a D-139 v6 plane FAILS; `mtu=1500` FAILS; hardcoded DC literal FAILS site-token check | +| C2 | #2 the (a) cross-DC host isolation control | offline fixture-file (`geneve-encap-assert` model); name minted with its SEC row | mechanism | Forward-accept between a dc0 and a dc1 bridge FAILS; **a host ADDRESS on a plane bridge FAILS (distinct failure mode)**; empty/missing ruleset FAILS (never clean); absent-interface-name rule FAILS (SEC-010 fail-open lesson); bare `policy drop` FAILS does-not-globalize | +| C3 | #1 teardown primitive (+#6 emergency lever rides the same fixture library) | stateful-fakebin (`phase-00-teardown-d061` model) | root-fork | A dc1 domain in a dc0-targeted set FAILS before any destroy; a shared-outer object leaking in FAILS; a missing expected domain FAILS completeness; EMPTY resolved set REFUSES (rc=2), never "nothing to do" success; failed canary blocks the group destroy | +| C4 | #11 power-key mitigation | offline rendered-ACL/`authorized_keys`/polkit parsing (`site-headend-install` grep-the-artifact model) | mechanism -- **and this harness is itself the precondition for the A4/A5/A9 cluster** | dc0 key reaching a dc1 domain FAILS; reaching `voffice1` FAILS; mirror both directions; a wildcard matching nothing FAILS as under-specified; excluding a dc0-own domain FAILS the must-not-over-restrict control | +| C5 | #13 D-131 retire-evidence checker | offline dig-capture fixtures; region IPs read from `lib-hosts`/`lib-net`, never duplicated | ready (fixtures exist in the two cited changelogs) | Resolution that SUCCEEDS via the forwarder alias IP FAILS (resolver IDENTITY asserted, not mere success); dc0's measured evidence PASSES; **dc1's current config is a standing-RED case until the live retirement happens** | +| C6 | #4 R7 credential-revocation checklist | offline; mock `vm-secret-locations` rows | ready | An unlisted/orphan row for the retiring host class that enumeration misses FAILS (SEC-027's "an unlisted location is not audited") | +| C7 | #5 MAAS record release/delete (+ rack-controller decommission) | fakebin `maas` (`dc-egress-check`/`phase-00` model) | ready | A surviving machine record OR the surviving rack-controller/`primary_rack`/DHCP reference after "release" must not report clean | + +**Rides (no build):** #6 -> C3's fixture library; #8 MAC re-measure -> C1's MAC invariant; +#10 geneve/jumbo -> existing `geneve-encap-assert` verbatim, new invocation point in the +runbook only. + +--- + +## 3. GATE-HOME MAP -- the three isolation controls + preflight/cloud-assert changes + +| Control / gate | Home(s) | Failable assertion (content-based, REFUSE-on-unrecognised) | +|---|---|---| +| **(i) cross-DC host isolation (the (a) control)** | **NEW Stage-1 gate** (host-scoped -- NOT a `DC=`-scoped preflight gate, which would double-run or silently check only the last DC); re-run at each per-DC apply's close and at **cloud-assert A11a** (periodic: post-deploy/restart/pre-change/post-incident) | `--check` on vcloud enumerates the live bridge set for BOTH DCs' six planes (from `lib-hosts.sh` conventions), asserts FORWARD denial between every dc0-tagged/dc1-tagged bridge pair, and **REFUSES if fewer than the full plane count resolves to a live interface**; rule COUNT/hash matched against the artifact's own expected state, not "a table named X exists". Ordering invariant: installed + verified **before the first flat apply of EITHER root** | +| **(ii) SEC-010 transit-drop successor** | **NEW preflight P10** (P10 confirmed next free; P6 reserved), DC-scoped, host-bound on the P7/`node_host_check()` model; one role-agnostic installer/checker covers BOTH ends (client VM + voffice1) | `nft` table present keyed to the CLIENT VM'S OWN re-measured transit interface (NIC-naming trap: dc0 live was `enp1s0`); mirrored voffice1 check; REFUSE if the keyed interface does not exist | +| **(iii) power-key mitigation** | **Split home:** (a) P4 dependency -- before P4 trusts a power-address literal, verify it matches the mitigation's issued shape (the A9 case); (b) **cloud-assert A11b** -- standing re-verification (the exposure exists whenever libvirtd is up, not only at deploy) | Verdict tied to a NEGATIVE test: connecting with either DC's region key CANNOT enumerate/control domains outside that DC's own roster (mechanism-dependent concrete form: forced-command rejects raw `virsh list --all`, or polkit ACL covers exactly the `vr1-dcN-*` roster); an existence-only key check is explicitly ruled out (GA-R6) | +| **rack-retirement / D-131 evidence** | ONE-TIME build-step evidence capture inside the promoted Part-A release/delete step (owed #5/#13) -- NOT a standing gate | The C5 checker: resolver-identity dig per fresh 10.13 region, captured into the build changelog; dc1's asymmetry is a standing-red case until closed | +| **P8 substrate drift** | Extend from the single hardcoded `opentofu/` path (`preflight.sh:284,298`) to **loop over every post-flatten root** (shared-outer + each per-DC-flat root under (B)), each printing which root it evaluated; the voffice1 "inner state lives elsewhere" WARN branch goes dead for the DC half | Existing pending-FAILs/refresh-WARNs logic is already content-based and stays. **Administrator amendment: the root list must be DECLARED, not globbed** -- a glob loop cannot fail on a missing/renamed root (it silently drops coverage); a declared list FAILS/WARNs. Blocked on root-naming ratification | +| **P5 creds-matrix** | DATA change, not code: `rack`-class register rows re-point to the `client` class; NEW rows owed for #3/#11 key material when minted -- else P5 silently stops covering the new credential-bearing hosts (SEC-022's exact coverage-gap class) | Already failable by construction (D-137 ruling 1); flagged as a delivery dependency on #3/#5/#11, not a script edit | +| **P4 / P9 / A11 placement** | P4: no body edit owed (measured clean, check 4) -- the A9 case is the net-new half. P9 (`dc-egress-check`): invocation-host literal only, re-points to the ruled B.5 host (client VM the structural candidate). **A11's home (cloud-assert vs a dedicated `isolation-assert.sh`) is recommendation-grade -- Phase 4 confirms** | -- | +| Unchanged | P1, P2 (D-143 axis only), P3, P6, P7 (already host-indirect via `host-identity` -- the pattern P4/P5 should imitate), cloud-assert A0-A10 (zero container-layer assumptions, verified by W3.2 full read) | -- | + +SEC rows: next-free **SEC-034** (check 6); 2-or-3 rows owed ((a) control, #11, and +concern (ii)'s disposition -- possibly a SEC-010 amendment); minted at delivery, next-free +re-grepped then. + +--- + +## 4. THE CRITICAL-PATH SEQUENCING CONSTRAINT (stated plainly) + +**Owed artifact #11 (the power-key blast-radius mitigation) is a BLOCKER for a +six-item cluster of test work.** Until its mechanism is chosen (Phase 4) and its artifact ++ C4 harness ship, NONE of the following may be written -- a guessed URI/key value here is +the false-green mint this repo's inferred-value rule exists to prevent: + +1. `lib-hosts.sh` `VIRSH_POWER_ADDRESS_FROM_OFFICE1`/`_FROM_DCREGION` re-derivation + (the tools-side edit the tests then pin); +2. `dc-selector` power-address rows (A4); +3. `maas-region-power-key` URI assertions (A5) -- A4/A5 edited TOGETHER, same session, + same ruled shape; +4. the new `tests/pre-flight-checks` P4 content case (A9); +5. cloud-assert A11b's concrete body; +6. the P5 register row for the mitigation's key material. + +Order: Phase-4 mechanism choice -> #11 artifact + C4 harness -> then the cluster, as one +delivery unit. **Interim state is deliberate:** A4/A5 go honestly RED the moment +`lib-hosts.sh` changes -- that fail-loud RED is wanted; do not "fix" it early. +Independent of this cluster (may proceed in parallel once their own blockers clear): +A1/A2 (root-naming), A7 (rack-ret), B1/B2 (ready), C1/C5/C6/C7 (ready/spec-ready). + +--- + +## 5. HARNESS-EDIT HAZARDS (delivery-time warnings, consolidated from W3.1 Sec 3) + +- **H1 -- the plausible-URI swap (sharpest risk in the survey):** swapping a + plausible-looking new power URI into A4/A5 ("just point it at vcloud") before #11's + mechanism exists produces a harness GREEN against a value nothing enforces -- worse + than the current honest RED. Only the ruled shape, only both harnesses together. +- **H2 -- the rc=2 coincidence:** if `--host-nodes` is REMOVED as a flag, a stale case + like `--host-nodes without --role rack -> 2` can keep passing on the + unrecognized-flag exit code, silently changing what the assertion proves. Re-derive + every expected rc from the NEW arg-parse contract; never assume rc=2 still means what + it meant. +- **H3 -- delete-to-go-green:** a stale hardcoded path in A1/A2 fails LOUD (safe); the + hazard is a sloppy fix that deletes the case instead of re-pointing it, quietly + dropping D-127 autostart-drift / MAC-count regression coverage (this class has gone + red twice before, `tests/node-vm/run-tests.sh:68-81`). Replace, never delete to go + green. +- **H4 -- the harness that stays green against a dead topology:** + `dc-dc-whole-host-budget`'s Model-A/B math keeps passing indefinitely after the + topology it models is gone; nothing forces a re-run until #7 ships. A8's extension + must land WITH the redeploy, and the old cases must not remain the sole coverage. +- **H5 -- the retirement residual gap:** retiring T13 leaves a stray `vvr1` autostart + re-add uncaught unless B1's negative-assertion case is added (recommended over + accept-and-declare). +- **H6 -- glob-based gate enumeration (administrator addition):** P8's root loop -- and + any future multi-root/multi-host gate -- must iterate a declared list; a glob cannot + fail on an absent member. + +--- + +## 6. THE PER-MODULE HARNESS CONTRACT (adopted from W3.3 Sec 4 -- the template every +Section-2 row follows) + +Adopted unchanged as the Phase-4 backbone; every principle grounded in an existing repo +precedent (cited in `pass3-w3-new-tests.md` Sec 4): **prove-it-can-fail** (every assertion +ships a fixture engineered to redden it); **assert the ARTIFACT, not the intent** (the +installed rule / rendered file / deployed value, never a comment or self-report); +**stage-assert-then-promote** (mutating subjects write to staging, assert staged, then +promote); **join the workspace to the deploy input** (byte/value equality between +generator output and what the consumer actually reads); **fail-closed on +absent/empty/unreachable input** (incl. the interface-name fail-open class); +**offline/fixture-driven by default** (fakebin or captured text; a separately-named LIVE +gate re-run proves the deployed state -- two tiers, never conflated); +**`$SITE`/`$DC`-parameterized** (fixtures exercise both DCs or a synthetic site); +**standard exit contract** (`lib-validate.sh` 0/1/2/3/4); **delivery discipline** (own +`tests//run-tests.sh`, changelog with revert, repo-lint clean -- no new-module +exception). + +--- + +## 7. Settled vs OPEN + +**SETTLED at this phase (do not re-open):** the unified change-set decomposition +(9 existing-change / 5 existing-retire verdict-blocks / 7 new-build harnesses / 3 rides; +count corrections in check 1); the gate-home map (Stage-1 for (i), P10 for (ii), P4+A11b +split for (iii), P8 declared-list loop, rack-evidence as one-time); the #11 sequencing +constraint (Section 4); SEC-034 next-free; the harness contract template; +`pre-flight-checks.sh` measured clean of direct container-layer literals; the +lib-identity handoff (covered by `tests/lib-validate`); `geneve-encap-assert` and +cloud-assert A0-A10 confirmed topology-agnostic. + +**OPEN -- carried to Phase 4 (rulings; operator rules, GA-R5):** +1. Root NAMING (blocks A1/A2 path re-points; pass2 Section 6 item 2). +2. Rack-controller retirement ratification + the owed live re-measure (blocks A7, B3; + shapes A6's concept question). +3. D-131 retirement ratification + live re-measure (blocks B5; C5's dc1 case stands RED + until the live retirement). +4. **The #11 mechanism choice** (restricted key / wrapper / polkit ACL) -- unblocks the + Section-4 cluster; its SEC row minted with it. +5. The (a) control's concrete nftables rule set + SEC number (C2's fixtures are + mechanism-agnostic until then). +6. A11's home: fold into cloud-assert vs a dedicated `isolation-assert.sh` (W3.2 Sec 4 + risk 5 -- recommendation-grade either way). +7. NetBox-migration design (owed #9) -- rename-in-place vs concept retirement for A6. +8. **`maas-fabric-prune` harness gap** (pre-existing; no Phase-3 worker addressed it, + check 5): build vs accept-as-named-exception -- a named Phase-4 decision. +9. D-127 client-VM autostart value (A1's new case needs the ruled value). +10. The teardown primitive's fixture shape rides the root-fork ratification (C3). + +**Doc-currency nits (ride other edits, no action here):** CLAUDE.md "98 harnesses" vs +103; `opentofu-validate` header "416 GiB" comment; `maas-node-power` cosmetic URI +fixture refresh. + +--- + +## 8. Verification note + +Author = "the administrator" (no model name asserted, operator instruction). Direct +measurements this session: `grep -oE 'SEC-[0-9]{3}' docs/security-ledger.md` (highest = +SEC-033) + repo-wide SEC-034 grep (sole hit = W3.2's own doc) + `bash +scripts/ledger-scan.sh` (D/DOCFIX/BUNDLEFIX next-free; confirms it does not compute SEC); +`grep -nE '' scripts/pre-flight-checks.sh` -> zero hits (rc=1) + +lib-hosts sourcing at `:44-45`; `ls tests/` + `wc -l tests/HARNESS-MANIFEST` (103; 104 +entries incl. the manifest); `tests/maas-fabric-prune` absent; `tests/pre-flight-checks` +and `tests/preflight` both exist (distinct). The three worker docs and three prior admin +reports were read in full; worker citations were spot-checked at their load-bearing +points, not re-derived wholesale. READ-ONLY; findings LOGGED only; nothing executed. diff --git a/docs/audit/container-elim-pass/pass3-w1-harnesses-gauntlet.md b/docs/audit/container-elim-pass/pass3-w1-harnesses-gauntlet.md new file mode 100644 index 0000000..f2bb53d --- /dev/null +++ b/docs/audit/container-elim-pass/pass3-w1-harnesses-gauntlet.md @@ -0,0 +1,143 @@ +# Pass 3 -- W3.1: harnesses + the gauntlet that assume the container layer + +**Author:** W3.1 (Phase-3 worker, `SCOPE-AND-EXECUTION-PLAN.md` Section 4). **Date:** 2026-08-09. +**Method:** surveyed all 103 harnesses in `tests/HARNESS-MANIFEST` / `tests/*/run-tests.sh` via +targeted greps (`vvr1`, `qemu+ssh`, `172.31`, `--host-nodes`, `expose_nested_virt`, `Model B`, +`node_host`, `inner root`, `bootstrap gate`) plus full reads of every harness a grep hit +implicated, cross-checked against `pass0-admin-report.md` / `pass1-admin-report.md` / +`pass2-admin-report.md` (which artifacts are owed/blocked/recommendation-grade). READ-ONLY; +no mutation; nothing here was executed. All CHANGE/RETIRE items are proposals feeding Phase 4; +several are explicitly CONTINGENT on Phase-4 rulings still open per pass2 Section 6/7 (root +naming, rack-controller retirement, D-131 retirement, concern-(iii) mitigation mechanism) -- +marked so below, not asserted as settled. + +--- + +## 1. Per-harness disposition table + +Only harnesses with a container-layer/two-root/qemu+ssh/`vvr1`/bootstrap-gate/depth-4 hit are +listed (13 of 103; the other 90 were grep-confirmed to have zero hits on the search terms above +and are not re-litigated here). "STAY" = the case's assertion target and value are unaffected +by Option 1 flattening. + +| Harness:case | Asserts about the container layer | Verdict | New invariant (if CHANGE) | +|---|---|---|---| +| `opentofu-validate:T13` (`:100`) | `opentofu/main.tf` autostart pin `# D-127: DC containment VM` for `vvr1-dc0`/`vvr1-dc1` (D-127 boot-matrix regression guard) | **RETIRE** | No successor pin -- the object is deleted, not renamed. Removing this case is itself the correct edit (not a silent drop: the D-127 boot-matrix comment block above it must say so) | +| `opentofu-validate:T14/T15` (`:101-102`) | `opentofu/vr1-dc0-substrate/main.tf` autostart pins: DC edge `=true`, node VMs `=false` | **CHANGE** | Same invariant (edge boots, node VMs stay MAAS-owned/manual), re-pointed at the flat root's file -- path depends on the OPEN root-naming ruling (pass2 Section 6.2: `-flat` vs reserving `-substrate`); do not pre-guess the filename. **Also ADD** a new case for the `vr1-dcN-client` VM's autostart value -- D-127's table predates the client VM and has no ruled value for it yet (OWED, not inferred) | +| `opentofu-validate:T8-T10` (S3 root-only blind spot) | Uses `modules/node-vm` + synthetic `modules/broken`/`modules/fine` fixtures to prove root-only `tofu validate` misses uncalled modules | **STAY** | Module-body level; unaffected (pass2: "zero module bodies need rewriting" except `wan-bridge`) | +| `opentofu-validate` header comment (`:92`) | "the 416 GiB nested container" | doc-currency nit, not a graded case | rides the T13 retirement edit | +| `node-vm:T8-T11` (`:47-98`) | Counts 12 `macs = [` lists / 72 pinned MAC literals in `opentofu/vr1-dc0-substrate/main.tf` (the INNER root, hardcoded path) | **CHANGE** | Re-point `INNER=` to the flat root's file (same open root-naming dependency as T14/T15 above). **Count invariant likely UNCHANGED** (12 nodes x 6 planes persists -- the client VM is a separate `cloudinit-vm` instance, not a `node-vm` instance, so it does not join this count per pass2 3.2's "client VM does NOT join `CARVE_AUX_HOSTS`") -- confirm at build, do not assume | +| `node-vm:T1-T7,T12-T15` | Module-body assertions on `opentofu/modules/node-vm` (MAC-pin variable shape, power-ownership `ignore_changes`) | **STAY** | Module body unchanged | +| `site-headend-install:` Section 8 (`:110-159`, the `--host-nodes` block) | `node_host_setup()`/`node_host_check()`: nested KVM, inner libvirt pool, D-125 WAN-bridge verify (`master $WAN_BRIDGE`, `br_netfilter` warning, uplink enslavement), the `vr1-dc0-substrate` hint-neutrality check, SEC-010 forward-drop keyed to `--transit-if` | **RETIRE (block), CHANGE (SEC-010 sub-case)** | `node_host_setup()`/`node_host_check()` and all D-125 WAN-bridge assertions RETIRE wholesale (pass2: dead code, delete; D-125 bridge-in has no successor -- edge WAN goes direct-NAT). The SEC-010 forward-drop assertions (`SEC-010`, `FORWARD-drop`, transit-if-exists fail-open check, nft declare-then-delete idempotency) **CHANGE**: re-target the extracted **role-agnostic installer subcommand** (pass2 Section 2.2/4.4 -- one artifact installing BOTH the client-VM end and the voffice1 end), asserting it runs on EITHER role, not gated behind `--host-nodes` | +| `site-headend-install:` Section 6 (`:80-102`, `--role rack`) | `maas init rack`, region+rack step suppression, enrollment-secret non-leak, for the standalone rack-controller role | **CONTINGENT CHANGE/RETIRE** -- blocked on Phase-4 ratifying pass2's rack-retirement recommendation (4.2(i), currently RECOMMENDATION-grade, not ruled; needs the "owed live re-measure" first). If ratified: `--role rack` retires (new builds `region+rack` directly on the region VM); this section's cases retire and a new case asserts `--role rack` is REFUSED/removed. If rejected: section stays as-is | New invariant depends on the ruling -- do not pre-pick | +| `site-headend-install:` Sections 1-5,7 (arg contract, `--compose-cidr`, dry-run mutate-nothing, region+rack default flow) | Nothing container-layer-specific | **STAY** | -- | +| `dc-selector:` power-address rows (`:204-247`) | Hard-pins `VIRSH_POWER_ADDRESS`/`_FROM_OFFICE1`/`_FROM_DCREGION` to `qemu+ssh://...@172.31.0.{2,6}/system` and `.../@10.12.{8,68}.2/system` -- i.e. the literal values that dial the CONTAINMENT VM's libvirtd via the transit/metal-admin legs | **CHANGE -- HIGHEST RISK, currently BLOCKED** | New values dial vcloud's OWN libvirtd instead; pass2 Section 2.3/3.2 has this **explicitly blocked** on the still-undesigned concern-(iii) power-key mitigation (owed artifact #11) -- URI shape, whether the FROM_OFFICE1/FROM_DCREGION split survives at all (both DCs may converge on one vcloud endpoint), are OPEN. Do not write a "corrected" literal into this harness before that mitigation is ruled -- that is exactly how a false-green gets minted (a guessed value that happens to match nothing live) | +| `maas-node-power:` URI fixture (`:60`) | `URI="qemu+ssh://jessea123@172.31.0.2/system"` exercised purely as an opaque pass-through arg (`maas-node-power.sh` takes the address as `$1`, no code change per pass2) | **STAY (functionally); doc-currency CHANGE optional** | Script logic is unaffected -- the fixture value doesn't need to be a real address. Consider refreshing the example post-rebuild so it doesn't read as a still-live containment address, but this is cosmetic, not a graded risk | +| `maas-region-power-key:` URI assertions (`:67,82,102,108`) | Pins the derived power-key install target as `qemu+ssh://...@10.12.8.2/system` / `...@10.12.68.2/system` -- the metal-admin `.2` address that is TODAY the rack containment VM's identity | **CHANGE -- BLOCKED, same dependency as `dc-selector`** | `maas-region-power-key.sh`'s own body is unchanged (pass2 3.3) but the `-` -> address DERIVATION it exercises resolves to a containment-VM identity that ceases to exist. Blocked on concern-(iii)'s mitigation design exactly like `dc-selector` -- these two harnesses should be updated TOGETHER, from the same ruled URI shape, or they will silently diverge | +| `dc-rack-net:` LEGS cases (T3,T5,T13,T15,T17) | Pins `vr1-dc0-metal-admin=10.12.8.2/22`, `vr1-dc0-provider-public=10.12.4.2/22`, dc1 equivalents, plus the systemd unit ordering for `${SITE}-rack-legs.service` -- the containment VM's OWN bridge-leg addressing | **RETIRE** (contingent, but on the STRONGER-settled side) | Per pass2 3.3: "Legs half RETIRES with the containment layer -- no flat VM is a libvirt host with its own bridges; `LEGS`/`br_of()` has no home to move to." Not blocked on a mitigation design the way concern-(iii) is -- this is a structural consequence of Option 1 itself | +| `dc-rack-net:` DNS-forwarder cases (T4,T9,T16 + the `DNS_UPSTREAM` literal) | D-131 forwarder confinement (no-resolv, single upstream, bind-interfaces) | **CONTINGENT RETIRE** -- blocked on Phase-4 ratifying D-131 retire-with-evidence (pass2 4.2(ii), RECOMMENDATION-grade, dc0 evidence-graded FUNCTIONAL not transcript, dc1 asymmetric and currently load-bearing -- "owed live re-measure" first). Note independently: `DNS_UPSTREAM="10.10.0.20"` is already flagged STALE for dc0 in pass2 3.3, a pre-existing defect unrelated to container-elim | If retired: this whole harness (all 18 cases, both LEGS and DNS halves) retires with `dc-rack-net.sh` itself. If D-131 is instead kept for a DC: the DNS half survives re-targeted at wherever component (i)'s host lands | +| `dc-rack-net:T2,T6,T7,T8,T11,T12,T14,T18` | Generic script-hygiene cases (`bash -n`, no `virbrN` literal, MEASURED-tag discipline, unknown-site/-mode refuse, `do_check` read-only) | rides the parent verdict | If the script retires wholesale these retire with it; they carry no container-layer-specific content of their own | +| `dc-rack-mgmt-import:` D-124 scheme pins (`:77-78`) | `"rack_dns": "vvr1-dc0"` / `"vvr1-dc1"` literal role-name pins inside `netbox/dc-rack-mgmt-import.py`'s SITES map | **CONTINGENT CHANGE/RETIRE** | Owed artifact #9 (pass1/pass2: NetBox DCIM migration -- decommission the `vvr1-dcN` device records, register the client VM + flat roster). If the rack-controller retirement (4.2 i) is also ratified, the whole "rack DNS device" concept this script imports may retire, not just get renamed -- do not pre-pick between rename-in-place and full replacement; that is a Phase-4/NetBox-migration-design call | +| `dc-rack-mgmt-import:` D-124 CONTAINER/transit-scheme pins (`:73,79` + `d124-transit-seed:42,48`) | `CONTAINER = "172.31.0.0/24"`, `.1`-gateway rejection, `/30`-or-`/31` shape, `dcim.site` scope | **STAY** | The transit supernet and its NetBox scheme are mesh-triangle facts, not containment facts (pass2 4.5: mesh triangle + all 3 `mesh-link` legs unchanged; only the office1-leg CONSUMER re-points from the containment VM's NIC1 to the client VM's transit NIC -- an endpoint change these value-pins don't encode) | +| `maas-profile-assert:` office1-profile fixture (`:41,68,72,80`) | Simulated MAAS machine list `[voffice1, vvr1-dc0, vvr1-dc1]` for the "office1" profile (the rack controllers as MAAS-visible machines) | **CONTINGENT CHANGE** | Fixture data, not logic. If rack-controller retirement (4.2 i) is ratified, `vvr1-dc0`/`vvr1-dc1` stop existing as MAAS machines under that profile -- update the fixture to the post-retirement roster. The client VM does NOT replace them here: per pass2 3.2 it is not MAAS-carved, so it does not appear in this profile's machine list either | +| `preflight:` pending-change fixture (`:89`) | Synthetic `tofu plan` line `module.vvr1_dc0.libvirt_domain.vm will be updated in-place`, used only to exercise the PENDING-change-detection regex | **STAY** | Fixture-only; the detection logic is module-name-agnostic. Cosmetic rename optional, non-load-bearing | +| `geneve-encap-assert:` (all cases) | Family-split (C1) / bracketed-encap-ofport (C2) OVN checks, driven entirely from input files | **STAY** | Confirmed family/topology-agnostic -- no `vvr1`/containment coupling found. This is the harness that will gate the STILL-OWED live geneve/jumbo assert on the vcloud-level planes post-build (pass0 Section 8 item 3) -- that live assert is a NEW USE of this same offline-tested tool, not a harness edit | +| `site-baseleg:` active-leg-row guard (`:69-73`) | Asserts NO un-commented DC supernet (`10.12.`/`172.31.`) appears in an active `LEGS` row -- i.e. structurally proves the script is STILL a no-op for DC legs | **STAY** | Already forward-compatible: this guard's job is to keep the script a no-op until a DC leg is explicitly measured and added, and Option 1 doesn't change that precondition (pass2 4.5: re-cite D-138 + the (a) control in the comment instead of the retired qemu+ssh premise -- doc-currency only, not a case change) | +| `cloudinit-vm:` (all cases, `T1-T6,T5a-T5d`) | D-130 lifecycle guard + MAC-pin shape on `opentofu/modules/cloudinit-vm` -- the module type the client VM will be a NEW INSTANCE of (pass2 3.1) | **STAY** | Module-body level, unaffected. No new W3.1 case needed for the client-VM instantiation itself (that is an apply-time/W3.2-W3.3 concern -- does the client VM need its own MAC pins once carved, etc. -- not a change to THIS harness's existing assertions) | +| `netem-link:` (all cases) | Header comment references "the outer root runs ON vcloud" (D-128 local-mode amendment) | **STAY** | Confirmed false-positive on the survey's "inner/outer root" grep -- this is the OUTER vcloud root's own local-vs-SSH execution mode, unrelated to the inner/outer CONTAINMENT split. Zero `vvr1`/qemu+ssh hits. Mesh triangle unchanged | + +**Harnesses grep-confirmed with NO container-layer coupling** (checked because their scripts were +named in pass2 as "no code change" and could plausibly hardcode a containment value, but do not): +`maas-role-tags`, `carve-host-interfaces`, `dc-node-carve`, `dc-egress-check`, `dc-plane-ipam`, +`dc-region-topology`, `provider-bundle-check`, `lib-validate` (the `lib-hosts.sh`/`lib-net.sh`/ +`lib-identity.sh` unit-test harness -- exercises `emit`/`vr_json`/env-scrub plumbing generically, +not the containment-keyed values themselves; those live in `dc-selector`, already covered above). + +--- + +## 2. Count and gauntlet-edit list + +**13 of 103 harnesses affected** (12.6%): 2 clean RETIRE, 2 contingent RETIRE (pending Phase-4 +rulings), 6 CHANGE (2 of them currently BLOCKED on an undesigned mitigation), 1 contingent CHANGE +(fixture only), plus module-body/fixture-only harnesses noted STAY for completeness +(`node-vm`'s module cases, `opentofu-validate`'s S3 cases, `cloudinit-vm`, `netem-link`, +`site-baseleg`, `geneve-encap-assert`, `preflight`, `d124-transit-seed`). + +Gauntlet cases needing edits, with the FAILING-DIRECTION fixture each needs (per repo discipline: +an assertion must be provably able to fail; a fix that makes a finding-string assertion stale +gets REPLACED with the new invariant, never deleted to go green): + +1. **`opentofu-validate` T13** -- delete the case; add a fixture that the D-127 boot-matrix + pin-table comment block is updated (`grep`-checkable: the comment must no longer describe a + containment-VM autostart row that doesn't exist). Failing direction: a stray re-add of a + `vvr1-dc0` autostart pin in `opentofu/main.tf` should have NO test catching it post-retire -- + flag this as an accepted residual gap, or add a negative-assertion case ("no `vvr1` domain + block exists in `opentofu/main.tf`") so a regression is still caught. +2. **`opentofu-validate` T14/T15 + `node-vm` T8-T11** -- re-point the hardcoded + `opentofu/vr1-dc0-substrate/main.tf` path to the ratified flat-root file once Phase 4 rules + root naming; keep the exact-count/exact-value assertions (12/72, edge=true/node=false) as the + failing-direction fixture (a wrong count or a flipped autostart bool must still redden). +3. **`site-headend-install`** -- the biggest single edit. Split into: (a) delete + `node_host_setup()`/`node_host_check()` assertions outright (RETIRE, ~15 grep cases); (b) + write NEW cases against the extracted role-agnostic SEC-010 installer subcommand, with a + fixture that installs it on a fake "client" role AND a fake "voffice1" role and asserts BOTH + ends get the drop rule, using the same fail-open discipline already proven here (verify the + named interface actually exists, not just that a rule loaded); (c) hold Section 6 (`--role + rack`) unedited until the rack-retirement ruling lands -- editing it now would be guessing + Phase 4's answer. +4. **`dc-selector` + `maas-region-power-key`** -- do NOT edit these yet. They are correctly RED + the moment `lib-hosts.sh`'s power-address derivation changes, which is exactly the + fail-loud behavior wanted while concern-(iii) is unmitigated. Edit them ONLY together with + the mitigation's ship (owed artifact #11), from the ruled URI/key shape -- editing either one + first, or guessing a value, is the false-green trap this repo's rule about inferred values + exists to prevent. +5. **`dc-rack-net`** -- do not edit pending the D-131/rack-retirement rulings; if both ratify, + RETIRE the harness file wholesale alongside `dc-rack-net.sh` (append-only bias: leave the + file in git history, remove from `tests/HARNESS-MANIFEST` via `--record-manifest`, log the + removal in a changelog with a revert per repo discipline). +6. **`dc-rack-mgmt-import`** -- hold for the NetBox-migration design (owed #9); do not rename the + `vvr1-dc0`/`vvr1-dc1` string pins speculatively. +7. **`maas-profile-assert`** -- fixture-only edit, lowest risk of the CHANGE set; safe to update + once rack retirement is ratified (drop the two `vvr1-dcN` rows from the office1-profile + fixture), independent of the other blocked items. + +--- + +## 3. Highest-risk invert-under-flattening cases (the ones most likely to false-green or +false-red if edited carelessly) + +1. **`dc-selector`'s power-address pins and `maas-region-power-key`'s derived-URI pins** -- + these are the sharpest risk in the whole survey. Both assert LITERAL containment-VM + addresses that MUST change, but the replacement values are explicitly BLOCKED on an + undesigned mitigation (concern iii, pass2 Section 2.3, owed artifact #11). The failure mode + to guard against is a well-meaning edit that swaps in a plausible-looking new URI (e.g. "just + point it at vcloud's own address") before the restricted-key/wrapper/ACL mechanism is + actually built -- that produces a harness that is GREEN against a value nothing enforces, + which is worse than the current honest RED. +2. **`site-headend-install`'s `--host-nodes` block** -- inverse risk: because `--host-nodes` + itself might simply be REMOVED as a flag (not just have its body gutted), a case like + `t "--host-nodes without --role rack -> 2"` could keep passing for the WRONG reason (an + unrecognized-flag exit code that happens to also be 2), silently changing what the assertion + proves without ever going red. Any edit here must re-derive the expected exit code from the + NEW arg-parse contract, not assume rc=2 still means what it meant before. +3. **`opentofu-validate` T13/T14/T15 and `node-vm` T8-T11** -- risk is a stale hardcoded PATH + silently reading as "file not found" -> FAIL, which is safe (loud, not silent), but a + sloppy fix that just deletes the whole case to make the gauntlet green again (rather than + re-pointing it) would violate the repo's "replace, never delete to go green" rule and quietly + drop the D-127 autostart-drift regression coverage this class of harness exists for (it has + gone red twice before on a missed re-run, per `node-vm`'s own commit-history commentary at + `tests/node-vm/run-tests.sh:68-81`). +4. **`dc-dc-whole-host-budget`** (flagged in Section 1's supplementary note, not the main table + since it asserts Model A/B math rather than a container-layer FACT directly) -- worth naming + here because it is the harness most likely to stay GREEN while testing something no longer + true: its Model-A/B RAM comparison (838 vs 822 GiB, "containment overhead = 2x16 GiB") will + keep passing indefinitely against the OLD script even after the topology it models is gone, + because nothing forces a re-run against a NEW `--model` flag until the FIT-calculator + extension (owed #7) actually ships. This is squarely W3.2/W3.3 territory (new-model-flag + test requirements) but is flagged here as the standing risk that motivates it. + +--- + +## 4. Durable-doc path + +`docs/audit/container-elim-pass/pass3-w1-harnesses-gauntlet.md` (this file). diff --git a/docs/audit/container-elim-pass/pass3-w2-gates.md b/docs/audit/container-elim-pass/pass3-w2-gates.md new file mode 100644 index 0000000..bdf843e --- /dev/null +++ b/docs/audit/container-elim-pass/pass3-w2-gates.md @@ -0,0 +1,187 @@ +# Pass 3 / W3.2 -- preflight / cloud-assert / stage-gate change table (flat topology) + +**Worker:** W3.2 (Phase 3 -- Tests review), container-layer-elimination pass. +**Date:** 2026-08-09. **READ-ONLY.** No mutation; findings LOGGED, per SCOPE Section 7/5.5. +**Inputs consumed in full:** `SCOPE-AND-EXECUTION-PLAN.md`, `pass0-admin-report.md`, +`pass1-admin-report.md`, `pass2-admin-report.md`. Baseline facts carried without +re-deriving: **Option 1 CONFIRMED** (flat node VMs on vcloud libvirt + one small +non-hypervisor `vr1-dcN-client` VM per DC, `.8`); cross-DC handling **(a) CONFIRMED** +(new vcloud-level host isolation control); MAAS region stays on `vr1-dcN-maas-01`; +**THREE isolation concerns**, each real, each open, none substituting for another +(pass2 Section 2): (i) cross-DC plane-bridge network adjacency -- the (a) control; +(ii) the SEC-010 transit-leg FORWARD-drop successor (client VM + voffice1); (iii) the +MAAS power-key blast radius (NEW, pass2 Section 2.3) -- owed artifact #11. Rack +controller: **RETIRE the standalone registration** (pass2 Section 4.2(i)), D-131 +**retire-with-evidence** (pass2 4.2(ii)). All riding D-143 (10.13 re-IP). + +Repo discipline applied throughout: assert on CONTENT not existence; an unrecognised +result REFUSES, never passes silently; a checker must be provably able to FAIL +(GA-R6 "checker that cannot fail is not a gate," `SKILL.md:321-333`); every gate cites +`docs/tool-index.md` before an operation is named in a runbook. + +--- + +## 1. `scripts/preflight.sh` -- per-gate change table + +Current gates read directly (`scripts/preflight.sh:1-386`): P1 repo-lint, P2 bundle +invariants (`provider-bundle-check.py`, DC-aware since F5), P3 channel assert, P4 live +pre-flight (MAAS/overlay/nodes), P5 credential matrix (D-137, hermetic tier1+tier2-local), +P7 octavia PKI (HEADEND-ONLY, `creds-manifests/host-identity` `headend` binding), P8 +substrate drift (outer tofu root, vcloud-only), P9 DC egress (rack-only), P6 stage-2 +reminders (not run here). + +| Gate | Current container-layer assumption | Option-1 change | New/changed assertion (failable?) | +|---|---|---|---| +| **P1 repo-lint** | None (static hygiene) | None | Unchanged. L11 (`*.original.md` residue) etc. stay as-is | +| **P2 provider-bundle-check.py** | DC-aware (`$DC` selects overlay set); no host-topology assumption in the bundle content itself | None to the bundle logic. Overlay literals (VIPs, machine placement) carry D-143's 10.13 addresses -- **address axis only**, not container-elim | Unchanged mechanism. `overlays/${DC}-machines.yaml` values change under D-143, tracked there not here | +| **P3 channel_assert.py** | None | None | Unchanged | +| **P4 pre-flight-checks.sh** | **YES -- the flagged assumption.** "live pre-flight (MAAS/overlay/nodes)" reads MAAS machine state, overlay/VIP data, and node readiness -- all of it currently reasoned about through the two-root/two-host Model-B shape implicitly (the runbooks it backs assume a `vvr1-dcN` rack host exists to enroll/carve against, per pass0 rows 4-7). Needs direct read of `pre-flight-checks.sh`'s body (not yet done this pass) to enumerate literal host/topology assumptions -- **flagged to W3.1's harness sweep**, since W3.2's charter is the gate SHAPE, not this script's full body | Node discovery/carve mechanism is **unchanged** (per-machine `power_type=virsh`, pass1 check 7) -- only the power-ADDRESS value re-derives, and that value is **BLOCKED on the concern-(iii) mitigation** (pass2 Section 2.3). P4 must not assert a specific power URI until #11 is designed | **Content assertion needed, new:** once #11 lands, P4 (or a new P-gate) should assert the LIVE power-address value in use matches the mitigation's issued form (e.g. `command=`-restricted key / wrapper), not the raw `qemu+ssh://` shape -- a regression-detector for the exact defect concern (iii) names. Failable: a stale containment-shaped URI in `lib-hosts.sh` or a live MAAS power-parameters read fails it | +| **P5 creds-matrix.py --tier2** | **YES.** D-137's register rows are keyed to host CLASSES including `rack` (SEC-028's "first `rack` rows," pass0 row 8) -- Option 1 retires the rack-controller registration (pass2 4.2(i)) and re-points those rows to `vr1-dcN-client`. Matrix content is data (rows), not code, so **no script change** -- but the row SET must be updated or P5 asserts against retired host classes and either false-FAILs (row references a host that no longer exists) or false-PASSes (a client-VM credential with no row at all, SEC-022's exact failure class) | Register rows: `rack` class rows -> `client` class rows (SEC-028/-029 residencies); **NEW rows owed for concern (ii) and (iii)'s minted artifacts** (the transit-drop installer's key material if any, the power-key mitigation's restricted key) | Already failable by construction (D-137 ruling 1: hard-fails on expected-but-absent/undeclared/asymmetric). **Change is DATA not CODE**: the matrix's row source must be updated when #11/#3/#5's artifacts are built, or P5 silently stops covering the new credential-bearing host (a coverage gap, not a false pass -- but coverage gaps are how SEC-022 happened, pass2 check 6, and D-137's founding incident). Not this pass's job to edit the matrix; flagged as a delivery dependency on owed artifacts #3/#5/#11 | +| **P7 octavia-pki.sh verify** | Host-bound to the **headend** via `creds-manifests/host-identity`'s `headend` row (D-109 note (b)) -- NOT container-layer-keyed at all; the PKI lives where D-138/D-128 put the deploy execution host, already re-pointed to the DC client VM by D-138 (pass0 check 3). **No change from container-elim**: P7 already reads "whichever host `host-identity` names," and that binding tracks D-138 independent of this pass | None -- confirms it is ALREADY flat-topology-correct by construction, a positive finding | Unchanged. Worth noting in the change-set as a NON-finding so it isn't re-litigated: P7's host-binding indirection is the pattern P4/P5 should imitate once concern (iii) picks a mechanism | +| **P8 substrate drift** | **YES, explicitly two-root shaped by naming.** Comment says "outer tofu root" and only evaluates `opentofu/terraform.tfstate` + `opentofu/.terraform` on vcloud; guards for "not the substrate host" (voffice1, which under Model B holds the INNER root's state) as a legitimate non-evaluating state (`:284-296`). Under Option 1 there is **no inner root and no voffice1-side state** for the DC substrate -- the guard's premise (the inner root lives elsewhere) is now false, and the shared-outer + per-DC-flat root fork (pass2 recommend (B), Phase-4-ratified) means potentially THREE state files (shared outer + `vr1-dc0-flat` + `vr1-dc1-flat`), not one | Must evaluate **every root that exists post-flatten**: the shared-outer root PLUS each per-DC-flat root (if (B) is ratified), each independently, each printing which root it evaluated. The voffice1 "not the substrate host, WARN not FAIL" branch becomes dead code for the DC-substrate half (voffice1 keeps only Plane-2/MAAS-NetBox duties per the D-128 amendment, pass1 check 6) | **Change needed, failable already in shape:** extend P8 to loop over a root LIST (`opentofu/` + `opentofu/vr1-dc0-flat/` + `opentofu/vr1-dc1-flat/`, names TBD at Phase-4 ratification) instead of the single hardcoded `opentofu/` path at `:284,298`. Each root keeps the existing pending-action-FAILs / refresh-only-WARNs / no-state-here-WARNs logic (content-based, already correct) -- only the root ENUMERATION is container-elim-shaped and needs the update. Root NAMING is Phase-4's call (pass2 Section 6.2); this gate cannot be finished until that's ratified | +| **P9 dc-egress-check.sh** | Runs **ON THE RACK** ("every value it uses describes the rack," `:335-336`); the rack IS `vvr1-dcN` today | The egress-testing HOST changes: under Option 1 there is no rack-as-libvirt-host to run this from. The natural new host is **`vr1-dcN-client`** (the DC-side VM with a transit+metal-admin leg, structurally the only candidate per pass2 Section 2.2's SEC-010-successor reasoning) OR a node itself. Needs a Phase-4 pick, same open item as B.5 placement (pass1 Section 7 item 2) | Mechanism (egress reachability from inside the DC) is unaffected; only the invocation-host literal in the warn/fail messaging (`:353-354`, `ssh <$DC rack>`) needs updating to name the new host once ruled. Already failable/content-based; no new assertion needed, a literal-currency edit only | +| **P6 stage-2 reminders** | None (printed only) | None | Unchanged | + +### 1.1 New P-gates this pass identifies as owed (not yet slotted with a number -- Phase-4/delivery mints) + +Three isolation controls' `--check` gates do **not** slot into the existing P1-P9 set +cleanly, because preflight's charter is **pre-deploy, per-DC** (`DC=` selects one site, +`:36-66`) while concern (i) and concern (iii) are **cross-DC, host-scoped, single vcloud +kernel** facts that exist independent of which DC is being gated. Recommendation: + +- **Concern (i) [the (a) control]:** belongs as a **new stage gate at Stage 1** (pass1 + Section 3's own recommendation, re-affirmed here), NOT a preflight P-gate -- + preflight runs `DC=`-scoped and re-running a cross-DC assertion once per DC either + duplicates work or silently only checks the last-invoked DC. A `--check` subcommand + of the new artifact (SEC-010's pattern) is the right shape; preflight's P4/P9-style + REFUSE-on-wrong-host guard applies (it must run on vcloud, the only host that can see + both DCs' bridges). +- **Concern (ii) [SEC-010 successor]:** **DOES fit as a preflight P-gate**, DC-scoped, + because the transit-leg drop is inherently per-DC (one client VM + voffice1, one pair + per DC) -- model it exactly on P7's shape (host-bound `--check`, REFUSE off-host, + content assertion on the nftables table + interface existing, matching SEC-010's own + `node_host_check()` pattern at `scripts/site-headend-install.sh:206-231`). Candidate + slot: **P10** (next free preflight letter; P6 is reserved as the non-executing + reminder block). +- **Concern (iii) [power-key blast radius]:** belongs as **P4's dependency**, not a + freestanding P-gate on its own merits (it gates whether P4's power-address literal is + trustworthy) -- but it is ALSO a standing, non-preflight assertion (the mitigation + must hold at all times libvirtd is up, not just at deploy time), so it needs BOTH a + preflight-time check (does the live power-address match the mitigation's issued + shape) and a `cloud-assert.sh`-time check (Section 2 below) that the mitigation is + still enforced. + +None of these can be given a fully concrete `--check` body in this pass -- their +underlying artifacts (#2, #3, #11) are UNDESIGNED (pass1/pass2 explicitly leave +mechanism open); this table gives each a **gate home + assertion SPEC**, not an +implementation (Section 3). + +--- + +## 2. `scripts/cloud-assert.sh` -- container-layer / two-host assumptions + +Direct read (`scripts/cloud-assert.sh:1-294`): sections A0-A10 are ALL post-deploy +juju-model/OpenStack-service behavioral checks (vault seal state, mysql cluster, +OVN uniformity/chassis, compute plane, octavia LBs, keystone/magnum, conductor graft, +vault-kv AppRole, HA arity). **Zero container-layer or two-host assumptions found** -- +every section either runs `juju ssh` into a unit or calls the OpenStack API; none reads +`vvr1-dcN`, dials qemu+ssh, or otherwise depends on the containment shape. This is +consistent with pass1/pass2's finding that D-140 keeps L4 (juju/openstack deploy) a +PROCEDURE layer sitting entirely above the substrate -- cloud-assert is an L4/L5 +artifact and the substrate reshape underneath it is, by design, invisible to it. + +**What a flat-topology cloud-assert must ADD (net-new, not a modification of A0-A10):** + +A flat topology's actual NEW risk is that isolation the containment layer provided +"for free" (pass0 Section 5: separate kernels per DC) must now be asserted explicitly. +None of A0-A10 tests network/credential isolation between DCs -- they test the OpenStack +control plane's OWN health, which is DC-scoped by the model they run against (`-m +"$MODEL"`, one DC's juju model per invocation). **Recommend a new section, `A11: cross-DC +isolation still enforced`**, appended to cloud-assert as the runtime-verification +half of concern (i) and (iii) (concern (ii)'s transit-drop is a preflight/deploy-time +concern, not an ongoing service-health one, though it could be echoed here too): + +- **A11a (concern i, network):** re-run the (a) control's `--check` from vcloud itself + (cloud-assert already assumes jumphost/headend execution context for some sections, + same class of host-binding as P7) -- content assertion: the nftables ruleset that + blocks inter-plane/inter-DC forwarding is LOADED and its rule COUNT/hash matches the + artifact's own expected state (not just "a table named X exists" -- SEC-010's own + fail-open lesson: keying to an absent interface loads clean but matches nothing, + `site-headend-install.sh:224`). This is the **B.7 re-run** pass1 already specifies + ("the (a) control's `--check` re-run now that both DCs' planes are actually + co-resident," pass1 Section 4 B.7) -- cloud-assert is the natural PERIODIC home for + that re-run (post-deploy, post-restart, pre-change baseline, post-incident -- exactly + cloud-assert's stated invocation points, `:6-9`), not a one-time deploy-gate. +- **A11b (concern iii, credential-scope):** assert the power-key mitigation is still in + effect -- e.g. if the mechanism is a `command=`-restricted key, grep the live + `authorized_keys` forced-command on vcloud for the expected wrapper/allowlist rather + than a bare key; if it is a polkit ACL, assert the ACL file's content matches the + per-DC scoping. **Content-based, failable**: an unrestricted key or a missing ACL + fails it. This directly operationalizes pass2's own warning that "a read of vcloud's + LIVE polkit/libvirt config is delivery work, not asserted here" (pass2 Section 2.3) -- + A11b IS that assertion, turned into a standing gate rather than a one-off read. + +**Framing note:** cloud-assert's own doc-comment (`:4-9`) explains it exists because +"juju status is BLIND to" certain classes of defect learned from incident history +(D-045/046/051/042). Concerns (i) and (iii) are the SAME shape of blind spot one layer +down the stack -- `juju status` and the OpenStack API are equally blind to a host-level +libvirt/nftables isolation failure. A11 is the structurally consistent place to add +them, not a bolt-on. + +--- + +## 3. The three isolation controls + rack-retirement evidence -- gate homes + failable assertion specs + +| # | Control | Gate home (recommendation) | Assertion spec (content-based, failable) | +|---|---|---|---| +| **(i)** | Cross-DC vcloud host-level isolation control (the "(a)" control) | **NEW Stage-1 gate** (pass1 Section 3's recommendation, confirmed here) -- `scripts/-check.sh` (name TBD, Phase-4/delivery), a `--check` subcommand run on vcloud; ALSO re-run at (per pass1 Section 3) each per-DC substrate apply's close AND at cloud-assert's `A11a` (Section 2 above) for ongoing verification | `nft list table inet ` must show a FORWARD-drop scoping EVERY pair of DC plane bridges (not just one -- the fail-open class SEC-010 already taught this repo: keying to one interface and silently matching nothing is not a control). Concretely: enumerate the live bridge set for both DCs' six planes (from `lib-hosts.sh`'s `NIC_PLANE_ORDER`/`BREX_PARENT_NIC` conventions, which persist per pass0 row 14), assert the ruleset denies forwarding between any bridge tagged `dc0` and any tagged `dc1`, and REFUSE (not pass) if fewer than the full plane count resolves to a live interface -- the SEC-010 fail-open lesson generalized. Ordering invariant (pass1 Section 3): must be installed + `--check`-verified **before the first flat substrate apply of EITHER root** | +| **(ii)** | SEC-010 transit-leg FORWARD-drop successor | **NEW preflight gate, P10** (Section 1.1), DC-scoped, host-bound to the client VM's own `--check` (mirroring `node_host_check()`, `site-headend-install.sh:205-231`) PLUS the voffice1-side install verified the same way. **One extracted role-agnostic installer/checker for BOTH ends** (pass2 Section 2.2's resolved spec) | `nft list table inet sec010` (or its successor table name) present on the client VM, keyed to the CLIENT VM'S OWN transit interface (re-measured, not assumed -- pass2 carries forward the "NIC-naming trap" lesson: dc0's live interface was `enp1s0` not the script default `mgmt`); mirrored check on voffice1's DC-facing transit leg. REFUSE if the keyed interface does not exist (exact SEC-010 fail-open precedent) | +| **(iii)** | MAAS power-key blast-radius mitigation | **Split gate, two homes:** (a) preflight P4 dependency -- before P4 asserts a power-address literal, verify it matches the mitigation's issued shape (Section 1's P4 row); (b) `cloud-assert.sh` **A11b** (Section 2) for standing/periodic re-verification, since the exposure exists any time vcloud's libvirtd is up, not only at deploy time | **What it must assert (pass2 Section 2.3's "consequence" made concrete):** that connecting with EITHER DC's region-VM power key to vcloud's `qemu:///system` endpoint CANNOT enumerate/control domains outside that DC's own set. Candidate concrete check (mechanism-dependent, Phase-4 picks the mechanism per pass2 Section 6 item 5): if a `command=`-restricted key, assert the `authorized_keys` forced-command wraps every DC-scoped invocation and REJECTS a raw `virsh -c qemu:///system list --all` from that key; if a polkit ACL, assert the ACL rule's domain-name-prefix match covers exactly that DC's `vr1-dcN-*` roster and denies the other DC's + voffice1's + the vcloud substrate's own domains. **A pass verdict must be tied to a NEGATIVE test** (the key CANNOT reach the wrong domains), not merely "the key exists" -- an existence-only check is exactly the class of non-gate GA-R6 rules out (`SKILL.md:321-333`) and exactly what SEC-012/-016's own ledger rows already flag as unaddressed ("blast radius... broader than the power verbs MAAS actually needs," `security-ledger.md` SEC-012) | +| **rack-retirement evidence** | D-131 retire-with-evidence (pass2 4.2(ii)) + MAAS rack-controller decommission (pass2 owed artifact #5, amended) | **NEW step in the promoted MAAS machine-record release/delete Part-A step** (pass1 owed artifact #5) -- not a standing gate, a ONE-TIME evidence capture at build time, per pass2's framing ("the retirement EVIDENCE step... is owed artifact #13") | The dc0 migration's own proof shape is the template (`docs/changelog-20260730-dc0-region-migration.md` Item 9, cited pass2 check 4): `dig` against the fresh region's own BIND from a node, asserting `dns_servers` resolves via the DC-LOCAL region (not a remote forwarder), ANSWER section non-empty, `flags: qr rd ra`. Run once per fresh 10.13 region at build time as evidence the D-131 forwarder is genuinely unneeded, captured into the build's changelog -- NOT wired as a recurring gate (the asymmetry pass2 found, dc0-load-bearing-false / dc1-load-bearing-true, was itself only caught by exactly this kind of direct dig-test, check 4) | + +--- + +## 4. Top risks / gaps this dimension surfaces + +1. **P4 and P5 are BLOCKED on undesigned artifacts.** Neither `pre-flight-checks.sh`'s + power-address assertions nor the creds-matrix row set can be finalized until + concern-(iii)'s mitigation mechanism is chosen (pass2 Section 6 item 5) -- this + gate work has a hard external dependency, already flagged upstream (pass2 Section + 2.3 "BLOCKED on this mitigation's design"). +2. **P8 (substrate drift) is currently single-root-hardcoded** (`opentofu/` literal at + `:284,298`) and will silently under-evaluate if the root-topology fork resolves to + (B) shared-outer + per-DC-flat (pass2's recommendation) without this loop-extension + landing first -- a real regression risk if the redeploy ships before this gate is + updated. +3. **`pre-flight-checks.sh`'s own body was not read this pass** (W3.2's charter names + it as the P4 delegate but the full literal-by-literal container-layer sweep of ITS + internals belongs to W3.1's harness-assumption charter, per the phase-prompt split) + -- flagged so Phase-3's administrator does not read this table as a complete P4 + audit. +4. **No `--check` body exists yet for any of the three isolation controls** -- this + table specifies WHAT each must assert (content, failable, REFUSE-on-ambiguous), not + HOW; per SCOPE Section 7 the pass plans, delivery builds, each with its own + `tests//run-tests.sh` harness (repo discipline, not yet started -- W3.3's + charter). +5. **A11's placement in cloud-assert is a recommendation, not a ratified decision** -- + Phase 4 should confirm cloud-assert (periodic/behavioral) vs. a dedicated new + `isolation-assert.sh` (single-purpose, callable independent of the full A0-A10 + sweep) is the right home; the case for folding in is cloud-assert's own stated + charter (catching what status/API checks are blind to) and its existing periodic + invocation points, not a structural necessity. + +--- + +## 5. Verification note + +Direct reads this session: `scripts/preflight.sh` (full, 386 lines), `scripts/ +cloud-assert.sh` (full, 294 lines), `scripts/site-headend-install.sh` (SEC-010 +`node_host_check()`/writer, lines ~205-320), `scripts/dc-egress-check.sh` (host-binding ++ REFUSE shape, lines ~22-60,335-354), `docs/security-ledger.md` (SEC-010/-012/-016 rows +in full, plus a sweep for the highest-numbered row = SEC-033, next-free SEC-034), +`docs/design-decisions.md` (highest-numbered decision = D-143, next-free D-144), +`docs/CURRENT-STATE.md:7786-7814` (G1-G18 in full), `.claude/skills/openstack-cloud-ops/ +SKILL.md:310-369` (GA-R6 gate discipline, deploy-loop pointers). No inferred values used; +every host/interface/mechanism cited as OWED or UNKNOWN where the source pass docs left +it open (concern-(iii) mechanism, root naming, P10's exact number, A11's exact name) is +marked as such rather than guessed. READ-ONLY; nothing executed; findings LOGGED only. diff --git a/docs/audit/container-elim-pass/pass3-w3-new-tests.md b/docs/audit/container-elim-pass/pass3-w3-new-tests.md new file mode 100644 index 0000000..b5a5b31 --- /dev/null +++ b/docs/audit/container-elim-pass/pass3-w3-new-tests.md @@ -0,0 +1,259 @@ +# Pass 3 -- W3.3: new per-module test-harness requirements (container-layer elimination) + +**Worker:** W3.3 (Phase 3, container-layer-elimination pass, `SCOPE-AND-EXECUTION-PLAN.md` +Section 4). **Date:** 2026-08-09. **Scope:** READ-ONLY -- specifies what harness each OWED +new/changed artifact must ship; writes and runs nothing. Baseline consumed: `pass0-admin-report.md` +Section 7a (Option 1 CONFIRMED; cross-DC handling (a) CONFIRMED), `pass1-admin-report.md` +Sections 3/5/6 (the (a)-control spec, the layer model, the owed-artifacts seed list), +`pass1-w4-module-planning.md` (layer model + design principles), `pass2-admin-report.md` +Sections 2/4/5 (the three isolation concerns separated; the 13 owed artifacts) and +`pass2-w4-module-decomposition.md` (the procedure-module contract, Section 3; the IaC<->procedure +boundary, Section 4). Harness-pattern precedent read live this session: +`tests/geneve-encap-assert/run-tests.sh`, `tests/site-headend-install/run-tests.sh`, +`tests/phase-00-teardown-d061/run-tests.sh`, `tests/dc-egress-check/run-tests.sh`, +`tests/dc-node-v6-verify/run-tests.sh`, `tests/opentofu-validate/run-tests.sh`, +`scripts/lib-validate.sh`, `docs/security-ledger.md:21` (SEC-010), `docs/changelog-20260730- +octavia-reissue-tool.md:158-167` (the stage-assert-promote / join-workspace-to-deploy-input +vocabulary), `opentofu/modules/node-vm/variables.tf:38-62` (MAC-pinning validation, reused not +reinvented). No inferred values; every invariant below cites the source line that establishes it. + +--- + +## 0. Method + +This pass does not design the mechanisms (nftables rule sets, wrapper shapes, dig-test +fixtures) -- several are explicitly OPEN pending Phase-4 ratification (`pass2-admin-report.md` +Section 6). It specifies, for each of the 13 owed artifacts (`pass2-admin-report.md` Section 5), +**what its harness must assert and, critically, the fixture that proves each assertion can turn +red** -- so that whichever mechanism Phase 4 picks, the harness's job is already scoped and the +delivery session cannot ship a checker that cannot fail (this repo's own recorded failure mode, +`instrument-currency-before-negatives` memory #13/#14/#16, and the D-061 pair's "a host that does +not resolve is a `note`, not a `fail`" defect, `docs/tool-index.md`). Six artifacts get full specs +(Section 2, the pass's assigned dimension); the remaining seven get a shorter table (Section 3) +so all 13 are accounted for and none is silently left harness-less. Section 4 is the general +contract every one of these harnesses -- and by extension every future module harness -- follows. + +--- + +## 2. Full harness specs -- the six artifacts named in this worker's charter + +### 2.1 `modules/dc-site` (new composing IaC module) -- owed artifact #12 + +**What it composes** (`pass2-admin-report.md` Section 3.1): storage pool + six planes + edge + +12 node VMs (9 D-121 role nodes + `juju-01`/`maas-01`/`tailscale-01`, `pass0-admin-report.md` +Section 1.1) + the new `vr1-dcN-client` VM (`.8`, `pass2-admin-report.md` Section 4.3) -- 13 +L1/L2 compute objects + 6 plane networks per DC, replacing the ~230-266-line copy-pasted +per-DC inner-root bodies. + +| Invariant | Source of truth | Failing-direction fixture | +|---|---|---| +| **Plane count = 6** per DC | `pass0-admin-report.md` Section 1.1 ("the SIX planes"); D-134 | a fixture module tree with a plane call REMOVED (5 planes) -> count assertion FAILS | +| **Node roster = 12** (9 role + 3 utility), correctly classed | `pass0-admin-report.md` Section 1.1 exact roster | a fixture with a role node dropped (11) OR a utility node duplicated (13 with 2x `maas-01`) -> roster assertion FAILS | +| **Client-VM presence = exactly 1** per DC | Option-1 gate confirmation, `pass0-admin-report.md` Section 7a | a fixture module tree with the `client_vm` call commented out -> presence assertion FAILS; a fixture with 2 client-VM calls -> also FAILS (exactly-one, not at-least-one) | +| **MAC pinning**: every `node-vm` (and the client-VM) call supplies non-empty `interface_macs`, one per NIC | `opentofu/modules/node-vm/variables.tf:38-62` -- the module ALREADY enforces "empty or exactly one MAC per network_names entry" and rejects partial pinning; `dc-site` must not construct a call that leaves this empty post-provisioning (VR-only trap noted at `:50`) | a fixture `dc-site` call passing `interface_macs = []` for one node (the pre-carve default, `variables.tf:39-40` -- "acceptable ONLY before a node is enlisted anywhere") on a tree tagged post-apply -> assertion FAILS; reuses `node-vm`'s own `error_message` string (`:57,:62`) rather than re-deriving the rule | +| **CIDR / address family** matches the ruled posture per plane (D-139 IPv6-primary, `ipv6-primary-posture` memory; D-143 octet-preserving 10.12->10.13) | D-139, D-143 | a fixture plane input carrying a v4-only CIDR on a plane D-139 rules v6 -> family assertion FAILS | +| **MTU**: `underlay_mtu=9000` threaded unchanged into every plane/edge input (`pass0-admin-report.md` Section 1.4: "removing [the containment hop] changes no byte budget -- do not let later phases imply an MTU benefit") | `variables.tf` `underlay_mtu`; D-101 | a fixture with `mtu=1500` on a plane call (silently regressing the jumbo budget) -> MTU assertion FAILS | +| **Site-token parameterization** -- no DC identity hardcoded in the module body (design principle 1, `pass1-w4-module-planning.md` Section 4 item 1) | existing repo norm, all 12 current modules | a fixture instantiating the SAME module body for `dc0` and `dc1` with only the `$DC` input changed must produce disjoint object names/addressing -- a hardcoded literal inside the body that fails to vary -> FAILS | + +**Ships-where:** `tests/dc-site/run-tests.sh`, offline/static, mirroring `tests/opentofu- +validate/run-tests.sh`'s `--static-only` fixture pattern (T3-T7: fixture `.tf` trees under +`tests/opentofu-validate/fixtures/`, no live provider dial). `dc-site` also automatically rides +the ONE shared IaC gate (`scripts/opentofu-validate.sh`, "validates EVERY module standalone", +`opentofu/README.md:5-6,40-41`, restated `pass1-w4-module-planning.md` Section 1.3) for its own +S1/S2 static guards (memory_unit/ACPI) at zero extra cost -- the dedicated harness above is for +`dc-site`'s OWN composition invariants (count/roster/MAC/family/MTU), which S1-S3 do not cover. +A live `tofu plan`-based resource-count assertion is a stretch addition, flagged OWED-AT-BUILD, +not specified here as fact -- `dc-site` does not exist yet and this pass does not infer its exact +resource graph (hard rule 2). **Dependency, not double-build:** the FIT-calculator extension +(owed #7) should consume `dc-site`'s node-class list once built, rather than re-deriving it. + +### 2.2 The (a) cross-DC host-isolation control -- owed artifact #2 (concern i) + +Spec settled at pass1 Section 3 / pass2 Section 2.1: a vcloud-level nftables artifact, +`--check` gate, own harness, SEC-NNN row, Stage-1 home, installed+verified before ANY flat +substrate apply. Mechanism (exact rule set) NOT yet designed -- specified here mechanism-agnostic. + +| Invariant | Failing-direction fixture | +|---|---| +| No FORWARD rule permits traffic between any two DC plane bridges (dc0<->dc1) | a fixture `nft list ruleset`-style capture WITH a forward-accept rule naming both a dc0 plane bridge and a dc1 plane bridge -> `--check` must FAIL, mirroring `tests/geneve-encap-assert/run-tests.sh`'s C1-family-split fixture shape (fixture text files fed via flags, no live `nft` call) | +| **No host address exists on any plane bridge** (the literal fixture named in this worker's charter) | a fixture `ip addr show`-style capture where a plane bridge (e.g. `br-vr1-dc0-metal-admin`) carries an assigned IP (not merely tap-enslaved interfaces) -> `--check` must FAIL -- this is a DISTINCT failure mode from the forward-rule case: an address ON the bridge lets the HOST itself route between planes even with FORWARD correctly scoped | +| Fail-closed on an absent/unloaded control (never silently PASS) | an EMPTY or missing ruleset capture -> `--check` must FAIL (rc != 0), not report clean -- mirrors `geneve-encap-assert.sh` T11/T12 "refuse-not-pass on empty inputs" | +| No fail-open via an interface-name mismatch (SEC-010's own recorded lesson) | a fixture where the rule's `oifname`/`iifname` targets an interface NAME that does not exist on the host -> `--check` must FAIL, per `docs/security-ledger.md:21`'s own hardening note ("hardened 2026-07-16 -- if the keyed transit interface does not EXIST, since an nftables oifname on an absent iface loads clean but matches nothing = fail-open") and `tests/site-headend-install/run-tests.sh`'s existing assertion of the same class (`ip link show "$TRANSIT_IF"` must be verified, not merely referenced) | +| The control does not globalize (does not break the DC edges' legitimate WAN/uplink egress) | a fixture ruleset using a bare `policy drop` on the whole FORWARD chain (not interface-scoped) -> a "does-not-globalize" assertion must FAIL, per D-125's br_netfilter constraint already recorded for SEC-010 (`site-headend-install.sh` comments ~`:273-296`) -- the SAME class of regression one layer up | + +**Ships-where:** `tests//run-tests.sh` (name minted with the SEC-NNN row), offline +fixture-file harness on the exact `geneve-encap-assert.sh` model (pre-captured text fed via +flags; `PASS=0; FAIL=0; run()` helper). **Two verification tiers, not one** (per `stage-assert- +then-promote`, Section 4): this offline harness proves the SCRIPT's logic; a SEPARATE live +`--check` re-run at B.3 (before any flat apply) and B.7 (once both DCs' planes are actually +co-resident, `pass1-admin-report.md` Part B.7) is the deploy-time gate proving the DEPLOYED +state, not this harness's job to fake. + +### 2.3 The teardown-primitive (module-scoped group-destroy) -- owed artifact #1 (+ #6) + +Re-earns D-122's one-command site-down for the flat shape (`pass1-admin-report.md` Section 6 +item 1). Root-topology (B) recommended -- shared-outer + per-DC-flat roots (`pass2-admin-report.md` +Section 4.1) -- so the primary primitive is a gated `cd && tofu destroy`; the emergency +lever (#6) is a scripted `virsh destroy` loop over that DC's domain set, roster-derived from +`lib-hosts.sh` (`pass1-admin-report.md` Section 6 item 6). Both share one failing-fixture class. + +| Invariant | Failing-direction fixture (the one named in this worker's charter) | +|---|---| +| The derived target/domain set for DC0's destroy contains **zero DC1 objects**, and vice versa | a fixture roster/state-list that (wrongly) includes **a domain from the other DC** (e.g. `vr1-dc1-node-05` appearing in a dc0-targeted destroy's resolved set) -> the target-set assertion must FAIL before any destroy call fires | +| The set also contains **zero non-DC objects** (voffice1, mesh legs, outer pools) under the recommended per-DC-flat root shape | a fixture where a shared-outer object (e.g. `office1_network`) leaks into a per-DC target list -> FAILS | +| The set is COMPLETE for that DC (no legitimate domain silently dropped -- the D-061 pair's own recorded defect: "a host that does not resolve is a `note`, not a `fail`", `docs/tool-index.md`) | a fixture roster missing one of the 13 expected objects for that DC -> a completeness assertion must FAIL, not silently proceed with a partial set | +| Refuses (rc=2), never destroys, against an EMPTY resolved target set (unreachable state/MAAS) | a fixture where the state-list/roster source returns nothing -> the primitive must REFUSE rather than report "nothing to do" as success, mirroring `tests/phase-00-teardown-d061/run-tests.sh`'s decompose-detection FAIL-LOUD pattern (`R3`: post-remove state dropping an expected host FAILS loud and blocks the destructive step) | +| A single-domain canary precedes the group destroy (D-061 precedent) | a fixture where the canary domain fails to actually stop/undefine -> the group destroy must NOT proceed, mirroring `phase-00-teardown-release.sh`'s `--canary` + decompose-check shape | + +**Ships-where:** `tests//run-tests.sh`, stateful-fakebin harness on the exact +`tests/phase-00-teardown-d061/run-tests.sh` model: a fake `tofu`/`virsh` served by fixture +JSON/text selected by a phase-state file; mutating subcommands LOG rather than execute; the +post-mutation re-read is asserted against fixture state, not live. **Note (not yet resolvable):** +the harness's exact fixture SHAPE (root-scoped `tofu state list` vs. `lib-hosts.sh`-derived +domain roster) depends on the still-OPEN root-topology fork (`pass1-admin-report.md` Section 7 +item 1, ratified Phase 4) -- flagged as a dependency, not guessed here (hard rule 2). Emergency +lever (#6) reuses the SAME fixture library/failing-direction class as a second entry point +(scripted `virsh destroy` vs. `tofu destroy`) -- not double-built, per `pass1-admin-report.md` +Section 6 item 6's own "distinct from #1 (emergency vs gated path)" framing. + +### 2.4 The power-key blast-radius mitigation -- owed artifact #11 (concern iii) + +**Verified real this session by pass2** (Section 2.3): each `vr1-dcN-maas-01` region VM holds a +live qemu+ssh virsh credential (`scripts/maas-region-power-key.sh`, SEC-012/SEC-016) that, once +re-pointed at vcloud's own `qemu:///system`, has NO per-domain scoping -- one connection reaches +every domain vcloud manages (both DCs' fleets + voffice1). Mechanism NOT yet chosen (`pass2- +admin-report.md` Section 2.3: "restricted key / wrapper / libvirt polkit ACL -- Phase-4 choice"). + +| Invariant | Failing-direction fixture (the one named in this worker's charter) | +|---|---| +| dc0's power key CANNOT reach any dc1 domain | a fixture allow-list/ACL/`command=`-restriction string that (wrongly) includes **a reachable dc1 domain name** -> the scope assertion must FAIL | +| dc0's power key CANNOT reach `voffice1` | a fixture allow-list including `voffice1` by name/UUID -> FAILS (this is the SECOND half of the named fixture -- "or voffice1" in the charter, not optional) | +| Symmetric for dc1's key against dc0 + voffice1 | mirror fixtures, both directions -- a mitigation validated in only one direction is unproven for the other (SEC-012/SEC-016 are explicitly per-DC, separate keys, `pass2-admin-report.md` Section 2.3 item 1) | +| **Positive coverage, not just absence-of-violation** -- the checker must enumerate what IS reachable and diff it against the DC's OWN roster (`lib-hosts.sh`-derived), not merely grep for known-bad names | a fixture allow-list using a WILDCARD/pattern that silently matches nothing (e.g. a typo'd site-token glob) -> the checker must FAIL this as under-specified/unverifiable, not pass it as "no bad match found" -- the exact fail-open shape SEC-010's own history warns against (`docs/security-ledger.md:21`, `site-headend-install.sh` harness item requiring the transit interface's EXISTENCE be checked, not just its rule) | +| **Negative control -- must NOT over-restrict.** dc0's key must still reach dc0's OWN roster (MAAS enlistment depends on this) | a fixture allow-list that (wrongly) excludes one of dc0's OWN domains -> a same-DC-reachability assertion must FAIL, catching a mitigation that breaks MAAS power control for its own fleet | + +**Ships-where:** `tests//run-tests.sh` (new SEC-NNN, `pass2-admin-report.md` Section 2.3), +offline fixture-file harness parsing a rendered ACL/`authorized_keys`/polkit-rule artifact against +known-good/known-bad domain-name fixtures -- no live libvirt/SSH dial, same shape as `tests/ +site-headend-install/run-tests.sh`'s grep-the-rendered-artifact pattern. **Blocking dependency, +stated in the harness's own header** (repo convention -- every gate names the root cause it +exists for, e.g. `geneve-encap-assert.sh`'s header): this harness's completion is a precondition +for the `lib-hosts.sh` `VIRSH_POWER_ADDRESS_FROM_OFFICE1`/`_FROM_DCREGION` re-derivation and +every `maas-node-power.sh` call-site literal (`pass2-admin-report.md` Section 3.2) -- the +mitigation's chosen mechanism determines the URI/key shape those edits need, so this harness +(and the artifact it tests) must land BEFORE those edits are written, not concurrently. + +### 2.5 SEC-010-successor consolidated installer -- owed artifact #3 (concern ii) + +Endpoints resolved: client VM (DC side) + voffice1 (Office1 side), `pass2-admin-report.md` +Section 2.2 / 4.4. Implementation shape: extract the SEC-010 nftables writer out of +`site-headend-install.sh`'s `node_host_setup()` (`:273-320`) into one role-agnostic subcommand +that installs BOTH ends, closing today's hand-mirrored voffice1 install. + +| Invariant | Failing-direction fixture | +|---|---| +| **FORWARD-drop lands on the client VM's transit leg**, scoped, not global | reuse `tests/site-headend-install/run-tests.sh`'s EXISTING `--transit-if` override + `ip link show "$TRANSIT_IF"` existence-check assertions (already proven failable there) -- must MIGRATE, not be dropped, when the code is extracted | +| **FORWARD-drop lands on voffice1's transit leg too** (the "right legs" invariant named in this worker's charter -- both ends, not one) | a fixture invoking the new subcommand in `voffice1` role mode -> assert a transit-scoped rule is emitted for voffice1's OWN interface, not a copy of the client VM's; a fixture invoking it with the WRONG role's default interface name (e.g. client-VM role using voffice1's leg name) -> the "right legs" assertion must FAIL | +| Never lands on a non-transit leg (metal-admin, WAN/uplink) | a fixture forcing the subcommand to target the client VM's metal-admin interface name -> the emitted ruleset must NOT scope FORWARD-drop to it; a positive check that metal-admin traffic is unaffected must FAIL if it is | +| Never globalizes (D-125 br_netfilter constraint, verbatim requirement already enforced for the OLD SEC-010 writer) | reuse `tests/site-headend-install/run-tests.sh`'s existing `br_netfilter`/"never global" grep assertion against the new extraction target -- migration-completeness, not a new invariant | +| Idempotent reload (declare-then-delete preamble; `nft -f` on a live table APPENDS otherwise) | reuse the existing `delete table inet sec010` presence assertion against the new location | + +**Ships-where:** if the extraction stays a subcommand of `site-headend-install.sh`, extends +`tests/site-headend-install/run-tests.sh`; if it becomes its own script, a new `tests// +run-tests.sh` inheriting EVERY SEC-010-related assertion already proven in the current harness +(item 8 in that file: transit-if override, node-host-mode presence, br_netfilter constraint, +idempotent-reload preamble) -- a **migration-completeness check** (grep the new location for +every trap-string the old harness asserted) is itself a required test, so the extraction cannot +silently drop a proven guard. + +### 2.6 D-131 retire-with-evidence step -- owed artifact #13 + +Retirement is the ruled 10.13 end state for BOTH DCs once D-132's per-DC regions remove the +forwarder's precondition (`pass2-admin-report.md` Section 4.2(ii)). dc0 already proves the end +state (dig-verified, `docs/changelog-20260730-dc0-region-migration.md` Item 9); **dc1's +forwarder is CURRENTLY load-bearing** (`docs/changelog-20260807-dc1-region-sequence.md:80-89`, +config "replicated verbatim", `dns_servers=10.12.68.3` the forwarder alias) -- this asymmetry +must not be assumed equal (`pass2-admin-report.md` check 4). + +| Invariant | Failing-direction fixture (the dc0/dc1 asymmetry named in this worker's charter) | +|---|---| +| Nodes resolve via the DC's OWN region BIND **directly**, not via the D-131 forwarder alias | a fixture `dig` capture that SUCCEEDS (answers correctly) but whose ANSWERING SERVER is the forwarder alias IP (`10.12.68.3`-class), not the region's own BIND (`10.12.68.6`-class per dc1's `.6` region VM) -> the checker must FAIL this, because a "did resolution succeed" test alone would PASS on dc1's still-load-bearing forwarder and falsely report retirement complete | +| dc0 passes (already the proven end state) | a fixture matching dc0's actual measured dig evidence (`dns_servers=10.12.8.6`, `flags: qr rd ra`, ANSWER: 9, `docs/changelog-20260730-dc0-region-migration.md` Item 9) -> must PASS, proving the checker is not just tuned to fail everything | +| **dc1 must be explicitly closed, not silently inherited as passing** | a fixture reproducing dc1's CURRENT (unretired) config verbatim (`docs/changelog-20260807-dc1-region-sequence.md:80-89`) -> the checker must FAIL dc1 today, and the harness's own dc1 test case must be RED until the live retirement actually happens -- this is the asymmetry as a standing red case, not a hypothetical | +| The checker asserts the resolver's IDENTITY, not merely that a name resolved | same fixture pair as row 1 -- restated because it is the entire point: a checker that only checks "resolution worked" is provably insufficient here and must not be shipped | + +**Ships-where:** `tests//run-tests.sh`, offline, fixture = captured `dig +short`/`dig ++stats`-style text comparing the answering resolver's IP against the DC's own region IP (read +from `lib-hosts.sh`/`lib-net.sh`, never a duplicated literal). Small, single-purpose L3 gate; +invoked per-DC at the retirement decision point and again at B.5's placement close-out +(`pass1-admin-report.md` Part B.5). + +--- + +## 3. The remaining seven owed artifacts -- harness coverage, concise + +| # | Artifact | Harness disposition | +|---|---|---| +| 4 | R7 credential-revocation checklist | NEW gate, offline: fixture `vm-secret-locations` register rows (mock) keyed to the retiring rack host class; asserts EVERY matching row is enumerated. Failing fixture: an unlisted/orphan row for that host class the enumeration misses -- the exact defect class SEC-027's ledger row names verbatim ("an unlisted location is not audited," `docs/security-ledger.md:80`) | +| 5 | MAAS machine-record release/delete step (+ rack-controller decommission) | NEW gate, fakebin `maas` on the `tests/dc-egress-check` / `tests/phase-00-teardown-d061` model; asserts post-release re-read (LENS-2) reaches zero AND the region-side rack-controller + `primary_rack`/DHCP reference are cleared. Failing fixture: a post-release fixture where one machine record OR the rack-controller record survives -- must not report clean | +| 6 | Emergency site-down lever (`virsh destroy` loop) | Shares Section 2.3's harness/fixture library (same cross-DC-domain failing fixture) -- explicitly not double-built (`pass1-admin-report.md` Section 6 item 6) | +| 7 | FIT-calculator extension + fresh capacity measure | EXTENDS the existing `tests/dc-dc-whole-host-budget/` harness (already YES, `pass2-w4-module-decomposition.md` Section 2) with the 3 utility-node classes + the artifact-service disk-sizing branch. Failing fixture: a roster total EXCEEDING the measured host budget must FAIL the calculator, not silently round or omit a class | +| 8 | MAC re-measurement pass (post-apply) | No new dedicated harness -- rides `dc-site`'s own MAC-pinning invariant (Section 2.1 row 4) once that module is built and applied; cross-reference only | +| 9 | NetBox DCIM migration (decommission `vvr1-dcN`; register client VM + roster) | EXTENDS the existing `netbox/dc-rack-mgmt-import.py` harness (`pass0-admin-report.md` row 11) with a new failing fixture: a stale `vvr1-dcN` device record surviving import must FAIL a decommission-completeness assertion. Not a new module | +| 10 | Post-build live asserts (geneve/jumbo re-verify; gap-#20 re-verify) | Rides the EXISTING `tests/geneve-encap-assert/run-tests.sh` verbatim (already fully fixture-proven, Section 2.2 confirms its shape) -- no new harness; only a new invocation point (post-flatten, vcloud-level planes) belongs in the runbook, not the test suite | + +All 13 owed artifacts are accounted for: 6 with full new-harness specs above, 5 extending an +existing harness with a new failing-direction fixture, 2 riding an existing invariant/harness +with no new build. + +--- + +## 4. The per-module harness CONTRACT -- the template every one of these follows + +Grounded entirely in patterns already proven in this repo (cited per row), not invented for +this pass -- consistent with `pass2-w4-module-decomposition.md` Section 3's procedure-module +contract and this pass's own charter (`SCOPE-AND-EXECUTION-PLAN.md` RULES). + +| Principle | What it requires | Repo precedent | +|---|---|---| +| **Prove-it-can-fail** | Every assertion ships with a fixture engineered to make it FAIL; a mutation pass that deletes the assertion must turn the suite red. A clean PASS with no paired failing fixture is not a gate. | `geneve-encap-assert`'s T5/T6/T11/T12 (family-split, ofport -1, empty-input refuse); `phase-00-teardown-d061`'s R3 decompose-detection; the repo's own named failure mode, `docs/tool-index.md`'s D-061 "note not a fail" defect | +| **Assert the ARTIFACT, not the config/intent** | Check the INSTALLED rule, the DEPLOYED overlay value, the RENDERED file -- never a comment, a doc string, or a generator's self-report of what it meant to do. | `docs/changelog-20260730-octavia-reissue-tool.md:158-167`, A17: "graded the workspace copy... the charm gets this one ($OVLCMP)" -- the consumed copy is what's graded | +| **Stage-assert-then-promote** | Any harness whose subject WRITES material (mints a credential, renders a config, applies an nftables table) writes to a STAGING path first, asserts the STAGED artifact, and only then promotes/applies -- never assert an in-place mutation after the fact with no rollback point. | `docs/changelog-20260730-octavia-reissue-tool.md:158-167` (four cited properties incl. stage-assert-promote); `scripts/octavia-pki.sh` staging-dir promotion flow | +| **Join the workspace to the deploy input** | When a generated artifact is separately CONSUMED downstream (an overlay literal, a `lib-hosts.sh` value, a bundle var), assert BYTE/VALUE EQUALITY between the generator's output and what the consumer actually reads -- not just that each was independently produced correctly. | `docs/changelog-20260730-octavia-reissue-tool.md:158-167`; `scripts/octavia-pki.sh` A17 (the exact defect: two things independently graded correct, nothing joined them, and the deployed copy diverged) | +| **Fail-closed on absent/empty/unreachable input** | An empty, missing, or unresolvable input REFUSES or FAILS -- it never reports a clean PASS by default. Includes the interface-name fail-open class: a rule referencing an absent name loads clean and matches nothing. | `docs/security-ledger.md:21` SEC-010's own 2026-07-16 hardening; `geneve-encap-assert` T11/T12; `site-headend-install`'s transit-interface-existence check | +| **Offline/fixture-driven by default** | No live cloud dependency in the harness itself -- a stateful fakebin or captured-text fixture stands in for the live system. A SEPARATE, explicitly-named live gate re-run (not this harness) proves the deployed artifact. | Every harness surveyed this session (`geneve-encap-assert`, `site-headend-install`, `phase-00-teardown-d061`, `dc-egress-check`, `dc-node-v6-verify`); `dc-node-v6-verify`'s own disclaimer: "WHAT THE GREEN BELOW DOES NOT PROVE: no node has ever been asserted by it" | +| **`$SITE`/`$DC`-parameterized, never hardcoded** | A fixture exercises BOTH dc0 and dc1 (or a synthetic third site) to prove the checker generalizes -- not that it happens to pass against dc0's literals. | `pass1-w4-module-planning.md` Section 4 item 1 (design principle 1); DOCFIX-151 `lib_net_select_dc`/`lib_hosts_select_dc` | +| **Standard exit contract** | Adopt `lib-validate.sh`'s 0/1/2/3/4 PASS/FAIL/HOLD/PASS_PENDING_MANUAL/SKIPPED vocabulary for any verify-mode gate, so it composes into the existing G-series/preflight aggregation. | `scripts/lib-validate.sh:16-24` | +| **Delivery discipline** | Ships with its own `tests//run-tests.sh`, a changelog entry with a revert, and is `repo-lint` clean -- no exception for a "new" module. | CLAUDE.md "Delivery"; `pass2-w4-module-decomposition.md` Section 3 item 7 | + +--- + +## 5. Open items (not resolved by this worker; feed the Phase-3 administrator / Phase 4) + +1. Four of the six full-spec harnesses (2.2 isolation control, 2.3 teardown primitive, 2.4 + power-key mitigation, 2.5 SEC-010 successor) specify invariants and failing fixtures + MECHANISM-AGNOSTICALLY because their underlying mechanism is still OPEN (`pass2-admin-report.md` + Section 6) -- the harness SHAPE (offline fixture-file, on the `geneve-encap-assert.sh`/`site- + headend-install.sh` model) is fixed; the exact fixture CONTENT is not writable until the + mechanism is ruled. +2. `modules/dc-site`'s live `tofu plan`-based resource-count check (Section 2.1) is a stretch + item flagged OWED-AT-BUILD, not specified as fact -- the module does not exist yet. +3. The teardown-primitive's fixture shape (Section 2.3) is explicitly contingent on the + root-topology fork ratification (Phase 4). +4. This worker did not review W3.1 (existing harnesses assuming the container layer) or W3.2 + (preflight/cloud-assert/stage gates for a flat topology) -- those are sibling workers' domains + per `SCOPE-AND-EXECUTION-PLAN.md` Section 4; only this worker's per-new-module dimension is + covered here. + +--- + +## 6. Verification note + +Author = the W3.3 worker (no model name asserted). Direct reads this session: the five admin/ +worker documents named in the header, in full; six existing `tests/*/run-tests.sh` harnesses +read in full or substantially; `scripts/lib-validate.sh` header; `docs/security-ledger.md:21` +(SEC-010) and `:79-80` (SEC-026/SEC-027) read directly; `docs/changelog-20260730-octavia- +reissue-tool.md:140-167` read directly for the stage-assert-promote / join-workspace vocabulary; +`opentofu/modules/node-vm/variables.tf:38-62` read directly for the MAC-pinning validation this +spec reuses. READ-ONLY; nothing executed; findings and specs are LOGGED only, per the pass's +charter. diff --git a/docs/audit/container-elim-pass/pass4-w1-master-change-inventory.md b/docs/audit/container-elim-pass/pass4-w1-master-change-inventory.md new file mode 100644 index 0000000..88a5520 --- /dev/null +++ b/docs/audit/container-elim-pass/pass4-w1-master-change-inventory.md @@ -0,0 +1,325 @@ +# Pass 4 / W4.1 -- MASTER CHANGE INVENTORY (container-layer elimination) + +**Author:** Worker W4.1 (Phase 4 -- change synthesis), multi-agent container-elim pass +(`SCOPE-AND-EXECUTION-PLAN.md` Section 4). **Date:** 2026-08-09. **Inputs:** +`pass0-admin-report.md` (baseline + confirmed target topology), `pass1-admin-report.md` +(planning change-set), `pass2-admin-report.md` (tools change-set), `pass3-admin-report.md` +(tests change-set) -- all four read in full. READ-ONLY synthesis; no mutation; nothing +here is executed. This document MERGES the planning + tools + tests change-sets into one +indexed table an execution session can work from directly. + +**Baseline carried in (do not re-derive):** Option 1 CONFIRMED (flat node VMs on vcloud +libvirt + one small non-hypervisor `vr1-dcN-client` VM per DC); cross-DC handling (a) +CONFIRMED (new vcloud-level host isolation control); MAAS region stays on +`vr1-dcN-maas-01`; root topology (B) shared-outer + per-DC-flat RECOMMENDED (Phase-4 +ratifies); THREE isolation controls confirmed distinct -- (i) the (a) cross-DC host +control, (ii) the SEC-010 transit-leg successor, (iii) the MAAS power-key blast-radius +mitigation; everything rides D-143 (10.12->10.13 re-IP); 13 owed artifacts (numbered #1-#13 +below); tests change-set = 9 existing-change verdict-blocks / 5 existing-retire +verdict-blocks / 7 new-build harnesses / 3 rides. + +**ID scheme:** `DEC-` decision (not a change; an open ruling) -- `TF-` tofu-module -- +`LB-` lib (`lib-hosts.sh`/`lib-net.sh`) -- `SC-` script/procedure -- `DC-` doc +(deployment-workflow.md / CURRENT-STATE.md prose, non-gate) -- `RB-` runbook -- `GT-` gate +(preflight/cloud-assert/G-series) -- `HN-` harness (`A#`/`B#`/`C#` tags preserved from +`pass3-w3-new-tests.md` for cross-reference) -- `SEC-` security-ledger row. +**Axis:** `[CE]` container-elim only, `[D-143]` re-IP only, `[both]` dual-labeled (per +pass1 check 5's four confirmed dual items + this pass's extensions). **Owed-artifact#** +refers to the 13-item list in `pass2-admin-report.md` Section 5 (also restated below). + +--- + +## 0. The 13 owed artifacts (for cross-reference; full spec in `pass2-admin-report.md` #5) + +1. Teardown primitive (root-scoped `tofu destroy` + emergency lever) -- 6. rides as the + emergency `virsh destroy` loop (distinct row, same fixture library) +2. The (a) cross-DC host isolation control (concern i) +3. SEC-010 transit-leg successor (concern ii) +4. R7 credential-revocation checklist +5. MAAS machine-record release/delete step (+ rack-controller decommission) +7. FIT-calculator extension + capacity measure (+ artifact-service sizing) +8. MAC re-measurement pass (post-apply) +9. NetBox DCIM migration +10. Post-build live asserts (geneve/jumbo) +11. Power-key blast-radius mitigation (concern iii) -- **critical path** +12. `modules/dc-site` +13. D-131 retirement-evidence step + +--- + +## 1. MASTER CHANGE-INVENTORY TABLE + +### 1.1 Decisions (open rulings -- not changes; gate the change rows below) + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| DEC-01 | decision | `docs/design-decisions.md` -- container-elim [ARCH] ruling (D-123 amendment vs new D-number) | new | none (root ruling; operator rules, GA-R5) | -- | [CE] | +| DEC-02 | decision | D-128 amendment ratification (Plane 2 shrinks to MAAS/NetBox; substrate build becomes wholly Plane 1) | new | DEC-01 | -- | [CE] | +| DEC-03 | decision | D-125 bridge-in retirement note (rides DEC-01) | new | DEC-01, TF-03 | -- | [CE] | +| DEC-04 | decision | D-138 concrete-host change (client VM replaces `vvr1-dcN` as the concrete host) | new | DEC-01 | -- | [CE] | +| DEC-05 | decision | D-122 site-down re-earn note (one-command site-down lost; re-earned via SC-09) | new | DEC-01 | -- | [CE] | +| DEC-06 | decision | D-124 sizing-void re-cause note (rack-addressing vars deleted with TF-01) | new | DEC-01 | -- | [CE] | +| DEC-07 | decision | D-132-addendum premises note (hypervisor-fate rationale moot under Option 1) | new | DEC-01, DEC-08 | -- | [CE] | +| DEC-08 | decision | Rack-controller retirement ratification (+ live re-measure of `primary_rack` both DCs) | new | live measurement (delivery-time, owed) | -- | [CE] | +| DEC-09 | decision | D-131 forwarder retire-with-evidence ratification (per-DC; dc1 asymmetry) | new | SC-15 (#13), DEC-08 | 13 | [CE] | +| DEC-10 | decision | Artifact-service (`.4`) placement + sizing decision | new | SC-16 (#7 FIT ext w/ mirror sizing) | 7 | [CE] | +| DEC-11 | decision | Root topology ratification: (B) shared-outer + per-DC-flat vs merged single root | new | none (Phase 4 ratifies recommendation) | -- | [CE] | +| DEC-12 | decision | Root naming (`vr1-dcN-flat` vs reserving `-substrate`) | new | DEC-11 | -- | [CE] | +| DEC-13 | decision | Client-VM octet `.8` + name `vr1-dcN-client` into D-134 standing map | new | none (recommended) | -- | [CE] | +| DEC-14 | decision | (a) control's concrete mechanism (nftables rule set / check shape / SEC-NNN) | new | none | 2 | [CE] | +| DEC-15 | decision | Concern-(iii) power-key mitigation mechanism choice (restricted key / wrapper / polkit ACL) + SEC-NNN -- **CRITICAL PATH** | new | none | 11 | [CE] | +| DEC-16 | decision | SEC-010 successor SEC-row disposition (new row vs amendment); endpoint ratification (client VM + voffice1, already recommended) | new | none | 3 | [CE] | +| DEC-17 | decision | `wan-bridge` module directory: delete vs leave-unreferenced (append-only bias) | new | DEC-01, TF-03 | -- | [CE] | +| DEC-18 | decision | SEC-013 `maas-vm-host` retire-or-keep (flagged to its owner, not this pass) | new | none | -- | [CE] | +| DEC-19 | decision | `maas-fabric-prune.sh`/`maas_fabric_classify.py` harness gap: build vs accept-as-named-exception (pre-existing, container-elim-adjacent only) | new | none | -- | [CE-adjacent] | +| DEC-20 | decision | A11's home: fold into `cloud-assert.sh` vs a dedicated `isolation-assert.sh` | new | DEC-14 | -- | [CE] | +| DEC-21 | decision | NetBox-migration design: rename-in-place vs concept retirement (for HN-A6/#9) | new | none | 9 | [CE] | +| DEC-22 | decision | D-127 client-VM autostart value (needed for HN-A1's new case) | new | none | -- | [CE] | +| DEC-23 | decision | State-blast-radius weighing (rides DEC-11) | new | DEC-11 | -- | [CE] | + +### 1.2 OpenTofu roots/modules + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| TF-01 | tofu-module | `opentofu/main.tf` `module vvr1_dc0/_dc1` + sizing/rack-addressing/pubkey vars (`variables.tf:137-156,175-194,196-244`) | retire | DEC-01, DEC-11 | -- | [CE] | +| TF-02 | tofu-module | `opentofu/vr1-dc0-substrate/`, `vr1-dc1-substrate/` (whole inner roots + states) | retire (as roots; module bodies re-home) | DEC-11, TF-12, TF-13 | -- | [CE] | +| TF-03 | tofu-module | `modules/wan-bridge` (+ `vr1_dcN_wan` calls, IP-less uplink NIC, `br-vr1-dcN-wan` netplan) | retire | DEC-01, DEC-17 | -- | [CE] | +| TF-04 | tofu-module | `modules/site-wan` output rewire (feeds DC edge directly, no bridge-in) | change | TF-03 | -- | [CE] | +| TF-05 | tofu-module | `modules/cloudinit-vm` (loses 2 containment calls, gains the client-VM call) | re-home | TF-12, DEC-13 | -- | [CE] | +| TF-06 | tofu-module | `modules/dc-planes` (6 planes re-homed to vcloud level; same CIDRs/families/MTU) | re-home | TF-12 | -- | [CE] (shape only; values D-139/D-143-owned) | +| TF-07 | tofu-module | `modules/dc-storage-pool` (2-per-DC collapses to 1) | re-home | TF-12 | -- | [CE] | +| TF-08 | tofu-module | `modules/node-vm` x12/DC (unchanged body, re-homed call site) | re-home | TF-12 | -- | [CE] | +| TF-09 | tofu-module | `modules/opnsense-edge` (one input re-pointed to TF-04's direct NAT) | change | TF-04 | -- | [CE] | +| TF-10 | tofu-module | `modules/base-image` (re-homed call site, no logic change) | re-home | TF-12 | -- | [CE] | +| TF-12 | tofu-module | **NEW** `modules/dc-site` (composes pool + 6 planes + edge + 12 node VMs + client VM; replaces the ~230-266-line copy-pasted per-DC inner-root bodies) | new | DEC-11, DEC-13 | 12 | [CE] | +| TF-13 | tofu-module | **NEW** per-DC flat root files (shared-outer + per-DC-flat, 3 roots total) invoking `modules/dc-site` | new | DEC-11, DEC-12, TF-12 | -- | [CE] | +| TF-14 | tofu-module | D-124 transit-leg re-point (Office1-leg consumer: `vvr1-dcN` NIC1 -> client-VM transit NIC; `mesh-link`/`netem-link` bodies unchanged) | change | TF-12/TF-13, DEC-13 | -- | [both] (octet math D-143, bearer host CE) | + +Note: `modules/office1-network`, `mesh-link` (x3), `netem-link` are CONFIRMED UNCHANGED +(pass2 check 3/4.5) -- not itemized as rows. `modules/maas-vm-host` is dead/orthogonal, +never instantiated -- see DEC-18, not itemized as a change row. + +### 1.3 `lib-hosts.sh` / `lib-net.sh` + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| LB-01 | lib | `scripts/lib-hosts.sh` `VIRSH_POWER_ADDRESS_FROM_OFFICE1`/`_FROM_DCREGION` (`:212-213,246-251`) | change | DEC-15 -- **BLOCKED** | feeds 11 | [CE] | +| LB-02 | lib | `scripts/lib-hosts.sh` `REGION_HOST_SUFFIX` comment (`:95-100`) | change | none | -- | [CE] (comment-currency, low priority) | +| LB-03 | lib | `scripts/lib-net.sh` (whole file) | change | D-143 ruling (separate axis) | -- | [D-143] (ZERO container-elim edits, grep-verified; noted here only so the axis is not conflated) | + +Note: `CARVE_AUX_HOSTS`, `NIC_PLANE_ORDER`, `BREX_PARENT_NIC`, `HOST_OCTET` maps/suffixes, +`HOST_TAG`, resolver fns -- UNCHANGED under container-elim (octet maps change under D-143 +only). The client VM does NOT get a `lib-hosts.sh` row (resolved: it is L1 `cloudinit-vm`, +not MAAS/virsh-power-managed; identity lives in tofu + NetBox). + +### 1.4 Scripts + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| SC-01 | script | `scripts/maas-node-power.sh` invocation-site/runbook literals (no code change to the script itself -- address is `$1`) | change | LB-01, DEC-15 -- **BLOCKED** | feeds 11 | [CE] | +| SC-02 | script | `scripts/dc-rack-net.sh` -- LEGS/`br_of()` half retires; DNS-forwarder half depends on D-131 | change/retire (split) | DEC-09, TF-01 | -- | [CE] | +| SC-03 | script | `scripts/site-headend-install.sh` `node_host_setup()`/`node_host_check()` `--host-nodes` (~134 lines) | retire | DEC-01 | -- | [CE] | +| SC-04 | script | `scripts/site-headend-install.sh` SEC-010 writer extraction (`:273-320`) into a role-agnostic subcommand installing BOTH ends | change | DEC-16 | 3 | [CE] | +| SC-05 | script | `scripts/site-headend-install.sh` `--role rack` (Section 6) | retire (contingent) | DEC-08 | -- | [CE] | +| SC-06 | script | `scripts/dc-mirror.sh` / `dc-cache-proxy.sh` / `dc-snap-proxy.sh` -- new host + explicit disk sizing | change | SC-16 (#7), DEC-10 | rides 7 | [CE] | +| SC-07 | script | `scripts/maas-region-power-key.sh` (body unchanged; URI/key shape it installs re-derives) | change | DEC-15 -- **BLOCKED** | feeds 11 | [CE] | +| SC-08 | script | `scripts/site-baseleg.sh` comment block (re-cite D-138 + the (a) control, not the retired qemu+ssh premise) | change | none | -- | [CE] (doc-currency; stays a no-op) | +| SC-09 | script | **NEW** teardown primitive (module/root-scoped `tofu destroy` procedure) | new | DEC-11 | 1 | [CE] | +| SC-10 | script | **NEW** (a) cross-DC host isolation control (nftables artifact) | new | DEC-14 | 2 | [CE] | +| SC-11 | script | **NEW** power-key blast-radius mitigation (restricted key / wrapper / polkit ACL) -- **CRITICAL PATH** | new | DEC-15 | 11 | [CE] | +| SC-12 | script | **NEW** R7 credential-revocation checklist (enumerate every `vm-secret-locations` row keyed to the rack host class) | new | none (ready) | 4 | [both] | +| SC-13 | script | **NEW** MAAS machine-record release/delete step (+ `maas rack-controller delete` decommission + region+rack runbook note) | new | DEC-08 (decommission half) | 5 | [D-143 primary, CE ride] | +| SC-14 | script | **NEW** emergency site-down lever (`virsh destroy` loop over the DC root's domain set, roster from `lib-hosts.sh`) | new | DEC-11, SC-09 | 6 | [CE] | +| SC-15 | script | **NEW** D-131 retirement-evidence checker (dig test against each fresh region's own BIND) | new | DEC-09 | 13 | [CE] | +| SC-16 | script | `scripts/dc-dc-whole-host-budget.py` FIT-calculator extension (3 utility-node classes + artifact-service disk-sizing branch) + fresh vcloud capacity measurement | change | none (ready) | 7 | [both] | +| SC-17 | script | MAC re-measurement pass (post-apply, before B.6 trusts any MAC -- likely force-replace) | change (procedure) | TF-13 | 8 | [CE] | +| SC-18 | script | `netbox/dc-rack-mgmt-import.py` (decommission `vvr1-dcN` DCIM records; register client VM + flat roster) | change | DEC-21 | 9 | [CE] | +| SC-19 | script | `scripts/geneve-encap-assert.sh` -- new invocation point post-build (no code change; verbatim re-run) | change (new invocation only) | TF-13 | 10 | [both] (MTU budget analytically unchanged, live assert still owed) | + +Note: `dc-node-carve.sh`, `dc-node-v6-carve.py`, `carve-host-interfaces.sh`, +`maas-role-tags.sh`, `maas-profile-assert.sh`, `maas-role-tags.sh`, `dc-egress-check.sh` +logic bodies -- NO CODE CHANGE (grep-verified zero containment hits; pure MAAS-API, +``-parameterized); only invocation-host currency (D-128-amendment territory) and doc +comments naming `vvr1-dcN` need updating -- LOW priority, not itemized as separate rows. + +### 1.5 Doc (workflow doc + gate-table prose, non-runbook) + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| DC-01 | doc | `docs/dc-dc-deployment-workflow.md` Stage 3 (Build/Gate/Owns/Reuse-vs-new lines) | change | DEC-01, DEC-11 | -- | [CE] | +| DC-02 | doc | `docs/dc-dc-deployment-workflow.md` Stage 4 gate-line + G17 literal | change | DEC-08 (rack placement), D-143 address | -- | [both] | +| DC-03 | doc | `docs/dc-dc-deployment-workflow.md` Stage 5 literals (transit-IP re-point, e.g. `docs/CURRENT-STATE.md:7829` "openstackclient ... ON THE dc0 RACK (172.31.0.2)") | change | DEC-13 | -- | [CE] | +| DC-04 | doc | `docs/dc-dc-deployment-workflow.md` gap register: NEW entry for the (a) control | new | SC-10 | 2 | [CE] | +| DC-05 | doc | `docs/dc-dc-deployment-workflow.md` gap register: #2 reshapes, #17 closing-mechanism note goes historical, #20 verdict re-verify (its own expiry clause triggers) | change | TF-13 | -- | [both] | +| DC-06 | doc | `docs/dc-dc-deployment-workflow.md` Stage 2 -- explicit two-containment-patterns-distinction note (D-114 KEPT vs D-123 RETIRED, so name-similarity does not sweep Stage 2 in) | new | none | -- | [CE] | + +Note: Stages 1, 6, 7 and the `dc-dc-office1-service-reip.md` / `dc-dc-phase0-vcloud-prep.md` +/ `dc-dc-phase1-office1-standup.md` runbooks are CONFIRMED NO CHANGE / OUT OF SCOPE (D-114, +zero containment hits) -- not itemized. + +### 1.6 Runbooks + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| RB-01 | runbook | `runbooks/dc-dc-teardown-rollback.md` | change (rewrite) | DEC-11, SC-09, SC-13 | rides 1,5 | [both] | +| RB-02 | runbook | `runbooks/dc-dc-phase2-tofu-dc-substrate.md` | change (heaviest rewrite) | TF-12, TF-13, SC-10 | rides 2,12 | [CE] | +| RB-03 | runbook | `runbooks/dc-dc-phase3-maas-enlist-deploy.md` (SSH-jump-target lines `:424,430`) | change (low delta) | DEC-08 (rack/placement ruling) | -- | [CE] | +| RB-04 | runbook | `runbooks/dc-dc-phase4-juju-bundle-per-dc.md` (RUN-LOCATION table 3rd correction) | change | DEC-13 | -- | [both] (overlay literals D-143, execution-host CE) | +| RB-05 | runbook | `runbooks/dc-dc-phase6-designate-cos-magnum.md` (`:437-444` pre-existing stale-D-138 defect) | change (ride-along fix, not a container-elim delta) | none | -- | [pre-existing; rides RB-04's sweep] | + +Note: `dc-dc-phase5-dr-failover-drill.md` has NO DIRECT CHANGE -- it inherits RB-04's table; +not itemized as its own row. + +### 1.7 Gates (preflight / cloud-assert / G-series) + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| GT-01 | gate | **NEW** Stage-1 gate for the (a) control -- `--check` enumerates the live bridge set for both DCs' six planes, asserts FORWARD denial between every dc0-tagged/dc1-tagged bridge pair, REFUSES if fewer than the full plane count resolves | new | SC-10, DEC-14 | 2 | [CE] | +| GT-02 | gate | `cloud-assert.sh` A11a (periodic re-verify of the (a) control: post-deploy/restart/pre-change/post-incident) | new | GT-01 | -- | [CE] | +| GT-03 | gate | `preflight.sh` P10 (SEC-010 successor / concern ii, DC-scoped, host-bound on the P7 model; one installer/checker covers both ends) | new | SC-04, DEC-16 | 3 | [CE] | +| GT-04 | gate | `preflight.sh` P4 dependency + `cloud-assert.sh` A11b (power-key concern-iii verification -- negative test: a DC's region key cannot reach domains outside its own roster) -- **CRITICAL PATH** | new/change | SC-11, DEC-15 | 11 | [CE] | +| GT-05 | gate | `preflight.sh` P8 substrate-drift loop -- extend from one hardcoded path to a **DECLARED list** of every post-flatten root (never a glob -- administrator amendment, a glob cannot fail on a missing/renamed root) | change | DEC-12 | -- | [CE] | +| GT-06 | gate | `preflight.sh` P5 creds-matrix register rows re-point (`rack`-class -> `client`-class); NEW rows for SC-04(#3)/SC-11(#11) key material when minted | change (data) | SC-04, SC-11 | -- | [both] | +| GT-07 | gate | `preflight.sh` P9 (`dc-egress-check` invocation-host literal, re-points to the ruled B.5 host) | change | DEC-08 (B.5 placement ruling) | -- | [CE] | +| GT-08 | gate | `docs/CURRENT-STATE.md` G9/G10 successor -- single apply-and-verify gate (substrate apply + (a) `--check` + depth-2 boot proof + direct-NAT egress test), replacing the outer/inner pair | change | TF-13, GT-01 | -- | [CE] | +| GT-09 | gate | `docs/CURRENT-STATE.md` G17 -- the one dual-cause gate-literal edit (new address family D-143 + new host container-elim, in one edit) | change | DEC-08, D-143 ruling | -- | [both] | +| GT-10 | gate | `docs/CURRENT-STATE.md` G14 (indirect -- residency re-points + >=1 new SEC row are count-affecting; flag for the next `ledger-scan.sh` reader, instrument-currency lesson #25) | change (flag only) | SEC-01, SEC-02 | -- | [CE] | + +Note: G12 is CLOSED/historical, read as "the shape being replaced," not touched further. +G18, G1-G8, G11, G13, G15, G16 and cloud-assert A0-A10 -- CONFIRMED no container-layer +dependency (verified per-gate by W1.2/W3.2); not itemized. + +### 1.8 Harnesses (tests/) -- A/B/C tags preserved from `pass3-w3-new-tests.md` + +**A: existing -- change (9 verdict-blocks)** + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| HN-A1 | harness | `tests/opentofu-validate` T14/T15 (autostart pins) + new `vr1-dcN-client` case | change | DEC-12 (root-naming), DEC-22 (D-127 client-VM value) | -- | [CE] | +| HN-A2 | harness | `tests/node-vm` T8-T11 (hardcoded `INNER=` path) | change | DEC-12 | -- | [CE] | +| HN-A3 | harness | `tests/site-headend-install` Section-8 SEC-010 sub-case (= owed #3's harness half) | change | SC-04 | 3 | [CE] | +| HN-A4 | harness | `tests/dc-selector` power-address rows (`:204-247`) | change | DEC-15, SC-11 -- **BLOCKED** | feeds 11 | [CE] | +| HN-A5 | harness | `tests/maas-region-power-key` URI assertions (`:67,82,102,108`) -- edited TOGETHER with HN-A4, same session | change | DEC-15, SC-11, HN-A4 -- **BLOCKED** | feeds 11 | [CE] | +| HN-A6 | harness | `tests/dc-rack-mgmt-import` vvr1 pins (`:77-78`, = owed #9's harness half) | change | DEC-21, DEC-08 | 9 | [CE] | +| HN-A7 | harness | `tests/maas-profile-assert` office1-profile fixture (`:41,68,72,80`) | change | DEC-08 | -- | [CE] | +| HN-A8 | harness | `tests/dc-dc-whole-host-budget` (= owed #7's harness; the one universe-boundary crossing) | change | SC-16 | 7 | [both] | +| HN-A9 | harness | `tests/pre-flight-checks` -- NEW case post-#11 (live power-address must match the mitigation's issued shape) | new (case) | SC-11, GT-04 -- **BLOCKED** | 11 | [CE] | + +**B: existing -- retire (5 verdict-blocks)** + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| HN-B1 | harness | `tests/opentofu-validate` T13 (D-127 containment autostart pin) | retire | TF-01 | -- | [CE] | +| HN-B2 | harness | `tests/site-headend-install` `--host-nodes` block (~15 cases; excludes the SEC-010 sub-case = HN-A3) | retire | SC-03 | -- | [CE] | +| HN-B3 | harness | `tests/site-headend-install` Section 6 (`--role rack`) | retire (contingent) | DEC-08 | -- | [CE] | +| HN-B4 | harness | `tests/dc-rack-net` LEGS cases (T3,T5,T13,T15,T17) | retire | SC-02 | -- | [CE] | +| HN-B5 | harness | `tests/dc-rack-net` DNS-forwarder cases (T4,T9,T16) + 8 hygiene cases | retire (contingent) | DEC-09 | -- | [CE] | + +**C: new-build (7 harnesses, one per owed artifact)** + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| HN-C1 | harness | `tests/dc-site/run-tests.sh` (static fixture `.tf` trees) | new | TF-12 | 12 | [CE] | +| HN-C2 | harness | (a) control offline fixture harness | new | SC-10, DEC-14 | 2 | [CE] | +| HN-C3 | harness | Teardown-primitive fixture library (+ #6 emergency lever rides the same library) | new | SC-09, DEC-11 | 1, 6 | [CE] | +| HN-C4 | harness | Power-key mitigation harness -- **itself a precondition for HN-A4/HN-A5/HN-A9** | new | SC-11, DEC-15 | 11 | [CE] CRITICAL | +| HN-C5 | harness | D-131 retire-evidence checker harness (fixtures exist in the two cited changelogs) | new | SC-15 | 13 | [CE] (fixture-ready) | +| HN-C6 | harness | R7 credential-revocation checklist harness | new | SC-12 | 4 | [both] (ready) | +| HN-C7 | harness | MAAS record release/delete (+ rack decommission) harness | new | SC-13 | 5 | [D-143/CE] (ready) | + +Rides (no separate build): #6 -> HN-C3's fixture library; #8 MAC re-measure -> HN-C1's MAC +invariant; #10 geneve/jumbo -> `geneve-encap-assert` verbatim, new invocation point only +(SC-19). `tests/opentofu-validate` T8-T10, `node-vm` T1-T7/T12-T15, `site-headend-install` +Sections 1-5/7, `maas-node-power` (opaque pass-through arg), `preflight` pending-change +fixture, `geneve-encap-assert` (all cases), `site-baseleg`, `cloudinit-vm`, +`d124-transit-seed`, `netem-link` (declared grep false-positive), and the 90 no-hit +harnesses are CONFIRMED STAY -- not itemized. + +### 1.9 Security-ledger (SEC) rows + +| ID | category | artifact | change | depends-on | owed-artifact# | axis | +|---|---|---|---|---|---|---| +| SEC-01 | SEC | NEW SEC-NNN row for the (a) cross-DC host isolation control (concern i); next-free confirmed **SEC-034** as of 2026-08-09, re-grep at mint time | new | DEC-14 | 2 | [CE] | +| SEC-02 | SEC | NEW SEC-NNN row for the power-key mitigation (concern iii) -- **CRITICAL PATH** | new | DEC-15 | 11 | [CE] | +| SEC-03 | SEC | SEC-010 disposition: new row vs amendment (concern ii) | new/change | DEC-16 | 3 | [CE] | +| SEC-04 | SEC | `vm-secret-locations` register rows (SEC-026/SEC-028/SEC-029) re-point rack-class -> client-class; rotation triggers ("if the rack is rebuilt") FIRE on this change | change | DEC-13, SC-12 | -- | [both] | + +--- + +## 2. Critical-path dependency chains + +### Chain A -- Power-key blast-radius mitigation (owed #11; the single largest blocker) +`DEC-15` (mechanism choice: restricted key / wrapper / polkit ACL) -> `SEC-02` (mint the +SEC row) -> `SC-11` (build the #11 artifact) -> `HN-C4` (harness -- itself a precondition, +not just coverage) -> `LB-01` (`lib-hosts.sh` power-address re-derivation) + `SC-07` +(`maas-region-power-key.sh` URI/key shape) -> `SC-01` (`maas-node-power.sh` call-site +literals) -> `HN-A4` + `HN-A5` (`dc-selector` / `maas-region-power-key` assertions, +**edited together, same session** -- H1 hazard: a plausible-looking URI swapped in before +the mechanism exists produces a false-green harness) -> `HN-A9` (`pre-flight-checks` P4 +case) + `GT-04` (cloud-assert A11b) -> `GT-06` (P5 register row for the new key material). +Six downstream test/tool edits are frozen until `DEC-15` rules (pass3 Section 4); the +interim RED on `HN-A4`/`HN-A5` once `lib-hosts.sh` changes is the DESIRED fail-loud state, +never something to "fix" early. + +### Chain B -- The (a)-control-before-any-flat-apply invariant (owed #2) +`DEC-14` (mechanism) -> `SEC-01` (SEC row) -> `SC-10` (#2 artifact) -> `HN-C2` (harness) -> +`GT-01` (Stage-1 gate, installed + `--check`-verified) -> **MUST PRECEDE** -> `TF-13` +(first flat substrate apply of EITHER per-DC root, under whatever `DEC-11` root shape +lands -- fork-robust: a merged single root's FIRST apply can create both DCs' planes at +once, so "before the second DC's apply" is not sufficient, only "before ANY flat apply" +survives the fork) -> `GT-02` (A11a re-verify at each apply's close) -> `GT-08` (folds into +the G9/G10 successor gate) -> re-verified again at Stage-5 live traffic (the first point +the claim is actually tested). + +### Chain C -- R7 + MAAS-release before destroy (Part A of the teardown sequence) +`SC-12` (#4 R7 credential-revocation checklist, enumerated from every `vm-secret-locations` +row keyed to the rack host class) + `HN-C6` (harness) -> **run BEFORE any substrate +destroy** (revoking after the hosts are gone degrades to "assume it's moot") -> `SC-13` +(#5 MAAS machine-record release/delete + rack-controller decommission + region-side +`primary_rack`/DHCP-reference cleanup) + `HN-C7` (harness) -> `TF-01`/`TF-02` destroy +applies (inner roots first from voffice1, then outer from vcloud) -> `RB-01` (teardown +runbook rewrite encodes this exact order). This chain governs the CURRENT 10.12 checkpoint +teardown and is largely independent of `DEC-11`'s root-shape ruling for the NEW build. + +### Cross-cutting: the root-topology fork (`DEC-11`) +Gates `TF-02`, `TF-12`, `TF-13`, `SC-09`, `SC-14`, `HN-C3`, `RB-01`, `GT-05`, `HN-A1`, +`HN-A2`, `DEC-12`, `DEC-23` -- the sequence itself is invariant to the fork (per pass1 +check 2), but the teardown primitive's exact wording, the state blast radius, and every +root-naming literal are NOT. Ratify `DEC-11` early; it unblocks the largest single cluster +of "ready once ratified" rows. + +--- + +## 3. Count summary by category + +| Category | Rows | +|---|---| +| decision (DEC) | 23 | +| tofu-module (TF) | 13 | +| lib (LB) | 3 | +| script (SC) | 19 | +| doc (DC) | 6 | +| runbook (RB) | 5 | +| gate (GT) | 10 | +| harness (HN) | 21 (9 change + 5 retire + 7 new-build) | +| SEC | 4 | +| **Total** | **104** | + +Cross-check against source counts: harness total (21) matches pass3 Section 2's +9-existing-change + 5-existing-retire + 7-new-build decomposition exactly; owed-artifact +references (13 distinct #-tags) all appear at least once across TF/SC/GT/HN/SEC rows, with +no double-counting (rack decommission folds into SC-13/#5; artifact-service sizing rides +SC-16/#7; the SEC-010-writer extraction IS SC-04/#3's implementation shape -- all per +pass2 Section 5's explicit "not double-counted" note, carried forward here). + +--- + +## 4. Verification note + +Author = W4.1 (no model name asserted). This document is a MERGE of `pass1-admin-report.md` +Sections 2-6, `pass2-admin-report.md` Sections 3-6, and `pass3-admin-report.md` Sections +2-5 -- no new repo reads were performed beyond the four admin reports and their stated +verification notes; every row's artifact path/line traces to a citation already verified +in one of those four reports (see each report's own Section verifying "Verification note" +/ "Adversarial-check results" for the underlying grep/read evidence). READ-ONLY; nothing +executed; findings LOGGED only. diff --git a/docs/audit/container-elim-pass/pass4-w2-module-workflow-design.md b/docs/audit/container-elim-pass/pass4-w2-module-workflow-design.md new file mode 100644 index 0000000..e26b700 --- /dev/null +++ b/docs/audit/container-elim-pass/pass4-w2-module-workflow-design.md @@ -0,0 +1,480 @@ +# Pass 4 / W4.2 -- THE LAYERED MODULE-WORKFLOW DESIGN (container-layer elimination) + +**Worker:** W4.2 (Phase 4, container-layer-elimination pass, `SCOPE-AND-EXECUTION-PLAN.md` +Section 4 -- "the pass's headline deliverable"). **Date:** 2026-08-09. **Scope:** READ-ONLY +planning. Produces the concrete design of the layered module system the redeploy becomes; +performs no mutation. + +**Baseline consumed in full, this session:** `SCOPE-AND-EXECUTION-PLAN.md`, +`pass0-admin-report.md` (Option 1 CONFIRMED, Section 7a), `pass1-admin-report.md` + +`pass1-w4-module-planning.md` (the L0-L5 layer model, adopted here as the backbone), +`pass2-admin-report.md` + `pass2-w1-tofu-modules.md` + `pass2-w4-module-decomposition.md` +(the module change-set, the `dc-site` IaC module, the IaC<->procedure boundary, the 13 owed +artifacts), `pass3-admin-report.md` (the unified tests change-set, the gate-home map, the +per-module harness contract). Every module/artifact below is grounded in one of these +documents or the live repo; anything not yet built is marked **[NEW]** and cites which pass +scoped it. No inferred values -- unresolved design points (root naming, the (a)/(iii) +mechanisms, rack-controller placement) are carried as OPEN, matching their source reports. + +--- + +## 0. What this document is (and is not) + +This is the CONCRETE module design that `pass1-w4-module-planning.md`'s L0-L5 model +scaffolds and `pass2`/`pass3`'s change-sets populate. It answers three questions: + +1. **Per layer, what are the modules, and what is each one's contract?** (Section 2) +2. **How does a full per-DC redeploy compose these modules into one ordered + invocation, parameterized by site token?** (Section 3) +3. **What design principles make this repeatable and Roosevelt-transferable, and + which layers/modules actually carry over to the pre-Roosevelt bare-metal test + unchanged?** (Sections 4-5) + +It does NOT re-rule anything already settled (Option 1, handling (a), MAAS-region +placement) or re-open anything explicitly OPEN in the source reports (root naming, the +(a)/(iii) mechanisms, rack-controller-remainder placement, root topology's final +ratification). Where a module's shape depends on an OPEN item, this document states the +dependency and defers -- it does not invent a value (hard rule 2). + +--- + +## 1. The layer model, restated as the design's backbone + +Adopted verbatim from `pass1-w4-module-planning.md` Section 2, consolidated at +`pass1-admin-report.md` Section 5: + +``` +L0 Host & inter-site substrate (IaC) -- Stage 1 +L1 Site/edge nodes (IaC) -- Stage 2 (Office1) / Stage 3 (DC, flattened in) +L2 DC substrate: planes + node VMs (IaC) -- Stage 3 +L3 Enlist/commission (procedure) -- Stage 4 +L4 Juju/OpenStack deploy (procedure) -- Stage 5 (+6 DR, +7 Designate/COS/Magnum, additive) +L5 Verify/gate (cross-cutting procedure, re-invoked at every boundary) +``` + +**Layer-boundary rule** (the design's one hard invariant, `pass1-w4-module-planning.md` +Section 2, restated by `pass2-w4-module-decomposition.md` Section 4 as "identity, not +orchestration"): each layer's input is the layer directly below's OUTPUT only. OpenTofu +(L0-L2) owns up to *"a booted libvirt domain exists, with its network identity (MAC per +NIC, and any statically-assigned IP) correctly wired to the right plane bridges."* The +procedure layer (L3+) begins at the first live dial into that object and **never reads +OpenTofu state** -- it re-derives everything from live, independently observable identity +(MAC, hostname, a fresh API/SSH probe), per `lib-hosts.sh:6-11`'s own stated design rule. + +**The container layer was the one place this rule was violated**, and it is the reason +this pass exists: the inner root (`opentofu/vr1-dc0-substrate/main.tf`) dialed OUT to a +`qemu+ssh` provider INTO the outer root's own `vvr1_dc0` output -- IaC reaching into IaC +across a live-dial boundary that only a procedure module should cross +(`opentofu/main.tf:28-29`; `pass2-w4-module-decomposition.md` Section 4). Option 1 removes +this violation structurally: with one flat per-DC root, L2 is IaC end to end, and the +FIRST live dial into anything L2 produced is L3's MAAS commissioning -- exactly where the +boundary rule says it should be. Every module table below is built so this stays true. + +--- + +## 2. Per-layer module design + +Legend: **kind** = IaC module / procedure module / library / gate. **[EXISTS]** = re-homed +or unchanged artifact, cited to its current path. **[NEW]** = owed artifact, cited to the +pass that scoped it (`pass1` #N / `pass2` #N referring to the numbered owed-artifact lists +in `pass1-admin-report.md` Section 6 / `pass2-admin-report.md` Section 5). + +### L0 -- Host & inter-site substrate (IaC, Stage 1) + +**Contract:** given vcloud's `qemu:///system` connection and the site tokens +(`office1`/`vr1-dc0`/`vr1-dc1`), produce the cross-site fabric every other layer plugs +into: the mesh triangle, per-site storage pools, the Office1 L2 network, and the +inter-DC netem link. No DC-specific node/plane content lives here. + +| Module | Kind | Inputs | Outputs | Current/planned artifact | Harness | +|---|---|---|---|---|---| +| Mesh triangle | IaC | 3 site pairs, MTU 9000 | 3 mesh-link bridges (dc0<->dc1, dc0<->office1, dc1<->office1) | `modules/mesh-link`, outer `main.tf:125-141` **[EXISTS, unchanged]** -- confirmed non-consumer of the container layer (`pass2-w1-tofu-modules.md` row `mesh-link`) | `tests/opentofu-validate/` (module-standalone validate) | +| Netem link | IaC | dc0<->dc1 mesh bridge (`virbr5`) | tc-shaped DR-drill link | `modules/netem-link`, `main.tf:353-358` **[EXISTS, unchanged]** -- targets the bridge directly, no dependency on either DC's internal shape | `tests/opentofu-validate/` | +| Per-site storage pools | IaC | host path per site | `office1_storage` pool (DC pools re-home to L2, Section below) | `modules/dc-storage-pool`, `office1_storage` call **[EXISTS, unchanged]** | `tests/opentofu-validate/` | +| Office1 L2 network | IaC | -- | `office1-network` | `modules/office1-network`, `main.tf:76-80` **[EXISTS, unchanged, D-114 territory, out of scope]** | `tests/opentofu-validate/` | +| Shared base image | IaC | source image | `ubuntu_noble_base`, consumed by every `cloudinit-vm` instance | `modules/base-image`, `main.tf:165-173` **[EXISTS, unchanged, gains client-VM consumers]** | `tests/opentofu-validate/` | +| Cross-DC host isolation control **[NEW]** | Gate (procedure, L5 kind, L0-scoped invocation) | live bridge enumeration on vcloud, both DCs' plane tags | PASS/REFUSE; nftables rules asserting no inter-plane/inter-DC forwarding | `pass1` #2 / `pass2` Sec 2.1 -- concern (i); spec: Section 6 below | **[NEW]** `tests//run-tests.sh`, offline fixture model (`pass3` C2) | + +**Composes onto:** nothing below it -- L0 is the floor. **Stage owner:** Stage 1 +(`docs/dc-dc-deployment-workflow.md`, gate content unchanged per `pass1-admin-report.md` +Section 2.2). The (a) control's recommended home is also Stage 1 -- host-scoped, not +per-DC-apply-scoped (`pass1-admin-report.md` Section 3). + +### L1 -- Site/edge nodes (IaC, Stage 2 Office1 / Stage 3 DC-flattened) + +**Contract:** given an L0 network + pool, produce a booted, network-identified +non-hypervisor VM: Office1's headend, each DC's edge, and (Option 1's new instance) +each DC's client VM. All three are the SAME module type -- `cloudinit-vm` -- differing +only by inputs. + +| Module | Kind | Inputs | Outputs | Current/planned artifact | Harness | +|---|---|---|---|---|---| +| Office1 headend (`voffice1`) | IaC | L0 office1-network + pool | booted VM, MAAS-composed LXD host | `modules/cloudinit-vm`, `main.tf:175` **[EXISTS, unchanged -- D-114, out of scope]** | `tests/opentofu-validate/`, `tests/node-vm/` (adjacent) | +| DC edge (OPNsense) | IaC | L0 pool; L2 planes (WAN leg); `uplink_network_name` string (an L0 `site-wan` output) | booted edge VM, WAN attached directly to the per-DC NAT (no `wan-bridge`) | `modules/opnsense-edge`, re-homed into `dc-site` (Section below) **[EXISTS body, re-homed]** -- input `wan_network_name` rewires from `module.vr1_dc0_wan.network_name` (dead) to `module.vr1_dc0_uplink.network_name` (`pass2-w1-tofu-modules.md` row `opnsense-edge`) | `tests/opentofu-validate/` | +| **Client VM** (`vr1-dcN-client`) **[NEW instance, existing module]** | IaC | L0 pool + mesh-triangle transit leg; NetBox-assigned `metal_admin_ip`/`transit_ip`; `expose_nested_virt=false`; ~4 vCPU/8192 MiB/80 GiB | booted, non-hypervisor VM carrying the D-138 client role + SEC-028/029 credential residencies + (open) rack/D-131/mirror placement | `modules/cloudinit-vm`, NEW call inside `dc-site` (`pass0-admin-report.md` Option 1; `pass1-w4-module-planning.md` Section 3; `pass2-w1-tofu-modules.md` row `cloudinit-vm`) -- **duty roster SHRUNK to client role + credentials + transit leg** once rack-controller retirement is adopted (`pass2-admin-report.md` Section 4.2, 4.5) | `tests/opentofu-validate/` T14/T15 (autostart -- **new case owed, D-127 value not yet ruled**) | +| Site-WAN NAT (L0/L1 boundary) | IaC | -- | per-DC egress NAT, now the edge's DIRECT upstream | `modules/site-wan`, `main.tf:379-396` **[EXISTS, unchanged, gains a direct consumer]** | `tests/opentofu-validate/` | +| SEC-010 transit-leg successor **[NEW]** | Procedure, installed on L1 objects | the client VM's + voffice1's re-measured transit interface names | nftables FORWARD-drop on both transit endpoints | `pass1` #3 / `pass2` Sec 2.2 -- concern (ii); ONE role-agnostic installer subcommand, extracted from `site-headend-install.sh`'s `node_host_setup()` (`:273-320`), replacing today's hand-mirrored voffice1 install | `pass3` A3 -- extends `tests/site-headend-install/` (migrates every proven SEC-010 assertion) | + +**Composes onto:** L0's network + pool outputs only. **Stage owner:** Stage 2 (Office1, +untouched) / Stage 3 (DC edge + client VM, now co-located with L2's per-DC apply -- +Section 2.4 of `pass2-w1-tofu-modules.md`: teardown symmetry + SEC-026 state isolation + +D-138 fidelity all argue for grouping the client VM's APPLY with its DC's flat root even +though its module TYPE is L1). + +### L2 -- DC substrate: planes + node VMs (IaC, Stage 3) + +**Contract:** given the L0 pool + L1 client VM's transit identity, and a site token + +D-121/R-3 node roster + D-134 octet map, produce the DC's six plane networks and its 12 +node-VM libvirt domains (9 role nodes + 3 utility nodes), each with a MAC-pinned network +identity MAAS can discover. **This is the layer the container layer's boundary +violation lived in and Option 1 collapses it to a single root/state.** + +| Module | Kind | Inputs | Outputs | Current/planned artifact | Harness | +|---|---|---|---|---|---| +| `dc-planes` | IaC | `planes` map (name->CIDR), `mtu`, `domain_suffix` | 6 isolated plane bridges, MTU 9000 (D-101/D-139 own the CIDR values) | `modules/dc-planes` **[EXISTS body, re-homed]** from inner root (qemu+ssh) to the per-DC-flat root (`pass2-w1-tofu-modules.md` row `dc-planes`) | `tests/opentofu-validate/` | +| `dc-storage-pool` (per-DC) | IaC | host path | one per-DC pool (collapses the outer "containment-VM-disk" pool + the inner "node-disk" pool into ONE) | `modules/dc-storage-pool` **[EXISTS body, COLLAPSES 2-per-DC -> 1-per-DC]** (`pass2-w1-tofu-modules.md` row `dc-storage-pool`) | `tests/opentofu-validate/` | +| `node-vm` (x12/DC) | IaC | `nodes` map (vcpu/mem/disk/osd_gib?/macs), attaches to `dc-planes` outputs | 12 MAC-pinned libvirt domains: 9 D-121 role nodes + `vr1-dcN-juju-01` (.5) + `vr1-dcN-maas-01` (.6) + `vr1-dcN-tailscale-01` (.7) | `modules/node-vm` **[EXISTS body, re-homed]**, `for_each` map moves verbatim from the inner root (`pass2-w1-tofu-modules.md` row `node-vm`) | `tests/node-vm/` T1-T15 (T8-T11 re-point the hardcoded `INNER=` path; T1-T7/T12-T15 STAY) | +| **`modules/dc-site`** **[NEW, highest-leverage new artifact]** | IaC (composing module) | `site_token`, `domain_suffix`, `underlay_mtu`, `pool_path`, `planes`, `opnsense_base_path`, `uplink_network_name` (string), `nodes` (12-entry roster), `client_vm` (object: sizing, macs, `transit_network_name`, `metal_admin_ip`/`transit_ip`) | plane `network_names`, node `domain_names`/`ids`, client-VM `domain_id`, edge `domain_id` | Replaces the ~230-266-line copy-pasted per-DC inner-root bodies with ONE module both per-DC-flat roots call once; composes `dc-storage-pool` -> `dc-planes` -> `opnsense-edge` -> `node-vm` `for_each` -> `cloudinit-vm` (client), in that order (`pass2-w1-tofu-modules.md` Section 3.1) | **[NEW]** `tests/dc-site/run-tests.sh`, static-fixture `.tf`-tree model (`pass3` C1: 5-plane tree FAILS count; 11/13-node roster FAILS; 2 client-VM calls FAILS exactly-one; `mtu=1500` FAILS; hardcoded DC literal FAILS site-token check) | +| Per-DC-flat root (`vr1-dcN-flat/`, name OPEN) | IaC (root, not a module) | one `provider "libvirt" { }` block pointed at vcloud's own `qemu:///system` (no keyfile/sshauth/qemu+ssh trap class); one `module "site" { source = "../modules/dc-site" }` call; one tfvars file | the DC's entire flat substrate, one apply cycle, one state file | **[NEW root]**, direct successor to today's inner root (`pass2-w1-tofu-modules.md` Section 3.2) -- **root topology (B) shared-outer + per-DC-flat RECOMMENDED, Phase-4 ratifies; naming OPEN (avoid unqualified `-substrate` if reserved for Roosevelt)** | `tests/opentofu-validate/` (extend P8-equivalent to a DECLARED root list, not a glob -- `pass3-admin-report.md` Section 3 "H6") | +| `wan-bridge` | IaC | -- | -- | `modules/wan-bridge` **[COLLAPSES/DELETED]** -- D-125 bridge-in dead; edge WAN -> direct NAT (`pass2-w1-tofu-modules.md` row `wan-bridge`) | retired with its parent | + +**Composes onto:** L0's pool (via `pool_path`) and L1's client-VM transit identity (via +`client_vm.transit_network_name`, a cross-root STRING reference -- the same pattern +`office1_opnsense`'s `wan_network_name` literal already uses, `main.tf:112`). **Does NOT +reach into L3** -- `dc-site`'s outputs are network_names/domain_ids/MACs only; it never +calls MAAS or juju. **Stage owner:** Stage 3, the HEAVIEST rewrite +(`pass1-admin-report.md` Section 2.2): Step B (bootstrap gate) eliminated wholesale, Step +C (inner apply) merges into Step A -- one apply, one root-scope, one host, one state. + +### L3 -- Enlist/commission (procedure, Stage 4) + +**Contract:** given L2's MAC-pinned node VMs and a reachable MAAS region (L1's +`vr1-dcN-maas-01`, D-132 addendum, unchanged by flattening), produce READY, carved, +tagged MAAS machines and the per-DC artifact-mirror/DNS-forwarder services those +machines need during commissioning. **The handoff artifact from L2 is a MAC address, +observed independently on both sides -- no tofu state crosses this boundary** +(`pass2-w4-module-decomposition.md` Section 4). + +| Module | Kind | Inputs | Outputs | Current/planned artifact | Harness | +|---|---|---|---|---|---| +| `dc-node-carve.sh` | Procedure | ``, live MAAS API | v4 NIC/br-ex carve | **[EXISTS, no code change]** -- grep-verified zero containment hits, pure MAAS-API (`pass2-admin-report.md` Sec 3.3) | `tests/dc-node-carve/` | +| `dc-node-v6-carve.py` | Procedure | ``, live MAAS API, MAAS tag `openstack-` | v6 static assignment | **[EXISTS, no code change]** | `tests/dc-node-v6-carve/` | +| `maas-node-power.sh` | Procedure | power address as `$1` (topology-agnostic) | `power_type=virsh` set, MAC-matched | **[EXISTS, no code change to the script body]** -- every invocation-site/runbook LITERAL updates, **BLOCKED on the power-key mitigation (#11) mechanism** (Section 6) | `pass3` A4/A5 -- **BLOCKED on #11**, edited together, same session | +| `maas-role-tags.sh` | Procedure | `` | per-role MAAS tags the bundle constrains on | **[EXISTS, no code change]** | `tests/maas-role-tags/` | +| `maas-region-power-key.sh` | Procedure | region host | installs/verifies the per-DC MAAS->libvirt power key (SEC-012 dc0 / SEC-016 dc1) | **[EXISTS body unchanged]** -- the KEY it installs IS the power-key blast-radius object; URI/key shape re-derives per the #11 mitigation | `pass3` A5 -- **BLOCKED on #11** | +| `dc-region-topology.sh` | Procedure (+ gate mode) | `` | region fabric/space/subnet/tag topology | **[EXISTS, no change]** | `tests/dc-region-topology/` | +| `dc-plane-ipam.sh` | Procedure (+ gate mode) | `` | plane IPAM state, D-134's executable gate | **[EXISTS, no change]** | `tests/dc-plane-ipam/` | +| `dc-mirror.sh` / `dc-cache-proxy.sh` / `dc-snap-proxy.sh` | Procedure | target host, `` | per-DC apt/UCA mirror or cache, snap proxy | **[EXISTS, host RE-TARGET only]** -- neither Option-1 VM fits dc0's mirror as authored (maas-01: 150 GiB earmarked, no spare; client VM: ~80 GiB) -- **sizing decision rides the FIT-calculator extension (#7)** | rides `tests/dc-mirror/` etc., extended sizing case | +| `site-headend-install.sh --role rack` remainder | Procedure | target host | MAAS `--role rack` enrollment (contingent) | **[EXISTS, CONTINGENT RETIRE]** -- `--host-nodes` (~134 lines, the bootstrap-gate duty) DELETES wholesale; `--role rack` retires IF rack-controller retirement is ratified (both maas-01 VMs already run `region+rack`; measured, `pass2-admin-report.md` Sec 1.2/Sec 4.2) | `pass3` B2 (retire), B3 (contingent) | +| `dc-rack-net.sh` (D-131 forwarder half) | Procedure | rack host | node-DNS forwarder | **[EXISTS, LEGS half retires structurally (no flat VM is a libvirt host with own bridges); forwarder half CONTINGENT on D-131 retire-with-evidence]** | `pass3` B4 (retire), B5 (contingent) | +| Cross-DC isolation control's `--check` re-verify | Gate | live bridge state, post-apply | PASS/REFUSE | Re-invocation of the L0-homed (a) control at each per-DC apply's close (`pass1-admin-report.md` Section 3) | shares C2's harness | +| MAAS machine-record release/delete + rack-controller decommission **[NEW, teardown-adjacent]** | Procedure | MAAS API, region-side residue | zero-record confirmation | `pass1` #5, AMENDED by `pass2` Sec 4.2(i) to include `maas rack-controller delete` + `primary_rack`/DHCP-reference cleanup | **[NEW]** `pass3` C7, fakebin `maas` model | +| D-131 retirement-evidence step **[NEW, one-time]** | Procedure (evidence capture, not a standing gate) | fresh region's own BIND | dig-proven resolver identity | `pass2` #13 | **[NEW]** `pass3` C5, offline dig-capture fixtures; **dc1's current config is a standing-RED case until retirement is live** | +| Power-key blast-radius mitigation **[NEW]** | Procedure (credential-scope control) | vcloud's libvirtd connection scoping mechanism (candidate: restricted SSH key / per-DC virsh wrapper / polkit ACL) | a power key that can control ONLY its own DC's domain set | `pass2` Sec 2.3, concern (iii) -- **BLOCKS `maas-node-power.sh`/`maas-region-power-key.sh` literal updates and the `lib-hosts.sh` power-address re-derivation** | **[NEW]** `pass3` C4, offline rendered-ACL/`authorized_keys`/polkit parsing model; itself the precondition for A4/A5/A9 | + +**Composes onto:** L2's MAC-pinned domains + L1's reachable MAAS region only; never +dials L0 or reads tofu state. **Stage owner:** Stage 4 (`dc-dc-phase3-maas-enlist- +deploy.md`, LOW DELTA -- only the two SSH-jump-target literals change, per +`pass1-admin-report.md` Section 2.1). + +### L4 -- Juju/OpenStack deploy (procedure, Stage 5 + additive Stage 6/7) + +**Contract:** given L3's READY machines, a site token, `bundle.yaml` + per-DC overlays, +and a Vault root, stand up an independently-running OpenStack cloud per DC. D-140 PINS +this as a procedure layer for THIS redeploy -- settled, not folded into IaC +(`pass1-admin-report.md` Section 5). + +| Module | Kind | Inputs | Outputs | Current/planned artifact | Harness | +|---|---|---|---|---|---| +| Juju bootstrap + `bundle.yaml` deploy | Procedure | L3's READY machines, `$DC` | running controller + bundle | `runbooks/dc-dc-phase4-juju-bundle-per-dc.md` **[EXISTS, MODERATE change]** -- RUN-LOCATION table row 1 (juju/openstack CLI) re-points `rack` -> `vr1-dcN-client`; row 3 ("never vcloud") doctrine intact | -- (runbook-level, gated by `preflight.sh` below) | +| `render-dc-overlays.py` | Procedure | `` | deterministic per-DC bundle-overlay files | **[EXISTS, no change -- derive/render split already site-parameterized]** | `tests/render-dc-overlays/` | +| VR0 template chain (`phase-00`..`phase-07-*.sh`) | Procedure | `` | admin creds, network stand-up, Octavia amphora pipeline, Magnum/CAPI stack, Ceph rbd-mirror/radosgw-multisite | **[EXISTS, no change]** -- run twice, once per DC, per `pass1-w4-module-planning.md` Section 1.2 "the VR0 template it runs twice" | Each ships its own `tests/phase-0N-*/` dir (`pass2-w4-module-decomposition.md` Section 2) | +| Preflight gate | Gate | live cloud state | PASS/FAIL, P1-P10 | `scripts/preflight.sh` **[EXISTS, EXTENDED -- new P10 for concern (ii); P4 gains an #11-dependent content case]** | `tests/preflight/` | +| Credential minting (SEC-026/028/029) | Procedure (data, not a module body) | the client VM's identity | freshly-minted per-DC credentials | Register rows re-point (`vm-secret-locations`); no material migrates -- freshly minted on the new host (`pass1-admin-report.md` Section 1 check 7) | `pass3` C6 (R7 revocation checklist, rides teardown) | + +**Composes onto:** L3's READY-machine handoff only. **Stage owner:** Stage 5 (primary), +Stage 6 (DR drill, additive per `vr0-to-vr1-is-additive`), Stage 7 (Designate/COS/Magnum, +additive) -- Stages 6-7 have NO container-layer dependency, verified by W1.2 +(`pass1-admin-report.md` Section 2.2). + +### L5 -- Verify/gate (cross-cutting procedure, invoked at every layer boundary) + +**Contract:** at every layer's declared-done state, produce a failable PASS/FAIL (never +existence-only, GA-R6) plus, at milestones, a committed BOM. L5 is not owned by one +stage; it re-runs at each stage's close. + +| Gate | Scope | Home / re-invocation points | Current/planned artifact | Harness | +|---|---|---|---|---| +| `opentofu-validate.sh` | L0-L2, every module standalone + every root | Delivery-time + Stage 1-3 close | **[EXISTS, extended]** -- root enumeration moves to a DECLARED list (not a glob, `pass3` "H6"); T13 (containment autostart) RETIRES with a negative-assertion case | `tests/opentofu-validate/` | +| `preflight.sh` | L3->L4 boundary, P1-P10 | Before every L4 deploy step | **[EXISTS, extended]** -- new P10 (concern ii), P4 content case (concern iii, `#11`-blocked) | `tests/preflight/` | +| `cloud-assert.sh` (+ `--capture`) | L4, behavioral | Post-deploy/restart/pre-change/post-incident (periodic) | **[EXISTS, extended]** -- new A11a (concern i re-verify) and A11b (concern iii standing re-verify) | `tests/cloud-assert/` | +| `geneve-encap-assert.sh` | L2/L4, OVN geneve family/tunnel health | Post-build on the vcloud-level planes (live assert OWED -- `pass2` #10) | **[EXISTS, unchanged body, NEW invocation point only]** -- MTU/geneve budget analytically unaffected by flattening (`pass0-admin-report.md` Section 1.4) | `tests/geneve-encap-assert/` | +| `dc-egress-check.sh` | L3, DC-egress probe | Invocation-host literal re-points to the ruled B.5 host | **[EXISTS, host literal only]** | `tests/dc-egress-check/` | +| `dc-node-v6-verify.sh` | L2/L3, v6 statics + forwarding | unchanged | **[EXISTS, no change]** | `tests/dc-node-v6-verify/` | +| Budget calculators (`dc-dc-mtu-geneve-budget.sh`, `dc-dc-ceph-disk-budget.sh`, `dc-dc-whole-host-budget.py`) | L0/L2 | Delivery-time, before any FIT verdict enters the change-set | **[EXISTS, the whole-host one EXTENDED]** -- `pass2` #7: add the 3 utility-node classes + artifact-service disk-sizing branch; the Model-A/B comparison cases must not remain the SOLE coverage against a retired topology (`pass3-admin-report.md` "H4") | `tests/dc-dc-whole-host-budget/` | +| `maas-profile-assert.sh` | L3 | Region-profile resolution proof | **[EXISTS, fixture change]** -- drop the two `vvr1-dcN` rows from the simulated roster (rack-retirement-contingent) | `tests/maas-profile-assert/` | +| Cross-DC host isolation control **[NEW, concern i]** | L0-scoped, cross-DC | Stage 1 install; re-verify at each per-DC apply's close; re-verify at Stage-5 live verify (A11a) | `pass1` #2 -- Section 6 below | **[NEW]** `pass3` C2 | +| SEC-010 transit-drop successor **[NEW, concern ii]** | L1-scoped, both transit endpoints | preflight P10 | `pass1` #3 (amended) -- Section 2 L1 table | extends `tests/site-headend-install/` | +| Power-key mitigation gate **[NEW, concern iii]** | L3/L4-scoped, credential blast radius | preflight P4 dependency; cloud-assert A11b standing | `pass2` #11 | **[NEW]** `pass3` C4 | + +**Composes onto:** any layer's declared-done state; L5 gates never mutate by default +(gate = "its sole job is verify," `pass2-w4-module-decomposition.md` Section 2 legend). + +--- + +## 3. THE COMPOSITION MECHANISM -- one ordered invocation, parameterized by site token + +A full per-DC redeploy is the following ordered sequence. `$SITE` is the site token +(`vr1-dc0` / `vr1-dc1`, D-119); steps marked **(once)** run before the per-DC loop and are +NOT re-run per DC; steps marked **(per $SITE)** run once for each DC, in either order +(root shape (B) makes them structurally independent -- Section 3.1 below). Axis tags +`[D-143]`/`[CE]`/`[both]` carried from `pass1-admin-report.md` Section 4 so the two +riding changes (the 10.13 re-IP and the container-elim) stay distinguishable in this +same sequence. + +``` +STAGE 1 (once, L0 + L5) + 1.1 [unchanged] tofu apply: outer/shared root -- mesh triangle, per-site pools, + office1-network, netem link, base-image (L0) + 1.2 [CE, NEW] install + --check the cross-DC host isolation control (L5) + -- MUST precede 1.4/2.x for EITHER DC (fork-robust invariant, + `pass1-admin-report.md` Section 3) + +STAGE 2 (once, L1) + 2.1 [unchanged] Office1 headend (voffice1) standup -- D-114, untouched (L1) + + --- per-$SITE loop begins; (B) root shape makes each iteration state-isolated --- + +STAGE 3 (per $SITE, L1 + L2) + 3.1 [both] tofu apply: vr1-$SITE-flat root -> module "site" = dc-site( + site_token=$SITE, planes=..., nodes=..., client_vm=...) + composes, IN ORDER: dc-storage-pool -> dc-planes -> opnsense-edge + -> node-vm[12] -> cloudinit-vm(client) (L1+L2) + 3.2 [CE] re-verify the (a) control's --check now that this DC's planes + exist (post-apply close) (L5) + 3.3 [NEW] MAC re-measurement pass (owed #8) -- confirm every domain's MAC + before Stage 4 trusts one (L2/L3 handoff) + +STAGE 3.5 (per $SITE, L1-hosted procedure installs) + 3.5.1 [CE] SEC-010 successor installer, BOTH ends (client VM + voffice1) (L1) + 3.5.2 [CE, cont.] site-headend-install.sh --role rack (IF rack-retirement is + NOT ratified; else SKIPPED) (L1/L3) + 3.5.3 [both] maas-region-power-key.sh -- key shape per the #11 mitigation (L3) + +STAGE 4 (per $SITE, L3) + 4.1 [D-143] dc-region-topology.sh --commit (L3) + 4.2 [D-143] dc-plane-ipam.sh --commit (L3) + 4.3 [unchanged*] dc-node-carve.sh / dc-node-v6-carve.py -- per-machine + power_type=virsh; *only the power-ADDRESS value re-derives (L3) + 4.4 [both] maas-node-power.sh -- value BLOCKED on #11 (L3) + 4.5 [unchanged] maas-role-tags.sh (L3) + 4.6 [CE, cont.] dc-rack-net.sh (D-131 forwarder) -- IF retire-with-evidence + is REJECTED for this DC; else the D-131 evidence step (4.7) (L3) + 4.7 [CE, cont.] D-131 retirement-evidence step -- IF retirement adopted (L3) + 4.8 [CE] dc-mirror.sh / dc-cache-proxy.sh / dc-snap-proxy.sh, sized + per the FIT-calculator extension (#7) (L3) + +STAGE 5 (per $SITE, L4 + L5) + 5.1 [gate] preflight.sh (P1-P10) (L5) + 5.2 [unchanged] juju bootstrap + bundle.yaml + overlays, FROM vr1-$SITE-client (L4) + 5.3 [both] SEC-026/028/029 credentials freshly minted on the client VM (L4) + 5.4 [gate] cloud-assert.sh (incl. A11a/A11b) + geneve-encap-assert.sh (L5) + + --- per-$SITE loop ends --- + +STAGE 6 (once, cross-DC, L4, additive) + 6.1 [unchanged] DR/failover drill: netem-link mechanism + Ceph rbd-mirror / + radosgw-multisite (D-108) (L4) + +STAGE 7 (once per $SITE, L4, additive) + 7.1 [unchanged] Designate / COS / Magnum (D-106/D-105, `vr0-to-vr1-is-additive`) (L4) + +CLOSE-OUT (per $SITE, L5 + record-keeping) + C.1 [CE] NetBox DCIM: decommission vvr1-$SITE, register client VM+roster (record) + C.2 [D-143/CE] the container-elim [ARCH] ruling itself is NOT a gate -- it is a + precondition the operator must have already ruled before Stage 3 + runs for the FIRST time (Section 7) +``` + +### 3.1 Why the loop is safe to run per-DC in either order (root shape (B)) + +Under the recommended root shape -- shared-outer + per-DC-flat roots +(`pass2-w1-tofu-modules.md` Section 2, RECOMMENDED, Phase-4 ratifies) -- DC1's Stage-3 +apply cannot create or touch any DC0 resource: it is not in DC0's state file, full stop. +This makes the composition mechanism's ordering invariant reduce to a single, simple +contract: **"the (a) control's `--check` must pass before the first Stage-3 apply of +EITHER per-DC-flat root"** (Step 1.2), rather than depending on `-target` discipline +being followed correctly on every apply (the risk a merged single root would carry). +This is why Stage 1's step 1.2 is drawn OUTSIDE the per-`$SITE` loop and BEFORE it in the +sequence above -- it is a precondition for the loop, not a loop step. + +### 3.2 Teardown primitive's place in the same composition + +The teardown/site-down lever is the composition mechanism's INVERSE, not a separate +design: under (B), site-down for `$SITE` is `cd opentofu/vr1-$SITE-flat/ && tofu destroy` +(gated) or a scripted `virsh destroy` loop over that root's own state-listed domains +(emergency) -- touching only that DC's root, never the shared-outer root or the other +DC's root (`pass2-w1-tofu-modules.md` Section 2.3). This is owed artifact `pass1` #1 (+ +#6 the emergency lever, riding the same fixture library per `pass3-admin-report.md` +Section 2 rides list) and sits at the **L2/L5 boundary**: it is a procedure module that +WRAPS an IaC destroy, verified by an L5-style completeness check (a dc1 domain leaking +into a dc0-targeted set FAILS; an empty resolved set REFUSES rather than reporting +"nothing to do" success -- `pass3-admin-report.md` C3). + +--- + +## 4. Design principles that make this repeatable + Roosevelt-transferable + +Adopted from `pass1-w4-module-planning.md` Section 4, sharpened by `pass2`/`pass3`'s +findings. Each principle is grounded in an EXISTING repo pattern, not invented for this +pass (`pass1-w4-module-planning.md` Section 1's framing: "the container-elim does not +need to invent module mechanics, only re-home"). + +1. **Site-token parameterization, never hardcoded identity.** The `$DC`/`$SITE` + selector (D-119, DOCFIX-151, `lib_net_select_dc`/`lib_hosts_select_dc`) already does + this for every procedure module; the IaC layer already does it structurally (no + module BODY names a DC). Every new artifact this pass introduces takes the same + token as an input rather than being written DC0-specific and copy-pasted for DC1: + `dc-site`'s `site_token` input (Section 2, L2), the (a) control's per-bridge-tag + enumeration (Section 2, L0/L5), the power-key mitigation's per-DC domain-set scoping + (Section 2, L3). This is the exact anti-pattern + `docs/dc-dc-deployment-workflow.md`'s gap-register item 1 was created to close. + +2. **Every module ships its tested harness, no exception.** `tests//run-tests.sh` + for procedure modules, `opentofu-validate.sh`'s standalone-module coverage for IaC + (CLAUDE.md "Delivery"; `pass2-w4-module-decomposition.md` Section 1.3). Of the 34 + procedure modules surveyed at Phase 2, 32 already comply; the two pre-existing gaps + (`maas-fabric-prune.sh` / `maas_fabric_classify.py`) are NOT silently waved through -- + they are routed as a named Phase-4 decision (build vs. accept-as-named-exception, + `pass3-admin-report.md` Section 7 item 8), because this pass's own harness-discipline + principle would otherwise be inconsistent about them. Every one of the 7 net-new + artifacts in Section 2/6 ships its `tests//run-tests.sh` FROM THE SAME COMMIT + that ships the artifact, per the per-module harness contract adopted at + `pass3-admin-report.md` Section 6: prove-it-can-fail, assert the artifact not the + intent, offline/fixture-driven by default with a separately-named LIVE re-run, + `$SITE`-parameterized fixtures, standard exit contract, delivery discipline. + +3. **Idempotence at every layer.** L0-L2 get this from `tofu apply`'s own semantics; + L3/L4 procedure modules stay re-run-safe via the existing `check`/`apply --commit` + split (`dc-region-topology.sh`, `dc-plane-ipam.sh`, `phase-00-maas-standup.sh`) and + the MAAS "READY not deployed" handoff that `preflight.sh`'s drift-detection gate + already enforces at the L3/L4 boundary. The teardown primitive (Section 3.2) and + every new L5 gate inherit the same "safe to re-run, safe to re-verify" contract -- + the (a) control's `--check` and cloud-assert's A11a/A11b are explicitly RE-INVOKED + at multiple points in Section 3's sequence, not one-shot. + +4. **The strict IaC<->procedure boundary is the load-bearing rule, not a preference.** + Restated from Section 1: the boundary is identity (a MAC, an IP, a hostname), never + orchestration (a state file read, a cross-host provider dial). This is the rule the + container layer violated and the ONE thing Option 1 fixes structurally. Every new + module in Section 2/6 is placed on the correct side of this line by construction: the + (a) control and the power-key mitigation are procedure/L5, NOT tofu resources, even + though both are "about" IaC-produced objects (`pass1-admin-report.md` Section 3; + `pass2-w1-tofu-modules.md` Section 3.4) -- because both are LIVE verifications of a + running kernel/credential state, not declarations of desired infrastructure shape. + +5. **No layer reaches past the one directly below it.** L1 does not dial L3; L3 does not + dial L0's infrastructure directly; a Stage-4 script never reads `opentofu/*.tfstate` + (`pass2-w4-module-decomposition.md` Section 4's boundary-summary table). This is what + makes the per-`$SITE` loop in Section 3 safe to reorder, parallelize, or re-run a + single stage in isolation without re-deriving the whole sequence's state by hand. + +6. **Findings/design stay logged at their true layer, not folded upward.** The + rack-controller-remainder/D-131/mirror placement question is explicitly an L3 + placement decision (Section 2, L3 table) that this design does NOT resolve by + picking a value -- it states the dependency (Stage 3.5/4.6-4.8's conditional steps) + and defers to the ruling. Mirrors `pass1-w4-module-planning.md` Section 4 item 5's + own discipline. + +7. **D-140 is a distinct, future axis -- noted, not folded in.** D-140 (PINNED, not + ruled) would eventually make L4 IaC-managed too, consuming `dc-site`'s outputs from a + SEPARATE later root (e.g. `opentofu/vr1-dcN-juju/`) without reshaping L0-L3 + (`pass2-w1-tofu-modules.md` Section 4). The root-naming convention this design + recommends (`-`, Section 2 L2) stays generic enough to admit a + future `-juju` root per DC without a rename. + +--- + +## 5. THE ROOSEVELT-TRANSFER STORY -- per layer, not per artifact + +Per `pass1-w4-module-planning.md` Section 4 item 7's stated lens ("Roosevelt-transfer +judged per layer, not per artifact"): the pre-Roosevelt bare-metal test replaces the +VIRTUALIZATION substrate (vcloud's libvirt) with physical hosts. The question for each +layer is whether its module bodies assume libvirt/qemu, or whether they already operate +purely on live-observed identity (MAC, IP, hostname) that is substrate-agnostic. + +| Layer | Transfers to bare metal? | Why / why not | +|---|---|---| +| **L0** | **NO -- does not transfer** | The mesh triangle (dark-fiber `mesh-link` legs) and the netem-link DR-drill mechanism are virtualization-only shims with no bare-metal analog (`pass1-w4-module-planning.md` Section 4 item 7, citing `docs/dc-dc-deployment-workflow.md:11-14`'s own shim register). A physical multi-site test either has real inter-site links (no `mesh-link` needed) or none at all -- this layer's IaC bodies are simply not reused. | +| **L1** | **YES -- the client-VM PATTERN is the direct pre-Roosevelt deliverable** | The client VM ("the cloud-facing client lives IN the DC," D-138, `docs/design-decisions.md:7118-7124`) is explicitly the D-138 Roosevelt bastion analog, "rehearsed early" (`pass0-admin-report.md` Option-1 "For" bullet). On bare metal this becomes a physical or minimally-virtualized bastion host in the same role -- same D-138 principle, same L1 contract (booted, network-identified, non-hypervisor), different provisioning mechanism underneath. voffice1's D-114 pattern is separate and out of this pass's scope either way. | +| **L2** | **NO, as an IaC mechanism -- but the CONTRACT transfers** | `node-vm`/`dc-planes`'s libvirt-domain bodies have no bare-metal analog (physical hosts are not libvirt domains) -- this is the same "node-VM shim" `pass1-w4-module-planning.md` Section 4 item 7 names as non-transferring. What DOES transfer is the L2 CONTRACT itself: "produce a booted object with correct MAC-per-NIC network identity wired to the right planes" is exactly what a physical host's out-of-band provisioning (PXE + BMC) must also satisfy before L3 can start -- the CONTRACT is substrate-agnostic even though `dc-site`'s OpenTofu implementation is not. | +| **L3** | **YES -- unchanged, verbatim** | This is the clean transfer point: `dc-node-carve.sh`, `maas-node-power.sh`, `maas-role-tags.sh`, `dc-region-topology.sh`, `dc-plane-ipam.sh` already work by re-deriving everything from LIVE identity (`lib-hosts.sh:6-11`'s own design rule) and take the power/commissioning TARGET as an input, never a baked value. The one substitution point is `maas-node-power.sh`'s power backend (`power_type=virsh` -> a physical BMC/IPMI power type) -- a parameter change to the SAME script, not a rewrite, exactly the class of change D-143's address-substitution axis already demonstrates this repo tolerates cleanly. | +| **L4** | **YES -- wholly unchanged; this IS "the good test of the module deployment project"** | `bundle.yaml`, the per-DC overlays (`render-dc-overlays.py`), the VR0 template chain (`phase-00`..`phase-07-*.sh`), and `preflight.sh`/`cloud-assert.sh` have zero dependency on the compute substrate below L3's MAAS handoff -- they consume READY machines by API, not by libvirt identity. This is precisely the operator's own framing of the bare-metal test: *"a good test of the module deployment project we are developing during the teardown and redeploy"* (`SCOPE-AND-EXECUTION-PLAN.md` Section 1) -- L4 (and L3) are the module deployment project's actual payload; L0/L2's virtualization substrate is scaffolding around it. | +| **L5** | **MOSTLY YES, with one exception** | `preflight.sh`, `cloud-assert.sh`, `geneve-encap-assert.sh` are topology-agnostic (verified clean of container-layer assumptions, `pass1-admin-report.md` Section 2.3 / `pass3-admin-report.md` Section 1 check 4). The cross-DC host isolation control (concern i) and the power-key blast-radius mitigation (concern iii) are BOTH specific to vcloud's single-libvirtd co-residency problem -- on physical hosts there is no shared hypervisor kernel for two DCs' domains to leak across, so neither control's PROBLEM exists in the same shape on bare metal (a bare-metal equivalent, if any, would be a physical-network isolation control -- explicitly out of this design's scope, a Roosevelt-time question). | + +**Bottom line, matching the operator's own framing:** the layers that DO NOT transfer +(L0's mesh/netem shim, L2's libvirt-domain mechanics) are exactly the layers that exist +ONLY because vcloud is virtualized hardware standing in for real DCs -- not because they +encode anything about how OpenStack gets deployed. The layers that DO transfer (L1's +client-VM pattern, L3's MAAS-enlist procedures, L4's Juju/bundle procedures) are the +actual "module deployment project" the operator named -- rehearsing them on the 10.13 +flat topology now is the direct dry run for Roosevelt. + +--- + +## 6. Where the three isolation controls + the client VM + the teardown primitive sit + +Consolidated view (each already placed in Section 2/3 above; gathered here per the +task's explicit requirement): + +| Object | Layer | Kind | Why there | +|---|---|---|---| +| **Client VM** (`vr1-dcN-client`) | **L1** (IaC instantiation, same `cloudinit-vm` module type as voffice1/edges) | IaC module instance | It is a booted, network-identified, non-hypervisor VM -- exactly L1's contract. Its APPLY groups with L2's per-DC-flat root (teardown symmetry + SEC-026 state isolation + D-138 fidelity, `pass2-w1-tofu-modules.md` Section 2.4) even though its module TYPE is L1 -- layer classification and apply grouping are orthogonal, both hold (`pass1-admin-report.md` Section 1 check 7). | +| **Cross-DC host isolation control** (concern i, the "(a)" control) | **L0-scoped, L5 kind** (a procedure/verify artifact ABOUT an L0-level fact -- vcloud's own kernel-level forwarding) | Gate (procedure), NOT a tofu module | Confirmed explicitly: "SEC-010's actual pattern... NOT an OpenTofu module" (`pass1-admin-report.md` Section 3). Installed and `--check`-verified BEFORE the first Stage-3 (L2) apply of either per-DC root (Section 3.1's fork-robust invariant) -- it gates the layer above it without being IN that layer. | +| **SEC-010 transit-leg successor** (concern ii) | **L1-scoped** (installed on two L1 objects: the client VM + voffice1) | Procedure, one role-agnostic installer for both ends | It protects the transit LEG -- an L1-object property (both endpoints are `cloudinit-vm` instances) -- not a DC-substrate (L2) or commissioning (L3) fact. Homed at preflight P10, DC-scoped (`pass3-admin-report.md` Section 3). | +| **Power-key blast-radius mitigation** (concern iii) | **L3-scoped** (the credential itself is installed and used at L3 -- `maas-region-power-key.sh` / `maas-node-power.sh`), with its **gate split across L3/L4** (preflight P4 dependency + cloud-assert A11b standing re-verify, L5) | Procedure (credential-scope control) + L5 gate | It is fundamentally about WHO can dial WHICH domains via libvirt -- an L3 commissioning-credential fact, not an L2 substrate fact or an L0 network fact (`pass2-admin-report.md` Section 2.3's own distinction: "two network controls and one credential-scope control"). | +| **Teardown primitive** | **L2/L5 boundary** | Procedure that WRAPS an IaC destroy, verified by an L5-style completeness check | It operates on L2's state (a `tofu destroy` scoped to one DC's flat root, or a `virsh destroy` loop over that root's state-listed domains) but its CORRECTNESS property (no cross-DC leakage, no incomplete teardown) is verified the same way an L5 gate verifies anything else -- hence "L2/L5 boundary," matching `pass1-w4-module-planning.md` Section 3's own row for this artifact. | + +--- + +## 7. Open design questions (not resolved by this worker -- carried to the operator / delivery) + +Every item below is already OPEN in an earlier pass report; restated here only because +it bears directly on this module design's final shape. None is invented or pre-picked. + +1. **Root naming** (`vr1-dcN-flat` vs. reserving `-substrate`) -- blocks A1/A2 test + path re-points; Phase-4's call (`pass2-admin-report.md` Section 6 item 2). +2. **Root topology ratification** -- (B) is RECOMMENDED (Section 3.1), not yet ruled; + the whole composition sequence's per-DC independence claim depends on it. +3. **Rack-controller-remainder + D-131 forwarder + artifact-service (`.4`) placement** + -- gates Stage 3.5/4.6-4.8's conditional steps in Section 3; "THE highest-leverage + open item" per `pass1-admin-report.md` Section 7. +4. **The (a) control's concrete nftables rule set** -- kind/home/ordering are settled + (Section 2/6), the RULE SET itself is not. +5. **The power-key mitigation's mechanism** (restricted key / per-DC virsh wrapper / + polkit ACL) -- blocks `maas-node-power.sh`/`maas-region-power-key.sh` literals and + the `lib-hosts.sh` power-address re-derivation (Section 2, L3; the H1 + "plausible-URI-swap" hazard, `pass3-admin-report.md` Section 5). +6. **D-127's client-VM autostart value** -- needed for the new `opentofu-validate` T14/ + T15 case (Section 2, L1 table). +7. **The container-elim [ARCH] ruling itself** (D-123 amendment vs. new D-number, + + ride-alongs: D-125 retirement, D-138 concrete-host, D-128 amendment, D-122 site-down + re-earn, D-124 sizing-void) -- this is a PRECONDITION for Stage 3 ever running for + the first time (Section 3, C.2), framed by Phase 4's W4.4 dimension, ruled by the + operator (GA-R5), not by this document. +8. **`maas-fabric-prune.sh` harness gap** -- pre-existing, named Phase-4 decision + (Section 4 item 2). + +--- + +## 8. Verification note + +Author = "the worker" (no model name asserted, operator instruction). This document +performs no live measurement of its own -- every module/artifact claim traces to a +direct read of `pass0`-`pass3`'s administrator reports and their cited worker docs +(`pass1-w4-module-planning.md`, `pass2-w1-tofu-modules.md`, +`pass2-w4-module-decomposition.md`), plus `docs/tool-index.md`'s stated discipline +(consulted, not newly operated against). Every **[NEW]** artifact cites the pass/section +that scoped it; every **[EXISTS]** artifact cites its current path. No inferred values -- +every OPEN item in Section 7 is carried as OPEN, not resolved here. READ-ONLY; no +mutation performed. diff --git a/docs/audit/container-elim-pass/pass4-w3-execution-sequencing.md b/docs/audit/container-elim-pass/pass4-w3-execution-sequencing.md new file mode 100644 index 0000000..b3eac36 --- /dev/null +++ b/docs/audit/container-elim-pass/pass4-w3-execution-sequencing.md @@ -0,0 +1,231 @@ +# Pass 4 / W4.3 -- Execution sequencing into the teardown -> redeploy (module invocations) + +READ-ONLY planning artifact. Container-layer-elimination pass, Phase 4, Worker 3. +Confirmed ground truth consumed: `pass0-admin-report.md` Section 7a (**Option 1**; +cross-DC handling **(a)**; MAAS region stays on `vr1-dcN-maas-01`); `pass1-admin-report.md` +Sections 3-7 (planning change-set, the (a)-control spec, the L0-L5 layer model); `pass1-w3- +sequencing.md` (the base ordered sequence, adversarially corrected by the pass1 admin); +`pass2-admin-report.md` Sections 2-5 (the THREE isolation concerns, the 13 owed artifacts, +Phase-2 recommendation-grade resolutions); `pass3-admin-report.md` Sections 2-4 (the unified +tests change-set, the gate-home map, the #11 critical-path constraint). No mutation +proposed or performed here -- this is the ORDERED PLAN, expressed as module/artifact +invocations, for later gated execution. + +**Layer model used (adopted from `pass1-admin-report.md` Section 5, W1.4's design -- +NOT redefined here; W4.2 owns its full specification):** L0 host/inter-site substrate -> +L1 site/edge nodes (incl. the client VM as a `cloudinit-vm` instance) -> L2 DC substrate +(planes + node VMs; the two-root split collapses into L2) -> L3 enlist/commission -> L4 +Juju/OpenStack deploy (D-140 pins as procedure) -> L5 verify/gate (cross-cutting). + +Axis tags throughout: **[D-143]** (value/address substitution) / **[CE]** (container-elim, +shape change) / **[both]** (four confirmed dual-cause items, per pass1 check 5) / +**[unchanged]** (persists across the pivot, neither axis's content). + +--- + +## 0. Four critical-path invariants -- where each lands, verified against the sequence below + +| # | Invariant (source) | Lands at | Verified honored | +|---|---|---|---| +| 1 | **(a)-control BEFORE any flat substrate apply** (`pass1-admin-report.md` Section 3; `pass2-admin-report.md` Section 2.1) | **B.3, before B.4** | YES -- B.3 gate-closes before B.4.0/B.4.1 open; B.4 is the first step that can co-locate two DCs' plane bridges on vcloud's kernel | +| 2 | **Power-key mitigation (#11) BEFORE lib-hosts power-address re-derivation** -- blocks a 6-item test cluster (`pass3-admin-report.md` Section 4) | **B.6.1, before B.6.2** | YES -- B.6.2 (the `lib-hosts.sh` edit) is sequenced strictly after B.6.1's mechanism ruling + C4 harness; B.6.2's own gate lists the same 6-item cluster as blocked-together | +| 3 | **R7 revocation + MAAS record-release BEFORE substrate destroy** (`pass1-w3-sequencing.md` Part A steps 3-4; Part E risk 4) | **A.3-A.4, before A.5-A.7** | YES -- A.3/A.4 precede the destroy-plan/apply steps A.5-A.7 in Part A below, unchanged from pass1's ordering | +| 4 | **D-143 value-substitution kept attributable, separate from container-elim shape-change** (`SCOPE-AND-EXECUTION-PLAN.md` Section 7; pass1 check 5) | **Every row's tag column** | YES -- every step below carries its axis tag; the four genuinely dual-cause items (B.1.4, B.4/B.6.2 power-address, R7/A.3, B.7.1/B.7.4) are marked `[both]`, not silently folded into one axis | + +No step below violates any of the four; the invariant table above is the single +cross-check a future session should re-run if this document is ever amended. + +--- + +## PART A -- TEARDOWN of the current 10.12 Model-B checkpoint + +Baseline unchanged from `pass1-w3-sequencing.md` Part A (pass1-admin-report.md consumed it +without correction) -- today's tooling (`runbooks/dc-dc-teardown-rollback.md`) tears down +today's actually-deployed two-root shape; the flat-topology target does not change what +commands remove the OLD shape. + +| # | Module/artifact invoked | Layer (old shape) | Tag | Depends-on | Gate that closes it | +|---|---|---|---|---|---| +| A.1 | `tofu state pull` / backup, BOTH roots BOTH hosts (outer state on vcloud; each DC's inner state on the Office1 headend, D-128 Plane 2) | L2 (old two-root) | [D-143] | -- | backup file exists for both roots, both hosts, verified present before proceeding | +| A.2 | MAAS machine census, from voffice1 -- LENS 1 (every `power_type=virsh` record) + LENS 2 (corroborate vs `lib-hosts.sh` pinned boot-MAC roster) | L3 | [D-143] | A.1 | both lenses succeed; LENS 2 returns zero unattributed records | +| A.3 | **R7 credential-revocation checklist** (owed artifact #4) -- SEC-026 client cred, SEC-028 service cred, SEC-029 rack-local PKI copy (shred), D-126 per-env SSH keypair (clean retirement, no successor), MAAS rack enroll-secret residue; enumerated from EVERY `vm-secret-locations` row keyed to the rack/`vr1-dcN-client` host class, not only the named examples | L5 (credential gate) | **[both]** -- owed BY D-143's re-IP ruling (owed-execution item 5); the credentials revoked are container-elim-eliminated host classes | A.2 (hosts still up and reachable) | every row individually confirmed revoked/shredded/rotated with captured command output (never asserted from memory); rows marked RETIRED, append-only | +| A.4 | MAAS machine-record release/delete [MUTATION, operator-gated per record class] (owed artifact #5, amended by pass2 to include `maas rack-controller delete` decommission of `vvr1-dcN`'s Office1 registration + the region's `primary_rack`/DHCP reference) | L3 | [D-143] | A.3 (revoke before the record that names the credential-bearing host disappears) | LENS 2 (A.2's check) re-run and now returns zero; region-side rack-controller/`primary_rack` residue confirmed clean | +| A.5 | Plan destroy, INNER root (`opentofu/vr1-dcN-substrate/`), from voffice1; capture pre-destroy `virsh list/net-list/pool-list --all` baseline while the containment VM is still reachable | L2 (old inner root) | [D-143] | A.4 | plan reviewed; baseline captured | +| A.6 | Plan destroy, OUTER root (vcloud): `-target=module.vvr1_dcN -target=module.vr1_dcN_uplink -target=module.vr1_dcN_storage` only -- **do NOT target mesh/netem modules** (mesh triangle survives the pivot, confirmed at Phase 2/pass2 4.5) | L1 (old containment shape) | [D-143] | A.5 | plan reviewed; mesh/netem legs confirmed excluded from the target set | +| A.7 | Apply destroys, INNER then OUTER (`tofu destroy` plans from A.5/A.6); verify each half from the host that can see it (inner via the containment VM's virsh before it's gone; outer via vcloud's own virsh after) | L2 then L1 | [D-143] | A.6; each destroy individually operator-approved (hard rule 3 -- destructive steps never batched) | inner half destroyed + verified; outer half destroyed + verified | +| A.8 | Untargeted `tofu plan` drift gate (outer root; equivalent check on any surviving DC's inner root) | L5 | [D-143] | A.7 | zero unexpected drift | +| A.9 | NetBox DCIM decommission of the `vvr1-dcN` device record(s) (`netbox/dc-rack-mgmt-import.py`) | system-of-record (cross-layer) | [D-143] | A.7 confirmed gone | record marked decommissioned in live NetBox | +| A.10 | Repeat A.1-A.9 per DC (Path A) or run the Path B batched form for both DCs at once | -- | [D-143] | A.1-A.9 per DC | both DCs' checkpoints torn down and drift-gated | + +**A.10 note (site-down primitive):** D-122's one-command site-down (`virsh destroy +vvr1-dcN`) is still literally available during Part A (the shape being torn down still has +it) but is NOT used raw here -- module-scoped `tofu destroy` is the reviewed, gated path +(CLAUDE.md hard rule 4, the 2026-08-03 incident this rule exists to prevent). The +PRIMITIVE'S REPLACEMENT for the new flat shape is a Part-B design note (B.0), not built or +exercised during this teardown. + +--- + +## PART B -- REDEPLOY on flat 10.13 (Option 1), as module invocations + +### B.0 -- site-down primitive replacement (design note only; not an executable step) + +Under Option 1 there is no single VM whose destruction equals "the DC is gone." The +target-state primitive is **one module/root-scoped `tofu destroy` against the flat root's +`vr1_dcN_*` module set** (owed artifact #1, root-shape-(B)-dependent; owed artifact #6 is +the emergency `virsh destroy` loop over the same root's state-listed domains as a distinct +non-gated lever). Neither is built during the redeploy itself -- noted here so B.4/B.8 do +not read as having silently re-earned it without a delivered artifact. + +### B.1 -- prerequisites (run once, before any apply; GATES B.4) + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate that closes it | +|---|---|---|---|---|---| +| B.1.1 | **`dc-dc-whole-host-budget.py` FIT-calculator extension + fresh vcloud capacity re-measurement** (owed artifact #7, amended to include the 3 utility-node classes + artifact-service disk sizing) | L5 | **[both]** -- D-143 needs it for rebuild sizing; container-elim needs it to convert "~176 GiB freed" from directional to measured | -- | **>>> HARD GATE, closes before B.4's first apply <<<** live-measured host budget + extended FIT verdict for the flat 12+1-VM/DC roster, both DCs; a roster total exceeding the measured budget must FAIL, not be rounded or omitted (test A8) | +| B.1.2 | NetBox apex re-carve: mint B2 role `Cloud -- VR1 rebuild` owning `10.13.0.0/16`; re-carve VR1 prefixes/VIPs/ranges octet-preserving (D-143 owed-execution item 1) | system-of-record | [D-143] | -- | 10.13 prefixes exist under the new role; `Cloud` (10.12) role confirmed untouched | +| B.1.3 | `scripts/lib-net.sh`: flat defaults stay 10.12 (vr0-dc0 no-op preserved); `vr1-dc0`/`vr1-dc1` gain full 10.13 literal blocks; `:124-134` stale "inherits VR0" comment corrected (D-143 owed-execution item 2, F13) | L0 library | [D-143] | B.1.2 | `tests/lib-validate` green against the new literals | +| B.1.3a | 10.13 naming-collision DOCFIX (D-143 owed-execution item 3): `netbox/README.md:49` and `tests/dc-rack-mgmt-import/test_logic.py:265` reconciled so the /16 adoption doesn't leave a rack-IP test fixture inside the new Cloud space | doc/test currency | [D-143] | B.1.2 | repo-lint clean; no residual `10.13.0.5`-class fixture collision | +| B.1.4 | D-124 transit routes re-point + DC-side static shift (12->13) -- octet-preserving math is D-143; the BEARER host changes from `vvr1-dcN`'s rack-transit IP to the new `vr1-dcN-client` VM's transit leg (CE) | L1 (client-VM identity) / L0 (transit) | **[both]** | B.1.2 | NetBox shows the client VM's transit `/30` address; carried into B.4's client-VM module inputs | + +### B.2 -- Office1 + vcloud prep (persists, off the container layer entirely) + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.2.1 | vcloud host prep (Phase-0 runbook) | L0 | [unchanged] | -- | existing Phase-0 runbook gate | +| B.2.2 | Office1 headend standup (Phase-1 runbook), if not already up from the dc0 checkpoint -- voffice1 remains MAAS region host (D-132 addendum, unaffected) and remains Plane-2 host for whatever still needs a remote dial | L0/L1 | [unchanged] | -- | existing Phase-1 runbook gate | + +### B.3 -- the NEW cross-DC isolation control (concern i) -- MUST precede B.4 + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.3.1 | Build + install the vcloud-level host isolation control (owed artifact #2): nftables artifact asserting no inter-plane/inter-DC forwarding on vcloud's own kernel; new SEC-NNN row (next-free SEC-034 per pass3 check 6); harness `tests/` (C2, offline fixture model) | **NEW L5 artifact** (SEC-010's pattern, one layer up) | **[CE]** | B.2.1 (vcloud host prep) | **>>> HARD GATE, closes before B.4 <<<** `--check` enumerates the live bridge set for BOTH DCs' six planes and REFUSES if fewer than the full plane count resolves; must be installed and verified **before the first flat apply of EITHER per-DC root** (fork-robust invariant #1, honored regardless of root shape) | + +### B.4 -- the flat apply (formerly outer + bootstrap-gate + inner), COLLAPSED into L1/L2 + +Root shape recommended (B) shared-outer + per-DC-flat roots (pass2 Section 3.1, +recommendation grade, Phase-4 ratifies). One apply CYCLE per DC. + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.4.0 | Shared-outer root apply: mesh triangle legs (x3, MTU 9000) + uplink NAT `/24` -- already exist at vcloud level, UNCHANGED by flattening | L0 | [unchanged] | B.3 gate closed | existing mesh/uplink module state clean | +| B.4.1 | Per-DC-flat root apply invoking **`modules/dc-site`** (owed artifact #12, new IaC composition unit): composes storage pool (collapsed from 2-per-DC to 1) + six planes (`dc-planes`, unchanged CIDRs/MTU/family, D-139 unaffected) + DC edge (`opnsense-edge`, WAN input re-pointed to attach directly to `vr1_dcN_uplink` NAT, no `wan-bridge`) + 12 role/utility node VMs (`node-vm` x9 + juju-01 + maas-01 + tailscale-01) | L2 | **[CE]** structural; address VALUES inside these modules are **[D-143]** via B.1.3's lib-net literals | B.3 gate closed; B.1.1 FIT gate closed; B.1.3 lib-net literals present; B.1.4 client-VM address minted | apply succeeds (harness C1, `tests/dc-site`: 6-plane tree, 13-node roster incl. client VM = exactly one client-VM call, `mtu=1500` tenant / 9000 underlay, site-token parameterized, no hardcoded DC literal) | +| B.4.2 | Same `dc-site` composition, `client_vm` sub-call: the new **`vr1-dcN-client` VM** (`cloudinit-vm` module type, ~4/8192/80, octet `.8` metal-admin + D-124 transit leg per B.1.4) -- Option 1's non-hypervisor D-138 client host | L1 | **[CE]**, no D-143 content | B.4.1 (same apply) | client VM boots; identity confirmed live (not `lib-hosts.sh`/`CARVE_AUX_HOSTS`-tracked -- pass2 3.2 resolution) | +| B.4.3 | **MAC re-measurement pass** (owed artifact #8) -- 24 node MACs + 2 edge MACs move provider under the re-home; treat every MAC as unmeasured until re-pinned (hard rule 2) | L2 verify | [CE] | B.4.1 apply complete | live MAC roster captured and matches `dc-site`'s MAC invariant (rides C1's harness, pass3 3-rides list) -- **must close before B.6 trusts any MAC** | +| -- | *Vanished from the module set entirely (not re-homed, not a step):* `vvr1_dc0`/`vvr1_dc1` containment modules + sizing/rack-addressing/pubkey vars; the two `opentofu/vr1-dcN-substrate/` roots as separate roots/states; the qemu+ssh provider dial + D-126 keys (already revoked at A.3, no successor); `modules/wan-bridge` + `br-vr1-dcN-wan` netplan; the bootstrap gate's `--host-nodes` duty in `site-headend-install.sh` | -- | [CE] | -- | (deletion confirmed by B.4.1's module inventory containing none of these) | + +### B.5 -- surviving duties, re-targeted (Phase-2 recommendation-grade placement; Phase-4 ratifies) + +Pass1 left this OPEN ("THE highest-leverage open item"); pass2's W2.3 measurement +**largely dissolves it** -- carried here as pass2's recommended placement, explicitly +flagged not-yet-ratified. + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.5.1 | **Rack-controller: RECOMMEND RETIRE the standalone Office1 enrollment entirely** (pass2 4.2(i)) -- both `vr1-dcN-maas-01` VMs already run `region+rack` (dc1 transcript-grade, dc0 functional-grade, `pass2-admin-report.md` check 2); NO `--role rack` re-enrollment step runs against the client VM. **A current-day live re-measure of both DCs' `primary_rack`/rackd state is OWED before this is treated as settled** (instrument-currency discipline) | L3 (was; now dissolved) | [CE] | B.4.1 (maas-01 exists) | live re-measure confirms `region+rack` on both fresh `-maas-01` VMs; A.4's decommission of the OLD `vvr1-dcN` rack record already closed this in Part A | +| B.5.2 | **D-131 forwarder: RECOMMEND RETIRE-WITH-EVIDENCE** for both DCs (pass2 4.2(ii)) -- each fresh 10.13 region sets its own BIND from the start; retirement evidence = a dig test against the fresh region's own BIND (owed artifact #13, harness C5). dc1's asymmetry (its forwarder was load-bearing pre-teardown) is a standing-RED case in C5 **until this live retirement is confirmed on the new build** | L3 verify (one-time evidence capture, not a standing service) | [CE] | B.4.1 (region exists), B.5.1 | dig-resolver-identity test PASSES against each fresh region's BIND (both DCs); captured into the build changelog | +| B.5.3 | Artifact-service (`.4`) placement -- **OPEN, sizing decision, not resolved here.** Neither the client VM (~80 GiB) nor `vr1-dcN-maas-01` (150 GiB earmarked, no spare) fits dc0's full mirror as authored; resolved WITH NUMBERS by B.1.1's extended FIT calculator (dc0 full-mirror vs dc1 cache per D-135) | L2/L3 storage | [CE] | **B.1.1 gate** (numbers), B.4.1 | a right-sized, explicitly-provisioned home is chosen and confirmed before B.6.6 (artifact-source verify) depends on it | +| B.5.4 | **SEC-010 transit-leg FORWARD-drop successor (concern ii)** -- endpoints RESOLVED by pass2 recommendation: client VM (DC side) + voffice1 (Office1 side); ONE extracted role-agnostic installer subcommand (out of `site-headend-install.sh`'s `node_host_setup()`) installs BOTH ends (owed artifact #3) | **NEW L5 artifact**, distinct from B.3 | [CE] | B.4.2 (client VM exists, its transit NIC re-measured -- NIC-naming trap: dc0 live was `enp1s0`, not `mgmt`) | harness A3 (`site-headend-install` Section-8 sub-case, migrated): both-ends fixture passes; wrong-role interface-name fixture FAILS; new SEC-row disposition (possibly a SEC-010 amendment) minted | + +### B.6 -- MAAS enlist / commission / carve (largely unchanged mechanics; power-address value is BLOCKED) + +**>>> Invariant #2, the load-bearing one for this section: owed artifact #11 (power-key +blast-radius mitigation) MUST close, with its harness green, before B.6.2 -- B.6.2 is the +edit that a 6-item test cluster is pinned to (pass3 Section 4). A guessed URI/key value +here is exactly the false-green mint the inferred-value rule (hard rule 2) exists to +prevent. <<<** + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.6.1 | **Choose the concern-(iii) power-key mitigation mechanism** (owed artifact #11: `command=`-restricted SSH key / per-DC-scoped virsh wrapper / libvirt polkit ACL) -- required because BOTH DCs' region VMs + voffice1 now dial the SAME vcloud `qemu:///system` connection with no per-domain scoping (verified this pass, `pass2-admin-report.md` Section 2.3); root-shape (B) does NOT mitigate this (state isolation, not connection scoping) | **NEW L5 artifact** | [CE] | B.4.1 (flat topology exists to design against); B.2.1 (vcloud host prep) | **>>> HARD GATE <<<** harness C4 (offline rendered-ACL/`authorized_keys`/polkit parsing) green: dc0's key reaching a dc1 domain FAILS; reaching voffice1 FAILS (mirrored both directions); new SEC-NNN row minted with the mechanism | +| B.6.2 | `scripts/lib-hosts.sh` `VIRSH_POWER_ADDRESS_FROM_OFFICE1`/`_FROM_DCREGION` re-derivation to vcloud's own libvirtd (whether the FROM_OFFICE1/FROM_DCREGION split survives is itself a #11 design output) | L0 library | **[both]** -- D-143-shaped value change, CE-shaped mechanism dependency | **B.6.1 gate closed** | the SAME edit session closes A4 (`dc-selector`) + A5 (`maas-region-power-key` URI assertions) together -- editing one without the other is the named hazard H1 (a plausible-but-unenforced URI swap) | +| B.6.3 | `scripts/maas-node-power.sh` invocation-site/runbook literal updates -- no code change (address is `$1`, topology-agnostic) | L3 | [D-143 value] | B.6.2 | every invocation-site/runbook example carries the new address | +| B.6.4 | MAAS discovers node VMs via per-machine `power_type=virsh` (mechanism unchanged, D-123 amendment 2026-07-20 -- the pod/`vm-host` mechanism was already refuted before this pass; only the VALUE re-derives) | L3 | [D-143 value] | B.6.2, B.6.3 | discovery succeeds against the new power addresses | +| B.6.5 | Commission each node (Phase-3 Step 3); tag; nodes stay Ready; Pattern-A interface carve via `scripts/dc-node-carve.sh` (no code change, grep-verified zero containment hits, pure MAAS-API, ``-parameterized) | L3 | [unchanged] | B.6.4, B.4.3 (MACs re-measured) | commission + carve succeed; single virsh hop, no "which host can see this domain" ambiguity (the old containment-layer confusion is structurally gone) | +| B.6.6 | PXE/boot-fabric verify (Step 6); per-DC artifact-source + time-authority verify (Step 7, target host = wherever B.5.3 lands); topology consistency check (Step 8) | L3 | [D-143 value for target host] | B.6.5, **B.5.3 resolved** | verify steps pass against the ruled artifact-service host | + +### B.7 -- Juju + OpenStack bundle deploy (mechanically unchanged, host renamed) + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.7.1 | Juju controller bootstrap on a tagged machine, run FROM the `vr1-dcN-client` VM (D-138 execution host, "never the vcloud jumphost" doctrine intact) -- same commands, renamed host from `vvr1-dcN` | L4 | **[both]** -- mechanics [D-143] unchanged, execution-host identity [CE] changed | B.4.2 (client VM up), B.6.6 (enlist/commission/carve complete) | `preflight.sh` gate (incl. new P10 for B.5.4, P4 dependent on B.6.1's mitigation, P8's declared-root-list loop replacing the single hardcoded `opentofu/` path) | +| B.7.2 | model/spaces setup, DC egress gate (`dc-egress-check.sh`, invocation-host literal re-points to the client VM), `juju deploy bundle.yaml` + overlays, dry-run first, mid-deploy watch | L4 | [unchanged mechanics] | B.7.1 | Phase-4 runbook's own dry-run + watch gates | +| B.7.3 | SEC-026/028/029 credentials (re-)MINTED on the client VM -- Part A step A.3 already revoked the old copies; this is a clean mint, not a migration (pass1 check 7 reconciliation) | L4 credential | **[both]** | B.7.1 | ledger register rows (P5 creds-matrix) re-point to the `client` host class; new rows added for B.3/B.5.4/B.6.1's own key material | +| B.7.4 | vault bring-up (Step 5), IPv6 family-matrix overlay (Step 6), phase-03/04/05 steps (7-9), `cloud-assert.sh --capture` (Step 10), controller backup (Step 11) | L4/L5 | [D-143 for literal addresses; mechanics unchanged] | B.7.2 | existing Phase-4 gates, unchanged content | +| B.7.5 | **VERIFY-LIVE gates** (Step 12): Ceph-over-v6 bind check (D-143 new literals); `geneve-encap-assert.sh` (owed artifact #10 -- MTU budget analytically unchanged by container-elim, live assert still owed as a NEW invocation point, not a script edit); **PLUS the B.3 (a)-control's `--check` RE-RUN**, now that both DCs' planes are actually co-resident on vcloud -- the first point in the whole sequence where the isolation claim is tested against real traffic (cloud-assert A11a periodic re-verify begins here) | L5 | [D-143] for Ceph literals; **[CE]** for the (a)-control re-run and the geneve re-invocation point | B.7.4 | Ceph bind check passes; geneve-encap-assert passes; (a)-control `--check` passes with BOTH DCs' planes live | + +### B.8 -- close-out (both axes) + +| # | Module/artifact invoked | Layer | Tag | Depends-on | Gate | +|---|---|---|---|---|---| +| B.8.1 | NetBox DCIM: register the new `vr1-dcN-client` VM + the flat node roster device records (owed artifact #9; mirrors A.9's decommission, opposite direction) | system-of-record | [CE] | B.4.1, B.4.2 | roster fully registered in live NetBox | +| B.8.2 | **Enter the container-elim [ARCH] decision record** (D-123 amendment vs. new D-number; ride-alongs: D-125 bridge-in retirement, D-138 concrete-host change, D-128 amendment -- Plane 2 shrinks to MAAS/NetBox, D-122 site-down re-earn, D-124 sizing-void re-cause, D-132-addendum premises note, Stage-3 Owns/Reuse-vs-new rewrite, the client-VM `.8` octet into the D-134 map) | -- (decision-record, not a build step) | [CE] | all of Part B substantively complete | operator ruling recorded (GA-R5) -- **the redeploy is NOT "done" while this is owed**, per SCOPE Section 7 | + +--- + +## Bare-metal-test plug-in points (pre-Roosevelt, hardware specs owed "within days") + +Per the operator's own framing and D-138's history (the client VM is explicitly "the D-138 +Roosevelt bastion analog, rehearsed early," `pass0-admin-report.md` Section 4), the specs +plug into **sizing and provider targets inside existing steps**, not a new sequence. Split +by whether a step is hardware-PARAMETERIZED (re-run with new numbers/roster once specs +land) or hardware-AGNOSTIC (unaffected either way): + +**Hardware-parameterized (specs plug in directly):** +- **B.1.1** -- capacity/FIT math re-derives from real host specs instead of vcloud's + measured budget; the node-class roster may shrink ("a smaller set of hardware"). +- **B.4.1/B.4.2** -- `dc-site`'s module bodies (planes x6, edge, node roster, client VM) + are the SAME composition intended to transfer; the specs let sizing be re-derived for + bare metal without changing which modules exist or their call order -- only the provider + target (bare libvirt host vs vcloud-hosted VM) and concrete sizing numbers change. +- **B.4.3 / B.6.4** -- MAC handling: bare metal uses real NIC MACs discovered via MAAS + enlistment, not libvirt-injected MACs from a tofu module -- the RE-MEASUREMENT discipline + transfers, the injection mechanism does not. +- **B.5.3** -- artifact-service sizing is explicitly hardware-dependent (disk on real + spinning/SSD capacity vs a vcloud-hosted volume). + +**Hardware-agnostic (unaffected; the reusable asset IS the sequence/module shape):** +- **B.1.2/B.1.3** -- NetBox/lib-net addressing is pure IPAM, runs identically regardless of + substrate. +- **B.6.5** -- carve scripts (`dc-node-carve.sh` etc.) are already ``-parameterized, + pure MAAS-API; grep-verified zero containment hits (pass2 3.3). +- **B.7** -- Juju/bundle deploy mechanics are unchanged by substrate class (D-140 pins L4 as + procedure regardless). +- **The module INVOCATION ORDER itself (B.1 -> B.8)** is the reusable deliverable -- the + "layered module system" the operator asked this pass to produce IS the thing rehearsed + early, independent of whether L1/L2 target vcloud-libvirt VMs or bare hosts. + +**Two controls that likely do NOT transfer as-is (flag explicitly, do not let a future +session assume parity):** +- **B.3's (a) cross-DC isolation control** is an artifact of hosting BOTH DCs' plane + bridges on ONE vcloud-libvirt kernel. Genuinely separate physical hardware per DC at + Roosevelt would not have this specific exposure -- the CONTROL PATTERN (host-level + forwarding assertion) may still be good practice, but the gap it closes here may not + exist at bare metal. Phase-4 should note this, not assume the artifact ships unchanged. +- **B.6.1's power-key mitigation (#11)** is a libvirt/`qemu:///system`-connection-scoping + problem specific to multiple VMs dialing one shared hypervisor's virsh. Bare-metal power + management (IPMI/Redfish, not virsh) is a structurally different credential-scope + problem; the MITIGATION DISCIPLINE (per-DC credential separation, verified by a negative + test) transfers, the concrete artifact does not. + +**Nesting-depth honesty (carried from `pass1-w3-sequencing.md` Part D, not restated as new +here):** flat-Option-1 = `vcloud (libvirt) -> node VM -> nova KVM guest`, depth 2 +(VR0-proven). Roosevelt bare metal = `bare host -> KVM guest`, depth 1. The flattening +rehearses the MODULE SHAPE and the single-apply/no-bootstrap-gate WORKFLOW, not the exact +nesting depth. + +**B.5's placement should be RE-DECIDED, not carried forward blindly, once hardware specs +exist** -- a bare-metal build may have a materially different answer for where MAAS-rack/ +D-131/artifact-service duties land than a VM-hosted `vr1-dcN-client` does (this echoes +`pass1-w3-sequencing.md` Part D verbatim; not re-derived, cited). + +--- + +## Sources read this session + +`SCOPE-AND-EXECUTION-PLAN.md` (full), `pass0-admin-report.md` (full), `pass1-admin-report.md` +(full), `pass1-w3-sequencing.md` (full), `pass2-admin-report.md` (full), `pass3-admin-report.md` +(full), plus `docs/design-decisions.md:8083-8182` (D-143 full ruling + owed-execution list, +re-read directly this session to confirm the five owed-execution items and their exact +wording). No inferred values used; every module name, owed-artifact number, and gate +condition traces to the cited pass document. Items marked OPEN/RECOMMENDATION-GRADE above +are stated as such, not silently upgraded to settled. diff --git a/docs/audit/container-elim-pass/pass4-w4-decision-framing.md b/docs/audit/container-elim-pass/pass4-w4-decision-framing.md new file mode 100644 index 0000000..032a109 --- /dev/null +++ b/docs/audit/container-elim-pass/pass4-w4-decision-framing.md @@ -0,0 +1,305 @@ +# Pass 4 -- Worker W4.4: the owed [ARCH] decision framing (container-elim + layered module workflow) + +**Author:** W4.4 (Phase 4, `SCOPE-AND-EXECUTION-PLAN.md` Section 4). **Date:** 2026-08-09. +**Inputs read in full:** `SCOPE-AND-EXECUTION-PLAN.md`, `pass0-admin-report.md`, +`pass1-admin-report.md`, `pass2-admin-report.md`, `pass3-admin-report.md`; `docs/design-decisions.md` +entries D-122, D-123 (+ both amendments), D-124 (+ 3 amendments), D-125, D-126, D-127, D-128, +D-131, D-132 (+ 2026-07-30 amendment + addendum), D-134 (+ 5 amendments), D-138, D-143; the GA-R3 +admission test (`SKILL.md:406-410`). **READ-ONLY.** This document FRAMES a decision package for the +operator (GA-R5); it rules nothing. Every recommendation below is graded "recommend," never "ruled." + +--- + +## 0. What this pass confirmed is already SETTLED and is not re-opened here + +Per pass0 Section 7a (operator-confirmed at the Phase-0 gate, a DIRECTIONAL PLANNING +CONFIRMATION -- explicitly **not** a GA-R5 [ARCH] ruling): target topology = **Option 1** (flat +node VMs on vcloud libvirt + one small non-hypervisor `vr1-dcN-client` VM per DC); cross-DC +adjacency handling = **(a)** (accept co-residency + a new vcloud-level host isolation control); +MAAS region stays on `vr1-dcN-maas-01`. Phases 1-3 planned, tooled, and tested against this. **The +formal [ARCH] ruling this document frames is the one thing the Phase-0 gate explicitly deferred** +(SCOPE Section 7: "the container-elim itself is an OWED [ARCH] decision... Phase 4 FRAMES it... +it is ruled by the OPERATOR"). + +--- + +## 1. THE CORE QUESTION: D-123 amendment, or a new D-number? + +### 1.1 Apply the GA-R3 admission test verbatim (`SKILL.md:406-410`) + +> "a new D-number needs architectural consequence beyond the stage, a Roosevelt-delta (the A1 +> test: a Roosevelt build session would grep it before touching a built surface), or +> supersession. Everything else is OPERATIONAL... doubt resolves DOWN to OPS." + +All three admission triggers fire, independently: + +- **Architectural consequence beyond the stage.** Confirmed across every phase: the change + restructures Stage 3 wholesale (pass1 Section 2.2), amends D-128's own definition (pass1 + check 6), voids/re-causes a D-124 clause (pass0 row 6, Section 6 item 8), terminates D-125 + (Section 3 below), reopens D-131's standing-pattern status (pass2 4.2(ii)), touches D-132's + addendum premise (pass2 4.2(i)), forces a D-134 map addition (pass2 4.3), and changes D-138's + concrete host (pass0 check 3). No single existing entry's scope contains all of that. +- **Roosevelt-delta (A1 test).** Explicit and operator-stated: SCOPE Section 1 quotes the + operator directly -- the redeploy is "a good test of the module deployment project we are + developing during the teardown and redeploy," and SCOPE Section 7 names the layered module + workflow a "Roosevelt deliverable... design it to transfer to the pre-Roosevelt bare-metal + test." A future Roosevelt-adjacent session would grep this decision before laying out any + future site's substrate shape or module structure -- textbook A1. +- **Supersession.** D-123's core adopted ruling -- Model B, 4-level nesting, node VMs live + INSIDE `vvr1-dcN`, single-`virsh-destroy` site-down (`design-decisions.md:4929-4939`) -- is + directly reversed. Option 1 does the opposite of what D-123 rules: nodes return to being flat + vcloud-level siblings (D-123's own rejected "Model A" shape, now readopted with 2026-07-30+ + rulings layered on). + +### 1.2 The load-bearing precedent already in this ledger: D-143 + +D-143 is the closest analog and it is a **directly on-point precedent, not an analogy**: D-143 +"AMENDS D-115 premise" and "TERMINATES D-101... clause" (`design-decisions.md:8083`) -- exactly +the shape of relationship the container-elim has with D-123/D-124/D-125/D-128 -- **and it was +minted as its OWN new D-number**, not appended as a D-115 or D-101 amendment. D-143's +Reconciliation section (`:8124-8147`) is the template this document's Section 3 follows: a new +entry states its verb (SUPERSEDES / AMENDS / TERMINATES / CONSISTENT-WITH / PRESERVES) against +each affected decision, rather than the affected decision being edited to absorb the reversal. + +### 1.3 Why "D-123 amendment" under-fits + +D-123 already carries two prior AMENDMENT entries: the 2026-07-16 ruling that ADOPTED Model B +over the recommended Model A, and the 2026-07-20 correction that Model B's virsh-pod MECHANISM +was refuted (per-machine power replaces it) while Model B's SHAPE stood. A third amendment that +reverses D-123's own central ruling -- the shape itself, the thing the decision exists to answer +-- would have D-123 amend itself out of existence: a future reader grepping "D-123" for "what +shape is a DC site" would need to read three superseding layers to reach a NO that contradicts +the entry's own header. GA-R1's append-only discipline is better served by leaving D-123's history +intact (Model A recommended -> Model B ruled -> mechanism refuted) and recording the reversal as a +fresh, separately-dated decision that names D-123 as SUPERSEDED -- exactly D-143's pattern against +D-101/D-115. + +### 1.4 RECOMMENDATION + +**Mint a new D-number** (next-free = **D-144**, per `ledger-scan.sh` run this session inside +pass3 -- `pass3-admin-report.md` Section 1 item 6: "D next-free=144" -- re-verify at mint time per +repo numbering discipline, do not assign here). Title shape (for the operator's edit, not a +pre-write): *"D-144: VR1 container-layer elimination -- flat per-DC substrate + per-DC client VM, +layered module workflow (SUPERSEDES D-123 Model B)"*. This is a **recommendation**, presented +alongside the D-123-amendment alternative below for the operator to weigh; GA-R5 forbids the pass +from picking. + +**Alternative presented for completeness (not recommended): amend D-123 in place.** Would keep +one canonical "DC site shape" entry instead of a reader needing D-123 -> D-144 to reach current +truth. Weighed against: it fights the D-143 precedent, and D-123's entry would need heavy internal +surgery (its own "Model A" section would need to become "Model A, re-adopted under D-144" -- +essentially rewriting the entry to point at a new one anyway, which is the same reader cost with +none of the whole-history discoverability of a fresh entry). + +--- + +## 2. RIDE-ALONG DECISIONS -- each framed ([ARCH]/[OPS], reconciliation, Roosevelt-delta) + +Seven items, per the assignment; doubt resolves DOWN to OPS (GA-R3). + +### (1) D-128 amendment -- the Plane-1/Plane-2 execution-host split + +**[ARCH].** D-128 is itself tagged [OPS] in the ledger, but this specific edit meets the GA-R3 +bar on its own: it changes a RULED decision's own DEFINITION (pass1 check 6, verified direct +read `design-decisions.md:5354-5364`) -- Plane 2 currently includes "the INNER `tofu` root +(`opentofu/vr1-dc0-substrate/`, `qemu+ssh` FROM Office1 into `vvr1-dc0`)"; under Option 1 that +object ceases to exist, so the substrate build becomes wholly Plane 1 (vcloud-local) and Plane 2 +shrinks to MAAS/NetBox (juju/openstack already moved off Plane 2 by D-138). **Reconciliation:** +AMENDS D-128 -- narrow, single-entry scope change, not a reopening of the two-plane MODEL itself +(Claude-stays-on-jumphost, workstation-is-human-path all stand unchanged). **Roosevelt-delta:** +indirect -- D-128 states explicitly it has "No Roosevelt analog" for Plane 1 (throwaway simulation +scaffolding); the amendment doesn't change that, it only shrinks what Plane 2 covers. +**Recommendation:** own dated "D-128 -- AMENDMENT" entry, riding the same operator exchange as the +core ruling (it is a direct, mechanical consequence of Option 1, not an independent question) but +recorded against D-128, not folded into D-144's body. + +### (2) D-125 bridge-in retirement -- within the core ruling, or its own note? + +**[ARCH]** (D-125 is itself [ARCH]) **but with no independent architectural surface once +separated from the core question.** D-125 exists SOLELY to solve OBS-3 -- the egress-dead-end +created by Model B's WAN NAT living inside a nested containment VM with no direct route out +(`design-decisions.md:5091-5096`). Remove the nesting and OBS-3's precondition disappears; there +is nothing left for D-125 to fix. This is not a parallel decision riding alongside the core +ruling -- it is a direct, total consequence OF it (pass2 Section 3.1: `wan-bridge` "collapses," +edge WAN reattaches to the vcloud-level NAT directly -- literally D-122's ORIGINAL pre-D-125 +intent, restored). **Roosevelt-delta:** none of its own -- D-125's bridge-in mechanism was always +VR1-only rehearsal scaffolding for a nesting pattern that itself has no Roosevelt analog. +**Recommendation:** do NOT mint or independently amend D-125 as a ride-along item; record its +TERMINATION as a named consequence inside the core new entry's reconciliation ledger (Section 3 +below), the same way D-143 named D-101's clause TERMINATED inside D-143's own body rather than as +a separate D-101 amendment entry. + +### (3) The D-134 `.8` octet addition for `vr1-dcN-client` + +**[ARCH]-adjacent by D-134's own established pattern, but narrow.** D-134 is tagged [ARCH] and +has FIVE prior dated amendments, three of which are single-slot octet-map extensions with the +identical shape to this one (`.5` juju 2026-07-29, `.6` MAAS region via the D-132 addendum, +`.7` tailscale 2026-08-07 -- `design-decisions.md:5947-6007`). Each was minted as its own dated +"D-134 -- AMENDMENT" entry, never folded into the decision that motivated the new host class +(D-104 for juju, D-132 for MAAS, D-129(iii) for tailscale). The client VM is `.8`, next free slot +(pass2 Section 4.3, W2.2 verified against the standing table which enumerates through `.7`). +**Reconciliation:** AMENDS D-134 (adds one row; does not touch the CIDR/band structure itself). +**Roosevelt-delta:** real and explicit -- D-134's own amendment text states "A future DC standup +carves `.N` for its [service] by this standard" (`:6006`); a Roosevelt build session greps D-134 +for the utility-band map before assigning any new per-DC service an address, exactly the A1 test. +**Recommendation:** own dated "D-134 -- AMENDMENT" entry, following the established pattern +exactly (motivated by/cited from the core ruling, minted separately) -- consistent, not a new +D-number, not folded into D-144's body. + +### (4) Root-topology (B) + root naming + +**[OPS].** Shared-outer + per-DC-flat tofu roots vs a single merged root is a **tofu-state +organization choice**: both alternatives deploy the identical physical/network topology (pass2 +check 3, verified: "root-shape (B) mitigates NONE of" the three isolation concerns -- it is about +state blast radius, destroy scoping, and apply ordering, not what gets built). It does not change +what exists on the wire or in the hypervisor; it changes how the tofu STATE that describes it is +partitioned. Root NAMING (`-flat` vs reserving `-substrate`) is purely a repo convention. +**Reconciliation:** none against an existing D-number -- no prior decision rules tofu root +topology at this granularity. **Roosevelt-delta:** weak/indirect -- Roosevelt has no "tofu roots" +concept for physical hardware it doesn't create; the delta that DOES transfer (module composition, +`dc-site` as a reusable unit) is captured by the module-workflow design itself, not by the root +split. **Recommendation:** ratify as part of the Phase-4 module-workflow design / delivery +change-set (an OPS-graded record: changelog entry + `docs/dc-dc-deployment-workflow.md` Stage-3 +Build-line text), not a D-number. + +### (5) The three isolation controls' SEC rows -- power-key mitigation especially + +Three distinct controls (pass2 Section 2, verified this session's cross-reads); each graded +separately: + +- **(i) cross-DC plane-bridge network adjacency (the "(a)" control) and (ii) the SEC-010 + transit-leg successor: [OPS].** Both are new/re-authored SECURITY MITIGATIONS -- nftables + artifacts with `--check` gates and SEC-NNN ledger rows -- of the same kind SEC-010 itself was, + and SEC-010 was never a D-number (D-125 only cross-references it). GA-R3's OPS bucket is + precisely "runbook edit / session-changelog line / as-built row" for this class of work; a SEC + row plus a harness is the established mechanical pattern (pass3 confirms next-free is + **SEC-034**, computed by direct ledger grep since `ledger-scan.sh` does not compute SEC numbers). + **Recommendation:** SEC-ledger rows, not D-numbers; the REQUIREMENT that they exist and their + ordering invariant (installed before any flat apply) is recorded as an owed artifact inside the + core new entry's execution notes, mirroring how D-125 recorded SEC-010's constraint without + itself being a SEC entry. + +- **(iii) the MAAS power-key blast radius -- the hardest of the three, argued in detail.** This + is the item the prompt specifically asks whether it "also warrants a D-number given it makes + SEC-012/016 per-DC separation vacuous" (pass2 Section 2.3, verified this session's citations: + `maas-region-power-key.sh` installs a per-DC key SEC-012/SEC-016 assume stays scoped; + `maas-node-power.sh:28-30` confirms the REGION dials power; `opentofu/main.tf:175` confirms + voffice1 is a sibling domain on the SAME outer libvirtd that will also hold every flattened DC + node -- so under Option 1, each DC's region-resident power key opens a `qemu:///system` + connection with virsh control over EVERY domain vcloud manages, not just its own DC). **Grading + reasoning:** the mechanism CHOICE (restricted key / wrapper / polkit ACL) is squarely OPS -- a + SEC-NNN mitigation like (i)/(ii). But the FINDING itself -- that flattening (the core ruling's + own structural consequence) silently VOIDS a previously-relied-upon per-DC credential-isolation + guarantee that SEC-012/SEC-016 were written to provide -- is architectural framing content: it + is a tradeoff the operator is accepting BY ruling the core question, not an independent decision + with its own Roosevelt-delta or supersession target. It has no existence apart from the core + ruling (unlike D-125, above, it doesn't supersede or terminate any EXISTING decision by name -- + SEC-012/016 aren't D-numbers to supersede -- it just makes them functionally moot). **Net + classification: [ARCH]-adjacent finding, [OPS] mitigation.** **Recommendation:** do NOT mint a + separate D-number for this. Record the finding and the accepted tradeoff explicitly INSIDE the + core new entry's body (so a future reader sees "flattening was known to weaken per-DC power-key + scoping, mitigated by SEC-034[+1]" in the same place they see the topology ruling) -- this is + the one item where recommending "fold into the core entry's text, not its own line-item" matters + most, because burying it only in a SEC row (which nobody greps before touching architecture) + would repeat exactly the failure class GA-R3's A1 test exists to prevent. + +### (6) The rack-controller retirement + D-131 retire-with-evidence + +**Split, two different triggers, two different answers -- do not conflate them.** + +- **Rack-controller retirement itself: [OPS].** Both DCs' `region+rack` MAAS controllers already + measurably run all rackd duty (pass2 check 2, direct reads: `changelog-20260807-dc1-region- + sequence.md:80-89` dc1 transcript-grade; `changelog-20260730-dc0-region-migration.md:~532-549` + dc0 process-measured, `pgrep -c dhcpd`=0 on the rack). Retiring `vvr1-dcN`'s vestigial Office1 + rack registration is a decommission step (a `maas rack-controller delete` + runbook note) -- it + does not change any RULED architecture, it cleans up a registration nobody is using. This is the + MAAS machine-record release/delete class of work already scoped as owed artifact #5. +- **D-131 retire-with-evidence: [ARCH]-touching, but AMENDS D-131, does not need a new D and is + NOT itself created by the container-elim.** D-131 sub-decision 1 explicitly RULED the forwarder + as "the STANDING per-DC pattern... applied to dc0 and part of every future DC standup's + definition-of-done" (`design-decisions.md:5658-5661`) -- retiring that changes an ARCH decision's + forward applicability, which is a real edit, not a nit. But the TRIGGER is D-132's already-ruled + per-DC-region architecture (2026-07-30) removing D-131's own stated precondition ("rack-only + controller, remote region") -- container-elim did not create this fact, it merely removes the + vestigial host that made the asymmetry easy to miss (dc0 already proves the end state; dc1's + forwarder is still load-bearing today, pass2 check 4 -- a real, unequal-grade asymmetry, not + assumed equal). **Reconciliation:** AMENDS D-131 sub-decision 1 (own dated entry, D-131's + existing amendment-free structure notwithstanding -- this would be its first). **Roosevelt- + delta:** real -- D-131 sub-decision 4 is explicitly "PINNED... to be executed at next-deployment + design time," i.e. this retirement IS that next-deployment design point arriving early. + **Recommendation:** own dated "D-131 -- AMENDMENT" entry, gated on the owed live re-measure + (current-day `primary_rack` state, both DCs) landing BEFORE the ruling is asked for, not a + D-144 ride-along and not folded into D-144's body (its trigger is independent of container-elim). + +### (7) Artifact-service re-homing/sizing + +**[OPS].** A capacity/placement decision resolved by measurement (the FIT-calculator extension, +owed artifact #7/#13 amendment) against two concrete VM sizings that are both too small for dc0's +full mirror as authored (pass2 Section 3.3, verified sizing cites: `vr1-dc0-substrate/main.tf: +176-181` maas-01's 150 GiB earmarked "no spare"; client VM ~80 GiB). No architectural principle is +at stake -- it is "which host gets a bigger disk," decided with numbers, not a ruling about shape. +**Reconciliation:** none against an existing D-number. **Roosevelt-delta:** none identified -- +mirror/cache sizing is a per-deployment capacity fact (D-135 already owns the dc0-full-mirror / +dc1-cache-proxy split as an ARCH decision; this ride-along is purely which HOST realizes it here). +**Recommendation:** resolved via the FIT-calculator numbers at delivery, recorded as a changelog +entry, no D-number. + +### Ride-along count / [ARCH]-[OPS] split + +**7 ride-along items.** Classification: **2 carry [ARCH] weight** (item 1, D-128 amendment; +item 3, D-134 amendment) that get their OWN dated amendment entries against existing D-numbers; +**1 is [ARCH]-in-substance but has no independent existence apart from the core ruling** (item 2, +D-125 -- folds into D-144's reconciliation ledger, not its own entry) **plus a second such item** +(item 5-iii, the power-key finding -- folds into D-144's BODY as an accepted tradeoff, its +mitigation MECHANISM graded OPS/SEC-row); **1 splits into an [OPS] cleanup + a separate [ARCH] +amendment** (item 6: rack retirement OPS, D-131 amendment ARCH but independently triggered); +**2 are cleanly [OPS]** (item 4 root-topology, item 7 artifact-service sizing); item 5-i/5-ii (the +two network isolation controls) are OPS/SEC-row work referenced from D-144's text. **Net: 0 new +D-numbers beyond the core D-144; 3 existing-decision amendments (D-128, D-134, D-131) filed +separately; 2 findings folded into D-144's own body; the remainder is SEC-ledger/changelog/runbook +work.** + +--- + +## 3. THE RECONCILIATION LEDGER -- what the container-elim (D-144, proposed) does to each existing decision + +Verb vocabulary matches D-143's own precedent (`design-decisions.md:8124-8147`): **SUPERSEDES** +(replaces the cited ruling outright), **AMENDS** (a factual premise or scoped clause changes, the +rest of the ruling holds), **TERMINATES** (a clause/mechanism is retired with no successor), +**PRESERVES** (the ruling's principle stands unchanged; at most its concrete realization updates), +**CONSISTENT-WITH** (untouched, cited for completeness). + +| D-number | Verb | What changes / what does not | +|---|---|---| +| **D-122** (site shape) | **AMENDS** | Intent preserved (each site still has a distinct containment/entry-point identity -- now the client VM); the "single `virsh destroy ` = site-down" LITERAL claim regresses to a root/module-scoped group-destroy (pass0 row 9) -- a real capability loss D-144 must state honestly, not silently drop. Dark-fiber/per-site-ISP/edge-shape clauses: untouched. | +| **D-123** (Model B containment) | **SUPERSEDED** | The core ruling -- nodes nested inside `vvr1-dcN`, depth-4, single-VM destroy -- is reversed. D-123's own history (Model A recommended -> Model B ruled 2026-07-16 -> mechanism refuted 2026-07-20) stays intact, append-only; D-144 records the supersession rather than editing D-123. | +| **D-124** (region<->rack transit addressing) | **AMENDS** (a clause reverses again) | The Scheme-A transit addressing (`172.31.0.0/24`, region/rack `/30`s) SURVIVES -- the client VM keeps the transit leg. The 2026-07-16 "sizing VOID under Model B" amendment (rack must hold a full node fleet, ~416/480 GiB) itself gets RE-CAUSED: under Option 1 the client VM is NOT a node-fleet host, so sizing returns toward something near D-124's small original proposal (4 vCPU/8192 MiB/80 GiB) -- exact figure OWED, not re-derived here (pass0 Section 6 item 8, "D-124 sizing-void re-cause"). | +| **D-125** (bridge-in egress) | **TERMINATES** | Its sole reason to exist (OBS-3, nested-WAN-NAT-with-no-egress) disappears with the nesting. `modules/wan-bridge` retires; edge WAN reattaches directly to the vcloud-level per-DC uplink NAT -- D-122's ORIGINAL pre-D-125 realization, restored (see item (2) above; folded into D-144's body, not filed separately). | +| **D-125's `wan-bridge` fallback note (double-NAT)** | moot alongside D-125 | Never adopted; moot with the primary mechanism gone. | +| **D-126** (rootless SSH access convention) | **PRESERVES** | The general decision (Option A, systemd --user local-forward, per-env key convention) is a reusable ACCESS PATTERN used by multiple site-service VMs, not solely the qemu+ssh dial into `vvr1-dcN`. Only ONE consumer of the pattern retires: the qemu+ssh dial's per-env key, which has "no successor" (pass1 Section 4, Part A step 3) because its object (the inner root's cross-host provider) is gone. D-126 itself is unchanged; note the retired consumer as a changelog line, not a D-126 edit. | +| **D-127** (VM autostart policy) | **AMENDS** | The explicit classification row "`vvr1-dc0` (and future vr1-dc1 containment) -> MANUAL" (`:5307-5311`, itself cross-referencing D-123) has no successor object. The client VM needs its OWN classification against D-127's stated rule of thumb (foundational/no-fragile-boot = autostart; deliberate/gated/resource-heavy = manual) -- pass3 confirms this is an OWED ruling, not inferable (A1's new test case cannot be written until it lands). Recommend a short D-127 amendment adding the client-VM row once ruled; the POLICY framework itself is unchanged. | +| **D-128** (two-plane operating model) | **AMENDS** | Real, substantive (item (1) above): Plane 2's own definition shrinks (loses the inner-root qemu+ssh object); the two-plane MODEL and Claude's-jumphost-residency are unaffected. | +| **D-131** (node-facing DNS strategy) | **AMENDS** (independently triggered, rides alongside) | Sub-decision 1's "standing per-DC pattern" status is reopened by D-132's already-ruled per-DC regions removing D-131's own stated precondition; container-elim only removes the vestigial host, it doesn't create the trigger. Filed as its own D-131 amendment (item (6) above), gated on the owed live re-measure. | +| **D-132 addendum** (region in own VM at `.6`) | **PRESERVES; premise-note only** | No change -- MAAS region stays on `vr1-dcN-maas-01` (operator-confirmed at the Phase-0 gate). Flag only: the addendum's original "hypervisor-fate" rationale (why the region got its own VM rather than co-locating) becomes MOOT under Option 1 (nothing is co-locating with a hypervisor anymore) -- a premise-currency note inside D-144's text, not a reversal (pass2 4.2(i)). | +| **D-134** (octet bands / standing cross-DC map) | **AMENDS** (new row, existing pattern) | Existing bands/CIDRs unchanged; adds `.8` = client VM, following the identical amendment shape as `.5`/`.6`/`.7` (item (3) above). Filed separately, per D-134's own established convention. | +| **D-138** (client lives IN the DC) | **PRESERVES; concrete host updates** | The PRINCIPLE -- cloud-facing tools (Juju, `openstack` CLI) dial the cloud from inside the DC, not from `voffice1` -- is UNCHANGED and is exactly what Option 1's client VM realizes. Only the CONCRETE HOST changes: from `vvr1-dcN` (a DC-node-fleet-containing hypervisor) to the new small non-hypervisor `vr1-dcN-client`. Credential-residency consequences (SEC-026/028/029 migration, pass0 row 8) are a direct, named consequence to record in D-144's reconciliation, not a D-138 rewrite. | + +**Not superseded/amended by this change (verified, named to prevent scope creep):** D-114 +(Office1's OWN containment VM, `voffice1` -- a structurally distinct, KEPT pattern; pass1 check 1 +is explicit that Stage 2 is untouched and the two "containment VM" patterns must be named as +DIFFERENT going forward so name-similarity doesn't sweep D-114 in); D-133 (flat per-NIC plane +carve); D-139 (IPv6-only east-west); D-140 (Juju-as-tofu, pinned for a LATER redeploy, explicitly +not this one); D-143 (the re-IP this pass rides -- additive, kept distinguishable per axis +throughout Phases 1-3, `[D-143]`/`[CE]`/`[both]` tagging). + +--- + +## 4. Summary for the operator (what D-144, if ruled, would need to say) + +If the operator adopts the "new D-number" recommendation, D-144's body would need, at minimum: +(a) the core supersession of D-123's Model B; (b) the D-125 termination and D-124 sizing +re-cause, stated as direct consequences, not separate rulings; (c) the accepted power-key +blast-radius tradeoff (Section 2 item 5-iii) named explicitly, with its SEC-row mitigation as the +mitigating control, not the whole answer; (d) pointers to the three amendments filed alongside it +(D-128, D-134, D-131) so a reader lands on the complete picture from one grep. This document does +not draft that text -- GA-R5 reserves the ruling, and the ruling's exact wording, to the operator.