Newer
Older
openstack-caracal-dc-dc / docs / audit / container-elim-pass / pass3-w1-harnesses-gauntlet.md

Pass 3 -- W3.1: harnesses + the gauntlet that assume the container layer

Author: W3.1 (Phase-3 worker, SCOPE-AND-EXECUTION-PLAN.md Section 4). Date: 2026-08-09. Method: surveyed all 103 harnesses in tests/HARNESS-MANIFEST / tests/*/run-tests.sh via targeted greps (vvr1, qemu+ssh, 172.31, --host-nodes, expose_nested_virt, Model B, node_host, inner root, bootstrap gate) plus full reads of every harness a grep hit implicated, cross-checked against pass0-admin-report.md / pass1-admin-report.md / pass2-admin-report.md (which artifacts are owed/blocked/recommendation-grade). READ-ONLY; no mutation; nothing here was executed. All CHANGE/RETIRE items are proposals feeding Phase 4; several are explicitly CONTINGENT on Phase-4 rulings still open per pass2 Section 6/7 (root naming, rack-controller retirement, D-131 retirement, concern-(iii) mitigation mechanism) -- marked so below, not asserted as settled.


1. Per-harness disposition table

Only harnesses with a container-layer/two-root/qemu+ssh/vvr1/bootstrap-gate/depth-4 hit are listed (13 of 103; the other 90 were grep-confirmed to have zero hits on the search terms above and are not re-litigated here). "STAY" = the case's assertion target and value are unaffected by Option 1 flattening.

Harness:case Asserts about the container layer Verdict New invariant (if CHANGE)
opentofu-validate:T13 (:100) opentofu/main.tf autostart pin # D-127: DC containment VM for vvr1-dc0/vvr1-dc1 (D-127 boot-matrix regression guard) RETIRE No successor pin -- the object is deleted, not renamed. Removing this case is itself the correct edit (not a silent drop: the D-127 boot-matrix comment block above it must say so)
opentofu-validate:T14/T15 (:101-102) opentofu/vr1-dc0-substrate/main.tf autostart pins: DC edge =true, node VMs =false CHANGE Same invariant (edge boots, node VMs stay MAAS-owned/manual), re-pointed at the flat root's file -- path depends on the OPEN root-naming ruling (pass2 Section 6.2: -flat vs reserving -substrate); do not pre-guess the filename. Also ADD a new case for the vr1-dcN-client VM's autostart value -- D-127's table predates the client VM and has no ruled value for it yet (OWED, not inferred)
opentofu-validate:T8-T10 (S3 root-only blind spot) Uses modules/node-vm + synthetic modules/broken/modules/fine fixtures to prove root-only tofu validate misses uncalled modules STAY Module-body level; unaffected (pass2: "zero module bodies need rewriting" except wan-bridge)
opentofu-validate header comment (:92) "the 416 GiB nested container" doc-currency nit, not a graded case rides the T13 retirement edit
node-vm:T8-T11 (:47-98) Counts 12 macs = [ lists / 72 pinned MAC literals in opentofu/vr1-dc0-substrate/main.tf (the INNER root, hardcoded path) CHANGE Re-point INNER= to the flat root's file (same open root-naming dependency as T14/T15 above). Count invariant likely UNCHANGED (12 nodes x 6 planes persists -- the client VM is a separate cloudinit-vm instance, not a node-vm instance, so it does not join this count per pass2 3.2's "client VM does NOT join CARVE_AUX_HOSTS") -- confirm at build, do not assume
node-vm:T1-T7,T12-T15 Module-body assertions on opentofu/modules/node-vm (MAC-pin variable shape, power-ownership ignore_changes) STAY Module body unchanged
site-headend-install: Section 8 (:110-159, the --host-nodes block) node_host_setup()/node_host_check(): nested KVM, inner libvirt pool, D-125 WAN-bridge verify (master $WAN_BRIDGE, br_netfilter warning, uplink enslavement), the vr1-dc0-substrate hint-neutrality check, SEC-010 forward-drop keyed to --transit-if RETIRE (block), CHANGE (SEC-010 sub-case) node_host_setup()/node_host_check() and all D-125 WAN-bridge assertions RETIRE wholesale (pass2: dead code, delete; D-125 bridge-in has no successor -- edge WAN goes direct-NAT). The SEC-010 forward-drop assertions (SEC-010, FORWARD-drop, transit-if-exists fail-open check, nft declare-then-delete idempotency) CHANGE: re-target the extracted role-agnostic installer subcommand (pass2 Section 2.2/4.4 -- one artifact installing BOTH the client-VM end and the voffice1 end), asserting it runs on EITHER role, not gated behind --host-nodes
site-headend-install: Section 6 (:80-102, --role rack) maas init rack, region+rack step suppression, enrollment-secret non-leak, for the standalone rack-controller role CONTINGENT CHANGE/RETIRE -- blocked on Phase-4 ratifying pass2's rack-retirement recommendation (4.2(i), currently RECOMMENDATION-grade, not ruled; needs the "owed live re-measure" first). If ratified: --role rack retires (new builds region+rack directly on the region VM); this section's cases retire and a new case asserts --role rack is REFUSED/removed. If rejected: section stays as-is New invariant depends on the ruling -- do not pre-pick
site-headend-install: Sections 1-5,7 (arg contract, --compose-cidr, dry-run mutate-nothing, region+rack default flow) Nothing container-layer-specific STAY --
dc-selector: power-address rows (:204-247) Hard-pins VIRSH_POWER_ADDRESS/_FROM_OFFICE1/_FROM_DCREGION to qemu+ssh://...@172.31.0.{2,6}/system and .../@10.12.{8,68}.2/system -- i.e. the literal values that dial the CONTAINMENT VM's libvirtd via the transit/metal-admin legs CHANGE -- HIGHEST RISK, currently BLOCKED New values dial vcloud's OWN libvirtd instead; pass2 Section 2.3/3.2 has this explicitly blocked on the still-undesigned concern-(iii) power-key mitigation (owed artifact #11) -- URI shape, whether the FROM_OFFICE1/FROM_DCREGION split survives at all (both DCs may converge on one vcloud endpoint), are OPEN. Do not write a "corrected" literal into this harness before that mitigation is ruled -- that is exactly how a false-green gets minted (a guessed value that happens to match nothing live)
maas-node-power: URI fixture (:60) URI="qemu+ssh://jessea123@172.31.0.2/system" exercised purely as an opaque pass-through arg (maas-node-power.sh takes the address as $1, no code change per pass2) STAY (functionally); doc-currency CHANGE optional Script logic is unaffected -- the fixture value doesn't need to be a real address. Consider refreshing the example post-rebuild so it doesn't read as a still-live containment address, but this is cosmetic, not a graded risk
maas-region-power-key: URI assertions (:67,82,102,108) Pins the derived power-key install target as qemu+ssh://...@10.12.8.2/system / ...@10.12.68.2/system -- the metal-admin .2 address that is TODAY the rack containment VM's identity CHANGE -- BLOCKED, same dependency as dc-selector maas-region-power-key.sh's own body is unchanged (pass2 3.3) but the <region>-<dc> -> address DERIVATION it exercises resolves to a containment-VM identity that ceases to exist. Blocked on concern-(iii)'s mitigation design exactly like dc-selector -- these two harnesses should be updated TOGETHER, from the same ruled URI shape, or they will silently diverge
dc-rack-net: LEGS cases (T3,T5,T13,T15,T17) Pins vr1-dc0-metal-admin=10.12.8.2/22, vr1-dc0-provider-public=10.12.4.2/22, dc1 equivalents, plus the systemd unit ordering for ${SITE}-rack-legs.service -- the containment VM's OWN bridge-leg addressing RETIRE (contingent, but on the STRONGER-settled side) Per pass2 3.3: "Legs half RETIRES with the containment layer -- no flat VM is a libvirt host with its own bridges; LEGS/br_of() has no home to move to." Not blocked on a mitigation design the way concern-(iii) is -- this is a structural consequence of Option 1 itself
dc-rack-net: DNS-forwarder cases (T4,T9,T16 + the DNS_UPSTREAM literal) D-131 forwarder confinement (no-resolv, single upstream, bind-interfaces) CONTINGENT RETIRE -- blocked on Phase-4 ratifying D-131 retire-with-evidence (pass2 4.2(ii), RECOMMENDATION-grade, dc0 evidence-graded FUNCTIONAL not transcript, dc1 asymmetric and currently load-bearing -- "owed live re-measure" first). Note independently: DNS_UPSTREAM="10.10.0.20" is already flagged STALE for dc0 in pass2 3.3, a pre-existing defect unrelated to container-elim If retired: this whole harness (all 18 cases, both LEGS and DNS halves) retires with dc-rack-net.sh itself. If D-131 is instead kept for a DC: the DNS half survives re-targeted at wherever component (i)'s host lands
dc-rack-net:T2,T6,T7,T8,T11,T12,T14,T18 Generic script-hygiene cases (bash -n, no virbrN literal, MEASURED-tag discipline, unknown-site/-mode refuse, do_check read-only) rides the parent verdict If the script retires wholesale these retire with it; they carry no container-layer-specific content of their own
dc-rack-mgmt-import: D-124 scheme pins (:77-78) "rack_dns": "vvr1-dc0" / "vvr1-dc1" literal role-name pins inside netbox/dc-rack-mgmt-import.py's SITES map CONTINGENT CHANGE/RETIRE Owed artifact #9 (pass1/pass2: NetBox DCIM migration -- decommission the vvr1-dcN device records, register the client VM + flat roster). If the rack-controller retirement (4.2 i) is also ratified, the whole "rack DNS device" concept this script imports may retire, not just get renamed -- do not pre-pick between rename-in-place and full replacement; that is a Phase-4/NetBox-migration-design call
dc-rack-mgmt-import: D-124 CONTAINER/transit-scheme pins (:73,79 + d124-transit-seed:42,48) CONTAINER = "172.31.0.0/24", .1-gateway rejection, /30-or-/31 shape, dcim.site scope STAY The transit supernet and its NetBox scheme are mesh-triangle facts, not containment facts (pass2 4.5: mesh triangle + all 3 mesh-link legs unchanged; only the office1-leg CONSUMER re-points from the containment VM's NIC1 to the client VM's transit NIC -- an endpoint change these value-pins don't encode)
maas-profile-assert: office1-profile fixture (:41,68,72,80) Simulated MAAS machine list [voffice1, vvr1-dc0, vvr1-dc1] for the "office1" profile (the rack controllers as MAAS-visible machines) CONTINGENT CHANGE Fixture data, not logic. If rack-controller retirement (4.2 i) is ratified, vvr1-dc0/vvr1-dc1 stop existing as MAAS machines under that profile -- update the fixture to the post-retirement roster. The client VM does NOT replace them here: per pass2 3.2 it is not MAAS-carved, so it does not appear in this profile's machine list either
preflight: pending-change fixture (:89) Synthetic tofu plan line module.vvr1_dc0.libvirt_domain.vm will be updated in-place, used only to exercise the PENDING-change-detection regex STAY Fixture-only; the detection logic is module-name-agnostic. Cosmetic rename optional, non-load-bearing
geneve-encap-assert: (all cases) Family-split (C1) / bracketed-encap-ofport (C2) OVN checks, driven entirely from input files STAY Confirmed family/topology-agnostic -- no vvr1/containment coupling found. This is the harness that will gate the STILL-OWED live geneve/jumbo assert on the vcloud-level planes post-build (pass0 Section 8 item 3) -- that live assert is a NEW USE of this same offline-tested tool, not a harness edit
site-baseleg: active-leg-row guard (:69-73) Asserts NO un-commented DC supernet (10.12./172.31.) appears in an active LEGS row -- i.e. structurally proves the script is STILL a no-op for DC legs STAY Already forward-compatible: this guard's job is to keep the script a no-op until a DC leg is explicitly measured and added, and Option 1 doesn't change that precondition (pass2 4.5: re-cite D-138 + the (a) control in the comment instead of the retired qemu+ssh premise -- doc-currency only, not a case change)
cloudinit-vm: (all cases, T1-T6,T5a-T5d) D-130 lifecycle guard + MAC-pin shape on opentofu/modules/cloudinit-vm -- the module type the client VM will be a NEW INSTANCE of (pass2 3.1) STAY Module-body level, unaffected. No new W3.1 case needed for the client-VM instantiation itself (that is an apply-time/W3.2-W3.3 concern -- does the client VM need its own MAC pins once carved, etc. -- not a change to THIS harness's existing assertions)
netem-link: (all cases) Header comment references "the outer root runs ON vcloud" (D-128 local-mode amendment) STAY Confirmed false-positive on the survey's "inner/outer root" grep -- this is the OUTER vcloud root's own local-vs-SSH execution mode, unrelated to the inner/outer CONTAINMENT split. Zero vvr1/qemu+ssh hits. Mesh triangle unchanged

Harnesses grep-confirmed with NO container-layer coupling (checked because their scripts were named in pass2 as "no code change" and could plausibly hardcode a containment value, but do not): maas-role-tags, carve-host-interfaces, dc-node-carve, dc-egress-check, dc-plane-ipam, dc-region-topology, provider-bundle-check, lib-validate (the lib-hosts.sh/lib-net.sh/ lib-identity.sh unit-test harness -- exercises emit/vr_json/env-scrub plumbing generically, not the containment-keyed values themselves; those live in dc-selector, already covered above).


2. Count and gauntlet-edit list

13 of 103 harnesses affected (12.6%): 2 clean RETIRE, 2 contingent RETIRE (pending Phase-4 rulings), 6 CHANGE (2 of them currently BLOCKED on an undesigned mitigation), 1 contingent CHANGE (fixture only), plus module-body/fixture-only harnesses noted STAY for completeness (node-vm's module cases, opentofu-validate's S3 cases, cloudinit-vm, netem-link, site-baseleg, geneve-encap-assert, preflight, d124-transit-seed).

Gauntlet cases needing edits, with the FAILING-DIRECTION fixture each needs (per repo discipline: an assertion must be provably able to fail; a fix that makes a finding-string assertion stale gets REPLACED with the new invariant, never deleted to go green):

  1. opentofu-validate T13 -- delete the case; add a fixture that the D-127 boot-matrix pin-table comment block is updated (grep-checkable: the comment must no longer describe a containment-VM autostart row that doesn't exist). Failing direction: a stray re-add of a vvr1-dc0 autostart pin in opentofu/main.tf should have NO test catching it post-retire -- flag this as an accepted residual gap, or add a negative-assertion case ("no vvr1 domain block exists in opentofu/main.tf") so a regression is still caught.
  2. opentofu-validate T14/T15 + node-vm T8-T11 -- re-point the hardcoded opentofu/vr1-dc0-substrate/main.tf path to the ratified flat-root file once Phase 4 rules root naming; keep the exact-count/exact-value assertions (12/72, edge=true/node=false) as the failing-direction fixture (a wrong count or a flipped autostart bool must still redden).
  3. site-headend-install -- the biggest single edit. Split into: (a) delete node_host_setup()/node_host_check() assertions outright (RETIRE, ~15 grep cases); (b) write NEW cases against the extracted role-agnostic SEC-010 installer subcommand, with a fixture that installs it on a fake "client" role AND a fake "voffice1" role and asserts BOTH ends get the drop rule, using the same fail-open discipline already proven here (verify the named interface actually exists, not just that a rule loaded); (c) hold Section 6 (--role rack) unedited until the rack-retirement ruling lands -- editing it now would be guessing Phase 4's answer.
  4. dc-selector + maas-region-power-key -- do NOT edit these yet. They are correctly RED the moment lib-hosts.sh's power-address derivation changes, which is exactly the fail-loud behavior wanted while concern-(iii) is unmitigated. Edit them ONLY together with the mitigation's ship (owed artifact #11), from the ruled URI/key shape -- editing either one first, or guessing a value, is the false-green trap this repo's rule about inferred values exists to prevent.
  5. dc-rack-net -- do not edit pending the D-131/rack-retirement rulings; if both ratify, RETIRE the harness file wholesale alongside dc-rack-net.sh (append-only bias: leave the file in git history, remove from tests/HARNESS-MANIFEST via --record-manifest, log the removal in a changelog with a revert per repo discipline).
  6. dc-rack-mgmt-import -- hold for the NetBox-migration design (owed #9); do not rename the vvr1-dc0/vvr1-dc1 string pins speculatively.
  7. maas-profile-assert -- fixture-only edit, lowest risk of the CHANGE set; safe to update once rack retirement is ratified (drop the two vvr1-dcN rows from the office1-profile fixture), independent of the other blocked items.

3. Highest-risk invert-under-flattening cases (the ones most likely to false-green or

false-red if edited carelessly)

  1. dc-selector's power-address pins and maas-region-power-key's derived-URI pins -- these are the sharpest risk in the whole survey. Both assert LITERAL containment-VM addresses that MUST change, but the replacement values are explicitly BLOCKED on an undesigned mitigation (concern iii, pass2 Section 2.3, owed artifact #11). The failure mode to guard against is a well-meaning edit that swaps in a plausible-looking new URI (e.g. "just point it at vcloud's own address") before the restricted-key/wrapper/ACL mechanism is actually built -- that produces a harness that is GREEN against a value nothing enforces, which is worse than the current honest RED.
  2. site-headend-install's --host-nodes block -- inverse risk: because --host-nodes itself might simply be REMOVED as a flag (not just have its body gutted), a case like t "--host-nodes without --role rack -> 2" could keep passing for the WRONG reason (an unrecognized-flag exit code that happens to also be 2), silently changing what the assertion proves without ever going red. Any edit here must re-derive the expected exit code from the NEW arg-parse contract, not assume rc=2 still means what it meant before.
  3. opentofu-validate T13/T14/T15 and node-vm T8-T11 -- risk is a stale hardcoded PATH silently reading as "file not found" -> FAIL, which is safe (loud, not silent), but a sloppy fix that just deletes the whole case to make the gauntlet green again (rather than re-pointing it) would violate the repo's "replace, never delete to go green" rule and quietly drop the D-127 autostart-drift regression coverage this class of harness exists for (it has gone red twice before on a missed re-run, per node-vm's own commit-history commentary at tests/node-vm/run-tests.sh:68-81).
  4. dc-dc-whole-host-budget (flagged in Section 1's supplementary note, not the main table since it asserts Model A/B math rather than a container-layer FACT directly) -- worth naming here because it is the harness most likely to stay GREEN while testing something no longer true: its Model-A/B RAM comparison (838 vs 822 GiB, "containment overhead = 2x16 GiB") will keep passing indefinitely against the OLD script even after the topology it models is gone, because nothing forces a re-run against a NEW --model flag until the FIT-calculator extension (owed #7) actually ships. This is squarely W3.2/W3.3 territory (new-model-flag test requirements) but is flagged here as the standing risk that motivates it.

4. Durable-doc path

docs/audit/container-elim-pass/pass3-w1-harnesses-gauntlet.md (this file).