Authored 2026-07-18 by the grounding-audit Phase-2 ground-truth agent (charter: docs/audit/grounding-audit-charter.md, section 3) at repo HEAD e999b03 on branch dc-dc-stage3-phase2-dc-substrate. Every claim below carries its evidence (path:line, quoted command output, or commit hash). Claims that could not be evidenced read-only are marked UNKNOWN with what would resolve them. Nothing here is guessed.
STATUS OF THIS DOCUMENT: SIGNED by the operator 2026-07-19 (section 11; charter Phase 6 item 5). STANDING RULE (GA-R1, RATIFIED 2026-07-18 with amendments C1+C2 -- docs/audit/ga-rulings.md): no status claim is hand-written anywhere else; other documents point HERE; this document cites captured command output, and measurement always wins over it (C2). Other status surfaces are pointers or history; where one still carries a claim, it is a defect (Phase 1 proved they contradict: docs/audit/record-inventory.md, 12 groups; findings GA-F01..GA-F15).
docs/design-decisions.md:1946).docs/dc-dc-deployment-workflow.md:148; runbook runbooks/dc-dc-phase2-tofu-dc-substrate.md) -- CLOSED 2026-07-21 for its vr1-dc0 scope (operator-ruled "Close and merge"; dc1 was the stage's designed HELD remainder, gate G12 -- now also CLOSED 2026-07-23: dc1 substrate built + commissioned 9/9, merged to main, branch retired; see the G12 gate row). Close-out set: gauntlet ALL GREEN + repo-lint 0-fail + this consolidation commit + GA-R7 memory review + merge of dc-dc-stage3-phase2-dc-substrate to main (merge commit) + branch retirement; stage record docs/archive/stage-records/vr1-stage3-record.md. Stages 0-2 precede it; stages 5-7 are authored, not executed (workflow doc:873).runbooks/dc-dc-phase3-maas-enlist-deploy.md): CLOSED 2026-07-27 (operator-gated "Merge it and retire the branch"). MERGED to main as merge commit 6f5701d (2 parents, NOT squashed; 77 commits from branch dc-dc-stage4-phase3-maas-deploy, opened 2026-07-23 off post-merge main), branch RETIRED local + remote after confirming containment via git branch --merged main. Post-merge verification ON main: gauntlet ALL GREEN (81), repo-lint 0-fail. Stage record: docs/archive/stage-records/vr1-stage4-record.md. Delivered: 18 nodes READY (NOT Deployed -- MAAS-deploy SKIPPED per DOCFIX-200), carved across 10 named plane fabrics with pinned MACs, tags 9+9 verified live, jammy images synced, and two deliberately different per-DC artifact strategies proven (dc0 full mirror / dc1 caching proxy). DoD bullets 1-5 met; bullet 6 STRUCK (DOCFIX-204, unsatisfiable under D-129(iv)); node-side half split to gate G17. TRAVELLING FORWARD by design: G17, 21 open SEC rows, and 7 residual credential-register findings incl. the dc0 edge-API re-mint. The NEXT stage branches off post-merge main. POST-CLOSE DURABILITY SWEEP 2026-07-27 (operator-requested before clearing the session; capture docs/audit/queued-findings-20260727.txt, precedent queued-findings-20260726.txt). Four transcript-only items were landed on surfaces, of which one was consequential: dc-mirror.sh's dc1 site row carried no warning that dc1 is no longer a mirror site, so dc-mirror.sh install dc1 would have silently rebuilt the whole removed apparatus -- enabled daily debmirror timer included, and a fresh ~950G pull -- which is exactly what someone would reach for on treating a failing check dc1 as a regression. The row now states dc1 is proxy-only by the D-135 amendment, that check dc1 FAILS BY DESIGN, and that install dc1 is a deliberate strategy change and never a repair; a RUNTIME guard is queued, not built (hard rule 1). Also: a WRONG causal claim in stage4-mirror-gate-20260727.txt was superseded by an appended correction (reset-failed cannot re-arm an inactive timer -- the vector was a REBOOT, via Persistent=yes firing immediately on a missed window); the generalisable lesson is now platform-traps section 5 ("stopped is not dormant across a reboot"; and a oneshot with RemainAfterExit=yes does not undo its work on stop, so stopping it proves NOTHING about dependents -- the reason the dc1 teardown test deleted the live addr/route instead); and dc-cache-proxy.sh's two-net-units-coexist claim is now marked REASONED, NOT MEASURED (hard rule 2 -- nobody has run both on one host). Dangling-reference sweep: every path this session introduced or cited RESOLVES; the pre-existing dangles are all legitimate (deleted-as-history, not-yet-built, or the deliberately-absent octavia PKI overlay that preflight P4 fails on). ledger-scan reconciled: 21 open SEC, next-free D 138 / DOCFIX 205 / BUNDLEFIX 053. History of the stage while it was open follows. Precondition PASS + discovery CAPTURED read-only (docs/audit/stage4-discovery-20260723.txt): both DC racks enrolled (vvr1-dc0 7chphy, vvr1-dc1 nmpcq4); 9 nodes per DC ALL Ready, power=virsh, dc0 boot fabric-4 / dc1 fabric-142, zone/pool both default (no zone/pool convention -- fabric + MAC prefix is the grouping signal). Runbook deltas vs reality: commissioning (Step 3) already satisfied by Stage 3; order is carve-then-deploy per the runbook's own Ready-state sequencing note; VR1 nodes have 6 flat per-plane NICs (enp1s0 metal-admin/boot, enp2s0 provider-public, enp3s0 metal-internal, enp4s0 data-tenant, enp5s0 storage, enp6s0 replication). D-133 ADOPTED 2026-07-23 (flat per-NIC carve for VR1; next deployment rehearses hardware-faithful bonds/trunks). D-134 ADOPTED + AMENDED 2026-07-23 (contiguous bands, Option B: .4-.49 utility, .50-.99 VIP, nodes one run .100-.200 (control .100-.119, compute .120-.149, storage .150-.200), dynamic MOVES to .201-.254 -- the carve's first gated mutation per DC. VR1 nodes: control .100-.102, compute .120-.121, storage .150-.153, keyed by tofu name / pinned boot MAC). CARVE EXECUTED 2026-07-23, BOTH DCs (logged window stage4-carve; capture docs/audit/stage4-carve-verify-20260723.txt): dynamic ranges .201-.254 in place (office1 untouched); 10 named plane fabrics; 90 NIC re-homes (18 nodes, read-back each, pinned MACs matched); 2 provider subnets moved off fabric-5 (+ ruled .1 gateways) + 8 plane subnets created; 6 spaces (bundle binding names) with 12 VLANs assigned; 18 nodes' statics + br-ex OVS (D-133 flat, D-100 provider-raw). ALL 18 nodes still Ready. Flagged residue (cleanup gated at stage close): 192.168.1.0/24 + emptied fabric-5, ~90 empty auto-created fabrics. Post-carve rulings, all 2026-07-23 (this entry lands one commit late -- an owned L10 defect, changelog item 8): deploy model = READY handoff (MAAS-deploy SKIPPED; Juju provisions at Stage 5 per the phase-01 precondition; DOCFIX-200 rewrote the runbook's Step 4, which would have broken Stage 5); jammy boot images SYNCED to the region (selection id=2; ubuntu/jammy reads Synced); placement tags = openstack-vr1-dc0/-dc1 APPLIED 9+9 (D-119 naming). Mirror gate: rehearsal exception REJECTED -- D-135 ADOPTED (per-DC mirror BUILT on the rack hosts at the D-134 utility .4 address; item 1 apt+UCA debmirror+nginx builds in-stage; items 2-3 pinned to Stage 5; egress narrowing is the closing mutation). dc-mirror.sh SHIPPED 2026-07-23 (site-keyed check/install/sync; D-134 utility .4 + edge default route -- egress path MEASURED working, edge DNS answers; debmirror GPG-verified jammy triple + UCA caracal; harness 18/18; gauntlet ALL GREEN (77), docs/audit/gauntlet-20260723-stage4-dcmirror.txt). INSTALLED BOTH RACKS 2026-07-23 (check PASS 15/15 each; mirrors answer 200 on 10.12.8.4 / 10.12.68.4; initial syncs RUNNING -- HOME-under- systemd fix shipped same-hour, harness 19/19). INCIDENT resolved same-hour (changelog item 9): the fresh dc1 edge passed NO LAN traffic -- pf ruleset never regenerated after v4 addressing (the set-interface script reloads the interface, not the filter; dc0 was masked by its 07-20 plugin work). Fix: configctl filter reload; dc1 rack egress now 0% loss / HTTP 301. Delivery batch LANDED 2026-07-23 (changelog item 11): lib-hosts VR1 arms populated (tofu-name keys, pinned MACs, D-134 octets, per-DC power/tags, new host_sysid_by_bootmac resolver; dc-selector 48 PASS), D-133 guards in reenroll-hosts + carve-host-interfaces (VR0 flows refuse vr1-), lib-net VID caveat, appendix-A edge-pf entry, DoD items 8-9 + item-4 range correction; gauntlet ALL GREEN (77). Still QUEUED to stage close: set-interface-v4 reload amendment (option, unruled); permission-rules prune. D-135 AMENDED 2026-07-24 (operator-ruled -- DC0/DC1 artifact-path split): DC0 = the full-mirror deployment under test (debmirror RUNNING, measured ~70% / 665 GiB of ~948 at 2026-07-24T15:35Z); DC1 debmirror PAUSED at ~330 GiB (measured: the vcloud uplink was NOT the shared bottleneck -- solo dc0 did not speed up after pausing dc1; kept dormant as the fallback) and switched to an INTERIM apt CACHING PROXY (apt-cacher-ng, proxy mode at utility .4:3142) to unblock the DC1 build now, consumed via juju apt-http-proxy; DC1 proxy REPLACED by a full mirror afterward -- that clause is SUPERSEDED by the D-135 AMENDMENT 2026-07-27: the per-DC split is the DELIBERATE EXPERIMENT (dc0 tests the full mirror, dc1 tests the proxy), there is no fallback and no later convergence, and dc1's partial mirror + apparatus were REMOVED that day. scripts/dc-cache-proxy.sh SHIPPED (site-keyed check/install; harness 15/15) + INSTALLED + verified on the DC1 rack 2026-07-24 (check PASS 8/8 -- apt-cacher-ng 3.7.4-1ubuntu5.24.04.1 active on .4:3142; proxy serves archive + UCA 200). DC1 build consumes the proxy (juju apt-http-proxy=http://10.12.68.4:3142). *DC1 STAGE-5 (proxy-method deploy) STARTED 2026-07-24, then BLOCKED on a bundle-rework gap (committee-reviewed):
juju add-cloud vr1-maas (region 10.10.0.20:5240) + add-credential both DCs; VIP overlay overlays/vr1-dc1-vips.yaml (11 apps -> dc1 bands); phase-4 Step 2.0 (MAAS-cred deployment task) added to the runbook.juju bootstrap fetches the juju agent stream + juju/juju-db SNAPS pre-apt (not covered by the apt proxy; D-135 items 2-3 agent/snap mirror NOT built) -- it works ONLY while the DC1 edge egress is OPEN. Do NOT apply the D-135/D-107 egress-narrowing until AFTER bootstrap (or mirror agents/snaps first). EGRESS MEASURED OPEN 2026-07-27 (Stage-5 grounding audit; previously this entry carried the requirement but no measurement). Probed from the dc1 rack directly with --noproxy "*" so an apt-cacher-ng hit could not fake it: https://streams.canonical.com/juju/tools/ -> 200 (the juju agent stream the bootstrap fetches), https://api.snapcraft.io/v2/snaps/info/juju reachable (HTTP 400 = endpoint answered), http://archive.ubuntu.com/ubuntu/dists/jammy/Release -> 200, ping 1.1.1.1 0% loss, default route via 10.12.64.1. The bootstrap window is therefore OPEN as of this date. This measurement has a shelf life -- re-probe immediately before bootstrap, because the D-135/D-107 narrowing is the closing mutation and nothing prevents it being applied first.docs/audit/committee-20260724-track2-bundle-render.md). The committed bundle.yaml is still VR0's 4-node HYPERCONVERGED layout (machines "8"-"11", tags=openstack); VR1 dc1 is 9 ROLE-SEPARATED nodes (3 control/2 compute/4 storage, tag openstack-vr1-dc1, D-121 Option C). Architecture DECIDED; deploy ARTIFACTS not yet rendered. A 4-lens committee reviewed it and every load-bearing claim was re-verified against the repo -- TWO corrected the brief: designate IS in-bundle (bundle.yaml:812/840/922, D-106 supersedes D-019), so the phase-4 "ships NO designate" (line 192) is STALE and part of the de-stale work; and the inner caller is for_each-keyed (opentofu/vr1-dc1-substrate/main.tf:140) so a 10th node is ADDITIVE (1 add/0/0), not a substrate reopening. Correcting an earlier characterization: dc-ha-scaleup.yaml DELIBERATELY excludes ceph-osd/nova-compute (scale-out, not control containers) -- it is NOT "stale ceph-osd 3"; its real gaps are the dc0-named tokens + that the machines block belongs in bundle.yaml.juju-controller-vr1-dc1, no role tag). Capacity GATE PASS (measured): scripts/dc-dc-whole-host-budget.py 10-node x 2-DC = RAM 854/1024 GiB = 83%, FIT, 170 GiB headroom (baseline 3+2+4 = 838/82%); capture docs/audit/stage5-controller-capacity-20260724.txt. Fork 2 (hand-render NOW + EXTEND provider-bundle-check.py with placement/anti-affinity/count assertions; DEFER a renderer tool) and Fork 3 (per-role MAAS tags + re-enlist-runbook wiring) = OPS, committee-UNANIMOUS, doubt-resolves-DOWN (GA-R3) -- executed under gating; the Fork-3 tag mutation and the 10th-VM apply stay individually operator-gated.ceph-rbd-mirror carried a duplicate bindings: (an orphan at the old line ~907) that YAML take-last absorbed, clobbering its D-108 replication bindings (ceph-local/ceph-remote) + injecting a phantom dashboard binding -> juju deploy reject. Orphan removed; yaml.safe_load confirms the correct set restored; provider-bundle-check.py PASS. Pre-existing latent defect (app was prep-only, never deployed).72cad5f; validator spine c3970a9 (provider-bundle-check overlay-merge + per-DC bands + placement invariants, harness 15/15); the 9-node role-separated bundle RENDER (bundle.yaml machines "0"-"8" role+DC tags, all to: re-homed, nova-compute 3->2, ceph-osd on 4 storage, ovn-chassis -> the 2 dc1 compute provider MACs; new overlays/vr1-dc1-machines.yaml retag overlay; dc-ha-scaleup tokens rendered to logical 0/1/2). VALIDATED: base + full dc1 deploy input (3 overlays, --dc vr1-dc1) PASS, harness 15/15, gauntlet ALL GREEN (78), repo-lint 0-fail. Deployment-expansion review (3 read-only reviewers) recorded docs/audit/stage5-expansion-review-20260724.md; it confirmed relations/bindings/subordinates HOLD under role-sep (uniform 6-NIC, SEC-011) and no channel/pin change is needed.vault-hacluster cluster_count:3 but NO vip -> vault:secrets consumers hit a unit address). D-020 (not D-036) lists vault among the apps carrying provider+metal VIPs. Needs a per-DC VIP in the symmetric overlay shape (ruling 3; metal-admin+metal-internal only, no provider leg; octet .61 band-legal per D-134-amended .50-.99) + the vault-HA VERIFY-LIVE. OPS (D-020 conformance repair), not a new D; (b) address family: RULED 2026-07-25 (D-101 CONFIRMED via a dated ruling note under
vip per app per DC before deploy -- no longer deferrable. Octavia family remains an OPEN sub-ruling (see D-136 open question 2). (c) per-DC artifact shape: RULED 2026-07-25 (ruling 3) -- bundle.yaml becomes VIP-free (topology only); every per-DC value arrives via overlay, dc0 included; dc0's VIPs extract from base into overlays/vr1-dc0-vips.yaml (2 commits: neutral MOVE proven via provider-bundle-check --overlay --dc vr1-dc0, then the dual-stack ADD; MIGRATE the inline VIP-line operational comments). FOLLOW-UP scripts/gates (2026-07-25 pack-review blockers fold in here) -- preflight.sh runs the checker BARE (must pass the dc1 overlays
juju bootstrap needs a controller-tag constraint; phase-4 Step 4 de-stale. VERIFY-LIVE gates: vault 1.8 HA-on-MySQL@3 (D-121), hacluster cluster_count:3 semantics, Ceph pools size=3/min_size=2/failure-domain=host, machines-block overlay merge (--dry-run); keystone policyd-override (RULED 2026-07-25): juju status shows PO: on EVERY keystone unit after scale-up, never PO (broken): -- a broken override is atomic (whole policy discarded, silently reverting every tenant domain-manager to a plain user; protects D-051/D-064).netbox/sandbox-fidelity-check.py (blind to D-124/D-134 today). The v6 subcarve is NOT an open gate (D-111 ADOPTED; values imported both DCs). The target OPERATING MODEL (NetBox draft -> approve-on-rendered-diff -> apply -> bounded self-healing) is PINNED for a planning session (docs/audit/operating-model-plan-20260725.md), not ratified.docs/audit/decision-recon-20260725.md). office1-netbox holds the entire IP PREFIX layer (all 12 planes + transit + uplink + v6 GUA/ULA), DRIFT-FREE across netbox/lib-net/artifacts, and is correctly VLAN-empty (untagged-per-fabric, D-133 -- VLANs are a MAAS construct, not NetBox). The GAP is sub-prefix: the D-134 per-DC bands + the 33 VIP addresses live only in prose/overlays, NOT the apex -- the exact thing the NetBox->deploy coupling must close. Node/MAC/power config is drift-free; Fork-3 role tags + the 10th VM are pending gated (deploy-blocking); live MAAS tag state is NEEDS-LIVE-VERIFY (no tag column in the discovery capture). Decision-status integrity flagged STALE surfaces (fixes QUEUED, operator-deferred): this doc's section 4 "RULED-BUT-NOT-BUILT" is comprehensively stale (HIGH -- 4/5 bullets contradicted by section 1 + on-disk; do NOT read section 4 as current), G14 SEC count 12 -> 15, phase-4 runbook:192 "no designate". The reconciliation backlog (the "what else to queue" answer, P0-P2) is in the record; extend netbox/sandbox-fidelity-check.py FIRST (it gates every apex recon).docs/audit/maas-admin-recovery-20260725.txt). The 2026-07-25 ledger claim that MAAS web-GUI login "does NOT exist / was never minted" is FALSIFIED: the admin superuser's password was minted 2026-07-13 by site-headend-install.sh:452 and is MEASURED working (login 204 with the stored value vs 400 wrong-password control; has_usable_password=True for all accounts). The real defect was narrower -- it sat root-only at voffice1:/root/maas-secrets/admin.pass, un-consolidated (the SEC-009 miss class). Operator-ruled scope "New account + consolidate only": admin password CONSOLIDATED byte-identical (sha256-verified) to ~/vr1-office1-creds/maas-admin-password -- the VM copy REMAINS source-of-record; any future rotation must update both copies or the VM copy becomes a stale trap; NEW superuser operator minted for human GUI login (password via stdin, never argv). NO existing password rotated -- juju-vr1-dc0/dc1 stay random+unstored BY RULING (their SEC-018/019 API keys are load-bearing for the bootstrap below, and passwords are independent of API keys). MAAS + maas-init-node are MAAS-internal, untouched. creds-audit now CLEAN on all THREE sites (was RED: dc0 3
admin doubles as the automation identity (its API key drives the maas admin CLI profile across 19 call sites in 4 scripts), so automation never needed the password and nothing forced it into the folder -- proposable as the next-free D-number [ARCH], NOT assigned.juju bootstrap (controller tag, egress OPEN) -> deploy. Separately: DC0 full-mirror completes -> reachability gate -> stage close-out.runbooks/dc-dc-phase3-maas-enlist-deploy.md:472-485), four are met and captured (nodes READY+carved+tagged, six planes per node, provider NIC raw, PXE v4). The remaining two are BOTH defective as written:
dc-mirror.sh check asserted only that last-sync.status EXISTS, printing its contents behind an unconditional OK, so it could not fail. MEASURED both racks 2026-07-27: dc0 read FAIL ... ubuntu=255, dc1 read a FOUR-DAY-STALE RUNNING 2026-07-23T21:49:35Z left by the debmirror the D-135 amendment killed -- both printed OK and PASSED. FIXED (status word now case-analysed; RUNNING is an explicit UNKNOWN cross-checked against the unit; absent and unrecognised both refuse; harness 24/24, was 19). Re-run now reports the true state: dc0 FAIL / dc1 FAIL, capture docs/audit/stage4-mirror-gate-20260727.txt. Substance: dc0's mirror CONTENT is intact (949G+342M, all three dists + pool) and only the overnight INCREMENTAL failed, on a transient upstream 500 read timeout for dists/jammy/Release; dc1's PROXY -- its RULED artifact path per the D-135 amendment -- checks PASS genuinely (apt-cacher-ng on .4:3142 serving archive + UCA 200), while its dormant fallback debmirror is dormant only ACCIDENTALLY (timer enabled with EMPTY next-elapse because the unit sits in failed/Result=signal, so a reset-failed re-arms a 330G->949G pull; the ruling is not enforced by anything). The node-side half of bullet 5 is SPLIT OUT to new gate row G17 by operator ruling 2026-07-27 (GA-R5, utterance quoted in the G17 row) -- NOT closed conditionally, which GA-R6 E3 forbids. Nodes are powered off in Ready by the READY-handoff ruling, so no node-side probe can run inside Stage 4; G17 carries it to Stage 5 first boot. What REMAINS in Stage 4 for bullet 5 is the RACK-side half only: the artifact source answers on its own address with an attested-current sync. dc1's proxy already satisfies that (PASS). dc0's sync was RE-RUN 2026-07-27 (gated) and completed in 20s: OK 2026-07-27T08:43:46Z ubuntu=0 uca=0; the fixed dc-mirror.sh check dc0 now reports PASS -- a PASS that is trustworthy precisely because the same check FAILED on the same rack twenty minutes earlier. BULLET 5 IS THEREFORE MET FOR BOTH DCs (dc0 mirror PASS, dc1 proxy PASS), with the node-side half at G17.dc-cache-proxy.sh did not own its utility net layer, so three files named dc1-mirror-* were load-bearing FOR THE PROXY -- the recorded removal procedure would have killed the dc1 apt path, and the proxy could not be rebuilt afterward without reinstalling the whole mirror apparatus (enabled daily timer included), so teardown-and-rebuild did not converge. FIXED FIRST: the net layer moved into dc-cache-proxy.sh as <site>-cache-proxy-net.service + apply helper + its own resolved drop-in (harness 20/20), installed on the dc1 rack, and INDEPENDENCE PROVEN by deleting the live address and default route outright and applying the proxy-owned unit ALONE -- check dc1 PASS with dc1-mirror-net.service disabled. THEN removed wholesale: all mirror units, the timer, the sync helper, the nginx vhost, the resolved drop-in, the debmirror package, and the 330G of partial mirror (33,037 files, rm -rf 47.6s; rack disk 339G -> 9.0G). Zero mirror-named residue; proxy PASS throughout, never down. nginx LEFT installed deliberately (generic, now serving only its stock vhost). Capture docs/audit/dc1-mirror-teardown-20260727.txt.docs/dc-dc-deployment-workflow.md:206, docs/dc-dc-buildout-design.md:120, runbooks/dc-dc-phase3-maas-enlist-deploy.md:412,484, runbooks/dc-dc-phase4-juju-bundle-per-dc.md:26) and needs a DOCFIX re-expressing it as MAAS-hierarchy time verification before it can be checked at all.docs/audit/stage4-carve-residue-cleanup-20260727.txt): 108 fabrics -> 17. Deleted subnet id=8 192.168.1.0/24 (the superseded OPNsense FACTORY LAN -- both edges were re-addressed to 10.12.4.1 / 10.12.64.1), then fabric-5, then the 90 auto-created empties (ids 6..95). The audit caught an ORDERING dependency the original flag did not state: the 192.168.1.0/24 subnet was the ONLY occupant of fabric-5, so deleting the fabric first would have cascaded the subnet away. Emptiness was PROVEN per fabric (zero subnets, zero ipranges, zero node interfaces across all its VLANs) against a fresh occupancy snapshot taken immediately before the batch, and the cascade check was re-run AFTER -- the inverse of the 2026-07-21 pod-delete incident, where the association check ran too late and cost 9 machine records. POST-STATE: 18 Ready + 2 Deployed office1 guests, all 18 still power_type=virsh, 7 interfaces each (6 flat planes per D-133 + br-ex), ZERO orphaned interfaces, placement tags 9+9 intact, both artifact paths re-verified PASS. The 17 survivors are all load-bearing and enumerated in the capture (office1 base+GUA+compose, both transits, both metal-admin/boot fabrics, libvirt default, and the 10 named plane fabrics). dc1's leftover nginx also removed the same day (purged nginx + nginx-common, /etc/nginx gone, nothing listening on :80, proxy still PASS on :3142; it had been a MIRROR prereq only -- dc0 keeps its nginx and is unaffected).set-interface-v4 RELOAD AMENDMENT RULED + SHIPPED 2026-07-27 (operator utterance "a" to the three presented options; OPS under GA-R3, governed by D-113 -- no new D-number). --commit now runs configctl filter reload and reads the automatic NAT rules back via pfctl -s nat, so the appendix-A "edge passes no LAN traffic after v4 addressing" defect cannot be reached through the normal addressing path. Prose-only prevention was rejected on evidence: DoD item 8 existed and did NOT fire at dc1. PLACEMENT IS THE SUBSTANCE: the reload runs over a FRESH connection to the post-move address, NOT on the next line of the interface reconfigure heredoc -- re-addressing the interface you arrived on drops the session DURING the reconfigure, so a reload there would never execute in exactly the dc1 scenario that caused the defect. The NAT read-back is REPORTED not gated (a first addressing may legitimately have no gateway yet); the hard gate stays address-on-the-kernel. Harness 59/59 (was 53), cases 15-15f. No live edge touched -- both DC edges are already addressed, so this affects the NEXT one.docs/ 25 -> 18, stage record docs/archive/stage-records/vr1-stage4-record.md; 12 stale docs/changelog-* paths in live surfaces rewritten to the archive, including some already-dangling from the G12 close). Skill sweep done -- three new INVARIANTS folded in (per-DC artifact delivery is a per-DC STRATEGY and a utility service owns its own net prerequisites; Stage 4 hands off READY nodes not deployed ones; "a checker that cannot fail is not a gate", with the assert-on-content / enumerate-what-exists rules) plus two routing rows. Dated snapshot REGENERATED as .claude/skills/openstack-cloud-ops-consolidated-20260727.md (1589 lines, ASCII/LF verified) and the superseded 20260725 one REMOVED -- a stale snapshot being uploaded is the exact failure docs/audit/skill-divergence-20260725.md records, and it is a derived artifact regenerable from any commit. GA-R7 memory review done: NO new memory (everything durable graduated to the skill/repo, which is the correct GA-R7 outcome); both existing entries verified against the repo and extended with the two reasoning traps this stage produced. Gauntlet ALL GREEN (81), repo-lint 0-fail. REMAINING: the operator-gated merge of dc-dc-stage4-phase3-maas-deploy to main as a MERGE commit (not squash), then branch retirement (local + remote). Every substantive in-stage item is closed or split to gate row G17.lib-hosts.sh, D-134 statics perfect, 17 fabrics, zero orphaned interfaces, both artifact paths serving, gauntlet ALL GREEN (81). What is NOT ready is the layer between the substrate and the deploy. Headline blockers, each measured: the Office1 headend clone -- the D-128 Plane-2 host Stage 5 EXECUTES from -- is 105 commits behind main on a branch retired four days ago, with both dc1 overlays ABSENT and bundle.yaml still the VR0 4-node hyperconverged layout; the openstack client is installed on NEITHER host while ten Stage-5/6/7 scripts invoke it; the D-104-amendment 10th controller VM is UNAUTHORED (no OpenTofu resource anywhere) and its MAAS tag does not exist, while dc1 has exactly 9 Ready nodes for a 9-machine bundle; ceph-osd targets /dev/vdb and every node has only vda (found independently by two lenses using different methods); per-role MAAS tags are consumed by the machines block and authored nowhere; there is ZERO IPv6 in the DC substrate while D-101 RULED dual-stack for both DCs this deployment; and Step 4's "follow phase-01 verbatim" points at a runbook whose VIP guard ABORTS for dc1 (and will abort for dc0 once the ruled VIP extraction lands). Several gates that should have caught these CANNOT FAIL -- repo-lint returns PASS (0 fail, 0 warn) over ZERO files on a one-character typo of its flag or root; provider-bundle-check passes decorative HA because cluster_count is checked nowhere in scripts/ or tests/; preflight's aggregator ignores any sub-gate exit code that is not 1 or 2; and P3 verified ZERO of 33 charm-channel pins because juju is not on the host's PATH. DELIVERABLES (this doc stays the status authority; those are findings and questions, not status): ordered precondition checklist docs/audit/stage5-readiness-20260727.md (READ FIRST); verbatim committee record docs/audit/stage5-committee-raw-20260727.md; 11 Stage-5-blocking + 4 standing questions awaiting GA-R5 rulings, one exchange each, in docs/audit/queued-rulings-20260727.md -- NONE are adopted; measurements docs/audit/stage5-live-measurement-20260727.txt; charter docs/audit/stage5-grounding-audit-scope-20260727.md. The DOCFIX remediation batch (21 items, Phase 3 of the readiness doc) is LOGGED NOT EXECUTED -- nearly every runbook fix interlocks with an unanswered ruling, so landing them now would encode assumptions. The sole mechanical fix taken this session is the G3 row correction above. RULINGS IN PROGRESS (operator returned 2026-07-27; ONE exchange each per GA-R5). R1 RULED 2026-07-27 -- exact utterance "Add an OSD volume to node-vm (Recommended)", recorded as a D-121 AMENDMENT (docs/design-decisions.md is the ruling authority). The ceph-osd data device becomes a REAL second block device on the four storage nodes per DC -- 8 volumes, not 18, since ceph-osd is placed on machines 5-8 only. D-121's own capacity re-validation had already budgeted it ("Ceph disk re-run for 4 storage/DC = PASS 5.31 TiB"): the disk was budgeted and never built. The APPLY is a SEPARATE operator-gated step and is NOT authorised by this ruling -- four preconditions are recorded in the amendment, including that it deliberately spends the inner roots' ZERO-DIFF property, and that whether re-commissioning preserves the D-134 statics and pinned MACs must be verified BEFORE the apply (the 2026-07-20 MAC-regeneration incident is the precedent). R2 RULED 2026-07-27 -- exact utterance "Carve v6 and deploy dual-stack as ruled (Recommended)", recorded as a D-101 RULING NOTE (re-confirmation, no amendment; design-decisions.md is the authority). CORRECTION, same session, before dependent work: the v6 literals are ASSIGNED, not pending. This entry first claimed D-101's "Remaining open item" (org ULA /48 + per-DC GUA carve) had become a Stage-5 precondition. It had not: D-111 ADOPTED them 2026-07-11 and the apex carries them -- ULA fd50:840e:74e2::/48 with DC0 planes at :220/:221/:230/:240/:250::/64 and DC1 at :320/:321/:330/:340/:350::/64, GUA provider-public DC0 2602:f3e2:f02:10::/64 + VIP f02:11::/64 and DC1 2602:f3e2:f03:10::/64 + VIP f03:11::/64 (measured from netbox/draft/vr1-office1-current-20260725.json; line 264 of THIS document already recorded the apex as holding the v6 GUA/ULA). Sub-question R2a is WITHDRAWN as never open. The real precondition is PROPAGATION and needs NO ruling: the ratified values are absent from scripts/lib-net.sh (no v6 arm at all) and from MAAS (no v6 on any of the 12 DC plane fabrics), and both are mechanical copies from an authoritative source -- moved to the Phase-3 mechanical batch. R9 and R11 still inherit dual-family; the L3-9 overlay collision must still be reconciled BEFORE either authority location is populated (the dangerous merge order is the one that PASSES -- it silently drops every v6 leg); R8 is still NOT resolved. NEW DOCFIX-class finding: D-101's own "Remaining open item" paragraph is STALE (still says "pending NetBox assignment" for literals D-111 adopted on 07-11) -- that stale prose is what caused this error, and it is queued in Phase 3. Remaining: R3-R11 blocking, R12-R15 standing, in docs/audit/queued-rulings-20260727.md. UNMEASURED-GAP SWEEP 2026-07-27 (operator challenge: did the committee actually measure live state, or take shortcuts?). Register: docs/audit/stage5-unmeasured-register-20260727.md. Honest accounting -- the apex was never polled by ANY lens (lens 2 declared it out of scope, correctly and without overclaiming, but it is a key system it had reach to), and the synthesis then used a two-day-old repo dump instead of polling live. The sharpest instance of the general problem: lens 6 declared juju restore-backup unverifiable because "no Juju client exists on this host" while juju 3.6.27 was installed on voffice1 and lens 5 was successfully running juju help against it in the same session. Fourteen deferred items were CLOSED by the sweep. Consequential results: the LIVE apex is identical to the dump (139 prefixes / 103 IPv6, zero drift -- the conclusion was right, the method was not); juju restore-backup DOES NOT EXIST on 3.6.27, promoting L6-14 from RISK to CONFIRMED DEFECT and making Stage 6 Step 9's D-104 restore drill unsatisfiable as written; dc0's compute provider MACs measured, CONFIRMING L3-8; curl present on both racks so L4-13's hole is theoretical; and the pinned charm channels DO resolve (2024.1, 2.4, squid all present via juju info on voffice1), proving P3's 33 warns are purely the missing juju binary on vcloud. NEW FINDING: preflight's "MAAS unreachable" is a MISDIAGNOSIS -- the maas binary is simply ABSENT on vcloud. With the absent juju and absent openstack, THREE separate preflight/deploy failures on this jumphost are all "the client is not installed" and each is reported as something else. SECOND SWEEP 2026-07-27 (operator: "Close the remaining gaps") -- U15-U17 closed. U15 voffice1 transit addressing is REBOOT-DURABLE (positive result): live enp2s0 172.31.0.1/30 + enp3s0 172.31.0.5/30, both netplan-persistent via /etc/netplan/60-transit.yaml and 61-transit-dc1.yaml; no leg row is owed. U16 RETROFIT_WAIT=30m has NO recorded provenance -- traced to a single bulk commit with no rationale, and NO constant anywhere in scripts/ is documented as nested-virt calibrated, so it is an inherited default that has never been validated against the depth-4 nested I/O it will run on; separately, that script's preconditions require BOTH the openstack and juju clients and NO host has both. U17 the DC data path carries NO IPv6 at any layer -- on BOTH racks, zero global v6 on any plane bridge, no v6 default route, accept_ra=1 with nothing arriving. With the MAAS and node measurements that is a THREE-LAYER confirmation, and it WIDENS the R2 propagation task: the rack bridges need v6 too, not just MAAS. STILL OPEN, with cause: the two DC edges' own interface-level v6 config -- dc0 is blocked by SEC-021(a) (no opnsense-api.txt in ~/vr1-dc0-creds/, visible on disk; the re-mint is a live edge mutation deliberately excluded from the 07-27 batch) and dc1's API is not reachable from vcloud (measured timeout; the path runs from the rack, where the creds are correctly not staged per SEC-015). U17 already answers the substantive question from the rack side. repo-lint/gauntlet ON voffice1 remain deliberately deferred until precondition 0.1 advances that 105-commit-stale clone -- running them today would measure a stale tree. R3 RULED 2026-07-27 -- exact utterance "Raise the two lagging segments to 9000 (Recommended)", recorded as a D-101 RULING NOTE (D-102 is merged into D-101 and directs amendments there). The question was re-framed by measurement before it was put: scripts/dc-dc-mtu-geneve-budget.sh had NEVER been run to a recorded verdict despite D-101 calling the measured underlay MTU a Phase-0 gate. Run both ways this session (capture docs/audit/mtu-budget-20260727.txt): underlay 9000 -> tenant MTU stays 1500; underlay 1500 -> tenant MTU 1444 requiring ovn geneve + tenant-network + amphora to agree permanently, which NOTHING in this repo checks. Measured underlay: every vcloud MESH leg is ALREADY 9000 including the inter-DC virbr5, as are all six plane bridges on both racks; the four 1500 legs are the D-125 SIMULATED-ISP uplinks and must STAY 1500. Exactly TWO segments lag -- the rack transit NIC enp1s0 in both containment VMs, and all 17 MAAS VLAN records (the silent one: MAAS renders VLAN MTU into node netplan, so a jumbo bridge under a 1500 record still yields 1500 node interfaces). Coupled to R2: the 56-byte budget overhead is the IPv6 figure and applies BECAUSE dual-stack was ruled (v4-only would have been 42 / 1458). Execution is a SEPARATE gated step; the verification owed is a BEHAVIOURAL large-frame test with DF set across the inter-DC path, not a reading of interface MTUs.docs/audit/outer-plan-20260719-postA-converged.txt). Deploy step B (bootstrap) COMPLETE 2026-07-20 in the same logged dc0-deploy window: transit reach established (voffice1 holds 172.31.0.1/30; reach = ssh -J voffice1 w/ dc0 key -- no vcloud host leg, item-20 disposition), rack ENROLLED to the Office1 region, node-host ready (libvirt + nested KVM + inner pool), SEC-010 applied+verified BOTH transit ends (row CLOSED), OPNsense 26.7 nano base staged (operator ruling; step-C boot REVALIDATES the D-112/D-113 path on 26.7). Named gate check EXIT 0: docs/audit/stepB-check-20260720-final.txt. Deploy step C (inner apply) COMPLETE 2026-07-20: executed FROM voffice1 (D-128 Plane 2 -- tofu 1.12.4 + repo clone + dc0 key staged there), 28/28 resources, inner plan CONVERGED zero diff (docs/audit/inner-converge-20260720-stepC.txt); 10/10 domains RUNNING inside vvr1-dc0 (9 nodes + edge); edge = fresh 26.7 nano, serial log at the FreeBSD login prompt (D-112 boot path first-datapoint PASS on 26.7). The INNER tfstate lives ON voffice1 (vr1-dc0-substrate/terraform.tfstate -- new state-of-record location; add to the site backup set). ACTIVE gate: G10 remaining. Edge bootstrap DONE 2026-07-20: D-112(c) console bootstrap complete (key-only root SSH proven) and the D-113(a2) API key minted via the vendor model -- GET core/firmware/status 200 with CORE_ABI 26.7, the first proof the API path works on 26.7. Measured: edge vtnet0 = LAN (provider-public), vtnet1 = WAN; edge still on its FACTORY LAN 192.168.1.1/24. Rack legs 10.12.4.2/22 + 10.12.8.2/22 added INTERIM (non-persistent ip addr; script support is a queued finding), plus a temporary 192.168.1.2/22 to reach the factory LAN. D-125 egress isolation gate: PASS / CLOSED 2026-07-20 (executed as written -- throwaway VM on br-vr1-dc0-wan; two identical consecutive runs: gateway ping 0, internet ping 0, curl 1.1.1.1 301, curl archive.ubuntu.com 200). Bridge-in is PROVEN end to end and the double-NAT fallback is NOT needed. Captures: docs/audit/d125-egress-gate-20260720{,-matrix}.txt. One earlier run failed ICMP-to-internet on the same path and is recorded UNEXPLAINED in the session changelog (start there if a DC edge shows first-boot egress failure). Edge ADDRESSED 2026-07-20 via the NEW operator-ruled opnsense-set-interface-v4 pair (D-113 amendment re-measured and still true on 26.7 -- base-iface addressing is not REST-covered): WAN 172.30.2.2/24 + default gw 172.30.2.1 (was dhcp, which could never work on a /24 with no DHCP server), LAN 192.168.1.1/24 -> 10.12.4.1/22 (ruled provider-public gateway). Verified on the kernel; the edge itself egresses to 1.1.1.1 at 0% loss, and the API answers at the new LAN address. Interim bootstrap address removed; virbr5 now carries only the ruled 10.12.4.2/22. D-129 edge profile APPLIED 2026-07-20 on 26.7 (operator-ruled): expose_qga_channel shipped in modules/opnsense-edge (opt-in, default OFF; dc0 true) and applied as an IN-PLACE domain update; os-qemu-guest-agent + os-iperf installed for real and the agent ANSWERS -- guest-ping -> {"return":{}} and domifaddr --source agent reports both legs. Note this run also exposed and fixed a false-success bug: opnsense-plugins.sh apply had ALWAYS dry-run (see session changelog item 12), so any prior "applied" claim from that script is void. Step D part 1 DONE 2026-07-20: rack registered (7chphy, rackd running), metal-admin dynamic range 10.12.8.100-.200 created (operator-ruled D-120 inheritance), VLAN 5005 dhcp_on=true primary_rack=7chphy verified by read-back. INCIDENT RESOLVED 2026-07-20 (operator-approved region restart): dhcpd now RUNNING on both controllers (verified by process, not service status), and all 9 DC0 nodes ENLISTED in MAAS with shapes exactly matching D-121 Option C (3x16cpu/64GiB + 2x12cpu/48GiB + 4x8cpu/24GiB) -- docs/audit/stepD-enlistment-20260720.txt. The G10 depth-4 nested boot gate is therefore PASS: node VMs inside vvr1-dc0 PXE-booted from the Office1 region across the transit and run MAAS's ephemeral kernel. The incident as originally found: MAAS 3.7 drives DHCP via Temporal, and Temporal is wedged on the region ("Not enough hosts to serve the request", 2807 retries), so no dhcpd runs on EITHER controller -- including voffice1 itself, whose compose net reads dhcp=True with no dhcpd process. Predates and is NOT caused by this deploy (almost certainly since the 2026-07-17 host reboot); unnoticed because both Office1 VMs were already Deployed. Any "Office1 MAAS DHCP working" claim is currently FALSE. Proposed gated remedy: restart MAAS on the region -- DONE, and it fixed BOTH sites, confirming a single root cause. Details + the queued detection-gap finding (cloud-assert trusts MAAS's self-report and missed a dead DHCP server): session changelog items 13-14; appendix-A entry queued. Step D part 2 BLOCKED on a ruling (2026-07-20): the nine nodes fell back to New with no power_type -- commissioning cannot finish without power control. New root opentofu/vr1-dc0-maas/ is shipped and its plan is clean, but the apply FAILED: Failed talking to pod: Failed to login to virsh console. MEASURED cause -- the MAAS snap is confined, gets Permission denied on /var/run/libvirt/libvirt-sock, and snap connections maas lists NO libvirt interface, so a LOCAL qemu:///system pod is IMPOSSIBLE with snap MAAS. This refutes the mechanism stated in D-123 Model B and in modules/maas-vm-host's header (intent survives, mechanism does not); both need an amendment once the replacement is ruled. The qemu+ssh replacement was then wired with an operator-ruled DEDICATED key and PROVEN reachable from both snaps -- but the pod apply failed again, finally on domblkinfo ... missing storage backend for 'volume' storage, REPRODUCED LOCALLY on the rack with an active pool. So MAAS virsh pods are incompatible with modules/node-vm's pool+volume disk refs; the pod would require converting node-vm to file-path disks and re-applying all nine domains. The pod is however UNNECESSARY -- its D-103 job was DISCOVERY, already done via PXE -- and per-machine power_type=virsh is MEASURED WORKING (query-power-state -> {"state":"off"} on the canary), which is also the Roosevelt shape (per-node IPMI). RULED 2026-07-20: per-machine virsh power. STEP D IS COMPLETE: shipped scripts/maas-node-power.sh + harness (24/24; gauntlet now 72 ALL GREEN), MAC-matched (MAAS renames machines at enlistment), dry-by-default, each write verified by a real query-power-state. All 9 nodes have power (docs/audit/stepD-power-20260720.txt) and commissioning works end to end -- 3 Ready / 6 Commissioning at time of writing, shapes still exact to D-121 Option C. opentofu/vr1-dc0-maas/ is retained but UNUSED (the pod route is refuted); retire-or-keep is a stage-close question, as is the D-103/D-123 amendment text. Session changelog items 15-17. INCIDENT 2026-07-21 (pod-delete cascade): RESOLVED SAME-DAY -- all 9 nodes READY again. During the operator-ruled retire of opentofu/vr1-dc0-maas, deleting the stale pod object (id=4 vr1-dc0-inner) cascaded to the nine machine records the failed 2026-07-20 pod refresh had silently linked to it -- the association check was run AFTER the delete (agent process error, owned; capture docs/audit/incident-20260721-pod-delete-cascade.txt). Substrate was measured intact throughout (10/10 domains, MAC pins config-carried, rack services untouched); only MAAS records were lost. Operator-ruled recovery executed immediately: virsh power-on -> PXE re-enlist (9/9 in ~2 min, pinned MACs) -> maas-node-power.sh dry+commit (9/9, power verified) -> re-commission -> ALL 9 READY in ~3 min, shapes exact to D-121 Option C, power=virsh (docs/audit/incident-20260721-recovery-verify.txt; note the MAAS hostnames are NEW random names -- any doc quoting the old ones is history). The retire-fully ruling is now FULLY EXECUTED: repo root removed, stale pod gone, voffice1 tfstate remnants + the SEC-013 on-disk key file deleted (absence verified; SEC-013 row narrowed to CLI-profile-only). Lesson shipped to appendix-A: read a pod's machine list BEFORE vm-host delete; non-empty = STOP. COMMISSIONING RESOLVED 2026-07-21: all 9 nodes Ready (logged window ops-commissioning-diag; adjudication docs/audit/commissioning-diag-20260721.txt; session changelog 2026-07-21). TWO stacked faults, both measured: (1) the 2026-07-20 in-place serial-console apply REGENERATED all 9 node NIC MACs (tofu-reported 0/9/0 in-place), so MAAS's records went stale and every post-apply boot was an unknown node -- no PXE event, no tag kernel_opts, silent 30-min timeout; repaired operator-ruled via per-machine boot-interface MAC update (mark-broken/update/mark-fixed where needed), read-back verified 9/9. (2) Beneath it, the MAAS 3.7 RACK-ONLY agent resolver SERVFAILs every query on an internet-isolated rack (walks public root hints even for its own authoritative maas-internal zone; ignores resolv.conf), so cloud-init's cloud-config-url never resolved and nodes booted to a login prompt without ever fetching commissioning scripts. Office1/VR0 were immune (co-located region BIND owns node DNS) -- this surface is FIRST EXERCISED in VR1; LP report queued. Operator-ruled workaround, live and proven: dc0-node-dns.service on the rack (dnsmasq on virbr2 alias 10.12.8.3 forwarding to region BIND over the rack's OWN transit connection; SEC-010 re-verified enforced and untouched) + metal-admin subnet dns_servers=10.12.8.3, allow_dns=false. PROOF: canary Ready in ~3 min after seven consecutive 30-min failures, commissioning scripts visible on serial; fleet of 8 re-commissioned concurrently, ALL 9 READY in ~4 min, shapes exact to D-121 Option C. Committee record closed by addendum (its mechanisms were wrong; its instrument found the cause). D-131 PARTIALLY RULED (sub-1 RULED 2026-07-21: the forwarder is the STANDING per-DC pattern, repo-carried + part of DC standup definition-of-done; sub-2 RULED 2026-07-21: metal-admin-only scope; sub-3 RESOLVED 2026-07-21 by measurement: no dhcpd option-6 defect, stale read, no second LP; sub-4 OPEN + pinned DNS architectural review -- status line in design-decisions.md is the authority). SEC-014 OPENED (rack cluster secret exposure during diagnosis). Queued delivery: incident docs SHIPPED 2026-07-21 (two appendix-A entries, platform-traps 1e second corollary + index row, LP draft docs/audit/lp-draft-20260721-maas-agent-resolver.md -- operator to file). Still queued: stale pod object cleanup (stage close, with SEC-013). Forwarder + rack-legs persistence SHIPPED 2026-07-21 as scripts/dc-rack-net.sh (D-131 sub-1 delivery; harness 14 cases; gauntlet 74 ALL GREEN) and INSTALLED on the rack 2026-07-21 (operator-approved): install EXIT 0, self-check PASS 10/10 (docs/audit/dc-rack-net-install-20260721.txt), post-install behavioral probe = forwarder answers authoritative maas-internal SOA. The three rack bridge legs are now reboot-persistent (dc0-rack-legs.service); the hand-placed interim state is fully superseded. MAC pinning SHIPPED 2026-07-21 (54 MACs measured via virsh domiflist + pinned in modules/node-vm + vr1-dc0-substrate; harness 15 cases; gauntlet 73 ALL GREEN) together with an operator-ruled power-ownership guard (ignore_changes = [running] -- MAAS owns node power; the pin-adoption plan had carried 9 out-of-band power-ons). Verification plan captured (docs/audit/inner-plan-20260721-macpin.txt: 0/9/0, 54 mac adoptions, ZERO replaces); guarded re-plan zero power flips (docs/audit/inner-plan-20260721-macpin-guarded.txt); APPLIED 2026-07-21 (operator-approved) from voffice1 via saved plan, exact 0/9/0, convergence zero diff (docs/audit/inner-apply-20260721-macpin.txt); post-apply verified all 9 domains still shut off, MACs unchanged. Node NIC MACs are now config-pinned end to end. History of the diagnosis (superseded; kept for the audit trail): the 2026-07-20 state read "3 nodes Ready, 6 timed out." Established: PXE and the ephemeral handoff WORK, and the ephemeral OS boots with working networking (nodes hold leases and do NTP to the rack) -- it simply never completes. Ruled out by measurement: memory, rack boot-image sync, DHCP, and node shape. The node->region path (SEC-010) is SUSPECTED but UNCONFIRMED (those rules carry no counters). A serial console was added to modules/node-vm and applied in-place to all 9, but the logs stay empty -- firmware writes to VGA, so serial alone does NOT make a PXE-booting node observable (correction queued). Two of the agent's own isolation experiments were INVALID and must not be cited (other nodes were still running; and a re-commission did not restart MAAS's timer) -- so contention remains a LIVE hypothesis, not a refuted one. CLEAN experiment now RUN (item 20): a genuinely isolated node still failed at 1770s (~29.5 of 30 min) -- that refutes CONTENTION but is consistent with INHERENTLY SLOW, and the batch pattern 3-pass/6-fail-at-the-mark is the signature of a MARGINAL 30-min timeout over slow depth-4 nested I/O. 3 nodes reached Ready on this exact rack/subnet/metadata path, so metadata is NOT globally broken (rack :5248 up, rack->region 301). LEADING HYPOTHESIS + cheap decisive test, needing an operator decision (MAAS-wide config): raise node_timeout and commission one node. DIAGNOSTIC COMMITTEE run 2026-07-20 (4 independent reviewers, docs/audit/commissioning-committee-20260720.md) REFUTED that hypothesis 4/4 -- 30 min of SILENCE is a hang, not slow progress; a longer clock cannot fix a hang, and the proposed one-node test was CONFOUNDED (changed timeout + concurrency together). Post-committee reads: MTU branch EXONERATED (metal-admin MAAS VLAN MTU is 1500, so the guest never goes jumbo); region healthy at rest. STILL-LIVE causes, both needing observation DURING a run: region Temporal starvation, and a commissioning-only script hang on nested-virt hardware. Decisive gated test (supersedes node_timeout): one commission with console=ttyS0 on the kernel + a full-window, lease-IP-keyed capture on virbr2 + enp1s0. Failed commissioning is re-runnable; nothing is lost. The committee record (docs/audit/commissioning-committee-20260720.md) is the durable authority for this diagnosis and its ranked live hypotheses. STEP E (netem) DONE 2026-07-21 -- G10 CLOSED. The sudo mechanism: operator-ruled scoped NOPASSWD, fragment SHIPPED (gauntlet 75 ALL GREEN) and INSTALLED on vcloud (operator-run; verified 0440 root:root, byte-identical, sudo -n -l exit 0 -- docs/audit/netem-sudo-install-20260721.txt). Wiring: modules/ netem-link amended with a LOCAL execution mode (empty ssh target = bare sudo tc; the module's Office1-era always-SSH assumption is refuted by D-128 -- the outer root runs ON vcloud, and a self-hop would have needed a new standing credential; NEW tests/netem-link harness 12 cases, gauntlet 76 ALL GREEN). Target = the dc0<->dc1 mesh leg virbr5 (re-measured at wire time via virsh net-info; the runbook Step-11 text targeting the office1 leg is a FLAGGED divergence, DOCFIX queued -- that leg now carries the live rack<->region transit, netem there would perturb operations). The wire plan came back 1/1/0 = STOP (section 5): the extra in-place change is the office1 edge picking up D-129's channels = [] state-schema reconcile (commit f5510c7; benign in config terms but an in-place update against the LIVE unpinned-MAC office1 edge -- the 07-20 MAC-regen class). RULED 2026-07-21: targeted netem apply (question + selection in session changelog). Applied via saved -target plan, exact 1/0/0 (docs/audit/outer-{plan,apply}-20260721-netem*.txt); placeholder profile LIVE on virbr5: netem delay 3ms 1ms loss 0.01% (PROVISIONAL -- S6 same-metro lean; D-100 gap #11 final numbers remain unruled), virbr7/virbr3 untouched (docs/audit/stepE-netem-20260721.txt). Convergence re-plan = 0/1/0, exactly the office1 residual -- split per E3 into NEW gate G16 (office1 channels state reconcile), which then CLOSED 2026-07-21 by operator-ruled state surgery (outer plan back to ZERO DIFF, section 5; gate table row G16). ACTIVE: the stage-close set only (GA-R2 consolidation, skill sweep, final gauntlet, operator-gated merge to main).148dcef; rulings docs/audit/ga-rulings.md; the Phase-5 sweep ran as six operator-gated batches in one session; exit runs docs/audit/phase6-exit-runs-20260719.md). The FREEZE is LIFTED -- normal change discipline (this document + the GA rulings) governs.docs/dc0-deploy-readiness.md:107) HAS happened: host rebooted ~2026-07-17 23:39, both guests self-recovered via autostart (docs/audit/env-snapshot-20260718.md:10-16; re-measured this session, section 2.2 below).tofu -chdir=opentofu state list,run 2026-07-18, 20 resources)
Office1 site (live, load-bearing):
module.voffice1.libvirt_domain.vm + .libvirt_volume.disk + .libvirt_volume.seed + .libvirt_cloudinit_disk.seed (BUT see divergence 2.3-i: the cloudinit staging ISO no longer exists live)module.office1_opnsense.libvirt_domain.vm + .libvirt_volume.diskmodule.office1_network.libvirt_network.office1_localmodule.office1_storage.libvirt_pool.dcmodule.ubuntu_noble_base.libvirt_volume.baseInter-site fabric and DC scaffolding:
module.mesh_vr1_dc0_vr1_dc1.libvirt_network.link, module.mesh_vr1_dc0_office1.libvirt_network.link, module.mesh_vr1_dc1_office1.libvirt_network.link (the D-100 mesh triangle)module.vr1_dc0_planes.libvirt_network.plane["data-tenant" | "metal-admin" | "metal-internal" | "provider-public" | "replication" | "storage"] -- applied and live, but REMOVED from config (see 2.3-iii)module.vr1_dc0_storage.libvirt_pool.dc, module.vr1_dc1_storage.libvirt_pool.dchostname -> vcloud; uname -r -> 6.8.0-136-generic.virsh list --all -> exactly two domains, both running: voffice1 (Id 1), office1-opnsense (Id 2).virsh dominfo -> Autostart: enable on BOTH domains.ssh voffice1 'snap list maas lxd; uname -r' -> maas 3.7.2-17972-g.35e297c4d rev 41649 (3.7/stable), lxd 5.21.5-f2a1a0e rev 40074 (5.21/stable, held), guest kernel 6.8.0-136-generic.ssh office1-netbox 'curl -s -o /dev/null -w "netbox=%{http_code}" http://localhost:8000/' -> netbox=302 (service up, redirecting to login).ssh office1-tailscale 'tailscale status | head -1' -> 100.64.0.53 office1-tailscale ... linux - (subnet-router VM up).scripts/site-baseleg.sh check office1 passed at the Phase-1 snapshot (docs/audit/env-snapshot-20260718.md:16-17); not re-run this session.docs/vr1-office1-as-built.md:42, updated 2026-07-18). NOT re-measured this session -- measuring requires the gated API credential path; see section 7./etc/nftables-sec010.nft on voffice1 and BOTH DC racks now carries the idempotent declare-then-delete preamble -- sec010-fw double-restart converged with no rule duplication (4/2/2 drops; pre/post capture docs/audit/sec010-reassert-20260723.txt; per-host backup .pre-reassert-20260723 on each host). Generator fixed the same day (queue-pass changelog items 3/8).i. module.voffice1.libvirt_cloudinit_disk.seed is IN STATE but its staging ISO was deleted by the reboot -- it still plans as a benign re-create (1 of section 5's 6 adds). The FORCED REPLACEMENT it used to force on libvirt_volume.seed (the GA-F01 defect that stopped the apply) is FIXED: D-130 ADOPTED (a) + implemented 2026-07-19, verified by the v8/v7 captures (gate rows G4/G5). Mechanism history: docs/finding-20260718-voffice1-cloudinit-seed-replace.md:186-228. ii. Autostart: RESOLVED 2026-07-19 by the G6 state surgery (operator- ruled (ii), gate row G6): state now records autostart = true on both domains (state show | grep -c autostart -> 1 each); the 2 in-place changes are gone from the plan (section 5 capture). Guests were never touched. iii. The six vr1-dc0 plane networks exist live and in state but are REMOVED from config -- the INTENDED Model B relocation (planes get recreated inside vvr1-dc0 by the inner root, opentofu/vr1-dc0-substrate/main.tf:28). Their emptiness (0 leases, 0 attached domains) was verified in a PRIOR session (docs/dc0-deploy-readiness.md:43-45) and must be re-verified in the same session as any apply (finding doc:259-265). iv. RESOLVED 2026-07-19 (Batch 2, GA-F02): the readiness doc's falsified deploy-ready banner and its three contradictory plan counts are demoted -- status and the expected triple point HERE; the fresh- session banner points at the G9 canonical entry doc.
(2026-07-20) The step-B transit-reach work is APPLIED and verified -- voffice1 holds the region end 172.31.0.1/30 on enp2s0 (plus its in-guest drop-in /etc/netplan/60-transit.yaml); vvr1-dc0 answers at 172.31.0.2 (netplan set-name root cause fixed, kernel names enp1s0/enp2s0 kept); ssh -J voffice1 with the dc0 key works; rack->region ping 10.10.0.20 0% loss. Kea reservation re-keyed to the regenerated voffice1 MAC (incident, session changelog item 3). Consequence for the G10 bootstrap: call site-headend-install.sh with --transit-if enp1s0 --uplink-if enp2s0.
module "vvr1_dc0" (opentofu/main.tf:360) -- the DC0 containment VM (416 GiB / 108 vCPU, D-121/D-123 sizing) + its disk, seed volume, and cloudinit seed. 4 of the 5 committed DC0 creates in the plan capture.
module "vr1_dc0_uplink" (opentofu/main.tf:341) -- the D-125 simulated ISP NAT network 172.30.2.0/24 (capture lines 226-247). The 5th create.opentofu/vr1-dc0-substrate/ (main/variables/ versions.tf; no state file exists in that directory) -- inner storage pool, the six relocated planes, bridge-in WAN, and the rest of the Model B step-C build.module "vr1_dc0_planes" from the outer config (the 6 intended destroys; see 2.3-iii).autostart = true on voffice1/office1-opnsense in config (D-127) -- live-true but state-absent (2.3-ii).scripts/opnsense-plugins.sh + tests/opnsense-plugins/ (D-129 profile-installer; commit 4cefa8b). BUILT and green; the live apply against the edge is an operator-gated firmware mutation, NOT run (docs/design-decisions.md:4050-4056).scripts/site-headend-install.sh --host-nodes writing the transit FORWARD-drop) -- COMMITTED, applies on vvr1-dc0 at deploy step B; ledger row stays OPEN until applied+verified (bash scripts/ledger-scan.sh output, SEC-010 row).runbooks/dc-dc-phase3..6-*.md) -- written, not executed (docs/dc-dc-deployment-workflow.md:873).docs/design-decisions.md:3291,3474,3543,3635,3713): the DC0 deploy sequence they rule (steps A-E, docs/dc0-deploy-readiness.md:170-179) exists only as config + runbook. Nothing DC0 is built.vr1-dc1: topology ruled (D-100/D-101) and addressing RATIFIED 2026-07-21 (D-124 amendment: planes contiguous in 10.12.64.0/19, transit 172.31.0.4/30, uplink 172.30.3.0/24; apex confirm-free at authoring). Still UNBUILT: no vr1_dc1_rack_* variables, no dc1 substrate root; only its storage pool and mesh legs exist (state list, 2.1). Gate row G12 carries the remaining [V] leg.opentofu/modules/netem-link) but HELD as a comment in the root (opentofu/main.tf:309-319); placeholder parameters ruled for the rehearsal (readiness doc:73-75); final parameters unruled (section 8).os-smart, os-nut | os-apcupsd, microcode, os-lldpd) -- recorded, inert until the Roosevelt edge build (docs/design-decisions.md:4030-4031).docs/design-decisions.md:4041-4047,4057-4063).The true EXPECTED outer plan count is currently NOBODY'S:
docs/audit/outer-plan-20260718.txt, line 428: "Plan: 7 to add, 2 to change, 7 to destroy." -- the ONLY citable plan-count source);docs/dc0-deploy-readiness.md:59, docs/session-ledger.md:278).The EXPECTED outer plan is ZERO DIFF ("no differences"), re-recorded 2026-07-22 with its evidencing capture (docs/audit/outer-plan-20260722-postdc1-converged.txt) after the G12 dc1 substrate step-A apply (saved plan 5/0/0 exact -- vvr1-dc1 + vr1-dc1-uplink adds only, zero touches to live resources). A future outer plan showing ANY diff is a STOP (investigate drift before touching anything). History: 7/2/7 post-reboot symptom -> 6/2/6 post-D-130 -> 6/0/6 post-G6-reconcile -> applied exact -> zero diff -> 1/1/1 (voffice1 transit, ruled+applied) -> 2/0/2 (rack netplan fix, applied) -> zero diff converged 2026-07-20 -> 1/1/0 netem-wire STOP -> targeted netem apply 1/0/0 exact -> 0/1/0 office1 residual -> G16 state surgery -> zero diff converged 2026-07-21 -> 5/0/0 dc1 substrate adds (G12 step A, saved-plan applied exact 2026-07-22) -> zero diff converged -> 0/1/0 office1 qga channel (G13 bundle, saved-plan applied exact 2026-07-23) -> zero diff converged (docs/audit/outer-plan-20260723-postqga-converged.txt, this entry) -> RE-CONFIRMED ZERO DIFF 2026-07-27 by the Stage-5 grounding audit, and for the FIRST TIME across ALL THREE roots in one session: the outer root on vcloud (docs/audit/outer-plan-20260727-stage5-audit.txt, "No changes") AND both inner roots on voffice1 (vr1-dc0-substrate and vr1-dc1-substrate, each "No changes", exit 0). Validity of the inner pair was established FIRST by proving the voffice1 clone's inner roots are byte-identical to main despite that clone being 105 commits behind (git diff --stat 61c416e..main -- opentofu/vr1-dc0-substrate/ opentofu/vr1-dc1-substrate/ opentofu/modules/ -> EMPTY); had that diff been non-empty the inner plans would have been UNMEASURED, not green. Full capture: docs/audit/stage5-live-measurement-20260727.txt.
Owner legend: operator (human ruling/approval), session (agent work under gating), external (outside this repo/track). Type legend (GA-R6/E1): [V] = verification-type (closes on its named executable check); [R] = ruling-type (closes per a GA-R5 recorded ruling).
| # | Gate | What closes it | Owner | Evidence of current state |
|---|---|---|---|---|
| G1 | Audit Phase 3: fresh-agent grounding test | [V] 3 clean-context probes score the 7-question set against this doc; holes map made | session | CLOSED 2026-07-18: 3 probes, 21/21 PASS, holes H1 (amended into G9) + H2 (no action) -- docs/audit/phase3-grounding-test-20260718.md |
| G2 | Audit Phase 4: GA-R1..R7 structural rulings + the stage-status vocabulary A/B | [R] ruling-type gate (GA-R6 rule 6): closes when every item carries a GA-R5 Status block | operator | CLOSED 2026-07-18: all seven GA-R + vocabulary (Option A + H1) RATIFIED, utterances quoted (docs/audit/ga-rulings.md, through commit fe4f1c4 + this one) |
| G3 | Audit Phase 5: repair sweep of GA-F01..F15 (incl. memory hygiene GA-F05..F08, skill sweep) | [R] operator-gated fix batches, each commit naming its GA-F | operator + session | Batch 0 OPENED by operator 2026-07-19; items 0.1 (repo-lint L10, GA-R1/C1), 0.2 (SEC repoint, GA-R4/F3), 0.3 (counter hardening, GA-F15), 0.4 (extractor vocab scan, GA-F10/H1) landed; Batch 0 CLOSED (verification passed 2026-07-19); Batches 0-4 CLOSED 2026-07-19 (Batch 4: GA-R4 ledger rotation 1187->131 lines, F1 cap now enforceable; 96 changelogs + 24 history docs consolidated to docs/archive/ with 4 stage records + per-stage manifest commits; top-level docs/ = 16 files < 25; live-surface refs rewritten); Batch 5 CLOSED 2026-07-19 (skill sweep: GA section added reconciled against ratified text, stale phase/UNVALIDATED claims demoted, bookends + stage-close rewritten; checklist docs/audit/skill-sweep-checklist-20260719.md); Batch 6 OPEN: exit runs 1/2/4/5 PASS (adjudication + captures: docs/audit/phase6-exit-runs-20260719.md); exit item 3 PASS 2026-07-19 (fresh trio 21/21, 7/7 all three -- exit record); CLOSED 2026-07-19 -- CORRECTED 2026-07-27 (Stage-5 grounding audit, finding L1-7). This cell previously read "Batch 6 OPEN ... PENDING only item 6 (operator re-read + re-sign) ... FREEZE holds for un-gated surfaces". That was contradicted by its OWN cited evidence file: docs/audit/phase6-exit-runs-20260719.md:84 records "## 6. Operator re-read + re-signature -- SIGNED 2026-07-19 ... PASS" and :91-92 states "ALL SIX EXIT RUNS PASS. The grounding audit is EXITED; G3 + G11 CLOSED". Corroborated by row G11 (CLOSED, re-signed 2026-07-19) and by section 11's recorded utterance "Reviewed, approved, continue." No ruling was required to correct this -- two surfaces already declared it closed. The consequential half was the FREEZE clause, not the state token: left standing it would have blocked the very DOCFIX remediation batch the 2026-07-27 audit queues. The freeze was lifted at audit exit 2026-07-19 (section 1) and normal change discipline governs |
| G4 | The two D-130 verifications | [V] run them, capture output | session | CLOSED 2026-07-19: v8 suppression CONFIRMED (7/2/7 -> 6/2/6, zero forces-replacement; docs/audit/outer-plan-20260719-v8-ignorechanges.txt + -baseline.txt); v7 no-bounce under running domain, zero residue (docs/audit/throwaway-v7-20260719.txt) |
| G5 | D-130 mechanism ruling (seed-volume durable fix) | [R] operator rules in Phase 5, quoting G4's captured output | operator | CLOSED 2026-07-19: D-130 ADOPTED (a) ignore_changes (docs/design-decisions.md D-130, GA-R5 utterance quoted); implemented in modules/cloudinit-vm + tests/cloudinit-vm |
| G6 | State reconcile of autostart + seed WITHOUT bouncing guests | [R] gated mechanism, operator-ruled (S3) | operator | CLOSED 2026-07-19: ruled (ii) state surgery (GA-R5); pull -> inject autostart:true on both domains -> push (serial 22->23, backup terraform.tfstate.pre-G6-20260719); guests never touched (ids 1/2 unchanged, running) |
| G7 | New captured plan == the expected triple recorded in section 5 | [V] re-plan to a capture file after G5+G6 | session | CLOSED 2026-07-19: capture docs/audit/outer-plan-20260719-postG6.txt = 6/0/6, equals section 5 exactly |
| G8 | Same-session pre-apply re-verify: 6 planes still empty | [V] run in the SAME session as the apply | session | CLOSED 2026-07-19: verified in the apply session itself (all six 0 leases; only office1 nets attached) immediately before step A |
| G9 | DC0 outer apply (deploy step A) | [V] operator-gated, logged (run-logged.sh), after G1-G8; audit exit criteria met (charter Phase 6). SEC pre-apply dependency (S2): SEC-010's transit FORWARD-drop is applied+verified at deploy step B via site-headend-install.sh --host-nodes --check on vvr1-dc0 (gate G10) -- the ONLY SEC row gated on this apply (register of record: security-ledger). CANONICAL ENTRY DOC (probe hole H1): runbooks/dc-dc-phase2-tofu-dc-substrate.md, with docs/dc0-deploy-readiness.md section E as the step table |
operator | CLOSED 2026-07-19: G8 same-session planes check passed (6x 0 leases, 0 attachments); saved plan == 6/0/6 applied in the logged dc0-deploy window; convergence re-plan = no differences; vvr1-dc0 running, prior guests untouched |
| G10 | Deploy steps B-E in-sequence gates: SEC-010 --host-nodes --check on vvr1-dc0; depth-4 nested boot; D-125 foreign-MAC egress test; MAAS reachability + TF_VAR_maas_api_key before step D; netem placeholder step E |
[V] exercised during the gated deploy | session (each mutation operator-approved) | Step B DONE 2026-07-20 (--check EXIT 0 incl. SEC-010, docs/audit/stepB-check-20260720-final.txt; interfaces enp1s0/enp2s0). Depth-4 nested boot DONE (10 domains running inside vvr1-dc0). D-125 egress isolation test PASS 2026-07-20 (docs/audit/d125-egress-gate-20260720-matrix.txt), and the edge itself now egresses 0% loss after the v4 addressing. Step D COMPLETE incl. commissioning: ALL 9 NODES READY 2026-07-21 (two stacked faults diagnosed + fixed -- docs/audit/commissioning-diag-20260721.txt; section 1). Step E (netem) DONE 2026-07-21: sudo fragment installed+verified, module local-mode amendment, targeted apply 1/0/0 exact (operator-ruled at the 1/1/0 STOP), placeholder profile live on virbr5, virbr7/virbr3 untouched (docs/audit/stepE-netem-20260721.txt + outer-{plan,apply}-20260721-netem*.txt). G10 CLOSED 2026-07-21 |
| G11 | Operator signs THIS document | [R] read top-to-bottom; discrepancies resolved in the document | operator | CLOSED: RE-SIGNED 2026-07-19 at audit exit, section 11 (replaces the 2026-07-18 signature) |
| G12 | vr1-dc1 build |
[R] operator rules dc1 transit/rack addressing; then vars + substrate authored | operator + session | CLOSED 2026-07-23 (operator-ruled "Merge to main + full close"; commissioning 9/9 READY, merge commit on main, branch retired) -- [R] leg CLOSED 2026-07-21: addressing RATIFIED (D-124 amendment 2026-07-21, utterance quoted). [V] leg IN PROGRESS (branch dc-dc-g12-dc1-substrate): apex confirm-free DONE 2026-07-21 -- planes/uplink already assigned+consistent, transit 172.31.0.4/30 + rack 10.12.68.2 FREE (docs/audit/dc1-apex-confirm-20260721.txt); importer per-site dc1 support shipped (harness 117/117) with live dry-run preflight PASS (docs/audit/dc1-rack-import-dryrun-20260721.txt). vars + substrate root + lib-net dc1 arm COMMITTED 2026-07-22 (successor session landed the disconnected item 3 + the harness reconcile as changelog item 4): six harnesses reconciled to the ratified dc1 arm, phase-00 PLANES parity guard added, rbd-mirror/radosgw cross-DC reminder fixed; gauntlet ALL GREEN (76) (docs/audit/gauntlet-20260722-g12-reconcile.txt), repo-lint 0-fail. Apex --commit EXECUTED 2026-07-22 (operator-gated): 172.31.0.4/30 + 10.12.68.2/22 CREATED, post-commit read-back idempotent (docs/audit/dc1-rack-import-commit-20260722.txt). dc1 svc key minted (creds-audit CLEAN), tfvars authored (local), outer step-A apply DONE 2026-07-22: saved plan 5/0/0 exact, converged ZERO DIFF (section 5), vvr1-dc1 RUNNING, prior guests untouched (as-executed log dc1-deploy; changelog-20260722-g12-dc1-build.md). Step B COMPLETE 2026-07-22: cloudinit-vm interface_macs port + voffice1 dc1-transit NIC (0/2/0 exact, MACs pinned both domains, post-bounce battery ALL PASS, converged zero diff -- docs/audit/outer-plan-20260722-voffice1-dc1nic.txt), transit LIVE (voffice1 .5/30 <-> rack .6/30, dc1-key ssh proven), rack ENROLLED (region lists vvr1-dc1 nmpcq4), SEC-010 applied+verified BOTH ends, OPNsense 26.7 base staged via hash-verified copy of dc0's proven artifact; named gate EXIT 0 docs/audit/dc1-stepB-check-20260722-final.txt (changelog-20260722 items 5-8, three queued findings). Step C COMPLETE 2026-07-22: inner apply FROM voffice1 -- plan 28/0/0 exact (54 pinned MACs verified in-capture), one fix-forward (serial-log staging dir absent on dc1; queued to standup DoD), resume 10/0/0 exit 0; 28/28 in state, convergence ZERO DIFF (docs/audit/inner-converge-20260722-dc1-stepC.txt), 10/10 domains RUNNING inside vvr1-dc1, edge at the 26.7 FreeBSD login prompt (D-112 datapoint #2); dc1 inner tfstate ON voffice1 (site backup set). D-125 egress gate PASS 2026-07-22 (two identical runs, dc0 criteria exact, isolation confirmed -- docs/audit/d125-egress-gate-20260722-dc1.txt). Edge bootstrap + v4 addressing COMPLETE 2026-07-23 (changelog-20260723-g12-dc1-edge.md): D-112(c) console bootstrap done (SSH + dc1 edge key materialized; payload needed util.inc/shell_safe() -- dc0 lesson iv the .b64 artifact lacked), key-only SSH VERIFIED (15.1-RELEASE-p1); D-113(a2) API key MINTED via the vendor model + smoke test GET core/firmware/status exit 0 product_abi 26.7 (second 26.7 datapoint); edge ADDRESSED -- WAN 172.30.3.2/24 gw 172.30.3.1 (egress 1.1.1.1 0% loss), LAN 192.168.1.1 -> 10.12.64.1/22 (ruled provider-public gw), API answers at the new LAN; interim reach leg removed, rack provider-public leg 10.12.64.2/22 LIVE on virbr4. Creds consolidated to ~/vr1-dc1-creds/opnsense-api.txt (creds-audit CLEAN, 5 entries); rack edge-key copy shredded (SEC-015 transient, remediated). Two queued findings: bootstrap .b64 missing util.inc; opnsense-bootstrap-apikey.sh scp had a transient post-restart-sshd failure (readiness-wait/retry candidate). Rack standup + region MAAS config DONE 2026-07-23 (changelog-20260723 items 7-11): dc-rack-net.sh dc1 arm shipped (harness 18/18, gauntlet 76 GREEN) + INSTALLED on the rack (check 10/10, forwarder answers authoritative maas-internal SOA -- D-131 fix; docs/audit/dc1-rack-net-install-20260723.txt); region MAAS on metal-admin subnet 11 -- D-120 range 10.12.68.100-.200, D-131 dns_servers=10.12.68.3 allow_dns=false, DHCP dhcp_on=true primary_rack=nmpcq4 (dhcpd verified RUNNING on virbr6, no Temporal incident); dc1 enlistment PROVEN via canary (machines 11->12 in ~2 min). SEC-016 RULED + WIRED 2026-07-23 (operator: "Mint a dedicated dc1 power key" -- per-DC isolation; dedicated key authorized on the rack + installed in the region MAAS snap with per-host ssh config, dc0's SEC-012 key untouched). COMMISSIONING 9/9 READY 2026-07-23 (docs/audit/dc1-commissioning-verify-20260723.txt): all 9 nodes PXE-enlisted by pinned 52:54:01:d1 MACs, power_type=virsh set + verified by real query-power-state (SEC-016 path proven), commissioned to ALL 9 READY in ~3.5 min (no timeout, no SERVFAIL), shapes EXACT to D-121 Option C (3x16cpu/64GiB + 2x12cpu/48GiB + 4x8cpu/24GiB). dc0's two stacked faults pre-empted by pinned MACs + the dc-rack-net forwarder. G12 [V] leg (the dc1 build) is COMPLETE. NEXT: G12 close-out only -- consolidate this session's changelogs (GA-R2), final gauntlet + repo-lint, GA-R7 memory review, skill sweep, operator-gated merge of dc-dc-g12-dc1-substrate -> main (merge commit), branch retirement; then G12 CLOSES. NOTE open SEC rows now include SEC-014/-015/-016 (G14 row count stale -- reconcile in the close). |
| G13 | D-129 residuals | [R] operator-gated live plugin install on office1-opnsense; qga channel retrofit at that edge's next scheduled restart. All 4 sub-decisions RULED 2026-07-21 (D-129 Status line) -- only the two execution items remain | operator | CLOSED 2026-07-23 (operator-approved full maintenance bundle, logged window ops-sec010-reassert): qga channel retrofitted via outer tofu saved-plan apply 0/1/0 exact (docs/audit/outer-plan-20260723-office1-qga.txt; the apply's edge bounce = the ruled "next scheduled restart"; MACs were pinned 07-22 so the in-place-update trap class was closed); edge updated 26.7 -> 26.7.1 via REST (no reboot required; os-iperf had been REFUSED on 26.7 pending exactly this update); both plugins installed=1 by firmware-info read-back, guest-ping -> {"return":{}}, agent reports both legs, egress 0% loss, outer plan re-converged ZERO DIFF (docs/audit/outer-plan-20260723-postqga-converged.txt). Named close capture: docs/audit/g13-close-20260723.txt |
| G14 | 12 OPEN SEC rows (SEC-001, -003..-008, SEC-012, -013, -014, plus SEC-015 + SEC-016 opened 2026-07-23 for dc1 credentials; SEC-010/-011 CLOSED) | [R] per-row: rotations/flips at v1 close (external to VR1 track); SEC-012/-016 carry the same libvirt-group SCOPE hardening question; SEC-016 also a snap-refresh re-assert (queued to DC standup DoD) | operator / external | docs/security-ledger.md (register of record, GA-R4/F3). COUNT RECONCILED 2026-07-27: 21 open, measured bash scripts/ledger-scan.sh (SEC-001, -003..-008, -012..-025; SEC-024 opened 2026-07-26, SEC-025 opened 2026-07-27 for the consolidated NetBox GUI admin password). Earlier figures in this cell (19 at 2026-07-25) are history. The row's own title text ("12 OPEN SEC rows") is the 2026-07-23 figure and is SUPERSEDED by this cell -- the gate is the ledger, not the count. Since 07-23: SEC-017 (caveman supply-chain), -018/-019 (per-DC MAAS API keys), -020 (MAAS region superuser passwords), and -021/-022/-023 opened 2026-07-25 from the D-137 credential research -- dc0 custody defects (a consolidated credential ABSENT from its recorded location + per-DC power-key divergence), UNAUDITED shadow *-creds/ stores on voffice1 (a scope gap in creds-audit itself, which has no remote capability), and sprawl-glob blind spots incl. a PREDICTED Stage-5 ~/admin-openrc exposure. All three are logged-not-actioned (hard rule 1); remediation is coupled to the unruled D-137 forks. Creation-point research capture: docs/audit/creds-creation-points-20260725.md -- 55 MINT sites inventoried, and 12 declared secrets have NO mint command anywhere in the repo (ssh-keygen returns ZERO hits repo-wide; six SSH keypairs + the OPNsense root password/hash are operator-terminal mints recorded only in a manifest comment, i.e. NOT reproducible if the jumphost is rebuilt -- a Roosevelt-transfer defect, not just hygiene). It also names three credential DIRECTORIES outside the SEC-009 *-creds/ convention and outside creds-audit entirely: ~/vault-init/ (Vault 5 unseal shares + root token), ~/octavia-pki/ (8 PKI artifacts incl. CA private keys), ~/tenant-<client>/; plus overlays/octavia-pki.yaml, which lands a CA key + plaintext passphrase INSIDE the repo clone (gitignored -- and SEC-004 says the repo is still PUBLIC) |
| G15 | D-068 / D-071 rulings | [R] operator rules (section 8); neither blocks the VR1 substrate | operator | D-071 ADOPTED 2026-07-21 (all four points); D-068 items 2-3 RULED 2026-07-21; item 1: plan DRAFTED + Q1/Q2-structure/Q3 ALL RULED 2026-07-23 (three amendments, utterances quoted; monthly-review lines delivered). Sole D-068 remainder: Q2 path selection at Roosevelt Vault design time -- G15 is otherwise decision-complete |
| G16 | office1 edge channels = [] state reconcile (the D-129 module-schema residual) |
[R] operator rules the mechanism; then [V] the converged re-plan capture | operator + session | CLOSED 2026-07-21: RULED "State surgery (Recommended)" (GA-R5, session changelog item 16); executed per G6 precedent -- channels null -> [] injected, serial 29 -> 30, backup kept, guests untouched (office1-opnsense Id 2 running throughout); convergence = ZERO DIFF (docs/audit/outer-plan-20260721-postG16-converged.txt); section 5 re-recorded |
| G17 | Per-DC artifact source reachable FROM A NODE -- the node-side half of Stage 4 DoD bullet 5, split out of Stage 4 by operator ruling rather than closed conditionally | [V] a NAMED executable check run from a node that has actually booted an OS on its real NICs, per DC, each capturing: dc0 -> curl -sI http://10.12.8.4/ returns 200 from the node (the D-135 item-1 full mirror); dc1 -> the node resolves and fetches through the apt proxy at 10.12.68.4:3142 (the D-135-AMENDED ruled artifact path -- dc1 has NO node-facing mirror, so checking it as one would fail by design). The natural trigger is Stage 5 first boot, when Juju provisions the nodes and they run apt for real; a gated MAAS rescue-boot is the alternative if it must be answered sooner |
session (each boot operator-approved) | OPEN 2026-07-27. WHY THIS EXISTS: the DoD bullet reads "per-DC mirror reachable from nodes", but the READY-handoff ruling (2026-07-23, DOCFIX-200) leaves all 18 nodes powered off in Ready -- MAAS-deploy is SKIPPED and Juju provisions at Stage 5 -- so no node-side probe can run inside Stage 4 at all. GA-R6 E3 forbids a conditional close, so the remainder splits here. RULING (GA-R5). Question as presented 2026-07-27: "The node-side half of bullet 5. Nodes are powered off by the READY-handoff ruling, so no node-side probe can run as things stand. Either a gated rescue-boot check on one node per DC now (closes it inside Stage 4), or split it into its own gate row targeted at Stage 5 first boot (GA-R6 E3 explicitly permits this; a conditional close is not permitted)." Operator answer, exact utterance: "split it into its own gate row". SCOPE NOTE: what stays in Stage 4 is the RACK-side half -- the artifact source answers on its own address with an attested-current sync -- which is what dc-mirror.sh check / dc-cache-proxy.sh check verify (both fixed this session to stop false-greening; capture docs/audit/stage4-mirror-gate-20260727.txt). G17 is NOT a Stage-5 precondition and must not be conflated with one: Stage 5's own bootstrap needs OPEN edge egress for the juju agent stream + snaps (D-135 items 2-3 unbuilt), which is a different path from the apt artifact source this gate covers. |
| Component | Measured value | Command (run 2026-07-18) | Where measured |
|---|---|---|---|
| OpenTofu | v1.12.4 | tofu version |
vcloud (also docs/audit/env-snapshot-20260718.md:26) |
| libvirt provider | dmacvicar/libvirt 0.9.8 (pinned) | grep -A2 'provider' opentofu/.terraform.lock.hcl |
repo lock file |
| MAAS provider | canonical/maas 2.7.2 (pinned) | same | repo lock file |
| MAAS | 3.7.2-17972-g.35e297c4d (3.7/stable) | ssh voffice1 'snap list maas' |
voffice1 |
| LXD | 5.21.5-f2a1a0e (5.21/stable, held) | ssh voffice1 'snap list lxd' |
voffice1 |
| Kernel (host) | 6.8.0-136-generic | uname -r |
vcloud |
| Kernel (voffice1) | 6.8.0-136-generic | ssh voffice1 'uname -r' |
voffice1 |
| OPNsense edge | 26.7.1 (FreeBSD base 15.1) | MEASURED 2026-07-23 via the gated API (GET core/firmware/status -> product_version 26.7.1, capture docs/audit/g13-close-20260723.txt); updated 26.7 -> 26.7.1 in the G13 bundle. DC edges (vr1-dc0/dc1) remain 26.7 |
office1-opnsense |
| NetBox (Office1 apex) | 4.6.4 per as-built docs/vr1-office1-as-built.md:44; service UP verified (HTTP 302) this session |
ssh office1-netbox 'curl ... localhost:8000' |
office1-netbox |
| Juju | 3.6.27 (rev 35621, 3/stable) MEASURED 2026-07-27 on voffice1 -- the headend is the D-128 Plane-2 execution host and this is the client that will bootstrap the controller. Supersedes the 3.6.25 figure recorded 2026-07-24 (also at line 162, kept there as history): an in-channel patch refresh, which D-071 ADOPTED 2026-07-21 explicitly permits (patch-only jumps, in-channel-only refreshes), so this is policy-compliant drift and NOT an incident. The jumphost has NO juju client (measured ABSENT). |
ssh voffice1 'snap list' (capture docs/audit/stage5-live-measurement-20260727.txt) |
voffice1 |
| OpenStack client | ABSENT ON BOTH HOSTS, MEASURED 2026-07-27 -- command -v openstack returns nothing on vcloud AND on voffice1, and voffice1's snap list carries only core24/juju/lxd/maas/postgresql/snapd. Ten Stage-5/6/7 scripts invoke it (phase-03-admin-openrc.sh, phase-04-network-{create,verify}.sh, phase-04-internal-cert-san-verify.sh, phase-05-{amphora-pipeline,octavia-verify}.sh, phase-06-{bootstrap,capi-stack,mgmt-vm,net-setup}.sh), so every post-deploy step from Stage 5 Step 7 onward has no client to run. Recorded here because it is a measured Stage-5 precondition that no prior surface carried. |
command -v openstack; ssh voffice1 'command -v openstack; snap list' |
vcloud + voffice1 |
The known-stale pin sites this table used to enumerate (the GA-F03/F04/ F05 tofu, OPNsense, and jumphost-name values -- stated token-free here so the scan does not count them) were ALL fixed or demoted to pointers in sweep Batches 2-3, 2026-07-19 (session changelog).
docs/design-decisions.md D-130 (question + utterance + captures).docs/D-068-vault-migration-plan-draft.md). Q1 RULED 2026-07-23 (rehearsal-scoped EOL risk-acceptance, posture 1b -- amendment in design-decisions.md is the authority). Q2 structural assumption RULED 2026-07-23 (Roosevelt baselines on 1.8/stable + a FUNDED remediation track; path 2a/2b/2c + deadline stay OPEN to Roosevelt design time on re-verified V1-V5). Q3 RULED 2026-07-23 (monthly-review lines delivered into ops-update-procedure 0c; design-time re-verify trigger). SOLE remainder on item 1: Q2 path selection (2a/2b/2c) + deadline, at Roosevelt Vault design time (item 1 Status line: OPEN on exactly that remainder -- reworded at the 2026-07-23 close for scan attribution).scripts/opnsense-plugins.sh apply vr1-edge)
creds-audit vr1-office1 read CLEAN on 2026-07-15 while four region-VM secrets minted 2026-07-13 sat undeclared, and admin.pass surfaced only 2026-07-25 (SEC-020). It is also wired in exactly ONE place and as PROSE (phase-3 runbook:498), not in preflight/cloud-assert/repo-lint/gauntlet -- and that line did not fire at EITHER DC standup (dc0 3 + dc1 1 undeclared at this session's open). THREE forks await a ruling, one exchange each (GA-R5): enforcement strength (advisory / preflight-blocking / plus a PreToolUse guard); --remote discovery scope (declared-directories vs broader sweep, a tenant-isolation concern); policy home (this D as authority with SEC-009 demoted to a pointer, vs policy stays in the ledger). NOT implemented -- PROPOSED means present options, never build. SUB-RULING 1 RULED 2026-07-25 (GA-R5, utterance quoted in the D-137 Status block): "Blocking in preflight" -- the check lands as a new Pn in scripts/preflight.sh and HARD-FAILS on any expected-but-absent / undeclared / per-DC-asymmetric credential. The PreToolUse guard and advisory-only were NOT adopted. SUB-RULING 2 RULED 2026-07-25: "Derive manifests from matrix" -- the matrix is SINGLE SOURCE, --render regenerates creds-manifests/*.manifest, gauntlet fails on rendered-vs-checked-in drift; accepted cost is that manifests become generated (their governance prose must become matrix fields, not be dropped). SUB-RULING 3 RULED 2026-07-26: "Declared locations only" -- a creds-manifests/vm-secret-locations list bounds --remote absolutely (no tenant surface touched, D-069 preserved). To be faithful the list must include the headend shadow stores (SEC-022), the region maas-secrets dir, and the three dirs outside the SEC-009 convention found by the creation-point inventory. SUB-RULING 4 RULED 2026-07-26: "D-137 is the authority" -- D-137 becomes the credential-lifecycle policy authority and the SEC-009 convention block demotes to a pointer (its founding history stays in the ledger as history); gates may then cite a D-number instead of an exposure register. SUB-RULING 5 RULED 2026-07-26: "Fold in as a D-137 invariant" -- the invariant is ONE IDENTITY SERVES ONE PRINCIPAL TYPE, enforced by the matrix principal column, so the SEC-020 conflation becomes machine-detectable and needs no separate D-number. ALL FIVE SUB-RULINGS RULED -> D-137 is ADOPTED 2026-07-26 and implementation is UNBLOCKED, with the build spec at docs/D-137-implementation-plan.md (Status line in design-decisions.md is the authority). Expect the first run to be RED by design: admin serves both a human and a service row, which is the defect the invariant names. Research capture: docs/audit/creds-creation-points-20260725.md -- its APPENDIX now carries the FULL per-row inventory (55 MINT rows with file:line, host, destination, stage and human/service classification), transcribed in-repo 2026-07-26 so the matrix SEED does not depend on a session transcript. TIER 1 BUILT 2026-07-26 (offline/STATIC half only; tiers 2-3 NOT built, and the preflight Pn of ruling 1 is NOT wired -- see the sequencing question below). Shipped: creds-matrix.tsv (72 rows), creds-matrix-notes.md, scripts/creds-matrix.py (S1 schema / S2 manifest coverage both-bounds / S3 render drift / S4 mint-ref resolution / S5 per-DC symmetry / S6 ruling-5 principal invariant / S7 notes integrity), harness tests/creds-matrix/run-tests.sh 24/24; gauntlet ALL GREEN (80), repo-lint 0-fail. SCHEMA AMENDED 10 -> 12 columns (operator-ruled 2026-07-26, "go with the 12-column amendment"): custody + notes-ref added and rows re-keyed to (credential, location), because a read-first round-trip check proved the 10-column form could not carry what ruling 2 forbids dropping -- and because ruling 3 puts four deliberate, reasoned credential copies INSIDE --remote's declared locations, where creds-audit.sh:63-67 would report every one as UNDECLARED. Detail + rationale in docs/D-137-implementation-plan.md; OPS under GA-R3, the five sub-rulings are untouched. RED BY DESIGN AND CORRECTLY SO -- 5 findings, do not "fix" by deleting rows: the ruling-5 identity conflation on maas-region-admin (its own SEC-020 defect), 3x EXPECTED-BUT-ABSENT for SEC-021 (dc0 declares neither an edge API credential nor a jumphost-consolidated power key, both of which dc1 declares), and the S5 asymmetry of dc0's divergently-named headend power key. 27 rows carry mint-ref=operator-terminal = the research FINDING 1 reproducibility debt, admitted and counted, not faulted. Acceptance test PARTIALLY met (corrected from the plan's original text): SEC-021's DECLARATION half is tier-1 detectable and is reproduced; SEC-022/-023 and SEC-021's on-disk half need tier 2. OPEN SEQUENCING QUESTION for the operator: ruling 1 lands the check as a BLOCKING preflight Pn and the first run is red, so wiring it hard-fails preflight.sh until the credential defects are remediated; wiring-as-ruled is the default and deferring until after per-row remediation is the departure. RESOLVED + EXECUTED 2026-07-26. Question as presented: wire the blocking Pn as ruled, or defer it until after per-row remediation. Operator answer, exact utterance: "wire tier 1 as the blocking preflight Pn". TIER 1 IS NOW WIRED as preflight.sh P5, blocking, ahead of the stage-2 reminders block so its verdict participates in the deploy decision; it FAILS CLOSED if the checker is absent (a missing file made python3 exit 2, which note maps to WARN -- so deleting the gate would have downgraded it to a warning; harness T9 encodes this). Tier 2's Pn is NOT wired and remains a separate decision: it needs --remote/--privileged and a caller-supplied --pending-stage, none of which belong in an unattended gate. CONSEQUENCE, STATED PRECISELY: preflight.sh exits 1 and P5 is one of the reasons -- but preflight was ALREADY exiting 1 before this change (P4: overlays/octavia-pki.yaml absent, and MAAS unreachable from the jumphost). P5 adds a fifth reason to an already-red gate; it did NOT flip preflight from pass to fail, and no deploy path that was open is closed by it. TIER 2 TOOLING BUILT 2026-07-26, NOT YET RUN LIVE. Shipped: creds-manifests/vm-secret-locations (ruling 3's absolute bound -- jumphost creds folders, the headend shadow stores of SEC-022, the region secrets dir of SEC-020, the three dirs outside the SEC-009 convention, the in-clone PKI overlay, and the DOCFIX-175 plaintext tfstate; tenant dirs are LOCAL-only so no tenant surface is reachable and D-069 holds by construction); creds-matrix.py --tier2 [--remote] (E1 expected-but- absent / E2 mode / E3 undeclared-at-a-declared-location, with an unreachable host SKIPPED explicitly because "could not look" must never read as "nothing there", and a --pending-stage selector so a not-yet-reached mint stage defers instead of failing -- the CALLER supplies it, the script carries no status claim, GA-R1); creds-audit.sh sprawl globs WIDENED for the SEC-023 blind spots (admin.pass, *.apikey, *.key, *.pem, *_ed25519, *_rsa, *openrc*) -- the old six patterns could not have seen the SEC-020 secret or the predicted Stage-5 ~/admin-openrc. Harnesses creds-matrix 33/33 and creds-audit 13/13 (was 7); gauntlet ALL GREEN (80), repo-lint 0-fail. LIVE TIER-2 SWEEP RUN 2026-07-26, read-only, no sudo (capture docs/audit/d137-tier2-sweep-20260726.txt; jumphost local + voffice1 + office1-netbox, stat over ssh, metadata only). SEC-021's ON-DISK half is now REPRODUCED as a named failure -- vr1-dc0-maas-power_ed25519{,.pub} are genuinely ABSENT from the dc0 jumphost creds folder, not merely undeclared. The sweep ALSO corrected two of its own false-greens, both found by running it: (a) an unprivileged [ -d ] on a root-owned directory is indistinguishable from absent, so /root/maas-secrets and /root/netbox-secrets first reported "does not exist" -- they are now correctly reported UNREADABLE ("could not look" is never "nothing there"), which is a FAIL, not a skip; (b) a role with any unprobed location no longer lets its other locations' listings manufacture false EXPECTED-BUT-ABSENT findings -- 14 headend/netbox rows are explicitly NOT JUDGED instead. STILL OUTSTANDING: the two root-owned directories need a privileged read, so SEC-022's shadow-store verification and the SEC-020 region secrets remain unconfirmed; that run is a remote-sudo shape and is operator-gated. ESCALATION PATH WIRED BUT BLOCKED 2026-07-26: --privileged adds a sudo -n retry attempted ONLY where an unprivileged probe returned unreadable, so the privileged surface stays as small as the ruling-3 bound keeps the search surface (still metadata only, stat, never content). The operator APPROVED the run, but the Claude Code AUTO-MODE CLASSIFIER denied the remote-sudo shape -- the same wall recorded at the 2026-07-23 close, whose noted fix is manual permission mode (a targeted ask rule is the alternative). NOT worked around. PRIVILEGED SWEEP COMPLETED 2026-07-26 (capture docs/audit/d137-tier2-privileged-20260726.txt; supersedes the unprivileged capture). ROOT CAUSE of the block was NOT the classifier overriding a rule: Bash(ssh * sudo *) was ALREADY in the project ask list, but the pattern needs a literal space before sudo and the command was ssh <host> 'sudo ...' -- the quote meant NO rule matched, so it fell through to the classifier. Fixed by adding the quoted variants (ssh *'sudo *, ssh *"sudo *, ssh *sudo -n *) plus a targeted ask rule for the privileged invocation; all are ask, never allow. Bash(ssh * virsh *) carried the IDENTICAL latent gap; it was flagged-not-fixed at discovery (hard rule 1) and then FIXED 2026-07-26 under operator direction (commit 6d43619, quoted variants added). MEASURED RESULTS -- capture docs/audit/d137-location-listing-20260726.txt (a stat listing; the tier-2 capture cited here previously is a checker VERDICT file and contains none of these values -- a GA-R1 rule 2 defect found by the committee and corrected in this commit). Region secrets dir: admin.apikey, admin.pass, db.pass, lxd-trust.pass (0600) -- precisely the SEC-020(i) carve-out list. Netbox dir: admin.pass, api.token, secret_key (0600). SEC-022 shadow stores confirmed. The claim that "ZERO undeclared files remain, so SEC-020/SEC-022 are fully accounted for" is WITHDRAWN -- it rested on a site-blind check (see the committee block below). INFERRED-FILENAME MISS (hard rule 2), corrected count: the session first reported SIX; the committee measured ~21, because only the rows the sweep physically touched were re-measured. .maas.cli was itself never measured and is WRONG -- the MAAS snap CLI stores its profile at ~/snap/maas/current/.maascli.db (measured this session). Matrix 77 rows at that point. The 9 findings then reported were true so far as they went, but the run that produced them is superseded below. COMMITTEE AUDIT 2026-07-26 (6 independent read-only lenses: correctness, coverage, claim-verification, ruling-fidelity, record-integrity, Roosevelt-transfer). VERDICT: the register's DESIGN holds, but TIER 2's VERDICTS ARE NOT TRUSTWORTHY as delivered and this document overstated what was verified. Every defect below was REPRODUCED, and FOUR lenses converged independently on the first one. Superseding figures (these are current; earlier figures in this entry are history): matrix 77 rows, creds-matrix harness 35/35, creds-audit 13/13, gauntlet ALL GREEN (80). The earlier "14 rows NOT JUDGED" was from the unprivileged run; the privileged run reports 2. The widened sprawl glob is *.pass, not admin.pass. Confirmed false greens (a real missing or misplaced credential passes a green sweep): (1) tier 2 is SITE-BLIND -- observed/declared are keyed by host-role only, so a file in the wrong site's folder satisfies the row; this masked SEC-021's opnsense-api.txt on-disk absence, so "SEC-021's on-disk half REPRODUCED" is corrected to the power-key artifacts ONLY; (2) probe_remote reports UNREACHABLE for a location it successfully read (empty-but-readable dir, or absent literal-file path), which gates the whole role and converts every absence FAIL at that role into an [ok]; (3) literal-file locations get no absent/unreadable detection at all -- the T33 fix covered only dir/* patterns; (4) an empty locations file bypasses ruling 3's refusal and affirms existence over zero locations; (5) no non-empty floor -- a 0-row matrix passes every check; (6) a mint-ref pointing at a directory crashes with exit 1, indistinguishable from findings, and S5/S6/S7 never run; (7) E2 MODE is custody-gated so 43 of 77 rows are never mode-checked; (8) S5's dc1->dc0 direction is untested -- deleting it leaves the harness green (the in-repo both-bounds precedent, in the symmetry check itself). Also outstanding: ruling 4's SEC-009 demotion is NOT done and its stated trigger (tier 2 built) has passed; the 12-column amendment is recorded in no D-numbered surface; sec-ref mis-attribution recurs (both juju-maas-user rows cite SEC-020; correct is SEC-018/-019); the libvirt SSH power password (standing rotation obligation, reenroll-hosts.sh:22-25) has no row; and T24 asserts literal finding strings, so remediating the ruling-5 conflation would turn the GAUNTLET red -- the test punishes the fix it exists to protect. Remediation is IN PROGRESS under operator direction ("Do as many as you can autonomously"); this block is the authority on what is fixed. REMEDIATION COMPLETE 2026-07-26 (phases 1-3), capture docs/audit/d137-tier2-postcommittee-20260726.txt (supersedes the pre-committee privileged capture, whose verdicts predate the site-scoping fix and are NOT comparable). ALL EIGHT false greens are FIXED and individually regression-locked; harness 44/44 (was 35), creds-audit 15/15, gauntlet ALL GREEN (80), repo-lint 0-fail. PROOF THE SITE FIX WORKS: E1 EXPECTED-BUT-ABSENT: dc0-edge-api 'opnsense-api.txt' now appears. That on-disk absence was masked in EVERY prior run, so SEC-021's on-disk half is NOW genuinely reproduced in full (the earlier withdrawal stands as history). The ~21 inferred filenames were re-measured against their mint-refs (octavia's 8 real basenames + its three SUBDIRECTORIES, which the bare ~/octavia-pki/* pattern matched none of; vault's init.txt; the tenant rows' <client>- instance prefix, now supported by a placeholder matcher); sec-ref mis-attributions corrected (both juju-maas-user rows SEC-020 -> SEC-018/-019; tfstate SEC-009 -> none); admin.pass's mint-ref moved from the consumer (:452) to the mint (:450); the libvirt power password and vault-ca-root added as rows. Matrix 81 rows. Findings 35 -> 13, all TRUE: SEC-021 (3x S2 + 3x E1), 3x S5 power-key asymmetry, the ruling-5 conflation, one SEC-022 shadow-store gap, SEC-024, and 2 rows disclosed as UNCHECKABLE. NEW EXPOSURE FOUND BY THE FIXED CHECKER -- SEC-024 OPENED: opentofu/terraform.tfstate.backup is mode 0664 (group- and world-readable) and carries the Office1 MAAS API key in plaintext per DOCFIX-175; the live state file is correctly 0600. Invisible to every prior control because the world-readable check was custody-gated and the siblings were undeclared. REMEDIATED 2026-07-26 (mode): operator-approved chmod 600 on terraform.tfstate.backup and .pre-G16-20260721, read-back verified -- all four state files now 0600 and the E2 finding cleared. SEVERITY CORRECTED first: the row's original "group- and world-readable" overstated it -- opentofu/ is 0700 and ~ is 0750, so no other account could traverse to it, and neither file is git-tracked, so SEC-004 was never implicated. It was defence-in-depth, not live exposure. The RETENTION question is RULED 2026-07-27 (GA-R5). Question as presented: whether to delete both pre-* state-surgery snapshots (cleanest, closes SEC-024 fully), keep both and let P5 keep watching them, or keep pre-G6 and delete pre-G16; noting both gates are CLOSED, both carry the MAAS API key in plaintext per DOCFIX-175, and deletion is irreversible with no other copy of that pre-surgery state. Operator answer, exact utterance: "Keep both". So terraform.tfstate.pre-G6-20260719 and .pre-G16-20260721 are RETAINED as the only record of what the state looked like before the two direct tfstate edits; both are 0600 and neither is git-tracked. SEC-024 stays OPEN as a standing WATCH rather than an open remediation: its mode defect is remediated, but the umask CAUSE was out of ruled scope, so a future apply may rewrite 0664 -- P5 is the detection. The key reaching state files in plaintext at all remains DOCFIX-175, rotation owed under SEC-018/-019. STILL OUTSTANDING (not done, not silently dropped): ruling 1's TIER 2 remote Pn (tier 1 + tier-2-local ARE wired -- see the tier-2 gate block below); --render's source-field derivation and the manifest flip to generated output; and the cardinality field remains largely inert with a one-token S5 bypass (per-DC -> singleton). (A dangling fragment here, left by an earlier edit in this same session, was repaired at session close -- noted rather than silently fixed, since CURRENT-STATE is the status authority and its defects are worth seeing.) CONSOLIDATION BATCH EXECUTED 2026-07-27 (operator question: is there a consolidated set of login creds on vcloud for every account that exists; operator direction after the audit: "clear the whole consolidation batch first"). Capture docs/audit/creds-consolidation-audit-20260727.txt; detail docs/archive/changelogs/changelog-20260727-creds-consolidation.md. Superseding figures: matrix 82 rows, creds-matrix harness 60/60 (was 56; V2 had shipped with ZERO cases), creds-audit CLEAN on all three sites, gauntlet ALL GREEN (81), repo-lint 0-fail, findings 13 -> 7. VERIFIED POSITIVE, both previously only asserted: the MAAS account set is COMPLETE -- all 6 accounts enumerated live (maas admin users read) are accounted for (admin + operator passwords on vcloud, juju-vr1-dc0/dc1 random+unstored BY RULING with their API keys present, MAAS/maas-init-node MAAS-internal) -- and tier-3 V1 now MEASURES maas-admin-password byte-identical to the headend source-of-record, so the stale-trap risk SEC-020 records is clear as of this date. DONE: dc0's SEC-012 power key consolidated to vcloud + .pub DERIVED (SEC-021(b) as written; measured first -- the headend maas-virsh_ed25519 and the snap's id_ed25519 are the SAME key, it IS dedicated (distinct from the dc0 service key), and dc0 using the snap default identity is SEC-016's ruled design, so NO re-mint and no live power path touched); dc1 svc .pub backfilled to the headend; NetBox web-GUI admin password consolidated (SEC-025 OPENED for the at-rest exposure the copy creates -- open rows 20 -> 21); V2 taught the ruled-deferral state so SEC-006's standing "revoke at completion of this deployment" ruling is ACKNOWLEDGED (still naming the credential as live and exposed) instead of failing every run -- reissuing it would have CONTRAVENED that ruling. Tier 3 is not in preflight P5, so the blocking gate's behaviour is unchanged. RESIDUAL 7 findings, expected, NOT green: dc0-edge-api x2 (the opnsense-api.txt re-mint is a live edge mutation, deliberately EXCLUDED from the batch -- sole remaining SEC-021(a) item), S5 x3 (RULED by SEC-016, the register needs a ruled-exception mechanism -- operator decision), S6 conflation x1 (the SEC-020 defect), E4 uncheckable x2 (Stage-5/6 rows). NEW FINDINGS LOGGED NOT ACTIONED: (i) NO registered root/console credential at EITHER DC edge -- measured absence of row, manifest entry and SEC row; what those passwords ARE is UNKNOWN and deliberately unprobed (hard rule 2), vector is the LAN-reachable GUI + serial console, not SSH (key-only, proven); (ii) two structural blind spots that let (i) hide -- S5 compares only cardinality=per-DC while all six per-site rows are office1-only, and vm-secret-locations declares no rack/edge/cloud/unit/client location though the checker accepts them (SEC-015 was rack-resident, so the class is real). The lesson generalises D-137's founding argument one level up: absence of a ROW is invisible to the register, so enumerating what EXISTS is a distinct control from auditing what is declared. RE-RUN 2026-07-27 (Stage-5 grounding audit), operator-authorized privileged sweep: python3 scripts/creds-matrix.py --tier2 --remote --privileged -> exit 1, capture docs/audit/stage5-creds-privileged-20260727.txt. This is a RE-CONFIRMATION of the 2026-07-26 privileged run against the current 82-row matrix, NOT a newly-closed item. Result: still exactly 7 findings, the same set (S2 dc0-edge-api, S5 x3 power-key asymmetry, S6 conflation, E1 dc0-edge-api on-disk, E4 x2 uncheckable) -- no drift in a day. All three root-owned locations read successfully via ESCALATION (sudo -n, metadata only): /root/maas-secrets/*, /var/snap/maas/current/root/.ssh/*, /root/netbox-secrets/*. ZERO E3 findings -- no undeclared file at any declared location. Note the checker's own honesty on the V1 arm, worth preserving: "V1 provenance: no identity had two digestible copies -- NOTHING was verified here; this is a skip, not a pass." QUEUED FINDINGS CAPTURED 2026-07-26 at session close: docs/audit/queued-findings-20260726.txt -- an end-of-session sweep for content that existed ONLY in the session transcript. Part A: the three secrets-storage items NOT already repo-carried (Tang/Clevis as the no-HSM unseal mechanism; MAAS 3.7's Vault integration MEASURED status: disabled, with the MAAS/Vault circular dependency that must be designed around before enabling it on bare metal; a Vault SSH CA to retire the static keypairs) plus sequencing advice -- framed so a future session does not re-propose what D-068's analysis and D-137 item 2(a) ALREADY carry. Part B: ten committee findings ACKNOWLEDGED but deliberately NOT acted on, the most consequential being that mint-ref line pins rot SILENTLY (S4 checks existence and EOF, never content, so every pin becomes wrong-but-passing when the runbooks are rewritten for Roosevelt) and that ruling-5 REMEDIATION and ruling-5 EVASION are indistinguishable to S6. Part C: items deferred by ruling, recorded so they are not later mistaken for oversights. NONE of it is ruled or built; cardinality (B6) needs an operator ruling.2026-07-19 -- wrap-aware exclusion, GA-F15; history below)
ledger-scan.sh's mention-derived next-free counters (DOCFIX, BUNDLEFIX) are SELF-INFLATED by any doc that quotes a "next-free" value and lets the hyphenated token wrap onto a line without the words "next-free" -- the per-line exclusion filter (scripts/ledger-scan.sh:121-129, the grep -viE 'next[- ]free' at :124) then counts the quote as a real assignment; the script's own CAUTION comment (:112-120) documents exactly this failure class. It happened TWICE inside the audit itself on 2026-07-18: the Phase-1 env snapshot's wrapped next-free line inflated the BUNDLEFIX counter (051 -> reported 052), and this document's own first draft of this very section inflated the DOCFIX counter the same way while asserting DOCFIX was unaffected. Both audit surfaces were reworded token-free the same day and the counters re-verified at their true values (D=130, DOCFIX=197, BUNDLEFIX=052 -- see GA-F15). The D counter is header-authoritative and was never affected. Batch 0.3 hardened the scanner: an excluded line now also suppresses the immediately following line (the wrap case); counters re-verified unchanged post-fix. The authoring discipline stands regardless: never write a hyphenated register-token quote of a next-free value into any doc; state the numbers token-free as this section does.
Run read-only, from the repo root:
git rev-parse HEAD; git status --short (this doc was authored at e999b03, clean tree).tofu -chdir=opentofu state list (expect the 20 resources in section 2.1); virsh list --all; virsh dominfo voffice1 | grep -i autostart (and office1-opnsense).grep -n '^module ' opentofu/main.tf (12 blocks; vvr1_dc0 + vr1_dc0_uplink absent from state list); ls opentofu/vr1-dc0-substrate/ (no *.tfstate).docs/audit/outer-plan-20260718.txt line 428. Do NOT re-run tofu plan casually against live state; if a fresh capture is taken, it must be written to a new dated capture file and cited here.tofu version; ssh voffice1 'snap list maas lxd; uname -r' </dev/null; uname -r; grep -A2 provider opentofu/.terraform.lock.hcl.bash scripts/ledger-scan.sh (BUNDLEFIX caveat: section 9); decision status lines: grep -n '^## D-' docs/design-decisions.md then read each Status line -- a decision's Status line in that file is the ONLY ruling authority.docs/audit/grounding-audit-charter.md; docs/audit/grounding-audit-20260718.md for GA-F01..F14.ssh office1-netbox 'curl -s -o /dev/null -w "%{http_code}" http://localhost:8000/' </dev/null (expect 302); ssh office1-tailscale 'tailscale status | head -1' </dev/null.What this document is NOT built from and you must not rebuild it from: the prose of the 95 docs/changelog-*.md files, the docs/session-ledger.md narrative, or auto-memory -- all proven to carry false status (GA-F14, GA-F06..F08).
SIGNED 2026-07-19 (re-signature at audit exit; REPLACES the 2026-07-18 signature per GA-R1 rule 7 -- git history keeps it). Question as presented (Batch 6 item 6, 2026-07-19): read this document top to bottom, then provide the signature statement. Operator answer, exact utterance: "Reviewed, approved, continue." This document is the signed status authority; charter Phase 6 item 5 MET at this baseline (repo HEAD at signing recorded in the close commit).