Session changelog part 3 (part 1 = changelog-20260730-docfix205-d117-annotation.md, part 2 = changelog-20260730-octavia-reissue-tool.md). Branch dc-dc-stage5-preconditions. Status claims live ONLY in docs/CURRENT-STATE.md.
Trigger. The standing operator directive recorded at the 2026-07-30 part-2 close: "we have to continue to juju deployment next session no matter what". This session opens Stage 5 and runs the deployment.
No new D-number (GA-R3). Opening a stage is OPS; the P5 acceptance is an operational gate disposition against existing SEC rows, not architecture. Next-free UNCHANGED: D 138 / DOCFIX 206 / BUNDLEFIX 053.
docs/CURRENT-STATE.md section 1 (the status authority). The six enumerated findings are accepted, known, pre-existing exposure; preflight.sh continues to exit FAIL on P5 for the stage's duration and that RED is ruled-accepted. The acceptance covers those six and nothing else.docs/CURRENT-STATE.md section 1. The executed bootstrap carries BOTH --bootstrap-constraints and --constraints. D-104 is NOT amended; only the flag implementing it is clarified.What. docs/CURRENT-STATE.md section 1 gains the Stage-5 OPEN entry: the branch decision and why, the measured entry gate, the P5 ruling with the question and the operator's exact utterance, and one logged-not-executed finding. New capture docs/audit/stage5-preflight-dc0-20260730.txt (237 lines, DC=vr1-dc0 bash scripts/preflight.sh run ON voffice1, exit 1).
Why (evidence). GA-R1/C1 puts the status change and the document update in one commit; GA-R5 requires the ruling committed and pushed before dependent work. Four read-only checks were made BEFORE putting the question, so it was asked once and asked grounded:
c58bf95, same branch, clean. voffice1 was found on a 105-commit-stale retired branch at the 2026-07-27 close and the "back to main at merge" follow-up never fired, because nothing merged. Verified rather than assumed: a stale clone silently deploys the wrong bundle.bundle.yaml VR1 9-node role-separated; both per-DC -vips/-machines/-octavia-pki overlays present, the PKI pair 0600 and gitignored). This is the standing "RULED IS NOT BUILT -- check the artifact" rule applied to the 2026-07-24 committee record, which is now superseded by measurement.CURRENT-STATE.md:986). There is no ruled DC ordering; dc1-first artifacts are not a divergence.Revert. git revert <this commit> -- it removes the Stage-5 OPEN entry, the recorded ruling and the capture reference. The capture file itself can be deleted separately; it is evidence, not configuration, and nothing reads it. Reverting the ruling does NOT un-ask the question: re-asking would need a fresh GA-R5 exchange.
ceph-osd carries the stale VR0 constraint tags=openstackWhat. bundle.yaml:592. Recorded in docs/CURRENT-STATE.md section 1; no edit made.
Why (evidence). ceph-osd is the ONLY application of 56 carrying a tag constraint -- every other reads arch=amd64 alone (parsed from bundle.yaml, not grepped). The tag is measured ABSENT from the VR1 region: maas admin tags read returns virtual, pod-console-logging, serial-console, openstack-vr1-dc0, openstack-vr1-dc1, control, compute, storage, juju-controller-vr1-dc0, juju-controller-vr1-dc1 -- no bare openstack. Neither machines overlay overrides it (vr1-dc0-machines.yaml is applications:-only and says so).
What is NOT measured, and is labelled as such. ceph-osd has explicit placement (to: ["5","6","7","8"]), so the initial deploy is EXPECTED to place by machine id regardless. The reasoned-not-measured exposure is a later unplaced juju add-unit ceph-osd matching no machine. The real impact is observable at Step 4.2's --dry-run and nowhere earlier, which is why it is recorded now and decided there. Hard rule 1 forbids fixing it mid-step, and the standing rule is that a finding is an observation, not a conclusion.
Revert. Nothing to revert -- no artifact was changed.
What. Not yet assigned a number (next-free DOCFIX is 206; assign at the point of delivery, per the standing numbering rule -- do not write the token above the high-water mark in prose before it exists). Two surfaces state the bootstrap machine is targeted by juju bootstrap --constraints tags=juju-controller-$DC: runbooks/dc-dc-phase4-juju-bundle-per-dc.md:277-278 and the D-104 amendment's "Distinct tag, no role tag" bullet in docs/design-decisions.md.
Why (evidence). Juju 3.6.27's own help, read ON voffice1 rather than from memory, assigns machine-targeting to --bootstrap-constraints and model-defaulting to --constraints (both sentences quoted verbatim in docs/CURRENT-STATE.md section 1). The operator ruled "Use both flags", so the live command is correct; the RUNBOOK is what still reads wrong, and a future DC standup following it literally would get the single-flag form. No gate reads prose, so nothing in this repo can catch it -- the same class as the citation defect the 2026-07-27 audit recorded.
Scope note. The DOCFIX corrects the FLAG only. D-104's decision -- a dedicated per-DC controller VM carrying its own tag and none of the role tags -- is untouched and was independently confirmed live this session: moved-troll (7n87bt) carries juju-controller-vr1-dc0 and no openstack-vr1-dc0.
Revert. N/A -- nothing delivered yet; this item is the record that it is owed.
What. New capture docs/audit/stage5-bootstrap-reachability-20260730.txt. Status recorded in docs/CURRENT-STATE.md section 1. No reachability change was made.
Why (evidence). juju bootstrap selected the correct machine and MAAS deployed jammy end to end (Image Deployed -- deployed ubuntu/jammy/amd64/generic), then juju could not SSH it and released the machine. voffice1 has no route to any DC node plane, and two independent, DELIBERATE controls forbid creating one: SEC-010's transit FORWARD-drop on the rack, and libvirt's blanket reject into the isolated plane bridges. The DC edge has no metal-admin leg. This is a contradiction between ruled surfaces -- SEC-010/D-052 make metal-admin DC-local and forbid the region routing to 10.12.8.0/22, while D-100 says the fiber carries Juju traffic and D-128 puts the Juju client on voffice1 -- so it needs a ruling, not a firewall edit.
The generalisable lesson, and it is the reason this was invisible for weeks. SEC-010's cost was assessed with the sentence "a MAAS rack proxies at the application layer and needs no kernel forwarding, so pinning is free." That is true of every Plane-2 tool THEN in use -- MAAS, the inner tofu root over qemu+ssh, NetBox -- and false of the one tool that had not been run yet. A security control priced against the tools you have is not priced against the tools the next stage introduces. scripts/site-baseleg.sh:42,47 recorded the same blind spot from the other side, deferring the DC rows because "the DCs nest inside vvr1-dc0 and are reached by qemu+ssh" with a # MEASURE first note that was never executed.
Two claim-discipline corrections made to the capture before it was committed, both worth keeping as examples: (i) the post-failure machine fields (osystem:"", status:Ready/Released) were first read as "the machine never booted" -- those fields are CLEARED BY THE RELEASE and describe the aftermath, not the attempt; the event log is the authority and it shows a clean full deploy. (ii) The capture initially said the both-flags form "targeted it deterministically"; one run cannot attribute the selection to either flag, only show that A constraint preferred the tagged machine over nine role-tagged candidates.
Revert. The capture and the CURRENT-STATE entry are records of a live event; reverting them would delete evidence, not undo a change. Nothing was mutated on the cloud by this item.
What. vr1-dc0-juju-01 (7n87bt, iface 419) and vr1-dc1-juju-01 (p8tdwg, iface 426) converted from a single AUTO v4 link to STATIC dual-stack: 10.12.8.5 + fd50:840e:74e2:220::5, and 10.12.68.5 + fd50:840e:74e2:320::5. Read back and asserted on content.
Why (evidence). Operator ruling 2026-07-30, exact utterance "Fix now: static .5 + v6, re-bootstrap (Recommended)". Both controllers held AUTO, v4-only addresses while all nine role nodes per DC hold STATIC dual-stacked addresses in their D-134 bands -- the controller VMs were added on 2026-07-29, after both the D-134 carve and the IPv6 carve, so neither reached them. The D-134 amendment of 2026-07-29 ruled .5 "on every plane it is attached to" and made the octet map a STANDING cross-DC standard; it was ruled-but-not-built until this item, the same class the 2026-07-27 audit found twice.
AUTO was the live hazard, not the octet. An auto address is not pinned: a future redeploy of the controller machine could return a different address, and the controller address is written into every unit's agent.conf, so every agent in the cloud would lose its controller at once.
Values were derived, not typed. The v6 prefixes were confirmed by VLAN pairing -- dc0 v4 subnet 6 and v6 subnet 21 are both vlan 5005/fabric-4; dc1 v4 subnet 11 and v6 subnet 27 are both vlan 5143/fabric-142 -- combined with the 2026-07-27 ruling that the v6 host part mirrors the v4 octet. Inferring the prefix from a single node's address would have been the hard-rule-2 violation.
Revert. Per controller: maas admin interface unlink-subnet <sys> <iface> id=<link> for each of the two new links, then maas admin interface link-subnet <sys> <iface> mode=AUTO subnet=<v4 subnet> to restore the original single auto link. dc0 = 7n87bt iface 419 subnet 6; dc1 = p8tdwg iface 426 subnet 11. Both machines are Ready and powered off, so this is safe at any time before they are allocated.
What. New ## D-138 in docs/design-decisions.md (ARCH); new SEC-026 row in docs/security-ledger.md; docs/CURRENT-STATE.md section 1 updated in the same commit (GA-R1/C1). Counters after: 22 open SEC, next-free D 139 / DOCFIX 206 / BUNDLEFIX 053, reconciled by bash scripts/ledger-scan.sh.
Why (evidence). Operator ruling 2026-07-30, exact utterance "Move the cloud-facing client into the DC (Recommended)", taken on the Stage-5 bootstrap failure. Admitted as a D-number under GA-R3's A1 test: it decides where the Juju client lives at every future DC standup, and it AMENDS D-128. Full question, options and consequences are in the D-138 entry.
The ruling costs nothing in security posture, which is why it is the smaller change. SEC-010, D-052's DC-local invariant and D-125's proven egress isolation are all untouched; no rule was punctured and no plane was opened. The alternative would have opened two planes per DC across a dozen ports and required D-125 to be re-tested.
SEC-026 is a consequence, not an afterthought. D-138 puts a MAAS ADMIN-scoped key on a DC-local host, and the MAAS region is shared across both DCs -- so that host has region-wide MAAS admin including over the other DC's hardware. The row's load-bearing control is isolation: each DC's client host gets ONLY its own credential. Copying voffice1's whole credentials.yaml (which holds both vr1-dc0-cred and vr1-dc1-cred) to a DC rack would, in one act, destroy the per-DC isolation SEC-018/-019 exist to create.
repo-lint L10 did its job. The first attempt to land D-138 alone FAILED the pending change-set check -- a Status-bearing surface changed without CURRENT-STATE.md in the same commit. Recorded because it is a gate proving it can fail, on a real commit rather than a fixture.
Revert. git revert <this commit> removes D-138, SEC-026 and the CURRENT-STATE paragraph together. Nothing was executed on the cloud under this ruling yet, so no live state depends on it at the time of writing.
What. vr1-dc0-maas-01 and vr1-dc1-maas-01 added to their substrate roots (4 vCPU / 8 GiB / 150 GiB). dc0 APPLIED: 2 added, 0 changed, 0 destroyed; MACs pinned from measurement and applied; converged zero diff (tofu plan -detailed-exitcode -> 0); virsh power set; enlisted as hot-kid / tw7ptw and commissioning. dc1 is AUTHORED ONLY, not applied.
Why (evidence). D-132's amendment + addendum, both RULED 2026-07-30. Sizing is the shape the capacity gate actually modelled -- dc-dc-whole-host-budget.py --containment-overhead-mem-gib 32 --containment-overhead-vcpu 12 -> RAM 870/1024 = 85%, FIT, 154 GiB headroom; a 16 GiB variant also fits (886/1024 = 87%). Disk is 150 GiB rather than the controller's 100 because this VM holds the region's PostgreSQL AND its boot-image set. The plan was asserted on CONTENT, not on its summary line: tofu show <plan> | grep "will be (created|destroyed|updated|replaced)" returned exactly two lines, both vr1-dc0-maas-01.
A count correction, made against measurement rather than inherited prose. Both new entries' comments initially claimed the addition "plans 1 add / 0 change / 0 destroy", carried from the D-104 amendment's text. MEASURED: it is 2 to add -- modules/node-vm creates a libvirt_domain AND a libvirt_volume per node. The load-bearing half of that figure is 0 change / 0 destroy, which held. D-104's "1 add" figure appears to be wrong at resource level for the same reason; flagged, not edited, since it is not this item's surface.
THE MAC-PIN BOUNCE INTERRUPTED COMMISSIONING -- a real ordering trap, recorded because the runbook rule as written cannot avoid it. The standing rule is: apply with macs = [], then pin from measurement "immediately after first apply, before enlistment". But the first apply sets running = true, so the VM PXE-boots and BEGINS ENLISTING AND COMMISSIONING on its own within ~90 seconds. The pin is therefore always racing an auto-commission, and applying it is an in-place libvirt_domain update -- which bounces the guest (Still modifying... 32s elapsed; platform-traps already records "in-place bounces the guest"). Measured result: domain shut off, MAAS stuck at Loading ephemeral, commissioning stalled. Recovery, clean and cheap: maas admin machine abort -> New, then maas admin machine commission -> MAAS powers it on via the virsh power type just set. The generalisable rule: after pinning MACs on a freshly-applied VM, ASSUME commissioning was interrupted and re-commission deliberately. Do not read a stalled Commissioning as a fault. The alternative -- pinning before the VM ever boots -- is impossible while the module creates domains running = true.
Verification that the pin is real, not just written. Live virsh domiflist after the in-place update returns all six MACs byte-identical to the pinned values, and the root re-plans to ZERO DIFF. Both were checked because an in-place domain update is precisely where this project regenerated MACs and stranded nine nodes on 2026-07-20.
Revert. tofu destroy -target='module.vr1_dc0_node["vr1-dc0-maas-01"]' in opentofu/vr1-dc0-substrate removes the VM and its volume; then delete the MAAS machine record (maas admin machine delete tw7ptw) and git revert this commit to drop both roots' entries. Nothing else in either root is touched -- the plan proved 0 change / 0 destroy.
What. New capture docs/audit/queued-findings-20260730-stage5.txt (F1-F12), and an logs/as-executed-index.md row for the stage5-dc0-deploy window.
Why (evidence). A close-time sweep grepped each finding raised in-session against docs/CURRENT-STATE.md and this changelog. Three lived ONLY in the transcript and would have been lost: F1 the absent openstack CLI on the DC client host (blocks Step 7+), F2 the metal-admin v6 rack-leg gap, F3 the plane-fabric mtu=1500 vs libvirt mtu 9000 mismatch. Two more were partial: F5's run-location half and F6's as-executed gap. The rest are cross-references so the sweep is complete rather than selective.
The one that most needed writing down is F6 -- the as-executed log is PARTIAL. The harness classifier refused several script -aqe ... -c wrapped forms while the plain ssh voffice1 'maas admin ...' form passed, so those calls ran unwrapped. Every action is in the changelog and CURRENT-STATE with read-backs, but the log itself must not be read as a complete record of this window, and the index row now says so. A log that silently looks complete is worse than one that declares its gap.
F9 collects six tool traps measured today -- ssh -n eating a heredoc (and its mirror-image, an inner ssh without -n eating the rest of the script, which silently skipped a credential shred), | grep -q under pipefail, snap /tmp confinement, a MAAS write returning a full JSON object while changing nothing, the space-aligned .tsv, and boot-resources import racing a fresh selection. They are candidates for references/platform-traps.md, queued not executed.
Revert. Records only; nothing executable changed. git revert drops the capture, the index row and this item.