Newer
Older
openstack-caracal-dc-dc / docs / changelog-20260730-stage5-open.md

Changelog 2026-07-30 (part 3) -- STAGE 5 OPENED: the Juju deployment

Session changelog part 3 (part 1 = changelog-20260730-docfix205-d117-annotation.md, part 2 = changelog-20260730-octavia-reissue-tool.md). Branch dc-dc-stage5-preconditions. Status claims live ONLY in docs/CURRENT-STATE.md.

Trigger. The standing operator directive recorded at the 2026-07-30 part-2 close: "we have to continue to juju deployment next session no matter what". This session opens Stage 5 and runs the deployment.

No new D-number (GA-R3). Opening a stage is OPS; the P5 acceptance is an operational gate disposition against existing SEC rows, not architecture. Next-free UNCHANGED: D 138 / DOCFIX 206 / BUNDLEFIX 053.


Operator rulings recorded (GA-R5, one exchange each, verbatim)

  1. P5 GATE -- "Accept and proceed to deploy (Recommended)". Question as presented and the full consequence text are in docs/CURRENT-STATE.md section 1 (the status authority). The six enumerated findings are accepted, known, pre-existing exposure; preflight.sh continues to exit FAIL on P5 for the stage's duration and that RED is ruled-accepted. The acceptance covers those six and nothing else.
  2. BOOTSTRAP CONSTRAINT FLAG -- "Use both flags". Question as presented and the full consequence text are in docs/CURRENT-STATE.md section 1. The executed bootstrap carries BOTH --bootstrap-constraints and --constraints. D-104 is NOT amended; only the flag implementing it is clarified.

Items

1. Stage 5 OPENED; the P5 acceptance ruling recorded; entry gate captured

What. docs/CURRENT-STATE.md section 1 gains the Stage-5 OPEN entry: the branch decision and why, the measured entry gate, the P5 ruling with the question and the operator's exact utterance, and one logged-not-executed finding. New capture docs/audit/stage5-preflight-dc0-20260730.txt (237 lines, DC=vr1-dc0 bash scripts/preflight.sh run ON voffice1, exit 1).

Why (evidence). GA-R1/C1 puts the status change and the document update in one commit; GA-R5 requires the ruling committed and pushed before dependent work. Four read-only checks were made BEFORE putting the question, so it was asked once and asked grounded:

  • Both clones at c58bf95, same branch, clean. voffice1 was found on a 105-commit-stale retired branch at the 2026-07-27 close and the "back to main at merge" follow-up never fired, because nothing merged. Verified rather than assumed: a stale clone silently deploys the wrong bundle.
  • Preflight run ON voffice1, not here. P5 was found probing the wrong host's filesystem on 2026-07-30 (34 findings vs the true 7), P7 is headend-only, and one gate was found RED on the only host that deploys. A vcloud reading is not the gate reading.
  • Deploy artifacts verified as FILES (bundle.yaml VR1 9-node role-separated; both per-DC -vips/-machines/-octavia-pki overlays present, the PKI pair 0600 and gitignored). This is the standing "RULED IS NOT BUILT -- check the artifact" rule applied to the 2026-07-24 committee record, which is now superseded by measurement.
  • The deploy-order ruling was READ, not taken from its summary (CURRENT-STATE.md:986). There is no ruled DC ordering; dc1-first artifacts are not a divergence.

Revert. git revert <this commit> -- it removes the Stage-5 OPEN entry, the recorded ruling and the capture reference. The capture file itself can be deleted separately; it is evidence, not configuration, and nothing reads it. Reverting the ruling does NOT un-ask the question: re-asking would need a fresh GA-R5 exchange.

2. LOGGED, NOT EXECUTED -- ceph-osd carries the stale VR0 constraint tags=openstack

What. bundle.yaml:592. Recorded in docs/CURRENT-STATE.md section 1; no edit made.

Why (evidence). ceph-osd is the ONLY application of 56 carrying a tag constraint -- every other reads arch=amd64 alone (parsed from bundle.yaml, not grepped). The tag is measured ABSENT from the VR1 region: maas admin tags read returns virtual, pod-console-logging, serial-console, openstack-vr1-dc0, openstack-vr1-dc1, control, compute, storage, juju-controller-vr1-dc0, juju-controller-vr1-dc1 -- no bare openstack. Neither machines overlay overrides it (vr1-dc0-machines.yaml is applications:-only and says so).

What is NOT measured, and is labelled as such. ceph-osd has explicit placement (to: ["5","6","7","8"]), so the initial deploy is EXPECTED to place by machine id regardless. The reasoned-not-measured exposure is a later unplaced juju add-unit ceph-osd matching no machine. The real impact is observable at Step 4.2's --dry-run and nowhere earlier, which is why it is recorded now and decided there. Hard rule 1 forbids fixing it mid-step, and the standing rule is that a finding is an observation, not a conclusion.

Revert. Nothing to revert -- no artifact was changed.

3. OWED -- DOCFIX for the Step 2 bootstrap command and D-104's mechanism sentence

What. Not yet assigned a number (next-free DOCFIX is 206; assign at the point of delivery, per the standing numbering rule -- do not write the token above the high-water mark in prose before it exists). Two surfaces state the bootstrap machine is targeted by juju bootstrap --constraints tags=juju-controller-$DC: runbooks/dc-dc-phase4-juju-bundle-per-dc.md:277-278 and the D-104 amendment's "Distinct tag, no role tag" bullet in docs/design-decisions.md.

Why (evidence). Juju 3.6.27's own help, read ON voffice1 rather than from memory, assigns machine-targeting to --bootstrap-constraints and model-defaulting to --constraints (both sentences quoted verbatim in docs/CURRENT-STATE.md section 1). The operator ruled "Use both flags", so the live command is correct; the RUNBOOK is what still reads wrong, and a future DC standup following it literally would get the single-flag form. No gate reads prose, so nothing in this repo can catch it -- the same class as the citation defect the 2026-07-27 audit recorded.

Scope note. The DOCFIX corrects the FLAG only. D-104's decision -- a dedicated per-DC controller VM carrying its own tag and none of the role tags -- is untouched and was independently confirmed live this session: moved-troll (7n87bt) carries juju-controller-vr1-dc0 and no openstack-vr1-dc0.

Revert. N/A -- nothing delivered yet; this item is the record that it is owed.

4. Bootstrap ATTEMPTED and FAILED; root cause captured; Stage 5 blocked on a ruling

What. New capture docs/audit/stage5-bootstrap-reachability-20260730.txt. Status recorded in docs/CURRENT-STATE.md section 1. No reachability change was made.

Why (evidence). juju bootstrap selected the correct machine and MAAS deployed jammy end to end (Image Deployed -- deployed ubuntu/jammy/amd64/generic), then juju could not SSH it and released the machine. voffice1 has no route to any DC node plane, and two independent, DELIBERATE controls forbid creating one: SEC-010's transit FORWARD-drop on the rack, and libvirt's blanket reject into the isolated plane bridges. The DC edge has no metal-admin leg. This is a contradiction between ruled surfaces -- SEC-010/D-052 make metal-admin DC-local and forbid the region routing to 10.12.8.0/22, while D-100 says the fiber carries Juju traffic and D-128 puts the Juju client on voffice1 -- so it needs a ruling, not a firewall edit.

The generalisable lesson, and it is the reason this was invisible for weeks. SEC-010's cost was assessed with the sentence "a MAAS rack proxies at the application layer and needs no kernel forwarding, so pinning is free." That is true of every Plane-2 tool THEN in use -- MAAS, the inner tofu root over qemu+ssh, NetBox -- and false of the one tool that had not been run yet. A security control priced against the tools you have is not priced against the tools the next stage introduces. scripts/site-baseleg.sh:42,47 recorded the same blind spot from the other side, deferring the DC rows because "the DCs nest inside vvr1-dc0 and are reached by qemu+ssh" with a # MEASURE first note that was never executed.

Two claim-discipline corrections made to the capture before it was committed, both worth keeping as examples: (i) the post-failure machine fields (osystem:"", status:Ready/Released) were first read as "the machine never booted" -- those fields are CLEARED BY THE RELEASE and describe the aftermath, not the attempt; the event log is the authority and it shows a clean full deploy. (ii) The capture initially said the both-flags form "targeted it deterministically"; one run cannot attribute the selection to either flag, only show that A constraint preferred the tagged machine over nine role-tagged candidates.

Revert. The capture and the CURRENT-STATE entry are records of a live event; reverting them would delete evidence, not undo a change. Nothing was mutated on the cloud by this item.

5. D-134 amendment BUILT on both controllers (was ruled-but-not-built)

What. vr1-dc0-juju-01 (7n87bt, iface 419) and vr1-dc1-juju-01 (p8tdwg, iface 426) converted from a single AUTO v4 link to STATIC dual-stack: 10.12.8.5 + fd50:840e:74e2:220::5, and 10.12.68.5 + fd50:840e:74e2:320::5. Read back and asserted on content.

Why (evidence). Operator ruling 2026-07-30, exact utterance "Fix now: static .5 + v6, re-bootstrap (Recommended)". Both controllers held AUTO, v4-only addresses while all nine role nodes per DC hold STATIC dual-stacked addresses in their D-134 bands -- the controller VMs were added on 2026-07-29, after both the D-134 carve and the IPv6 carve, so neither reached them. The D-134 amendment of 2026-07-29 ruled .5 "on every plane it is attached to" and made the octet map a STANDING cross-DC standard; it was ruled-but-not-built until this item, the same class the 2026-07-27 audit found twice.

AUTO was the live hazard, not the octet. An auto address is not pinned: a future redeploy of the controller machine could return a different address, and the controller address is written into every unit's agent.conf, so every agent in the cloud would lose its controller at once.

Values were derived, not typed. The v6 prefixes were confirmed by VLAN pairing -- dc0 v4 subnet 6 and v6 subnet 21 are both vlan 5005/fabric-4; dc1 v4 subnet 11 and v6 subnet 27 are both vlan 5143/fabric-142 -- combined with the 2026-07-27 ruling that the v6 host part mirrors the v4 octet. Inferring the prefix from a single node's address would have been the hard-rule-2 violation.

Revert. Per controller: maas admin interface unlink-subnet <sys> <iface> id=<link> for each of the two new links, then maas admin interface link-subnet <sys> <iface> mode=AUTO subnet=<v4 subnet> to restore the original single auto link. dc0 = 7n87bt iface 419 subnet 6; dc1 = p8tdwg iface 426 subnet 11. Both machines are Ready and powered off, so this is safe at any time before they are allocated.