Newer
Older
openstack-caracal-dc-dc / docs / archive / session-ledger-rotated-20260802b.md

Session-ledger rotation 2026-08-02 (b)

Rotated VERBATIM out of docs/session-ledger.md at the second 2026-08-02 close (GA-R4 rule 3 / F1). The live ledger stood at 294 lines and this close's bounded summary would have breached the 300-line cap.


SESSION CLOSE 2026-07-30 (part 3) -- STAGE 5 OPENED; three bootstraps; D-138 + D-132 ruled; per-DC MAAS region LIVE at dc0 (bounded, GA-R4)

  • Branch dc-dc-stage5-preconditions, 12 commits pushed (726127d..). STAGE 5 OPENED, not closed. Scan: 3 decisions, SEC 23 (SEC-026, -027 opened), D 139 / DOCFIX 206 / BUNDLEFIX 053.
  • 5 RULINGS (GA-R5, quoted): P5 accepted; bootstrap flags "Use both flags"; D-138 cloud-facing client moves INTO the DC; D-132 q1 per-DC MAAS region; its placement VM at utility .6 (extends the D-134 octet map: .4 artifact / .5 juju / .6 region).
  • THREE BOOTSTRAP ATTEMPTS, each failing one layer deeper -- every failure a real defect no record would have shown. (1) no juju-client->node-plane path anywhere (SEC-010 + isolated libvirt nets; never existed). (2) D-138 fixed the client, then the CONTROLLER VM had NO DEFAULT ROUTE -- under-carved, added after both Stage-4 carves. (3) route fixed, agent fetched on attempt 1, then jujud could not reach the MAAS region API. SEC-010 is INCOMPATIBLE with D-104-in-DC-controller + Office1 region; per-DC regions REMOVE the requirement rather than excepting it, so SEC-010 stays UNAMENDED.
  • BUILT: D-134 .5+v6 on BOTH controllers (was ruled-but-not-built); region VM applied dc0 (2 add/0/0, MACs pinned, converged ZERO DIFF); dc0 region LIVE -- noble + MAAS 3.7.2 + PostgreSQL 16.14 matching Office1, API 200 on 10.12.8.6:5240, dhcpd INACTIVE (no DHCP conflict), jammy+noble Synced.
  • G17 dc0 CAPTURED, both assertions PASS -- the one-shot window opened during a FAILED bootstrap and was taken. Found G17's own text names chronyc, absent on the MAAS jammy image, so the gate as written could only REFUSE.
  • CREDENTIALS (operator-directed): 3 region secrets minted on-VM, never printed, consolidated to ~/vr1-dc0-creds/ (sha256-verified identical), SEC-027 opened; matrix now EXPECTS them at BOTH DCs (5 rows/DC, new region host-role enum, 4 vm-secret-locations rows). dc1's correctly FAIL S2 -- the forward register making absence detectable. Harness 65/65.
  • MEASURED, DO NOT CONFLATE: MAAS boot images come from images.maas.io, NOT the DC mirror (404s on every simplestreams path -- it mirrors apt only). Three artifact classes, ONE local. The D-135/D-107 narrowing would BREAK image sync AND bootstrap; nothing sequences it after them (queued F10 -- highest-consequence unruled item).
  • OWNED: ran the DC checkers on the wrong host; then re-probed egress FROM THE RACK and called the window open when the NODE does the fetching -- same wrong-host class, made right after writing that note; scoped D-138 without enumerating that jujud is itself a MAAS client, which is why a 4th blocker appeared after the 3rd was fixed. All corrected on-surface.
  • NOTHING IS HALF-APPLIED. Office1 still owns all 9 dc0 nodes, correctly carved and Ready. Pre-migration carve captured (dc0-maas-carve-premigration-20260730.txt, 179 lines) -- MAAS cannot move machines between regions, so the migration must recreate it.
  • AS-EXECUTED LOG IS PARTIAL -- the classifier refused several script -aqe wrapped forms; index row says so (F6). A log that looks complete is worse than one declaring its gap.
  • NEXT: recreate topology on the new region -> DHCP handover (two servers on one segment is the hazard) -> re-enrol/re-commission 9 nodes -> re-apply statics/br-ex/tags -> re-point Juju at 10.12.8.6:5240. dc1's region VM authored, NOT applied. Start with fresh context: the destructive half is nine nodes.
  • Sweep: docs/audit/queued-findings-20260730-stage5.txt (F1-F12). Bodies: docs/changelog-20260730-stage5-open.md. Status ONLY in CURRENT-STATE.md.

SESSION CLOSE 2026-07-30 (part 4) -- dc0 region topology BUILT 40/0; cutover BLOCKED on a permission wall (bounded, GA-R4)

  • Branch dc-dc-stage5-preconditions, 8 commits pushed (0ec9c97..b1afa42). NO stage opened/closed. Scan unchanged: 3 decisions, SEC 23, D 139 / DOCFIX 206 / BUNDLEFIX 053.
  • STAGE 5 REMAINS BLOCKED -- now on the DHCP cutover, refused by the harness classifier and NOT retried in an altered shape. Exact 4-command sequence, every parameter re-resolved live, is in changelog-20260730-dc0-region-migration.md item 10. NOTHING half-applied (Office1 verified read-only still dhcp_on=True primary_rack=7chphy).
  • NEW dc0 REGION TOPOLOGY LIVE: dc-region-topology.sh check = 40 assertions / 0 failed, EXIT 0. 5 named plane fabrics, 6 spaces, 6 v4 subnets + gateways, 6 VLAN->space bindings, site tag, metal-admin dns/dynamic range. Power key installed + proven (9/9, real virsh enumerating 12 domains -- the snap ssh dir did not exist at all, so commissioning could not have powered a node on).
  • THREE TOOLS THE REPO NEVER HAD: maas-profile-assert.sh (region identity by RACK IDENTITY -- a machine count is not proof), maas-region-power-key.sh, dc-region-topology.sh. A survey found the named fabrics, v4 subnets and site tag were built AD-HOC in the Stage-4 window, logged only to an as-executed file NOT in the repo -- there was nothing to re-run. dc1 now rebuilds from tools.
  • I INTRODUCED A DEFECT THAT PASSED MY OWN GATE 39/39. The first apply CREATED the provider-public fabric and MOVED the subnet; MAAS does not bring interface links along, leaving the region VM's enp2s0 on a subnet-less VLAN -- the same under-carve class that cost three bootstrap attempts. Live connectivity was unaffected so nothing flagged it. Fixed live + in the tool (RENAME, never move) + in the gate (stranded-link assertion).
  • A GATE WAS ALREADY RED AT HEAD: tests/node-vm T8/T9 (11/66 vs assertions reading 10/60) -- commits 086c827+447315f added the region VMs and the gauntlet was never re-run. PROVEN pre-existing by stashing this session's work. Re-pointed to the new invariant.
  • NEARLY RECORDED A FALSE FINDING: a +time=3 +tries=1 probe said the new region's BIND could not resolve externally, which would have forced keeping the cross-fiber DNS dependency. Cold-cache timeout; at +time=5 +tries=3 it resolves everything. Node DNS now points at the DC-local 10.12.8.6.
  • MEASURED, load-bearing: VLAN row id 5005 is dc0 metal-admin in Office1 and vr1-dc0-data-tenant in the new region; hot-kid is c3aqh8 there and tw7ptw in Office1. Ids do not cross regions. NIC->plane order recorded in lib-hosts.sh and is NOT PLANE_CIDRS order (walking that positionally strands PXE on provider-public).
  • dc-node-v6-carve.py FIXED, not just noted -- it was env-blind (export MAAS_PROFILE silently hit OFFICE1) and tracebacked on an absent CLI. Harness 14/14.
  • DELIBERATELY NOT WRITTEN: dc-node-carve.sh (60 NIC re-homes + 9 br-ex + 54 statics). Both defects found today were caught by LIVE runs, not fixtures, and it cannot be exercised until the nodes are re-enrolled. Its two hard inputs are now settled and recorded.
  • OWED: wire maas-profile-assert.sh into dc-plane-ipam.sh / maas-role-tags.sh / maas-node-power.sh (all still default to admin, where dc1's nine nodes live); rack DNS forwarder upstream; remove the vr1-dc0-region profile from voffice1 (SEC-026) and Office1's dc0 power-key copy.
  • Gauntlet ALL GREEN (92); repo-lint 0 fail. AS-EXECUTED LOG NOT USED -- run-logged.sh needs an interactive shell; captures went to docs/audit/* and this is declared. Body: docs/changelog-20260730-dc0-region-migration.md. Status ONLY in CURRENT-STATE.md.