diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index a72aec5..b558dfd 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -511,6 +511,47 @@ RE-POINTED to the new invariant (11 nodes x 6 planes) with the reason recorded in-file and re-proven able to fail. dc1 reads 11 lists / 60 literals -- EXPECTED, since `vr1-dc1-maas-01` is authored with `macs = []` and not yet applied. + **MIGRATION PREREQ 5 DONE 2026-07-30 -- THE dc0 REGION'S PLANE TOPOLOGY IS BUILT AND + VERIFIED. Named check `dc-region-topology.sh check vr1-dc0` = 39 assertions / 0 failed, + EXIT 0** (capture `docs/audit/dc0-region-topology-20260730.txt`). `scripts/dc-region-topology.sh` + is NEW (harness 39/39, 4 mutations each killed tests) and closes the largest of the + no-tool gaps. The apply created 5 named plane fabrics, 6 spaces and 4 plane subnets, + MOVED `10.12.4.0/22` off the auto `fabric-1` onto `vr1-dc0-provider-public`, bound all 6 + VLANs to their spaces, and created `openstack-vr1-dc0`. `--profile` has NO DEFAULT and + every mutating path runs `maas-profile-assert.sh` first. **metal-admin DELIBERATELY keeps + MAAS's own auto-created fabric** -- it is a discovery artifact of rack registration, and + Juju binds SPACES not fabrics; the check asserts only that it does not share a fabric with + a named plane. **CONCRETE PROOF THAT ROW IDS ARE NOT PORTABLE, measured in passing: VLAN + row id 5005 is dc0's metal-admin VLAN in the OFFICE1 region and `vr1-dc0-data-tenant` in + the NEW one** -- same integer, different DC plane. + **metal-admin service config SET on the new region:** `dns_servers=10.12.8.6`, + `allow_dns=false`, D-134 dynamic range `.201-.254` (subnet resolved BY CIDR). **The DNS + target was MEASURED and my first reading of it was WRONG:** a probe at + `+time=3 +tries=1` suggested the new region's BIND could not resolve external names; that + was a COLD-CACHE TIMEOUT, not a capability gap. Re-probed at `+time=5 +tries=3` it answers + `archive.ubuntu.com` / `streams.canonical.com` / `api.snapcraft.io`, with `flags: qr rd ra` + and `ANSWER: 9`. Pointing node DNS at the DC-LOCAL region is correct: it is authoritative + for the zone the migrated nodes will live in (own `maas-internal`, SOA serial 16) and drops + the cross-fiber dependency D-132 q1 exists to remove. **FOLLOW-UP LOGGED, NOT ACTIONED:** + the live rack forwarder `/etc/dnsmasq-dc0-node.conf` is `no-resolv` + `server=10.10.0.20`, + so after migration its upstream is a region that no longer owns dc0's nodes + (`dc-rack-net.sh` carries `DNS_UPSTREAM="10.10.0.20"` for BOTH sites at `:66`/`:82`; dc1's + is still correct, so this is a per-site cutover change). The live unit also carries the + D-119-retired BARE token (`dc0-node-dns.service`) while the script now generates `vr1-dc0-*`. + **>>> STAGE 5 IS BLOCKED ON A PERMISSION WALL, NOT ON A DECISION OR A DEFECT. <<<** + The next step is the one-way DHCP cutover -- `vlan update 4 0 dhcp_on=false` at Office1, + verify `dhcpd` STOPPED **by process** (MAAS self-report is not evidence -- the 2026-07-20 + Temporal incident class), then `vlan update 0 0 dhcp_on=true primary_rack=` on the + new region and verify RUNNING by process. **The harness classifier REFUSED it. NOT retried + in an altered shape**, per the precedent set by the four MAAS carve calls refused earlier + the same day. **NOTHING IS HALF-APPLIED, verified read-only after the refusal:** Office1 + still reads `10.12.8.0/22 dhcp_on=True primary_rack=7chphy` and the rack still runs exactly + ONE `dhcpd` on `virbr2`. Measured boundary state, so the next session does not re-derive it: + dc0 metal-admin is **fabric-4 / vid 0 / vlan row 5005** in Office1 and **fabric-0 / vid 0 / + vlan row 5001** in the new region; Office1 racks are `mtstwf`/`7chphy`/`nmpcq4`, the new + region's sole rack is `hot-kid`. All 10 dc0 machines are `Ready` and POWERED OFF, so the + DHCP gap is free to take -- two servers on one segment is the hazard the off-then-on + ordering avoids, which is why it is ONE operation. - Project: Omega Cloud, VR1 DC-DC rehearsal -- a two-DC + Office1-headend virtual rehearsal on KVM (vcloud host), rehearsing the future bare-metal diff --git a/docs/audit/dc0-region-topology-20260730.txt b/docs/audit/dc0-region-topology-20260730.txt new file mode 100644 index 0000000..a2b2f7e --- /dev/null +++ b/docs/audit/dc0-region-topology-20260730.txt @@ -0,0 +1,53 @@ +dc0 PER-DC REGION -- PLANE TOPOLOGY REBUILT (D-132 q1 migration) +Captured 2026-07-30. Target region: http://10.12.8.6:5240/ (profile vr1-dc0-region). +Run FROM voffice1 through the rack tunnel -L 127.0.0.1:5241 -> 10.12.8.6:5240. +================================================================ + +-- region identity proven FIRST (a machine count is not proof) -- +OK profile 'vr1-dc0-region' -> region with racks [hot-kid] (as expected) + +-- named executable check AFTER the apply -- +$ dc-region-topology.sh check vr1-dc0 --profile vr1-dc0-region --expect-rack hot-kid +OK profile 'vr1-dc0-region' -> region with racks [hot-kid] (as expected) +dc-region-topology check: site=vr1-dc0 profile=vr1-dc0-region + [ok] fabric 'vr1-dc0-provider-public' exists + [ok] fabric 'vr1-dc0-metal-internal' exists + [ok] fabric 'vr1-dc0-data-tenant' exists + [ok] fabric 'vr1-dc0-storage' exists + [ok] fabric 'vr1-dc0-replication' exists + [ok] space 'provider-public' exists + [ok] space 'metal-admin' exists + [ok] space 'metal-internal' exists + [ok] space 'data-tenant' exists + [ok] space 'storage' exists + [ok] space 'replication' exists + [ok] subnet 10.12.4.0/22 (provider-public) exists + [ok] subnet 10.12.4.0/22 is on fabric 'vr1-dc0-provider-public' (got 'vr1-dc0-provider-public') + [ok] subnet 10.12.4.0/22 VLAN is bound to space 'provider-public' (got 'provider-public') + [ok] subnet 10.12.4.0/22 gateway '10.12.4.1' (got '10.12.4.1') + [ok] subnet 10.12.8.0/22 (metal-admin) exists + [ok] metal-admin 10.12.8.0/22 is on MAAS's own fabric 'fabric-0' (by design) + [ok] subnet 10.12.8.0/22 VLAN is bound to space 'metal-admin' (got 'metal-admin') + [ok] subnet 10.12.8.0/22 gateway 'none' (got 'none') + [ok] subnet 10.12.12.0/22 (metal-internal) exists + [ok] subnet 10.12.12.0/22 is on fabric 'vr1-dc0-metal-internal' (got 'vr1-dc0-metal-internal') + [ok] subnet 10.12.12.0/22 VLAN is bound to space 'metal-internal' (got 'metal-internal') + [ok] subnet 10.12.12.0/22 gateway 'none' (got 'none') + [ok] subnet 10.12.16.0/22 (data-tenant) exists + [ok] subnet 10.12.16.0/22 is on fabric 'vr1-dc0-data-tenant' (got 'vr1-dc0-data-tenant') + [ok] subnet 10.12.16.0/22 VLAN is bound to space 'data-tenant' (got 'data-tenant') + [ok] subnet 10.12.16.0/22 gateway 'none' (got 'none') + [ok] subnet 10.12.32.0/22 (storage) exists + [ok] subnet 10.12.32.0/22 is on fabric 'vr1-dc0-storage' (got 'vr1-dc0-storage') + [ok] subnet 10.12.32.0/22 VLAN is bound to space 'storage' (got 'storage') + [ok] subnet 10.12.32.0/22 gateway 'none' (got 'none') + [ok] subnet 10.12.36.0/22 (replication) exists + [ok] subnet 10.12.36.0/22 is on fabric 'vr1-dc0-replication' (got 'vr1-dc0-replication') + [ok] subnet 10.12.36.0/22 VLAN is bound to space 'replication' (got 'replication') + [ok] subnet 10.12.36.0/22 gateway 'none' (got 'none') + [ok] metal-admin carries node-facing dns_servers (got '10.12.8.6') -- D-131 + [ok] metal-admin allow_dns is False (got 'False') -- D-131 + [ok] metal-admin has a DHCP dynamic range (got 1) + [ok] tag 'openstack-vr1-dc0' exists +dc-region-topology: 39 passed, 0 failed + EXIT = 0 diff --git a/docs/changelog-20260730-dc0-region-migration.md b/docs/changelog-20260730-dc0-region-migration.md index f0164da..2e7273d 100644 --- a/docs/changelog-20260730-dc0-region-migration.md +++ b/docs/changelog-20260730-dc0-region-migration.md @@ -296,7 +296,113 @@ --- -## Item 8 -- as-executed log NOT used for this window, and why +## Item 8 -- `scripts/dc-region-topology.sh` + harness (NEW), and APPLIED at dc0 + +**What.** `check|apply --profile

--expect-rack [--commit]` over the five +named plane fabrics, the six spaces, the six v4 plane subnets on the right VLANs with the +ruled gateways, the VLAN->space bindings, and the `openstack-` tag. This is the tool +item 6(c) found missing. + +**Design decisions worth stating:** +- **`--profile` has NO default.** The repo-wide default `admin` resolves to the Office1 + region, where this `apply` would be an idempotent no-op that prints PASS. Every mutating + path is gated by `maas-profile-assert.sh` FIRST. +- **metal-admin deliberately keeps MAAS's own auto-created fabric.** It is a discovery + artifact created when the rack registers, before anything in this repo could name it; + Juju binds SPACES, not fabrics. The check asserts only that it does not share a fabric + with a named plane (which would collapse two L2 domains). +- **No numeric MAAS id is persisted.** Everything resolves per-run by name or CIDR. + +**LIVE at dc0** (capture `docs/audit/dc0-region-topology-20260730.txt`). The apply created +5 fabrics, 6 spaces, 4 subnets, MOVED `10.12.4.0/22` off the auto `fabric-1` onto +`vr1-dc0-provider-public`, bound 6 VLANs to their spaces, and created the site tag. Final +named check: **39 passed, 0 failed, exit 0.** + +**A CONCRETE ILLUSTRATION OF WHY ROW IDS ARE NOT PORTABLE, measured in passing:** VLAN row +id **5005** is dc0's metal-admin VLAN in the Office1 region and `vr1-dc0-data-tenant` in +the new region. Same integer, different object, different DC plane. Anything that carried +`5005` across would have mis-targeted silently. + +**Harness 39/39; 4 mutations each killed tests** (region gate removed; profile defaulted to +`admin`; space-binding assertion neutered; fabric-placement assertion neutered). +**Fixtures were corrected against a live read:** MAAS reports an unbound VLAN as the STRING +`"undefined"`, not `null`, so the first fixtures graded a payload MAAS never emits. + +**Revert.** The apply is additive. To undo: delete the five `vr1-dc0-*` fabrics and the six +spaces on the new region, and move `10.12.4.0/22` back to `fabric-1`. Nothing in the +Office1 region was touched by this item. + +--- + +## Item 9 -- metal-admin service config set on the new region (LIVE) + +`dns_servers=10.12.8.6`, `allow_dns=false`, and the D-134 dynamic range +`10.12.8.201-10.12.8.254`. Subnet resolved BY CIDR (-> id 1), never from the capture. + +**The DNS target was MEASURED, not carried over, and my first reading of it was WRONG.** +The old region set `dns_servers=10.12.8.3` (the D-131 rack forwarder). A first probe +suggested the new region's own BIND at `10.12.8.6` could not resolve external names, which +would have forced keeping the forwarder. That was a **cold-cache timeout at +`+time=3 +tries=1`, not a capability gap** -- re-probed with `+time=5 +tries=3`, `10.12.8.6` +answers `archive.ubuntu.com`, `streams.canonical.com` and `api.snapcraft.io` correctly, and +a full `dig` shows `flags: qr rd ra` with `ANSWER: 9`. I nearly recorded a false finding on +one under-timed query. Pointing node DNS at the DC-LOCAL region is the correct end state: +it is authoritative for the zone the migrated nodes will live in (its own `maas-internal`, +SOA serial 16) and removes the cross-fiber dependency that D-132 q1 exists to remove. + +**FOLLOW-UP RECORDED, NOT ACTIONED (hard rule 1):** the live rack forwarder is +`/etc/dnsmasq-dc0-node.conf` with `no-resolv` + `server=10.10.0.20` -- it sends EVERYTHING +to the Office1 region. After dc0's nodes migrate, that upstream points at a region that no +longer owns them. `dc-rack-net.sh` carries `DNS_UPSTREAM="10.10.0.20"` for BOTH sites +(`:66`, `:82`), and dc1's is still correct, so this is a per-site change at each DC's +cutover. **Also noted: the live unit is named with the BARE `dc0` token** +(`dc0-node-dns.service`, `/etc/dnsmasq-dc0-node.conf`) while the current script generates +`vr1-dc0-*` names -- a D-119 retired-token residue from the 2026-07-21 install. + +**Revert.** `maas vr1-dc0-region subnet update 1 dns_servers=10.12.8.3 allow_dns=false` and +delete the dynamic iprange. + +--- + +## Item 10 -- BLOCKED: the DHCP handover (task 5) needs an operator decision + +**What was attempted.** With the topology verified complete (39/0), the next step is the +one-way DHCP cutover: + + # 1. off at Office1 (measured target, resolved live, not from memory) + maas admin vlan update 4 0 dhcp_on=false + # 2. verify dhcpd actually STOPPED on the rack, BY PROCESS + # (MAAS self-report is not evidence -- the 2026-07-20 Temporal incident class) + # 3. on at the new region, primary_rack = hot-kid + maas vr1-dc0-region vlan update 0 0 dhcp_on=true primary_rack= + # 4. verify dhcpd RUNNING by process + +**It was REFUSED by the harness permission classifier.** Per the standing rule -- and the +precedent recorded in `CURRENT-STATE` for the four MAAS carve calls refused on 2026-07-30 -- +it was **NOT retried in an altered shape.** + +**NOTHING IS HALF-APPLIED, verified read-only after the refusal:** Office1 still reads +`10.12.8.0/22 dhcp_on=True primary_rack=7chphy`, and the rack still runs exactly one +`dhcpd` on `virbr2`. The refusal landed before any mutation. + +**Measured state at the boundary, so the next session does not re-derive it:** + +| | Office1 region | new dc0 region | +|---|---|---| +| dc0 metal-admin | fabric-4 / vid 0 / vlan row 5005 | fabric-0 / vid 0 / vlan row 5001 | +| dhcp_on | **True**, primary_rack `7chphy` (vvr1-dc0) | False, primary_rack None | +| dns_servers | `10.12.8.3` | `10.12.8.6` | +| rack system_ids | mtstwf / 7chphy / nmpcq4 | hot-kid | + +**Why the gap is safe to take:** all 10 dc0 machines are `Ready` and POWERED OFF, so no +node is mid-DHCP. Two DHCP servers on one segment is the hazard the ordering avoids, which +is why off-then-on is one operation and not two independent ones. + +**No revert needed** -- nothing was changed. + +--- + +## Item 11 -- as-executed log NOT used for this window, and why `scripts/run-logged.sh` opens an INTERACTIVE `script(1)` subshell and is unusable from a non-interactive agent session. F6 (2026-07-30) already records the harness classifier