# Pass-5 W2 -- METAL + TENANT/DATA plane address families under the flat Option-1 topology

**READ-ONLY planning follow-on to D-144 (container-layer elimination).** Scope: the METAL
(metal-admin/PXE) and TENANT/DATA planes (the six D-052/D-101/D-139 planes + the uplink NAT +
lb-mgmt) -- the surfaces the D-144 flatten does **NOT** redesign. D-144 re-homes these planes
onto vcloud libvirt with **zero body changes to family/CIDR/MTU**
(`FINAL-PLAN.md:29-31`: *"plane CIDRs/families/MTU (D-139/D-143 own the values -- the removal
changes NO byte budget)"*). This document enumerates what family each plane carries TODAY (built)
and under RULING (target), what forces v4 where it persists, and separates two items that are
**NOT touched or unlocked by the collapse**: the LXD API-charm container v6 gap, and the
geneve-over-v6 rebuild items.

**A note on "current family" citations, honestly stated up front:** `scripts/lib-net.sh`
(`PLANE_CIDRS`/`PLANE_NAME`, lines 24-32) carries **v4 literals only** -- there is no v6 arm in
the file today. D-139's own execution list (design-decisions.md:7442) still owes *"update
`lib-net.sh`'s v6 arm and its harness"*. So every v4 cell below cites `lib-net.sh` directly; every
v6 cell cites the D-139 ruling tables / CURRENT-STATE's measured build state instead, because
there is nothing else on disk to cite.

**A note on current-vs-ruled, also stated up front:** D-139 ruling A's four now-v6-only planes
(`metal-internal`, `data-tenant`, `storage`, `replication`) are **currently dual-carved in MAAS**
-- the v4 subnets are removed **LAST**, "after each is proven" (design-decisions.md:7443). So a
live `maas admin subnets read` today will show v4 present on all six planes; the family column
below is the **ruled target**, with a build-status note distinguishing it from the live state.

---

## 1. metal-admin -- the PXE/commissioning + DHCP plane

| Item | Value |
|---|---|
| Current family (lib-net) | `10.12.8.0/22` (`lib-net.sh:26`), gw `10.12.8.1` at vr0-dc0 only -- **VR1 has NO router on this plane, measured** (`lib-net.sh:135-146`: `maas admin subnets read` returns gateway `none`, `ping 10.12.8.1` 100% loss / incomplete ARP, both DCs) |
| Ruled family | **dual-stack** (D-139 ruling A, design-decisions.md:7327) -- v4 RETAINED + v6 GUA added (`f0X:20::/60` parent, ruling B, :7357) |
| forced-v4 / already-v6 / dual | **dual, by ruling** -- but the PXE/DHCP sub-function inside it is **v4-only by design**, not merely v4-heavy |
| What forces the v4 half | Three independent, repo-stated reasons (D-139 ruling A's own framing, :7308-7309): (1) PXE/commissioning; (2) the MAAS region API `jujud` dials for life; (3) the entire v4-only D-134 utility band (`.4` mirror / `.5` juju / `.6` MAAS region / `.7` tailscale / `.8` client -- D-134 AMENDMENT 2026-08-10, :6240); (4) "a rack with NO global v6 at all" -- confirmed by measurement, not assumption: CURRENT-STATE's 2026-07-31 sweep found `ip -6 -o addr show scope global` returns EMPTY on the rack, so the D-135 mirror, D-131 node-DNS forwarder and MAAS rack agent are v4-only by construction |
| **PXE-over-v6 capability -- VERIFY or mark UNVERIFIED (the anchor question)** | **UNVERIFIED as a platform-capability claim; RULED+BUILT as a posture.** Nothing read in this repo *measures* whether MAAS 3.7 can commission a node IPv6-only -- there is no such test, capture, or MAAS-source citation on disk. What IS on record: (a) D-101 **rules** "PXE is v4-first" (design-decisions.md:2307-2308) -- a ruling, not a measurement; (b) the D-134 node-v6-acquisition ruling (:6138-6144, operator: "MAAS static assignment is correct") leaves `dhcpd6` **off by design** -- "MAAS starts it only for a v6 dynamic range, and under this ruling none should exist" (:6151); (c) commissioning today allocates from the v4 `dynamic .201-.254` range, and every v6 `/64` carries **zero ip ranges** (CURRENT-STATE measured sweep, "ABSENT on v6" item i); (d) the rack itself has no global v6 (item ii above). **Net: this repo has built and ruled a v4-first PXE path and never built or tested a v6-only alternative to compare against -- do not import outside knowledge of MAAS's IPv6 PXE support; state the capability question as open.** |
| Tooling status | v4 carve/verify: `scripts/dc-node-carve.sh`, `scripts/carve-host-interfaces.sh` (BUILT, live). v6 carve/verify: `scripts/dc-node-v6-carve.py`, `scripts/dc-node-v6-verify.sh` (gate **G19**, header cites D-139/D-101/D-134) -- **BUILT, on disk**, but the v6 execution list items ("MAAS carve per DC", "re-carve 54 node v6 statics") are marked NOT DONE at D-139 (:7439-7443). A Stage-5-first-boot G17 v6-static assertion is PROPOSED, not adopted (:7157-7163 style, D-101 node-IPv6 note). |
| Governing D | D-101 (family + PXE v4-first rule, :2307-2308), D-139 ruling A (dual-stack retained, :7327), D-134 (utility octet band, v4-only by construction) |

---

## 2. The six D-052/D-101/D-139 planes -- ruled target vs. built state

| Plane | Current CIDR (lib-net, v4) | Ruled family (D-139 ruling A) | forced-v4/already-v6/dual | What forces it | Governing D |
|---|---|---|---|---|---|
| `provider-public` | `10.12.4.0/22` (`lib-net.sh:25`), gw `10.12.4.1` (the ONE routed VR1 plane, measured) | **dual-stack** (:7326); GUA `f0X:10::/60`, :7355 | dual | Tenant FIPs against "a still-substantially-v4 internet", external API VIPs, Octavia tenant LB VIPs, the edge default route (:7310-7311) | D-101, D-139 ruling A |
| `metal-admin` | `10.12.8.0/22` (`lib-net.sh:26`) | dual-stack (:7327) | dual (PXE sub-forces v4; see Sec 1) | see Sec 1 | D-101, D-139 ruling A |
| `metal-internal` | `10.12.12.0/22` (`lib-net.sh:27`), gw `none` (VR1) | **IPv6-only** (:7328) -- SUPERSEDES D-101's earlier "datastore east-west v4-bound" clause (:7335-7336) | already-v6 (ruled); **v4 still present live** -- removed LAST after proof (:7443) | Nothing forces v4 here per D-139's own audit: D-101's v4 lean was "a JUDGEMENT... not a protocol requirement" (:7312-7313), reversed by D-139 | D-101 (superseded clause), D-139 ruling A |
| `data-tenant` | `10.12.16.0/22` (`lib-net.sh:28`) | **IPv6-only** (:7329) -- geneve underlay | already-v6 (ruled); v4 still present live | East-west only, no external clients, on-link, no gateway needed (:7314-7315) | D-101, D-139 ruling A |
| `storage` | `10.12.32.0/22` (`lib-net.sh:29`) | **IPv6-only** (:7330) -- Ceph public | already-v6 (ruled); v4 still present live | Same as above; `ceph-mon`/`ceph-osd` are in `PREFER_IPV6_CHARMS` (:7432) | D-101, D-139 ruling A |
| `replication` | `10.12.36.0/22` (`lib-net.sh:30`) | **IPv6-only** (:7331) -- Ceph cluster incl. cross-DC leg | already-v6 (ruled); v4 still present live; **cross-DC leg has no v6 route at all** (build-constraints bullet, :7431, OPEN) | Same on-link reasoning; the cross-DC hop is the one exception D-139 itself leaves open | D-101, D-139 ruling A, D-108 (replication) |
| `lb-mgmt` | n/a (never had a v4 carve) | **IPv6-only, NEW plane** (:7332) | **two distinct v6 objects -- do not conflate (G18 warning)**: (1) octavia's **charm-created ULA `fc00::/64`** (R8-ruled, live once Octavia deploys); (2) the **D-139 apex GUA `f0X:80::/64`** (dc0 `2602:f3e2:f02:80::/64`) carved as a MAAS-underlay plane with **NO charm consumer** -- octavia binds no `lb-mgmt` space -- kept **`reserved`** per the G18 ruling (option b, CURRENT-STATE:7825 cell, RULED 2026-08-08) | Object (1): charm default, no CIDR/family option exposed by the charm. Object (2): an architectural allocation with intent recorded, deliberately unconsumed | D-101 (original ULA rule), D-139 (apex GUA carve + annotation, :7287-7296), G18 ruling (CURRENT-STATE:7825) |

**Tooling status (all six planes):** carve/verify built (`scripts/dc-node-v6-carve.py`,
`scripts/dc-node-v6-verify.sh`, gate G19); IPAM incl. ULA-retirement bookkeeping in
`scripts/dc-plane-ipam.sh` (D-139 ruling B retires the VR1 ULA `/48` in favor of full GUA, so this
script's ULA-retirement arm is directly relevant, `dc-plane-ipam.sh:368-372` also holds the MAAS
"zero ip-ranges != zero availability" fact that overturned the original mis-diagnosed root cause).
D-139's own execution list (apex push, MAAS carve, re-carve 54 node statics, reissue Octavia v6
SANs, `lib-net.sh` v6 arm, `lb-mgmt` VLAN/space/subnet carve, v4 subnet removal LAST) is **entered
nowhere as done** -- design-decisions.md:7437 states plainly "none of it is done."

---

## 3. The uplink/ISP NAT (`172.30.x`) -- persists through the collapse, stays v4

| Item | Value |
|---|---|
| Module evidence | `opentofu/modules/site-wan/main.tf` -- `resource "libvirt_network" "site_wan"` takes **one** `var.cidr`, one `forward.mode="nat"`, and a **single-entry** `ips = [{ address = cidrhost(var.cidr, 1), prefix = ... }]` block. No v6 CIDR variable, no second `ips` entry, no v6 anywhere in the module -- single-stack v4 by construction, not by omission of a flag. |
| Live values | `172.30.2.0/24` (vr1-dc0, `opentofu/main.tf:382`), `172.30.3.0/24` (vr1-dc1, `:394`) -- D-115-assigned, RULED literals (D-125 addressing note, design-decisions.md:5151-5159) |
| forced-v4 / already-v6 / dual | **forced-v4**, single-stack |
| What forces it | It simulates a v4 ISP hand-off: OPNsense WAN keeps its static v4 `.2` (D-113 edge config, unchanged by D-125), and D-101's own governing rationale explicitly defers external v6: *"External IPv6 routing is deliberately NOT added this deployment; it arrives with the next full deployment"* (design-decisions.md:2300-2301) -- **NAT64/DNS64 was considered as a v6-egress simulant and REJECTED** ("a shim with NO Roosevelt analog... testing a translation path that will never run in production", :2295-2299). Roosevelt itself will have "full IPv4 and IPv6 edge transport" (operator, :2274-2276) -- VR1's v4-only edge is a deliberate rehearsal simplification, not a capability gap. |
| Effect of D-144 | **Persists unchanged in family.** D-144's own text: *"edge WAN attaches directly to the per-DC `site-wan` NAT"* (:8256-8257) -- the module and its cidr are UNCHANGED; only the WAN bridge indirection (`modules/wan-bridge`, D-125's bridge-in plumbing that existed solely to give the now-eliminated containment VM egress) is removed. D-125 itself is marked **TERMINATED by D-144** (design-decisions.md:5091-5094) but the uplink NAT it wired through is explicitly retained. |
| Governing D | D-101 (external-v6-deferred rationale), D-113 (OPNsense static WAN), D-115 (the `/24` literals), D-125 (TERMINATED by D-144, history only), D-144 (persistence + re-wire) |

---

## 4. lb-mgmt (already v6) -- summary

See Section 2 row above for the full two-object treatment. Short form: **already v6**, both
objects, but the live/consumed one is the charm-created ULA `fc00::/64` (R8); the D-139 apex GUA
`/64` is a reserved architectural placeholder with no consumer and must stay that way per the G18
ruling -- pointing octavia at it would reopen R8.

---

## 5. NOT unlocked by the collapse -- two SEPARATE layers, explicitly out of D-144's scope

**Primary source for this separation, cited verbatim:** `docs/audit/container-elim-pass/pass0-admin-report.md:57-60` --

> "The `vvr1-dcN` VMs + their inner libvirtd + the inner roots + the bootstrap gate's node-host
> mode + the D-125 bridge-in plumbing... and the qemu+ssh provider dial and its D-126 keys. **It is
> NOT: the six planes, the node VMs, the DC edge, the mesh triangle, the uplink NATs, or the LXD
> API-charm containers on deployed nodes** (all of which persist; the first three re-home)."

### 5a. The LXD API-charm container v6 gap (CURRENT-STATE ~1456-1730)

This is a **juju-side container-provisioning defect on the deployed OpenStack NODE VMs** --
completely different hardware/software layer from the `vvr1-dcN` containment VM D-144 eliminates.
Summary of the mechanism, as measured and recorded (CURRENT-STATE.md, the 2026-07-31 D3 block):

- The HOST (a node VM) is fully dual-stacked: every plane bridge on it carries a global v6
  address (`br-ex 2602:f3e2:f02:10::100/64`, etc.).
- The LXD **containers** juju provisions on that host for the API charms (keystone, cinder,
  glance, neutron-api, nova-cloud-controller, openstack-dashboard, ceph-radosgw, ...) come up
  **v4-only**, because juju's `EthernetDeviceForBridge` (`state/linklayerdevices.go`, pinned
  `v3.6.27`) takes `addrs[0]` from an **unsorted** mongo query and derives exactly ONE subnet --
  "zero family-awareness anywhere on that path" (D-139, design-decisions.md:7401-7406). This is
  **LP #1723240** (`Triaged`/`Low`, open since 2017), not a VR1-specific defect.
- D-139's own ruling A response to this measurement was to make the four internal planes
  **IPv6-only** (dual-stack is "NOT EXPRESSIBLE" for juju containers, so make v6 the only option
  where possible, :7410-7412). `metal-admin` **stays dual-stack and carries containers**, so its
  containers land on ONE family by mongo order (today v4) -- a **CARRIED RISK**, with a detection
  gate ("every container's `metal-admin` leg is IPv4") explicitly **owed, not built** (:7413-7417).
- Separately, `prefer-ipv6` is currently **set false on all seven charms that declare it**
  (D-101 ruling note (b), 2026-07-31) until IPv6 is proven operational end-to-end; every dual-family
  VIP leg (13 apps x v6) is kept regardless.

**D-144's relationship to this gap: NONE.** The client VM / flat-topology change touches
substrate provisioning (tofu roots, the qemu+ssh dial, the WAN bridge); it does not touch juju's
container-network-allocation code path, MAAS's v6 ip-range population, or the `prefer-ipv6` charm
config. Nothing in D-144 changes any input to LP #1723240's failure mode.

### 5b. The geneve-over-v6 rebuild items (D-139, 2026-08-09 amendment)

Also a **different layer** -- OVN/OVS chassis config on the deployed compute/network nodes, not
the containment VM. Confirmed viable 2026-08-09 (real VM-to-VM cross-compute ping over v6 geneve,
8/8, 0% loss on OVS 3.3.0 / OVN 24.03.2 / kernel 5.15.0-186; design-decisions.md:7960-8001). Two
conditions, both still owed at rebuild time and neither touched by D-144:

1. Containerized OVN chassis (the octavia / ovn-chassis-octavia LXD units -- itself an instance of
   the 5a auto-pick problem) must take a **carved**, not auto-picked, v6 data-tenant address.
2. `ovn-encap-ip` must reach OVS **unbracketed** -- ovn-chassis 24.03 sets it bracketed
   (`"[2602:...]"`), which OVS's geneve rejects (`ofport -1`); fix is a fixed charm revision or a
   persistent `ovs-vsctl` override.

Plus `neutron overlay_ip_version=6` for the correct tenant MTU (~1422 vs 1500, v6 geneve being 20
bytes heavier than v4 geneve). **Gate:** `scripts/geneve-encap-assert.sh` (built 2026-08-09,
harness 16/16) checks both encap-family AND per-tunnel `ofport>=0`, wired into
`runbooks/dc-dc-phase4-juju-bundle-per-dc.md` Step 12.2. **D-144's relationship: NONE** -- these
are OVN chassis / neutron config items that apply identically whether the node VMs sit under a
containment VM or flat on vcloud libvirt; the fix targets the chassis-carve and the charm
revision, not the virtualization depth.

**Why this separation matters for the redeploy:** D-144's `FINAL-PLAN.md` two-axis discipline
(`[D-143]` value-substitution vs `[CE]` shape-change) explicitly excludes both 5a and 5b from
either axis -- `lib-net.sh carries ZERO container-elim edits` (:8323) and neither gap is named in
D-144's reconciliation ledger (:8270-8294). Crediting the flatten with fixing (or being blocked
by) either gap would be a manufactured link between two independently-tracked defects.

---

## 6. Risk notes carried forward (not new findings -- surfaced because they sit at the W2 boundary)

- **Replication's cross-DC leg** is ruled IPv6-only (Sec 2) but has **no v6 route today**
  (D-139 build-constraints bullet, OPEN) **and** D-144's DEC-24 leaves the post-flatten mesh-leg /
  netem attachment for that same carrier **undefined** (`FINAL-PLAN.md` / D-144 :8312-8317,
  "the pass reassigned the Office1-transit legs to the client VM cleanly but left this leg's
  post-flatten attachment + netem routing undefined"). This is the one point where the W2 family
  ruling and the D-144 collapse genuinely intersect -- an unresolved carrier for an unresolved
  route.
- **PREFER_IPV6_CHARMS reopen-per-app:** `ceph-mon`, `ceph-osd`, `mysql-innodb-cluster` are all in
  this set; converting `storage`/`replication`/`metal-internal` to v6-only (Sec 2) REOPENS the
  2026-07-31 "set false on all seven" ruling note per-app as each plane converts (D-139
  build-constraints bullet, :7432-7435) -- not yet re-ruled.
- **metal-admin container-family-flip risk** (Sec 5a) has a named, owed, unbuilt detection gate.

---

## Sources read in full for this pass

`scripts/lib-net.sh` (all 264 lines); `docs/design-decisions.md` D-101 (:2243-2510ish, incl. the
2026-07-25/07-27 ruling notes), D-125 (:5089-5170, incl. the D-144 termination annotation), D-134
(all amendments through the 2026-08-10 `.8` client-VM entry), D-139 (:7280-8001, both rulings + the
2026-08-01 reconsideration note + the 2026-08-09 geneve amendment), D-141 (header only, cross-ref),
D-144 (:8210-8329, full entry); `docs/CURRENT-STATE.md` :1440-1740 (the LXD container v6 gap, D1-D3
+ the systematic v6 sweep + both rulings) and the G18 gate cell (:7825); `docs/tool-index.md` (full,
for artifact lookups); `docs/audit/container-elim-pass/FINAL-PLAN.md` (Sections 1-2);
`docs/audit/container-elim-pass/pass0-admin-report.md` (:1-90, the container-layer scope
definition); `opentofu/modules/site-wan/main.tf` (full); `opentofu/main.tf` (uplink module
instantiations, :369-394).
