# QUEUED RULINGS -- Stage-5 grounding audit (2026-07-27)

**Nothing here is adopted.** Every item is PRESENTED, never picked (hard rule 1;
GA-R5: PROPOSED means present options). Each question is a SEPARATE exchange --
GA-R5 rule 1 makes a batch adoption INVALID, so answering "yes to all" rules
NOTHING and the session that reads this must stop and re-ask.

**How to use this file.** Answer questions ONE AT A TIME. Write your exact words
on the `OPERATOR UTTERANCE:` line. A ruling exists only once its Status block
quotes the question AND your exact utterance, dated, committed and pushed --
before any dependent work starts. An ambiguous or template answer rules nothing.

Evidence for every claim below is in `docs/audit/stage5-committee-raw-20260727.md`
(verbatim lens returns) and `docs/audit/stage5-live-measurement-20260727.txt`
(this session's own measurements). Finding IDs are cited so nothing here has to
be taken on trust.

Ordering note: **R1, R2 and R3 change what gets deployed.** They should be
answered before the rest, because several later questions have different right
answers depending on them.

---

# PART A -- STAGE-5 BLOCKING

These must be answered before `juju deploy`. Each one, left unanswered, either
stops the deploy or bakes in a state that is expensive to reverse on a live cloud.

---

## R1. The Ceph OSD device does not exist on any node

**Finding:** L2-1 (measured twice -- `virsh domblklist` on both racks AND MAAS
`physicalblockdevice_set` on all 18 nodes). Verified independently by this session.

`bundle.yaml:560` sets `osd-devices: /dev/vdb`. Every one of the 18 VR1 node VMs
has exactly ONE block device, `vda`. `opentofu/modules/node-vm/main.tf:104-125`
declares a single disk and neither substrate root has any OSD-disk variable. The
`/dev/vdb` line is a VR0 as-built comment ("libvirt-attached, MAAS-untracked") --
VR0 hosts had an attached second disk; VR1 nodes never did.

Consequence if unanswered: `ceph-osd` deploys onto four storage nodes and finds
no device. Ceph never forms, and everything storage-backed behind it stalls.

**Options:**
- **(a)** Add an OSD volume to `modules/node-vm` and re-apply BOTH substrate
  roots. Closest to Roosevelt (real machines have real disks), but it re-opens
  the substrate on 18 running-but-powered-off nodes and both roots currently plan
  ZERO DIFF -- that property is deliberately being spent.
- **(b)** Re-point `osd-devices` at a directory or partition path on the existing
  `vda`. Cheapest, no substrate change, but it rehearses a Ceph layout Roosevelt
  will not use, which cuts against MINIMIZE DELTA TO ROOSEVELT.
- **(c)** Re-shape the storage role (e.g. fewer, larger storage nodes with a
  dedicated disk each).

OPERATOR UTTERANCE: **"Add an OSD volume to node-vm (Recommended)"** -- RULED
2026-07-27. **R1 IS CLOSED.** Recorded as a D-121 AMENDMENT (the governing decision
already ratifies `modules/node-vm` sizing, and its own capacity re-validation already
budgeted "Ceph disk re-run for 4 storage/DC = PASS 5.31 TiB" -- the disk was budgeted
and never built). Two scope corrections landed with the ruling: it is **8 volumes, not
18** (ceph-osd is placed on the four storage nodes per DC only), and the apply is a
SEPARATE operator-gated step with four preconditions, including verifying that
re-commissioning preserves the D-134 statics and pinned MACs BEFORE the apply -- the
2026-07-20 MAC-regeneration incident is the precedent. Full text: `docs/design-decisions.md`,
"D-121 AMENDMENT 2026-07-27".

---

## R2. There is no IPv6 anywhere in the DC substrate, but dual-stack is RULED

**Finding:** L2-3 (measured). D-101's RULING NOTE of 2026-07-25 records your exact
words -- "Dual stack to be used where IPv4 is required, IPv6 where IPv6 only makes
sense" and "Dual stack deployment for DC0 and DC1. This is the deployment when the
dual-stack is added" -- with the stated effect that "the v4-only phasing option ...
is CLOSED, for BOTH DC0 and DC1 in this deployment."

Measured reality: exactly ONE IPv6 subnet exists cloud-wide
(`2602:f3e2:f01:100::/64` on the Office1 base fabric). None of the 12 DC plane
fabrics carries an IPv6 subnet. Zero IPv6 links across all 18 nodes; every node
reports `default_gateways.ipv6 = NONE`. D-101's own family matrix requires ULA on
data-tenant, storage and replication and a ULA leg on metal-admin/metal-internal --
none of which has any MAAS v6 presence to bind against. D-101's own "Remaining open
item" is the un-assigned NetBox literals: the org ULA /48 and the per-DC GUA carve.

This is the single largest fork in the audit. It is BLOCKING because addresses
become as-built Keystone endpoints and Vault-issued cert SANs at deploy time.

**Options:**
- **(a)** Assign the ULA /48 and per-DC GUA carve in NetBox, carve them into MAAS,
  and deploy dual-stack as ruled. Honours D-101; adds real work before Stage 5.
- **(b)** Deploy v4-only now and add v6 legs post-deploy. **This CONTRADICTS a
  ruling you already made** -- it would need an explicit amendment, not a silent
  choice, and re-addressing endpoints and re-issuing SANs afterwards is the
  expensive path.
- **(c)** Deploy dc1 v4-only as a deliberate, recorded rehearsal exception while
  dc0 goes dual-stack, making the v4/v6 delta itself the experiment (the D-135
  per-DC-difference pattern applied to address family).

OPERATOR UTTERANCE: **"Carve v6 and deploy dual-stack as ruled (Recommended)"** --
RULED 2026-07-27. **R2 IS CLOSED.** Recorded as a D-101 RULING NOTE (re-confirmation,
no amendment -- `docs/design-decisions.md` is the authority). **Consequence that
changes the plan: D-101's "Remaining open item" (the org ULA /48 and per-DC GUA
carve), until now carried as non-blocking "pending NetBox assignment", is now a
STAGE-5 PRECONDITION** -- dual-stack cannot deploy against literals that do not
exist. R9 and R11 both inherit "dual-family" from this. The L3-9 overlay collision
must be reconciled BEFORE either authority location is populated; note the dangerous
direction is the one that PASSES (vips overlay last silently drops every v6 leg and
reports green). R8 is NOT resolved by this ruling. **R2a WITHDRAWN -- see below.**

---

## R2a. WHICH IPv6 literals? -- WITHDRAWN 2026-07-27, NEVER OPEN

**This question was asked in error and is withdrawn before any ruling.** The operator
asked whether the NetBox apex had been polled. It had not. It has now been read, and
the literals are ASSIGNED, RATIFIED AND RECORDED -- and have been since 2026-07-11
under **D-111 (ADOPTED)**.

MEASURED from `netbox/draft/vr1-office1-current-20260725.json` (139 prefixes, 103
IPv6), every row tagged `D-101/D-111`: ULA `fd50:840e:74e2::/48` with DC0 planes at
`:220`/`:221`/`:230`/`:240`/`:250::/64` and DC1 at `:320`/`:321`/`:330`/`:340`/`:350::/64`;
GUA provider-public DC0 `2602:f3e2:f02:10::/64` + VIP `f02:11::/64`, DC1
`2602:f3e2:f03:10::/64` + VIP `f03:11::/64`.

**What is actually owed is PROPAGATION, and it needs no ruling.** The ratified values
are absent from the two places Stage 5 reads: `scripts/lib-net.sh` carries no v6 arm at
all, and MAAS carries no v6 on any of the 12 DC plane fabrics. Both are mechanical
copies from an authoritative source. This moves from Part A (needs a decision) to the
Phase-3 mechanical batch in the readiness doc.

**Why the error happened, recorded because it is the audit's own failure mode:** the
audit's lens 2 explicitly listed the apex as UNMEASURED and warned "D-101's literals
may exist in NetBox and simply not be carved into MAAS -- L2-3's claim is scoped to
MAAS + node reality and does not assert the apex is empty." That warning was correct
and available, and a question was put to the operator anyway, on the strength of
D-101's own stale "Remaining open item" prose. Trusting stale decision prose over an
available measurement is precisely what this audit was convened to catch.

**Consequential side-finding, logged not fixed (hard rule 1):** D-101's "Remaining open
item" paragraph still reads "pending NetBox assignment (gap #3)" for literals D-111
adopted on 2026-07-11. That is a DOCFIX-class contradiction of a later ruling --
DOCFIX-200/204 class -- and it is what misled this session. Added to the Phase-3 batch.

OPERATOR UTTERANCE: *(none required -- withdrawn, not ruled)*

R2 directs that the org ULA /48 and per-DC GUA carve be assigned; it does not assign
them. This is an apex assignment and under D-136's unruled coupling the working apex
is `office1-netbox` (10.10.1.10), with `netbox.baldurkeep.com` a read-only v1
reference.

What is already measured: the only IPv6 in the whole cloud today is
`2602:f3e2:f01:100::/64` on the Office1 base fabric, so a GUA allocation of
`2602:f3e2:f01::/48` demonstrably exists and is partly in use. The ULA side has no
existing assignment at all.

I am deliberately NOT proposing specific prefixes -- picking your address space is
yours, and hard rule 2 forbids me inventing a literal. What needs deciding is the
SHAPE, after which the actual values are a mechanical carve:

**Options:**
- **(a)** ULA per RFC 4193 (pseudo-random /48), with a per-DC /56 out of it and a
  per-plane /64; GUA carved per-DC /56 out of the existing `2602:f3e2:f01::/48`.
  Symmetric with the D-134 v4 band discipline.
- **(b)** ULA only on the v6-only planes (data-tenant, storage, replication) with
  GUA reserved for provider-public, per D-101's family matrix read strictly.
- **(c)** Assign in NetBox first and let the carve fall out of the apex record --
  couples this to the D-136 render-pipeline question rather than pre-empting it.

OPERATOR UTTERANCE:

---

## R3. The MTU budget is in neither of D-101's two permitted states

**Finding:** L2-5 (measured on both racks and across all 17 MAAS VLANs).

D-101's Tenant/MTU sub-policy allows exactly two shapes: raise the underlay to
jumbo (9000) end-to-end so tenant MTU stays 1500, OR accept 1500 and pin tenant
MTU to about 1444 consistently across ovn geneve, tenant-network MTU and amphora.
It also says: "The measured underlay MTU is a Phase-0 gate -- do not assume jumbo."

Measured: the six plane bridges on both racks are MTU **9000**; every MAAS VLAN
record (17/17) says **1500**; the rack transit leg `enp1s0` is **1500**, so
cross-DC replication is not jumbo end-to-end; and `grep -i mtu bundle.yaml
overlays/*.yaml` returns NOTHING. That is neither branch. `scripts/dc-dc-mtu-geneve-budget.sh`
exists and correctly refuses to guess (`FAIL: --underlay-mtu is REQUIRED`), but has
never been run to a recorded verdict.

Consequence if unanswered: the deploy SUCCEEDS and the failure appears later as
tenant/geneve blackholing -- D-101 names this "the classic nested-OpenStack failure
mode". It is a bundle option that must be set BEFORE deploy.

**Options:**
- **(a)** Raise the rack transit and the MAAS VLAN records to 9000 so the underlay
  is genuinely jumbo end-to-end, and keep tenant MTU at 1500.
- **(b)** Accept a 1500 underlay and pin ~1444 consistently across ovn geneve,
  tenant-network MTU and amphora, set explicitly in the bundle.

OPERATOR UTTERANCE: **"Raise the two lagging segments to 9000 (Recommended)"** --
RULED 2026-07-27. **R3 IS CLOSED.** Recorded as a D-101 RULING NOTE (D-102, the
original MTU sub-policy, is merged into D-101 and directs amendments there).
**The question was re-framed by measurement before it was put:** the budget script
had never been run to a verdict, and running it showed the jumbo branch is nearly
complete already -- every vcloud MESH leg including the inter-DC `virbr5` is 9000,
as are all six plane bridges on both racks. Only TWO segments lag: the rack transit
NIC `enp1s0` inside both containment VMs, and the 17 MAAS VLAN records. The four
1500 legs are the D-125 simulated-ISP uplinks and must STAY 1500. Capture:
`docs/audit/mtu-budget-20260727.txt`. Coupled to R2: the 56-byte overhead is the
IPv6 figure and applies because dual-stack was ruled. Execution is a separate gated
step; the verification owed is a behavioural large-frame test with DF set across the
inter-DC path, not a reading of interface MTUs.

---

## R4. D-134's reserved address bands exist in prose only

**Finding:** L2-2 (measured).

MAAS holds exactly THREE ipranges cloud-wide and ALL THREE are `type=dynamic`.
There are ZERO `type=reserved` ranges anywhere. `maas admin subnet
unreserved-ip-ranges` reports the `.50-.99` VIP band as allocatable on **12 of 12**
DC plane subnets. Only the node statics and the `.201-.254` dynamic ranges are
protected. The tool that creates these reservations, `phase-00-maas-standup.sh`,
REFUSES to run for any non-VR0 DC (`:136`), so no VR1 path to create them exists.

Consequence: nothing stops MAAS handing a VIP-band address to a Juju/LXD container
during the deploy, and `phase-04-network-verify.sh:100` already hard-fails if the
FIP pool is not a reserved iprange.

**Options:**
- **(a)** Create the reserved ipranges on all 12 subnets AND carve the FIP pool
  now, as one gated mutation batch before deploy.
- **(b)** Create the VIP-band reservations now; defer the FIP pool to phase-04
  where its own verifier expects it.
- **(c)** Accept the bands as unreserved for the rehearsal and rely on static
  assignment, recording the risk explicitly.

OPERATOR UTTERANCE: **"Build a DC-aware tool; full v4 scheme + FIP now, v6 bands
after the carve (Recommended)"** -- RULED 2026-07-27. **R4 IS CLOSED.** Recorded as a
**D-134 AMENDMENT (2026-07-27)**. Re-measured before presenting, which sharpened it
considerably: the collision is QUANTIFIED -- dc1 metal-admin's lowest free span is
`.5-.99` (95 addrs), exactly the utility+VIP bands, against **27 LXD units in the base
bundle rising to ~55** with the HA overlay; zero `10.12.*` addresses are allocated today
so nothing has collided YET. Confirmed there is genuinely NO VR1 path (only
`site-headend-install.sh` and `phase-00-maas-standup.sh` can create ipranges, and the
latter correctly REFUSES non-VR0). **New architectural content: D-134's bands are
v4-only, so R2's dual-stack ruling left the v6 planes with no band discipline at all --
the amendment establishes that they inherit an equivalent scheme.** The v6 pass is
FORCED to follow the carve (a range cannot be reserved on a subnet that does not exist),
not deferred by choice.

---

## R5. Designate ships in the bundle, so Stage 5 will deploy it -- and that breaks Stage 7's gate

**Finding:** L6-1, corroborated by L3-2.

`bundle.yaml` carries live `designate`, `designate-bind`, `designate-mysql-router`
and `designate-hacluster` blocks plus 8 relations (DOCFIX-167 closed that on
2026-07-10). So `juju deploy ./bundle.yaml` deploys Designate at Stage 5. But
Stage 5's own runbook says the bundle "explicitly ships NO designate", and Stage 7
Step 5's gate requires "the diff shows ONLY the new designate/... applications
being added" -- which can never be true if they are already there.

**Options:**
- **(a)** Suppress Designate for the Stage-5 deploy (a subtractive overlay, or a
  rendered per-stage bundle) and keep Stage 7's incremental-add shape intact.
- **(b)** Accept Designate landing at Stage 5 and rewrite Stage 7 Steps 1/5 into a
  "configure, not deploy" shape. ~~**Caveat worth weighing:** D-106's own bootstrap
  order puts `os-public-hostname` + FQDN-SAN certificates BEFORE Designate;
  option (b) inverts that order.~~ **THIS CAVEAT WAS WRONG AND IS STRUCK.** D-106's
  order is a CONFIGURATION sequence (hostname -> certs -> zones -> neutron); it governs
  when the DNS wiring happens, not when the charm is installed. Deploying the app
  earlier does not invert it -- the zones still follow the certs. See the correction
  recorded in the D-106 RULING NOTE 2026-07-27.

OPERATOR UTTERANCE: **"Accept at Stage 5; rewrite Stage 7 Step 5 to configure-not-deploy
(Recommended)"** -- RULED 2026-07-27. **R5 IS CLOSED.** Recorded as a **D-106 RULING NOTE
(2026-07-27)**, amending nothing in D-106's order. Measured before presenting: all four
designate apps and all EIGHT relations are deploy-ready at Stage 5 (every peer --
mysql-innodb-cluster, keystone, rabbitmq-server, vault, memcached, designate-bind -- is
created by Stage 5), but they will be FUNCTIONALLY INERT because `os-public-hostname` is
set in no deploy artifact. **New surface defect found: `dc-dc-phase6-designate-cos-magnum.md`
CONTRADICTS ITSELF** -- `:172` says no designate application block exists anywhere,
`:178` says designate is deployed in-bundle. The bundle settles it (DOCFIX-167, 2026-07-10),
and the runbook needs rewriting regardless of this ruling. Option (c) -- setting
`os-public-hostname` at Stage 5 -- was refused as the one branch that genuinely DOES
collide with D-106: it recreates the D-019 root cause (metal-only charms pulling a public
FQDN endpoint they cannot resolve) before FQDN-SAN certs exist.

---

## R6. Whether to apply the HA scale-up overlay at Stage 5

**Finding:** L6-4.

`overlays/dc-ha-scaleup.yaml` sets `ceph-radosgw` and `designate` to `num_units: 3`
with `cluster_count: 3`. But Stage 6's radosgw multisite procedure and the
DOCFIX-165 script behind it are single-unit-shaped: measured,
`dc-dc-radosgw-multisite.sh --help` exposes `master-init ... --unit U`, one unit,
with no all-units mode, and the runbook says `juju run ceph-radosgw/0 restart`.
Realm/period membership would land on one of three gateways and Stage 6's Step-4
gate could false-green.

Scaling 1 -> 3 after the fact is a live-cloud change, which is why this is a
Stage-5-time decision.

**Options:**
- **(a)** Apply the HA overlay at Stage 5 and fix the Stage-6 radosgw path to be
  multi-unit-aware first.
- **(b)** Deploy single-unit at Stage 5 and scale up after Stage 6's multisite
  work, accepting a live-cloud scale-out later.

OPERATOR UTTERANCE: **"Close the two VIP gaps first, then apply the overlay whole
(Recommended)"** -- RULED 2026-07-27. **R6 IS CLOSED.** Recorded as a **D-121 RULING
NOTE (2026-07-27)**. Measured before presenting, which reframed the question: the
overlay scales 14 apps and moves all 12 base hacluster subordinates to
`cluster_count: 3`, but **eleven of twelve principals already carry proper VIPs** --
the gap is exactly TWO apps, and they differ in kind. designate is in the base bundle
with an hacluster and no vip; **vault has no hacluster in base AT ALL** -- the overlay
INTRODUCES the `vault-hacluster` application, its `cluster_count: 3` and the `vault:ha`
relation, and still no vip. **23 relations consume `vault:certificates`**, so
remediating on a live 3-unit vault would mean re-pointing all of them and re-issuing
SANs on a running cloud. **Consequence for sequencing: R11 is now a hard Stage-5
precondition ordered BEFORE the overlay**, not a parallel item. The Stage-6 radosgw
single-unit path and the absence of any `cluster_count` check remain tracked separately.

---

## R7. Do the two DCs share one Octavia CA, or get independent trust domains?

**Finding:** L7-6. The runbook itself flags this correctly and says it must not
silently default to reuse -- but the call has never been made.

`overlays/octavia-pki.yaml` is a single unscoped path holding CA private keys plus
a plaintext issuing-CA passphrase. Reusing it across both DCs puts one amphora
control-plane CA private key across two clouds that D-100 defines as independent.
Note the contrast one step later in the same runbook: per D-109 each DC's Vault is
its OWN independent root CA, no regional root-of-trust.

For a commercial multi-tenant cloud with hard tenant isolation, shared-CA is the
weaker posture -- but it is your call, and the rehearsal cost differs.

**IMPORTANT -- the runbook offers you a choice that is currently impossible on one
side.** Lens 5 (L5-3) measured the generator: `runbooks/phase-01-bundle-deploy.md:369-373`
reads the octavia VIP out of `bundle.yaml` and hard-gates it with
`grep -qE '^10\.12\.4\.[0-9]{1,3}$' || { echo "FAIL: implausible VIP -- stop"; exit 1; }`.
dc1's octavia VIP is `10.12.64.57` and lives in `overlays/vr1-dc1-vips.yaml:37`, not
in `bundle.yaml` at all. Its CN and SAN are hardcoded
`octavia-controller.omega.dc0.vr0.cloud.neumatrix.local` and the CA subject is
`/CN=VR0 DC0 Omega Cloud Octavia Controller CA`. So option (a) requires a generator
fix first, and option (b) bakes a dc0 CN/SAN into dc1's Octavia trust domain.
Also note `runbooks/phase-01-bundle-deploy.md:144-145` hard-ABORTS the deploy if the
overlay is absent -- which it currently is.

**Options:**
- **(a)** Regenerate fresh per-DC Octavia PKI -- independent trust domains per DC,
  consistent with D-109's per-DC Vault root. **Requires fixing the generator's
  dc0-frozen VIP gate and CN/SAN literals first** (a DOCFIX, no choice in it).
- **(b)** Reuse the existing CA across both DCs for the rehearsal, recording both
  the divergence from D-109's per-DC posture AND the dc0 CN/SAN in dc1's chain.
- **(c)** Fix the generator now and defer the trust-domain decision until it can
  actually be executed either way.

OPERATOR UTTERANCE: **"Per-DC independent Octavia PKI; fix the generator first
(Recommended)"** -- RULED 2026-07-27. **R7 IS CLOSED.** Recorded as a **D-109 AMENDMENT
(2026-07-27)** extending per-DC cryptographic independence from Vault roots to the
Octavia amphora control-plane PKI. Measured refinement: the generator is dc0-frozen in
TWO ways of DIFFERENT severity -- the CA SUBJECT is a baked `VR0 DC0` literal, but the
controller cert's SAN is **already DERIVED per-DC by design** (DOCFIX-067, "never a baked
literal"), so only the subject and the `^10\.12\.4\.` VIP gate need work. Reuse was
refused on posture: the overlay carries CA private keys plus a plaintext passphrase in a
repo SEC-004 records as PUBLIC, so sharing one amphora CA across two clouds D-100 defines
as independent would widen an existing exposure. Roosevelt analog: same "no cross-DC
shared secret" principle already ruled for per-DC MAAS power keys (SEC-012/-016).
**Caveat carried forward:** the generator's VIP gate must read the MERGED deploy input,
not `bundle.yaml`, or it breaks again the moment ruling-3's VIP extraction and R11's
`.61`/`.62` land.

---

## R8. Octavia `lb-mgmt-net` address family

**Finding:** L6-9. Register item 13 carries the same open fork.

Stage 5's runbook states plainly that "Octavia's `lb-mgmt-net` IPv6 support is a
real, open risk, not resolved" and that it must be decided before the
Ceph-over-v6 / geneve-over-v6 gate is declared closed. Stage 6's ENTRY condition
requires that gate to have passed, yet Stage 5's own text sanctions recording
"blocked on Step 6" -- so Stage 5 can close without producing Stage 6's
precondition, which GA-R6 E3 forbids resolving by conditional close.

This is downstream of R2: if R2 goes v4-only, this question largely dissolves.

**Options:**
- **(a)** Pin `lb-mgmt-net` to IPv4 for this deployment regardless of R2, and
  record it as a scoped exception to the dual-stack ruling.
- **(b)** Attempt IPv6 `lb-mgmt-net` and make it a named Stage-5 verification.
- **(c)** Defer until R2 is answered, then re-present.

OPERATOR UTTERANCE:

---

## R9. Where do dc1's OpenStack-layer network literals live?

> **ANSWER R2 FIRST.** R2 decides what the VIP literals ARE, and this question only
> asks where they live. If R2 goes v4-only, `overlays/dc-dc-ipv6-family-matrix.yaml`
> drops out of the deploy input and this becomes materially simpler. If it goes
> dual-stack, the per-key `vip:` REPLACE collision that L3-9 measured -- one overlay
> order hard-fails with ten "vip not a triple" errors, the reverse order silently
> drops EVERY v6 leg and still reports PASS -- has to be solved BEFORE either
> authority location is populated, or you will populate it with the wrong values.

**Finding:** L1-1 (with an explicit guard), extended by L7-10 and L7-3.

`scripts/lib-net.sh`'s `vr1-dc1` arm `unset`s `VIP_PREFIX_*`, `FIP_POOL_*`,
`VIP_COUNT_EXPECT` and `KEYSTONE_VIP_DEFAULT` on the stated grounds that they are
"NOT yet ruled/measured". They have since been ruled (D-134 amendment, 2026-07-23)
and built (`overlays/vr1-dc1-vips.yaml`). Any Stage-5 script that correctly calls
the selector now dies under `set -u` on a value that exists.

**GUARD -- do not let this be "fixed" mechanically.** That arm unsets TWO groups
for TWO different reasons. `METAL_INTERNAL_VID` and `METAL_INTERNAL_IFACE` are
**correctly** unset: D-133 abolished the VLAN-103 / `br-internal` stack for VR1 and
those facts genuinely do not exist. Only the VIP/FIP/keystone group is superseded.

Compounding context (L7-3, measured): of 27 `lib-net.sh` consumers, only 6 call
`lib_net_select_dc` at all. The rest source it unconditionally and silently get
VR0/dc0's literals -- so Stage 5 Steps 7-9 would write dc0's `10.12.4/8/12` values
against a DC whose planes are `10.12.64-84`. Whichever option you pick, that
consumer sweep is the larger half of the work.

**Options:**
- **(a)** Populate `lib-net.sh`'s dc1 arm from the D-134 amendment -- `lib-net.sh`
  stays the single authority for network literals.
- **(b)** Re-point the `phase-0*` scripts at `overlays/vr1-dc1-vips.yaml` as the
  authority, leaving `lib-net.sh` for substrate facts only -- closer to where
  D-136 would eventually take this.

OPERATOR UTTERANCE:

---

## R10. On what basis does Stage 5 start against a red preflight?

**Finding:** L1-9, with this session's measurement.

Stage 5's stated entry gate is "`preflight.sh` PASS". Measured today, preflight
exits 1, and the red set is exactly the known one: P4's missing
`overlays/octavia-pki.yaml` (a gitignored secret, absent by design), P4's "MAAS
unreachable from the jumphost" (expected -- the region is on voffice1), and P5's 7
credential findings. Nothing new has joined. But "PASS" is unreachable as written,
so the gate as stated can never authorise Stage 5.

Note this interacts with L4-2: P3 currently verifies ZERO of 33 charm-channel pins
because `juju` is not installed on the host preflight runs on.

**Options:**
- **(a)** Re-express the Stage-5 entry gate as a NAMED subset that must be green
  (e.g. P1, P2 and P5-with-known-residuals), with the known-red items listed as
  accepted preconditions.
- **(b)** Fix the reds first -- place the octavia overlay, run preflight from the
  headend where MAAS is reachable, remediate the 7 credential findings.
- **(c)** Record a dated, explicit exception basis for this stage only.

OPERATOR UTTERANCE:

---

## R11. Vault and Designate both have HA intent and no VIP

> **ANSWER R2 FIRST.** Whether the VIPs you add here are single-family or
> dual-family follows directly from R2. Adding v4-only VIPs and then re-doing them
> as dual-family means re-issuing certificate SANs against changed endpoints on a
> live cloud.

**Finding:** the vault half was already recorded; L3-2 found the SECOND case.

`grep -n 'vip' bundle.yaml` returns exactly 11 lines -- none for vault, none for
designate. Both are scaled to 3 with an `hacluster` subordinate related and
`cluster_count: 3`. `provider-bundle-check.py:133-135` skips any app with no `vip`
(`if not vip: continue`), which is why neither has ever been flagged.

Compounding (L4-3, verified by this session): `cluster_count` is checked NOWHERE in
`scripts/` or `tests/` -- a 3 -> 1 rewrite of all 20 occurrences yields a
byte-identical PASS. So neither before nor after the deploy does anything assert
that HA is real (L4-10: `cloud-assert.sh` reports "Cluster ID uniform across units"
over a SINGLE unit).

**Options:**
- **(a)** Add per-DC VIPs for both vault and designate in the symmetric overlay
  shape, and extend the checker to fail on any app with an hacluster relation and
  no VIP, plus any `cluster_count` that does not match its principal's `num_units`.
- **(b)** VIPs for both, checker work deferred to a follow-up.
- **(c)** Deploy them without VIPs deliberately (recording why consumers reaching a
  unit address is acceptable here).

OPERATOR UTTERANCE: **"Both full triples (.61 vault, .62 designate), dual-family, and
fix the gate (Recommended)"** -- RULED 2026-07-27. **R11 IS CLOSED.** Recorded as a
**D-020 AMENDMENT (2026-07-27)**. Key finding from measuring first: **vault was ALREADY
RULED and never built** -- D-020 enumerates it by name as carrying both a provider and a
metal VIP, and base vault has an EMPTY options block. Second ruled-but-unbuilt decision
this audit found, after D-134's bands. designate is genuinely new (absent from D-020) and
is ADDED. Octets `.61`/`.62` continue a consecutive map ending at ceph-radosgw `.60`.
Refused (b) vault-metal-only: contradicts D-020's enumeration and makes vault the only
non-triple. Mechanical: `OCTET_LO/HI` and `VIP_OCTET_MAX` widen to `.99` (two separately
named constants, two files), `VIP_COUNT_EXPECT` 11 -> 13. Gate hardening ruled IN SCOPE --
the checker learns to fail on hacluster-without-VIP.

---

# PART B -- STANDING / NOT STAGE-5 BLOCKING

Real, evidenced, and safe to answer after Stage 5 starts. Kept separate so the
eleven above are not diluted.

---

## R12. G17's scope: does it carry the node time-source check?

**Finding:** L1-8. `docs/CURRENT-STATE.md:852` lists only the artifact-reachability
commands, while `docs/dc-dc-deployment-workflow.md:206` says "Node-side
reachability **and the node time source** are gate G17" and the phase-4 runbook
requires `chronyc sources` to show the MAAS-served source, not the DC edge
(D-129(iv)). CURRENT-STATE also self-contradicts on whether DoD bullet 6 is STRUCK
or awaiting a DOCFIX. The observation window is one-time -- first boot.

Separately (L4-7), G17's check as written cannot fail: `curl -sI` exits 0 on
404/500 and the dc0 URL is an autoindex root that answers 200 with nothing behind it.

**Options:** (a) fold time verification into G17's `[V]` text and fix the check to
assert content with an exit-code predicate; (b) give time verification its own gate
row; (c) confirm it is STRUCK and remove the two conflicting surfaces.

OPERATOR UTTERANCE:

---

## R13. Credential reproducibility -- convert mint-refs now, or build `creds-mint.sh` first?

**Finding:** L7-7 (measured: 30 rows across 16 ids carry `mint-ref=operator-terminal`;
`grep -rnI "ssh-keygen" .` returns ZERO hits repo-wide).

Lens 7's assessment is worth quoting because it changes the shape of the fix: the
sharp edge is VR1-PRESENT, not Roosevelt-future -- SEC-007/-015 make edge SSH the
only management path, so a jumphost rebuild locks you out of both DC edges TODAY.
And the minimum fix needs no new tool: record each mint invocation as a numbered
runbook step and flip those rows' `mint-ref` from `operator-terminal` to
`runbook:<path>:<step>`, which the existing S4 check already resolves.
`creds-mint.sh` is orthogonal -- it prevents the NEXT unregistered mint; it does not
make an existing key reproducible.

**Options:** (a) convert the six edge/service/power key rows to `runbook:` refs
before Stage 5 (they are the unrecoverable-in-place ones); (b) build and rule
`creds-mint.sh` first, since Stage 5 is the largest minting event; (c) both, in
that order.

OPERATOR UTTERANCE:

---

## R14. The register cannot express a RULED exception

**Finding:** carried from the 2026-07-27 close and re-measured today -- 3 of the 7
standing credential findings are S5 power-key asymmetries that SEC-016 RULED to be
correct by design. The register has no way to say "this asymmetry is ruled", so it
reports a permanent red that a reader learns to ignore. That is how a real finding
gets lost.

**Options:** (a) add a ruled-exception field to the matrix, citing the SEC/D number,
which the checker honours and prints; (b) leave it red and rely on prose; (c) rework
the S5 rule so a ruled per-DC divergence is representable.

OPERATOR UTTERANCE:

---

## R15. Should the gauntlet and repo-lint pin a floor?

**Finding:** L4-8 and L4-1, both verified by this session.

`run-tests-all.sh` counts what it DISCOVERS and compares that count to nothing, so a
renamed or deleted harness is neither run nor failed and the gauntlet still prints
ALL GREEN. `repo-lint` reports `PASS (0 fail, 0 warn)` over ZERO files given a
one-character typo. Both are the gates every stage close cites. The fixes are small
and mechanical, but they change what "green" means, so they are worth your explicit
sign-off rather than my assumption.

**Options:** (a) add a floor to both (minimum harness count; refuse a non-directory
root and a zero-file scan); (b) floor on the gauntlet only; (c) leave as-is and rely
on the operator noticing a changed count.

OPERATOR UTTERANCE:
