Nothing here is adopted. Every item is PRESENTED, never picked (hard rule 1; GA-R5: PROPOSED means present options). Each question is a SEPARATE exchange -- GA-R5 rule 1 makes a batch adoption INVALID, so answering "yes to all" rules NOTHING and the session that reads this must stop and re-ask.
How to use this file. Answer questions ONE AT A TIME. Write your exact words on the OPERATOR UTTERANCE: line. A ruling exists only once its Status block quotes the question AND your exact utterance, dated, committed and pushed -- before any dependent work starts. An ambiguous or template answer rules nothing.
Evidence for every claim below is in docs/audit/stage5-committee-raw-20260727.md (verbatim lens returns) and docs/audit/stage5-live-measurement-20260727.txt (this session's own measurements). Finding IDs are cited so nothing here has to be taken on trust.
Ordering note: R1, R2 and R3 change what gets deployed. They should be answered before the rest, because several later questions have different right answers depending on them.
These must be answered before juju deploy. Each one, left unanswered, either stops the deploy or bakes in a state that is expensive to reverse on a live cloud.
Finding: L2-1 (measured twice -- virsh domblklist on both racks AND MAAS physicalblockdevice_set on all 18 nodes). Verified independently by this session.
bundle.yaml:560 sets osd-devices: /dev/vdb. Every one of the 18 VR1 node VMs has exactly ONE block device, vda. opentofu/modules/node-vm/main.tf:104-125 declares a single disk and neither substrate root has any OSD-disk variable. The /dev/vdb line is a VR0 as-built comment ("libvirt-attached, MAAS-untracked") -- VR0 hosts had an attached second disk; VR1 nodes never did.
Consequence if unanswered: ceph-osd deploys onto four storage nodes and finds no device. Ceph never forms, and everything storage-backed behind it stalls.
Options:
modules/node-vm and re-apply BOTH substrate roots. Closest to Roosevelt (real machines have real disks), but it re-opens the substrate on 18 running-but-powered-off nodes and both roots currently plan ZERO DIFF -- that property is deliberately being spent.osd-devices at a directory or partition path on the existing vda. Cheapest, no substrate change, but it rehearses a Ceph layout Roosevelt will not use, which cuts against MINIMIZE DELTA TO ROOSEVELT.OPERATOR UTTERANCE: "Add an OSD volume to node-vm (Recommended)" -- RULED 2026-07-27. R1 IS CLOSED. Recorded as a D-121 AMENDMENT (the governing decision already ratifies modules/node-vm sizing, and its own capacity re-validation already budgeted "Ceph disk re-run for 4 storage/DC = PASS 5.31 TiB" -- the disk was budgeted and never built). Two scope corrections landed with the ruling: it is 8 volumes, not 18 (ceph-osd is placed on the four storage nodes per DC only), and the apply is a SEPARATE operator-gated step with four preconditions, including verifying that re-commissioning preserves the D-134 statics and pinned MACs BEFORE the apply -- the 2026-07-20 MAC-regeneration incident is the precedent. Full text: docs/design-decisions.md, "D-121 AMENDMENT 2026-07-27".
Finding: L2-3 (measured). D-101's RULING NOTE of 2026-07-25 records your exact words -- "Dual stack to be used where IPv4 is required, IPv6 where IPv6 only makes sense" and "Dual stack deployment for DC0 and DC1. This is the deployment when the dual-stack is added" -- with the stated effect that "the v4-only phasing option ... is CLOSED, for BOTH DC0 and DC1 in this deployment."
Measured reality: exactly ONE IPv6 subnet exists cloud-wide (2602:f3e2:f01:100::/64 on the Office1 base fabric). None of the 12 DC plane fabrics carries an IPv6 subnet. Zero IPv6 links across all 18 nodes; every node reports default_gateways.ipv6 = NONE. D-101's own family matrix requires ULA on data-tenant, storage and replication and a ULA leg on metal-admin/metal-internal -- none of which has any MAAS v6 presence to bind against. D-101's own "Remaining open item" is the un-assigned NetBox literals: the org ULA /48 and the per-DC GUA carve.
This is the single largest fork in the audit. It is BLOCKING because addresses become as-built Keystone endpoints and Vault-issued cert SANs at deploy time.
Options:
OPERATOR UTTERANCE: "Carve v6 and deploy dual-stack as ruled (Recommended)" -- RULED 2026-07-27. R2 IS CLOSED. Recorded as a D-101 RULING NOTE (re-confirmation, no amendment -- docs/design-decisions.md is the authority). Consequence that changes the plan: D-101's "Remaining open item" (the org ULA /48 and per-DC GUA carve), until now carried as non-blocking "pending NetBox assignment", is now a STAGE-5 PRECONDITION -- dual-stack cannot deploy against literals that do not exist. R9 and R11 both inherit "dual-family" from this. The L3-9 overlay collision must be reconciled BEFORE either authority location is populated; note the dangerous direction is the one that PASSES (vips overlay last silently drops every v6 leg and reports green). R8 is NOT resolved by this ruling. R2a WITHDRAWN -- see below.
This question was asked in error and is withdrawn before any ruling. The operator asked whether the NetBox apex had been polled. It had not. It has now been read, and the literals are ASSIGNED, RATIFIED AND RECORDED -- and have been since 2026-07-11 under D-111 (ADOPTED).
MEASURED from netbox/draft/vr1-office1-current-20260725.json (139 prefixes, 103 IPv6), every row tagged D-101/D-111: ULA fd50:840e:74e2::/48 with DC0 planes at :220/:221/:230/:240/:250::/64 and DC1 at :320/:321/:330/:340/:350::/64; GUA provider-public DC0 2602:f3e2:f02:10::/64 + VIP f02:11::/64, DC1 2602:f3e2:f03:10::/64 + VIP f03:11::/64.
What is actually owed is PROPAGATION, and it needs no ruling. The ratified values are absent from the two places Stage 5 reads: scripts/lib-net.sh carries no v6 arm at all, and MAAS carries no v6 on any of the 12 DC plane fabrics. Both are mechanical copies from an authoritative source. This moves from Part A (needs a decision) to the Phase-3 mechanical batch in the readiness doc.
Why the error happened, recorded because it is the audit's own failure mode: the audit's lens 2 explicitly listed the apex as UNMEASURED and warned "D-101's literals may exist in NetBox and simply not be carved into MAAS -- L2-3's claim is scoped to MAAS + node reality and does not assert the apex is empty." That warning was correct and available, and a question was put to the operator anyway, on the strength of D-101's own stale "Remaining open item" prose. Trusting stale decision prose over an available measurement is precisely what this audit was convened to catch.
Consequential side-finding, logged not fixed (hard rule 1): D-101's "Remaining open item" paragraph still reads "pending NetBox assignment (gap #3)" for literals D-111 adopted on 2026-07-11. That is a DOCFIX-class contradiction of a later ruling -- DOCFIX-200/204 class -- and it is what misled this session. Added to the Phase-3 batch.
OPERATOR UTTERANCE: (none required -- withdrawn, not ruled)
R2 directs that the org ULA /48 and per-DC GUA carve be assigned; it does not assign them. This is an apex assignment and under D-136's unruled coupling the working apex is office1-netbox (10.10.1.10), with netbox.baldurkeep.com a read-only v1 reference.
What is already measured: the only IPv6 in the whole cloud today is 2602:f3e2:f01:100::/64 on the Office1 base fabric, so a GUA allocation of 2602:f3e2:f01::/48 demonstrably exists and is partly in use. The ULA side has no existing assignment at all.
I am deliberately NOT proposing specific prefixes -- picking your address space is yours, and hard rule 2 forbids me inventing a literal. What needs deciding is the SHAPE, after which the actual values are a mechanical carve:
Options:
2602:f3e2:f01::/48. Symmetric with the D-134 v4 band discipline.OPERATOR UTTERANCE:
Finding: L2-5 (measured on both racks and across all 17 MAAS VLANs).
D-101's Tenant/MTU sub-policy allows exactly two shapes: raise the underlay to jumbo (9000) end-to-end so tenant MTU stays 1500, OR accept 1500 and pin tenant MTU to about 1444 consistently across ovn geneve, tenant-network MTU and amphora. It also says: "The measured underlay MTU is a Phase-0 gate -- do not assume jumbo."
Measured: the six plane bridges on both racks are MTU 9000; every MAAS VLAN record (17/17) says 1500; the rack transit leg enp1s0 is 1500, so cross-DC replication is not jumbo end-to-end; and grep -i mtu bundle.yaml overlays/*.yaml returns NOTHING. That is neither branch. scripts/dc-dc-mtu-geneve-budget.sh exists and correctly refuses to guess (FAIL: --underlay-mtu is REQUIRED), but has never been run to a recorded verdict.
Consequence if unanswered: the deploy SUCCEEDS and the failure appears later as tenant/geneve blackholing -- D-101 names this "the classic nested-OpenStack failure mode". It is a bundle option that must be set BEFORE deploy.
Options:
OPERATOR UTTERANCE: "Raise the two lagging segments to 9000 (Recommended)" -- RULED 2026-07-27. R3 IS CLOSED. Recorded as a D-101 RULING NOTE (D-102, the original MTU sub-policy, is merged into D-101 and directs amendments there). The question was re-framed by measurement before it was put: the budget script had never been run to a verdict, and running it showed the jumbo branch is nearly complete already -- every vcloud MESH leg including the inter-DC virbr5 is 9000, as are all six plane bridges on both racks. Only TWO segments lag: the rack transit NIC enp1s0 inside both containment VMs, and the 17 MAAS VLAN records. The four 1500 legs are the D-125 simulated-ISP uplinks and must STAY 1500. Capture: docs/audit/mtu-budget-20260727.txt. Coupled to R2: the 56-byte overhead is the IPv6 figure and applies because dual-stack was ruled. Execution is a separate gated step; the verification owed is a behavioural large-frame test with DF set across the inter-DC path, not a reading of interface MTUs.
Finding: L2-2 (measured).
MAAS holds exactly THREE ipranges cloud-wide and ALL THREE are type=dynamic. There are ZERO type=reserved ranges anywhere. maas admin subnet unreserved-ip-ranges reports the .50-.99 VIP band as allocatable on 12 of 12 DC plane subnets. Only the node statics and the .201-.254 dynamic ranges are protected. The tool that creates these reservations, phase-00-maas-standup.sh, REFUSES to run for any non-VR0 DC (:136), so no VR1 path to create them exists.
Consequence: nothing stops MAAS handing a VIP-band address to a Juju/LXD container during the deploy, and phase-04-network-verify.sh:100 already hard-fails if the FIP pool is not a reserved iprange.
Options:
OPERATOR UTTERANCE: "Build a DC-aware tool; full v4 scheme + FIP now, v6 bands after the carve (Recommended)" -- RULED 2026-07-27. R4 IS CLOSED. Recorded as a D-134 AMENDMENT (2026-07-27). Re-measured before presenting, which sharpened it considerably: the collision is QUANTIFIED -- dc1 metal-admin's lowest free span is .5-.99 (95 addrs), exactly the utility+VIP bands, against 27 LXD units in the base bundle rising to ~55 with the HA overlay; zero 10.12.* addresses are allocated today so nothing has collided YET. Confirmed there is genuinely NO VR1 path (only site-headend-install.sh and phase-00-maas-standup.sh can create ipranges, and the latter correctly REFUSES non-VR0). New architectural content: D-134's bands are v4-only, so R2's dual-stack ruling left the v6 planes with no band discipline at all -- the amendment establishes that they inherit an equivalent scheme. The v6 pass is FORCED to follow the carve (a range cannot be reserved on a subnet that does not exist), not deferred by choice.
Finding: L6-1, corroborated by L3-2.
bundle.yaml carries live designate, designate-bind, designate-mysql-router and designate-hacluster blocks plus 8 relations (DOCFIX-167 closed that on 2026-07-10). So juju deploy ./bundle.yaml deploys Designate at Stage 5. But Stage 5's own runbook says the bundle "explicitly ships NO designate", and Stage 7 Step 5's gate requires "the diff shows ONLY the new designate/... applications being added" -- which can never be true if they are already there.
Options:
os-public-hostname + FQDN-SAN certificates BEFORE Designate; option (b) inverts that order.OPERATOR UTTERANCE:
Finding: L6-4.
overlays/dc-ha-scaleup.yaml sets ceph-radosgw and designate to num_units: 3 with cluster_count: 3. But Stage 6's radosgw multisite procedure and the DOCFIX-165 script behind it are single-unit-shaped: measured, dc-dc-radosgw-multisite.sh --help exposes master-init ... --unit U, one unit, with no all-units mode, and the runbook says juju run ceph-radosgw/0 restart. Realm/period membership would land on one of three gateways and Stage 6's Step-4 gate could false-green.
Scaling 1 -> 3 after the fact is a live-cloud change, which is why this is a Stage-5-time decision.
Options:
OPERATOR UTTERANCE:
Finding: L7-6. The runbook itself flags this correctly and says it must not silently default to reuse -- but the call has never been made.
overlays/octavia-pki.yaml is a single unscoped path holding CA private keys plus a plaintext issuing-CA passphrase. Reusing it across both DCs puts one amphora control-plane CA private key across two clouds that D-100 defines as independent. Note the contrast one step later in the same runbook: per D-109 each DC's Vault is its OWN independent root CA, no regional root-of-trust.
For a commercial multi-tenant cloud with hard tenant isolation, shared-CA is the weaker posture -- but it is your call, and the rehearsal cost differs.
IMPORTANT -- the runbook offers you a choice that is currently impossible on one side. Lens 5 (L5-3) measured the generator: runbooks/phase-01-bundle-deploy.md:369-373 reads the octavia VIP out of bundle.yaml and hard-gates it with grep -qE '^10\.12\.4\.[0-9]{1,3}$' || { echo "FAIL: implausible VIP -- stop"; exit 1; }. dc1's octavia VIP is 10.12.64.57 and lives in overlays/vr1-dc1-vips.yaml:37, not in bundle.yaml at all. Its CN and SAN are hardcoded octavia-controller.omega.dc0.vr0.cloud.neumatrix.local and the CA subject is /CN=VR0 DC0 Omega Cloud Octavia Controller CA. So option (a) requires a generator fix first, and option (b) bakes a dc0 CN/SAN into dc1's Octavia trust domain. Also note runbooks/phase-01-bundle-deploy.md:144-145 hard-ABORTS the deploy if the overlay is absent -- which it currently is.
Options:
OPERATOR UTTERANCE:
lb-mgmt-net address familyFinding: L6-9. Register item 13 carries the same open fork.
Stage 5's runbook states plainly that "Octavia's lb-mgmt-net IPv6 support is a real, open risk, not resolved" and that it must be decided before the Ceph-over-v6 / geneve-over-v6 gate is declared closed. Stage 6's ENTRY condition requires that gate to have passed, yet Stage 5's own text sanctions recording "blocked on Step 6" -- so Stage 5 can close without producing Stage 6's precondition, which GA-R6 E3 forbids resolving by conditional close.
This is downstream of R2: if R2 goes v4-only, this question largely dissolves.
Options:
lb-mgmt-net to IPv4 for this deployment regardless of R2, and record it as a scoped exception to the dual-stack ruling.lb-mgmt-net and make it a named Stage-5 verification.OPERATOR UTTERANCE:
ANSWER R2 FIRST. R2 decides what the VIP literals ARE, and this question only asks where they live. If R2 goes v4-only,
overlays/dc-dc-ipv6-family-matrix.yamldrops out of the deploy input and this becomes materially simpler. If it goes dual-stack, the per-keyvip:REPLACE collision that L3-9 measured -- one overlay order hard-fails with ten "vip not a triple" errors, the reverse order silently drops EVERY v6 leg and still reports PASS -- has to be solved BEFORE either authority location is populated, or you will populate it with the wrong values.
Finding: L1-1 (with an explicit guard), extended by L7-10 and L7-3.
scripts/lib-net.sh's vr1-dc1 arm unsets VIP_PREFIX_*, FIP_POOL_*, VIP_COUNT_EXPECT and KEYSTONE_VIP_DEFAULT on the stated grounds that they are "NOT yet ruled/measured". They have since been ruled (D-134 amendment, 2026-07-23) and built (overlays/vr1-dc1-vips.yaml). Any Stage-5 script that correctly calls the selector now dies under set -u on a value that exists.
GUARD -- do not let this be "fixed" mechanically. That arm unsets TWO groups for TWO different reasons. METAL_INTERNAL_VID and METAL_INTERNAL_IFACE are correctly unset: D-133 abolished the VLAN-103 / br-internal stack for VR1 and those facts genuinely do not exist. Only the VIP/FIP/keystone group is superseded.
Compounding context (L7-3, measured): of 27 lib-net.sh consumers, only 6 call lib_net_select_dc at all. The rest source it unconditionally and silently get VR0/dc0's literals -- so Stage 5 Steps 7-9 would write dc0's 10.12.4/8/12 values against a DC whose planes are 10.12.64-84. Whichever option you pick, that consumer sweep is the larger half of the work.
Options:
lib-net.sh's dc1 arm from the D-134 amendment -- lib-net.sh stays the single authority for network literals.phase-0* scripts at overlays/vr1-dc1-vips.yaml as the authority, leaving lib-net.sh for substrate facts only -- closer to where D-136 would eventually take this.OPERATOR UTTERANCE:
Finding: L1-9, with this session's measurement.
Stage 5's stated entry gate is "preflight.sh PASS". Measured today, preflight exits 1, and the red set is exactly the known one: P4's missing overlays/octavia-pki.yaml (a gitignored secret, absent by design), P4's "MAAS unreachable from the jumphost" (expected -- the region is on voffice1), and P5's 7 credential findings. Nothing new has joined. But "PASS" is unreachable as written, so the gate as stated can never authorise Stage 5.
Note this interacts with L4-2: P3 currently verifies ZERO of 33 charm-channel pins because juju is not installed on the host preflight runs on.
Options:
OPERATOR UTTERANCE:
ANSWER R2 FIRST. Whether the VIPs you add here are single-family or dual-family follows directly from R2. Adding v4-only VIPs and then re-doing them as dual-family means re-issuing certificate SANs against changed endpoints on a live cloud.
Finding: the vault half was already recorded; L3-2 found the SECOND case.
grep -n 'vip' bundle.yaml returns exactly 11 lines -- none for vault, none for designate. Both are scaled to 3 with an hacluster subordinate related and cluster_count: 3. provider-bundle-check.py:133-135 skips any app with no vip (if not vip: continue), which is why neither has ever been flagged.
Compounding (L4-3, verified by this session): cluster_count is checked NOWHERE in scripts/ or tests/ -- a 3 -> 1 rewrite of all 20 occurrences yields a byte-identical PASS. So neither before nor after the deploy does anything assert that HA is real (L4-10: cloud-assert.sh reports "Cluster ID uniform across units" over a SINGLE unit).
Options:
cluster_count that does not match its principal's num_units.OPERATOR UTTERANCE:
Real, evidenced, and safe to answer after Stage 5 starts. Kept separate so the eleven above are not diluted.
Finding: L1-8. docs/CURRENT-STATE.md:852 lists only the artifact-reachability commands, while docs/dc-dc-deployment-workflow.md:206 says "Node-side reachability and the node time source are gate G17" and the phase-4 runbook requires chronyc sources to show the MAAS-served source, not the DC edge (D-129(iv)). CURRENT-STATE also self-contradicts on whether DoD bullet 6 is STRUCK or awaiting a DOCFIX. The observation window is one-time -- first boot.
Separately (L4-7), G17's check as written cannot fail: curl -sI exits 0 on 404/500 and the dc0 URL is an autoindex root that answers 200 with nothing behind it.
Options: (a) fold time verification into G17's [V] text and fix the check to assert content with an exit-code predicate; (b) give time verification its own gate row; (c) confirm it is STRUCK and remove the two conflicting surfaces.
OPERATOR UTTERANCE:
creds-mint.sh first?Finding: L7-7 (measured: 30 rows across 16 ids carry mint-ref=operator-terminal; grep -rnI "ssh-keygen" . returns ZERO hits repo-wide).
Lens 7's assessment is worth quoting because it changes the shape of the fix: the sharp edge is VR1-PRESENT, not Roosevelt-future -- SEC-007/-015 make edge SSH the only management path, so a jumphost rebuild locks you out of both DC edges TODAY. And the minimum fix needs no new tool: record each mint invocation as a numbered runbook step and flip those rows' mint-ref from operator-terminal to runbook:<path>:<step>, which the existing S4 check already resolves. creds-mint.sh is orthogonal -- it prevents the NEXT unregistered mint; it does not make an existing key reproducible.
Options: (a) convert the six edge/service/power key rows to runbook: refs before Stage 5 (they are the unrecoverable-in-place ones); (b) build and rule creds-mint.sh first, since Stage 5 is the largest minting event; (c) both, in that order.
OPERATOR UTTERANCE:
Finding: carried from the 2026-07-27 close and re-measured today -- 3 of the 7 standing credential findings are S5 power-key asymmetries that SEC-016 RULED to be correct by design. The register has no way to say "this asymmetry is ruled", so it reports a permanent red that a reader learns to ignore. That is how a real finding gets lost.
Options: (a) add a ruled-exception field to the matrix, citing the SEC/D number, which the checker honours and prints; (b) leave it red and rely on prose; (c) rework the S5 rule so a ruled per-DC divergence is representable.
OPERATOR UTTERANCE:
Finding: L4-8 and L4-1, both verified by this session.
run-tests-all.sh counts what it DISCOVERS and compares that count to nothing, so a renamed or deleted harness is neither run nor failed and the gauntlet still prints ALL GREEN. repo-lint reports PASS (0 fail, 0 warn) over ZERO files given a one-character typo. Both are the gates every stage close cites. The fixes are small and mechanical, but they change what "green" means, so they are worth your explicit sign-off rather than my assumption.
Options: (a) add a floor to both (minimum harness count; refuse a non-directory root and a zero-file scan); (b) floor on the gauntlet only; (c) leave as-is and rely on the operator noticing a changed count.
OPERATOR UTTERANCE: