| 2026-07-27 |

R3 RULED: finish the jumbo underlay; first recorded MTU budget verdict
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Raise the two lagging segments to 9000 (Recommended)".
Recorded as a D-101 RULING NOTE. D-102 is the original MTU sub-policy but is
MERGED INTO D-101 and its body directs amendments there.
MEASURED BEFORE ASKING, which materially re-framed the question. D-101 requires
"the measured underlay MTU is a Phase-0 gate -- do not assume jumbo", and
scripts/dc-dc-mtu-geneve-budget.sh had existed without ever being run to a
recorded verdict. Running it both ways, plus measuring the host bridges, showed
the jumbo branch is nearly complete already rather than a large project:
- Every vcloud MESH leg is ALREADY 9000, including mesh-vr1-dc0-vr1-dc1
(virbr5), the inter-DC path itself. All six plane bridges on both racks: 9000.
- The four 1500 legs are the D-125 SIMULATED-ISP uplinks. They model the
internet, D-125's egress gate is defined against them, and they must STAY
1500. Not a defect.
- Exactly TWO segments lag: enp1s0 inside both containment VMs, and all 17 MAAS
VLAN records.
The MAAS record is the one that bites silently: MAAS renders VLAN MTU into node
netplan, so a jumbo bridge beneath a 1500 record still yields 1500 node
interfaces. Jumbo bridges alone do not deliver a jumbo underlay.
Why (a): tenant MTU stays 1500, so no per-charm MTU coordination is needed at
all. Option (b) required ovn geneve + tenant-network MTU + amphora to agree
exactly and permanently across both DCs, and NOTHING in this repo checks MTU --
neither cloud-assert.sh nor provider-bundle-check.py has any MTU assertion -- so
drift would be silent, which is precisely what D-101 calls "the classic
nested-OpenStack failure mode". Option (c) was the worst: jumbo at the plane
layer with a still-capped transit throttles D-108 rbd-mirror and radosgw
multisite behind a chokepoint invisible where an operator would look.
COUPLING TO R2, recorded because it changes the arithmetic: the 56-byte overhead
is the IPv6 figure and applies BECAUSE dual-stack was ruled. Under v4-only it
would have been 42, and the 1500-underlay tenant MTU 1458. R2 and R3 are not
independent.
Execution is a SEPARATE gated step and is NOT authorised by this ruling. The
verification owed is BEHAVIOURAL: an end-to-end large-frame test with DF set
across the inter-DC path. An interface claiming 9000 is not proof a 9000-byte
frame survives the path -- the assert-on-content rule applied to MTU.
Also worth recording: invoked bare, the budget script correctly REFUSES ("FAIL:
--underlay-mtu is REQUIRED -- no default, measure it this session"). It does not
guess. That discipline is why this verdict is trustworthy.
Revert: git revert this commit; the ruling note and capture are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

CORRECTION: the v6 literals are ASSIGNED (D-111), not pending -- R2a withdrawn
...
The operator asked whether the Office1 NetBox apex had actually been polled. It
had not. It has now been read, and the answer inverts a claim I made one commit
ago and a question I put in front of them.
WHAT IS TRUE: the v6 literals are assigned, ratified and recorded, and have been
since 2026-07-11 under D-111 (ADOPTED). Measured from
netbox/draft/vr1-office1-current-20260725.json -- 139 prefixes, 103 IPv6, every
relevant row tagged D-101/D-111: ULA fd50:840e:74e2::/48 with DC0 planes at
:220/:221/:230/:240/:250::/64 and DC1 at :320/:321/:330/:340/:350::/64; GUA
provider-public DC0 2602:f3e2:f02:10::/64 + VIP f02:11::/64, DC1
2602:f3e2:f03:10::/64 + VIP f03:11::/64.
WHAT I GOT WRONG: the R2 ruling note claimed D-101's "Remaining open item" (the
org ULA /48 and per-DC GUA carve) had become a Stage-5 precondition because
"dual-stack cannot deploy against literals that do not exist". They exist. R2a
("which literals") is WITHDRAWN as never having been open; no operator utterance
is owed on it.
THE REAL PRECONDITION IS NARROWER AND BETTER: propagation, not assignment, and it
needs no ruling. The ratified values are absent from the two places Stage 5
actually reads -- scripts/lib-net.sh carries no v6 arm at all, and MAAS carries
no v6 on any of the 12 DC plane fabrics. Both are mechanical copies from an
authoritative source, so both move from Part A (needs a decision) to the Phase-3
mechanical batch.
Note CURRENT-STATE's "DRIFT-FREE across netbox/lib-net/artifacts" is NOT
contradicted: drift-free means values appearing in more than one place agree, not
that coverage is complete. lib-net.sh simply never gained a v6 arm.
PROCESS FAILURE, OWNED AND RECORDED IN THE DECISION ITSELF: the audit's own lens 2
listed the apex as UNMEASURED and warned, in terms, that "D-101's literals may
exist in NetBox and simply not be carved into MAAS". That warning was correct and
available, and I put a question to the operator anyway on the strength of D-101's
stale prose. Trusting stale decision prose over an available measurement is
exactly what this audit was convened to catch.
NEW DOCFIX-CLASS FINDING, logged not fixed (hard rule 1): D-101's "Remaining open
item" paragraph still reads "pending NetBox assignment (gap #3)" for literals
D-111 adopted on 2026-07-11 -- a decision's own prose contradicted by a later
ruling, DOCFIX-200/204 class, and the direct cause of this error. Queued in the
Phase-3 batch.
The R2 ruling itself STANDS -- dual-stack as ruled. Only its stated effect changed.
Revert: git revert this commit to restore the prior (incorrect) framing; the R2
ruling in the preceding commit is independent and should not be reverted with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R2 RULED: dual-stack as originally decided; the v6 literals move onto the critical path
...
GA-R5: question and exact utterance quoted in the Status block, dated, pushed
before dependent work.
Operator utterance: "Carve v6 and deploy dual-stack as ruled (Recommended)".
Recorded as a D-101 RULING NOTE -- a re-confirmation, amending nothing, in the
same shape as the 2026-07-25 note it follows.
WHY IT WAS A LIVE QUESTION AT ALL: the 07-25 ruling was measured against the
substrate for the first time by this audit and found unimplemented. Exactly ONE
IPv6 subnet exists cloud-wide and it is on the Office1 base fabric; none of the
12 DC plane fabrics carries one; zero v6 links across all 18 nodes; every node
reports default_gateways.ipv6 = NONE. D-101's matrix requires ULA on
data-tenant/storage/replication plus legs on metal-admin/metal-internal, with
nothing to bind against.
THE CONSEQUENTIAL EFFECT, and the reason this commit matters beyond recording a
choice: D-101's own "Remaining open item" -- the org ULA /48 and per-DC GUA carve
-- has been carried since authoring as "pending NetBox assignment ... not a
ratification question", i.e. non-blocking. It is now a STAGE-5 PRECONDITION.
Dual-stack cannot be deployed against literals that do not exist.
Downstream, all now inheriting dual-family: R9 (where dc1's literals live), R11
(vault/designate VIPs, which must be dual-family from the outset or their cert
SANs get re-issued on a live cloud), and the L3-9 overlay collision, which must
be reconciled BEFORE either authority location is populated. Worth restating: of
the two merge orders the audit measured, the DANGEROUS one is the one that
PASSES -- vips-overlay-last silently replaces every v6 leg and reports green.
R8 (octavia's family) is explicitly NOT resolved by this ruling.
NEW SUB-QUESTION OPENED, queued not ruled: R2a, which literals. The ruling
directs that they be assigned; it does not assign them. I have deliberately not
proposed prefixes -- that is the operator's address space, and hard rule 2
forbids inventing a literal. Only the SHAPE is presented.
Revert: git revert this commit; the ruling note and R2a are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R1 RULED: ceph-osd gets a real second disk on the storage nodes (D-121 amendment)
...
GA-R5: question as presented and the operator's exact utterance are both quoted
in the Status block, dated, and pushed BEFORE any dependent work.
Operator utterance: "Add an OSD volume to node-vm (Recommended)".
Recorded as a D-121 AMENDMENT rather than a new D-number. D-121 is the decision
that ratifies modules/node-vm sizing/count, and its own capacity re-validation
already records "Ceph disk re-run for 4 storage/DC = PASS 5.31 TiB" -- the second
disk was BUDGETED and never BUILT. GA-R3: architectural consequence, but it
amends a ruled node layout rather than establishing new ground.
TWO SCOPE CORRECTIONS landed with the ruling, both of which narrow it:
- It is EIGHT volumes, not eighteen. ceph-osd is placed on the four storage
nodes per DC (bundle.yaml:556-557, to: ["5","6","7","8"]). The other five
nodes per DC neither run ceph-osd nor need a second device. My earlier framing
said "all 18 nodes" -- true of the measured defect, wrong about the fix.
- Rack headroom measured before asserting feasibility: dc0 2.0T available, dc1
2.9T available under /var/lib/libvirt, thin-provisioned qcow2.
Why (a) over (b): MINIMIZE DELTA TO ROOSEVELT -- Roosevelt storage nodes have
real dedicated disks. Option (b) was also not safe to pick as presented: whether
ceph-osd at the pinned squid/stable supports a directory- or partition-backed
OSD was explicitly UNRESEARCHED, and the audit refused to assert it either way.
THE APPLY IS NOT AUTHORISED BY THIS RULING. Four preconditions recorded, the
sharpest being that re-commissioning is required for MAAS to see the new device,
and whether that preserves the D-134 statics and pinned MACs is UNVERIFIED --
the 2026-07-20 MAC-regeneration incident is the precedent for exactly that class.
modules/node-vm is shared by both DCs, so the change must be additive/opt-in or
it plans against all 18 domains.
Revert: git revert this commit; the amendment is additive and no artifact changed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Stage-5 grounding audit: post-review corrections + dc1 egress MEASURED
...
Three fixes from a final review pass, one of which is a new measurement.
MEASUREMENT -- dc1 edge egress is OPEN, probed rather than assumed. Lens 5 had
correctly listed this as UNMEASURED (it rested on a record). Probed directly
from the dc1 rack with --noproxy so an apt-cacher hit could not fake it:
streams.canonical.com/juju/tools/ -> 200 (the agent stream bootstrap actually
needs), api.snapcraft.io reachable, archive.ubuntu.com -> 200, ping 1.1.1.1 0%
loss, default route via 10.12.64.1. So Phase 5.2 stays a standing caveat rather
than becoming a blocker -- but it is now measured, and the row says to re-probe
immediately before bootstrap because the measurement has a shelf life.
DELIVERABLE CORRECTIONS:
- The readiness doc said "Eleven decisions are yours", which reads as the total.
It is 11 blocking (R1-R11) plus 4 standing (R12-R15) = 15, now stated as such,
with the GA-R5 one-at-a-time rule restated at the point of use.
- Phase 3's intro still said the runbook fails "in at least five distinct ways"
from before lens 5 added ten more rows. Corrected to 21 listed corrections and
nine independent failure paths, with the rows named.
- R9 and R11 are NOT independent of R2 and were presented as if they were. Both
now carry an explicit "ANSWER R2 FIRST" dependency, matching R8. R2 decides
what the VIP literals ARE; R9 only asks where they live, and R11's VIPs are
single- or dual-family depending on it. Under dual-stack, the L3-9 overlay
collision must be solved BEFORE either authority is populated.
Also recorded for future sessions: there is no Bash(timeout *) allow rule, which
is why lens 1's `timeout 240 bash scripts/preflight.sh` was denied while the bare
command is allowed -- a failed-to-MATCH, same class as the 2026-07-26 quoted-sudo
finding, not a classifier override.
Revert: git revert this commit; all changes are to audit artifacts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Stage-5 grounding audit: lens 5 + the three deliverables
...
VERDICT: Stage 5 would NOT run error-free today, and would fail early. The
substrate underneath is excellent -- all three OpenTofu roots ZERO DIFF, 18
nodes Ready exact to D-121 Option C, MACs and power addresses matching
lib-hosts, D-134 statics perfect, 17 fabrics, zero orphaned interfaces, both
artifact paths serving, gauntlet ALL GREEN (81). What is not ready is the
layer between the substrate and the deploy.
LENS 5 (relaunched after an API error) returned the ordered precondition list
and the largest runbook defect found: Step 4 says "follow phase-01 verbatim",
and phase-01 ACTIVELY REFUSES dc1 -- its VIP guard greps bundle.yaml for
eleven 10.12.4.x VIPs, dc1's live in an overlay on 10.12.64/68/72, so it
takes the "ABORT: VIP guard failed" branch. After the ruled VIP extraction it
aborts for dc0 too. It also carries hardcoded VR0 system_ids, a jumphost-local
libvirt loop over disks that do not exist, and a 4-machine plan gate against a
9-machine bundle.
Three Stage-5 gate commands cannot execute at all, measured against the juju
actually installed (3.6.27): juju run used for a shell command when it is the
action runner; download-backup given a backup-id when it takes a controller
path; and the geneve gate grepping ovn-central for a config key the same
runbook says does not exist. Four of the five VERIFY-LIVE gates the record
says Stage 5 owes have NO step in the runbook -- including the keystone
policyd-override check, which is RULED.
Two convergent confirmations raise confidence in the whole set: the Ceph OSD
blocker was found independently by two lenses using different methods, and the
stale-clone blocker independently by this session and lens 5.
DELIVERABLES:
- docs/audit/stage5-readiness-20260727.md -- ordered precondition checklist,
READ FIRST. 5 phases, each row with status/evidence/what-breaks.
- docs/audit/stage5-committee-raw-20260727.md -- all 7 lenses verbatim.
- docs/audit/queued-rulings-20260727.md -- 11 Stage-5-blocking + 4 standing
questions, GA-R5 shape, one exchange each, blank utterance lines. NONE
adopted; a batch answer rules NOTHING.
MECHANICAL FIX TAKEN (exactly one, deliberately): the G3 gate row read OPEN
with a standing FREEZE while its own cited evidence file records G3 CLOSED and
the freeze lifted. Left standing, that clause would have blocked the very
DOCFIX batch this audit queues. No ruling was required -- two surfaces already
declared it closed.
The 21-item DOCFIX remediation batch is LOGGED NOT EXECUTED. Nearly every
runbook fix interlocks with an unanswered ruling, so landing them now would
encode assumptions about questions the operator has not answered.
Next-free numbers unchanged (D 138 / DOCFIX 205 / BUNDLEFIX 053) -- no number
was assigned, correctly, since nothing was remediated.
Revert: git revert this commit; the deliverables are new files and the
CURRENT-STATE edits are additive plus the one G3 correction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|