| 2026-07-27 |

R7 RULED: per-DC independent Octavia PKI (D-109 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Per-DC independent Octavia PKI; fix the generator first
(Recommended)".
Recorded as a D-109 AMENDMENT. D-109 established per-DC INDEPENDENT Vault roots
with a DR-honest rationale but never mentioned the Octavia amphora
control-plane PKI, which is a SEPARATE trust domain with its own CA generated
outside Vault by phase-01 step 1.0-GEN. This extends the same posture to it.
MEASURED REFINEMENT THAT NARROWS THE WORK. The generator is dc0-frozen in two
ways of DIFFERENT severity, and only one is a design problem:
- The CA SUBJECT is a baked literal, "/CN=VR0 DC0 Omega Cloud Octavia Controller
CA/O=Neumatrix". Reusing it roots dc1's amphora chain in a VR0-DC0-named CA.
- The controller cert's SAN is ALREADY DERIVED per-DC by design -- the step
states it carries the controller FQDN, octavia API FQDN and the Octavia API VIP
"derived from the bundle at generation time -- DOCFIX-067; never a baked
literal". That half needs no change at all.
- The actual blocker is the VIP gate, grep -qE '^10\.12\.4\.', which accepts only
dc0's band while dc1's octavia VIP is 10.12.64.57 and lives in an overlay.
A DOCFIX was owed regardless of the ruling: the generator cannot produce a dc1
artifact as written, and phase-01:144-145 hard-ABORTS the deploy when
overlays/octavia-pki.yaml is absent, which it is. Option (c) would therefore have
deferred only the part D-109 already answers by precedent while leaving the
required work untouched.
REUSE REFUSED ON POSTURE. The overlay carries CA private keys plus a plaintext
issuing-CA passphrase inside the repo clone, and SEC-004 records the repo as
still PUBLIC. Sharing one amphora control-plane CA private key across two clouds
D-100 defines as independent would widen an existing exposure rather than
contain it, in a commercial multi-tenant cloud with hard tenant isolation.
Independent per-DC CAs keep that blast radius to one DC.
CAVEAT CARRIED FORWARD, because it will bite otherwise: the generator's VIP gate
must be re-pointed at the MERGED deploy input rather than bundle.yaml. Left
reading the base bundle it breaks again the moment ruling-3's VIP extraction and
R11's new .61/.62 allocations land.
Roosevelt analog: per-DC amphora CAs match per-DC IPMI/BMC credentials and the
per-DC MAAS power keys of SEC-012/-016 -- the same no-cross-DC-shared-secret
principle already ruled twice in this deployment.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R11 RULED: vault .61 and designate .62, dual-family triples (D-020 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Both full triples (.61 vault, .62 designate), dual-family,
and fix the gate (Recommended)".
THE FINDING THAT MEASURING FIRST PRODUCED: vault was ALREADY RULED and never
built. D-020's decision text enumerates vault BY NAME among the clustered
applications carrying both a provider and a metal VIP; measured, base vault is
num_units:1 with an EMPTY options block. This is a conformance repair of a
2026-era decision, not a new choice -- and it is the SECOND ruled decision this
audit has found unimplemented, after D-134's address bands (R4). Worth stating
plainly: the pattern is ruled-but-never-built, and nothing in the repo could
detect either case.
designate is genuinely new -- absent from D-020's enumeration -- so it needed a
ruling rather than a repair, and this amendment adds it. Its dnsaas endpoint is
already bound provider-public, so a provider leg is coherent for it.
Shape is the ESTABLISHED triple, not a new form. The measured octet map is
consecutive: keystone .50 through ceph-radosgw .60, all provider/admin/internal
triples. vault takes .61, designate .62, both dual-family per R2 -- adding
v4-only now and re-doing them later would mean re-issuing certificate SANs on a
live cloud, and for vault that is 23 certificate relations.
Option (b), vault metal-only per the 2026-07-25 expansion review, was refused and
the conflict is RECORDED so that proposal is not later mistaken for the ruled
position: it has a real technical argument (all 23 vault consumers are internal)
but it contradicts D-020's own enumeration and would leave vault the single
non-triple in the bundle, a permanent special case for the checker.
Mechanical consequences, no choice in them: .61/.62 are legal under D-134's
amended .50-.99 band but REJECTED by the gate today -- provider-bundle-check.py
holds OCTET_LO/HI = 50,60 and lib-net.sh holds VIP_OCTET_MAX=60. These are
SEPARATELY NAMED constants in two files, so widening the band is a two-file
change (L3-7). VIP_COUNT_EXPECT moves 11 -> 13.
Gate hardening ruled IN SCOPE rather than deferred: the checker learns to FAIL on
an application with an hacluster relation and no vip. Justification is measured --
grep -rn cluster_count scripts/ tests/ returns NOTHING, and a constructed overlay
rewriting all 20 cluster_count values 3 -> 1 yields a byte-identical PASS (L4-3).
That is exactly why decorative HA was found by an audit instead of by a gate.
Per the R6 ruling, these VIPs land BEFORE dc-ha-scaleup.yaml is applied.
Execution is a separate gated step.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R6 RULED: close the two VIP gaps, then apply the HA overlay whole (D-121 note)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Close the two VIP gaps first, then apply the overlay whole
(Recommended)".
THE TENSION IT RESOLVES. D-121 is titled "VR1 makes HA real -- scale the
decorative single-unit control plane to 3". Deferring the overlay would have
deployed VR1 in exactly the shape D-121 was written to retire. Applying it as-is
would have shipped, for vault, precisely the defect D-121 exists to remove -- a
pacemaker cluster with nothing to manage. Only the ruled sequence satisfies the
decision rather than half of it.
MEASURED BEFORE PRESENTING, and it reframed the question. The overlay scales 14
applications and moves all 12 base hacluster subordinates to cluster_count 3, but
ELEVEN OF TWELVE principals already carry proper VIP triples. The gap is exactly
two applications, and they are not the same shape:
- designate is in the base bundle WITH an hacluster and an ha binding, no vip.
- vault has NO hacluster in base at all and an EMPTY options block. The overlay
INTRODUCES the vault-hacluster application, its cluster_count 3, and the
vault:ha relation -- and still no vip.
Vault is the deciding case because of blast radius: 23 relations consume
vault:certificates, plus barbican-vault:secrets-storage. With no VIP every
consumer binds a unit address and there is nothing for a failover to move.
Remediating that on a live 3-unit vault means re-pointing 23 certificate
relations and re-issuing SANs on a running cloud, which is why the faster option
was refused.
SEQUENCING CONSEQUENCE: R11 becomes a HARD Stage-5 precondition ordered BEFORE
the overlay, not a parallel item.
Scope is tractable and the pattern is already in-repo: octavia carries a correct
triple (10.12.4.57 10.12.8.57 10.12.12.57), so R11 has a shape to copy rather
than design.
NOT resolved here, tracked separately: the Stage-6 radosgw multisite path is
single-unit-shaped while this scales ceph-radosgw to 3; and
provider-bundle-check.py checks cluster_count NOWHERE, which is precisely why
decorative HA was found by an audit instead of by a gate.
Address demand is not an argument against this -- R4's band ruling already covers
the overlay's growth from 27 to ~55 LXD units.
Execution is a separate gated step, not authorised by this ruling.
Revert: git revert this commit; the ruling note is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R5 RULED: designate deploys at Stage 5, is configured at Stage 7 (D-106 note)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Accept at Stage 5; rewrite Stage 7 Step 5 to
configure-not-deploy (Recommended)".
CORRECTION RECORDED IN THE DECISION ITSELF, not just this message. When R5 was
first put to the operator, the audit claimed this option "inverts D-106's
bootstrap order, which puts os-public-hostname + FQDN-SAN certs BEFORE
Designate". That was WRONG. D-106's order is a CONFIGURATION sequence -- static
hosts, then os-public-hostname, then Vault FQDN-SAN certs, then zones and A/AAAA,
then neutron, then tenant subnets. It governs when the DNS wiring happens, not
when the charm is installed. Deploying the app earlier does not invert it: the
zones still follow the certs. Conflating "install the charm" with "do D-106's
work" nearly cost a decision made on a false constraint, so the correction is
recorded in D-106 where a future session will meet it.
MEASURED BEFORE PRESENTING. All four designate applications and all EIGHT
relations are deploy-ready at Stage 5 -- every relation peer (mysql-innodb-cluster,
keystone, rabbitmq-server, vault, memcached, designate-bind) is created by Stage
5. But they will be FUNCTIONALLY INERT: os-public-hostname is set in NO deploy
artifact, and bundle.yaml:11 records the current posture as IP-ONLY with the dual
VIPs as the catalog endpoint. The charm runs; the feature does not. That is
exactly the state Stage 7 exists to resolve.
NEW SURFACE DEFECT FOUND BY THE MEASUREMENT: the phase-6 runbook contradicts
ITSELF. Line 172 states "There is no designate: or designate-bind: application
block anywhere in ..."; line 178 states "Designate is deployed in-bundle in each
DC". Both cannot be true. The bundle settles it -- DOCFIX-167 put designate there
on 2026-07-10 -- and the runbook needs rewriting regardless of this ruling.
Option (c), setting os-public-hostname at Stage 5, was refused and the reason is
worth keeping: it is the ONE branch that genuinely collides with D-106. Publishing
a public FQDN endpoint before Vault has issued FQDN-SAN certs recreates the exact
D-019 root cause D-106 was written to remove -- metal-only charms pulling a public
FQDN endpoint they cannot resolve, which is also the D-021 amphora constraint.
Option (b), suppressing designate at Stage 5, was refused because it needs
machinery that does not exist and partially undoes DOCFIX-167 -- building a
mechanism to satisfy a stale gate rather than fixing the gate.
Execution is a DOCFIX against the phase-6 runbook (Step 1 inventory prose, Step 5
verb and gate), logged to the Phase-3 batch, not executed under this ruling.
Revert: git revert this commit; the ruling note is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R4 RULED: reserved bands become enforced in MAAS via a DC-aware tool (D-134 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Build a DC-aware tool; full v4 scheme + FIP now, v6 bands
after the carve (Recommended)".
MEASURED BEFORE ASKING, which sharpened the question considerably.
The defect is worse than "bands not enforced": D-134's band table has never
existed anywhere but prose. maas admin ipranges read returns THREE ranges
cloud-wide, ALL type=dynamic, ZERO reserved.
The collision is now QUANTIFIED rather than hypothetical. MAAS's own
subnet unreserved-ip-ranges for dc1 metal-admin reports its lowest free span as
10.12.68.5-.99 (95 addrs) -- precisely the .4-.49 utility band and the whole
.50-.99 VIP band. Against that, the dc1 bundle places 27 LXD units in base form
and ~55 once dc-ha-scaleup scales 14 applications, each needing a MAAS-allocated
address, and the VIP band is where the 33 per-DC VIPs live. Zero 10.12.*
addresses are allocated today, so this is a PRE-EMPTION and not an incident.
MAAS's exact allocation ORDER is deliberately NOT asserted -- lens 2 was right to
refuse that claim and I have not added it.
Confirmed there is genuinely NO VR1 path: only site-headend-install.sh
(office1-only) and phase-00-maas-standup.sh can create an iprange, and the latter
REFUSES for any non-VR0 DC. That refusal is CORRECT -- it is the cross-DC mixing
guard lib-net's selector exists to enforce. The gap is that nothing replaced it.
NEW ARCHITECTURAL CONTENT, and it exists because of R2: D-134's bands are v4-only.
Ruling dual-stack made the 33 VIPs dual-family, which left the v6 planes with no
band discipline at all and the same collision waiting in the other family. The
amendment establishes that the v6 planes inherit an equivalent scheme. The exact
v6 octet mapping is NOT ruled -- mirroring the v4 host-part layout is the obvious
mechanical default, not a fresh design question.
The v6 pass following the carve is FORCED SEQUENCING, not a deferral: a reserved
range cannot be created on a subnet that does not yet exist, and the v6 plane
subnets are not in MAAS.
Worth noting what the tool buys beyond the reservation itself: a `check` arm gives
D-134 an EXECUTABLE gate. Today nothing in the repo can detect that the bands are
unenforced, which is why this survived from 2026-07-23 to now.
Execution is a separate gated step. Standard delivery discipline applies to the
tool (harness green, gauntlet, repo-lint, changelog with revert); the 12-subnet
reservation pass is an operator-gated live MAAS mutation and is not batched with it.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R3 RULED: finish the jumbo underlay; first recorded MTU budget verdict
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Raise the two lagging segments to 9000 (Recommended)".
Recorded as a D-101 RULING NOTE. D-102 is the original MTU sub-policy but is
MERGED INTO D-101 and its body directs amendments there.
MEASURED BEFORE ASKING, which materially re-framed the question. D-101 requires
"the measured underlay MTU is a Phase-0 gate -- do not assume jumbo", and
scripts/dc-dc-mtu-geneve-budget.sh had existed without ever being run to a
recorded verdict. Running it both ways, plus measuring the host bridges, showed
the jumbo branch is nearly complete already rather than a large project:
- Every vcloud MESH leg is ALREADY 9000, including mesh-vr1-dc0-vr1-dc1
(virbr5), the inter-DC path itself. All six plane bridges on both racks: 9000.
- The four 1500 legs are the D-125 SIMULATED-ISP uplinks. They model the
internet, D-125's egress gate is defined against them, and they must STAY
1500. Not a defect.
- Exactly TWO segments lag: enp1s0 inside both containment VMs, and all 17 MAAS
VLAN records.
The MAAS record is the one that bites silently: MAAS renders VLAN MTU into node
netplan, so a jumbo bridge beneath a 1500 record still yields 1500 node
interfaces. Jumbo bridges alone do not deliver a jumbo underlay.
Why (a): tenant MTU stays 1500, so no per-charm MTU coordination is needed at
all. Option (b) required ovn geneve + tenant-network MTU + amphora to agree
exactly and permanently across both DCs, and NOTHING in this repo checks MTU --
neither cloud-assert.sh nor provider-bundle-check.py has any MTU assertion -- so
drift would be silent, which is precisely what D-101 calls "the classic
nested-OpenStack failure mode". Option (c) was the worst: jumbo at the plane
layer with a still-capped transit throttles D-108 rbd-mirror and radosgw
multisite behind a chokepoint invisible where an operator would look.
COUPLING TO R2, recorded because it changes the arithmetic: the 56-byte overhead
is the IPv6 figure and applies BECAUSE dual-stack was ruled. Under v4-only it
would have been 42, and the 1500-underlay tenant MTU 1458. R2 and R3 are not
independent.
Execution is a SEPARATE gated step and is NOT authorised by this ruling. The
verification owed is BEHAVIOURAL: an end-to-end large-frame test with DF set
across the inter-DC path. An interface claiming 9000 is not proof a 9000-byte
frame survives the path -- the assert-on-content rule applied to MTU.
Also worth recording: invoked bare, the budget script correctly REFUSES ("FAIL:
--underlay-mtu is REQUIRED -- no default, measure it this session"). It does not
guess. That discipline is why this verdict is trustworthy.
Revert: git revert this commit; the ruling note and capture are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

CORRECTION: the v6 literals are ASSIGNED (D-111), not pending -- R2a withdrawn
...
The operator asked whether the Office1 NetBox apex had actually been polled. It
had not. It has now been read, and the answer inverts a claim I made one commit
ago and a question I put in front of them.
WHAT IS TRUE: the v6 literals are assigned, ratified and recorded, and have been
since 2026-07-11 under D-111 (ADOPTED). Measured from
netbox/draft/vr1-office1-current-20260725.json -- 139 prefixes, 103 IPv6, every
relevant row tagged D-101/D-111: ULA fd50:840e:74e2::/48 with DC0 planes at
:220/:221/:230/:240/:250::/64 and DC1 at :320/:321/:330/:340/:350::/64; GUA
provider-public DC0 2602:f3e2:f02:10::/64 + VIP f02:11::/64, DC1
2602:f3e2:f03:10::/64 + VIP f03:11::/64.
WHAT I GOT WRONG: the R2 ruling note claimed D-101's "Remaining open item" (the
org ULA /48 and per-DC GUA carve) had become a Stage-5 precondition because
"dual-stack cannot deploy against literals that do not exist". They exist. R2a
("which literals") is WITHDRAWN as never having been open; no operator utterance
is owed on it.
THE REAL PRECONDITION IS NARROWER AND BETTER: propagation, not assignment, and it
needs no ruling. The ratified values are absent from the two places Stage 5
actually reads -- scripts/lib-net.sh carries no v6 arm at all, and MAAS carries
no v6 on any of the 12 DC plane fabrics. Both are mechanical copies from an
authoritative source, so both move from Part A (needs a decision) to the Phase-3
mechanical batch.
Note CURRENT-STATE's "DRIFT-FREE across netbox/lib-net/artifacts" is NOT
contradicted: drift-free means values appearing in more than one place agree, not
that coverage is complete. lib-net.sh simply never gained a v6 arm.
PROCESS FAILURE, OWNED AND RECORDED IN THE DECISION ITSELF: the audit's own lens 2
listed the apex as UNMEASURED and warned, in terms, that "D-101's literals may
exist in NetBox and simply not be carved into MAAS". That warning was correct and
available, and I put a question to the operator anyway on the strength of D-101's
stale prose. Trusting stale decision prose over an available measurement is
exactly what this audit was convened to catch.
NEW DOCFIX-CLASS FINDING, logged not fixed (hard rule 1): D-101's "Remaining open
item" paragraph still reads "pending NetBox assignment (gap #3)" for literals
D-111 adopted on 2026-07-11 -- a decision's own prose contradicted by a later
ruling, DOCFIX-200/204 class, and the direct cause of this error. Queued in the
Phase-3 batch.
The R2 ruling itself STANDS -- dual-stack as ruled. Only its stated effect changed.
Revert: git revert this commit to restore the prior (incorrect) framing; the R2
ruling in the preceding commit is independent and should not be reverted with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R2 RULED: dual-stack as originally decided; the v6 literals move onto the critical path
...
GA-R5: question and exact utterance quoted in the Status block, dated, pushed
before dependent work.
Operator utterance: "Carve v6 and deploy dual-stack as ruled (Recommended)".
Recorded as a D-101 RULING NOTE -- a re-confirmation, amending nothing, in the
same shape as the 2026-07-25 note it follows.
WHY IT WAS A LIVE QUESTION AT ALL: the 07-25 ruling was measured against the
substrate for the first time by this audit and found unimplemented. Exactly ONE
IPv6 subnet exists cloud-wide and it is on the Office1 base fabric; none of the
12 DC plane fabrics carries one; zero v6 links across all 18 nodes; every node
reports default_gateways.ipv6 = NONE. D-101's matrix requires ULA on
data-tenant/storage/replication plus legs on metal-admin/metal-internal, with
nothing to bind against.
THE CONSEQUENTIAL EFFECT, and the reason this commit matters beyond recording a
choice: D-101's own "Remaining open item" -- the org ULA /48 and per-DC GUA carve
-- has been carried since authoring as "pending NetBox assignment ... not a
ratification question", i.e. non-blocking. It is now a STAGE-5 PRECONDITION.
Dual-stack cannot be deployed against literals that do not exist.
Downstream, all now inheriting dual-family: R9 (where dc1's literals live), R11
(vault/designate VIPs, which must be dual-family from the outset or their cert
SANs get re-issued on a live cloud), and the L3-9 overlay collision, which must
be reconciled BEFORE either authority location is populated. Worth restating: of
the two merge orders the audit measured, the DANGEROUS one is the one that
PASSES -- vips-overlay-last silently replaces every v6 leg and reports green.
R8 (octavia's family) is explicitly NOT resolved by this ruling.
NEW SUB-QUESTION OPENED, queued not ruled: R2a, which literals. The ruling
directs that they be assigned; it does not assign them. I have deliberately not
proposed prefixes -- that is the operator's address space, and hard rule 2
forbids inventing a literal. Only the SHAPE is presented.
Revert: git revert this commit; the ruling note and R2a are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R1 RULED: ceph-osd gets a real second disk on the storage nodes (D-121 amendment)
...
GA-R5: question as presented and the operator's exact utterance are both quoted
in the Status block, dated, and pushed BEFORE any dependent work.
Operator utterance: "Add an OSD volume to node-vm (Recommended)".
Recorded as a D-121 AMENDMENT rather than a new D-number. D-121 is the decision
that ratifies modules/node-vm sizing/count, and its own capacity re-validation
already records "Ceph disk re-run for 4 storage/DC = PASS 5.31 TiB" -- the second
disk was BUDGETED and never BUILT. GA-R3: architectural consequence, but it
amends a ruled node layout rather than establishing new ground.
TWO SCOPE CORRECTIONS landed with the ruling, both of which narrow it:
- It is EIGHT volumes, not eighteen. ceph-osd is placed on the four storage
nodes per DC (bundle.yaml:556-557, to: ["5","6","7","8"]). The other five
nodes per DC neither run ceph-osd nor need a second device. My earlier framing
said "all 18 nodes" -- true of the measured defect, wrong about the fix.
- Rack headroom measured before asserting feasibility: dc0 2.0T available, dc1
2.9T available under /var/lib/libvirt, thin-provisioned qcow2.
Why (a) over (b): MINIMIZE DELTA TO ROOSEVELT -- Roosevelt storage nodes have
real dedicated disks. Option (b) was also not safe to pick as presented: whether
ceph-osd at the pinned squid/stable supports a directory- or partition-backed
OSD was explicitly UNRESEARCHED, and the audit refused to assert it either way.
THE APPLY IS NOT AUTHORISED BY THIS RULING. Four preconditions recorded, the
sharpest being that re-commissioning is required for MAAS to see the new device,
and whether that preserves the D-134 statics and pinned MACs is UNVERIFIED --
the 2026-07-20 MAC-regeneration incident is the precedent for exactly that class.
modules/node-vm is shared by both DCs, so the change must be additive/opt-in or
it plans against all 18 domains.
Revert: git revert this commit; the amendment is additive and no artifact changed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Stage-5 grounding audit: post-review corrections + dc1 egress MEASURED
...
Three fixes from a final review pass, one of which is a new measurement.
MEASUREMENT -- dc1 edge egress is OPEN, probed rather than assumed. Lens 5 had
correctly listed this as UNMEASURED (it rested on a record). Probed directly
from the dc1 rack with --noproxy so an apt-cacher hit could not fake it:
streams.canonical.com/juju/tools/ -> 200 (the agent stream bootstrap actually
needs), api.snapcraft.io reachable, archive.ubuntu.com -> 200, ping 1.1.1.1 0%
loss, default route via 10.12.64.1. So Phase 5.2 stays a standing caveat rather
than becoming a blocker -- but it is now measured, and the row says to re-probe
immediately before bootstrap because the measurement has a shelf life.
DELIVERABLE CORRECTIONS:
- The readiness doc said "Eleven decisions are yours", which reads as the total.
It is 11 blocking (R1-R11) plus 4 standing (R12-R15) = 15, now stated as such,
with the GA-R5 one-at-a-time rule restated at the point of use.
- Phase 3's intro still said the runbook fails "in at least five distinct ways"
from before lens 5 added ten more rows. Corrected to 21 listed corrections and
nine independent failure paths, with the rows named.
- R9 and R11 are NOT independent of R2 and were presented as if they were. Both
now carry an explicit "ANSWER R2 FIRST" dependency, matching R8. R2 decides
what the VIP literals ARE; R9 only asks where they live, and R11's VIPs are
single- or dual-family depending on it. Under dual-stack, the L3-9 overlay
collision must be solved BEFORE either authority is populated.
Also recorded for future sessions: there is no Bash(timeout *) allow rule, which
is why lens 1's `timeout 240 bash scripts/preflight.sh` was denied while the bare
command is allowed -- a failed-to-MATCH, same class as the 2026-07-26 quoted-sudo
finding, not a classifier override.
Revert: git revert this commit; all changes are to audit artifacts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Stage-5 grounding audit: lens 5 + the three deliverables
...
VERDICT: Stage 5 would NOT run error-free today, and would fail early. The
substrate underneath is excellent -- all three OpenTofu roots ZERO DIFF, 18
nodes Ready exact to D-121 Option C, MACs and power addresses matching
lib-hosts, D-134 statics perfect, 17 fabrics, zero orphaned interfaces, both
artifact paths serving, gauntlet ALL GREEN (81). What is not ready is the
layer between the substrate and the deploy.
LENS 5 (relaunched after an API error) returned the ordered precondition list
and the largest runbook defect found: Step 4 says "follow phase-01 verbatim",
and phase-01 ACTIVELY REFUSES dc1 -- its VIP guard greps bundle.yaml for
eleven 10.12.4.x VIPs, dc1's live in an overlay on 10.12.64/68/72, so it
takes the "ABORT: VIP guard failed" branch. After the ruled VIP extraction it
aborts for dc0 too. It also carries hardcoded VR0 system_ids, a jumphost-local
libvirt loop over disks that do not exist, and a 4-machine plan gate against a
9-machine bundle.
Three Stage-5 gate commands cannot execute at all, measured against the juju
actually installed (3.6.27): juju run used for a shell command when it is the
action runner; download-backup given a backup-id when it takes a controller
path; and the geneve gate grepping ovn-central for a config key the same
runbook says does not exist. Four of the five VERIFY-LIVE gates the record
says Stage 5 owes have NO step in the runbook -- including the keystone
policyd-override check, which is RULED.
Two convergent confirmations raise confidence in the whole set: the Ceph OSD
blocker was found independently by two lenses using different methods, and the
stale-clone blocker independently by this session and lens 5.
DELIVERABLES:
- docs/audit/stage5-readiness-20260727.md -- ordered precondition checklist,
READ FIRST. 5 phases, each row with status/evidence/what-breaks.
- docs/audit/stage5-committee-raw-20260727.md -- all 7 lenses verbatim.
- docs/audit/queued-rulings-20260727.md -- 11 Stage-5-blocking + 4 standing
questions, GA-R5 shape, one exchange each, blank utterance lines. NONE
adopted; a batch answer rules NOTHING.
MECHANICAL FIX TAKEN (exactly one, deliberately): the G3 gate row read OPEN
with a standing FREEZE while its own cited evidence file records G3 CLOSED and
the freeze lifted. Left standing, that clause would have blocked the very
DOCFIX batch this audit queues. No ruling was required -- two surfaces already
declared it closed.
The 21-item DOCFIX remediation batch is LOGGED NOT EXECUTED. Nearly every
runbook fix interlocks with an unanswered ruling, so landing them now would
encode assumptions about questions the operator has not answered.
Next-free numbers unchanged (D 138 / DOCFIX 205 / BUNDLEFIX 053) -- no number
was assigned, correctly, since nothing was remediated.
Revert: git revert this commit; the deliverables are new files and the
CURRENT-STATE edits are additive plus the one G3 correction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|