| 2026-07-27 |

Phase 0 delivery cleanup: session changelog with reverts, and a mis-filed finding corrected
...
Three gaps in the Phase 0 delivery, plus one new measured finding.
1. SESSION CHANGELOG ADDED (GA-R2/D1) -- docs/changelog-20260727-stage5-phase0.md.
It was missing, and it is the only revert surface: no logged window was
opened for this session, so the changelog + capture are the entire
as-executed record of five live mutations (two branch deletions on origin,
two on voffice1, a 68-package install on the region host). Each item now
carries WHAT / WHY / HOW TO REVERT with the pre-state hashes recorded
(39e8988, 61c416e, 57836b1) and the apt purge scoped with a caution that
68 packages were pulled and some are shared.
2. CURRENT-STATE section 7 no longer carries two rows for one component. The
prior "ABSENT ON BOTH HOSTS" state is folded into the single pin row as
history, following the Juju row's in-row supersession precedent. A pin
table asserting both presence and absence of the same client is the same
stale-surface class the audit close-out sweep found the audit itself
creating.
3. A FINDING I MIS-FILED IS CORRECTED BY MEASUREMENT. P4's
"metal-admin gateway=none (want 10.12.8.1)" was dismissed as covered by
readiness item 3.7. It is not -- 3.7 is specifically the VID-103 /
br-internal assertions. Measured from the dc0 rack: 10.12.8.1 answers 0/2
pings with an INCOMPLETE ARP entry (never resolved -- nothing holds it),
against a control ping to the edge at 10.12.4.1 at 0% loss. So MAAS is
right to carry no gateway and lib-net.sh:37's expectation is the defect, a
VR0 inheritance; the D-134 carve ruled .1 gateways for the PROVIDER subnets
only. Logged as P0-5, not fixed (hard rule 1). The dc1 arm (lib-net.sh:149)
carries the same shape at 10.12.68.1.
Also queued for one GA-R5 exchange, no D-number assigned: which host is
authoritative for the gates. R10 (preflight on voffice1) and R15(3) (preflight
stops failing open) executed together make preflight-on-voffice1 permanently
unpassable given P0-2, and R15(2)'s manifest does not reach P0-1 -- so the R15
execution scope needs this answered first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Stage-5 Phase 0 EXECUTED: clone advanced, refs pruned, openstack client installed
...
Readiness-doc preconditions 0.1/0.2/0.3, operator-approved step by step
("Approve A, B, and C, and retire the audit branch"). NO STAGE OPENED.
0.1 voffice1 is at main (HEAD == origin/main == 6495cfb, asserted). The real
risk was that both DCs' inner tfstate -- the substrate's state-of-record --
lives inside that working tree; every state artifact was proven gitignored and
origin/main proven to track an identical file set at those paths BEFORE the
switch. Both tfstate sha256s are byte-identical after. The two dc1 overlays are
now present and bundle.yaml is the 9-node role-separated layout.
0.3 three stale remote-tracking refs pruned, two stale local branches deleted
(the record named one), containment proven first. The audit branch was retired
on origin BEFORE the fetch so the clone could not be handed a fresh stale ref.
0.2 python3-openstackclient 6.6.0-0ubuntu2 on voffice1, verified behaviourally.
The snap was refuted by measurement: no Caracal channel (newest stable zed,
2023), plus the home-only confinement already recorded at design-decisions:638.
Two NEW findings from the payoff runs, logged not fixed (hard rule 1):
- the gauntlet is HOST-DEPENDENT (2/81 on voffice1 vs ALL GREEN on vcloud, same
commit); both failures are host-portability defects, one of them an aggregate
FAIL over all-passing sub-checks
- preflight P5 is HOST-BLIND (34 findings vs 7), and between the two hosts there
is no single host on which preflight is currently correct
R10's ruled consequence is confirmed: P3 now verifies all 33 charm-channel pins
(previously zero) and P4 reports MAAS reachable.
Capture: docs/audit/stage5-phase0-20260727.txt
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
Record the grounding-audit merge to main (607813b) in CURRENT-STATE
...
GA-R1/C1: the merge commit itself cannot carry a doc edit, so the merge is
recorded here in a follow-up commit -- the same shape Stage 4 used (merge
6f5701d recorded by 1023596).
States explicitly that no stage opened or closed, cites the post-merge
gauntlet (81) and repo-lint, and flags that both figures are read under R15:
the harness count is compared against the record by hand because the manifest
R15(2) rules is not built, and repo-lint emits no files-scanned figure, which
is R15(1)'s finding observed rather than assumed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Merge branch 'dc-dc-stage5-grounding-audit' into main
...
Stage-5 grounding audit: 7-lens read-only committee + live measurement sweep,
run before Stage 5 opens. NO STAGE OPENED OR CLOSED on this branch.
Brings onto main:
- 14 GA-R5 rulings (R1-R13, R15), 2 withdrawn as raised in error (R2a, R14),
and gate G18 opened (blocking, deferred until the cloud is live).
Authority for each is its D-number Status block in docs/design-decisions.md;
questions + exact utterances in docs/audit/queued-rulings-20260727.md.
- Audit deliverables under docs/audit/: stage5-readiness-20260727.md (the
ordered precondition checklist), stage5-committee-raw-20260727.md,
stage5-live-measurement-20260727.txt, the unmeasured-gap register, and the
R6-R15 measurement captures.
- Gate-integrity findings: three gates that could not fail (repo-lint over
zero files, the gauntlet's uncompared harness count, preflight failing open
on rc not in {1,2}) -- ruled fixed under R15, NOT yet built.
- Two decisions found RULED-BUT-NEVER-BUILT: D-134's per-DC bands and D-020's
vault VIP.
Execution of every ruling remains LOGGED-NOT-EXECUTED. Merged under operator
direction 2026-07-27 ("Merge to main, then start Phase 0").
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Session close: GA-R4 bookend, skill sweep, ledger rotation, permission block removed
...
CLOSE SET, all executed and verified:
BOOKEND (GA-R4). The STAGE-5 GROUNDING AUDIT close bookend had already landed
while the operator was away, by design. They then returned and worked the queued
rulings, so this session's second entry is recorded as a POST-CLOSE ADDENDUM
rather than a second close -- the 2026-07-25/07-26 precedent, which states the
convention in terms: "The close bookend landed early by design; this addendum
records the work that followed it rather than re-opening the entry." Catching
that avoided leaving TWO close entries for one session. Addendum is 12 bullets
against the 15 cap.
LEDGER ROTATION (GA-R4 rule 3 / F1). The addendum took the live ledger to 321
lines against the 300 cap. One oldest-first pass left it at 307, still over, so a
SECOND pass was taken in the same close: both 2026-07-25 summaries (handoff-pack
execution, and MAAS admin-account recovery) moved VERBATIM to
docs/archive/session-ledger-rotated-20260727.md. Live ledger now 292 lines, five
entries, from the 2026-07-26 D-137 addendum onward.
SKILL SWEEP, done although NO STAGE CLOSED, on operator direction. Three new
invariants folded into SKILL.md, each earned this session:
- A FINDING IS AN OBSERVATION, NOT A CONCLUSION -- measure before putting it to
the operator. Four of the first six questions changed shape on measurement and
three were withdrawn or re-scoped.
- A CITATION IS AN EXISTENCE CLAIM; ONLY ITS CONTENT IS EVIDENCE -- both bugs
behind the v4-only lb-mgmt draft failed on reading and the conclusion inverted.
No gate reads prose, so nothing in a repo can catch this.
- RULED IS NOT BUILT -- D-134's bands and D-020's vault VIP were both properly
ruled and never implemented, and passed every gate for weeks.
Into references/script-authoring.md: the repo-lint L5 heading trap (a heading
LEADING with a D-number needs AMENDMENT or RESOLVED on the same line -- it bit
this session twice and was on no author-facing surface), the commit-gated-on-lint
pattern that then caught its second occurrence, and the space-aligned .tsv.
Into references/operating-discipline.md: the probe PATH determines the answer --
the office1 VMs read "unreachable" via voffice1 and NO-CLONE from vcloud, and only
the second is a finding.
Snapshot regenerated as openstack-cloud-ops-consolidated-20260727.md: 1695 lines,
ASCII/LF byte-verified, all five new items present. A stale snapshot is the
recorded failure mode of docs/audit/skill-divergence-20260725.md.
CHANGELOG extended with the full ruling table (14 ruled, 2 withdrawn, G18 opened)
and the sweep/skill sections, each carrying its revert.
PERMISSION BLOCK REMOVED as promised at grant time: allow 303 -> 248, ask 24 -> 7.
All 55 session-scoped allow rules deleted and the five original broad
`ssh voffice1 "maas admin <noun> *` ask rules restored in place of the 23
verb-scoped ones. JSON validated. The file is gitignored, so the scope doc's
section 6 and this changelog are its only durable record.
FINAL VERIFICATION: repo-lint 0 fail / 1 warn (the L1 legacy carve-out, expected);
gauntlet ALL GREEN (81 harnesses); ledger-scan reconciles at 21 open SEC rows and
next-free D 138 / DOCFIX 205 / BUNDLEFIX 053 -- UNCHANGED, correctly, because
fourteen rulings were recorded and nothing was remediated.
Revert: git revert this commit; the skill edits are additive, the snapshot is
derived, and the ledger rotation is reversible from the archive file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Session-close sweep: three stale surfaces the audit itself created, plus transcript-only facts
...
Operator direction before the bookend: "complete a session sweep for anything
missed that should be committed that would be lost on session end". Precedent:
queued-findings-20260726.txt and -20260727.txt. Capture:
docs/audit/queued-findings-20260727-stage5-audit.txt.
THE SWEEP FOUND THREE STALE SURFACES THIS AUDIT ITSELF CREATED, which is the
exact defect class it was convened to find:
1. docs/audit/stage5-readiness-20260727.md -- the "read this first" artifact --
still carried THIRTEEN NEEDS-RULING markers after every one of R1-R15 had been
ruled or withdrawn. At that moment it was the single most misleading surface in
the repo. Fixed with a supersession banner that names the three authorities in
order and states what in the document is STILL true and worth reading: the
ordered precondition sequence, the evidence column, the what-breaks column. The
row-level markers were deliberately NOT rewritten -- they are the record of what
was owed at the time, and rewriting them would destroy that record.
2. queued-rulings-20260727.md still opened with "Nothing here is adopted" after
fourteen adoptions. Fixed with a STATUS block naming design-decisions.md and
CURRENT-STATE.md as the ruling authorities rather than that file; the original
sentence is struck through, not deleted.
3. SIX sections of that file had BLANK OPERATOR UTTERANCE lines -- R8, R9, R10,
R12, R13, R15 -- for rulings that WERE properly recorded in design-decisions
and CURRENT-STATE. GA-R5 was satisfied; the question sheet was not, and a
future session reading only that file would have believed six questions were
still open. All six filled with the exact utterance and where the ruling lives.
R14-ORIGINAL's blank line is CORRECT and left alone -- it is superseded text
kept for the audit trail.
TRANSCRIPT-ONLY OPERATIONAL FACTS NOW CAPTURED: the repo-lint L5 heading trap (a
heading LEADING with a D-number counts as a second definition unless the line also
contains AMENDMENT or RESOLVED -- it bit this session twice and is on no
author-facing surface); the commit-gated-on-lint discipline adopted after a red
commit was pushed, which then caught the second occurrence before it landed; the
read-only live-apex poll procedure that keeps the token in env and prints nothing;
that the office1 VMs answer from vcloud but NOT via voffice1, so "unreachable" from
the wrong host must never be recorded as "absent"; and that creds-matrix.tsv is
SPACE-aligned despite the .tsv extension, so awk -F'\t' returns zero rows and looks
like a real result.
META-FINDING RECORDED FOR THE NEXT COMMITTEE: the lenses' OBSERVATIONS were
reliable, their CONCLUSIONS repeatedly were not. Three of this audit's own premises
needed correcting by measurement -- R2a (literals already assigned under D-111), R9
(one failure mode not two, two scripts not twenty), R14 (the register already
attributed and explained the findings, and already warned against the suppression
the question contemplated) -- and R8's entire evidence base collapsed on reading
the cited bugs. A lens finding is an OBSERVATION, not a conclusion; measure before
putting it to the operator as a question.
VERIFIED DURING THE SWEEP, after an API disconnect mid-turn: nothing was lost or
corrupted. The earlier lens-5 API casualty was fully recovered -- its 18 findings
and section header are in the committed record, and all seven lenses are present.
The disconnect this turn fell cleanly between edits and a commit; the two modified
files parsed intact and lint was red only for the L10 this commit satisfies.
NOT DONE AND NOT OWED YET: the skill sweep. This session closed no STAGE, so the
stage-close fold-in and snapshot regeneration are not due. B1/C1/C2 in the capture
are the candidates when Stage 5 closes.
Revert: git revert this commit; the capture is new and the two repairs are
additive (banner, status block, filled utterance lines).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R15 RULED: all three gate floors, with a harness manifest rather than a count (OPS)
...
GA-R5: question and exact utterance quoted, dated, pushed. Operator utterance:
"All three, with a harness MANIFEST rather than a count (Recommended)". OPS under
GA-R3 -- three script fixes, no D-number.
IT OPERATIONALISES GA-R6. That ruling lets a stage close only on a NAMED
EXECUTABLE CHECK. These three checks were not capable of failing, so the closes
they certified were weaker than they read.
1. repo-lint gains a valid-root check and a files-scanned floor.
repo_lint.py:132-135 strips only the two KNOWN flags, so any other --flag
becomes argv[0] i.e. the ROOT, with no is_dir() check and no floor. A
one-character typo of either the flag or the path resolves to a nonexistent
directory, rglob yields nothing, and it reports PASS (0 fail, 0 warn) over
ZERO files -- reproduced twice this session. This is the gate whose "0-fail"
every GA-R6 stage close in this project's history cites, and it cannot
distinguish a clean repo from an unexamined one.
2. The gauntlet pins a checked-in MANIFEST of harness NAMES, drift-checked, not
a count. run-tests-all.sh:35 already has a zero-floor; what is absent is any
pin on WHICH harnesses ran, and 81 exists only as prose. The manifest choice
is the substance of this ruling: a bare count does NOT catch a rename, since
adding one harness while removing another holds the count. It mirrors two
proven in-repo patterns -- clientdocs/sweep-receipt.txt hash-pinning and the
D-137 creds-manifests derive-plus-drift gate.
3. preflight stops failing open. note() tests only rc -eq 1 and rc -eq 2, so
127/126/130 and any rc>=3 leave "PREFLIGHT: PASS -- clear to add-model /
deploy" -- measured with a sub-gate exiting 127. And rc=2, which is precisely
how these checkers signal "I could not evaluate anything", is remapped to WARN
in P1/P2/P3. The project already fixed exactly this for P5 (harness T9), so
this extends a proven pattern rather than inventing one. Preflight authorises
the deploy, which is why it was included rather than deferred.
ORDERING TRAP RECORDED FOR WHOEVER EXECUTES THIS: these changes alter what
"green" MEANS, so the gauntlet and lint runs that certify them must be read with
that in mind -- and a harness manifest introduced mid-session must be seeded from
a tree that is itself verified, not from whatever happens to be on disk. Seeding
it from an unverified tree would pin the very drift it exists to detect.
ALL QUEUED RULINGS ARE NOW CLOSED: R1-R13 and R15 ruled, R2a and R14 withdrawn as
raised in error, G18 opened deferred-and-gated.
Revert: git revert this commit; the CURRENT-STATE entry is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R13 RULED: register-first credential fixes, no new tool (D-137 sub-ruling 6)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Register-first, no new tool: fix the staging AND flip the 7
keypairs (Recommended)".
Measuring showed R13 was THREE problems, not the two the queued framing carried.
The third is the sharpest and had not been surfaced as a decision at all:
THE REGISTER MIS-STAGES WHAT STAGE 5 ACTUALLY MINTS. The vault-init, Octavia-PKI
and admin-openrc rows are cardinality=singleton with mint-stage in
vr0-phase01/02/03, and creds-manifests/stages-reached marks all three pending --
so flipping stage5 to reached never makes them expected. The measured consequence
is a FALSE GREEN over the largest minting event of the deployment: P5 reports
"[ok] E1 18 expected artifact(s) deferred as not-yet-minted". R7 made this MORE
urgent rather than less: per-DC independent Octavia PKI means two CA mints where
the register expects none.
The retroactive half needs no new tooling, which is what made register-first the
efficient answer: S4 ALREADY resolves runbook:<path>:<line> references, so
recording each mint as a numbered runbook step and flipping mint-ref from
operator-terminal converges the debt using machinery that exists and is tested.
30 rows / 16 ids carry that ref today and grep -rnI "ssh-keygen" returns ZERO
hits repo-wide.
Priority within that half is the SEVEN unrecoverable-in-place keypairs, where
irreproducibility is an operational risk rather than hygiene: both edge keys
(SEC-007/-015 make edge SSH the ONLY management path to the DC edges), both svc
keys, both power keys, and office1-svc-key. A jumphost rebuild today locks the
operator out of both DC edges.
creds-mint.sh STAYS QUEUED for its own ruling and is deliberately not built here.
It is orthogonal -- it would prevent the NEXT unregistered mint but makes no
existing key reproducible and fixes no staging. Bundling it would have made the
urgent no-tool work wait on unscoped tooling work.
Carried caveat for whoever executes this: the T24 finding-class baseline covers
TIER 1 ONLY, so a tier-2/3 regression introduced by these matrix edits would not
turn the gauntlet red on its own.
PROCESS NOTE, owned: the first attempt at this commit tripped repo-lint L5 by
titling the sub-ruling "### D-137 SUB-RULING 6" -- a heading LEADING with a
D-number counts as a second definition unless it also contains AMENDMENT or
RESOLVED. That is the same rule I hit on R5 and then documented. The lint gating
caught it before the commit landed this time, so nothing was pushed red; the
heading is now "### SUB-RULING 6 (D-137)", matching the convention.
Revert: git revert this commit; the sub-ruling is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R12 RULED and EXECUTED: G17 reshaped -- time check folded in, artifact check made falsifiable
...
GA-R5: question and exact utterance quoted, dated, pushed. Operator utterance:
"Fold time verification into G17 and fix the check to assert content
(Recommended)".
Unlike the other rulings this one is EXECUTED in the same commit, because what it
rules on IS a gate row, and gate rows live in CURRENT-STATE section 6 by GA-R6.
Recording the ruling without reshaping the row would have left the defect in place.
TWO DEFECTS IN ONE ROW, both now fixed:
1. THE CHECK COULD NOT FAIL. The prior text specified `curl -sI http://10.12.8.4/`.
Measured: curl -sI exits 0 on 404/403/500 -- a planted 404 printed "404 File not
found" with curl exit 0, while curl -fsI exited 22 -- and the dc0 URL is an nginx
autoindex root created EMPTY by dc-mirror.sh:330 before any sync, so / answers
200 whether or not last-sync.status says OK. It reintroduced the
existence-not-content class that dc-mirror.sh check was fixed for on the SAME
DAY. The dc1 half named no command at all, although the real one already existed
at dc-cache-proxy.sh:210-217. Now: a real package-path fetch with an exit-code
predicate for dc0, the existing named check for dc1, and an explicit refusal on
unrecognised or unreachable results.
2. THE SCOPE WAS WRONG. docs/dc-dc-deployment-workflow.md:206 and
runbooks/dc-dc-phase4-juju-bundle-per-dc.md:46 BOTH assign the node time source
to G17, while chronyc appeared ZERO times in CURRENT-STATE -- so the check two
surfaces required had no home in the gate meant to carry it. Now folded in.
Worth stating precisely because it was nearly misread: DOCFIX-204 struck "NTP
from the DC's own OPNsense edge" because D-129(iv) gave the edge no NTP role. It
did NOT strike time verification. Option (c) -- declaring it struck and deleting
the two conflicting surfaces -- would have discarded a check a recorded decision
deliberately kept.
Fixed in ONE edit rather than two because both defects live in the same row and
share the same ONE-TIME observation window at first boot; splitting them risked one
landing without the other, and the window does not come back.
chronyc mentions in CURRENT-STATE: 0 -> 3.
Revert: git revert this commit to restore the prior G17 [V] text.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R14 WITHDRAWN: the register already does everything the question assumed it could not
...
Withdrawn without a ruling, following the R2a precedent. R14 asked whether the
credential matrix needs a ruled-exception field, on the premise that three S5
power-key asymmetries were "ruled correct by SEC-016" and so reported a permanent
red a reader learns to ignore.
All three parts of that premise fail on measurement:
1. They are NOT ruled correct. SEC-016 ruled per-DC ISOLATION -- dc1 gets its own
dedicated power key rather than reusing dc0's -- which is satisfied. It never
blessed the filename, host and custody divergence. That is SEC-021(b), an OPEN
ledger defect whose own disposition reads "needs a naming/custody reconciliation
to the dc1 shape", and whose stated complaint was "Per-DC rows that should be
symmetric are not, and nothing compares them". S5 is the thing that now compares
them, so the finding is the register working as designed.
2. The register already ATTRIBUTES them: those rows carry sec-ref=SEC-021 and
notes-ref=n-dc0-power-key-divergence.
3. The register already EXPLAINS them: creds-matrix-notes.md carries the
n-dc0-power-key-divergence note describing the name and host-role divergence and
why the expected jumphost rows fail EXPECTED-BUT-ABSENT.
And the notes file already warns against the exact move this question contemplated.
The line immediately preceding that note reads: "failure, not a matrix error: do
not delete the row to make the checker green."
So no schema change is needed and none should be made. Adding a suppression
mechanism would have HIDDEN an open security-ledger item -- the opposite of what
the register exists for. The red clears when SEC-021(b) is remediated, which is the
intended behaviour.
This is the third premise this audit has had to correct in its own framing, after
R2a (literals already assigned under D-111) and R9 (one failure mode, not two).
Worth noting the pattern: the audit's lens findings were sound as observations and
repeatedly wrong as conclusions, and measurement caught it every time.
Residual recorded so it is not mistaken for an oversight: whether a ruled-exception
field is EVER needed is now hypothetical. If one arises, notes-ref and sec-ref are
the place to start, not a new column.
Revert: git revert this commit; the R14 entry is rewritten in place with the
original retained beneath as R14-ORIGINAL for the audit trail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R10 RULED: fix the stale clone and make all working directories current (OPS)
...
GA-R5: operator direction recorded verbatim -- "R10, we need to fix the stale
commit issue and make sure that all working directories are current." OPS under
GA-R3: an environment fix, no D-number assigned.
This resolves R10 by REMOVING the reds rather than recording an exception basis
for proceeding against them. The measurement showed that is the stronger answer:
the red set is substantially an artifact of running the gate on the wrong host,
not a set of conditions Stage 5 has to accept.
SCOPE ENUMERATED BY MEASUREMENT rather than assumed, because "all working
directories" could have been broader than the one clone already known:
- voffice1:~/openstack-caracal-dc-dc -- 105 behind origin/main on
dc-dc-g12-dc1-substrate, a branch DELETED upstream. THE ONLY STALE CLONE.
- vcloud:~/openstack-caracal-dc-dc -- current, on the audit branch, in sync.
- dc0 rack, dc1 rack -- NO CLONE. By design: the dc-* scripts are piped in over
ssh with 'sudo bash -s', never cloned to the rack.
- office1-netbox, office1-tailscale -- NO CLONE. Measured FROM VCLOUD after a
probe via voffice1 returned "unreachable"; unreachable is not absent, so it
was re-run from a host that can actually see them rather than recorded as a
negative finding.
- ~/ops-toolkit on vcloud is a DIFFERENT repository
(git.baldurkeep.com/git/ops/ops-toolkit.git, single commit) and is out of scope.
RECOVERY SHAPE MATTERS AND IS NOT THE OBVIOUS ONE: a plain `git pull` on voffice1
does NOT work, because its tracked branch no longer exists upstream. It needs
`git fetch origin && git switch main` with a HEAD-equals-origin/main assertion
afterward (finding L5-4). The untracked
opentofu/vr1-dc1-substrate/.terraform.lock.hcl there is not tracked on main, so
the checkout will not conflict.
CONSEQUENCES ONCE CURRENT: preflight becomes runnable on voffice1, which clears
P3's 33 warns and P4's "MAAS unreachable" -- both missing-binary artifacts, since
maas and juju are PRESENT on voffice1 and ABSENT on vcloud. The octavia-pki
absence and the 7 credential findings are host-independent and remain.
THEN OWED: repo-lint and the gauntlet ON voffice1. These were deliberately
deferred throughout this audit precisely because running them against a
105-commit-stale tree would have produced a number that meant nothing. The
deferral now discharges.
ALL STAGE-5 BLOCKING RULINGS (R1-R11) ARE NOW CLOSED. Remaining: R12-R15
standing, plus G18 deferred-and-gated.
Revert: git revert this commit; the CURRENT-STATE entry is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R9 RULED: derive lib-net's dc1 arm from the overlay with a drift check (D-119 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Derive lib-net's dc1 arm from the overlay, with a drift check
(Recommended)".
Recorded as a D-119 AMENDMENT -- D-119 owns the region-qualified per-DC arms in
lib-net.sh, so the question of how the dc1 arm is populated belongs there.
RULED SHAPE: lib-net.sh's vr1-dc1 arm is GENERATED from overlays/vr1-dc1-vips.yaml
with a render-drift check that fails the gauntlet on divergence. Consumers keep
sourcing lib-net unchanged -- no script learns to parse YAML -- and exactly one
authored copy of the values exists.
WHY THIS IS NOT NEW MACHINERY: D-137 sub-ruling 2 already ruled and built the same
shape for credentials, with creds-manifests DERIVED from creds-matrix.tsv and the
gauntlet failing on rendered-vs-checked-in drift. This applies a proven in-repo
pattern to network literals rather than inventing one.
IT DOES NOT PRE-EMPT D-136, and that was the deciding property. D-136 is PROPOSED
and unruled. This is the smaller compatible step: the overlay becomes the single
authored source for script-facing values. If D-136 is later adopted the apex
becomes the source and BOTH the overlay and this derived arm become generated --
a sub-case, not a competitor.
MEASURED BASIS, which inverts the intuitive reading: lib-net.sh:76-79 states the
backward-compatibility design in terms -- sourcing without the selector keeps
VR0/DC0 values "completely unchanged". So the unset block fires ONLY for
selector-callers. The 8 correctly-updated scripts break LOUDLY under set -u; the
20 never updated fail SILENTLY on dc0's literals. The silent half is the
dangerous half.
GUARD CARRIED FROM R11, and it matters for the generator's implementation: the
dc1 arm unsets NINE variables for TWO different reasons. METAL_INTERNAL_VID and
METAL_INTERNAL_IFACE are CORRECTLY unset because D-133 abolished the VLAN-103 /
br-internal stack for VR1 -- those facts genuinely do not exist. Only the
VIP/FIP/keystone group is superseded and derivable. A generator that populated
all nine would silently reintroduce a stack D-133 retired.
COUPLED: R2 makes the derivation dual-family, not v4-only. R11 moves
VIP_COUNT_EXPECT 11 -> 13 and widens the octet band to .99, and the two band
constants (OCTET_LO/HI in provider-bundle-check.py, VIP_OCTET_MAX in lib-net.sh)
are separately named in two files.
STILL OWED AND EXPLICITLY NOT COVERED: the consumer sweep. This ruling decides
where values live; it does not make the 20 non-selector scripts call the
selector. Stage-5 exposure is two of them (phase-03-core-verify.sh,
deploy-watch.sh); carve-host-interfaces.sh:48-67 (DOCFIX-166) is the in-repo
precedent for wiring one correctly.
Execution is a separate gated step under standard delivery discipline.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Final pre-ruling measurements R9/R10/R12-R15: two premises corrected
...
Operator direction: gather everything remaining so the last decisions can be
worked through without stopping to measure. Capture:
docs/audit/r9-r15-final-measurements-20260727.txt.
TWO OF THESE CHANGED MATERIALLY, and both are corrections to the audit's OWN
earlier framing rather than new findings.
R9 -- there is only ONE failure mode, not two, and the Stage-5 blast radius is
TWO scripts rather than twenty. lib-net.sh:76-79 states the design explicitly:
sourcing without the selector "continues to populate PLANE_CIDRS ... exactly as
above (VR0/DC0's real, measured values) -- every existing script that sources
lib-net.sh keeps working completely unchanged". So the unset block fires ONLY
for selector-CALLERS. The 20 non-selector consumers fail SILENTLY with dc0
literals; the 8 correctly-updated ones are the ones that break loudly under
set -u. That inverts the intuitive reading -- the scripts nobody fixed are the
dangerous ones. And grepping the two runbooks Stage 5 actually uses, only
phase-03-core-verify.sh and deploy-watch.sh are invoked there; the rest bite at
later stages.
R14 -- THE PREMISE IS WRONG and the question largely dissolves. I framed it as
"the matrix cannot express a RULED exception, so three S5 asymmetries ruled
correct by SEC-016 report as a permanent red". Measured: they are NOT ruled
correct. SEC-016 ruled per-DC ISOLATION -- dc1 gets its own dedicated power key
rather than reusing dc0's -- which is satisfied. It never blessed the filename,
host and custody divergence. THAT is SEC-021(b), an OPEN defect whose own
recorded disposition reads "needs a naming/custody reconciliation to the dc1
shape", and whose stated complaint was "Per-DC rows that should be symmetric are
not, and nothing compares them". S5 is the thing that now compares them. The
register is RIGHT and the finding is REAL. Suppressing it would have hidden an
open security-ledger item.
R15 refinement: the gauntlet ALREADY has a zero-floor -- run-tests-all.sh:35
exits 2 when RAN is 0. What is missing is a MINIMUM-count floor, since 81 is
pinned nowhere executable. repo_lint.py:132-135 has NEITHER a valid-root check
nor a files-scanned floor. The two gates need different fixes, which the earlier
framing conflated.
R12 and R10 confirmed as previously stated, with exact line citations.
Revert: git revert this commit; the capture is new and CURRENT-STATE additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

G18 OPENED: IPAM apex completeness for the Octavia lb-mgmt plane (blocking, deferred)
...
Operator direction, verbatim: "leave this as an open decision that will need a
ruling once we have the cloud live and we have a better read on the network and
how everything is functioning with the addition of the new IPv6 configurations.
Make this a gated decision so we cannot close the project (or whatever phase you
think it best ruled in) without a ruling on this item."
WHAT IT IS. R8 ruled that Octavia creates and owns its own IPv6 lb-mgmt network.
The measured consequence is that the octavia charm exposes NO CIDR, address-family
or router configuration option -- create-mgmt-network (default True) is the only
related option -- so the prefix is CHARM-GENERATED and cannot come from the D-111
carve. NetBox is therefore knowingly incomplete for exactly one plane. That is the
authority-inversion concern the audit's lens 7 raised: the apex being back-filled
to match a deploy rather than driving it.
WHY DEFERRED RATHER THAN RULED NOW. It is not answerable from artifacts -- it
needs the cloud live and an observed read on IPv6 behaviour. It is also adjacent
to the UNRULED D-136 NetBox-coupled render pipeline, which covers the same
apex-authority ground, so ruling it early would pre-empt that decision.
PLACEMENT. Recorded as gate G18 in CURRENT-STATE section 6, the gate authority.
[R] ruling-type: closes ONLY on a GA-R5 recorded ruling with the exact utterance,
dated, committed and pushed. ANSWERABLE from Stage 5 onward, since the prefix
exists once Octavia deploys. BLOCKING at the FINAL stage close / project close --
the deployment may not be declared complete while it is open. That placement gives
the widest window in which live evidence can accumulate while still guaranteeing
the question cannot be lost.
Options are pre-recorded in the gate row so a future session does not have to
re-derive them: (a) back-fill the charm-created prefix into NetBox post-deploy as
a documented record; (b) record the plane as charm-owned and explicitly out of
apex scope; (c) fold it into D-136's render-pipeline ruling if that is taken first.
ALSO CROSS-RECORDED, because it is exactly the kind of thing that gets "fixed" by
mistake: the absence of an lb-mgmt :x80 prefix in the VR1 ULA carve is CORRECT
under R8, not a gap. Noted in the D-101 R8 ruling note and in the G18 row.
Mirrored into docs/audit/queued-rulings-20260727.md as R16 (gated, deferred) with
CURRENT-STATE named as the authority, so the queue and the gate table agree.
Revert: git revert this commit; the gate row and cross-references are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R8a RULED: extend the o-hm0 verifier to compare MTUs (OPS, no D-number)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Extend the existing o-hm0 verifier to compare MTUs
(Recommended)".
Recorded as a SUB-RULING on the D-101 MTU ruling note rather than a new decision.
OPS under GA-R3 -- it is a script change, doubt resolves DOWN, so no D-number is
assigned. It discharges the LP #2018998 obligation that R8 carried forward.
MEASURED BEFORE PRESENTING, and it changed the option set. "Pre-emptive
configuration" was not actually available: the octavia charm exposes NO MTU
option, so nothing can be set in the bundle. The real choice was only how the
mismatch gets DISCOVERED.
The measurement also found a better option than either I had sketched:
scripts/phase-05-octavia-verify.sh:104 ALREADY inspects o-hm0, asserting
state != DOWN and the presence of an fc00::/ ULA. So it holds the handle already
and an MTU comparison is a natural extension rather than new machinery. Against
that, grep -i mtu across cloud-assert.sh, phase-05-octavia-verify.sh and
phase-04-network-verify.sh returns NOTHING -- MTU is asserted nowhere in any
post-deploy verifier, which is why this class would have surfaced as instability
rather than as a failed gate.
Rely-on-the-charm was refused: the 2025-12-31 recurrence against octavia 14.0.0 /
2024.1 stable is evidence the update-status self-heal is not reliable, and the
symptom -- spurious failovers under load -- reads as a Ceph, network or amphora
fault long before anyone suspects MTU. The unconditional ovs-vsctl mtu_request
workaround was refused: it fights the charm's own fix where that fix is working,
and would MASK a future regression instead of surfacing it. It stays available
as the remedy IF the check fires.
INCIDENTAL CORROBORATION OF R8, from this repo's own tooling: that o-hm0 verifier
is VR0-era and expects an fc00::/ ULA, which independently confirms the octavia
charm's default lb-mgmt-subnet is an IPv6 ULA. The as-built already assumed what
R8 ruled -- a second, internal line of evidence for a ruling whose external
evidence had to be rebuilt from scratch.
Delivery when executed: the assertion ships with its harness green, gauntlet ALL
GREEN, repo-lint 0-fail and a changelog entry carrying a revert.
Revert: git revert this commit; the sub-ruling is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R8 RULED: Octavia creates and owns its own IPv6 network (D-101 note; sub-ruling closed)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "I want to take a Octavia creates and owns its own IPv6
network." Supporting reasoning recorded because it is load-bearing: "We have to
research to troubleshoot if we run into MTU bug issues down the road. We have
already deployed using this topology in the v1 DC test deployments that got us to
this point."
This CLOSES the octavia-family sub-ruling that the 2026-07-25 D-101 note left
explicitly open.
THE RULING REQUIRES NO ARTIFACT CHANGE, and the operator's reasoning checks out on
the artifacts: create-mgmt-network is set NOWHERE in bundle.yaml or any overlay, so
the charm default True has always applied, and Octavia's whole options block is
debug, openstack-origin, amp-image-tag, vip. VR0 deployed Octavia on precisely this
shape. This ratifies the existing topology rather than changing it.
FAMILY AGREES THREE WAYS. The charm's default lb-mgmt-subnet is IPv6 (LP #1897418,
verbatim: "By default, Octavia charm uses ipv6 for its lb-mgmt-subnet") and it is a
ULA -- the amphora address quoted in LP #1911788 is fc00:fa21:3d5c:9cfd:..., i.e.
fc00::/7. That is exactly D-101's "IPv6-only ULA ... Octavia lb-mgmt ... Internal,
no external clients".
THE APPARENT VR0 GUA CONTRADICTION WAS A PAPER ALLOCATION, checked on operator
instruction: lib-net.sh gives VR0 six IPv4 planes and no lbaas plane; lbaas is
listed in STALE_SPACES; maas-as-built-reference records that NIC as "idle
(undefined; ex-lbaas), raw NIC, no link"; and D-101's own context says VR0 is
IPv4-only. The whole VR0 v6 tree in the apex is a design record.
CONSEQUENCE FOR D-111, worth recording so it is not later "fixed": the absence of
an lb-mgmt :x80 prefix in the VR1 ULA carve is CORRECT, not a gap. The charm
exposes no CIDR option, so the prefix is charm-generated and cannot come from the
apex.
STANDING OBLIGATION carried by the operator's own reasoning: LP #2018998 (o-hm0 vs
lb-mgmt-net MTU, charm-octavia, High) is Fix Released across our lineage but
recurred 2025-12-31 against octavia 14.0.0 / 2024.1 stable -- our exact pin. A
jumbo lb-mgmt-net beside a 1500 o-hm0 silently drops health messages over 1500
bytes and causes spurious failovers, which is the direct interaction with R3's
jumbo-underlay ruling. Owed at the Octavia step: verify o-hm0's MTU MATCHES
lb-mgmt-net's by measurement, not by trusting the charm.
NOT ASSERTED: whether the charm attaches an external gateway to the Neutron router
it creates. No config surface tells it to and ULA is not globally routable, but
isolation was not proven from the documentation. Settle by inspecting the router at
deploy time; queued as an observation, not a blocker.
Revert: git revert this commit; the ruling note and research sections are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R8 research: both in-repo citations fail; charm default is IPv6; new MTU risk found
...
Operator declined to rule on the overlay's own citations -- "I would like
documentation and vendor specific information rather than conjecture. Research is
cheap, guesswork is expensive" -- then directed a full read of two bugs weighed
against current versions. That was the right call and it inverted the answer.
BOTH CITATIONS IN overlays/dc-dc-ipv6-family-matrix.yaml FAIL ON INSPECTION:
- LP #1913409 is Fix Released (2021) against kolla-ansible, a DIFFERENT
INSTALLER. No bearing on a Juju/charm deployment.
- LP #1911788 is Incomplete and a DUPLICATE of LP #1896630, and is NOT an IPv6
defect. Its diagnosed cause is an OVN port-binding hostname mismatch --
binding_host_id carrying the shortname instead of the FQDN for LXD containers
on MAAS, giving binding_failed instead of ovs. The IPv6 address in the timeout
was incidental; the port never bound, so nothing would have worked over any
family.
VERSION WEIGHING, verified rather than asserted. #1896630's primary fix is
ovs-record-hostname.service, which the bug places in OVS 2.15 (Focal shipped
2.13). Rather than assume, I queried THIS DEPLOYMENT'S OWN dc0 mirror for what a
jammy node would actually install: openvswitch-switch 2.17.0-0ubuntu1 (jammy) and
2.17.9-0ubuntu0.22.04.2 (jammy-updates). 2.17.9 >> 2.15, so the fix is present on
every node this deployment will build. The bug's era was Juju 2.7.8, charms 20.08
Ussuri, Focal/OVS 2.13; we are Juju 3.6.27, charms 2024.1, jammy/OVS 2.17.9. The
charm-layer-ovn task -- the one closest to this repo's OVN usage -- is closed
Invalid.
AND THE DEFAULT RUNS THE OTHER WAY: LP #1897418 records, verbatim, "By default,
Octavia charm uses ipv6 for its lb-mgmt-subnet", with a reporter noting success
"using the charm default (ipv6) network but using ipv4 everywhere else". Upstream
Octavia documents that "IPv6 subnets can be used for the LB Network". So D-101's
IPv6-only placement of lb-mgmt is ALIGNED with the charm default, and a v4-only
lb-mgmt-net would be the DEPARTURE requiring justification -- the opposite of how
R8 was first framed.
NEW LIVE RISK FOUND, COUPLED TO THE R3 RULING: LP #2018998 "MTU mismatch between
o-hm0 and lb-mgmt-net" (charm-octavia, High, Edward Hope-Morley). A jumbo
lb-mgmt-net beside a 1500 o-hm0 silently drops health messages over 1500 bytes
and triggers SPURIOUS load-balancer failovers -- evidenced in the bug as "UDP,
length 1534 > 1500 and o-hm0 never receives them" and "o-hm0: dropped over-mtu
packet: 1744 > 1500". Fix Released across our lineage, BUT a recurrence was
reported 2025-12-31 against octavia 14.0.0 / 2024.1 stable, the exact channel
bundle.yaml pins. R3 ruled the underlay be finished to jumbo, which is precisely
this bug's precondition. OWED at the Octavia step of Stage 5: verify o-hm0's MTU
MATCHES lb-mgmt-net's after deploy, not merely that the charm claims to set it.
GENERALISABLE LESSON, recorded in the capture: a bug NUMBER in a comment is an
existence claim; what the bug SAYS is the content, and only content is evidence.
Same class as the two ruled-but-never-built decisions this audit already found --
a plausible in-repo statement that nothing verifies.
R8 is NOT ruled by this commit. It will be re-presented on corrected evidence.
Revert: git revert this commit; the capture is new and CURRENT-STATE additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R7 RULED: per-DC independent Octavia PKI (D-109 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Per-DC independent Octavia PKI; fix the generator first
(Recommended)".
Recorded as a D-109 AMENDMENT. D-109 established per-DC INDEPENDENT Vault roots
with a DR-honest rationale but never mentioned the Octavia amphora
control-plane PKI, which is a SEPARATE trust domain with its own CA generated
outside Vault by phase-01 step 1.0-GEN. This extends the same posture to it.
MEASURED REFINEMENT THAT NARROWS THE WORK. The generator is dc0-frozen in two
ways of DIFFERENT severity, and only one is a design problem:
- The CA SUBJECT is a baked literal, "/CN=VR0 DC0 Omega Cloud Octavia Controller
CA/O=Neumatrix". Reusing it roots dc1's amphora chain in a VR0-DC0-named CA.
- The controller cert's SAN is ALREADY DERIVED per-DC by design -- the step
states it carries the controller FQDN, octavia API FQDN and the Octavia API VIP
"derived from the bundle at generation time -- DOCFIX-067; never a baked
literal". That half needs no change at all.
- The actual blocker is the VIP gate, grep -qE '^10\.12\.4\.', which accepts only
dc0's band while dc1's octavia VIP is 10.12.64.57 and lives in an overlay.
A DOCFIX was owed regardless of the ruling: the generator cannot produce a dc1
artifact as written, and phase-01:144-145 hard-ABORTS the deploy when
overlays/octavia-pki.yaml is absent, which it is. Option (c) would therefore have
deferred only the part D-109 already answers by precedent while leaving the
required work untouched.
REUSE REFUSED ON POSTURE. The overlay carries CA private keys plus a plaintext
issuing-CA passphrase inside the repo clone, and SEC-004 records the repo as
still PUBLIC. Sharing one amphora control-plane CA private key across two clouds
D-100 defines as independent would widen an existing exposure rather than
contain it, in a commercial multi-tenant cloud with hard tenant isolation.
Independent per-DC CAs keep that blast radius to one DC.
CAVEAT CARRIED FORWARD, because it will bite otherwise: the generator's VIP gate
must be re-pointed at the MERGED deploy input rather than bundle.yaml. Left
reading the base bundle it breaks again the moment ruling-3's VIP extraction and
R11's new .61/.62 allocations land.
Roosevelt analog: per-DC amphora CAs match per-DC IPMI/BMC credentials and the
per-DC MAAS power keys of SEC-012/-016 -- the same no-cross-DC-shared-secret
principle already ruled twice in this deployment.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R11 RULED: vault .61 and designate .62, dual-family triples (D-020 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Both full triples (.61 vault, .62 designate), dual-family,
and fix the gate (Recommended)".
THE FINDING THAT MEASURING FIRST PRODUCED: vault was ALREADY RULED and never
built. D-020's decision text enumerates vault BY NAME among the clustered
applications carrying both a provider and a metal VIP; measured, base vault is
num_units:1 with an EMPTY options block. This is a conformance repair of a
2026-era decision, not a new choice -- and it is the SECOND ruled decision this
audit has found unimplemented, after D-134's address bands (R4). Worth stating
plainly: the pattern is ruled-but-never-built, and nothing in the repo could
detect either case.
designate is genuinely new -- absent from D-020's enumeration -- so it needed a
ruling rather than a repair, and this amendment adds it. Its dnsaas endpoint is
already bound provider-public, so a provider leg is coherent for it.
Shape is the ESTABLISHED triple, not a new form. The measured octet map is
consecutive: keystone .50 through ceph-radosgw .60, all provider/admin/internal
triples. vault takes .61, designate .62, both dual-family per R2 -- adding
v4-only now and re-doing them later would mean re-issuing certificate SANs on a
live cloud, and for vault that is 23 certificate relations.
Option (b), vault metal-only per the 2026-07-25 expansion review, was refused and
the conflict is RECORDED so that proposal is not later mistaken for the ruled
position: it has a real technical argument (all 23 vault consumers are internal)
but it contradicts D-020's own enumeration and would leave vault the single
non-triple in the bundle, a permanent special case for the checker.
Mechanical consequences, no choice in them: .61/.62 are legal under D-134's
amended .50-.99 band but REJECTED by the gate today -- provider-bundle-check.py
holds OCTET_LO/HI = 50,60 and lib-net.sh holds VIP_OCTET_MAX=60. These are
SEPARATELY NAMED constants in two files, so widening the band is a two-file
change (L3-7). VIP_COUNT_EXPECT moves 11 -> 13.
Gate hardening ruled IN SCOPE rather than deferred: the checker learns to FAIL on
an application with an hacluster relation and no vip. Justification is measured --
grep -rn cluster_count scripts/ tests/ returns NOTHING, and a constructed overlay
rewriting all 20 cluster_count values 3 -> 1 yields a byte-identical PASS (L4-3).
That is exactly why decorative HA was found by an audit instead of by a gate.
Per the R6 ruling, these VIPs land BEFORE dc-ha-scaleup.yaml is applied.
Execution is a separate gated step.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R6 RULED: close the two VIP gaps, then apply the HA overlay whole (D-121 note)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Close the two VIP gaps first, then apply the overlay whole
(Recommended)".
THE TENSION IT RESOLVES. D-121 is titled "VR1 makes HA real -- scale the
decorative single-unit control plane to 3". Deferring the overlay would have
deployed VR1 in exactly the shape D-121 was written to retire. Applying it as-is
would have shipped, for vault, precisely the defect D-121 exists to remove -- a
pacemaker cluster with nothing to manage. Only the ruled sequence satisfies the
decision rather than half of it.
MEASURED BEFORE PRESENTING, and it reframed the question. The overlay scales 14
applications and moves all 12 base hacluster subordinates to cluster_count 3, but
ELEVEN OF TWELVE principals already carry proper VIP triples. The gap is exactly
two applications, and they are not the same shape:
- designate is in the base bundle WITH an hacluster and an ha binding, no vip.
- vault has NO hacluster in base at all and an EMPTY options block. The overlay
INTRODUCES the vault-hacluster application, its cluster_count 3, and the
vault:ha relation -- and still no vip.
Vault is the deciding case because of blast radius: 23 relations consume
vault:certificates, plus barbican-vault:secrets-storage. With no VIP every
consumer binds a unit address and there is nothing for a failover to move.
Remediating that on a live 3-unit vault means re-pointing 23 certificate
relations and re-issuing SANs on a running cloud, which is why the faster option
was refused.
SEQUENCING CONSEQUENCE: R11 becomes a HARD Stage-5 precondition ordered BEFORE
the overlay, not a parallel item.
Scope is tractable and the pattern is already in-repo: octavia carries a correct
triple (10.12.4.57 10.12.8.57 10.12.12.57), so R11 has a shape to copy rather
than design.
NOT resolved here, tracked separately: the Stage-6 radosgw multisite path is
single-unit-shaped while this scales ceph-radosgw to 3; and
provider-bundle-check.py checks cluster_count NOWHERE, which is precisely why
decorative HA was found by an audit instead of by a gate.
Address demand is not an argument against this -- R4's band ruling already covers
the overlay's growth from 27 to ~55 LXD units.
Execution is a separate gated step, not authorised by this ruling.
Revert: git revert this commit; the ruling note is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Pre-ruling measurements for R6-R15: measure before asking, all ten
...
Operator: "Measure all the rest then we can work through them with relevant
data at hand." Taken directly by the session rather than delegated, after the
UNMEASURED-gap sweep showed lens findings need measuring before they become
questions. Capture: docs/audit/r6-r15-measurements-20260727.txt.
THE TWO RESULTS THAT CHANGE A QUESTION'S SHAPE:
R6/R11 -- vault's ENTIRE HA apparatus exists only in dc-ha-scaleup.yaml. Base
vault is num_units:1 with an EMPTY options block: no vip, no hacluster, no
relation. The overlay INTRODUCES the vault-hacluster application (not present
in base at all), sets cluster_count:3, adds the vault:ha relation -- and still
no vip. Applying the overlay therefore creates a 3-node pacemaker cluster with
nothing to manage. The blast radius is what makes it sharp: 23 relations
consume vault:certificates, plus barbican-vault:secrets-storage. Vault is the
CA for the whole cloud and every consumer binds a unit address. Of the 12 base
hacluster subordinates exactly ONE principal lacks a vip (designate), and
octavia carries a proper triple -- so this is a 2-app gap, not a pattern.
R10 -- the question largely DISSOLVES. Preflight has been run on the wrong
host. Measured: vcloud has neither maas nor juju nor openstack; voffice1 has
maas AND juju. So P3's 33 warns and P4's "MAAS unreachable" both clear simply
by running preflight on the D-128 Plane-2 host where it belongs. Only the
octavia-pki absence and the 7 credential findings are host-independent. Same
"tool absence reported as something else" class as U6, now three instances.
OTHERS: R7 the octavia CA subject is a baked VR0-DC0 literal while the SAN is
already derived by design (DOCFIX-067) -- the two halves differ in severity.
R8 the v4-only lb-mgmt shape is already pre-analysed in-repo with LP #1911788
and #1913409 cited and a drafted block ready (NOT verified upstream by me).
R9 only 8 of 28 lib-net consumers call the DC selector; the 20 that do not
include the entire phase-02..phase-06 family. R12 chronyc appears ZERO times
in CURRENT-STATE, confirming G17 omits the time check. R13 30 rows / 16 ids
are operator-terminal, of which SEVEN are keypairs whose loss is
unrecoverable-in-place. R14 the matrix has no exception field at all. R15 81
harnesses on disk and 81 reported, with 81 pinned nowhere executable.
METHOD NOTE, recorded because it nearly published a wrong figure: parsing
creds-matrix.tsv with awk -F'\t' returned ZERO operator-terminal rows,
contradicting lens 7's 30. The file is SPACE-ALIGNED, not tab-separated.
Re-measured correctly it is 30/16 and lens 7 was right. A disagreement with a
prior finding was treated as a reason to re-check my instrument rather than to
publish the new number.
Revert: git revert this commit; the capture is new and CURRENT-STATE additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Fix L5 collision introduced by the R5 ruling-note heading
...
repo-lint L5 failed on the previous commit and I pushed it anyway -- I ran lint
in the same compound command as the commit but did not GATE on its result.
Owned, and fixed forward rather than by rewriting pushed history.
Root cause, read from the rule rather than guessed: repo_lint.py:355 matches
`^##+\s+(D-\d{3,})\b(?!.*AMENDMENT|.*RESOLVED)` -- ANY heading level that LEADS
with a D-number counts as a definition unless the line also contains AMENDMENT
or RESOLVED. My heading "### D-106 RULING NOTE 2026-07-27 -- ..." led with the
D-number and contained neither, so it registered as a second definition of
D-106 and collided with "## D-106: Designate reactivation (VR1)".
The established convention already handles this and I should have followed it:
the existing D-101 ruling notes read "### RULING NOTE 2026-07-25 -- D-101 ...",
placing the D-number AFTER the words so the regex does not match. The heading is
now "### RULING NOTE 2026-07-27 -- D-106: ...", matching that pattern. Content
is unchanged.
Why the R1 and R4 amendments did not trip this: both contain the literal word
AMENDMENT, which the negative lookahead excludes.
Process fix for the rest of this session: commits are now gated on repo-lint
exiting clean, not merely preceded by a lint run whose output I read afterwards.
Revert: git revert this commit (restores the colliding heading and a red lint).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R5 RULED: designate deploys at Stage 5, is configured at Stage 7 (D-106 note)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Accept at Stage 5; rewrite Stage 7 Step 5 to
configure-not-deploy (Recommended)".
CORRECTION RECORDED IN THE DECISION ITSELF, not just this message. When R5 was
first put to the operator, the audit claimed this option "inverts D-106's
bootstrap order, which puts os-public-hostname + FQDN-SAN certs BEFORE
Designate". That was WRONG. D-106's order is a CONFIGURATION sequence -- static
hosts, then os-public-hostname, then Vault FQDN-SAN certs, then zones and A/AAAA,
then neutron, then tenant subnets. It governs when the DNS wiring happens, not
when the charm is installed. Deploying the app earlier does not invert it: the
zones still follow the certs. Conflating "install the charm" with "do D-106's
work" nearly cost a decision made on a false constraint, so the correction is
recorded in D-106 where a future session will meet it.
MEASURED BEFORE PRESENTING. All four designate applications and all EIGHT
relations are deploy-ready at Stage 5 -- every relation peer (mysql-innodb-cluster,
keystone, rabbitmq-server, vault, memcached, designate-bind) is created by Stage
5. But they will be FUNCTIONALLY INERT: os-public-hostname is set in NO deploy
artifact, and bundle.yaml:11 records the current posture as IP-ONLY with the dual
VIPs as the catalog endpoint. The charm runs; the feature does not. That is
exactly the state Stage 7 exists to resolve.
NEW SURFACE DEFECT FOUND BY THE MEASUREMENT: the phase-6 runbook contradicts
ITSELF. Line 172 states "There is no designate: or designate-bind: application
block anywhere in ..."; line 178 states "Designate is deployed in-bundle in each
DC". Both cannot be true. The bundle settles it -- DOCFIX-167 put designate there
on 2026-07-10 -- and the runbook needs rewriting regardless of this ruling.
Option (c), setting os-public-hostname at Stage 5, was refused and the reason is
worth keeping: it is the ONE branch that genuinely collides with D-106. Publishing
a public FQDN endpoint before Vault has issued FQDN-SAN certs recreates the exact
D-019 root cause D-106 was written to remove -- metal-only charms pulling a public
FQDN endpoint they cannot resolve, which is also the D-021 amphora constraint.
Option (b), suppressing designate at Stage 5, was refused because it needs
machinery that does not exist and partially undoes DOCFIX-167 -- building a
mechanism to satisfy a stale gate rather than fixing the gate.
Execution is a DOCFIX against the phase-6 runbook (Step 1 inventory prose, Step 5
verb and gate), logged to the Phase-3 batch, not executed under this ruling.
Revert: git revert this commit; the ruling note is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R4 RULED: reserved bands become enforced in MAAS via a DC-aware tool (D-134 amendment)
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Build a DC-aware tool; full v4 scheme + FIP now, v6 bands
after the carve (Recommended)".
MEASURED BEFORE ASKING, which sharpened the question considerably.
The defect is worse than "bands not enforced": D-134's band table has never
existed anywhere but prose. maas admin ipranges read returns THREE ranges
cloud-wide, ALL type=dynamic, ZERO reserved.
The collision is now QUANTIFIED rather than hypothetical. MAAS's own
subnet unreserved-ip-ranges for dc1 metal-admin reports its lowest free span as
10.12.68.5-.99 (95 addrs) -- precisely the .4-.49 utility band and the whole
.50-.99 VIP band. Against that, the dc1 bundle places 27 LXD units in base form
and ~55 once dc-ha-scaleup scales 14 applications, each needing a MAAS-allocated
address, and the VIP band is where the 33 per-DC VIPs live. Zero 10.12.*
addresses are allocated today, so this is a PRE-EMPTION and not an incident.
MAAS's exact allocation ORDER is deliberately NOT asserted -- lens 2 was right to
refuse that claim and I have not added it.
Confirmed there is genuinely NO VR1 path: only site-headend-install.sh
(office1-only) and phase-00-maas-standup.sh can create an iprange, and the latter
REFUSES for any non-VR0 DC. That refusal is CORRECT -- it is the cross-DC mixing
guard lib-net's selector exists to enforce. The gap is that nothing replaced it.
NEW ARCHITECTURAL CONTENT, and it exists because of R2: D-134's bands are v4-only.
Ruling dual-stack made the 33 VIPs dual-family, which left the v6 planes with no
band discipline at all and the same collision waiting in the other family. The
amendment establishes that the v6 planes inherit an equivalent scheme. The exact
v6 octet mapping is NOT ruled -- mirroring the v4 host-part layout is the obvious
mechanical default, not a fresh design question.
The v6 pass following the carve is FORCED SEQUENCING, not a deferral: a reserved
range cannot be created on a subnet that does not yet exist, and the v6 plane
subnets are not in MAAS.
Worth noting what the tool buys beyond the reservation itself: a `check` arm gives
D-134 an EXECUTABLE gate. Today nothing in the repo can detect that the bands are
unenforced, which is why this survived from 2026-07-23 to now.
Execution is a separate gated step. Standard delivery discipline applies to the
tool (harness green, gauntlet, repo-lint, changelog with revert); the 12-subnet
reservation pass is an operator-gated live MAAS mutation and is not batched with it.
Revert: git revert this commit; the amendment is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R3 RULED: finish the jumbo underlay; first recorded MTU budget verdict
...
GA-R5: question and exact utterance quoted, dated, pushed before dependent work.
Operator utterance: "Raise the two lagging segments to 9000 (Recommended)".
Recorded as a D-101 RULING NOTE. D-102 is the original MTU sub-policy but is
MERGED INTO D-101 and its body directs amendments there.
MEASURED BEFORE ASKING, which materially re-framed the question. D-101 requires
"the measured underlay MTU is a Phase-0 gate -- do not assume jumbo", and
scripts/dc-dc-mtu-geneve-budget.sh had existed without ever being run to a
recorded verdict. Running it both ways, plus measuring the host bridges, showed
the jumbo branch is nearly complete already rather than a large project:
- Every vcloud MESH leg is ALREADY 9000, including mesh-vr1-dc0-vr1-dc1
(virbr5), the inter-DC path itself. All six plane bridges on both racks: 9000.
- The four 1500 legs are the D-125 SIMULATED-ISP uplinks. They model the
internet, D-125's egress gate is defined against them, and they must STAY
1500. Not a defect.
- Exactly TWO segments lag: enp1s0 inside both containment VMs, and all 17 MAAS
VLAN records.
The MAAS record is the one that bites silently: MAAS renders VLAN MTU into node
netplan, so a jumbo bridge beneath a 1500 record still yields 1500 node
interfaces. Jumbo bridges alone do not deliver a jumbo underlay.
Why (a): tenant MTU stays 1500, so no per-charm MTU coordination is needed at
all. Option (b) required ovn geneve + tenant-network MTU + amphora to agree
exactly and permanently across both DCs, and NOTHING in this repo checks MTU --
neither cloud-assert.sh nor provider-bundle-check.py has any MTU assertion -- so
drift would be silent, which is precisely what D-101 calls "the classic
nested-OpenStack failure mode". Option (c) was the worst: jumbo at the plane
layer with a still-capped transit throttles D-108 rbd-mirror and radosgw
multisite behind a chokepoint invisible where an operator would look.
COUPLING TO R2, recorded because it changes the arithmetic: the 56-byte overhead
is the IPv6 figure and applies BECAUSE dual-stack was ruled. Under v4-only it
would have been 42, and the 1500-underlay tenant MTU 1458. R2 and R3 are not
independent.
Execution is a SEPARATE gated step and is NOT authorised by this ruling. The
verification owed is BEHAVIOURAL: an end-to-end large-frame test with DF set
across the inter-DC path. An interface claiming 9000 is not proof a 9000-byte
frame survives the path -- the assert-on-content rule applied to MTU.
Also worth recording: invoked bare, the budget script correctly REFUSES ("FAIL:
--underlay-mtu is REQUIRED -- no default, measure it this session"). It does not
guess. That discipline is why this verdict is trustworthy.
Revert: git revert this commit; the ruling note and capture are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Second gap sweep: U15-U17 closed; DC data path has no IPv6 at ANY layer
...
Operator: "Close the remaining gaps". All three measurable items are now
measured; what remains open is blocked by a known missing credential or is a
deliberate deferral, and both are stated with cause rather than left implicit.
U15 -- voffice1 transit addressing is REBOOT-DURABLE. A positive result, which
is worth recording as clearly as a defect: live enp2s0 172.31.0.1/30 (dc0 leg)
and enp3s0 172.31.0.5/30 (dc1 leg), both netplan-persistent via
/etc/netplan/60-transit.yaml and 61-transit-dc1.yaml. The D-128 Plane-2 path
survives a headend reboot and no site-baseleg row is owed for it.
U16 -- RETROFIT_WAIT=30m has NO recorded provenance. git log -S traces it to a
single bulk commit ("New Phase Scripts") with no rationale, and a repo-wide grep
finds no constant anywhere in scripts/ documented as nested-virt or depth-4
calibrated. It is an inherited default that has never been validated against the
nested I/O it will actually run on at Stage 5 Step 9. Log-only, but it should
not be mistaken for a tuned value. Bonus finding from reading the file: that
script's own preconditions require BOTH the openstack and juju clients, and no
host currently has both -- vcloud has neither, voffice1 has juju only. It cannot
run anywhere today, which reinforces S-1.
U17 -- the DC data path carries NO IPv6 at any layer measured. On BOTH racks:
zero global v6 addresses on any plane bridge, no v6 default route, and
accept_ra=1 (so they would accept a router advertisement; none is arriving).
Combined with the existing measurements -- no v6 subnet on any of the 12 MAAS
plane fabrics, no v6 link on any of the 18 nodes -- this is a THREE-LAYER
confirmation of L2-3. It also WIDENS the R2 propagation task: it is not only
"carve into MAAS", the rack bridges have no v6 either.
STILL OPEN, with cause rather than silence: the two DC edges' own interface-level
v6 configuration. dc0 is blocked by a known missing credential, not by a failure
to look -- ~/vr1-dc0-creds/ contains no opnsense-api.txt while ~/vr1-dc1-creds/
does, which is SEC-021(a) directly visible on disk, the residual item the
2026-07-27 consolidation batch deliberately excluded because the re-mint is a
live edge mutation. dc1's API is credentialed but measured unreachable from
vcloud (timeout; the path runs from the rack, where the edge creds are correctly
not staged per SEC-015). U17 answers the substantive question from the rack side.
repo-lint/gauntlet ON voffice1 remain deliberately deferred until precondition
0.1 advances that clone -- running them against a 105-commit-stale tree would
produce a number that means nothing. That is a judgement, not an omission, and
it is recorded as such.
Revert: git revert this commit; register and CURRENT-STATE changes are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

UNMEASURED-gap sweep: 14 deferred items closed; juju restore-backup CONFIRMED absent
...
Operator challenge: did the committee actually measure live state, or take
shortcuts? Answer: some of both. This commit is the accounting, plus the sweep
that closes what was closeable.
WHO DID WHAT, precisely, because the accounting matters:
- The apex was never polled by ANY lens. Lens 2 declared it out of scope and did
so HONESTLY -- its UNMEASURED note explicitly warned "D-101's literals may
exist in NetBox and simply not be carved into MAAS". No overclaim. But it is a
key system that lens had reach and time to measure. Under-scoped, not dishonest.
- The two-day-old dump was MINE, in the correction, not an agent's.
- Sharpest instance of the general problem: lens 6 declared juju restore-backup
unverifiable because "no Juju client exists on this host", while juju 3.6.27
was installed on voffice1 and lens 5 was running juju help against it in the
same session. One lens called impossible what another was measuring.
CLOSED BY THE SWEEP (14 items). The consequential ones:
- LIVE apex polled via netbox/office1-record-dump.py: 139 prefixes, 103 IPv6,
ZERO added, ZERO removed vs the dump. The conclusion was right; the method was
not, and that distinction is the whole point.
- juju restore-backup DOES NOT EXIST on 3.6.27 -- "not a juju command ... Did you
mean: create-backup". L6-14 goes from flagged RISK to CONFIRMED DEFECT: Stage 6
Step 9's D-104 restore drill and its DoD bullet are unsatisfiable as written.
- juju bind DOES exist as documented -- lens 6 was RIGHT to exclude it from L6-7
rather than assert it wrong. A correctly-handled uncertainty.
- dc0 compute provider MACs measured (52:54:00:1b:19:e6 / 52:54:00:18:ab:b4),
matching lib-hosts and CONFIRMING L3-8.
- curl present on both racks, so L4-13's hole is theoretical, not live.
- The pinned charm channels DO resolve (2024.1, 2.4, squid all present via juju
info on voffice1) -- proving L4-2's root cause and its fix: P3's 33 warns are
purely the missing juju binary on vcloud, not a charmhub problem.
NEW FINDING: preflight's "MAAS unreachable: 'maas admin subnets read' failed" is
a MISDIAGNOSIS -- the maas binary is simply ABSENT on vcloud. Together with the
absent juju (L4-2) and absent openstack (S-1), THREE separate preflight/deploy
failures on this jumphost are all "the client is not installed", and each is
reported as something else. A tool-absence must never be reported as a
target-unreachable. Its own DOCFIX.
STILL-OPEN gaps that are measurable and were NOT done are listed in the
register's section 3 explicitly, not buried: DC edge v6 state, repo-lint/gauntlet
on voffice1 (deferred on purpose -- that clone is 105 commits stale, so running
them today would measure a stale tree), voffice1 transit reboot-durability, and
the provenance of RETROFIT_WAIT=30m.
STANDING LESSON added for the next committee: before declaring an item
unmeasurable, VERIFY THE CONSTRAINT YOU ARE ASSERTING. The UNMEASURED escape
hatch is load-bearing for honesty and must not become a way to avoid looking.
Revert: git revert this commit; the register is new and the CURRENT-STATE
paragraph is additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

CORRECTION: the v6 literals are ASSIGNED (D-111), not pending -- R2a withdrawn
...
The operator asked whether the Office1 NetBox apex had actually been polled. It
had not. It has now been read, and the answer inverts a claim I made one commit
ago and a question I put in front of them.
WHAT IS TRUE: the v6 literals are assigned, ratified and recorded, and have been
since 2026-07-11 under D-111 (ADOPTED). Measured from
netbox/draft/vr1-office1-current-20260725.json -- 139 prefixes, 103 IPv6, every
relevant row tagged D-101/D-111: ULA fd50:840e:74e2::/48 with DC0 planes at
:220/:221/:230/:240/:250::/64 and DC1 at :320/:321/:330/:340/:350::/64; GUA
provider-public DC0 2602:f3e2:f02:10::/64 + VIP f02:11::/64, DC1
2602:f3e2:f03:10::/64 + VIP f03:11::/64.
WHAT I GOT WRONG: the R2 ruling note claimed D-101's "Remaining open item" (the
org ULA /48 and per-DC GUA carve) had become a Stage-5 precondition because
"dual-stack cannot deploy against literals that do not exist". They exist. R2a
("which literals") is WITHDRAWN as never having been open; no operator utterance
is owed on it.
THE REAL PRECONDITION IS NARROWER AND BETTER: propagation, not assignment, and it
needs no ruling. The ratified values are absent from the two places Stage 5
actually reads -- scripts/lib-net.sh carries no v6 arm at all, and MAAS carries
no v6 on any of the 12 DC plane fabrics. Both are mechanical copies from an
authoritative source, so both move from Part A (needs a decision) to the Phase-3
mechanical batch.
Note CURRENT-STATE's "DRIFT-FREE across netbox/lib-net/artifacts" is NOT
contradicted: drift-free means values appearing in more than one place agree, not
that coverage is complete. lib-net.sh simply never gained a v6 arm.
PROCESS FAILURE, OWNED AND RECORDED IN THE DECISION ITSELF: the audit's own lens 2
listed the apex as UNMEASURED and warned, in terms, that "D-101's literals may
exist in NetBox and simply not be carved into MAAS". That warning was correct and
available, and I put a question to the operator anyway on the strength of D-101's
stale prose. Trusting stale decision prose over an available measurement is
exactly what this audit was convened to catch.
NEW DOCFIX-CLASS FINDING, logged not fixed (hard rule 1): D-101's "Remaining open
item" paragraph still reads "pending NetBox assignment (gap #3)" for literals
D-111 adopted on 2026-07-11 -- a decision's own prose contradicted by a later
ruling, DOCFIX-200/204 class, and the direct cause of this error. Queued in the
Phase-3 batch.
The R2 ruling itself STANDS -- dual-stack as ruled. Only its stated effect changed.
Revert: git revert this commit to restore the prior (incorrect) framing; the R2
ruling in the preceding commit is independent and should not be reverted with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R2 RULED: dual-stack as originally decided; the v6 literals move onto the critical path
...
GA-R5: question and exact utterance quoted in the Status block, dated, pushed
before dependent work.
Operator utterance: "Carve v6 and deploy dual-stack as ruled (Recommended)".
Recorded as a D-101 RULING NOTE -- a re-confirmation, amending nothing, in the
same shape as the 2026-07-25 note it follows.
WHY IT WAS A LIVE QUESTION AT ALL: the 07-25 ruling was measured against the
substrate for the first time by this audit and found unimplemented. Exactly ONE
IPv6 subnet exists cloud-wide and it is on the Office1 base fabric; none of the
12 DC plane fabrics carries one; zero v6 links across all 18 nodes; every node
reports default_gateways.ipv6 = NONE. D-101's matrix requires ULA on
data-tenant/storage/replication plus legs on metal-admin/metal-internal, with
nothing to bind against.
THE CONSEQUENTIAL EFFECT, and the reason this commit matters beyond recording a
choice: D-101's own "Remaining open item" -- the org ULA /48 and per-DC GUA carve
-- has been carried since authoring as "pending NetBox assignment ... not a
ratification question", i.e. non-blocking. It is now a STAGE-5 PRECONDITION.
Dual-stack cannot be deployed against literals that do not exist.
Downstream, all now inheriting dual-family: R9 (where dc1's literals live), R11
(vault/designate VIPs, which must be dual-family from the outset or their cert
SANs get re-issued on a live cloud), and the L3-9 overlay collision, which must
be reconciled BEFORE either authority location is populated. Worth restating: of
the two merge orders the audit measured, the DANGEROUS one is the one that
PASSES -- vips-overlay-last silently replaces every v6 leg and reports green.
R8 (octavia's family) is explicitly NOT resolved by this ruling.
NEW SUB-QUESTION OPENED, queued not ruled: R2a, which literals. The ruling
directs that they be assigned; it does not assign them. I have deliberately not
proposed prefixes -- that is the operator's address space, and hard rule 2
forbids inventing a literal. Only the SHAPE is presented.
Revert: git revert this commit; the ruling note and R2a are additive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

R1 RULED: ceph-osd gets a real second disk on the storage nodes (D-121 amendment)
...
GA-R5: question as presented and the operator's exact utterance are both quoted
in the Status block, dated, and pushed BEFORE any dependent work.
Operator utterance: "Add an OSD volume to node-vm (Recommended)".
Recorded as a D-121 AMENDMENT rather than a new D-number. D-121 is the decision
that ratifies modules/node-vm sizing/count, and its own capacity re-validation
already records "Ceph disk re-run for 4 storage/DC = PASS 5.31 TiB" -- the second
disk was BUDGETED and never BUILT. GA-R3: architectural consequence, but it
amends a ruled node layout rather than establishing new ground.
TWO SCOPE CORRECTIONS landed with the ruling, both of which narrow it:
- It is EIGHT volumes, not eighteen. ceph-osd is placed on the four storage
nodes per DC (bundle.yaml:556-557, to: ["5","6","7","8"]). The other five
nodes per DC neither run ceph-osd nor need a second device. My earlier framing
said "all 18 nodes" -- true of the measured defect, wrong about the fix.
- Rack headroom measured before asserting feasibility: dc0 2.0T available, dc1
2.9T available under /var/lib/libvirt, thin-provisioned qcow2.
Why (a) over (b): MINIMIZE DELTA TO ROOSEVELT -- Roosevelt storage nodes have
real dedicated disks. Option (b) was also not safe to pick as presented: whether
ceph-osd at the pinned squid/stable supports a directory- or partition-backed
OSD was explicitly UNRESEARCHED, and the audit refused to assert it either way.
THE APPLY IS NOT AUTHORISED BY THIS RULING. Four preconditions recorded, the
sharpest being that re-commissioning is required for MAAS to see the new device,
and whether that preserves the D-134 statics and pinned MACs is UNVERIFIED --
the 2026-07-20 MAC-regeneration incident is the precedent for exactly that class.
modules/node-vm is shared by both DCs, so the change must be additive/opt-in or
it plans against all 18 domains.
Revert: git revert this commit; the amendment is additive and no artifact changed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|