| 2026-08-03 |

C.2 measured: the API-down teardown verb does not work; C.2b is the real path
...
Ran C.2's API-DOWN branch under operator authorisation. It FAILED, exit 1, after
~10 minutes (02:22:00Z -> 02:32:08Z):
Unable to open API: open connection timed out
ERROR cannot connect to model config API: unable to connect to API:
dial tcp 10.12.8.5:17070: connect: connection refused
Its help implies a direct-to-provider fallback ("Timeout before direct
destruction", -t default 5m0s) and the first version of Path C repeated that.
It does not reach the fallback -- it needs the MODEL CONFIG API to decide what to
destroy, so an unreachable API defeats it before the timeout matters. Corrected
in place rather than deleted, so the next session does not spend ten minutes
rediscovering it.
NOTHING WAS DESTROYED AND NOTHING CASCADED: the MAAS census is identical before
and after -- 10 machines, 9 Ready/owner=None, controller VM Deployed. Safe, just
useless.
Added C.2b, the manual pair, which is the D-061 coordination principle applied to
the controller -- clean up juju's view first, then MAAS's, never the reverse:
(1) juju unregister, client-side only, touches no cloud resource; (2) release the
controller VM in MAAS, identified BY TAG rather than hostname because MAAS
auto-generated hostnames do not match the ruled names. Gates on both sides: the
tag query must return exactly one machine and must not be a role node; after
release the count must be unchanged and the tag must survive, since the tag is
what C.6/C.7 bootstrap against.
Controller VM measured for the record: subtle-grouse, system_id arfr7p, Deployed,
owner juju-vr1-dc0, tags [virtual, juju-controller-vr1-dc0].
C.2b step (2) is a gated mutation and has NOT been run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Path C: full juju CONTROLLER teardown + rebuild, from the existing build sequence
...
Operator: "We already have a build sequence. Build out the full teardown and
rebuild steps." Assembled from phase-4's own steps rather than invented; Path C
cites them instead of duplicating their detail.
C.1 census first (record the MAAS machine count -- a drop is the 2026-07-21
cascade signature). C.2 the teardown, chosen by whether the API answers, which
you TEST rather than assume. C.3 verify release without cascade. C.4 client-side
unregister. C.5 the region-scoped credential gate (phase-4 Step 2.0 +
DOCFIX-206), including the point that a controller rebuild does NOT invalidate
the MAAS credential -- it belongs to the cloud definition in the client, not to
the controller -- but prove it anyway with the scoped login/read/logout
sequence. C.6 the controller-tag gate. C.7 bootstrap with BOTH ruled constraint
flags and --bootstrap-base, noting there is no --dry-run for bootstrap, which is
why C.5 and C.6 are gates rather than formalities. C.8 model-defaults, flagged
hardest: they live ON the controller, so a rebuild loses every one and nothing
carries over. C.9-C.10 hand back to phase-4 Step 3.5 / 3 / 3.9 then 4.2-4.4.
C.11 what a rebuild does not restore.
POLICY GAP LOGGED: the committed deny list gates the WEAKER controller-removal
verb while the STRONGER one is ungated. Path C says to treat both as
operator-gated regardless of the rule engine, per SEC-030's finding that the
presentation discipline, not the rule engine, is the real gate here.
Also reordered so M.6 stays with Path M rather than being stranded after C.11.
GUARD-HOOK NOTE: the first attempt to commit this was BLOCKED by
.claude/hooks/guard-destructive.py, which matched the controller-removal command
names appearing in the commit MESSAGE. The hook cannot distinguish documenting a
command from invoking one. Recorded as a finding; the guard behaved
conservatively and correctly, and the message is passed by file rather than on
the command line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

my reap poll reported REAPED while the controller was down -- absence-test defect
...
The background wait-loop polled 'juju models | grep -q name: admin/vr1-dc0' and
broke when the grep MISSED, reporting 'REAPED after ~2730s'. The model was still
there. With the API refused, juju models prints NOTHING, the grep matches
nothing, and the absence-test read empty output as 'gone'. Verified after: port
17070 still REFUSED, model still present.
This is verbatim the rule this repo already states -- 'could not look' is never
'nothing there' -- failing in a checker I wrote in the same session I wrote that
sentence into this very runbook. Same command also used 'cmd | head -3; rc=0',
which captures head's status, not the command's.
Path M gains a second instrument warning: poll on a POSITIVE signal -- require
juju models to actually return model lines, then ask whether the target is among
them; treat empty output or a non-zero juju exit as REFUSE, never as gone.
No mutating action taken. Controller remains down; awaiting the operator's
decision on the rebuild.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

model destroyed OK, then --force orphaned it and took the controller DOWN
...
destroy-model returned 'Model destroyed.' EXIT 0 with a clean progressive drain
(36/56 -> 26 -> 20 -> 16 -> 9 -> 0) and a correct MAAS release: census 10, count
UNCHANGED, nine nodes Ready/owner=None, controller VM still Deployed. Proxies
PASS. M.4 model-defaults corrected and read back.
Then add-model refused: model already exists, stuck at life: dead. Root cause
from the controller log -- the destroy left the model doc alive while its STATUS
doc was gone. undertaker crash-loops on 'cannot set status: model not found';
modelcache crash-loops on 'status doc <uuid>:e not found' (732 iterations in ~12
min). The API server depends on modelcache, so 17070 is connection refused and
THE CONTROLLER IS DOWN -- while the VM pings and systemctl reads active. A
jujud restart did not fix it. LP #1737487 class.
OWNED, two of my decisions are implicated: I carried --force --no-wait over from
the 07-31 stall where agents were STOPPED and force was genuinely required; this
run's agents were ALIVE and draining, so force was almost certainly unnecessary
and it is the documented cause of this inconsistency. And when the model would
not reap I re-issued destroy against a model already dead, which moved it back
to dying and re-armed the loop.
Path M amended: try without --force when agents are alive; reserve --force
--no-wait for the measured stall; never re-issue destroy on a dead model; if the
controller is already down on this symptom, rebuild the controller rather than
attempt state-DB surgery.
Not degrading: nine nodes Ready, MAAS healthy, both proxies PASS, model-defaults
now controller-level.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Path M: the juju MODEL teardown + release path, from the runs we already logged
...
Operator: "You have that information from previous teardowns for the juju
release path ... find the previously logged information in the repo. Log the
commands and steps for future reference ... this has bit in the past and will be
a reusable and needed set for the future."
I had said the release path was the one thing I could not prove in advance. It
was already proven and logged, in CURRENT-STATE under MODEL TEARDOWN 2026-07-31.
Two corrections to the command I proposed:
1. --no-wait BELONGS IN IT and I had excluded it, reasoning from juju help that
it was reckless. Measured 07-31: the plain destroy-model STALLED TERMINALLY --
'attempt 30 ... model not empty, found 26 machines, 37 applications', flat ~19
minutes, application set byte-identical, because ALL 26 agents were stopped so
no teardown hook could execute. --force --no-wait cleared it (18->5->2, then
Model destroyed.). The help text talks you out of the flag that works.
2. The release path is proven, not unknown: all nine role nodes came back
Ready/owner=None, zero stranded, no maas machine release needed or run.
Expected post-state is 9 Ready + 1 Deployed -- the controller VM stays
Deployed in the controller model, so expecting 10 raises a false alarm.
Also recovered, and live for this rebuild: destroy-model TAKES THE MODEL CONFIG
WITH IT. On 07-31 that silently removed apt-mirror and the spaces work and
nothing in the repo would have caught it. Measured today: this controller's
model-defaults carry apt-mirror=http://10.12.8.4/ubuntu and nothing else -- wrong
after the convergence -- while the three settings the deploy needs were set at
MODEL level and will be destroyed. Defaults apply to NEW models only, so they
must be fixed BEFORE add-model.
And the VR0 pod warning does not transfer: VR0's virsh-POD MAAS decomposes
pod-composed machines on destroy-model; VR1 uses per-machine power_type=virsh,
not pods, so there is nothing to decompose.
The teardown runbook documented only SUBSTRATE teardown -- the model layer had no
procedure, which is why this kept being re-derived. Now Path M, M.0-M.6.
Final command: juju destroy-model vr1-dc0 --force --no-wait --no-prompt
BLOCKED by the permission layer and NOT run; no workaround attempted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-08-02 |

WITHDRAWN: the designate "contradiction" was mine, not the bundle's
...
I raised, as an unresolved question for the operator, that the Stage-5 plan
deploys four designate applications while a phase-4 bullet says the bundle
ships "NO designate". There was nothing to resolve.
D-019 (v1 ships no cloud DNS) is SUPERSEDED by D-106, which REACTIVATES
Designate for VR1. DOCFIX-167 put all four designate applications into this
bundle.yaml on 2026-07-10, and phase-01-bundle-deploy.md:174 already says the
change "corrects this GATE's old 'NO designate (D-019)' text, since D-019 is
superseded". Stage 7 owns the DNS ACTIVATION -- per-DC zones + A/AAAA per the
D-008 bootstrap order, FQDN-SAN certs, the B5 os-public-hostname reversal --
explicitly not the charm deploy. Four designate applications in the plan is the
planned outcome; their absence would have been the defect.
The stale artifact was the phase-4 bullet alone, paraphrasing a gate text
phase-01 corrected 23 days earlier. Rewritten to state the D-106 reactivation
and the Stage-5/Stage-7 split. My annotation is deleted, not kept as history --
leaving it would mislead the next reader about a live deploy input. Step 4.4's
speculative designate caveat removed for the same reason.
Operator correction quoted in CURRENT-STATE and the changelog: VR0 -> VR1 is
ADDITIVE, and a manufactured contradiction costs next-step attention, which is
this project's named failure mode. Also recognised late: the dry-run's ignored
"name"/"variables" field warnings are the benign R11 pair phase-01 already
documents.
repo-lint 0 fail / 1 legacy warn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

DOCFIX-208 part 3: graduate the dry-run instrument finding; harden Step 4.4
...
The juju finding belongs above a changelog -- it generalizes to any bundle
deploy, not just phase-4's gate. platform-traps.md gains a Juju section: what
--dry-run prints, what it never prints at any verbosity, the two grep decoys
(controller API addresses look like DC-band VIPs; --debug echoes the command
line so overlay filenames hit), and the three consequences.
Step 4.4, written earlier this session and never run, gets two corrections
before its first execution: the --format=json read used a direct index, so a
missing key would print a traceback into a results column -- this repo has a
logged instance of an assertion satisfied by a traceback -- now .get() with a
visible <<UNSET-OR-APP-ABSENT>> sentinel documented as a failure. And the block
now states plainly that it is authored against a measured instrument but has
never been exercised, with the instruction to report a shape correction rather
than work around one.
repo-lint 0 fail / 1 legacy warn; gauntlet ALL GREEN (98).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

DOCFIX-208 part 2: the dry-run graded the gate, and three items could not fail
...
The corrected dc0 command dry-ran EXIT 0 from the dc0 rack against an empty
vr1-dc0 model, staged input sha256-verified against HEAD first. Nine machines
0-8, all nine constraints tags=openstack-vr1-dc0,<role> in the 3/2/4 split,
all three overlays consumed. Capture: docs/audit/stage5-dryrun-dc0-20260802.txt.
The run also measured what the command can and cannot show. juju deploy
--dry-run on a bundle prints NO application options and NO VIPs at any
verbosity; --debug adds only the per-machine constraint lines. Greps over the
--debug capture: bridge-interface-mappings 0, physnet1 0, openstack-origin 0,
cloud-archive 0, prefer-chassis 0.
So three Step-4.2 gate items could not be graded by the command the gate names:
item 1 (tags) needed --debug, which is now in the named command; the
PRE-EXISTING VIP item could never have failed because no VIP is ever in the
plan, and its property is relocated in writing to preflight P2; and the two
option-reading items I added earlier this session had the same defect and are
removed within the hour. Their properties become new Step 4.4 -- a juju config
read run the moment deploy returns, which is the only thing that can prove
juju merged the options map rather than replacing it.
Both machines overlays' VERIFY-LIVE headers said to assert this "at the live
dry-run". Measured impossible; corrected in place, pointed at 4.4.
Also settles the deferred ceph-osd tags=openstack item (placed by explicit
placement, bare tag absent from the plan) and logs a contradiction not resolved
here: the runbook says the bundle ships NO designate, the plan deploys four
designate applications.
The deploy itself is NOT run -- 4.3 is a gated mutation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

DOCFIX-208: the dc0 deploy command was missing a load-bearing overlay
...
The phase-4 runbook told the operator to deploy vr1-dc0 WITHOUT
overlays/vr1-dc0-machines.yaml, on the stated grounds that the file does
not exist. It has existed since 2026-07-29. preflight.sh P2 folds it into
the merged input it validates, so P2 was grading an input the deploy would
not have passed.
Re-measured by deep-merging bundle.yaml against the overlay: SIXTEEN
applications carry a delta, not the one this repo recorded on 2026-07-31.
ovn-chassis.bridge-interface-mappings (whose ONLY source is this overlay --
bundle.yaml:494 forbids re-adding it there) plus 15 mirror repoints from
the 2026-07-31 UCA ruling. Deploying as written would have taken two
injuries, neither surfacing as a deploy error: no br-ex mapping (dead
provider egress) and 15 charms pointed at a UCA measured unreachable from
a node under the D-107 airgap.
Corrected at 6 sites (4.1 table + prose, 4.2 dry-run, 4.3 deploy, Step 7
dry-run + its rules sentence, 4.2 VERIFY-LIVE pointer). Step 4.2's gate
could not have caught any of it -- its four items read the machines block,
machine count and VIPs, while the overlay's whole payload is under
applications: -- so gate items 5 and 6 were added to assert the merged
ovn-chassis options map and the reachability of every origin/source.
preflight.sh needed no change: it was already right.
CURRENT-STATE corrects its own narrower 2026-07-31 finding in place
(GA-R1 C2). repo-lint 0 fail / 1 legacy warn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Fold: open the register, close both Class-A rows (F1 phase4, F12 SKILL.md)
...
Operator ruled fold-before-exercise, start at Phase 2, all DCs come up in their own
region. Register: docs/runbook-fold-register.md, 12 rows classed A/B/C.
Measured gap: D-138 and D-139 appear in NO runbook. Nothing in the chain builds a per-DC
MAAS region, a snap proxy, or the GUA carve. Running it as written would rebuild the
pre-D-132/D-138/D-139 shape and fail where we already fixed things by hand.
F1 (Class A): phase4 asserted "Both DCs deploy from the Office1 headend by the SAME
procedure" and "Every juju and maas command runs there". The 2026-07-27 ruling it cited
was about DC ordering, not run location; D-138 reversed the location half. Replaced with
a tool-split table; superseded text struck in place, not deleted.
F12 (Class A, worse): SKILL.md -- loaded by every session before any runbook -- had ZERO
mentions of D-138 while stating Plane 2 executes on voffice1. Now carries the split, the
structural reason, and the measured test: juju controllers on voffice1 -> "No controllers
registered"; on the dc0 rack -> the live vr1-dc0-controller.
Both record the racks-have-no-repo-clone trap and the sha256-verify-staged-copies rule.
A row closes on the edit AND on being exercised by the dc1 Phase-2 rebuild; an edit alone
is a ruled-but-not-built claim.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Close both egress gaps: dc-egress-check.sh + preflight P9, wired into restart + phase-4
...
Gap 1: nothing tested DC egress. dc-rack-net.sh check PASSED through a 19-hour outage
because it asserts legs and units -- a leg is not a path -- and the mirror answered 200
from its own nginx, which says nothing about upstream.
Gap 2: the post-reboot check set was my judgement, and it omitted egress.
New scripts/dc-egress-check.sh: site-keyed, runs on the rack, LAYERED (route -> edge
answers -> traffic leaves, ICMP and TCP -> the three upstreams), reporting the FIRST
failure as the cause while the rest SKIP. A4 asserts the UPSTREAM the local artifact
services sync FROM, which is independent of whether they serve.
Proven live on two different failure modes: dc0 fails at A2 naming the dead edge; dc1
passes A2 and fails A3/A4 -- traffic not leaving a healthy edge.
preflight P9 WARNs off-rack with the exact command, FAILs on the rack when broken, WARNs
if the checker is absent. Gap 2 is closed by the WIRING: restart procedure Stage 0
(before anything fetches) and phase-4 Step 3.9 (before the deploy), both stating P9 is
not a substitute for running it on the rack.
Mutation pass caught a decorative test of mine: T11 never reached A4's unrecognised-code
branch, so flipping it to ok() left the suite green. T14 added to exercise A4 alone.
tests/dc-egress-check 14/14; tests/preflight 39 -> 43/43; 5 mutations across both, each
killing a named case, scripts restored sha256-identical. Gauntlet ALL GREEN (97),
manifest 96 -> 97; repo-lint 0 fail / 1 warn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-31 |

Juju re-pointed at the dc0 region: credential minted, proven, registered (SEC-028)
...
Operator-ruled: 'Mint juju-vr1-dc0 on the new region (Recommended)' and
'Add it to allow'. juju-vr1-dc0 minted in the dc0 region as a superuser,
mirroring the measured Office1 shape. Never printed; moved host-to-host with
all three sha256 digests compared; staging shredded.
Proven to AUTHENTICATE before bootstrap -- login + users/machines/rack reads
confirming superuser scope and DC-local region identity. Cloud re-pointed to
10.12.8.6:5240 and read back; the stale Office1-keyed credential removed first.
DOCFIX-206: Step 2.0's credential gate was not region-scoped, and I hit it --
all three checks passed while the user lived in Office1, so the SKIP branch
would have led to an auth failure that reads as a network fault.
Two snap-confinement traps recorded: private /tmp, and the home interface
refusing dot-directories. Neither error names confinement.
Registered BEFORE bootstrap: SEC-028, 8 matrix rows, dc0 manifest, and the
first 'rack' rows in vm-secret-locations. creds-matrix 65/65; repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-30 |

lib-identity.sh: the estate's name and DNS base in ONE place
...
Operator: 'The specific server and region names do not matter as they change
during every deployment and hardening of the deployment workflow.' Then, after
scoping the two options: 'Yes, take option 2.'
WITHDREW MY OWN FIRST RECOMMENDATION. I proposed retiring the vr0-dc0 selector
as removing 'the ambiguity class outright'. Measured, wrong on both halves:
D-119 already closed it by region-qualifying (the hazard was the BARE dc0, which
is rejected loudly today); lib-net.sh:110 documents that arm as a NO-OP over the
file's flat defaults; and the footprint is 43 files, mostly the separate VR0
phase-NN track the VR1 runbooks cite as precedent. Medium cost, near-zero
benefit. Logged, not executed.
NEW scripts/lib-identity.sh -- CLOUD_NAME + CLOUD_DOMAIN, no side effects,
env-overridable. With <dc> and <region> already derived from the site token,
these were the LAST typed identity in the shell surface. A rebuild that renames
the estate now edits one file and the certs, zones, SANs and P7 all follow.
DELIBERATELY NOT lib-net.sh: sourcing that bare populates a full flat plane/VIP
namespace (its own header says so). A certificate checker needs two strings, not
a network namespace, and octavia-pki.sh sources NOTHING today -- pulling in
VR0-shaped defaults it never asked for is the R9 hazard in miniature. Sourced
from SCRIPT_DIR (sibling, survives the harness REPO override) and FAILS CLOSED
with REFUSE 3 if absent: a cert gate that invents the estate's identity is worse
than none.
Consumers: octavia-pki.sh derive_zone(), phase-01 1.0-GEN.c, dc-dc-phase6 Step 0.
The tofu side was ALREADY parameterised (opentofu/variables.tf domain_suffix
passed explicitly into modules/dc-planes), so nothing there changed.
HARNESS 48/48 -> 51/51:
- T45 the centralisation must be REAL, not decorative -- overrides both values
and requires the zone to follow BOTH (acme.dc0.vr1.example.test). A constant
nothing can vary is indistinguishable from a literal.
- T46 a missing lib-identity.sh REFUSES rather than guessing.
- T47 SHELL vs OPENTOFU drift. HCL cannot source a shell file, so two copies of
one fact exist; if they part, the libvirt plane domains and the cert zones
describe different estates and nothing else would notice. PROVEN able to fail
by injecting a drifted value, then restored.
Also: repo-lint L6 flagged '. $REPO/scripts/...' as a bare invocation. Its own
docstring blesses a bash|source|python3 prefix and sourcing needs no exec bit,
so the RUNBOOK moved to the documented 'source' form -- the gate was not changed
for style. And ${REPO:?} is now guarded at its FIRST use in GEN.c, not only at
:538 below it.
STILL typed, deliberately: the site allowlists. One that accepts anything is how
a typo'd DC name becomes a wrong-target write.
gauntlet ALL GREEN (89); repo-lint 0 fail; GEN.c bash -n clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

F8 + F9 fixed AT SOURCE in 1.0-GEN.c -- the next DC standup will not recreate them
...
Everything before this repaired the two EXISTING DCs. The GENERATION path still
carried both defects, so the next DC standup would have minted them again.
F9 at source: GEN.c baked omega.dc0.vr0 as a literal in THREE places (CN + both
DNS SANs). It now DERIVES DC_ZONE from $DC with the same two expansions the tool
uses (${DC%%-*} / ${DC#*-}), so runbook and gate cannot disagree, and echoes the
zone for confirmation before the sign.
F8 at source: the cat > controller.cnf heredoc is replaced by one printf per
line, values passed as %s ARGUMENTS. The 2026-07-29 note deferred this as 'an
untested rewrite ... reasonable when someone can run a real generation'. That
condition is now MET: the identical shape was exercised end to end by two real
mints (both DCs) plus 48 harness cases.
Plus structural assertions BEFORE signing -- four sections present, subjectAltName
wired, CN equals the derived zone, exactly 2 DNS entries -- because printf removes
the paste hazard but not the failure CLASS: a typo'd section name still yields a
SAN-less cert behind a wall of OK output.
The stale 'NOT changed -- outside R7's ruled scope' note is superseded in place,
keeping the reasoning worth carrying: R7 recorded this cert's SAN as 'already
DERIVED per-DC by design', true of the IP SAN and NOT the DNS names -- a claim
accurate about one half of a field, read as covering both. That is how F9 survived.
F10 HANDLED DELIBERATELY. Editing GEN.c shifts every mint-ref anchored below it.
All 13 octavia anchors re-resolved BY MARKER in a SINGLE PASS keyed by row id --
never sequential seds, because a line number can be simultaneously an old value
for one row and a new value for another -- then each verified to point at its
correct command. 12 rows rewritten; creds-matrix S4 CLEAN; block bash -n checked.
octavia-pki 48/48; creds-matrix 65/65; creds-matrix 101 rows / 5 findings, all
pre-existing; gauntlet ALL GREEN (89); repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

octavia-pki: a hardened, tool-driven controller-cert REISSUE path (F9's remedy)
...
Operator: 'Let's reissue now. Engineer the artifact, variable, based workflow to
reduce the chance of errors.' Five GA-R5 rulings, one exchange each, quoted in
the changelog; ruling 5 was 'Full script minter now'.
THE TOOL IS BUILT; THE MINT IS NOT YET RUN (gated operator action, 1.0-REISSUE).
scripts/octavia-pki.sh gains 'reissue <site> [--force] [--dry-run]': controller
LEAF only, signed by the EXISTING controller CA (neither CA regenerated, amphora
trust domain untouched), fresh P-256 key, zone DERIVED from the site token and
never typed, printf-built config (the F8 payoff), stage-assert-promote, and a
single-value overlay surgery that never puts the bundle in argv.
SIX GATE DEFECTS FOUND BY VALIDATION AGENTS, ALL CLOSED. Mutation testing is
what earned this: deleting the three assertions I had just written left the
harness fully GREEN every time, so they were decoration until each got a
failing-direction fixture.
- A15 keyUsage/EKU on the ISSUED cert. A cert with NEITHER minted, promoted and
passed everything. keyUsage appeared once in the file: the printf writing it.
- A16 validity. openssl x509 -req defaults to -days 30, and checkend/notAfter/
enddate appeared ZERO times in verify -- a 30-day cert passed forever.
- A17 the overlay must decode to what the workspace holds. A10 graded shape,
A1-A16 graded the workspace, nothing joined them; a desynced overlay read
PASS 0-failed and the overlay is what reaches the charm.
- The re-run guard REFUSED to fix broken certs: names-correct + stale IP SAN
gave verify FAIL and reissue exit 4 'nothing to fix'. Now gated on verify.
- A8's negative proves its instrument loads first (openssl verify exits non-zero
for 'could not load the CA' too, so a corrupt CA passed it vacuously).
- Two pre-existing overlay defects moved pre-mint, so they stop producing a
false exit-5 with the workspace already promoted.
Harness 23/23 -> 48/48. Runbook Step 1.0-REISSUE appended AFTER GEN.e
deliberately (F10 line-anchored mint-refs; creds-matrix reports no S4 findings
after the append). Skill gains two standing invariants for future DC standups:
every per-DC secret needs a ROTATION tool not just a generation recipe -- a
REFUSE-IF-PRESENT generation gate makes the dangerous path the only path -- and
a hand-pasted openssl chain is a defect class, with F8 and F9 as the receipts.
Also corrected: D-137 fork 1 is NOT open (sub-ruling 1, 2026-07-25, ruled
'Blocking in preflight' and explicitly declined the PreToolUse guard), so the
queued Q2 premise is stale in the same way Q1 was.
gauntlet ALL GREEN (89); repo-lint 0 fail; creds-matrix 5 findings, all
pre-existing and named; ledger-scan unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

DOCFIX-205: Q1 withdrawn -- D-117 ruled it 2026-07-13; annotate the four it owed
...
The operator asked to rule on queued Q1 (is substrate vr1-dc1 'dc1' by token or
'dc2' by position, for cert/DNS identity). Measured before presenting options.
The question was WITHDRAWN, not ruled: D-117 (ADOPTED 2026-07-13, four days after
D-106) names the supersession in its own Status line -- "Supersedes ... the D-106
dc1/dc2 zone labels" -- and rules the replacement: "The repo's dc1/dc2 labels are
retired in favour of dc0/dc1". vr1-dc0 -> dc0, vr1-dc1 -> dc1.
WHY IT RESURFACED, and this is the reusable half: D-117 ruled that the ADOPTED
decision texts D-101/D-106/D-111/D-115 be ANNOTATED in place. Measured: ZERO of
the four carried any D-117 annotation. It stayed invisible because D-117's OWN
Status line claimed "FULLY EXECUTED BY D-119", while D-119 scopes its discharge
to the SELECTOR half only. A reader checking whether the annotation was owed was
told it was already done. Ruled is not built -- check the artifact.
Landed (operator-approved batch, "All four + code fixes"):
- All four decisions annotated in place, per D-117's own ruled treatment. D-101
gets the hardest one: it carries BOTH namespaces, so it is annotated BY DATE
(dc1 means different datacenters in its two halves -- D-117 TRAP 2).
- D-117's Status line corrected by measurement (GA-R1/C2) to "EXECUTED IN TWO
HALVES", with the standing lesson that a Status line is a CLAIM about
execution, not evidence of it.
- phase-6 EXECUTABLE DEFECT fixed: os-public-hostname was built as
"keystone.omega.${DC}.vr1..." with $DC the D-119 selector, expanding to
omega.vr1-dc0.vr1... -- the region TWICE -- and would have been baked into
Vault-issued SANs at Stage 7. Both labels are now DERIVED from the site token
(${DC%%-*} / ${DC#*-}); verified by running it, not by reading it.
- octavia-pki.sh A12: refusal -> assertion on the full derived zone; a SAN-less
cert now also fails. Harness 21/21 -> 23/23, with T21 REPLACED (not deleted)
and new T21b (the other DC's label in the right region FAILS -- the cross-DC
mix-up region-only checking passed) and T21c (per-DC derivation proven).
- CURRENT-STATE + ledger corrected in the same commit (L10). "BLOCKING" was
wrong as stated: A12 measures INERT and PASSES live on both DCs. F9's reissue
obligation stands.
gauntlet ALL GREEN (89); repo-lint 0 fail; ledger-scan unchanged (3 decisions,
21 SEC, D 138 / BUNDLEFIX 053; DOCFIX 205 -> next-free 206).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Pre-bookend sweep: the operator's working commands were NOT in the repo
...
Operator: "I'm worried that the fixes we made to the defective commands will be lost when we
close this session ... Complete a full sweep before the bookend."
THE CONCERN WAS WELL FOUNDED. The corrected Step 5/6 commands -- the ones that actually produced
both DCs' live PKI -- had been reviewed in conversation, judged equivalent, and never folded in.
Four items now landed:
1. Portable base64. GEN.d shipped the GNU-only `-w0` flag; what ran was
`base64 < f | tr -d '\r\n'`. The runbook described a command nobody executed.
2. The operator's VIP_OVERLAY:? guard in GEN.c.
3. The heredoc column-1 requirement, as a CAUTION heading rather than a footnote, because that
failure is silent: no alt_names section yields a certificate with NO SANs while every
openssl command still prints OK.
4. ${DC_LABEL:?} at both CA call sites -- F8's sibling, where an unset label bakes a DC-less
subject into a 10-year CA with no error anywhere.
Not done, with the reason recorded: the heredoc was NOT rewritten as printf lines. More
paste-proof, but an untested rewrite of a step that mints 10-year CA material and cannot be
exercised end-to-end from an agent session. Risk bounded instead by A9 plus harness T2.
F9 is now structural rather than remembered: A12 arms itself from os-public-hostname appearing as
a real option key. Live PASS 29/0 both DCs, harness 21/21.
A new ruling-shaped gap, now blocking rather than filed: D-008's shape plus D-106:2563's VR1
instantiation do not say whether substrate vr1-dc1 is dc1 by token or dc2 by position -- the
DC1/DC2 ambiguity item 3.1 retired elsewhere, here deciding certificate identity. Both live certs
carry dc0.vr0, the VR0 region, wrong under either reading. A12 refuses rather than picking.
A precedence bug the harness caught: REFUSE was checked before FAIL, so once A12 armed a
CONFIRMED wrong-region SAN reported as "could not evaluate". A known defect outranks an
unevaluated one.
Sweep capture docs/audit/queued-findings-20260730.txt carries F10-F13 and two ruling-shaped
questions. F10 is the one worth acting on: mint-ref line numbers drift silently and S4 cannot see
it, since it asserts only within-EOF. That bit three times in two sessions and was caught by hand
every time, never by a gate. All 24 octavia refs re-anchored in a single-pass mapping keyed by row
id -- 391 was simultaneously an old and a new value, so sequential seds would have corrupted it.
repo-lint 0 fail; octavia-pki 21/21; creds-matrix 65/65.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

tfstate backup path built; E2/E3 closed live; a gate disagreement found and fixed
...
BACKUP PROVEN BY RESTORE, not by existence. Both PKI archives were extracted and compared
byte-for-byte against the live tree: 12/12 sha256 MATCH across both DCs, covering both encrypted
CA keys, both passphrases, both CA serials and both controller bundles. Archives are 600 inside
700 folders and differ between DCs, so per-DC independence survives into the backup.
E2/E3 CLOSED LIVE. Operator ran the chmod and the intermediate removal; octavia-pki.sh verify
now reads PASS 26/0 on BOTH DCs, workspaces 12 -> 10 files, and P5 on the headend fell 18 -> 8.
A GATE DISAGREEMENT, recorded as the pattern rather than the incident. verify read 26/0 while
creds-matrix E2 still reported the CA serial mode 664 on both DCs. The chmod list predated the
serial being declared, and verify's mode assertions covered six private files and three certs --
the serial was in neither. Two gates that can both see a file must not disagree about it. Fixed
in three places: verify now mode-checks the serial (integrity, not secrecy -- rewriting it forces
the next issuance to reuse a serial), the generator's chmod includes it, and T18 pins it. The
harness then caught its own fixture creating the serial at the inherited umask; baseline went red
until the fixture matched the corrected generator. 18/18.
TFSTATE BACKUP as dc-dc-phase2 step 13, 2 rows, notes key n-tfstate-backup. Register 99 rows;
schema, render-drift, mint-ref and notes checks clean.
Three deliberate differences from the PKI backup: tfstate is DYNAMIC so the serial is recorded
both sides and the step must be re-run after every apply (restoring a stale state is actively
dangerous -- tofu then believes everything created since does not exist); the pull block proves
RESTORE by decompressing to check the archive parses as JSON and reports a serial, because a
truncated gzip passes a size check; and mint-stage is stage3, which is REACHED, so P5 correctly
reports these EXPECTED-BUT-ABSENT until the step runs rather than deferring them.
Classification recorded honestly: the outer state is credential-bearing (DOCFIX-175, MAAS API key
in plaintext); the inner states show no credential-shaped key names, consistent with the inner
root using only libvirt over qemu+ssh. Strong evidence, NOT proof -- a name-based absence cannot
exclude material inside a value -- so they are registered and stored as if sensitive. Uncertainty
resolves toward more auditing, and an unaudited backup directory beside audited ones is how the
SEC-022 shadow stores happened.
Two of my own defects caught by the gates: repo-lint L9 rejected a hardcoded clone name in the
new step (the repo has been renamed once, D-110) -- now $REPO; and two mint-refs had drifted when
a later edit added lines above their anchors, which S4 cannot catch because it only checks
line <= EOF. Both re-anchored.
Also corrected: phase-01:605 said the overlay must be in the backup set. It is derivable from the
archived workspace by 1.0-GEN.d, so the line implied a gap that does not exist.
Known, not fixed: both live inner states are 664, group-writable, and they are the authority tofu
trusts. Tightening needs its own gated change plus proof the provider preserves the mode.
creds-matrix 65/65, octavia-pki 18/18, repo-lint 0 fail / 621 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Build the PKI backup path: three surfaces asserted a backup set that did not exist
...
Operator direction: back up to the per-DC jumphost creds folder BEFORE the cert cleanup, and
record that the pinned secrets-storage solution must carry a certificate/credential backup
procedure with this step folded into it.
THE GAP, MEASURED. phase-01:594 said the workspace must be "backed up securely", :605 said it
"MUST be in the per-DC backup set", and CURRENT-STATE said the inner tfstate should be added to
"the site backup set" -- while cloud-snapshot.sh, the only candidate, is a juju-layer capture
that mentions octavia, tfstate and terraform ZERO times. Both DCs' 10-year amphora trust roots
therefore sat in one place on one VM with no copy. Losing the issuing CA key means Octavia can
never sign another amphora certificate; losing the controller CA key means the controller
certificate can never be reissued, which F9 says it must be at Stage 7.
SHAPE: one gzipped archive per DC rather than a file-by-file tree copy -- 2 register rows instead
of ~20, atomic (a half-copied tree is the dangerous state), and sha256-verifiable. The live tree
on the headend stays what octavia-pki.sh verify asserts; the archive is recovery only.
COST STATED RATHER THAN GLOSSED: a second at-rest copy of both encrypted CA keys beside their
plaintext passphrases, so the archive's contents defeat encryption-at-rest. That is the trade
D-109 option (b) was refused for; the distinction is deliberate, because that ruling governed
where the authoritative artifact lives and which host the deploy reads. A recovery copy is not a
second source of truth.
Delivered: phase-01 step 1.0-GEN.e (build on the headend, pull to the per-DC creds folder,
compare sha256 against the source BEFORE removing the staging copy -- a truncated scp would
otherwise leave a verified-looking backup of nothing); 2 register rows, custody=consolidated
since the creds folder is the SEC-009 location; notes key n-pki-backup.
THE REGISTER CAUGHT AN INCOMPLETE CHANGE OF MINE: adding rows raised S2 EXPECTED-BUT-ABSENT
twice, because the per-site manifests are DERIVED from the matrix and I had not re-rendered them.
Rendered rows appended to both; S3 render drift clean, findings back to the pre-existing 7.
Standing forward requirement recorded in the register rather than as a runbook comment: the
pinned secrets-storage solution must include a documented process and procedure for certificate
and credential backup, and this step is one of the steps that must be folded into it. The
jumphost creds folder is the interim home only.
Still not backed up and out of scope here: the two inner terraform.tfstate files on the headend
-- gitignored, untracked, state-of-record for 20 VMs, and named as owed to the same backup set
that has just been shown not to exist.
Register 97 rows / 24 octavia rows. Gauntlet ALL GREEN (89), repo-lint 0 fail / 621 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

E3 resolved: declare the two artifacts that persist, delete the two intermediates at source
...
Operator-approved disposition after review. The four undeclared per-DC generator outputs were
separated by whether anything will read them again.
DECLARED. The controller certificate, because octavia-pki.sh verify reads it for the subject,
chain and SAN assertions -- a gate that depends on an undeclared artifact is incoherent. And
the CA serial file, because it is issuance STATE rather than residue: delete it and the next
-CAcreateserial starts a fresh sequence, so the same CA can issue a duplicate serial.
Reissuance is scheduled, not hypothetical -- the controller cert is 2-year and F9's D-106 work
will want new SANs.
DELETED AT SOURCE. The signing request is spent the moment the cert is issued. The openssl
config is fully derived from $DC plus the VIP overlay, so the repo already determines it, and
verify now asserts the SAN set on the CERT, which is where the F8 failure actually shows.
WHY NOT DECLARE ALL FOUR, which was lower-risk and less work: a register that accumulates rows
for artifacts that should not exist trains the reader to add a row rather than ask whether the
file belongs -- the register-as-rubber-stamp failure, the same family as this project's
false-green problems. Concretely, both intermediates are 664, so declaring them would have
produced four MORE world-readable findings for files nothing reads.
Also fixed at source, beyond the three requested steps and flagged as such: the generator now
chmod 600s all three certificates, so the next DC is correct by construction instead of by a
remembered follow-up.
Register 95 rows / 22 octavia rows; schema, mint-ref and notes checks clean; no new findings
(still 7 on the jumphost, the same pre-existing set). verify A2 extended 9 -> 10 so the register
and the gate agree on what should exist.
Harness 17/17 (was 15). T16 proves a missing serial FAILS. T17 proves a workspace with the
intermediates DELETED still PASSES -- without it, adding the serial to A2 would have silently
failed every clean workspace.
Gauntlet ALL GREEN (89), repo-lint 0 fail / 621 files.
Still owed, operator-executed: the one-off chmod on the six existing certs and removal of the
four existing intermediates. The guard blocks both for the session -- including the chmod, which
is worth recording as fork-1 evidence since it blocks an operation that strictly improves
posture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-29 |

Execute the ruled generation host: Step 1.0-GEN and the register move to the headend
...
Dependent work for the D-109 ruling note (b), "Generate on voffice1 (Recommended)". Repo-side
only -- no PKI has been minted.
All seven RUN markers inside Step 1.0-GEN now say voffice1, with $REPO stated to mean the
HEADEND clone. The 18 Octavia register rows move to host-role=headend, the Octavia entries in
vm-secret-locations become headend, and host-identity binds headend -> voffice1.
The Octavia locations are declared `local` rather than with the voffice1 ssh-target on purpose:
that routes them through the F6 host binding, so the checker refuses to measure them from the
wrong machine instead of quietly probing whatever filesystem it is standing on.
Measured both ways:
on vcloud -> all 8 locations report NOT PROBED, naming the host they actually live on
on headend -> probed for real, reporting not-yet-minted
Verdict unchanged at 7 findings on vcloud; tests/creds-matrix 65/65.
Consequence recorded rather than discovered later: P5's octavia rows are now judged ONLY on
voffice1, so the credential half of the Stage-5 entry gate must be read on the headend -- the
same host P3/P4 already require. A narrowing of where preflight is authoritative, and a direct
consequence of the ruled generation host.
repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Dangling-reference sweep: one dead pointer, and the retired DC1/DC2 spelling in two places
...
Swept every repo-relative path referenced by this session's changed files against the tree.
Most hits were regex artefacts from wrapped prose; two were real.
bundle.yaml pointed the Octavia PKI at runbooks/01a-octavia-pki-generation.md, which does not
exist and has no history in this repo. The generator is phase-01-bundle-deploy.md Step
1.0-GEN, now run once PER DC with DC exported. A reader following the old pointer would have
found nothing and had no way to know whether the step existed elsewhere.
The phase-6 hostnames-overlay proposal still used the dc1/dc2 spelling that item 3.1 retired
this session -- the exact ambiguity 3.1 measured, where "DC1" meant vr1-dc0 on one surface and
vr1-dc1 on another. Region-qualified, with the reason recorded inline. The surrounding text is
deliberately left as-is otherwise: it correctly states the files do NOT exist and must never be
passed to --overlay, which is the honest framing for a proposed-but-unauthored pattern.
provider-bundle-check still PASSes for both DCs; repo-lint 0 fail / 618 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

3.8 fixed per-DC -- and the fix exposed two repo records naming the wrong NIC
...
bridge-interface-mappings is out of bundle.yaml and per-DC in overlays/vr1-dc0-machines.yaml
(new) and vr1-dc1-machines.yaml. The machines overlay was chosen over the vips overlay for a
measured reason: both *-vips.yaml are marked GENERATED / do-not-hand-edit, so a key added
there is clobbered on the next render, and preflight.sh already assembles ${DC}-machines.yaml.
Creating the dc0 file also closes NEW-3 from the other side -- preflight's [ -f ] guard was
silently skipping a file the deploy needs.
THE WORSE DEFECT THE FIX EXPOSED. dc0's compute provider-public MACs are 52:54:00:8c:2a:8c /
52:54:00:50:48:88 -- position [1] in each node's macs list. Audit register row U5 and the
D-124 correction both name the [0] values, which are the PXE/boot MACs, and D-124 further
tells the reader to source them "from lib-hosts.sh" -- verified independently: lib-hosts.sh
declares only HOST_BOOT_MAC and cannot supply a provider MAC at all. Following either record
would have bridged br-ex onto the boot NIC: not a missing mapping that fails loudly, but a
wrong one that comes up and misbehaves. Confirmed three ways -- substrate main.tf:109-116 list
order, the macpin plan capture, and live MAAS showing br-ex over enp2s0. The correct values
and the [0]-vs-[1] trap are now recorded in the overlay, next to the data.
NEW-1: 10.12.12.0/22 was named as the geneve/DATA plane; that is metal-internal, data-tenant
is 10.12.16.0/22. Deliberately not swept repo-wide -- the 10.12.12.5x VIP legs are correct
metal-internal values.
NEW-2 phantom vr0-dc0-testcloud.yaml removed. NEW-4/6 stale octavia-pki name and the phantom
${DC}-hostnames.yaml removed (juju errors on a missing overlay). NEW-7 scale-up removed from
the VIP deploy command per R6, ${DC1_MODEL} -> ${DC_MODEL}.
R5/D-106 BUILT: phase-6 Step 5 rewritten configure-not-deploy, the unsatisfiable diff gate
replaced by a presence-and-state assertion, and the "no designate block exists" claim
corrected against bundle.yaml. os-public-hostname stays unset, per the ruling's refused
option (c).
Verified independently rather than from reported output: provider-bundle-check PASSes for
both DCs in preflight P2's own overlay order. repo-lint 0 fail / 617 files.
Logged not fixed: provider-bundle-check treats an absent bridge-interface-mappings as a silent
skip and never checks the MAC prefix against --dc, so it could not fail on either variant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Items 3.9 + 3.10: the teardown false clear existed in three places, not one
...
3.9 was the readiness audit's most dangerous item -- the teardown runbook is what an operator
reaches for DURING a failed Stage 5, under time pressure.
The false clear was in THREE places, not the one the audit named. Besides Step 2's vm-host
read -- which can never return, since this repo uses per-machine power_type=virsh and
instantiates no maas_vm_host module -- the same wrong premise sat in the "READ BEFORE ANY DC
TEARDOWN" header telling the operator to remove the maas-vm-host record, and in the
Relationship-to-D-061 claim that no VR1 DC had reached Stage 4. The header instance is the
consequential one: it is read FIRST, so it bypassed any fix confined to Step 2.
Step 2 is now a two-lens MACHINE census run from the headend (maas is measurably absent on
vcloud): lens 1 enumerates what exists and ends in a countable RECORDS REQUIRING ATTRIBUTION,
lens 2 corroborates against lib-hosts pinned boot MACs and exits 1 on any hit. Demonstrated
three ways against a fixture -- records present -> exit 1, zero -> exit 0, empty roster ->
REFUSE -- so it is a gate rather than a formality. The 2026-07-21 pod-cascade precedent is
retained as the reason associations are read before any destroy.
Step 3's six phantom module.dc1_* targets are retired; all 8 targets now resolve, verified
independently against ^module "X" across all three roots. The real insight: scoping a DC is a
ROOT choice, not a -target choice, since each substrate root holds exactly one site.
Step 4's VERIFY moved to qemu+ssh from the headend behind a virsh version REFUSAL, because an
unreachable URI, a stopped VM and a bad key otherwise return the same empty result. The
decision tree gains "no branch reaches a destroy without Step 2 passing" -- it never mentioned
the MAAS gate -- and the virsh destroy vs tofu destroy verb distinction. Also fixed: Step 1
backed up the wrong state file, "Two paths" pointed twice at a nonexistent Step 6, and $REPO
silently meant two different clones.
3.10: item 17 CLOSED with measured evidence, and its own stated fix corrected (it closed by
D-125 bridge-in, not the replication the entry claimed). Item 19 disambiguated 19a/19b rather
than renumbered, because both are cited by number from outside the file. Item 20 MEASURED, not
asserted -- verdict no leg required, with the rule mismatch written in rather than resolved
silently, plus an expiry condition; the voffice1-side reboot durability recorded as UNMEASURED
with the commands that would resolve it.
repo-lint 0 fail / 615 files, both files ASCII+LF.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Phase-3 batch re-measured and 13 items delivered into the Stage-5 runbook
...
The 21-item Phase-3 remediation batch was written 2026-07-27 and logged-not-executed.
Re-verified against HEAD: 2 FIXED / 19 REMAIN / 0 SUPERSEDED, plus 5 NEW. Zero superseded
is load-bearing -- ruling 3 CREATED 3.3's hazard rather than retiring it, and the
no-DC-ordering ruling promoted 3.8 to live rather than closing it.
Two measurements corrected the readiness doc instead of inheriting it: 3.6's "21
non-selector consumers" is really 23 consumers / 8 selector-calling / 15 not / only 5
carrying DC-dependent values (the doc counted grep hits, including four files that never
source lib-net), and the D-133 guard is already satisfied at lib-net.sh:175.
Delivered into runbooks/dc-dc-phase4-juju-bundle-per-dc.md (434 -> 929 lines): 3.1 (the
DC1/DC2 namespace, which meant vr1-dc0 in one place and vr1-dc1 in another), 3.2, 3.4
(controller-tag constraint derived from maas-role-tags.sh, gated on exactly one matching
machine), 3.11 phase-4 side, 3.13, 3.14 (asserts the ovs-vsctl observable, refuses on
no-encap), 3.15, 3.16 (four VERIFY-LIVE gates), 3.17 (dry-run now precedes the deploy it
gates), 3.18, 3.19, 3.20, 3.21. Per-DC octavia overlay name threaded through to match F1.
Three items deliberately NOT claimed as fixed: the dc0 apt-mirror model-config key rests on
a single repo comment with no client here to verify it; the juju create-backup flag shape is
corrected only where established and marked unverified-at-authoring, with the gate moved onto
the resulting file; and whether --unit <app>/leader resolves for a subordinate could not be
established, so the ovn-chassis step probes a named unit per existing repo precedent.
NEW-6/7/8 logged not fixed. A human read of the expanded runbook is owed -- it is lint-clean
and its claims were re-grepped, but length is not correctness.
repo-lint 0 fail / 615 files, zero non-ASCII.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

F1+F4 BUILT: the Octavia PKI generator is per-DC end to end, and the gitignore glob follows it
...
F1. The D-109 amendment rules per-DC independent Octavia CAs. R7 parameterised the CA
subject and the VIP gate but left every PATH fixed, so generating a second DC's PKI
overwrote the first DC's issuing-CA key, controller-CA key and both passphrases, and left
one fixed-name overlay carrying the wrong DC's CA for any later redeploy to pick up.
Step 1.0-GEN is now per-DC end to end: DC/DC_LABEL/REPO/VIP_OVERLAY/OCTAVIA_PKI_OVERLAY
export once in a rewritten 1.0-GEN.0; WORKDIR="$HOME/octavia-pki/$DC"; every WORKDIR
re-derivation and the overlay write carry ${DC:?} so an unset DC refuses rather than
silently reusing the old shared path.
Two gates added that did not exist:
- REFUSE-IF-PRESENT: aborts if this DC already has PKI material. Regeneration invalidates
every amphora already issued against that CA, so it must be a deliberate act rather than
the default outcome of re-running a step.
- F4 gate: `git check-ignore -q "$OUT"` immediately before writing key material, asserting
the ACTUAL ignore decision for that path rather than a pattern someone has to remember.
F4. .gitignore:40 pinned the single exact path overlays/octavia-pki.yaml, so the per-DC
rename would have left a file holding CA key blobs plus a plaintext issuing-CA passphrase
UNIGNORED and committable, in a repo SEC-004 records as PUBLIC. Widened to
overlays/*octavia-pki.yaml and proven BOTH ways, not assumed: all per-DC and legacy names
read IGNORED, and a negative control (overlays/vr1-dc0-vips.yaml) reads NOT ignored, so the
glob is not swallowing tracked overlays.
Also fixed, same dc0-freeze: Step 1.3's VIP guard was a hand-rolled grep triple anchored on
the literal 10\.12\.4\. -- it hard-aborted on dc1, counted IPv4 only so R2's ruled v6 legs
were guarded by nothing (chain-audit finding 22), and its prose said 11/11/0 while its code
demanded 13. It now calls provider-bundle-check.py on the merged per-DC input, which already
encodes the bands, the count and the 2026-07-28 dual-family arity coupling.
The 2026-06-03 as-built line keeps the old command verbatim: it is history, not instruction.
repo-lint 0 fail / 615 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Chain-audit deploy path: DC-aware preflight, dc1's real overlay set, Fork-3 role tags
...
Findings 16-20 from the D-136 chain audit, actioned.
preflight P2 is DC-AWARE and validates the deploy's ACTUAL overlay set. It
hardcoded --dc vr1-dc0 with no dc1 branch while dc-dc-phase4 invokes it BARE
as the dc1 gate and tells the operator to expect PASS -- so a dc1 operator
got a green gate that validated dc0's numbers against dc0's bands and never
parsed a dc1 artifact. It also passed one overlay where the deploy passes
two. DC=vr1-dc1 now merges vips + machines for that site.
dc1's deploy command names ALL its overlays. phase-4 Step 4 named only the
vips overlay, so a dc1 deploy run as written merged a machines block still
saying tags=openstack-vr1-dc0 -- grep for vr1-dc1-machines across runbooks
and scripts returned ZERO, i.e. that overlay was in no executable path. The
step now carries the full command, states dc-ha-scaleup is deliberately
excluded per R6, and adds the --dry-run machines-merge confirmation the
overlay's own VERIFY-LIVE note required and no step performed.
phase-6 deployed with NO VIP overlay and named a file that never existed.
Both blocks -- including the real apply against a live model -- passed
overlays/${DC}-hostnames.yaml, never authored, and no VIP overlay;
post-ruling-3 that is a merged input with zero VIPs including the designate
those commands exist to add.
scripts/maas-role-tags.sh ships Fork 3 (harness 9/9). Every machine
constrains on two tags; only the site tag is ever created, so the three role
tags exist nowhere and juju's allocation constraint matches no machine.
Role is DERIVED from two signals that must agree -- the ruled D-134 octet
band and the host's own name token -- and disagreement REFUSES rather than
picking one. Nodes match by pinned boot MAC, never hostname. Dry by default,
every write read back, absent CLI distinguished from unreachable service.
apply --commit is a live MAAS mutation and is NOT run here.
OWNED, third appearance of the IFS trap: ROLES="control compute storage"
under IFS=$'\n\t' does not word-split, so the script reported one absurd tag
named 'control compute storage'. Caught by its own harness pre-ship. Same
class as commit 1's CHECK 1 death and phase-05's aborting captures -- these
strict-bash gates should use arrays, never space-separated string lists.
Also fixed: the role-derivation refusal swallowed its own diagnostic.
Gauntlet ALL GREEN (87) on vcloud; repo-lint 0 fail / 614 files; preflight
verdict unchanged in kind.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

D-109 ruling note: the Octavia controller cert gains IPv6 IP SANs
...
Operator ruling (GA-R5), exact utterance: "Yes, add the v6 IP sans."
Recorded as a D-109 RULING NOTE (2026-07-29); OPS under GA-R3 -- a
generator change extending an already-ruled principle to a surface D-109
governs -- so no new D-number and next-free stays 138.
The generator now emits IP.1 = this DC's provider v4 leg and IP.2 = its
provider v6 leg, both derived from the same per-DC overlay under R7's $DC
selector. Verified: dc0 -> 10.12.4.57 + 2602:f3e2:f02:11::57; dc1 ->
10.12.64.57 + 2602:f3e2:f03:11::57; and a v4-only control emits NO v6 SAN,
so the generator stays correct on a tree where the dual-stack ADD has not
landed. Admin and internal legs stay EXCLUDED exactly as before --
DOCFIX-067's design has always been provider-leg-only.
Two repo guards fired and both were right: repo-lint L5 rejected a heading
that LED with the D-number (it reads as a second definition -- the third
time this trap has bitten on this branch), and L10 rejected the change for
touching a design-decisions Status line without CURRENT-STATE in the same
commit.
Observation logged, not acted on: amphorae reach the controller on o-hm0's
charm-generated fc00::/64 ULA, not on any VIP leg, so neither SAN matches
that path. Measured upstream, the controller verifies the amphora by UUID;
the reverse direction was not established from source. Settle by inspection
at the Octavia step rather than assuming.
Gauntlet ALL GREEN (85) on vcloud; repo-lint 0 fail / 611 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Octavia: R8a BUILT (record was false), two stale surfaces, R7 generator de-frozen
...
Follows an upstream/vendor research pass answering "can Octavia run full
IPv6" -- yes; nothing inside Octavia requires IPv4. charm-octavia creates
the lb-mgmt subnet with ip_version: 6 and has no IPv4 code path at all; the
health manager is v6-capable both directions; the amphora agent binds '::'
and verifies certs by amphora UUID, not IP; and Nova metadata is not a
dependency (config_drive=True). Surviving v4 pressure is soft or external.
R8a BUILT. phase-05-octavia-verify.sh now compares o-hm0's MTU against
lb-mgmt-net's (resolved BY TAG, never by name), fails in either direction,
and REFUSES when a value cannot be read. This closes the LP#2018998
obligation, live on our exact pin after its 2025-12-31 recurrence.
THE DECISION RECORD HAD CLAIMED THIS SHIPPED. design-decisions.md carried a
present-tense "Delivery: the assertion ships in ..." written on the day of
the ruling, while the file held zero MTU references and its last commit
predated the ruling by a month. Corrected in place. Sharper than the usual
ruled-but-not-built class: the decision doc was the false witness.
TWO OWNED ERRORS:
- I concluded "no octavia harness exists" from a NAME grep. The script's
harness has always been tests/phase-05. Duplicate deleted, R8a cases
folded into the real one.
- My first MTU draft ABORTED the script: a $( ) under set -euo pipefail with
inherit_errexit exits the run rather than reaching the refusal branch. The
pre-existing harness I had just declared nonexistent caught it. Same class
as commit 1's IFS word-splitting trap; both now locked by cases.
Two live surfaces still contradicted R8 -- dc-dc-phase4 and the workflow doc
both called lb-mgmt IPv6 "a real, open risk" and recommended keeping it
v4-only, citing the two refuted bugs. A reader would have re-litigated a
closed ruling in the wrong direction. Both superseded in place.
R7 executed: phase-01 Step 1.0-GEN gains a $DC selector; the baked
/CN=VR0 DC0 CA subjects derive from it, and the VIP gate derives this DC's
provider prefix from the same overlay it read the VIP from -- R7's explicit
"read the MERGED input" caveat. dc1's 10.12.64.57 used to hard-ABORT, so no
dc1 artifact could be produced. Verified both DCs, negative control still
aborts.
Logged not built: the controller cert's CN/DNS SANs still carry dc0.vr0
(outside R7's ruled scope, inert while os-public-hostname is unset), and it
gains no IPv6 IP SAN though the VIP is dual-family -- a gap, not a break,
and adding v6 SANs needs its own ruling.
tests/phase-05 14 cases (was 9); gauntlet ALL GREEN (85) on vcloud;
repo-lint 0 fail / 611 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Ruling 3 commit 2: dual-stack ADD -- R2 and R11 are now BUILT, L3-9 resolved
...
Both per-DC VIP overlays re-rendered dual-family: every app carries
prefer-ipv6: true and a six-address vip, and vault .61 + designate .62 are
built. Measured both DCs: 13 clustered VIPs (13 dual-family) + 12 hacluster
principals all carrying a VIP -> PASS. With dc-ha-scaleup stacked: 13
principals, PASS -- R6's ruled ordering (VIPs before the HA overlay) is now
satisfied and executable. Two ruled-but-never-built decisions became
artifacts in one pass.
L3-9 RESOLVED. dc-dc-ipv6-family-matrix.yaml is narrowed to ceph-mon only;
its ten duplicate vip + prefer-ipv6 pairs (every value an unrendered token)
are gone, removed TOGETHER -- removing the vips alone would have recreated
the defect deliberately. Both merge orders now PASS, i.e. order-independent,
which is the actual fix. ceph-mon's ULA-only networks survive because they
exist in no other file. The two Launchpad citations arguing octavia lb-mgmt
IPv6 was an open risk are removed: both were read in full and both failed,
and R8 closed that question on 2026-07-27.
Also corrected at source: CURRENT-STATE described L3-9 as "the dangerous
merge order is the one that PASSES". Measured -- that stopped being true
when invariant 9 shipped on 2026-07-28. The merge semantics it describes
remain accurate; only the green-over-loss conclusion is retired.
Gate repairs the agents caught BEFORE the edit:
- phase-01's deploy guard would have ABORTED THE DEPLOY AGAIN -- its HI regex
(5[0-9]|60) excludes .61/.62, so it would have read 13/11/0 against a
required 11/11/0. Widened, counts moved to 13, verified 13/13/0 -> DEPLOY.
Second near-miss on this one guard across two commits.
- pre-flight-checks CHECK 1 hardcoded if(n!=3) and would have failed all 13.
Now takes a triple OR a sextet and asserts a sextet's last three legs are
really v6. Proven able to fail on both new paths.
Harnesses re-pointed, never deleted. T16/T17 asserted R11's gaps were OPEN;
inverted to the surviving invariant and PAIRED with new T16b/T17b that remove
a VIP and demand the check still fires. T20/T21 had become no-ops that would
have gone GREEN testing nothing -- inverted. A v4-only twin fixture keeps the
pre-dual-stack cases testing their own invariants. The renderer harness now
proves reproduction on both shapes, including new T3b against the live
dual-family overlay.
provider-bundle-check 33/33; render-dc-overlays 18/18; render-baseline 10/10;
preflight 16/16; gauntlet ALL GREEN (85) on vcloud; repo-lint 0 fail / 611.
PREFLIGHT: BOTH STANDING REDS CLEARED, 3 fatal -> 2 fatal. The remaining two
are the deliberately-absent octavia-pki overlay and MAAS unreachable from
vcloud, which R10 measured clears on voffice1.
FLAGGED, not hidden: octavia's dual-family API VIP is an inference from R2
applied evenly, not a quoted ruling. Worth one operator confirmation.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|