| 2026-08-02 |

Queued-findings backlog: sweep F1 + F6 CLOSED, F2 diagnosed, F3/F4/F5 graduated, DOCFIX-207
...
Works the 2026-08-02 close sweep (docs/audit/queued-findings-20260802-stage5-edge-fold.txt)
with three read-only agents in parallel. Cites the SWEEP register (F1-F6); the runbook fold
register has its own F1-F12 and is untouched.
sweep F1 CLOSED -- the dc0 rack's staged deploy input matches the repo again. One gated scp
of overlays/vr1-dc0-vips.yaml, ed19d989e80da8da -> 80d861560a6b3c52. A single-file copy was
provably sufficient because the WHOLE staging dir was enumerated first: 14 files, exactly 1
diverged, 0 missing from the Step-4 deploy closure (policies/overrides.zip present at the
repo digest), 0 orphans. All 14 re-verified against CURRENT HEAD after the copy. The 0600
octavia PKI overlay is untouched (same digest/mode/mtime) and was hashed, never read.
Repo-side correctness MEASURED not inherited: 0 ULA legs, 39 GUA, render-drift 4/4 naming
the file with a proof-of-teeth case.
sweep F6 CLOSED, PASS -- the dc0 MAAS region DB is proven uncorrupted. pg_dump read every
page of every table in maasdb: 23,878,796 bytes / 37,199 lines / exit 0 / empty stderr,
completion marker asserted separately. F6's own diagnosis was wrong: snap confinement does
not reproduce as ubuntu. The discriminators are ROLE (maas, not ubuntu) and TRANSPORT (over
TCP the role is password-challenged, over the unix socket it needs no credential). Both
identity values now measured -- maasdb had been prose. Dump streamed, nothing persisted.
sweep F2 DIAGNOSED, not fixed -- reading R1 is true, R2 refuted. debmirror prints "All done."
then exits non-zero; confirmed at vendor source and re-verified independently here
(debmirror 1:2.39ubuntu2, 1615 say("All done."), 1620 exit 1 if !$ignore_small_errors).
Cause: one 500 read timeout on a jammy-backports dep11 index. TWO corrections to the sweep:
its "the log says it succeeded" quotes all came from the PASSING UCA leg; and "stays RED"
overstates it -- measured 16 runs, 7 finished, 9 failed, with four fail-then-succeed pairs
hours apart. No remedy applied: debmirror's exit code conflates "nothing mirrored" with
"mirrored minus N transient files", so any tolerance change alters what the gate attests.
DOCFIX-207 -- preflight P6 quoted "50 apps / 97 relations"; measured is 56 / 108. The same
figures were corrected in phase-01's own gate on 2026-07-10 and this copy was missed.
sweep F3/F4/F5 graduated from the audit capture to durable homes: systemctl show fabricating
Result=success for a non-existent unit -> platform-traps 5c; assert the harness case count
moved, and two scripts probing one endpoint must share the probe definition -> script-authoring.
Logged NOT fixed (hard rule 1): preflight P2 validates a merged input including
vr1-dc0-machines.yaml while phase4:527-528 says the file does not exist and Step 4 does not
pass it -- a gate grading a different artifact than the deploy consumes, on the Step-4 path.
And security-ledger SEC-029(3) calls ~/repo-stage a nine-file copy; it is fourteen.
session-ledger machine-derived block re-seeded (was the 2026-07-27 seed: 21 SEC / D-138;
now 28 SEC / D-141 / DOCFIX-208, each added SEC row verified against the register).
Gates: repo-lint 0 fail / 1 warn (legacy carve-out); run-tests-all ALL GREEN (97 harnesses,
count unchanged, so nothing moved silently); tests/preflight 43/43; tests/render-drift 4/4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

SESSION CLOSE 2026-08-02 (GA-R4 bookend): dc0 edge rebuilt, D-139 steps 1-3, fold opened
...
Durability triad caught a live deploy hazard: the dc0 rack's ~/repo-stage VIP overlay is
STALE (ed19d989 vs repo 80d86156) because this session re-rendered it onto GUA and nothing
propagates to a directory with no git. Under D-138 that rack IS the deploy client, so a
deploy from it would place 26 VIP legs on the ULA prefixes D-139 retires. The 08-01 sweep
recorded these digests matching, so that note now misleads. Logged, not fixed -- first
item next session.
Gates: repo-lint 0 fail / 1 warn (654 files); gauntlet ALL GREEN (97); ledger-scan 3 open
decisions, 28 open SEC, next-free D-141 / DOCFIX-207 / BUNDLEFIX-053. Reconciled against
what the session claimed: D 140->141, SEC 26->28, manifest 96->97.
Mirror sync re-ran and the content IS current (both Release files rewritten today, UCA now
succeeds so egress is genuinely fixed) but records FAIL: the journal says "Everything OK...
All done." then systemd status=1/FAILURE. dc-mirror.sh check stays red on an exit code that
disagrees with the sync's own verdict. Not relaxed to clear it.
Sweep: docs/audit/queued-findings-20260802-stage5-edge-fold.txt, 6 FIRST SURFACE, each
grep-proven absent first. Memory review clean; instrument-currency gains instances ten
through twelve. Ledger rotated 283 -> 266, now 286; no orphaned session. No stage opened
or closed -- Stage 5 remains OPEN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Pre-deploy items 1-3: preflight re-run, mirror sync running, dnsmasq settled on measurement
...
(1) Preflight for dc0 with the region profile: EXIT 1, exactly 11 [FAIL] lines, all the
ruled-accepted P5 set. Zero new findings, so no fresh GA-R5 exchange is owed. P8 clean
after this morning's guard fix; P9 warns correctly off-rack.
(2) dc0-mirror-sync re-triggered and running. The 08-02 failure was upstream-unreachable
during the edge outage; content was verified intact throughout. dc-mirror.sh check stays
red until the status file reads OK -- the attested-currency gate working as designed.
(3) The edge dnsmasq question settled by measurement, not reasoning: vtnet0 carries
inet 10.12.4.1/22 and inet6 LINK-LOCAL ONLY, no radvd. The v4 192.168.1.x range is not on
the interface subnet, and the v6 constructor:vtnet0,slaac range has no global prefix to
construct from. Inert today.
The residual risk is the v6 half, not the stale v4 one: provider-public has a GUA prefix
OpenStack will use, so if vtnet0 ever gains a global v6 the constructor range
self-activates and the edge advertises SLAAC on the provider segment. A constructor range
follows the interface -- it is not stale config. Logged, not changed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

dc0 edge REBUILT, egress restored (8/8); SEC-032; two dc-egress-check defects fixed
...
Agent-executed rebuild. tofu -replace asserted on `tofu show -json`: exactly 2 non-no-op
changes, both delete,create, both naming vr1_dc0_opnsense, 0 not naming it -- and the
POSITIVE half asserted too, because a plan replacing only the DOMAIN would have reattached
the corrupt qcow2 and passed a "nothing extra" check. Artifact verified, not the log.
Convergence re-plan clean. pfctl -s nat explicitly verified to show real outbound NAT --
the exact defect dc1 is sitting in.
MY BRIEF CARRIED A WRONG BASE-IMAGE PATH. I cited the module default; d124-inner.auto.tfvars
overrides it to a path on voffice1. The apply destroys the volume before recreating, and
neither edge disk has a backingStore, so a wrong path would have left no disk and no
rollback. The agent verified before mutating. Same guess-instead-of-look failure the
operator called out earlier today, in a brief written to prevent exactly that.
dc-egress-check: A4's snap probe omitted `Snap-Device-Series: 16` (400 without it through
the proxy AND direct; 200 with) -- a false FAIL on a healthy proxy. A3 now names the
forwards-without-translating signature when A2 passed. Harness 14 -> 16; T15 asserts on
the CODE after a first cut passed on the comment explaining the header.
Three findings needing attention pre-Stage-5: the rebuilt edge runs dnsmasq on udp4 *:67
over vtnet0 with a stale 192.168.1.x range on the segment carrying juju .5 and MAAS .6
(measured; inertness reasoned, not measured); SEC-021(a) answered -- the off-jumphost copy
exists; and a rebuilt edge invalidates the rack's known_hosts, which no runbook mentions.
Gauntlet ALL GREEN (97); repo-lint 0 fail / 1 warn. Open SEC 28.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

dc1 edge: forwards but does not TRANSLATE; config intact; REPAIR not rebuild (SEC-031)
...
Read-only agent assessment, docs/audit/dc1-edge-assessment-20260802.txt (643 lines).
Decisive measurement: simultaneous tcpdump on BOTH edge taps while the rack pinged
1.1.1.1 -- the same packet appears on the LAN tap and 0.8 ms later on the WAN tap with
source STILL 10.12.64.2. vcloud's virbr1 masquerades only 172.30.3.0/24, so it is never
translated and no reply returns. Routing works, pf is not blocking, there is no outbound
NAT. My "config lost" hypothesis was half right and the half that mattered was wrong.
Config INTACT on positive evidence: fsck names every inode it deletes, all 50 enumerated,
/conf/config.xml not among them; the edge booted onto its ruled as-built addresses, not
the factory ones. Qualifier: 92 further casualties are inode-only, so strong not absolute.
What fsck destroyed is the FreeBSD base-system user DB -- /etc/master.passwd and
/etc/group, plus BOTH /var/backups copies. Hence "Configuring firewall.....failed." on
every boot since, no pf ruleset, and SEC-031: the edge is an open router serving its GUI
to the simulated ISP. Confirmed against a control (office1 edge: WAN ICMP 100% loss, GUI
000; dc1: 0% loss, GUI 200 in 0.014s). Bounded by lab topology; D-125 unaffected.
Repair not rebuild -- rebuilding discards a good config and re-incurs D-112(c)+D-113.
Blocked on there being no working credential path (sshd down, authed API hangs, console
login likely broken by the same passwd DB); repair needs boot -s, a mutation, so the
read-only agent stopped.
Two findings that outlive this: dc-egress-check's A2 is structurally blind to
forwards-but-does-not-translate and needs a positive "the edge TRANSLATES" assertion --
a defect in a gate built earlier the same day. And NEITHER edge had ever been
reboot-tested; 2026-08-01 was the first time dc1's as-built config ever booted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Fold: open the register, close both Class-A rows (F1 phase4, F12 SKILL.md)
...
Operator ruled fold-before-exercise, start at Phase 2, all DCs come up in their own
region. Register: docs/runbook-fold-register.md, 12 rows classed A/B/C.
Measured gap: D-138 and D-139 appear in NO runbook. Nothing in the chain builds a per-DC
MAAS region, a snap proxy, or the GUA carve. Running it as written would rebuild the
pre-D-132/D-138/D-139 shape and fail where we already fixed things by hand.
F1 (Class A): phase4 asserted "Both DCs deploy from the Office1 headend by the SAME
procedure" and "Every juju and maas command runs there". The 2026-07-27 ruling it cited
was about DC ordering, not run location; D-138 reversed the location half. Replaced with
a tool-split table; superseded text struck in place, not deleted.
F12 (Class A, worse): SKILL.md -- loaded by every session before any runbook -- had ZERO
mentions of D-138 while stating Plane 2 executes on voffice1. Now carries the split, the
structural reason, and the measured test: juju controllers on voffice1 -> "No controllers
registered"; on the dc0 rack -> the live vr1-dc0-controller.
Both record the racks-have-no-repo-clone trap and the sha256-verify-staged-copies rule.
A row closes on the edit AND on being exercised by the dc1 Phase-2 rebuild; an edit alone
is a ruled-but-not-built claim.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

GA-R5 x2: D-135 amended (dc0 converges on the proxy at rebuild); D-140 PINNED (tofu/juju)
...
D-135 amendment. Operator: "if we have to rebuild in dc0 for any reason we will be using
a proxy rather than a full mirror rebuild." The 953 GB mirror is not reconstructed; a
rebuild comes back as an apt caching proxy on the dc1 pattern. This qualifies D-135's
standing "no fallback and no later convergence" -- none while the mirror stands, but a
rebuild converges. Nothing changes today: the running mirror is verified intact and stays.
Owed at the rebuild: capture the mirror arm's outcome BEFORE dismantling it, or the
experiment D-135 existed to run produces nothing. Consequence for the fold: the phase
runbooks can carry ONE artifact strategy instead of branching per DC.
D-140 PINNED. Operator: "I accept the opentofu plan. Hardened dc0 deployment then
opentofu management with a successful tested deployment method from dc0. Something to pin
for the end of deployment, review and add opentofu management to the remaining tasks."
Order: dc0 deploys and hardens -> method tested -> proven method is the translation input
-> reviewed at end-of-deployment. Not converted now because the bundle deploy has never
completed end to end and changing the mechanism first means two unknowns at once.
Value is measurably in the model/config layer, not topology. Cost to price at review: the
provider is resource-shaped not bundle-shaped, against a measured 56 apps / 108 relations.
Provider capabilities must be read, never recalled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

CORRECTION: the edge break was UFS damage from an unclean power cut, not an update
...
My earlier "partially-applied OPNsense update" diagnosis was an inference from the shape
of the symptom and is refuted by the log: ZERO pkg-static / opnsense-update / firmware
lines in the entire 477 KB serial history.
What precedes the first missing-library error is an fsck salvage of a badly damaged UFS
root -- UNREF FILE x2533, UNEXPECTED SOFT UPDATE INCONSISTENCY x785, SALVAGE? yes x257,
then immediately ld-elf.so.1: libcrypto.so.17 not found. The libraries were fsck
casualties.
Causal chain: the 08-01 memory resize was an IN-PLACE tofu update, which this document
already records BOUNCES the guest (confirmed on dc1 as canary that day). A bounce of a
containment VM is a hard power cut to every inner VM. Under D-127 the DC edges are the
only inner VMs with autostart=true, so they were the only ones hard-cut and auto-
restarted onto damaged filesystems.
Confirmed by the dc1 control -- same event, different severity: dc0 2533/785 with ld-elf
errors, dc1 76/181 with zero. dc1's "healthy edge that does not forward", logged earlier
as a separate thread, is NOT separate -- same incident, lesser damage.
Durable finding is a procedure gap: the resize plan was asserted on content and passed a
capacity gate, and none of it covered the inner guests. "In-place" describes the tofu
resource, not the blast radius. Any future in-place change to a containment VM must shut
down its inner guests first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Close both egress gaps: dc-egress-check.sh + preflight P9, wired into restart + phase-4
...
Gap 1: nothing tested DC egress. dc-rack-net.sh check PASSED through a 19-hour outage
because it asserts legs and units -- a leg is not a path -- and the mirror answered 200
from its own nginx, which says nothing about upstream.
Gap 2: the post-reboot check set was my judgement, and it omitted egress.
New scripts/dc-egress-check.sh: site-keyed, runs on the rack, LAYERED (route -> edge
answers -> traffic leaves, ICMP and TCP -> the three upstreams), reporting the FIRST
failure as the cause while the rest SKIP. A4 asserts the UPSTREAM the local artifact
services sync FROM, which is independent of whether they serve.
Proven live on two different failure modes: dc0 fails at A2 naming the dead edge; dc1
passes A2 and fails A3/A4 -- traffic not leaving a healthy edge.
preflight P9 WARNs off-rack with the exact command, FAILs on the rack when broken, WARNs
if the checker is absent. Gap 2 is closed by the WIRING: restart procedure Stage 0
(before anything fetches) and phase-4 Step 3.9 (before the deploy), both stating P9 is
not a substitute for running it on the rack.
Mutation pass caught a decorative test of mine: T11 never reached A4's unrecognised-code
branch, so flipping it to ok() left the suite green. T14 added to exercise A4 alone.
tests/dc-egress-check 14/14; tests/preflight 39 -> 43/43; 5 mutations across both, each
killing a named case, scripts restored sha256-identical. Gauntlet ALL GREEN (97),
manifest 96 -> 97; repo-lint 0 fail / 1 warn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

dc0 edge root cause: partially-applied OPNsense update, exposed (not caused) by the reboot
...
Console read: the edge sits at a single-user prompt with libcrypto.so.17 and
libpython3.13.so.1.0 missing, 10-configd / 15-templates / 90-carp all failing. The
running system had survived on already-mapped libraries; the 05:49 boot is where it fell
over. Symptom appears nowhere in appendix-A -- a new class.
Scope measured, not assumed: dc0 only. dc1's edge log has zero matches and booted clean
to a login prompt with its WAN addressed. Upstream healthy -- vcloud and voffice1 both
reach the internet, both uplink nets are active NAT, both racks hold default routes.
Recovery material: opnsense-26.7-nano.qcow2 is in the same pool, but the broken disk has
an empty backingStore -- standalone, no CoW base to roll back to.
Logged not conflated: dc1's rack also lacks egress despite a healthy edge; a dc1
forwarding/NAT question, off the deploy path, not diagnosed here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Step 3.5 done; BLOCKER: the dc0 edge has been down since the 08-01 05:48 reboot
...
Step 3.5 complete: model vr1-dc0 created (credential vr1-dc0-cred), spaces gate PASS
0 fatal, apt-mirror verified at http://10.12.8.4/ubuntu. Spaces script run from the
07-30 staged copy on the rack, all three files sha256-verified against the repo first.
Blocker found by dc-mirror.sh check dc0 -- the content-assertion fix from 2026-07-27
earning its keep. The rack cannot reach archive.ubuntu.com, streams.canonical.com,
api.snapcraft.io via the snap proxy, 1.1.1.1, or its own default gateway 10.12.4.1.
ip neigh shows 10.12.4.1 FAILED while .5 and .6 on the same segment answer. The edge
VM is running with both NICs attached, qga not connected, qemu log last written at the
reboot. Last successful mirror sync predates it. Step 4 cannot proceed: the deploy
needs the juju agent stream and snap upstream, both dead.
Two gaps exposed: no gate anywhere tests DC egress (dc-rack-net check passes with it
dead), and this session's own post-reboot verification cited five true facts none of
which tested egress -- a listener is not a path, and a 200 from a local service says
nothing about upstream.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

D-139 steps 2 and 3 EXECUTED for vr1-dc0: the nodes are on GUA
...
Step 2: dc-plane-ipam.sh carve-v6 --commit, applied=5, each GUA /64 on the same vlan as
its v4 twin and read back. 2 errors are correct refusals (lbaas-mgmt and oob have no v4
plane to pair with). Independent read: 11 v6 subnets, 6 GUA + 5 ULA, every GUA paired.
Step 3: dc-node-v6-carve.py replace --v6-family gua --commit, applied=45 skipped=9
errors=0, READ-BACK 45/45. Independent query: PRE 54 v6 links (GUA 9 / ULA 45) -> POST 54
(GUA 54 / ULA 0), and ZERO NICs carry more than one global v6, so G19's sole-global
predicate holds. v4 was not touched -- the ordering step 3 exists to enforce.
Gates after: dc-node-v6-carve check --v6-family gua PASS (54 correct, 0 missing, 0
errors); dc-plane-ipam check 29 pass / 2 fail, both expected absences and neither new
(lbaas-mgmt is step 5, oob has no MAAS plane). Was 7 fail before step 2.
Reversible: replace --v6-family ula --commit. Nothing was deleted; both families remain
in MAAS until step 6 retires the ULA rows. vr1-dc1 untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-08-01 |

GA-R5: VPN (:e0) deferred to Roosevelt design time, recorded not left silent
...
Operator utterance: "Defer to Roosevelt design time".
Recorded as a DEFERRAL rather than left open, because an unrecorded hole is exactly how
the OOB omission survived until the operator caught it. Measured basis: VPN is a
RESERVATION at both precedents, never a carved plane -- VR0 DC0 and Willamette hold the
:e0::/60 parent ONLY, no /64 and no IPv4, unlike OOB which holds /60+/64 at both. VR1's
VPN is Tailscale under D-129(iii), addressed from fd7a:115c:a1e0::/48, which this
deployment does not carve, so a VR1 :e0::/60 would reserve space with no occupant.
Consequence stated plainly: VR1's octet map diverges from VR0 and Willamette on exactly
one hextet, by decision, and ruling B's "conforms to an existing org standard" claim
should be read with that documented exception. Nothing in D-139's execution list, the
Stage-5 deploy, or any gate depends on :e0. Trigger: Roosevelt VPN design, alongside
D-132 and D-131 sub-4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
CURRENT-STATE: record the OOB v4 ruling (d47ef8c landed without this half)
...
d47ef8c committed the design-decisions half of the 10.12.40.0/22 / 10.12.88.0/22 ruling
but not the CURRENT-STATE block: the editing step failed its own assertion and the shell
chain did not gate the commit on it -- the same masking as the earlier | tail red-lint
push. Corrected forward, not by rewriting pushed history.
Records the ruling, the measured pattern (the dc0->dc1 offset is a consequence of dc0's
historical 20-28 gap, not a rule), why that gap is unusable (20+60=80 and 20+48=68 both
allocated), and that the v4 half is ruled-and-not-built and gates nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

D-139 step 2 (input half): VIP overlay onto GUA; both apex readers learn v6 FAMILY
...
overlays/vr1-dc0-vips.yaml is now fully GUA -- 39/39 v6 VIP legs, zero fd50:. The
renderer reproduces it byte-for-byte and provider-bundle-check PASSES with 13
dual-family VIPs.
Both apex readers had to learn family first. Step 1 creates GUA alongside ULA and step
6 retires ULA, so mid-transition every v6-only plane has two /64s under one (role,kind)
key and both readers could only REFUSE -- measured live: derive --dual-family rc 2, and
the gate "cannot evaluate barbican's dual-family vip". render-dc-overlays gains
--v6-family (refuse on ambiguity kept); provider-bundle-check resolves per application
AND per leg, needing no flag at any call site.
A lossy path was found and NOT taken: a full derive drops the 13 per-app comment fields
(4368 -> 3593 bytes). The values file was edited surgically instead: 26 changed overlay
lines, every comment intact. derive being lossy against its own values file is logged.
PROPERTY TRADED, recorded as a loss not a win: the gate no longer catches a wrong-FAMILY
leg -- a ULA leg in the GUA dc0 overlay now PASSES, graded against the ULA band that
exists until step 6. Necessary (dc1 is legitimately ULA) but it leaves D-139 family
conformance checked by nothing. A ruled-table conformance gate is OWED.
Two bugs I introduced were caught by the harness, not review: family inferred once from
the provider leg (GUA in both worlds -- 7 dc1 cases red), then bands resolved once per
bundle instead of per application.
tests/render-dc-overlays 18 -> 23/23 (3 mutations, restore sha256-identical);
tests/provider-bundle-check 55/55; gauntlet ALL GREEN (96); repo-lint 0 fail / 1 warn.
vr1-dc1 untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
D-139 step 1 EXECUTED for vr1-dc0: apex GUA carve pushed, 139 -> 152 prefixes
...
netbox/d139-gua-carve.py --dc vr1-dc0 --commit, exit 0. CREATE 13 | EXISTS 3 |
RETIRE-REPORT 9; READ-BACK 13/13; errors=0.
Verified by an independent API query rather than the tool's own read-back: apex count
139 -> 152 exact, 17 GUA rows scoped vr1-dc0 (4 pre-existing + 13 new) with correct
roles including both oob rows. Nothing deleted or orphaned -- the 9 ULA prefixes remain
and all 26 ULA VIP legs still resolve; retirement is step 6.
Prefixes only: no MAAS subnet, node address or VIP was touched. vr1-dc1 not pushed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Conformance: dc0 deploy inputs vs the D-139 specs -- the VIP overlay FAILS
...
Every ruled row of D-139 ruling A, ruling B, the "B plus C" narrowing and the OOB
amendment compared to the ARTIFACT (live MAAS, apex, deploy overlays), not to another
document. Capture: docs/audit/d139-conformance-dc0-20260801.txt.
BLOCKER: overlays/vr1-dc0-vips.yaml carries an R2 sextet per service, and 26 of its 39
v6 VIP legs are ULA (13 metal-admin :220::, 13 metal-internal :221::). Those 26 are
exactly the 26 dependents the step-1 dry run reports inside the retiring ULA /64s --
the same objects, confirmed by description. Deploying after steps 1-3 alone would place
26 VIP legs on prefixes ruling B retires, while their GUA replacements exist in apex and
MAAS but appear nowhere in the deploy input. Overlay must be re-rendered before Step 4.
Also measured: lib-net.sh contains ZERO IPv6, so step 4's "update the v6 arm" has no arm
to update; and no gate compares D-139's tables to any artifact, so this drift class has
no detector. Provider-public conforms in both families; all nine nodes match MAAS (G19).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

GA-R5: D-139 gains an OOB plane, dual-stack; 10.12.60.0/22 ruled back into force
...
Operator utterances: "make sure to include oob in the dual stack recordings" and
"Dual-stack; rule 10.12.60.0/22 back into force for OOB".
Measured gap: D-139's CARVE table carried 10,11,20,21,30,40,50,80 and no 0xf0 (OOB)
or 0xe0 (VPN) -- design-decisions.md:3066 declared all three out of scope of the
six-plane tool, D-139 brought :80 back and left the other two. The step-1 push would
have written a carve incomplete against the org standard ruling B cites. Held, not run.
The ruling knowingly diverges from both precedents: VR0 and Willamette carry OOB
IPv6-only and a sweep of every apex prefix returns ZERO v4 rows with role oob at any
site. Dual-stack is a deliberate Roosevelt correction (BMC/IPMI is v4 in practice).
Supersedes the "OOB n/a / bare-metal-only concern" row at design-decisions.md:93.
D-058 remains SUPERSEDED in full; only that one row's VALUE is reinstated, under
D-139's authority. Block verified free before ruling. Per-DC v4 split is UNRULED and
deliberately not inferred; the v6 half is complete and the v6-only push is unblocked.
VPN (:e0) is the same omission, left open rather than folded in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
GA-R5 ruling: the D-139 GUA carve is sequenced BEFORE the Stage-5 deploy
...
Operator answer, exact utterance: "Apex push AND the MAAS/node GUA carve, then deploy".
D-139 execution steps 1-3 run before the bundle deploy, so the cloud is deployed once
on its final GUA addresses and nothing is re-addressed underneath it. The 2026-07-30
standing directive is satisfied in substance rather than literally.
Recorded, not inferred: steps 4-6 were NOT put, and step 6 has a deploy coupling this
ruling does not settle -- the per-DC VIP overlays carry v6 VIPs in the ULA range, so
deploying after steps 1-3 alone leaves GUA nodes with ULA VIPs. Own exchange owed
before Step 4. Neither ruling A nor B is touched; no new D-number (GA-R3: OPS).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
L10 correction: cite the preflight capture in CURRENT-STATE (cde7441 pushed lint-red)
...
cde7441 added docs/audit/stage5-preflight-dc0-20260801.txt without touching
CURRENT-STATE.md in the same commit, which L10 (GA-R1 rule 8 / C1) forbids. The lint
was piped through `tail` so its exit code never gated the push -- the same masking
recorded at the 2026-07-31 close. Corrected forward, not by rewriting pushed history.
Records the post-fix re-run: P8 WARNs off-substrate, [FAIL] count 12 -> 11, all eleven
the ruled-accepted P5 set; preflight still exits 1, which both P5 rulings anticipate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

P8 host guard: && -> || (live false FAIL on voffice1); dc0 MAAS path restored
...
P8's guard warned only when BOTH terraform.tfstate and opentofu/.terraform were
absent. voffice1 has no state file but does have .terraform -- D-128 has it run the
INNER tofu roots, and `tofu init` leaves a provider cache -- so neither warn branch
fired, P8 ran a real `tofu plan`, and it died on a gitignored tfvar -> exit 1 -> hard
FAIL on the very host the Stage-5 runbook designates for preflight. terraform.tfstate
is the OWNERSHIP marker; .terraform is only a cache. Both must be present to evaluate.
No test caught it because T38 removes the whole opentofu/ directory, so both markers
vanish together; the fixture itself sat in the defective state and passed only because
the && was wrong. Fixture now creates both markers; new T39 covers voffice1's real
shape (cache present, state absent -> WARN). Mutation: reverting || to && turns T39
and only T39 red; script restored sha256-identical.
Also: the dc0 MAAS path was down -- the vr1-dc0-region profile points at an SSH
forward (127.0.0.1:5241) that died with the 05:48 rack reboot. Restored and proven
(plan vr1-dc0 PASS exit 0, was REFUSE exit 3). Not an SEC-010 puncture: an SSH forward
originates on the rack. voffice1 pulled to HEAD. Preflight re-run: 12 [FAIL] lines,
11 = the ruled-accepted P5 set compared by identity, 1 = P8.
tests/preflight 38 -> 39/39; gauntlet ALL GREEN (96); repo-lint 0 fail / 1 warn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

CORRECTION: no autostart setting was changed, and D-127 already rules it
...
Operator asked directly whether I had changed autostart. I had NOT -- item 21
logged it under hard rule 1 and started the two VMs by hand. The question, asked
to scope a power-settings-matrix session, prompted the check that shows I FRAMED
THE FINDING WRONGLY.
D-127 ("VR1 host-level VM autostart policy") already rules this, autostart is
TOFU-MANAGED, and every measured value is an explicit per-instance decision with
its reason in-line: main.tf:413/:540 containment VMs false ("MANUAL (gated
bring-up), never on host boot"); main.tf:102/:178 office1-opnsense and voffice1
true ("foundational"); DC edges true ("comes up with its containment VM"); node
VMs false ("MAAS-power-controlled"). The module variable's own description forbids
relying on its default. Calling it "a standing exposure" implied nobody had
decided. A session told to "fix autostart" would fight a ruling -- and now that P8
exists, a virsh autostart change would also surface as substrate DRIFT.
WHAT SURVIVES, narrower and real: vr1-dc0-maas-01 and vr1-dc0-juju-01 are created
by the vr1_dc0_node for_each, so they inherit the NODE justification
"MAAS-power-controlled". TRUE for the juju controller (subtle-grouse, a dc0-region
machine, power_type=virsh). CIRCULAR for the MAAS region VM: hot-kid's power is
owned by the OFFICE1 region -- the region D-132 q1 is migrating away from. Retire
Office1 and nothing owns the dc0 region VM's power while it does not autostart.
That is an ordering exposure in the migration, not an autostart bug. LOGGED, NOT
FIXED.
WHY THE OPERATOR'S MATRIX IS THE RIGHT INSTRUMENT: autostart is a boolean that
cannot express WHO OWNS a VM's power or WHAT HAPPENS WHEN THAT OWNER GOES AWAY --
the questions that surface the circularity, and they recur at every DC standup.
Seed inventory measured: 4 outer domains, 12 dc0 inner, 11 dc1 inner; exactly four
carry autostart=enable (voffice1, office1-opnsense, both DC edges).
repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Both DC VMs 416 -> 480 GiB; 128 GiB swap; new preflight gate P8 (substrate drift)
...
POST-BOOKEND: the GA-R4 bookend (4b8ba3c) was committed BEFORE this work, so the
ledger close summary and the 08-01 sweep do not cover it. The ledger line is
AMENDED in this commit rather than left stale.
Operator: "Option 2 look sgood me. I approve the sequencing, process as
autonomously as possible" -- +64 GiB to EACH DC with swap first, over my
recommendation of +48. The deciding arithmetic, put up before the ruling: at +64
each, 994 GiB allocated against a 1007.4 GiB host = 13.4 GiB residual while the
host's OWN measured footprint is 18.0 GiB, with 1 GiB swap free. Swap is what
makes +64 safe, hence the sequencing.
PROBLEM FIXED, MEASURED: dc0's rack ran 402 GiB of inner guests in a 409 GiB host
(98.3%), leaving 7 GiB for the rack OS, mirror, snap proxy and page cache. Each
rack now has ~71 GiB.
SWAP: operator-run -- sudo -n is NOT available on vcloud so that half could not be
automated. /swap2.img 128 GiB, 600 root:root, live + fstab. 135 GiB total.
CHANGED THROUGH TOFU, the only correct place: a virsh setmaxmem would have been
reverted by the next apply. The derivation now lives IN the variable comment.
PLAN ASSERTED ON CONTENT (2026-07-20: an in-place apply silently regenerated 9
node MACs): per DC 1 in-place, 0 create/destroy/replace, 0 MAC changes, memory the
only changed attribute. A THIRD resource appeared and was resolved BEFORE
applying: module.office1_opnsense ~ id = 2 -> 11 sits under "has changed outside
of OpenTofu" (drift OBSERVED), not "will be updated in-place" (action PLANNED) --
libvirt's domain id is a runtime value. It was not touched.
dc1 FIRST AS CANARY, and it answered the open question: the in-place update
BOUNCES the guest (domain id 9 -> 12, uptime 0 min). Guest sees 472.2 GiB.
>>> FINDING: vr1-dc0-maas-01 (MAAS region) and vr1-dc0-juju-01 (juju controller)
have autostart=disable. <<< Only the edge auto-recovers. Checking first is the
only reason the dc0 bounce did not come back looking catastrophic. Both started by
hand and verified. LOGGED NOT FIXED: a host reboot leaves dc0 with no region and
no controller -- a standing exposure independent of this change.
RESULT: host used 809 -> 63 GiB, available 197 -> 944 GiB; the restarts also
released the stranded RSS (vvr1-dc1 354 -> 12 GiB), discharging 08-01 sweep F1 as
a side effect. Recovery verified end to end: MAAS region API 000 -> 502 -> 200 (a
real boot progression), snap proxy listening, mirror 200, juju controller model
"Last connection: just now".
GATE P8 -- substrate drift. Built because answering "does tofu need updates
logged" produced a measurement: preflight.sh and pre-flight-checks.sh contained
ZERO tofu plan checks and the opnsense drift had sat unseen. DESIGN POINT:
-detailed-exitcode returns 2 for BOTH a pending change and a harmless observation,
so P8 asserts on CONTENT -- pending ACTION FAILS, observed-only drift WARNS,
unrecognised shape REFUSES. A gate permanently red on benign drift gets ignored.
It WARNS rather than fails when it cannot look, since the outer root lives on
vcloud and "not the substrate host" is a legitimate state.
Harness 33 -> 38 (T34-T38). Baseline proven by stashing: my first draft's 7
failures were MINE, not pre-existing. Three mutations each killed a NAMED test,
script restored sha256-identical. T38 exercises the absence by removing the
fixture STATE, never the fake tofu from fakebin -- the real tofu is on the system
PATH and a "closed" PATH would fall through to the LIVE substrate. PROVEN LIVE:
tofu exit=0, 0 pending, 0 drift -> [ok] zero diff.
SCOPE: tofu owns the SUBSTRATE only. MAAS carves, rack services and the v6 node
carve are script-driven by design (Model B / D-123); P8 does not check them.
gauntlet ALL GREEN (96); repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

SESSION CLOSE 2026-08-01 (GA-R4 bookend): snap proxy LIVE, D-139, IPv6 proven
...
DURABILITY TRIAD: vcloud 0 uncommitted / 0 unpushed at 844b2e4. voffice1 was 36
commits BEHIND and has been brought current. The dc0 rack's ~/repo-stage (D-138
client input, NO git) was DIGEST-COMPARED not assumed: bundle.yaml and both
non-PKI overlays sha256-MATCH the repo, octavia-pki correctly absent (gitignored).
The provenance gap is real; the drift is not.
GATES: repo-lint 0 fail / 1 warn (legacy carve-out), 649 files; gauntlet ALL GREEN
(96 harnesses); ledger-scan 3 open decisions, 26 open SEC (none opened this
session), next-free D-140 / DOCFIX-207 / BUNDLEFIX-053. RECONCILED: D moved
139->140 matching the one D-number assigned; SEC unchanged matching zero opened.
SWEEP: docs/audit/queued-findings-20260801-stage5-ipv6-d139.txt -- SEVEN FIRST
SURFACE items, each grep-proven absent from every repo surface before writing.
F1 is highest-consequence: the capacity headroom the next deploy needs is HELD BY
TWO IDLE VMs (vvr1-dc0 434 GB RSS, vvr1-dc1 371 GB, ~805 of 1007 GB) with dc0's
nodes powered off and dc1 carrying no model -- and this document's own capacity
record was computed against ALLOCATION, not residency. Also: the voffice1 lag; the
host-lag investigation and its NEGATIVE result; the rack digest MATCH; that no
gitignored permission rules changed; a transient _fmtprobe.tf; and the PreToolUse
guard firing on PROSE containing a guarded command's name.
GA-R7 MEMORY REVIEW: CLEAN -- zero entries claiming operator policy, priority or
posture. One update: instrument-currency gains its NINTH instance in a new shape
-- a "closed" PATH that still held the binaries it was meant to hide. Generalised
as: when a negative depends on an ABSENCE, prove it rather than arranging it.
LEDGER ROTATED (GA-R4 rule 3): stood at 296 lines and this close would have
breached 300. The two oldest closed-session summaries moved VERBATIM to
docs/archive/session-ledger-rotated-20260801.md; ledger now 282. No orphaned
session, so rule 7 is a no-op. NO STAGE OPENED OR CLOSED -- Stage 5 remains OPEN
and this is a session bookend, not a GA-R6 stage close.
OWNED: my BUG-1 fix was wrong; I scoped the v6 experiment wrong (ceph couples
storage+replication); I stated an agent's input source wrongly; I scoped a research
agent with no repo path so 355 lines landed in /tmp and needed rescuing; and one
commit went red on ASCII-only em-dashes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Next steps: nodes released, D-139 execution list replaced, LP draft written
...
(1) TEST NODES RELEASED. t7ymp6 / fg6gxm -> Ready / owner=None, released
individually and read back. Nothing stranded. G19 was run live before the window
closed, which was the whole reason for holding them.
(2) D-139 GAINS A CORRECTION NOTE REPLACING ITS EXECUTION LIST. Neither ruling is
touched. The original list is preserved verbatim inside the note rather than
silently overwritten -- same reason the struck RFC 6724 rationale was kept: a
later reader must be able to see what was wrong or they will re-derive it.
SILENT FAILURE (why this could not wait): dc-node-v6-carve.py assigns v6 "on
every plane WHERE IT ALREADY CARRIES IPv4", keys the subnet off the same MAAS
vlan as that v4 link, derives the host part from the v4 octet, and its loop body
is `if not v4: continue`. Run after v4 removal it carves FOUR FEWER PLANES PER
NODE AND EXITS CLEAN -- and the original list's "remove v4 LAST" wording
actively invited that ordering.
UNSATISFIABLE: "remove v4 LAST, after each is proven" -- per-plane conversion is
atomic, so no such state exists.
UNSAFE STEP FUSION: apex CREATE and RETIRE in one step, with 52 dependent
ip-addresses (the v6 VIP legs) inside the retiring ULA /64s.
Replacement is 7 steps: CREATE-only apex push; MAAS GUA carve ALONGSIDE the ULA;
node statics re-carved WHILE v4 IS STILL PRESENT; Octavia SAN reissue +
lib-net.sh; lb-mgmt carve; re-home the 52 VIPs THEN retire; and v4 removal as a
SEPARATE experimental step scoped to storage + replication TOGETHER per "B plus
C" (together because ceph-osd's prefer-ipv6 sets ms_bind_ipv4=False globally).
Step 7's prerequisite -- network-get on a v6-only bound space -- is stated unmet.
(3) LP DRAFT WRITTEN, NOT FILED: docs/audit/lp-draft-20260801-ceph-osd-ipv6-
static.md. The "C" half of the ruling. A COMMENT on the EXISTING LP #2061836, not
a new bug -- the ceph-mon task is already Fix Committed and only ceph-osd remains
New. Every code claim quoted from the downloaded rev-953 artifact. Claude does not
post to Launchpad; the operator files it (lp-draft-20260721 precedent).
Two things written INTO the draft's operator notes rather than left implicit: do
NOT cite LP #1590598 (Fix Released, a different defect -- the inverted-citation
class); and this bug does NOT block us locally, since setting ceph-public-network
/ ceph-cluster-network to the v6 CIDRs bypasses the buggy call. Overstating our
dependence on an upstream fix would be a bad-faith way to raise its priority.
OWNED: first push attempt went red on L1 -- I used em-dashes in the changelog and
this repo is ASCII-only. Caught by lint before commit, fixed, re-run 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

G19 BUILT and PASSED LIVE; finds a latent defect in D-139's execution list
...
scripts/dc-node-v6-verify.sh (270 lines, 24-line header) + tests/dc-node-v6-verify
(55 cases); manifest 95 -> 96. All three subcommands exercised LIVE this session:
plan vr1-dc0 PASS exit 0 (9 machines, v6 plane set IDENTICAL across all 9, 6
planes, prefixes DERIVED); node vr1-dc0 <spec> on storage-01 against live peer
storage-02 PASS exit 0 with 12 assertions -- six address-presence, six
peer-reachability, every plane "up, global, DAD complete, sole global on the NIC"
and "replies, neighbour DELAY"; bridges vr1-dc0 reporting 7 bridges. Harness
55/55, gauntlet ALL GREEN (96), three mutations each turning a NAMED test red.
RUN AGAINST A REAL NODE while the boot window was open specifically so it would
not ship fixture-green -- the snap proxy's harness was green for a week before
anything was proven end to end.
NO v6 PREFIX LITERAL EXISTS IN THE FILE, and that was a constraint not a nicety:
D-139 retires the ULA /48 for GUA, so a baked table would be right today and wrong
on carve day. plan derives from live MAAS and emits a SPEC; node consumes it and
knows no prefixes. A harness case drives the same green path with a ULA fixture
AND a GUA fixture, so the property is EXECUTED, not asserted.
>>> LATENT DEFECT IN D-139's OWN EXECUTION LIST <<<
D-139 names dc-node-v6-carve.py to "re-carve 54 node v6 statics". VERIFIED against
that script's source: it assigns v6 "on every plane WHERE IT ALREADY CARRIES
IPv4", picks the v6 subnet "on the SAME MAAS vlan as that v4 link", derives the
host part from "the last octet of the node's own v4 address", and its loop body is
a bare `if not v4: continue`. Under D-139 five planes lose v4 and lb-mgmt never
had a v4 twin, so after v4 removal it would SILENTLY CARVE FOUR FEWER PLANES PER
NODE rather than fail. The v6 carve MUST run BEFORE v4 removal, or the tool must
be rewritten -- compounding the already-recorded fact that D-139's "remove v4
LAST" clause is unsatisfiable. LOGGED, NOT FIXED. My own agent mandate stated this
tool's input source WRONGLY (apex/D-136; it is live MAAS); the agent checked
rather than inheriting the error, which is why the defect surfaced.
THREE MORE FINDINGS: (i) the multicast reading is ESTATE-WIDE -- all six plane
bridges plus the WAN bridge at snooping=1/querier=0 on dc0 AND dc1, and the
jumphost's uplinks too, i.e. the default here, not a chosen setting; the
subcommand tags it SUSPECT and asserts no Linux behaviour, correctly, since
cold-start multicast ND was already measured working across it. (ii) plan's PEER
SELECTION IS DEPLOYMENT-BLIND -- it picked control-01, which is not deployed, so
the derived SPEC cannot pass in a PARTIAL deployment; worked around by
substituting the live peer, logged not fixed. (iii) a harness bug caught by the
instrument-currency rule, NINTH of its kind: absent-ip/absent-virsh REFUSE cases
were driven with PATH=/usr/bin:/bin, which still contains both, so they were red
for the WRONG REASON.
TWO OPEN CHECKS SETTLED on the live nodes: every NIC carries exactly ONE global
(sole-global predicate holds today), and an OVS-internal br-ex does report
LOWER_UP (predicate safe).
GA-R2: the build agent wrote its own changelog; folded VERBATIM into this
session's single changelog and removed. Crossing midnight does not start a new
session. repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

IPv6 PROVEN WORKING on the dc0 node planes -- and already load-bearing
...
Live experiment, operator-approved: two storage nodes MAAS-deployed (jammy) to
open a boot window, issued individually. Region asserted first (vr1-dc0-region ->
one rack controller hot-kid; admin -> the three Office1 racks); maas-profile-
assert.sh is MISSING on voffice1, so the property was asserted inline not skipped.
A1 ADDRESSES COME UP -- PASS. Six global v6 addresses live on the NICs, matching
MAAS exactly, 0 tentative / 0 dadfailed, on-link /64 routes, no v6 default
(expected). Asserted on the INTERFACE, never on MAAS -- 54 carved statics were
never evidence any existed on a NIC, and this is the first check on a ROLE NODE.
A2 THE PLANE CARRIES v6 -- PASS on all six planes, 0% loss, every neighbour
REACHABLE. THE RISK I FLAGGED IS REFUTED FOR THIS PATH: every plane bridge is
multicast_snooping=1 / querier=0, a known ND failure mode, and the neighbour
table was COLD, so the first solicitation went out as MULTICAST and resolved on
all six. HONEST RESIDUAL: proves ND from cold, not behaviour across long idle
periods where snooping entries age out.
A3 G17's ROLE-NODE HALF CAPTURED -- PASS. Open since 2026-07-30 (that capture was
taken on the CONTROLLER VM and said so). Mirror real package path: exit 0, HTTP
200, 269219 bytes, all four body fields. AND the 2026-07-31 deploy blocker is
confirmed fixed FROM A REAL NODE -- all four suites 200, jammy-backports included,
where it was 404 and put 22 units into hook failed: "install".
A4 HEADLINE, and it inverts this repo's framing: THE NODE'S TIME SOURCE IS ALREADY
IPv6. ServerName fd50:840e:74e2:220::6, ServerAddress family 10 = AF_INET6,
Stratum 3, root distance 57.327ms, poll backed off to 2min8s, ntp.ubuntu.com
unused as fallback. G17 assertion (2) passes, and separately IPv6 is not a future
state here -- it is a live operational dependency predating all this work.
F7 CLOSED IN PASSING: "still unverified that a redeploy applies the new hostname"
-- it does; both report their ruled names from `hostname` on the running OS.
SCOPE: these are ULA addresses, the PRE-D-139 carve, superseded by design. The
PROPERTY tested is family-agnostic and transfers. Nothing here tests a charm, a
container, network-get or Ceph -- those remain the untested rungs.
Nodes deliberately left Deployed so the G19 gate can run against a LIVE node
rather than ship fixture-green. Release owed. repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

RULED "B plus C": D-139 ruling A narrowed to an experiment + upstream fix
...
Operator, exact utterance: "B plus C". Ruling B's GUA carve is unaffected and
proceeds; ruling A's five-plane v6-only scope narrows to a bounded experiment;
metal-internal, data-tenant and lb-mgmt stay DUAL-STACK pending upstream fixes.
That last sentence is recorded as an INTERPRETATION, not as the utterance.
>>> I SCOPED THE EXPERIMENT WRONG AND THIS CORRECTS IT. <<<
MEASURED against shipped ceph-osd rev 953, ceph_hooks.py:544-553: prefer-ipv6 is
ONE switch that sets ms_bind_ipv4 = False GLOBALLY, so it governs Ceph's PUBLIC
network (storage) and CLUSTER network (replication) TOGETHER. There is no state
in which Ceph runs cluster on v6 and public on v4. I recommended "replication
alone, smallest blast radius" -- not achievable. The experiment is storage +
replication as a PAIR, with a larger blast radius than I represented.
A PATH THE EARLIER ANALYSIS MISSED: the buggy bare get_ipv6_addr() (LP #2061836,
dynamic/SLAAC-only, never fixed on ceph-osd, while our nodes hold MAAS statics) is
reached ONLY under `if not public_network:` / `if not cluster_network:`. Setting
ceph-public-network and ceph-cluster-network to the v6 CIDRs means it is never
used for address selection -- the defect becomes AVOIDABLE BY CONFIG. This is a
CODE READING, NOT A MEASUREMENT, and it does not address the get_mon_hosts()/
get_host_ip race (LP #2109798).
REVIEW PASS on netbox/d139-gua-carve.py (independent second agent, per operator
standing instruction). 4 real defects fixed, 4 reported, 1 charter miss upheld.
D1 is the mandate's own named class "a refusal that does not refuse": an
unreachable apex gave a raw traceback and EXIT 1 -- the tool's own header defines
exit 1 as a write error -- on a dry run that wrote nothing. _req caught only
HTTPError so URLError walked past the REFUSE handler. Fixed, REFUSE(2).
D2: a comment asserted ruling B was "FLAGGED" -- already false at commit time.
Replaced with a citation; prose in a script rots against governance state, a
citation cannot. D3/D4: two greps over the tool's OWN SOURCE and two vacuous
negatives replaced with behavioural coverage.
main() decomposition DECLINED on a reusable argument: the `if not a.commit:
return 0` guard is ADJACENT to the write block, and extracting commit_writes()
would split that safety proof in two. Line count is not the criterion.
R1 IS A REAL SECURITY FINDING, MEASURED: urllib carries the Authorization header
across a CROSS-HOST redirect (loopback probe -- the target received the token
verbatim). SANDBOX_HOSTS gates the supplied url, not a redirect from it. LOW
severity, logged not executed.
Harness 65/65 -> 71/71; all 5 original mutations RE-RUN and still standing, plus
3 new; gauntlet ALL GREEN (95); dry-run body diffed BYTE-FOR-BYTE identical to the
committed capture; apex 139 prefixes before and after. QUALIFIED: a dry run
executes zero POSTs, so D1's change to the WRITE path's failure semantics is
UNTESTED and is declared, not implied.
PROCESS FAILURE REPEATED AND OWNED: the review mandate required a docs/audit/
capture and forbade touching CURRENT-STATE.md -- jointly unsatisfiable against
L10. The agent correctly left the tree RED and said so rather than working around
it. Standing fix, now twice learned: an agent mandate must name a repo path for
findings AND leave the committer able to satisfy L10.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

D-139 ruling A NOT achievable at these charm revisions; apex carve tool built
...
TWO agent results, both recorded; nothing actioned, nothing pushed to NetBox.
>>> RULING A (v6-only on five planes) IS BLOCKED BY UPSTREAM CHARM DEFECTS <<<
RULING B (full GUA) IS ORTHOGONAL AND STANDS -- GUA-vs-ULA is independent of
whether a plane keeps v4, so the carve tool is unaffected and a GUA DUAL-STACK
carve is achievable today. This needs an operator ruling.
RE-VERIFIED INDEPENDENTLY against the shipped ceph-osd rev 953 artifact, because
it decides the ruling. hooks/utils.py:203-217 get_host_ip() returns
get_ipv6_addr()[0] when prefer-ipv6 is set, else socket.inet_aton(hostname) and
on failure dns.resolver.query(hostname, 'A') -- an IPv4-ONLY lookup -- under the
charm's OWN comment "...just let it kill the hook". So ceph-osd fails on a
v6-only plane in BOTH directions: false -> inet_aton raises on a v6 literal then
NXDOMAINs an IN A lookup (LP #2109798, New); true -> bare get_ipv6_addr() (also
ceph_hooks.py:547, WITHOUT dynamic_only=False) returns DYNAMIC/SLAAC only while
every VR1 node holds a MAAS STATIC v6 (LP #2061836 -- Fix Committed on ceph-mon,
New and never fixed on ceph-osd, confirmed absent from rev 953).
Reported from the capture and NOT re-verified here, recorded as such: mysql /
mysql-router bare user:pw@addr URIs invalid for an unbracketed v6 literal;
hacluster 2.4/stable rev 166 (released 2026-06-29) still hard-writes
ip_version: ipv4 though LP #2111852 is Fix Committed -- FIX COMMITTED IS NOT FIX
RELEASED; OVN documenting encap as "The IPv4 address of the encapsulation tunnel
endpoint" with zero v6 encap values in OVN 24.03's test suite.
Counted locally: metal-internal appears 273 times in bundle.yaml across ~44 apps;
rabbitmq-server is DEPLOYED by the bundle but ABSENT from PREFER_IPV6_CHARMS --
the only mismatch across all 33 bundle pairs, where invariant 9a would emit a
factually false message.
APEX CARVE TOOL (netbox/d139-gua-carve.py, 232 lines) + tests/d139-gua-carve
(65 cases, offline stub apex); manifest 94->95. Live dry run: CREATE 22 |
EXISTS 6 untouched | RETIRE-REPORT 18, and it surfaced 52 DEPENDENT ip-addresses
inside the retiring ULA /64s that a later delete would orphan. Proof no write
occurred: apex prefix count 139 before and after. 5 mutations, each turning a
NAMED test red, tool restored sha256-identical.
OWED: (a) the foundational measurement is STILL not taken -- does network-get
return v6 on a v6-only bound space? (b) D-139's "remove v4 LAST, after each is
proven" is UNSATISFIABLE as written; per-plane conversion is atomic.
(c) the rabbitmq-server mismatch.
PROCESS FAILURE, OWNED: the research was written to /tmp and would have been LOST
-- I scoped the agent read-only and gave it no repo path. Rescued, sha256-verified.
STANDING CONSEQUENCE: every agent mandate must NAME A REPO PATH for its findings;
"read-only" constrains what it may CHANGE, not where it may RECORD.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

The "juju client blocker" is NOT one -- it is D-138 working correctly
...
CORRECTS this repo's own 2026-07-31 finding 9d, which called it a blocker.
Stage 5 can reach `add-model` TODAY, from the dc0 rack. No key movement needed.
VERIFIED INDEPENDENTLY from the main session, not accepted on the agent's word:
juju 3.6.27-genericlinux-amd64 at /snap/bin/juju on the rack; controller
vr1-dc0-controller* admin/superuser on cloud vr1-maas; the controller model reads
"Last connection: just now"; client credentials list EXACTLY ONE entry,
vr1-maas -> vr1-dc0-cred.
ROOT CAUSE OF THE ssh REFUSAL, MEASURED ON THE MACHINE: the alias is correct in
every part -- hop chain, host, port, per-hop identity, host key. subtle-grouse's
authorized_keys holds exactly two lines, both Juju's, so the office1_svc key the
alias offers is genuinely absent and the refusal is right. The obvious hypothesis
was REFUTED rather than assumed: the machine booted AFTER the 2026-07-30 key
import and still got only Juju keys, because deploy-time cloud-config carries
Juju:juju-client-key alone. `juju ssh -m controller 0` already works.
SEC-026 control (1) DISCHARGED BY MEASUREMENT on both sides -- rack lists one
credential, voffice1 lists two, so the forbidden whole-store copy did not happen.
No new credential residency; no new security-ledger row owed.
ONLY HYGIENE OWED, NOT EXECUTED (hard rule 3): a dangling current-model pointing
at the model destroyed 2026-07-31 makes bare `juju status` error. Measured NOT to
be an add-model precondition.
FOUR GAPS LOGGED NOT FIXED, in consequence order:
G4 the `openstack` CLI is MISSING ON THE RACK -- D-138 definition-of-done gap
blocking phase-03+. SAME item as F1 from the 2026-07-30 sweep, now confirmed
on the D-138 host. It has survived two sweeps.
G3 preflight CANNOT PASS on vcloud by construction: the octavia-pki overlay is
gitignored PKI material absent from the vcloud tree, and
pre-flight-checks.sh:161 hard-fails on it. The host-dependence class again.
G2 the D-138 client host's deploy input has NO PROVENANCE -- ~/repo-stage on the
rack is a hand-staged 9-file copy with no .git (8/9 byte-identical by sha256).
A third copy beside the two clones verified at c58bf95.
G1 phase-4 runbook :379,398 still say "voffice1" for add-model/spaces, stale vs
D-138; voffice1 has the binary but no registered controller, so the runbook as
written fails. DOCFIX owed.
DECLARED UNMEASURED with the reason: whether MAAS user juju-vr1-dc0 has an
imported ssh key -- no usable maas CLI profile exists anywhere (voffice1's
~/.maascli.db is ZERO BYTES). Does not change the fix; the proximate cause was
measured directly. The PreToolUse guard refused a command that would print the
MAAS API key (DOCFIX-016) and it was NOT retried in an altered shape.
Staged by explicit path; agent files still in flight are deliberately excluded.
repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|