| 2026-07-21 |
Ruling 1: vr1-dc0-maas RETIRED (repo removed); INCIDENT: pod delete cascaded to 9 machine records
...
- git rm opentofu/vr1-dc0-maas (operator-ruled 'Retire fully'); validate PASS
- maas vm-host delete 4 removed the stale pod AND the nine linked DC0
machine records (association check ran after the delete -- agent error,
owned; capture docs/audit/incident-20260721-pod-delete-cascade.txt)
- Substrate measured intact (10/10 domains, MAC pins config-carried);
CURRENT-STATE corrected same commit (GA-R1 C2); recovery flow gated,
pending operator decision; SEC-013 file cleanup HELD
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Stage-close batch: ledger-scan PARTIAL fix; DOCFIX-197; D-103/D-123 pod-refutation amendments; GA-R7 memory review
...
- scripts/ledger-scan.sh: PARTIALLY-RULED decisions now surface (D-131
blind spot); harness regression fixture, 46 checks ALL PASS
- DOCFIX-197: phase2 runbook Step 11 leg selection corrected to executed
reality (dc0<->dc1 leg; office1 leg carries live transit)
- D-103/D-123 amendments + modules/maas-vm-host header: 2026-07-20
measured pod refutation + per-machine-virsh ruling RECORDED (cited to
captures; no new ruling made)
- GA-R7 memory review DONE: no memory-only facts remain
- Read-only: stale MAAS pod object confirmed (id=4 vr1-dc0-inner, virsh)
for the pending retire-or-keep ruling
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|

Step E netem DONE -- G10 CLOSED; netem-link local mode; G16 opened (office1 channels residual)
...
- modules/netem-link: optional local execution (empty ssh target -> bare
sudo tc; D-128 outer root runs ON vcloud); NEW tests/netem-link harness
(12 cases); gauntlet 76 ALL GREEN
- outer root: module netem_vr1_dc0_vr1_dc1 wired -- virbr5 (measured at
wire time), placeholder profile 'delay 3ms 1ms loss 0.01%' (S6 lean,
PROVISIONAL; D-100 gap #11 unruled); runbook Step-11 leg divergence
flagged, DOCFIX queued
- 1/1/0 STOP honored; operator RULED 'Targeted netem apply (Recommended)';
saved -target plan exact 1/0/0, applied; qdisc live on virbr5,
virbr7/virbr3 untouched (docs/audit/stepE-netem-20260721.txt +
outer-{plan,apply}-20260721-netem*.txt)
- Convergence 0/1/0 = office1 channels=[] residual only
(outer-plan-20260721-postE-residual.txt) -> split per E3 into NEW gate
G16; CURRENT-STATE section 5 expected plan re-recorded 0/1/0 (GA-R1 C1)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Disconnect collapse: predecessor bookend landed; netem-tc install VERIFIED (step E unblocked)
...
- docs/audit/netem-sudo-install-20260721.txt: read-only verification of the
operator-run install (0440 root:root, byte-identical, sudo -n -l exit 0)
- CURRENT-STATE step-E paragraph: install PENDING -> INSTALLED+VERIFIED,
remaining path stated (GA-R1 C1/C2, same commit)
- session-ledger: bounded SESSION CLOSE for the close-and-delivery session
- changelog-20260721-netem-install-verify.md: session changelog w/ reverts;
ledger-scan D-131 status-phrasing blind spot flagged as queued finding
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Step-E netem sudo: scoped NOPASSWD fragment shipped (operator-ruled)
...
scripts/sudoers.d/netem-tc grants exactly the netem-link module's two tc
verbs per measured mesh bridge (virbr5/7/3, re-measured this session;
drifting-ID caveat + fail-closed property documented). 9-case harness;
gauntlet 75 ALL GREEN. Install is operator-only (visudo-checked) and
pending; then the gated placeholder netem run closes G10 step E.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
Incident docs: 2 appendix-A entries, platform-traps MAC-regen corollary, LP draft
...
Closes the remaining documentation queue from the 2026-07-21
commissioning incident. Appendix-A gains the MAC-drift and
agent-resolver symptoms (verbatim, with checks + recorded fixes);
platform-traps 1e gains the unpinned-MAC regeneration corollary + index
row; LP draft ready for the operator to file (sanitization warning
included). Still queued: stale pod cleanup at stage close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
dc-rack-net installed on the rack: EXIT 0, check 10/10, SOA probe answers
...
Operator-approved install executed; capture committed. Rack legs now
reboot-persistent (dc0-rack-legs.service); forwarder regenerated and
proven (authoritative maas-internal SOA via 10.12.8.3). Interim
hand-placed state fully superseded.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
dc-rack-net.sh: D-131 sub-1 delivery -- rack legs + node-DNS forwarder, repo-carried
...
Site-keyed (dc0 rows measured 2026-07-21), runs on the rack via piped
ssh. Legs move from reboot-lost bare `ip addr` into a oneshot unit;
bridges resolved from libvirt NET NAMES at runtime (dc-planes does not
pin bridge names -- virbrN is a drifting ID); dns unit Requires the legs
unit. 14-case harness; gauntlet 74 ALL GREEN. Rack install stays gated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
D-131 sub-1 RULED: rack node-DNS forwarder is the STANDING per-DC pattern
...
GA-R5 record: question + exact utterance ("Standing per-DC pattern
(Recommended)") in the D-131 Status line. Delivery (site-keyed unit +
config + install/check script + harness, folding in rack-legs
persistence) follows in this session. Sub-decisions 2-4 remain OPEN.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
MAC-pin apply executed on voffice1: exact 0/9/0, convergence zero diff
...
Saved-plan apply (operator-approved): 54 MAC pins adopted into the inner
state; zero power flips (guard held), zero replaces. Post-apply verified
all 9 domains still shut off, MACs unchanged. Captures: guarded plan +
apply + convergence in docs/audit/. Node NIC MACs now config-pinned end
to end -- the 2026-07-20 drift class is closed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
node-vm: MAAS owns node power -- ignore_changes on running (operator-ruled)
...
The MAC-pin verification plan (docs/audit/inner-plan-20260721-macpin.txt,
0/9/0, 54 mac adoptions, zero replaces) exposed 9 entangled out-of-band
power-ons: the module re-asserts running=true against nodes MAAS holds
OFF. Operator-ruled: guard first, then apply.
- lifecycle ignore_changes = [running]: create still boots (PXE
enlistment); afterwards MAAS owns power (virsh here, Roosevelt IPMI).
- tests/node-vm +3 cases (guard present, no scope creep, cites MAAS);
15/15 green, gauntlet 73 ALL GREEN, repo-lint 0 fail.
- CURRENT-STATE: MAC-pinning delivery status + pending gated apply.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
node-vm: pin NIC MACs (54 measured) -- in-place applies can no longer drift them
...
Queued delivery item from the 2026-07-21 commissioning diagnosis: an
"in-place" apply (0/9/0) regenerated all 9 unpinned boot MACs on
2026-07-20 and stranded the fleet in MAAS.
- modules/node-vm: optional interface_macs (all-or-nothing + format
validations); mac = { address = ... } shape verified against the
dmacvicar/libvirt 0.9.8 schema dump this session.
- vr1-dc0-substrate: all 9 nodes x 6 planes pinned to values MEASURED
this session (virsh domiflist on vvr1-dc0); metal-admin entries
cross-checked vs MAAS boot-interface records 9/9.
- NEW tests/node-vm harness (12 cases). Gauntlet 72 -> 73 ALL GREEN;
repo-lint 0 fail; opentofu-validate PASS.
NOT YET APPLIED on voffice1 (inner root state) -- the gated apply's plan
must show NO replace; any replace is a STOP.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|

Session close 2026-07-21: commissioning RESOLVED 9/9 Ready; D-131 proposed; SEC-014 opened
...
Lands the ops-commissioning-diag session's deliverable, whose close
bookend was lost to a session disconnect. Landed same-day by the
successor session after read-only re-verification (9/9 Ready, forwarder
active+enabled, dhcpd on both controllers, Phase-7 tag cleanup done).
- CURRENT-STATE: commissioning resolved (two stacked faults: MAC drift
from the in-place apply; MAAS 3.7 rack-only agent resolver SERVFAIL);
G10 remainder = netem step E only.
- design-decisions: D-131 PROPOSED (rack-only node DNS strategy).
- security-ledger: SEC-014 OPENED (rack cluster secret exposure).
- committee doc: closed by addendum.
- changelog + adjudication capture + as-executed index row.
- session-ledger: bounded close entry + machine-derived block re-seeded
from the 2026-07-21 scan (D next-free 132, DOCFIX 197, BUNDLEFIX 052,
10 open SEC rows).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
| 2026-07-20 |
Fix committee doc: ASCII byte (em-dash in title) + CURRENT-STATE pointer (L10)
...
Prior commit 34c8770 pushed with a non-ASCII em-dash in the committee doc
title and a decoupled L10; both fixed here. Correct lint gate learned: repo-lint
exits 0 clean / 1 FAIL / 2 warn-only, so the right gate is exit != 1 (the
permanent legacy design-decisions WARN makes exit 2 the normal clean state) --
not the exit==0 I mis-used, and not the piped-tail that masked real exit-1 fails.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|

Diagnostic committee (4 reviewers) refutes marginal-timeout 4/4; MTU exonerated by measurement
...
Committee synthesis: docs/audit/commissioning-committee-20260720.md. Consensus:
30-min silence = a HANG not slow progress, so raising node_timeout is the wrong
knob; the proposed one-node test was confounded (timeout + concurrency changed
together). Process errors caught: measured rack not region, tcpdump doubly
mis-scoped (port + window), enlistment paradox resolved (MAAS logs script
COMPLETION so a hung commissioning-only script = identical silence). Pivotal
zero-cost reads: metal-admin MAAS VLAN MTU = 1500 -> guest never jumbo -> MTU
branch EXONERATED; region ample RAM, Temporal quiet at rest (175k wedge errors
were pre-restart). Live causes now: region Temporal starvation during a run,
and a commissioning-only script hang on nested-virt hardware. Decisive gated
test: one commission + console=ttyS0 kernel opt + full-window lease-IP capture
on virbr2+enp1s0.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
Clean commissioning experiment: contention refuted; leading hypothesis = marginal 30-min timeout over slow nested I/O
...
Genuinely isolated node (0 others running, 397GiB free, terminal start
state, read-only polling) still failed at 1770s = 29.5/30 min. Refutes
contention but consistent with INHERENTLY SLOW; batch 3-pass/6-fail-at-the-
mark is the classic marginal-timeout signature. 3 nodes reached Ready on the
same rack/subnet/metadata path, so metadata is not globally broken (:5248 up,
rack->region 301; a :5248 tcpdump was inconclusive -- window closed before
the fetch stage). Recommended test needs an operator decision (MAAS-wide
config): raise node_timeout, commission one node. Stopped live diagnosis per
the circuit-breaker.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
SEC-012 + SEC-013: ledger rows for the two credentials this deploy created (obligation was outstanding)
...
SEC-012: dedicated MAAS->libvirt service key. Held by a daemon on BOTH
controllers (MAAS dials the pod from the REGION, measured), authorizes a
libvirt-group user, so it can drive every inner node VM and the DC edge.
Rotation obligation + a SCOPE question: libvirt-group is broader than the
power verbs MAAS needs; tighter grant is the hardening candidate and is
Roosevelt-relevant (IPMI has the same power-only-vs-full-control question).
SEC-013: admin-scoped MAAS API key materialized to disk on the region;
surfaces include any maas-provider tfstate and a CLI profile. Verified by
format only, never printed. Tied to whether vr1-dc0-maas is retired.
Open SEC count 7 -> 9; G14 row updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|

node-vm console shipped (0/9/0 in-place); commissioning narrowed, UNRESOLVED; two of my own experiments disclosed as INVALID
...
Console added + log dir parameterized, but logs stay 0 bytes: firmware
writes to VGA, so serial alone does not make a PXE-booting node observable
-- correction queued (needs graphics + virsh screenshot, or a MAAS kernel
cmdline change). Solid: PXE + ephemeral handoff work; ephemeral OS boots
with working networking (leases + NTP to rack); no region/metadata traffic
seen; memory, image sync, DHCP and shape all ruled out. SEC-010 blocking
node->region is SUSPECTED but unreadable (rules lack counters).
INVALID EXPERIMENTS RECORDED: the 'single node alone' test was not alone
(5-6 nodes still up, proven by tcpdump), and the retry did not restart
MAAS's timer while I destroyed domains mid-run. Contention is therefore a
LIVE hypothesis again, not refuted. Stopped per the circuit-breaker rather
than iterate on contaminated evidence.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
OPEN: 6/9 nodes fail commissioning; contention refuted; blocked on node-vm having no console
...
3 Ready, 6 timed out at 30 min. Hypotheses tested and REFUTED by
measurement: contention (single node alone stuck in Loading ephemeral 20+
min with 385GiB free, load 0.04), rack boot-image sync (status synced, 10
images), DHCP (same nodes enlisted fine earlier). Real blocker is
observability: modules/node-vm defines no serial console, so a stuck node
is a sealed box -- the same gap already logged for cloudinit-vm, now biting
a live diagnosis. Fix is the opt-in serial+log pattern opnsense-edge
already carries; NOT applied, since it re-applies domains mid-diagnosis and
is an operator ruling. Step E prerequisites measured and recorded (dc0-dc1
mesh = virbr5 ON VCLOUD; netem-link assumes passwordless sudo vcloud lacks).
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
STEP D COMPLETE: per-machine power tooling shipped; commissioning works, nodes reaching Ready
...
Operator-ruled per-machine virsh power. maas-node-power.sh + harness (24/24,
gauntlet 72 ALL GREEN): MAC-matched because MAAS renames machines at
enlistment, dry-by-default, every write verified by a real query-power-state.
Region prereqs fixed (no virsh there; maas profile was root-only -- created
via stdin login, key never printed). All 9 nodes powered and commissioning
end-to-end: 3 Ready / 6 Commissioning, shapes exact to D-121 Option C.
vr1-dc0-maas root retained but unused pending the stage-close D-103/D-123
amendment. G10 remainder: netem only.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
CURRENT-STATE: record step-D pod findings + pending ruling (clears L10 on prior commit)
...
Prior commit landed an audit capture without the same-commit CURRENT-STATE
update that GA-R1/C1 requires; L10 caught it. Position now records: pods
incompatible with node-vm volume-ref disks (reproduced locally), pod
unnecessary since PXE already discovered the nodes, per-machine virsh power
measured working, ruling pending.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|

Pod path chased to root cause: MAAS virsh pods incompatible with node-vm volume-ref disks; per-machine power PROVEN working
...
Dedicated MAAS->libvirt key minted+wired (operator-ruled). Two further
measured findings: (1) MAAS dials the pod from the REGION, not the rack, so
the credential belongs where MAAS dials from; (2) each 503 created a broken
pod server-side outside tofu state (both deleted). Final blocker: domblkinfo
fails on pool+volume disk refs -- reproduced LOCALLY on the rack with an
active pool, so it is libvirtd, not snap/SSH. Making pods work would mean
converting node-vm to file-path disks + re-applying 9 domains. Not needed:
the pod's job was discovery, which PXE already did, and per-machine
power_type=virsh query-power-state returns state=off on the canary. Per
machine power is also the Roosevelt shape (IPMI per node). Awaiting ruling.
CURRENT-STATE untouched pending the ruling.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
Step D part 2: MAAS root shipped; apply refutes D-123's local qemu:///system pod mechanism
...
New opentofu/vr1-dc0-maas root (isolates MAAS creds from substrate roots
per DOCFIX-179's lesson); plan clean, apply failed 'Failed to login to
virsh console'. MEASURED root cause: confined MAAS snap gets Permission
denied on libvirt-sock and ships NO libvirt interface to connect -- a local
qemu:///system pod is architecturally impossible with snap MAAS. D-123
Model B and modules/maas-vm-host both state that mechanism; intent (no
cross-fiber dial) survives, mechanism does not. Amendment pending a ruling.
qemu+ssh replacement measured feasible (snap ships ssh; rack user in
libvirt group). Nothing was created by the failed apply.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
INCIDENT RESOLVED + G10 depth-4 boot gate PASS: all 9 DC0 nodes enlisted in MAAS
...
Region restart recovered Temporal; dhcpd verified by PROCESS on both
controllers (0->1 each), not by service_set. One restart fixed both sites,
confirming the single wedged-Temporal root cause. The identical canary that
had failed then enlisted immediately -- proof it was the real blocker.
All 9 nodes enlisted with shapes exactly matching D-121 Option C
(3x16/64, 2x12/48, 4x8/24): node VMs inside vvr1-dc0 PXE-booted from the
Office1 region across the transit and run the ephemeral kernel = real
nested KVM at depth 4, proven behaviorally. Appendix-A entry queued for
this new incident class.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
Step D part 1 done (metal-admin DHCP configured); INCIDENT: MAAS DHCP down region-wide (Temporal wedged)
...
Rack registered + ruled dynamic range 10.12.8.100-.200 + VLAN 5005
dhcp_on/primary_rack verified by read-back. Canary node did NOT enlist:
measured dhcpd off, no process, no generated config on the rack while the
API reports dhcp_on=True -- config and running state disagree. Region
journal shows Temporal wedged ('Not enough hosts to serve the request',
2807 retries), and MAAS 3.7 drives DHCP through Temporal workflows.
Scope is wider than this deploy: voffice1 has NO dhcpd process either, so
Office1's own DHCP has been down since ~the 2026-07-17 reboot, masked by
both service VMs already being Deployed. Remedy proposed, NOT run (gated).
Detection gap queued: cloud-assert trusts MAAS's self-report.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
D-129 edge profile APPLIED on 26.7: qga channel shipped, plugins installed, agent ANSWERS guest-ping
...
expose_qga_channel (opt-in, default OFF; dc0 true) applied as an in-place
domain update -- disk with the console/API bootstrap never at risk. Real
install (status ok + msg_uuid, not dry-run echoes); service started via the
plugin's own API; gate met to the runbook's bar: guest-ping -> {'return':{}}
and domifaddr --source agent reports both legs. Prior capture of the same
name recorded the false dry-run and is replaced. Security finding queued:
the plugin can disable guest-exec RPCs (remote command execution) -- propose
at the 26.7 hardening review.
CURRENT-STATE updated in this commit (GA-R1/C1, L10 satisfied).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
FIX false-success: opnsense-plugins.sh apply ALWAYS dry-ran (DRY_RUN:+ expands on 0)
...
Found on the first-ever live run (vr1-dc0 edge, 2026-07-20): ${DRY_RUN:+--dry-run}
expands whenever the var is set and NON-EMPTY, and DRY_RUN=0 is non-empty -- so
every apply silently dry-ran while printing 'OK: profile applied'. Nothing was
ever installed by this script. The harness missed it because every test drove
--dry-run, leaving the live path with zero coverage. Fix tests the VALUE; harness
gains two live-path tests via a stub api client (verified to FAIL 2/2 against the
buggy version before restoring). 20/20.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
DC edge ADDRESSED: WAN 172.30.2.2/24 + default gw, LAN -> 10.12.4.1/22; edge egresses 0% loss
...
Applied via the new operator-ruled v4 setter pair, WAN first per its own
ordering warning. Kernel-verified both legs; 3 interfaces intact; API
answers at the new LAN address (CORE_ABI 26.7). The edge had shipped with
WAN on dhcp -- unworkable on a /24 with no DHCP server (same fact behind
the earlier probe artifact). Interim bootstrap address removed.
CURRENT-STATE position + G10 row updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
D-125 egress isolation gate: PASS / CLOSED -- bridge-in proven end to end (2 identical runs)
...
Measurement disproved the missing-rule hypothesis first: operator-run nft
dumps show virbr4 structurally identical to the working virbr11, with all
three masquerade rules present INCLUDING the generic one, already fired --
so the pre-staged net-destroy/net-start was correctly not run. Matrix
re-probe: gateway ping 0, internet ping 0, curl 1.1.1.1 301, curl
archive.ubuntu.com 200, twice. Run-2's one-off ICMP failure recorded
UNEXPLAINED rather than swept. Side observation kept: the WAN segment
cannot reach the Office1 LAN through the ISP NAT (D-122/SEC-010 intent
holding) -- queued as an explicit hardening assertion.
CURRENT-STATE position + G10 row updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
CORRECTION (GA-R1/C2): D-125 rack-bridge follow-up was CONFOUNDED, not evidence
...
The rack has no default route (transit-only netplan; WAN bridge IP-less),
so 172.30.2.51-sourced traffic to 1.1.1.1 fails in the rack's own routing
table before the NAT is consulted. Its ICMP/TCP failures say nothing about
virbr4; the only clean signal is the guest's ICMP failure, and a guest TCP
datapoint is still missing. Measurement request widened accordingly: FULL
nft AND iptables-save dumps (backend unmeasured), diffing virbr4 vs
virbr11 across BOTH rule families -- POSTROUTING masquerade (subnet-keyed)
and FORWARD accept (interface-keyed, which a subnet grep would miss).
Fix-forward pre-staged: net-destroy/net-start reinstalls libvirt's rules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|