| 2026-07-21 |
G16 CLOSED: office1 channels residual reconciled by operator-ruled state surgery; outer plan ZERO DIFF
...
- G6-precedent surgery: channels null -> [] on the office1 edge state
entry, serial 29 -> 30, backup terraform.tfstate.pre-G16-20260721;
guests untouched (office1-opnsense Id 2 running throughout)
- Convergence: no differences
(docs/audit/outer-plan-20260721-postG16-converged.txt)
- CURRENT-STATE: section 5 expected plan back to ZERO DIFF; G16 row
CLOSED; ACTIVE = stage-close set only
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Recovery COMPLETE: 9/9 Ready again (shapes exact, power=virsh); retire-fully fully executed; SEC-013 narrowed
...
- Operator-ruled recovery: PXE re-enlist (~2 min) -> maas-node-power.sh
dry+commit (9/9 verified) -> re-commission -> ALL 9 READY ~3 min
(docs/audit/incident-20260721-recovery-verify.txt; new MAAS hostnames)
- voffice1 cleanup: vr1-dc0-maas dir + tfstate + SEC-013 key file removed,
absence verified; SEC-013 row narrowed to CLI-profile-only
- CURRENT-STATE incident paragraph RESOLVED same commit (GA-R1 C1/C2)
- NEW appendix-A entry: pod delete cascades to linked machine records --
read the pod's machine list BEFORE vm-host delete; non-empty = STOP
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Ruling 1: vr1-dc0-maas RETIRED (repo removed); INCIDENT: pod delete cascaded to 9 machine records
...
- git rm opentofu/vr1-dc0-maas (operator-ruled 'Retire fully'); validate PASS
- maas vm-host delete 4 removed the stale pod AND the nine linked DC0
machine records (association check ran after the delete -- agent error,
owned; capture docs/audit/incident-20260721-pod-delete-cascade.txt)
- Substrate measured intact (10/10 domains, MAC pins config-carried);
CURRENT-STATE corrected same commit (GA-R1 C2); recovery flow gated,
pending operator decision; SEC-013 file cleanup HELD
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Stage-close batch: ledger-scan PARTIAL fix; DOCFIX-197; D-103/D-123 pod-refutation amendments; GA-R7 memory review
...
- scripts/ledger-scan.sh: PARTIALLY-RULED decisions now surface (D-131
blind spot); harness regression fixture, 46 checks ALL PASS
- DOCFIX-197: phase2 runbook Step 11 leg selection corrected to executed
reality (dc0<->dc1 leg; office1 leg carries live transit)
- D-103/D-123 amendments + modules/maas-vm-host header: 2026-07-20
measured pod refutation + per-machine-virsh ruling RECORDED (cited to
captures; no new ruling made)
- GA-R7 memory review DONE: no memory-only facts remain
- Read-only: stale MAAS pod object confirmed (id=4 vr1-dc0-inner, virsh)
for the pending retire-or-keep ruling
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|

Step E netem DONE -- G10 CLOSED; netem-link local mode; G16 opened (office1 channels residual)
...
- modules/netem-link: optional local execution (empty ssh target -> bare
sudo tc; D-128 outer root runs ON vcloud); NEW tests/netem-link harness
(12 cases); gauntlet 76 ALL GREEN
- outer root: module netem_vr1_dc0_vr1_dc1 wired -- virbr5 (measured at
wire time), placeholder profile 'delay 3ms 1ms loss 0.01%' (S6 lean,
PROVISIONAL; D-100 gap #11 unruled); runbook Step-11 leg divergence
flagged, DOCFIX queued
- 1/1/0 STOP honored; operator RULED 'Targeted netem apply (Recommended)';
saved -target plan exact 1/0/0, applied; qdisc live on virbr5,
virbr7/virbr3 untouched (docs/audit/stepE-netem-20260721.txt +
outer-{plan,apply}-20260721-netem*.txt)
- Convergence 0/1/0 = office1 channels=[] residual only
(outer-plan-20260721-postE-residual.txt) -> split per E3 into NEW gate
G16; CURRENT-STATE section 5 expected plan re-recorded 0/1/0 (GA-R1 C1)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Disconnect collapse: predecessor bookend landed; netem-tc install VERIFIED (step E unblocked)
...
- docs/audit/netem-sudo-install-20260721.txt: read-only verification of the
operator-run install (0440 root:root, byte-identical, sudo -n -l exit 0)
- CURRENT-STATE step-E paragraph: install PENDING -> INSTALLED+VERIFIED,
remaining path stated (GA-R1 C1/C2, same commit)
- session-ledger: bounded SESSION CLOSE for the close-and-delivery session
- changelog-20260721-netem-install-verify.md: session changelog w/ reverts;
ledger-scan D-131 status-phrasing blind spot flagged as queued finding
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
Step-E netem sudo: scoped NOPASSWD fragment shipped (operator-ruled)
...
scripts/sudoers.d/netem-tc grants exactly the netem-link module's two tc
verbs per measured mesh bridge (virbr5/7/3, re-measured this session;
drifting-ID caveat + fail-closed property documented). 9-case harness;
gauntlet 75 ALL GREEN. Install is operator-only (visudo-checked) and
pending; then the gated placeholder netem run closes G10 step E.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
Incident docs: 2 appendix-A entries, platform-traps MAC-regen corollary, LP draft
...
Closes the remaining documentation queue from the 2026-07-21
commissioning incident. Appendix-A gains the MAC-drift and
agent-resolver symptoms (verbatim, with checks + recorded fixes);
platform-traps 1e gains the unpinned-MAC regeneration corollary + index
row; LP draft ready for the operator to file (sanitization warning
included). Still queued: stale pod cleanup at stage close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
dc-rack-net installed on the rack: EXIT 0, check 10/10, SOA probe answers
...
Operator-approved install executed; capture committed. Rack legs now
reboot-persistent (dc0-rack-legs.service); forwarder regenerated and
proven (authoritative maas-internal SOA via 10.12.8.3). Interim
hand-placed state fully superseded.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
dc-rack-net.sh: D-131 sub-1 delivery -- rack legs + node-DNS forwarder, repo-carried
...
Site-keyed (dc0 rows measured 2026-07-21), runs on the rack via piped
ssh. Legs move from reboot-lost bare `ip addr` into a oneshot unit;
bridges resolved from libvirt NET NAMES at runtime (dc-planes does not
pin bridge names -- virbrN is a drifting ID); dns unit Requires the legs
unit. 14-case harness; gauntlet 74 ALL GREEN. Rack install stays gated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
D-131 sub-1 RULED: rack node-DNS forwarder is the STANDING per-DC pattern
...
GA-R5 record: question + exact utterance ("Standing per-DC pattern
(Recommended)") in the D-131 Status line. Delivery (site-keyed unit +
config + install/check script + harness, folding in rack-legs
persistence) follows in this session. Sub-decisions 2-4 remain OPEN.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
MAC-pin apply executed on voffice1: exact 0/9/0, convergence zero diff
...
Saved-plan apply (operator-approved): 54 MAC pins adopted into the inner
state; zero power flips (guard held), zero replaces. Post-apply verified
all 9 domains still shut off, MACs unchanged. Captures: guarded plan +
apply + convergence in docs/audit/. Node NIC MACs now config-pinned end
to end -- the 2026-07-20 drift class is closed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
node-vm: MAAS owns node power -- ignore_changes on running (operator-ruled)
...
The MAC-pin verification plan (docs/audit/inner-plan-20260721-macpin.txt,
0/9/0, 54 mac adoptions, zero replaces) exposed 9 entangled out-of-band
power-ons: the module re-asserts running=true against nodes MAAS holds
OFF. Operator-ruled: guard first, then apply.
- lifecycle ignore_changes = [running]: create still boots (PXE
enlistment); afterwards MAAS owns power (virsh here, Roosevelt IPMI).
- tests/node-vm +3 cases (guard present, no scope creep, cites MAAS);
15/15 green, gauntlet 73 ALL GREEN, repo-lint 0 fail.
- CURRENT-STATE: MAC-pinning delivery status + pending gated apply.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
node-vm: pin NIC MACs (54 measured) -- in-place applies can no longer drift them
...
Queued delivery item from the 2026-07-21 commissioning diagnosis: an
"in-place" apply (0/9/0) regenerated all 9 unpinned boot MACs on
2026-07-20 and stranded the fleet in MAAS.
- modules/node-vm: optional interface_macs (all-or-nothing + format
validations); mac = { address = ... } shape verified against the
dmacvicar/libvirt 0.9.8 schema dump this session.
- vr1-dc0-substrate: all 9 nodes x 6 planes pinned to values MEASURED
this session (virsh domiflist on vvr1-dc0); metal-admin entries
cross-checked vs MAAS boot-interface records 9/9.
- NEW tests/node-vm harness (12 cases). Gauntlet 72 -> 73 ALL GREEN;
repo-lint 0 fail; opentofu-validate PASS.
NOT YET APPLIED on voffice1 (inner root state) -- the gated apply's plan
must show NO replace; any replace is a STOP.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|

Session close 2026-07-21: commissioning RESOLVED 9/9 Ready; D-131 proposed; SEC-014 opened
...
Lands the ops-commissioning-diag session's deliverable, whose close
bookend was lost to a session disconnect. Landed same-day by the
successor session after read-only re-verification (9/9 Ready, forwarder
active+enabled, dhcpd on both controllers, Phase-7 tag cleanup done).
- CURRENT-STATE: commissioning resolved (two stacked faults: MAC drift
from the in-place apply; MAAS 3.7 rack-only agent resolver SERVFAIL);
G10 remainder = netem step E only.
- design-decisions: D-131 PROPOSED (rack-only node DNS strategy).
- security-ledger: SEC-014 OPENED (rack cluster secret exposure).
- committee doc: closed by addendum.
- changelog + adjudication capture + as-executed index row.
- session-ledger: bounded close entry + machine-derived block re-seeded
from the 2026-07-21 scan (D next-free 132, DOCFIX 197, BUNDLEFIX 052,
10 open SEC rows).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STHXiHfoxHqq8fGRVvb66G
|
| 2026-07-20 |
Fix committee doc: ASCII byte (em-dash in title) + CURRENT-STATE pointer (L10)
...
Prior commit 34c8770 pushed with a non-ASCII em-dash in the committee doc
title and a decoupled L10; both fixed here. Correct lint gate learned: repo-lint
exits 0 clean / 1 FAIL / 2 warn-only, so the right gate is exit != 1 (the
permanent legacy design-decisions WARN makes exit 2 the normal clean state) --
not the exit==0 I mis-used, and not the piped-tail that masked real exit-1 fails.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|

Diagnostic committee (4 reviewers) refutes marginal-timeout 4/4; MTU exonerated by measurement
...
Committee synthesis: docs/audit/commissioning-committee-20260720.md. Consensus:
30-min silence = a HANG not slow progress, so raising node_timeout is the wrong
knob; the proposed one-node test was confounded (timeout + concurrency changed
together). Process errors caught: measured rack not region, tcpdump doubly
mis-scoped (port + window), enlistment paradox resolved (MAAS logs script
COMPLETION so a hung commissioning-only script = identical silence). Pivotal
zero-cost reads: metal-admin MAAS VLAN MTU = 1500 -> guest never jumbo -> MTU
branch EXONERATED; region ample RAM, Temporal quiet at rest (175k wedge errors
were pre-restart). Live causes now: region Temporal starvation during a run,
and a commissioning-only script hang on nested-virt hardware. Decisive gated
test: one commission + console=ttyS0 kernel opt + full-window lease-IP capture
on virbr2+enp1s0.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
Clean commissioning experiment: contention refuted; leading hypothesis = marginal 30-min timeout over slow nested I/O
...
Genuinely isolated node (0 others running, 397GiB free, terminal start
state, read-only polling) still failed at 1770s = 29.5/30 min. Refutes
contention but consistent with INHERENTLY SLOW; batch 3-pass/6-fail-at-the-
mark is the classic marginal-timeout signature. 3 nodes reached Ready on the
same rack/subnet/metadata path, so metadata is not globally broken (:5248 up,
rack->region 301; a :5248 tcpdump was inconclusive -- window closed before
the fetch stage). Recommended test needs an operator decision (MAAS-wide
config): raise node_timeout, commission one node. Stopped live diagnosis per
the circuit-breaker.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
appendix-A: new incident class -- MAAS says dhcp_on=True but no dhcpd runs (Temporal wedged)
...
Records the 2026-07-20 incident by verbatim symptom: config vs running state
disagree, MAAS's own service_set reports healthy, and the region journal
shows Temporal 'Not enough hosts to serve the request'. Check is pgrep -c
dhcpd on EVERY controller; fix is snap restart maas on the region, verified
behaviorally. Includes the scope trap that made it invisible for days: an
already-built site keeps working because deployed VMs hold their addresses,
so it only surfaces when new nodes try to PXE.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
opentofu-validate: cover EXTRA roots (vr1-dc0-substrate, vr1-dc0-maas) -- second blind spot closed
...
The gate validated the outer root + every module and NOTHING ELSE, so both
applied roots shipped while it reported ALL GREEN over them -- the same
false-confidence failure the module loop exists to prevent. Roots are
DISCOVERED (any dir with a *.tf declaring terraform{}), not hardcoded, so a
future vr1-dc1-* root is covered the day it appears. Validated in a copy of
the WHOLE tree, because these roots reference ../modules by relative path
(copying the root alone gives 'Unreadable module directory' -- measured).
Gitignored *.auto.tfvars are stripped from the temp tree (may hold secrets).
VERIFIED the new check can FAIL: a deliberately broken root produced
[FAIL] vr1-dc0-maas before restoring. Harness 14/14, lint 0 fail.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
SEC-012 + SEC-013: ledger rows for the two credentials this deploy created (obligation was outstanding)
...
SEC-012: dedicated MAAS->libvirt service key. Held by a daemon on BOTH
controllers (MAAS dials the pod from the REGION, measured), authorizes a
libvirt-group user, so it can drive every inner node VM and the DC edge.
Rotation obligation + a SCOPE question: libvirt-group is broader than the
power verbs MAAS needs; tighter grant is the hardening candidate and is
Roosevelt-relevant (IPMI has the same power-only-vs-full-control question).
SEC-013: admin-scoped MAAS API key materialized to disk on the region;
surfaces include any maas-provider tfstate and a CLI profile. Verified by
format only, never printed. Tied to whether vr1-dc0-maas is retired.
Open SEC count 7 -> 9; G14 row updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|

node-vm console shipped (0/9/0 in-place); commissioning narrowed, UNRESOLVED; two of my own experiments disclosed as INVALID
...
Console added + log dir parameterized, but logs stay 0 bytes: firmware
writes to VGA, so serial alone does not make a PXE-booting node observable
-- correction queued (needs graphics + virsh screenshot, or a MAAS kernel
cmdline change). Solid: PXE + ephemeral handoff work; ephemeral OS boots
with working networking (leases + NTP to rack); no region/metadata traffic
seen; memory, image sync, DHCP and shape all ruled out. SEC-010 blocking
node->region is SUSPECTED but unreadable (rules lack counters).
INVALID EXPERIMENTS RECORDED: the 'single node alone' test was not alone
(5-6 nodes still up, proven by tcpdump), and the retry did not restart
MAAS's timer while I destroyed domains mid-run. Contention is therefore a
LIVE hypothesis again, not refuted. Stopped per the circuit-breaker rather
than iterate on contaminated evidence.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
modules/node-vm: add serial console + boot log (parameterized dir)
...
Its absence blocked a live diagnosis: 6/9 dc0 nodes stall in MAAS 'Loading
ephemeral' and, with no console and no agent, a stuck node is unobservable.
Same pattern opnsense-edge already carries, for the same reason -- a
PXE-booting node emits everything interesting before anything sshable
exists. Unconditional (MAAS-managed nodes are rebuilt at will; no
live-router bounce risk like the D-129 qga channel). Log dir is a variable
rather than opnsense-edge's hardcoded vcloud literal, which is wrong on any
other libvirt host -- backport queued.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
OPEN: 6/9 nodes fail commissioning; contention refuted; blocked on node-vm having no console
...
3 Ready, 6 timed out at 30 min. Hypotheses tested and REFUTED by
measurement: contention (single node alone stuck in Loading ephemeral 20+
min with 385GiB free, load 0.04), rack boot-image sync (status synced, 10
images), DHCP (same nodes enlisted fine earlier). Real blocker is
observability: modules/node-vm defines no serial console, so a stuck node
is a sealed box -- the same gap already logged for cloudinit-vm, now biting
a live diagnosis. Fix is the opt-in serial+log pattern opnsense-edge
already carries; NOT applied, since it re-applies domains mid-diagnosis and
is an operator ruling. Step E prerequisites measured and recorded (dc0-dc1
mesh = virbr5 ON VCLOUD; netem-link assumes passwordless sudo vcloud lacks).
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
STEP D COMPLETE: per-machine power tooling shipped; commissioning works, nodes reaching Ready
...
Operator-ruled per-machine virsh power. maas-node-power.sh + harness (24/24,
gauntlet 72 ALL GREEN): MAC-matched because MAAS renames machines at
enlistment, dry-by-default, every write verified by a real query-power-state.
Region prereqs fixed (no virsh there; maas profile was root-only -- created
via stdin login, key never printed). All 9 nodes powered and commissioning
end-to-end: 3 Ready / 6 Commissioning, shapes exact to D-121 Option C.
vr1-dc0-maas root retained but unused pending the stage-close D-103/D-123
amendment. G10 remainder: netem only.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
NEW: maas-node-power.sh (+harness 24/24) -- per-machine virsh power, MAC-matched
...
Operator-ruled 2026-07-20 over the pod route. Sets power_type=virsh +
power_address/power_id on already-enlisted machines, matching domains BY MAC
(MAAS renames machines at enlistment, so name matching would map nothing or
the wrong node). Dry by default; every write verified by an actual
query-power-state, since a stored parameter proves nothing about reachability.
Carries the measured reasoning: pods fail on node-vm's pool+volume disks
(domblkinfo 'missing storage backend', reproduced locally), the pod's
discovery role was already done by PXE, and per-node power is the Roosevelt
(IPMI) shape. Gauntlet 72 harnesses ALL GREEN.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
CURRENT-STATE: record step-D pod findings + pending ruling (clears L10 on prior commit)
...
Prior commit landed an audit capture without the same-commit CURRENT-STATE
update that GA-R1/C1 requires; L10 caught it. Position now records: pods
incompatible with node-vm volume-ref disks (reproduced locally), pod
unnecessary since PXE already discovered the nodes, per-machine virsh power
measured working, ruling pending.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|

Pod path chased to root cause: MAAS virsh pods incompatible with node-vm volume-ref disks; per-machine power PROVEN working
...
Dedicated MAAS->libvirt key minted+wired (operator-ruled). Two further
measured findings: (1) MAAS dials the pod from the REGION, not the rack, so
the credential belongs where MAAS dials from; (2) each 503 created a broken
pod server-side outside tofu state (both deleted). Final blocker: domblkinfo
fails on pool+volume disk refs -- reproduced LOCALLY on the rack with an
active pool, so it is libvirtd, not snap/SSH. Making pods work would mean
converting node-vm to file-path disks + re-applying 9 domains. Not needed:
the pod's job was discovery, which PXE already did, and per-machine
power_type=virsh query-power-state returns state=off on the canary. Per
machine power is also the Roosevelt shape (IPMI per node). Awaiting ruling.
CURRENT-STATE untouched pending the ruling.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
Step D part 2: MAAS root shipped; apply refutes D-123's local qemu:///system pod mechanism
...
New opentofu/vr1-dc0-maas root (isolates MAAS creds from substrate roots
per DOCFIX-179's lesson); plan clean, apply failed 'Failed to login to
virsh console'. MEASURED root cause: confined MAAS snap gets Permission
denied on libvirt-sock and ships NO libvirt interface to connect -- a local
qemu:///system pod is architecturally impossible with snap MAAS. D-123
Model B and modules/maas-vm-host both state that mechanism; intent (no
cross-fiber dial) survives, mechanism does not. Amendment pending a ruling.
qemu+ssh replacement measured feasible (snap ships ssh; rack user in
libvirt group). Nothing was created by the failed apply.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
NEW root opentofu/vr1-dc0-maas: MAAS VM-host registration, isolated from substrate roots
...
Third root by design, applying DOCFIX-179's own lesson rather than
repeating it: a provider 'maas' block forces EVERY plan in its root to
demand maas_api_url + the sensitive key, so coupling it to the substrate
roots would make routine substrate plans require MAAS creds. This root
touches MAAS only; the substrate roots stay credential-free. Registers
the pod via modules/maas-vm-host (discover, never compose -- D-103) with
power_address local to the rack per Model B. Validated explicitly (the
validator covers only the outer root + modules -- extending it to all
roots stays queued).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|