diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 9a726cc..88bbbbf 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -131,7 +131,37 @@ to D-121 Option C. `opentofu/vr1-dc0-maas/` is retained but UNUSED (the pod route is refuted); retire-or-keep is a stage-close question, as is the D-103/D-123 amendment text. Session changelog items 15-17. - **COMMISSIONING PARTIAL (open, 2026-07-20): 3 nodes Ready, 6 timed out.** + **COMMISSIONING RESOLVED 2026-07-21: all 9 nodes Ready** (logged window + ops-commissioning-diag; adjudication + `docs/audit/commissioning-diag-20260721.txt`; session changelog + 2026-07-21). TWO stacked faults, both measured: (1) the 2026-07-20 + in-place serial-console apply REGENERATED all 9 node NIC MACs + (tofu-reported 0/9/0 in-place), so MAAS's records went stale and every + post-apply boot was an unknown node -- no PXE event, no tag kernel_opts, + silent 30-min timeout; repaired operator-ruled via per-machine + boot-interface MAC update (mark-broken/update/mark-fixed where needed), + read-back verified 9/9. (2) Beneath it, the MAAS 3.7 RACK-ONLY agent + resolver SERVFAILs every query on an internet-isolated rack (walks + public root hints even for its own authoritative maas-internal zone; + ignores resolv.conf), so cloud-init's cloud-config-url never resolved + and nodes booted to a login prompt without ever fetching commissioning + scripts. Office1/VR0 were immune (co-located region BIND owns node DNS) + -- this surface is FIRST EXERCISED in VR1; LP report queued. + Operator-ruled workaround, live and proven: `dc0-node-dns.service` on + the rack (dnsmasq on virbr2 alias 10.12.8.3 forwarding to region BIND + over the rack's OWN transit connection; SEC-010 re-verified enforced + and untouched) + metal-admin subnet dns_servers=10.12.8.3, + allow_dns=false. PROOF: canary Ready in ~3 min after seven consecutive + 30-min failures, commissioning scripts visible on serial; fleet of 8 + re-commissioned concurrently, ALL 9 READY in ~4 min, shapes exact to + D-121 Option C. Committee record closed by addendum (its mechanisms + were wrong; its instrument found the cause). D-131 PROPOSED (rack-only + node DNS strategy -- Roosevelt-relevant). SEC-014 OPENED (rack cluster + secret exposure during diagnosis). Queued delivery: MAC pinning in + modules/node-vm, forwarder + rack-legs repo-carried persistence, + appendix-A entries, LP report, stale pod object cleanup. + History of the diagnosis (superseded; kept for the audit trail): + the 2026-07-20 state read "3 nodes Ready, 6 timed out." Established: PXE and the ephemeral handoff WORK, and the ephemeral OS boots with working networking (nodes hold leases and do NTP to the rack) -- it simply never completes. Ruled out by measurement: memory, rack boot-image @@ -164,10 +194,10 @@ re-runnable; nothing is lost. The committee record (`docs/audit/commissioning-committee-20260720.md`) is the durable authority for this diagnosis and its ranked live hypotheses. - REMAINING IN G10: this diagnosis, then netem (step E, + REMAINING IN G10: netem only (step E, NOT started -- target is the dc0<->dc1 mesh = **virbr5 on vcloud**, measured; `modules/netem-link` assumes PASSWORDLESS SUDO which vcloud's - operator account lacks). Session changelog item 18. + operator account lacks). Session changelog item 18 (2026-07-20). - The grounding audit is COMPLETE and EXITED (2026-07-19): Phases 1-6 all closed (charter `148dcef`; rulings `docs/audit/ga-rulings.md`; the Phase-5 sweep ran as six operator-gated batches in one session; exit @@ -354,7 +384,7 @@ | G7 | New captured plan == the expected triple recorded in section 5 | [V] re-plan to a capture file after G5+G6 | session | CLOSED 2026-07-19: capture `docs/audit/outer-plan-20260719-postG6.txt` = 6/0/6, equals section 5 exactly | | G8 | Same-session pre-apply re-verify: 6 planes still empty | [V] run in the SAME session as the apply | session | CLOSED 2026-07-19: verified in the apply session itself (all six 0 leases; only office1 nets attached) immediately before step A | | G9 | DC0 outer apply (deploy step A) | [V] operator-gated, logged (`run-logged.sh`), after G1-G8; audit exit criteria met (charter Phase 6). SEC pre-apply dependency (S2): SEC-010's transit FORWARD-drop is applied+verified at deploy step B via `site-headend-install.sh --host-nodes --check` on vvr1-dc0 (gate G10) -- the ONLY SEC row gated on this apply (register of record: security-ledger). CANONICAL ENTRY DOC (probe hole H1): `runbooks/dc-dc-phase2-tofu-dc-substrate.md`, with `docs/dc0-deploy-readiness.md` section E as the step table | operator | CLOSED 2026-07-19: G8 same-session planes check passed (6x 0 leases, 0 attachments); saved plan == 6/0/6 applied in the logged dc0-deploy window; convergence re-plan = no differences; vvr1-dc0 running, prior guests untouched | -| G10 | Deploy steps B-E in-sequence gates: SEC-010 `--host-nodes --check` on vvr1-dc0; depth-4 nested boot; D-125 foreign-MAC egress test; MAAS reachability + `TF_VAR_maas_api_key` before step D; netem placeholder step E | [V] exercised during the gated deploy | session (each mutation operator-approved) | Step B DONE 2026-07-20 (`--check` EXIT 0 incl. SEC-010, `docs/audit/stepB-check-20260720-final.txt`; interfaces enp1s0/enp2s0). Depth-4 nested boot DONE (10 domains running inside vvr1-dc0). D-125 egress isolation test PASS 2026-07-20 (`docs/audit/d125-egress-gate-20260720-matrix.txt`), and the edge itself now egresses 0% loss after the v4 addressing. REMAINING: D-129 plugin profile on 26.7, MAAS reach + key + region-side metal-admin DHCP (step D), netem (step E) | +| G10 | Deploy steps B-E in-sequence gates: SEC-010 `--host-nodes --check` on vvr1-dc0; depth-4 nested boot; D-125 foreign-MAC egress test; MAAS reachability + `TF_VAR_maas_api_key` before step D; netem placeholder step E | [V] exercised during the gated deploy | session (each mutation operator-approved) | Step B DONE 2026-07-20 (`--check` EXIT 0 incl. SEC-010, `docs/audit/stepB-check-20260720-final.txt`; interfaces enp1s0/enp2s0). Depth-4 nested boot DONE (10 domains running inside vvr1-dc0). D-125 egress isolation test PASS 2026-07-20 (`docs/audit/d125-egress-gate-20260720-matrix.txt`), and the edge itself now egresses 0% loss after the v4 addressing. Step D COMPLETE incl. commissioning: ALL 9 NODES READY 2026-07-21 (two stacked faults diagnosed + fixed -- `docs/audit/commissioning-diag-20260721.txt`; section 1). REMAINING: netem (step E) | | G11 | Operator signs THIS document | [R] read top-to-bottom; discrepancies resolved in the document | operator | CLOSED: RE-SIGNED 2026-07-19 at audit exit, section 11 (replaces the 2026-07-18 signature) | | G12 | `vr1-dc1` build | [R] operator rules dc1 transit/rack addressing; then vars + substrate authored | operator | HELD (`docs/dc0-deploy-readiness.md:100-103`) | | G13 | D-129 residuals | [R] operator-gated live plugin install on the edge; qga channel retrofit at next scheduled edge restart; 4 sub-decisions (section 8) | operator | OPEN / PARTIALLY RULED (`docs/design-decisions.md:4017`) | diff --git a/docs/audit/commissioning-committee-20260720.md b/docs/audit/commissioning-committee-20260720.md index d1cacb8..16c38b0 100644 --- a/docs/audit/commissioning-committee-20260720.md +++ b/docs/audit/commissioning-committee-20260720.md @@ -139,3 +139,29 @@ Web refs: LP #1807252, LP #2084788, LP #1908452, LP #1403955, MAAS 3.7 release notes, maas.io controller-communication docs. + +## OUTCOME ADDENDUM (2026-07-21 -- diagnosis complete; see +## docs/audit/commissioning-diag-20260721.txt for the full chain) + +The prescribed instrument (full-window capture + console=ttyS0) was run and +FOUND the cause on its first execution. Scorecard against this record: + +- CORRECT: refuting the marginal-timeout hypothesis (nothing was slow); + "observe, don't change the clock"; the instrument prescription itself; + the confounded-test objection. +- WRONG (all four reviewers): every ranked mechanism. There was no hang -- + TWO stacked faults: (1) all 9 node NIC MACs regenerated by the + 2026-07-20 in-place serial-console apply, so MAAS treated every + post-apply boot as an unknown node (the "silence" was an unrecognized + machine, booting fine); (2) underneath, the MAAS 3.7 rack-only agent + resolver SERVFAILs all queries on an internet-isolated rack, so + cloud-init's maas-internal cloud-config-url never resolved (nodes booted + to a login prompt, never fetched scripts). Temporal was quiet the whole + instrumented window (136 lines, 0 errors). +- The "same-3 vs rotating-3" pivotal unknown resolved as: the 3 passed + BEFORE the MAC-drifting apply -- deterministic, but by timeline, not + per-node hardware. + +Both faults measured, ruled, and repaired/worked-around 2026-07-21 +(canary Ready in ~3 min after seven consecutive timeouts). This addendum +closes the committee record. diff --git a/docs/audit/commissioning-diag-20260721.txt b/docs/audit/commissioning-diag-20260721.txt new file mode 100644 index 0000000..d1028f0 --- /dev/null +++ b/docs/audit/commissioning-diag-20260721.txt @@ -0,0 +1,90 @@ +commissioning-diag-20260721 -- instrumented diagnosis adjudication +Logged window: ~/as-executed/2026-07-21-ops-commissioning-diag.log (all +commands + outputs verbatim there; this file is the citable summary). +Committee charter run: docs/audit/commissioning-committee-20260720.md. + +VERDICT: two stacked faults, both measured, both fixed/worked-around. + +FAULT 1 -- MAC drift (all 9 nodes) +- Instrumented run (07:17:49 -> 07:48:51 timeout): domain restarted + (libvirt Id 39 -> 40, power-cycle works) but PXE DHCP came from + 52:54:00:2b:ed:ab -- absent from MAAS's record for hardy-dove + (52:54:00:c7:da:bc). MAAS saw an unknown node: no "Performing PXE boot" + event, no diag-serial tag kernel_opts served (serial log 0 bytes all + window), no association -- 30-min timer expired on a machine record + whose real VM booted normally. +- Event-timeline boundary: 2026-07-20 17:40 + 18:16 runs show PXE events; + every run after the in-place serial-console apply to all 9 node-vm + domains (session changelog 2026-07-20 item 19; tofu plan said 0/9/0 + in-place) shows none. The apply regenerated every metal-admin NIC MAC. +- Full-fleet audit: 9/9 live MACs differed from MAAS boot-interface + records (pairing per docs/audit/stepD-power-20260720.txt). +- Re-frames: "3 pass / 6 fail" = the 3 finished BEFORE the apply; item-20 + "isolated 1770s failure" = first post-apply run. Committee's + timeout-refutation stands; its ranked mechanisms do not (Temporal: 136 + journal lines, 0 errors during the full window). +- REPAIR (ruling 2026-07-21, "Update MACs in place (Recommended)"): + boot-interface MAC updated per machine to the measured live value; + Failed-state machines via mark-broken -> interface update -> mark-fixed + (MAAS requires New/Ready/Allocated/Broken). Read-back verified 9/9. + +FAULT 2 -- MAAS 3.7 rack-only agent resolver defect (surfaced by fixing 1) +- Post-repair proof run: PXE recognized t+14s; Ubuntu 24.04.4 boots to + login in ~54s (squashfs URL is IP-based, transfers fine); then + cloud-init: "Did not find any data source, searched classes: ()". + cloud-config-url uses http://10-12-8-0--22.maas-internal:5248/...; + DHCP option 6 pointed nodes at the rack agent resolver (10.12.8.2:53). +- Agent SERVFAILs EVERY query (incl. its own authoritative zone), + correlated per-query in its journal: "Failed to handle authoritative + query error=dial udp :53: connect: network is + unreachable". Rack has no internet (by design). Agent ignores + /etc/resolv.conf (proven: symlink to real upstream list + snap restart + -> unchanged); resolver config is Temporal-pushed, nothing on disk. +- Office1 immune: agent binds no :53 there; region BIND owns node DNS on + co-located controllers. The rack-only resolver is FIRST EXERCISED in + VR1 and is defective on isolated racks. LP report queued. +- Region BIND (10.10.0.20, reachable over transit) answers + 10-12-8-0--22.maas-internal -> 10.12.8.2 and hardy-dove.maas correctly. +- SEC-010 re-verified ENFORCED mid-diagnosis: dedicated nft chain, + priority filter-10, blanket enp1s0 FORWARD drop both directions. This + rules out (by design) any remedy that forwards node traffic to the + region; rulings below stay inside the boundary. +- Interim ruling "Region upstream, runtime": resolvectl on rack enp1s0 + (kept, harmless for the rack's own resolution) -- did NOT affect the + agent. Superseded by: +- WORKAROUND (ruling 2026-07-21, "Rack dnsmasq forwarder (Recommended)"): + dc0-node-dns.service on the rack -- dnsmasq bound to virbr2 alias + 10.12.8.3 (ip addr replace in ExecStartPre; After=libvirtd), no-resolv, + server=10.10.0.20; rack OUTPUT traffic over transit, SEC-010 untouched. + metal-admin subnet (id 6): dns_servers=10.12.8.3, allow_dns=false. + NOTE: dhcpd still templates option 6 as "10.12.8.2, 10.12.8.3" (rack IP + prepended despite allow_dns=false); acceptable -- stub resolvers fail + over on SERVFAIL, proven by the green run below. Template behavior to + be investigated at proper delivery. + +PROOF +- Canary hardy-dove: Ready in ~3 minutes (after seven consecutive 30-min + failures); commissioning scripts visible on serial (maas-lshw etc., all + rc=0). Serial observability itself was enabled by the diag-serial tag + (kernel_opts console=ttyS0,115200) -- the committee's instrument. +- Fleet: remaining 8 re-commissioned (5 had never truly commissioned -- + mark-fixed does not commission; 3 carried stale interface data). + RESULT: see CURRENT-STATE + session changelog (recorded at terminal + state of the fleet run). + +OPEN / HONEST UNKNOWNS +- How the 3 nodes passed on 2026-07-20 evening: dhcpd.conf has been + regenerated since; that day's option-6 value is unrecoverable. Does not + affect the fixes; noted for completeness. +- Stale MAAS pod object "vr1-dc0-inner" (id 4) still registered (pod + route was refuted) -- stage-close cleanup question with SEC-013. + +SEC +- SEC-014 OPENED: /var/snap/maas//agent.yaml read into session + context (and as-executed log) included the rack cluster secret. + Rotation = rack re-enrollment; see docs/security-ledger.md row. + +Artifacts (host-side, not committed): rack /home/jessea123/ +diag-{metal,transit}-20260721.pcap, region /home/jessea123/ +diag-temporal-20260721.log, rack serial log +/var/lib/libvirt/vr1/staging/vr1-dc0-storage-04-serial.log. diff --git a/docs/changelog-20260721-commissioning-diag.md b/docs/changelog-20260721-commissioning-diag.md new file mode 100644 index 0000000..3588f9f --- /dev/null +++ b/docs/changelog-20260721-commissioning-diag.md @@ -0,0 +1,132 @@ +# Session changelog 2026-07-21 -- commissioning diagnosis + repair (logged: ops-commissioning-diag) + +One changelog per session (GA-R2/D1). Status lives ONLY in +docs/CURRENT-STATE.md; this file is narrative + reverts. Full adjudication: +docs/audit/commissioning-diag-20260721.txt. As-executed log: +~/as-executed/2026-07-21-ops-commissioning-diag.log. + +## 1. Instrumented commissioning run (committee-prescribed) EXECUTED + +- Phase 1 zero-cost reads, then gated: diag-serial MAAS tag + (kernel_opts console=ttyS0,115200) on canary sw6wqe/hardy-dove only; + serial-console tag temporarily DETACHED from the canary (alphabetical + tag concatenation would have left console=tty0 last, sending userspace + output to VGA). Instruments: MAC+broadcast tcpdump on virbr2, full + tcpdump on enp1s0, live Temporal journal follow -- all 45-min capped. +- One commission, hands-off, full window. Result: timeout at 30 min, + serial 0 bytes -- and the pcap held the answer (item 2). +- FINDING (record divergence): a serial-console tag with console + kernel_opts ALREADY existed (comment cites D-114 debug) contradicting + the committed "ephemeral kernel is not told console=ttyS0" claim. +- **Revert:** diag-serial tag delete + re-attach serial-console to sw6wqe + (Phase 7, item 8). + +## 2. FAULT 1 FOUND + REPAIRED: all 9 node MACs drifted (in-place apply) + +- PXE DHCP in the capture came from a MAC absent from MAAS's records; + fleet audit: 9/9 live metal-admin MACs differ from MAAS boot-interface + records. Boundary: the 2026-07-20 in-place serial-console apply + (tofu 0/9/0) regenerated NIC MACs. Full chain in the adjudication file. +- RULING (GA-R5, 2026-07-21): question = how to repair the records; + operator selection = "Update MACs in place (Recommended)". +- Executed: per-machine boot-interface MAC -> measured live value; + Failed-state machines needed mark-broken -> update -> mark-fixed. + Read-back verified 9/9. +- **Revert:** maas admin interface update + mac_address= per + machine (old values in the as-executed log); machines return to their + pre-session states via mark-broken/mark-fixed as needed. + +## 3. FAULT 2 FOUND: MAAS 3.7 rack-only agent resolver defect + +- Post-repair proof run booted fully; cloud-init found no datasource: + the agent resolver (node DNS via DHCP) SERVFAILs everything on an + isolated rack -- walks public root hints even for authoritative + maas-internal; ignores resolv.conf. Office1 immune (region BIND owns + node DNS there). First-exercised surface in VR1. LP REPORT QUEUED + (item 10). SEC-010 re-verified enforced during diagnosis. +- Interim RULING "Region upstream, runtime": resolvectl dns enp1s0 + 10.10.0.20 + domain ~. on the rack + /etc/resolv.conf relinked to + /run/systemd/resolve/resolv.conf. Did NOT affect the agent (kept -- + gives the rack itself working resolution via region). +- **Revert:** resolvectl revert enp1s0; ln -sf + ../run/systemd/resolve/stub-resolv.conf /etc/resolv.conf. + +## 4. FAULT 2 WORKAROUND: dc0-node-dns forwarder (operator-ruled) + +- RULING (GA-R5, 2026-07-21): question = node DNS remedy; operator + selection = "Rack dnsmasq forwarder (Recommended)". +- Shipped to the rack (NOT yet repo-carried -- item 9): + /etc/dnsmasq-dc0-node.conf + /etc/systemd/system/dc0-node-dns.service + (dnsmasq on virbr2 alias 10.12.8.3, no-resolv, server=10.10.0.20; + alias via ExecStartPre ip addr replace; After=libvirtd). Enabled. +- MAAS: subnet 6 (metal-admin) dns_servers=10.12.8.3 + allow_dns=false. + dhcpd still prepends the rack IP (option 6 = "10.12.8.2, 10.12.8.3") + -- works via SERVFAIL failover (proven); template behavior to + investigate at delivery. +- **Revert:** maas admin subnet update 6 dns_servers= allow_dns=true; + on rack: systemctl disable --now dc0-node-dns; rm + /etc/systemd/system/dc0-node-dns.service /etc/dnsmasq-dc0-node.conf; + ip addr del 10.12.8.3/22 dev virbr2. + +## 5. PROOF + fleet re-commission + +- Canary hardy-dove: Ready in ~3 min after seven consecutive 30-min + failures; commissioning scripts visible on serial, rc=0. +- Fleet: remaining 8 commissioned (5 never truly commissioned -- + mark-fixed does not commission; 3 carried stale interface data). + Outcome recorded in CURRENT-STATE (same commit). +- **Revert:** n/a (commissioning is re-runnable). + +## 6. SEC-014 OPENED: rack cluster secret exposed to session context + +- agent.yaml read (diagnosing the resolver) included the `secret` field. + Exposure scope: Claude session context, operator terminal, as-executed + log (0600, jumphost). Row + disposition: docs/security-ledger.md. + Rotation = rack re-enrollment; schedule at operator convenience. +- **Revert:** n/a (exposure event; row tracks rotation). + +## 7. Session tooling: targeted ask rules (classifier walls) + +- .claude/settings.local.json (untracked, global-gitignored): ask rules + added for wrapped/unwrapped `maas admin tags|tag|machine|interface| + subnet` via voffice1 and rack-sudo via the dc0 key -- per the standing + multi-workstation pattern (operator-approved this session). Lesson: + a `sleep N;` prefix breaks prefix-matching -- keep wrapped commands + rule-shaped. +- **Revert:** remove the ask entries from settings.local.json. + +## 8. Phase 7 cleanup (pending fleet-green) + +- diag-serial tag: detach from sw6wqe + delete; re-attach serial-console. + QUEUED OBSERVATION: serial-console's opts end in console=tty0, which + claims /dev/console for VGA -- fine for kernel logs, hides userspace; + revisit the tag's opts when serial observability next matters. +- **Revert:** n/a (is itself the revert). + +## 9. QUEUED delivery work (change-delivery loop, not this session) + +- modules/node-vm: PIN NIC MACs (set to current live values) so no apply + can drift them again; harness case for MAC stability. Roosevelt note: + metal has fixed MACs; the trap is VR-only. +- platform-traps entry: tofu/libvirt "in-place" 0/N/0 domain change can + REGENERATE NIC MACs (measured 2026-07-20/21). +- appendix-A entries: (a) "commissioning times out at exactly 30 min, + NO 'Performing PXE boot' event" -> diff live domain MACs vs MAAS + records; (b) "cloud-init: no datasource, searched classes ()" on + rack-only controllers -> query the rack agent resolver for + maas-internal; if SERVFAIL see D-131/forwarder. +- dc0-node-dns forwarder + rack-legs persistence: repo-carried + script/module + tests (site-keyed, per D-128 register item 20). +- D-131 PROPOSED (authored this session, design-decisions.md): node-facing + DNS strategy for rack-only controllers (Roosevelt-relevant). +- Committee record addendum: authored this session (outcome section). +- Stale MAAS pod vr1-dc0-inner (id 4): stage-close cleanup with SEC-013. + +## 10. LP report (queued for operator) + +- Title: "maas-agent resolver on rack-only controller SERVFAILs + authoritative maas-internal queries on internet-isolated racks (walks + root hints; ignores resolv.conf); breaks commissioning cloud-config-url" + MAAS 3.7.2 / snap 41649. Evidence: adjudication file + agent journal + excerpts in the as-executed log. diff --git a/docs/design-decisions.md b/docs/design-decisions.md index 8bd911e..4f3eb0a 100644 --- a/docs/design-decisions.md +++ b/docs/design-decisions.md @@ -4173,3 +4173,35 @@ **Delivery:** module edit + `tests/cloudinit-vm/run-tests.sh` harness (static lifecycle-guard assertions; `tofu validate` when the binary is present). **Revert:** remove the lifecycle block and the harness; the replace pair returns on the next post-reboot plan. + +## D-131: node-facing DNS strategy for rack-only controllers [ARCH] + +**Status:** PROPOSED (authored 2026-07-21; interim workaround operator-ruled and live -- see below). + +**Context (measured, 2026-07-21).** MAAS 3.7's rack-only controllers serve node DNS from the +`maas-agent` resolver, which on an internet-isolated rack SERVFAILs ALL queries -- including its +own authoritative `maas-internal` zone -- because it resolves by walking public root hints and +ignores `/etc/resolv.conf` (no on-disk config surface; Temporal-pushed). cloud-init's +commissioning `cloud-config-url` uses a `maas-internal` name, so commissioning fails silently +(node boots to login, never fetches scripts). Co-located region+rack controllers are immune +(region BIND owns node-facing DNS), which is why VR0 and Office1 never hit this. Full chain: +`docs/audit/commissioning-diag-20260721.txt`; LP report queued. + +**Roosevelt relevance (the A1 test).** Roosevelt DCs are rack-only sites behind an edge, exactly +this topology. Any Roosevelt build session must decide node DNS BEFORE first commissioning. + +**Interim ruling (2026-07-21, operator utterance "Rack dnsmasq forwarder (Recommended)"):** +`dc0-node-dns.service` on the rack -- dnsmasq bound to a metal-admin alias (10.12.8.3), +`no-resolv`, `server=`; rack OUTPUT traffic only, SEC-010 untouched; +metal-admin subnet `dns_servers=10.12.8.3`, `allow_dns=false`. Live and proven (canary Ready in +~3 min). NOT yet repo-carried. + +**Open sub-decisions:** +1. Adopt the forwarder as the STANDING per-DC pattern (repo-carried, site-keyed unit + module or + script delivery, harness) vs treat as temporary until the MAAS agent defect is fixed upstream. +2. Whether the forwarder should serve the edge/other planes too, or metal-admin only (current). +3. dhcpd still prepends the rack IP to option 6 despite `allow_dns=false` (relies on client + SERVFAIL failover -- works, proven, but is an extra moving part): investigate template + behavior; possibly LP material of its own. +4. Roosevelt: same forwarder shape, or region-DNS-over-WAN with a firewall exception, or wait + for the upstream fix and require a minimum MAAS version. Evidence needed: LP outcome. diff --git a/docs/security-ledger.md b/docs/security-ledger.md index 8ec10fd..cd96bfd 100644 --- a/docs/security-ledger.md +++ b/docs/security-ledger.md @@ -21,6 +21,7 @@ | SEC-010 | 2026-07-16 | **metal-admin DC-LOCAL invariant (D-052/D-100) is PRESERVED but NOT ENFORCED in the committed Stage-3 config.** The `vvr1_dc0` rack straddles metal-admin (10.12.8.0/22, DC-local) + the office1<->dc0 transit (crosses fiber). Its committed cloud-init pins static IPs only -- **no `net.ipv4.ip_forward=0` sysctl, no host firewall on the transit leg.** Cross-plane routing (metal-admin <-> the whole Office1 /22, bidirectional) is blocked ONLY by Ubuntu's distro default; the deferred MAAS-rack install could silently flip it. A MAAS rack proxies at the application layer and needs no kernel forwarding, so pinning is free. **MODEL B UPDATE (2026-07-16):** after the D-123 Model B reshape, the OUTER `vvr1-dc0` is single-leg (transit only); metal-admin + the other 5 planes are now INNER bridges on `vvr1-dc0`'s own libvirtd. The forwarding hardening is MORE critical (vvr1-dc0 now bridges ALL 6 inner planes + the transit) and belongs on the inner libvirt host, not just a 2-leg rack -- wire it in the C3 bootstrap (`site-headend-install.sh` node-host mode). | 2026-07-16 plane-segregation review + Model B reshape cross-check; `opentofu/main.tf` `module "vvr1_dc0"` | operator | **CLOSED 2026-07-20** (was: artifact COMMITTED Phase D 2026-07-16, gated on apply+verify -- both now done, see close note at end of this cell). Enforced via a FORWARD-drop across the transit leg, NOT a global `ip_forward=0`. `scripts/site-headend-install.sh --host-nodes` writes `/etc/nftables-sec010.nft` (drop FORWARD in+out the transit interface, default `mgmt`) + a boot-persistent `sec010-fw.service`; `--host-nodes --check` is the MECHANICAL pre-apply gate (fails if the rule is absent, AND -- hardened 2026-07-16 -- if the keyed transit interface does not EXIST, since an nftables oifname on an absent iface loads clean but matches nothing = fail-open). **D-125 UPDATE (2026-07-16):** the earlier rationale (a global `ip_forward=0` being unusable because the inner `vr1-dc0-wan` NAT forced `ip_forward=1`) NO LONGER applies -- under bridge-in the inner `vr1-dc0-wan` is a BRIDGE (no inner NAT). The scoped FORWARD-drop is RETAINED as-is (correct either way). **br_netfilter CONSTRAINT:** it MUST stay interface-scoped and NEVER be globalized -- with `br_netfilter` loaded, bridged WAN frames traverse the L3 FORWARD chain, so a global forward-drop would silently kill the `br-vr1-dc0-wan` WAN path. `--check` now also verifies the WAN bridge exists with its uplink enslaved (same fail-open class). Nothing routes across the fiber THROUGH vvr1-dc0; the rack proxies MAAS at the app layer (originated/terminated, not forwarded). Same pin follows onto `voffice1` when its transit leg is wired. Region route must target only the rack transit /30, never 10.12.8.0/22. **CLOSED 2026-07-20 (operator ruling, exact utterance: "Close SEC-010 (Recommended)"): applied + verified BOTH ends in the logged dc0-deploy step-B window.** vvr1-dc0: written by `site-headend-install.sh --role rack --host-nodes`, verified by the named `--check` EXIT 0 (`docs/audit/stepB-check-20260720-final.txt` -- rule present + `enp1s0` exists + `br-vr1-dc0-wan`/`enp2s0` enslaved; note the transit interface is `enp1s0`, not the script-default `mgmt` -- netplan set-name dropped, session changelog). voffice1 (the "same pin follows" clause -- its transit leg `enp2s0` was wired 2026-07-20): identical scoped artifact + `sec010-fw.service` enabled, `nft list table inet sec010` verified, proxy/region paths re-verified working post-pin. | | SEC-012 | 2026-07-20 | **New SERVICE credential: a dedicated MAAS -> libvirt SSH keypair.** Minted during the DC0 deploy so MAAS can power-control the inner node VMs (`power_type=virsh`). It is a SERVICE credential held by a daemon, not an operator key: the private half is installed inside the MAAS snap on BOTH controllers (region + DC rack) because MAAS dials the power address from whichever controller it chooses -- MEASURED to be the REGION, not the rack that owns the hardware. The public half authorizes a user that is in the rack's `libvirt` group, so anyone holding this key can drive libvirt on `vvr1-dc0` -- i.e. power/define/destroy every inner node VM and the DC edge. Blast radius is deliberately narrowed by being per-purpose (operator-ruled 2026-07-20: "Dedicated MAAS->libvirt key" over reusing the dc0 service key). | 2026-07-20 step-D power wiring; `scripts/maas-node-power.sh` header; session changelog items 15-17 | operator | **OPEN -- rotation obligation + a scope question.** (1) ROTATE at v1 close, or immediately if either controller is rebuilt or shared. (2) SCOPE: the key authorizes a `libvirt`-group user, which is broader than the power verbs MAAS actually needs; a tighter grant (dedicated unix user + polkit rule limiting to domain start/stop/status) is the hardening candidate and is Roosevelt-relevant, since bare metal replaces this with IPMI credentials that have the same "power-only vs full-control" question. Custody detail off-repo per D-069. | | SEC-013 | 2026-07-20 | **MAAS API key materialized to disk on the region.** An admin-scoped MAAS API key (`consumer:token:secret`) was placed by the operator in a 0600 file on `voffice1` so the `opentofu/vr1-dc0-maas` root could consume it via `TF_VAR_maas_api_key`. It grants FULL MAAS admin API access (machines, power, deploy, users). Two exposure surfaces beyond the file itself: (a) the OpenTofu **state file** of any root that uses the maas provider records it -- unavoidable with this provider, so that state inherits credential handling (0600, never committed); (b) a MAAS CLI profile was also created for the operator user from the same key. It was verified by FORMAT ONLY (71 bytes, 3 colon-separated parts) and never printed, echoed, or passed in argv. | 2026-07-20 step-D part 2; `opentofu/vr1-dc0-maas/main.tf` header; session changelog item 15 | operator | **OPEN -- rotation obligation.** Rotate at v1 close, or immediately if `voffice1` is rebuilt/shared or the tfstate leaves the host. NOTE: the pod route this key was placed for was subsequently REFUTED (see changelog item 16) and `opentofu/vr1-dc0-maas` is currently UNUSED -- if that root is retired at stage close, this key's on-disk copy and the tfstate should be removed with it, leaving only the CLI profile. Custody detail off-repo per D-069. | +| SEC-014 | 2026-07-21 | **Rack cluster secret exposed to session context.** During the commissioning diagnosis, a read of `/var/snap/maas//agent.yaml` on the DC0 rack (hunting the agent resolver's config surface) returned the rack's MAAS cluster `secret` into the Claude session context, the operator terminal scrollback, and the as-executed log (`~/as-executed/2026-07-21-ops-commissioning-diag.log`, 0600, jumphost-only). The secret authenticates rack<->region enrollment. The read was not anticipated to contain a credential (config file, not a key file); disclosed same-session. | 2026-07-21 session changelog item 6; docs/audit/commissioning-diag-20260721.txt | operator | **OPEN -- rotation obligation.** Rotate the MAAS shared secret (= rack re-enrollment for `7chphy`) at a convenient maintenance point, or immediately if session artifacts leave the jumphost. Process fix queued: add `agent.yaml` to the guard hook's never-read list alongside key/cred globs. | | SEC-011 | 2026-07-16 | **Node least-connectivity gap (not an L2 breach).** Under D-121 Option C role separation, all nodes get a uniform 6-plane NIC set, so a ceph-osd STORAGE node has a leg on provider-public (external/FIP) + data-tenant (tenant geneve) -- planes it never binds per D-052. Planes stay isolated L2 (no crosstalk). | 2026-07-16 plane-segregation review; `opentofu/main.tf` `local.vr1_dc0_node_nics` | operator | **CLOSED 2026-07-16 (operator ruling -- keep uniform 6-NIC).** Review R3-F10: A2's cross-examination refuted the attack-surface concern -- in the isolated-L2 sim the unbound vNICs have no reachability out, and pruning would INCREASE Roosevelt-delta (baremetal trunks all VLANs to every node on bonded NICs, so all planes are present regardless of L3 binding). Uniform 6-NIC is the more Roosevelt-faithful model. Accepted non-issue; no code change. | **STANDING CONVENTION (SEC-009, 2026-07-15): per-site credential/env consolidation.** ALL sensitive diff --git a/docs/session-ledger.md b/docs/session-ledger.md index b94adb4..ae0206f 100644 --- a/docs/session-ledger.md +++ b/docs/session-ledger.md @@ -30,24 +30,17 @@ ## Machine-derived (re-seed from `scripts/ledger-scan.sh`; do not hand-edit) -_Re-seeded from the 2026-07-17 scan. Re-run `bash scripts/ledger-scan.sh` to refresh._ +_Re-seeded from the 2026-07-21 scan. Re-run `bash scripts/ledger-scan.sh` to refresh._ -- **PROPOSED / OPEN decisions:** D-068 (Vault substrate hardening, Roosevelt -- 1.16 ruled OUT - per amendment; the off-EOL-1.8.8 path remains OPEN), D-071 (routine update cadence + Juju - controller patch policy -- AMENDED, still PROPOSED, awaiting operator ratification; gated on - the pre-DC-DC controller HA/backup session), D-129 (OPNsense edge plugin/add-on base profile -- - OPEN / partially ruled 2026-07-18: VR1 profile adopted, FOUR sub-decisions still OPEN; review in - `docs/opnsense-edge-addon-review-20260718.md`). -- **OPEN security rows:** 8 open per `bash scripts/ledger-scan.sh` (re-seeded 2026-07-19). The +- **PROPOSED / OPEN decisions:** D-068 (Vault substrate hardening, Roosevelt), D-071 (routine + update cadence + Juju controller patch policy), D-129 (OPNsense edge plugin/add-on base + profile -- OPEN / partially ruled, four sub-decisions), D-131 (node-facing DNS strategy for + rack-only controllers -- PROPOSED 2026-07-21; interim forwarder workaround operator-ruled and + live). Status lines in `docs/design-decisions.md` are the only ruling authority. +- **OPEN security rows:** 10 open per `bash scripts/ledger-scan.sh` (re-seeded 2026-07-21). The SEC register of record is `docs/security-ledger.md`; row-level dispositions live THERE only (GA-R4 amendment F3) -- this block carries pointer + count, never rows. -- **Next-free numbers:** D = **130**, DOCFIX = 196, BUNDLEFIX = 013. - (Since the last re-seed D-121..128 all ADOPTED -- D-128 = VR1 operating model (2026-07-17). Load-bearing correction to the prior block: - **D-123 = Model B** (nodes nest INSIDE `vvr1-dc0`; site-down = one `virsh destroy vvr1-dc0`) via - operator ruling R-1 2026-07-16 -- this SUPERSEDES the earlier "D-123 Model A" this block used to - record. Also: D-124 transit addressing PINNED (172.31.0.0/24) + imported to office1-netbox; D-125 - bridge-in single-NAT uplink; D-126 durable rootless site access; D-127 IaC autostart matrix. The - Stage-4 rack<->region addressing the prior block flagged as HELD is now assigned.) +- **Next-free numbers:** D = **132**, DOCFIX = 197, BUNDLEFIX = 052. - **Standing numbering rule:** never write an identifier-shaped token (D-/DOCFIX-/BUNDLEFIX-NNN) ABOVE the real high-water mark anywhere in `docs/` or `runbooks/` prose -- historically a decoy token in prose inflated the next-free counter (hardened in DOCFIX-174); the authoring discipline @@ -130,3 +123,20 @@ measure how reach works before wiring ssh; D-124 tfvars local+present; dc0 key minted). Fresh session: CURRENT-STATE first, then resume at runbook step B in the dc0-deploy logged window (same-day log appends). + +## SESSION CLOSE 2026-07-21 -- commissioning diagnosis + repair (logged: ops-commissioning-diag; bounded, GA-R4) + +- Session DISCONNECTED before its bookend; entry landed same-day by the successor session after + read-only re-verification (9/9 Ready, forwarder active+enabled, dhcpd both controllers). +- G10 commissioning RESOLVED: ALL 9 DC0 nodes Ready, shapes exact to D-121 Option C. +- Two stacked faults, both measured: (1) the 2026-07-20 in-place apply regenerated all 9 node NIC + MACs -- MAAS records repaired in place (operator-ruled), read-back 9/9; (2) MAAS 3.7 rack-only + agent resolver SERVFAILs everything on an internet-isolated rack -- dc0-node-dns dnsmasq + forwarder workaround (operator-ruled), live and proven (canary Ready ~3 min). LP report queued. +- D-131 PROPOSED (rack-only node DNS, Roosevelt-relevant); SEC-014 OPENED (rack cluster secret + exposure); committee record closed by addendum. Phase-7 tag cleanup verified done. +- QUEUED delivery: node-vm MAC pinning + harness; forwarder + rack-legs repo persistence; + platform-traps + appendix-A entries; stale-pod cleanup at stage close (with SEC-013). +- REMAINING in G10: netem step E only (netem-link passwordless-sudo blocker). +- Details: docs/changelog-20260721-commissioning-diag.md + docs/audit/commissioning-diag-20260721.txt. + Status lives ONLY in CURRENT-STATE.md. diff --git a/logs/as-executed-index.md b/logs/as-executed-index.md index e1d1812..ef011f3 100644 --- a/logs/as-executed-index.md +++ b/logs/as-executed-index.md @@ -13,3 +13,4 @@ | 2026-07-07 | ops-devteam-onboard | jesse.austin (Claude Code lane) | devteam onboarding window: stages 0-4 + tenant-assert + canary on 10.100.0.0/24 (D-074 fallback, first client-facing run); capi-test-1 delete ruled + staged (operator to run) | | 2026-07-08 | ops-decommission | jesse.austin (Claude Code lane, per-command wrap) | decommission window: lbtest LB reclaim + beta-cluster teardown + foil1 offboard (first offboard v2 --apply; E0 always-403 defect logged) + magnum orphan-sweep audit (clean); open/close cloud-assert PASS | | 2026-07-08 | ops-devteam-template | jesse.austin (Claude Code lane, per-command wrap) | deliver devteam-k8s cluster template (stage5, calico, hidden=False) after stuck flannel TestCluster2 confirmed gone + RAM reclaimed; closes the addendum-39 no-delivered-template gap | +| 2026-07-21 | ops-commissioning-diag | jesse.austin (Claude Code lane, per-command wrap) | G10 commissioning diagnosis: instrumented run per committee; MAC-drift fault (all 9) repaired in place (ruled); MAAS 3.7 rack-only agent resolver defect found, dc0-node-dns forwarder workaround (ruled); ALL 9 NODES READY; SEC-014 opened; D-131 proposed |