diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 1f87149..52836be 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -131,7 +131,17 @@ to D-121 Option C. `opentofu/vr1-dc0-maas/` is retained but UNUSED (the pod route is refuted); retire-or-keep is a stage-close question, as is the D-103/D-123 amendment text. Session changelog items 15-17. - REMAINING IN G10: netem (step E) only. + **COMMISSIONING PARTIAL (open, 2026-07-20): 3 nodes Ready, 6 timed out.** + Contention REFUTED by test (a single node alone sat in `Loading ephemeral` + 20+ min with 385 GiB free, load 0.04); rack boot-image sync and DHCP also + ruled out by measurement. Diagnosis is BLOCKED because `modules/node-vm` + defines NO serial console -- the same sealed-box gap already logged for + `modules/cloudinit-vm`, and the fix is the opt-in serial+log pattern + `modules/opnsense-edge` already has. Failed commissioning is re-runnable; + nothing is lost. REMAINING IN G10: this diagnosis, then netem (step E, + NOT started -- target is the dc0<->dc1 mesh = **virbr5 on vcloud**, + measured; `modules/netem-link` assumes PASSWORDLESS SUDO which vcloud's + operator account lacks). Session changelog item 18. - The grounding audit is COMPLETE and EXITED (2026-07-19): Phases 1-6 all closed (charter `148dcef`; rulings `docs/audit/ga-rulings.md`; the Phase-5 sweep ran as six operator-gated batches in one session; exit diff --git a/docs/changelog-20260719-dc0-deploy-stepB.md b/docs/changelog-20260719-dc0-deploy-stepB.md index 9c5f81f..be4ea24 100644 --- a/docs/changelog-20260719-dc0-deploy-stepB.md +++ b/docs/changelog-20260719-dc0-deploy-stepB.md @@ -665,6 +665,48 @@ - **Revert:** `maas admin machine update power_type=` (clears power) per machine; the script itself reverts by commit. +## 18. OPEN: 6 of 9 nodes fail commissioning ("timed out"); NOT contention; blocked +## on node observability. Step E not started. + +- After power was wired, all 9 were commissioned at once: **3 reached Ready, + 6 hit MAAS's 30-minute timeout** (`Marking node failed - Node operation + 'Commissioning' timed out after 30 minutes`). +- FIRST HYPOTHESIS (resource contention -- 9 nodes x depth-4 nested, 384 GiB + of 416 GiB, all pulling ephemeral images through one /30 transit) was + TESTED AND REFUTED: a single failed node re-commissioned ALONE sat in + `Loading ephemeral` for 20+ minutes with the rack at **385 GiB free and + load 0.04**. It is not contention and not memory. +- Also RULED OUT by measurement: rack boot-image sync (`list-boot-images` + -> `status: synced`, 10 images), and DHCP itself (these same nodes + enlisted fine over PXE earlier, which is what proved the depth-4 gate). +- **BLOCKER IS OBSERVABILITY, and it is the same gap twice:** + `modules/node-vm` defines **no serial console** (confirmed in the module + and in the live domain XML -- no serial/console devices), so a node stuck + in `Loading ephemeral` is a sealed box: no console, no agent, no way to + see the boot. This is the identical finding already logged against + `modules/cloudinit-vm` (item 4), now proven to bite a second module and an + actual live diagnosis. **The fix is the same opt-in serial+log pattern + `modules/opnsense-edge` already carries** -- that module's console log is + what made the 2026-07-12 boot bugs legible, and it is exactly what is + missing here. +- NOT DONE, deliberately: adding a console to node-vm means re-applying + domains mid-diagnosis, which is a change to a built surface and belongs to + an operator ruling, not a reflex. +- Current fleet: 3 Ready (commissioned, deployable), 6 Failed/retrying. + Nothing is lost -- failed commissioning is re-runnable at will. +- **Step E (netem) NOT STARTED.** Measured prerequisites for when it is: + the dc0<->dc1 mesh is **virbr5 ON VCLOUD** (note: virbr5 INSIDE the rack + is provider-public -- same number, different host, measured not assumed); + its qdisc is currently default (`noqueue`). D-100 rules the parameters as + a PLACEHOLDER for the rehearsal within a "same-metro dark-fiber, low + single-digit ms" lean. Applying to the dc0<->dc1 leg is both the + semantically correct link for D-100 and the SAFE one -- it carries no + traffic (dc1 does not exist), unlike the transit, which is currently + load-bearing for MAAS/PXE. **`modules/netem-link` cannot run unattended + from here: it shells out to `ssh 'sudo tc ...'` and assumes + PASSWORDLESS SUDO, which vcloud's operator account does not have** -- a + documented module assumption that this host violates. + ## 8. STAGE-CLOSE QUEUE ITEM (operator directive, 2026-07-20): OPNsense 26.7 review - Operator directive (verbatim intent): at the proper time/place, review the