# Diagnostic committee -- MAAS commissioning failure (2026-07-20)

Four independent reviewers (paper review, no live-cloud access) over one shared
evidence packet. Lenses: MAAS commissioning internals; nested-virt/boot
performance; network/MTU; adversarial assumption-breaker. Evidence packet:
scratchpad `commissioning-evidence.md`.

## CONSENSUS (all four agree)

1. **The agent's leading hypothesis -- "marginal 30-min timeout over slow
   nested I/O" -- is WRONG.** Refuted 4/4 (confidences: HIGH, ~75-80%, high,
   ~90%). *Slow is not silent.* A slow-but-progressing commission still emits a
   trickle of script-result events, just late; the symptom is ZERO events for
   30 minutes = the node never phoned home = a **hang/stall**, not slow
   progress. A longer clock does not fix a hang.
2. **The failure is in the POST-"Loading ephemeral" phase.** Not slow boot, not
   the CPU model, not nested-virt perf (the boots are genuinely KVM-accelerated
   -- a `type=kvm` domain cannot silently fall back to TCG, and all 9 enlisted),
   not memory, not contention (the isolated node failed at the full ~29.5 min
   with 397 GiB free).
3. **The decisive next step is to OBSERVE, not to change the clock.** Two
   instruments, both prescribed by multiple reviewers: (a) a FULL-WINDOW packet
   capture keyed on a failing node's **lease IP** (NOT port-scoped) across the
   entire 30 min, on BOTH `virbr2` (metal-admin) and `enp1s0` (transit); (b)
   `console=ttyS0,115200` on the MAAS commissioning kernel cmdline so the
   (currently VGA-only) boot becomes legible in the serial log the node-vm
   module already writes.
4. **The agent's proposed test is CONFOUNDED and must not be run as designed:**
   "raise node_timeout AND commission ONE node" changes TWO variables at once
   (timeout + concurrency 9->1). A one-node pass is equally predicted by
   concurrency-relief and by a timing race -- it would "confirm" the timeout
   theory while proving nothing.

## PIVOTAL ZERO-COST UNKNOWNS (flagged by 3 of 4)

- **The MAAS SUBNET MTU for metal-admin (DHCP option 26)** -- a MAAS setting
  INDEPENDENT of the libvirt network's 9000. It decides whether the guest
  ephemeral NIC even comes up jumbo. If 1500 -> the whole MTU branch largely
  collapses and attention shifts to Temporal/script-hang. If 9000 -> the MTU
  race becomes prime. Single read, reshapes the whole recommendation.
- **Same-3 vs rotating-3 on a re-run** -- from event history / a second batch.
  Same 3 pass every time => deterministic per-node cause; a rotating 3 => a race
  or region-side starvation. Splits the hypothesis space for free.

## DISAGREEMENT ON MECHANISM (ranked candidates, pooled)

- **Region-side Temporal starvation / partial wedge** (adversarial #1). MAAS
  3.5+ drives commissioning via Temporal workflows ON THE REGION; Temporal was
  region-wide WEDGED and only `snap restart`-recovered immediately before these
  runs. Explains every hard fact: total silence (workflow never dispatches the
  next activity), isolated-node-still-fails (isolating the NODE cannot relieve a
  REGION bottleneck), 3/6 with no shape correlation (worker/task-queue
  starvation lets a few through), all-9-enlist (enlistment is a lighter/earlier
  workflow). **The agent measured only RACK resources; region Temporal health
  during a run is UNMEASURED** -- and it is the box that was just unwedged.
  Refs: LP #2084788 (3.5.1 stuck-forever-commissioning), MAAS 3.7 release notes
  (Temporal heartbeat-timeout failures).
- **Commissioning-phase-only script hang on nested-virt hardware** (adversarial
  #1 co-lead). Commissioning runs scripts enlistment does NOT (storage
  `smartctl`, `lldpd` waiting for a neighbor an isolated libvirt switch never
  sends, `bmc-config` probing a BMC that doesn't exist on a VM). MAAS logs
  script COMPLETION, not start -- so a hung script produces IDENTICAL silence,
  cleanly resolving the enlistment paradox. Per-attempt-invariant -> survives
  the isolated-node test.
- **Jumbo/squashfs bulk-transfer black hole** (MAAS-internals #1; network LOW;
  adversarial #3). Near-exact upstream match LP #1807252: hang right after
  "Loading ephemeral" downloading the squashfs, NIC MTU flips 1500->9000 ~20s
  after boot mid-transfer, large frames black-hole, fixed by MTU 1500. The
  OBVIOUS version is REFUTED (enlistment pulls the same squashfs over the same
  node<->rack path and all 9 enlisted; and the rack PROXIES metadata, so the
  node never PMTUs across the 1500 transit). Stays alive only via the narrow
  path: guest vNIC MTU is never set in node-vm (main.tf:115-129) while the plane
  is 9000, and the ~20s post-boot MTU flip is a documented RACE that would
  explain 3/6. Network reviewer's strong counter: config is UNIFORMLY jumbo (not
  mismatched) and the squashfs is a single LOCAL hop (node->tap->virbr2->rack
  httpd on the same host), so a static mismatch would fail all 9, not 3/6.
- **SEC-010 FORWARD-drop catching non-proxied return traffic** (adversarial #4,
  low). Only bites if a commissioning step makes the node talk region-DIRECT
  (forwarded, not proxied) with no established/related accept above the drop.
  Check only if the capture shows region-direct attempts.

## PROCESS CORRECTIONS THE COMMITTEE CAUGHT (agent's own errors)

1. Leading hypothesis refuted 4/4.
2. Proposed test confounded (timeout + concurrency changed together).
3. Measured the WRONG BOX -- rack resources, never region Temporal during a run.
4. The earlier tcpdump was doubly mis-scoped: port 5248 AND a 220s window; must
   be lease-IP-keyed, full-window, both interfaces.
5. The "enlistment paradox" has a clean resolution the agent missed: MAAS logs
   script completion not start, so a commissioning-only script hang = identical
   silence.

## RECOMMENDED SEQUENCE (supersedes "raise node_timeout")

1. Zero-cost reads FIRST (read-only): metal-admin MAAS **subnet MTU**; region
   **Temporal** worker health + region resource headroom; current machine
   states.
2. Then the decisive instrument (gated): full-window, lease-IP-keyed capture on
   virbr2 + enp1s0 during one commission, PLUS `console=ttyS0` on the
   commissioning kernel to read the boot. Four mutually-exclusive capture
   outcomes each kill whole branches (squashfs stalls mid-transfer -> MTU;
   squashfs completes then silence -> script/cloud-init hang; metadata reaches
   region but no/slow answer -> Temporal/proxy; scripts fetch then silent ->
   specific script hang).
3. Only AFTER localization: the targeted fix (MTU clamp to 1500 / node_timeout /
   Temporal restart / script exclusion), each reversible.

## POST-COMMITTEE ZERO-COST READS (agent, read-only, 2026-07-20)

- **metal-admin MAAS VLAN MTU = 1500** (not 9000). The libvirt NETWORK is 9000
  but MAAS serves DHCP option 26 = 1500, so the guest ephemeral NIC comes up
  1500 and NEVER emits a jumbo frame. **The MTU/squashfs-black-hole branch is
  EXONERATED by measurement** -- the exact "if 1500, MTU is off the table" fork
  three reviewers named. (10.10.1.0/24 is likewise 1500.)
- **Region resources ample:** voffice1 = 16 vCPU, 31 GiB RAM (23 GiB
  available), load ~0.95 at a quiet moment. Not starved at rest.
- **Temporal:** `"Not enough hosts to serve the request"` = **175,201 since
  04:00** but **0 in the last 15 min**. Temporal was catastrophically wedged
  this morning (the DHCP incident), recovered after `snap restart maas`, and is
  quiet AT REST now. This neither confirms nor clears the "Temporal starvation
  DURING a commission" hypothesis -- the 175k is mostly pre-restart, and no
  commission is running now. Only a capture/log DURING an active commission
  settles it.

## NET EFFECT ON THE HYPOTHESIS RANKING

- OUT (refuted 4/4): marginal 30-min timeout -> do NOT raise node_timeout as the
  first move.
- OUT (measured): jumbo/MTU black hole -> subnet MTU is 1500.
- STILL LIVE, need observation DURING a run:
  1. Region Temporal starvation during commissioning (was wedged this AM;
     healthy at rest; unmeasured during a run).
  2. Commissioning-only script hang on nested-virt hardware (lldpd/smartctl/
     bmc-config probing hardware absent in a VM; MAAS logs script COMPLETION so
     a hang = silence).
- The decisive gated instrument is unchanged: one commission with
  `console=ttyS0` on the kernel + a full-window, lease-IP-keyed capture on
  virbr2 + enp1s0, watched live for Temporal errors on the region.

Web refs: LP #1807252, LP #2084788, LP #1908452, LP #1403955, MAAS 3.7 release
notes, maas.io controller-communication docs.

## OUTCOME ADDENDUM (2026-07-21 -- diagnosis complete; see
## docs/audit/commissioning-diag-20260721.txt for the full chain)

The prescribed instrument (full-window capture + console=ttyS0) was run and
FOUND the cause on its first execution. Scorecard against this record:

- CORRECT: refuting the marginal-timeout hypothesis (nothing was slow);
  "observe, don't change the clock"; the instrument prescription itself;
  the confounded-test objection.
- WRONG (all four reviewers): every ranked mechanism. There was no hang --
  TWO stacked faults: (1) all 9 node NIC MACs regenerated by the
  2026-07-20 in-place serial-console apply, so MAAS treated every
  post-apply boot as an unknown node (the "silence" was an unrecognized
  machine, booting fine); (2) underneath, the MAAS 3.7 rack-only agent
  resolver SERVFAILs all queries on an internet-isolated rack, so
  cloud-init's maas-internal cloud-config-url never resolved (nodes booted
  to a login prompt, never fetched scripts). Temporal was quiet the whole
  instrumented window (136 lines, 0 errors).
- The "same-3 vs rotating-3" pivotal unknown resolved as: the 3 passed
  BEFORE the MAC-drifting apply -- deterministic, but by timeline, not
  per-node hardware.

Both faults measured, ruled, and repaired/worked-around 2026-07-21
(canary Ready in ~3 min after seven consecutive timeouts). This addendum
closes the committee record.
