|
Step D part 1 done (metal-admin DHCP configured); INCIDENT: MAAS DHCP down region-wide (Temporal wedged)
Rack registered + ruled dynamic range 10.12.8.100-.200 + VLAN 5005
dhcp_on/primary_rack verified by read-back. Canary node did NOT enlist:
measured dhcpd off, no process, no generated config on the rack while the
API reports dhcp_on=True -- config and running state disagree. Region
journal shows Temporal wedged ('Not enough hosts to serve the request',
2807 retries), and MAAS 3.7 drives DHCP through Temporal workflows.
Scope is wider than this deploy: voffice1 has NO dhcpd process either, so
Office1's own DHCP has been down since ~the 2026-07-17 reboot, masked by
both service VMs already being Deployed. Remedy proposed, NOT run (gated).
Detection gap queued: cloud-assert trusts MAAS's self-report.
CURRENT-STATE updated in this commit (GA-R1/C1).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KiUu1oqt76tWvV4vEC3NAr
|
|---|
|
|
| docs/CURRENT-STATE.md |
|---|
| docs/changelog-20260719-dc0-deploy-stepB.md |
|---|