diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 2333676..df2098c 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -2278,6 +2278,58 @@ session** -- the ledger holds no in-flight section without a close bookend, so GA-R4 rule 7 is a no-op this close. **NO STAGE OPENED OR CLOSED**; Stage 5 remains OPEN and this is a session bookend, not a GA-R6 stage close. + **>>> POST-BOOKEND WORK 2026-08-01: BOTH DC CONTAINMENT VMs RESIZED 416 -> 480 GiB, 128 GiB + SWAP ADDED, AND A SUBSTRATE-DRIFT GATE (P8) BUILT. <<<** The GA-R4 bookend (`4b8ba3c`) was + committed BEFORE this, so **the ledger's close summary and the 08-01 sweep do NOT cover it**; + the ledger line is amended in the same commit rather than left stale. + **OPERATOR RULING, exact utterance: "Option 2 look sgood me. I approve the sequencing, + process as autonomously as possible"** -- +64 GiB to EACH DC with swap added first, over my + recommendation of +48. **The deciding arithmetic, put up before the ruling:** at +64 each, + allocation is 994 GiB against a 1007.4 GiB host = **13.4 GiB residual, while the host's OWN + measured footprint is 18.0 GiB** -- a 4.6 GiB deficit if every guest went fully resident, + with 1 GiB of swap free. Swap is what makes +64 safe, hence the sequencing. + **THE PROBLEM IT FIXES, MEASURED: dc0's rack ran 402 GiB of inner guests in a 409 GiB host + (98.3%), leaving 7 GiB** for the rack OS, mirror, snap proxy and page cache. Each rack now + has ~71 GiB. + **SWAP: operator-run** (`sudo -n` is NOT available on vcloud, so that half could not be + automated). `/swap2.img` 128 GiB, `600 root:root`, live + fstab. **135 GiB total swap.** + **CHANGED THROUGH TOFU, WHICH IS THE ONLY CORRECT PLACE** -- `var.vvr1_dc0_memory_mib` / + `vvr1_dc1_memory_mib`; a `virsh setmaxmem` would have been reverted by the next apply. The + derivation now lives IN the variable comment, so the config carries its own reason. + **PLAN ASSERTED ON CONTENT** (the 2026-07-20 in-place apply silently regenerated 9 node + MACs): per DC **1 in-place, 0 create/destroy/replace, 0 MAC changes**, `memory` the only + changed attribute. **A third resource appeared and was resolved BEFORE applying:** + `module.office1_opnsense ~ id = 2 -> 11` sits under **"has changed outside of OpenTofu"** + (drift OBSERVED) not **"will be updated in-place"** (action PLANNED) -- libvirt's domain id + is a runtime value that changes on every restart. It was NOT touched. + **dc1 FIRST AS CANARY, and it answered the open question: the in-place update BOUNCES the + guest** (domain id 9 -> 12, uptime 0 min) -- the recorded trap, confirmed. Guest sees + **472.2 GiB** (480 minus firmware reserve). + **>>> FINDING THAT WOULD HAVE MADE THE dc0 BOUNCE READ AS CATASTROPHIC: `vr1-dc0-maas-01` + (MAAS region) and `vr1-dc0-juju-01` (juju controller) have `autostart=disable`. <<<** Only + the edge auto-recovers; the rack's own units (`dc0-snap-proxy`, `-net`, `dc0-mirror-net`, + `nginx`) are all `enabled` and returned unaided. Both VMs were started by hand and verified. + **LOGGED, NOT FIXED: a host reboot leaves dc0 with no MAAS region and no juju controller.** + That is a standing exposure, independent of this change. + **RESULT: host used 809 -> 63 GiB, available 197 -> 944 GiB** -- the restarts also released + the stranded RSS (vvr1-dc1 354 -> 12 GiB, its 6.6 GiB swap freed), which is F1 of the 08-01 + sweep discharged as a side effect. Recovery verified END TO END: MAAS region API + **000 -> 502 -> 200** (a real boot progression, not a flat failure), snap proxy listening, + mirror 200, `juju models` showing the controller model *"Last connection: just now"*. + **GATE P8 ADDED TO PREFLIGHT -- substrate drift.** Built on the operator's question about + keeping the tofu config current, which produced a measurement: **`preflight.sh` and + `pre-flight-checks.sh` contained ZERO `tofu plan` checks**, and the office1-opnsense drift + had been sitting unseen. **DESIGN POINT: `-detailed-exitcode` returns 2 for BOTH a pending + change and a harmless observation**, so P8 asserts on CONTENT -- a pending ACTION FAILS, an + observed-only drift WARNS, an unrecognised shape REFUSES. A gate permanently red on benign + drift gets ignored, which is the failure being prevented. It WARNS rather than fails when it + cannot look, because the outer root lives on vcloud (D-128 Plane 1) and "not the substrate + host" is a legitimate state. **Harness 33 -> 38 (T34-T38), 3 mutations each killing a NAMED + test, script restored sha256-identical. PROVEN LIVE, not fixture-green: against the real + outer root `tofu exit=0`, 0 pending, 0 drift -> `[ok] zero diff`.** + **SCOPE NOTE, because it bounds what P8 can ever mean: tofu owns the SUBSTRATE only.** MAAS + carves, the rack services, and the v6 node carve are script-driven by design (Model B / + D-123). "Keep tofu current" never covered those and P8 does not check them. - Project: Omega Cloud, VR1 DC-DC rehearsal -- a two-DC + Office1-headend virtual rehearsal on KVM (vcloud host), rehearsing the future bare-metal Roosevelt deployment (D-100, `docs/design-decisions.md:1946`). diff --git a/docs/changelog-20260731-snap-proxy-apply-ipv6.md b/docs/changelog-20260731-snap-proxy-apply-ipv6.md index 339c255..6674a21 100644 --- a/docs/changelog-20260731-snap-proxy-apply-ipv6.md +++ b/docs/changelog-20260731-snap-proxy-apply-ipv6.md @@ -980,3 +980,106 @@ defective execution list -- note that reverting reinstates a silent-failure ordering); `git rm docs/audit/lp-draft-20260801-ceph-osd-ipv6-static.md`. The node release is reversed by re-deploying, not by git. + +--- + +## Item 21 -- POST-BOOKEND: both DC containment VMs 416 -> 480 GiB; 128 GiB swap; P8 drift gate + +**NOTE ON ORDERING:** the GA-R4 bookend (`4b8ba3c`) was committed BEFORE this work, so the +ledger's close summary and the sweep do NOT describe items 21. Recorded here and in +CURRENT-STATE; the ledger line is amended in the same commit rather than left stale. + +**Why now, operator's reasoning:** *"If there was any time to add more resources is now while +the cost and risk is very low. Once we bring systems online within the host then the risks +begin to expand."* Correct, and the measurement agreed -- dc0's rack was running its inner +guests at **402 GiB in a 409 GiB host, 98.3%**, leaving 7 GiB for the rack OS, mirror, snap +proxy and page cache. + +**SIZING RULED BY OPERATOR: +64 GiB each ("Option 2"), i.e. 416 -> 480 GiB per DC, WITH swap +added first.** I had recommended +48 and was overruled with a reason -- tenant testing is +coming. **The arithmetic I put up was the deciding input:** at +64 each, allocation is 994 GiB +against a 1007.4 GiB host, leaving **13.4 GiB** while the host's OWN measured footprint is +**18.0 GiB** -- a 4.6 GiB deficit if every guest went fully resident, with only 1 GiB of swap +free. Swap is what makes +64 defensible, which is why it was sequenced first. + +**SWAP (operator-run; `sudo -n` is NOT available on vcloud, so this half could not be +automated).** `/swap2.img` 128 GiB, `600 root:root`, `swapon` live, fstab line added. +Verified: **135 GiB total swap** (8 + 128), 0 used. Effective cushion 13.4 + 135 GiB. + +**TOFU IS THE ONLY CORRECT PLACE FOR THIS CHANGE.** `var.vvr1_dc0_memory_mib` / +`vvr1_dc1_memory_mib` (`main.tf:415`/`:542`). A `virsh setmaxmem` would have been reverted by +the next apply. **The derivation now lives IN the variable comment** -- 402-in-409, overhead +32 -> 96 GiB -- so the config explains its own number instead of the reason living only here. + +**PLAN ASSERTED ON CONTENT, not eyeballed, because of the 2026-07-20 precedent** (an in-place +apply silently regenerated 9 node MACs). Per DC: 1 in-place, **0 create/destroy/replace, 0 MAC +changes**, and `memory` the only changed attribute. + +**A THIRD RESOURCE APPEARED IN THE PLAN AND WAS RESOLVED BEFORE APPLYING, NOT AFTER.** +`module.office1_opnsense` showed `~ id = 2 -> 11`. The distinction that settles it is tense: +**"has changed outside of OpenTofu" (drift OBSERVED) versus "will be updated in-place" (action +PLANNED)**. libvirt's domain `id` is a runtime value that changes on every guest restart. +office1-opnsense was NOT touched by either apply. + +**dc1 FIRST AS THE CANARY** (idle; only its edge running) -- the repo's standing pattern. +It answered the open question: **the in-place update BOUNCES the guest.** Domain id 9 -> 12, +qemu restarted, guest uptime 0 min. Exactly the recorded "in-place bounces the guest" trap. +Guest sees **472.2 GiB** (480 minus firmware reserve). + +**>>> A FINDING THAT WOULD HAVE MADE THE dc0 BOUNCE LOOK CATASTROPHIC, CAUGHT BY CHECKING +FIRST: `vr1-dc0-maas-01` (the MAAS region) and `vr1-dc0-juju-01` (the juju controller) have +`autostart=disable`. <<<** Only the edge auto-recovers. Had I applied dc0 without checking, +the rack would have returned with no region and no controller. Both were started by hand and +verified. **The rack's own services (`dc0-snap-proxy`, `dc0-snap-proxy-net`, +`dc0-mirror-net`, `nginx`) are all `enabled` and did return unaided.** +**LOGGED, NOT FIXED (hard rule 1): a host reboot leaves dc0 with no MAAS region and no juju +controller.** That is a standing exposure, not a consequence of this change. + +**RESULT, MEASURED.** Both VMs at 480 GiB; each rack now has ~71 GiB for its OS instead of 7. +**Host: used 809 -> 63 GiB, available 197 -> 944 GiB** -- because restarting both guests also +released the stranded RSS (vvr1-dc1 354 -> 12 GiB, its 6.6 GiB of swap freed). Recovery +verified end to end: MAAS region API **000 -> 502 -> 200** (a genuine boot progression, not a +flat failure), snap proxy listening, mirror 200, and `juju models` reporting the controller +model with *"Last connection: just now"*. + +## Item 22 -- BUILD: preflight gate P8, substrate drift + +**What.** `scripts/preflight.sh` (new P8) + `tests/preflight/run-tests.sh` (a `tofu` fixture +seam + T34-T38). No new harness, so `HARNESS-MANIFEST` is unchanged. + +**Why.** Answering the operator's question -- does tofu need updates logged, or is it dynamic? +-- produced a measurement: **`preflight.sh` and `pre-flight-checks.sh` contained ZERO +`tofu plan` checks**, and a real drift had been sitting unseen until this session's plan +surfaced it. Substrate changes are genuinely RARE, which is exactly why nobody looks. + +**THE DESIGN POINT: `-detailed-exitcode` returns 2 for BOTH a pending change and a harmless +observation.** The exit code alone cannot decide, so P8 asserts on CONTENT -- a pending ACTION +FAILS, an observed-only drift WARNS. **A gate that is permanently red on benign drift gets +ignored, which is the failure being prevented.** + +**It REFUSES rather than fails when it cannot look.** The outer root lives on vcloud (D-128 +Plane 1), so "not the substrate host" is a legitimate state, not a missing gate -- P4/P5's +fail-closed shape would be wrong here. Absent tofu, or absent outer-root state, WARN. + +**Mutation-proven, three mutations, each killing a NAMED test; script restored and +sha256-verified byte-identical:** neutering the pending-action predicate -> T34 red; +making the unrecognised-shape branch pass -> T36 red; making the cannot-look branch silent -> +T35 AND T38 red. Note M1/M2 turned tests red on the MESSAGE while the exit code was still +right -- content-checking earning its place. + +**T38 EXERCISES THE ABSENCE THE SAFE WAY, and the comment says why.** It removes the fixture's +STATE, never the fake `tofu` from `fakebin` -- the real tofu is on the system PATH, so a +"closed" PATH would fall through to it and touch the LIVE substrate. That is this session's +ninth instrument-currency instance applied prospectively rather than after the fact. + +**PROVEN LIVE, NOT FIXTURE-GREEN:** run against the real outer root, `tofu exit=0`, +pending-actions=0, observed-drift=0 -> **P8 verdict `[ok] zero diff`**. The substrate is fully +converged after both applies, and the office1-opnsense drift is gone because the applies +refreshed state. + +**Harness 33 -> 38 PASS, 0 FAIL.** Baseline proven by stashing: the 7 failures my first draft +caused were MINE, not pre-existing (`PASS=33 FAIL=0` with the change stashed). + +**Revert.** `git checkout -- scripts/preflight.sh tests/preflight/run-tests.sh`. +For item 21: set both variables back to `425984` and re-apply (this bounces both guests +again); `sudo swapoff /swap2.img && sudo rm /swap2.img` and drop the fstab line. diff --git a/docs/session-ledger.md b/docs/session-ledger.md index e60612e..496e35f 100644 --- a/docs/session-ledger.md +++ b/docs/session-ledger.md @@ -279,4 +279,5 @@ - **D-139's OWN EXECUTION LIST WAS DEFECTIVE AND IS REPLACED.** `dc-node-v6-carve.py` pivots on IPv4 existing, so run after v4 removal it would carve four fewer planes per node **and exit clean**. - **OWNED:** my BUG-1 fix was wrong (a secondary alias is never the kernel's chosen source); I scoped the v6 experiment wrong (ceph couples storage+replication); I stated an agent's input source wrongly; I scoped a research agent with no repo path, so 355 lines landed in `/tmp` and needed rescuing; and one commit went red on ASCII-only em-dashes. - Gauntlet **ALL GREEN (96)**, repo-lint 0 fail, `d139-gua-carve` 71/71, `dc-node-v6-verify` 55/55. **voffice1's clone is 36 commits BEHIND** -- no loss, but a live hazard on the Plane-2 host. +- **AMENDED AFTER THE BOOKEND (2026-08-01):** both DC containment VMs resized **416 -> 480 GiB** through tofu (operator: *"Option 2 look sgood me"*), **128 GiB swap** added (operator-run), and preflight gained **gate P8, substrate drift** (harness 33 -> 38, proven live at zero diff). Host used **809 -> 63 GiB**. Found: dc0's MAAS region and juju controller have `autostart=disable`. Bodies: changelog items 21-22. - **NEXT:** apex CREATE-only push (tool built, independently reviewed, dry-run byte-identical), then the bundle deploy -- its blockers are cleared. `network-get` on a v6-only bound space is still unmeasured and gates the v4-removal experiment. Sweep: `docs/audit/queued-findings-20260801-stage5-ipv6-d139.txt` (**6 FIRST SURFACE**). Body: `docs/changelog-20260731-snap-proxy-apply-ipv6.md`. Status ONLY in CURRENT-STATE.md. diff --git a/opentofu/variables.tf b/opentofu/variables.tf index 6c92524..d4d2575 100644 --- a/opentofu/variables.tf +++ b/opentofu/variables.tf @@ -141,9 +141,12 @@ } variable "vvr1_dc0_memory_mib" { - description = "vvr1-dc0 RAM in MiB (Model B). Derived: 384 GiB node fleet + 32 GiB overhead = 416 GiB." + # RAISED 416 -> 480 GiB, 2026-08-01. MEASURED: the inner guest allocation is 402 GiB + # against a 409 GiB rack, leaving 7 GiB (98.3%) for the rack OS, mirror, snap proxy + # and page cache. Overhead 32 -> 96 GiB. Operator-approved for tenant testing. + description = "vvr1-dc0 RAM in MiB (Model B). Derived: 384 GiB node fleet + 96 GiB overhead = 480 GiB." type = number - default = 425984 + default = 491520 } variable "vvr1_dc0_disk_bytes" { @@ -176,9 +179,12 @@ } variable "vvr1_dc1_memory_mib" { - description = "vvr1-dc1 RAM in MiB (Model B). Derived: 384 GiB node fleet + 32 GiB overhead = 416 GiB." + # RAISED 416 -> 480 GiB, 2026-08-01, SYMMETRIC with dc0 -- dc1 hits the identical + # 402-in-409 squeeze the moment it deploys, and cross-DC divergence at the same + # position is a defect in this repo (D-134 precedent). Overhead 32 -> 96 GiB. + description = "vvr1-dc1 RAM in MiB (Model B). Derived: 384 GiB node fleet + 96 GiB overhead = 480 GiB." type = number - default = 425984 + default = 491520 } variable "vvr1_dc1_disk_bytes" { diff --git a/scripts/preflight.sh b/scripts/preflight.sh index 76fc95c..31f5c11 100644 --- a/scripts/preflight.sh +++ b/scripts/preflight.sh @@ -266,6 +266,55 @@ fi fi +echo "================ P8: substrate drift (outer tofu root) ================" +# Nothing detected substrate drift before this gate: measured 2026-08-01, preflight +# contained ZERO `tofu plan` checks, and a real drift (office1-opnsense's runtime id) +# had been sitting unseen. The outer root runs ON vcloud (D-128 Plane 1), so this +# REFUSES rather than fails when run anywhere else -- P4/P5's fail-closed shape is +# wrong here, because "not the substrate host" is a legitimate state, not a missing gate. +# +# `-detailed-exitcode` returns 2 for BOTH a pending change and a harmless refresh-only +# observation, so the exit code alone cannot decide. Asserted on CONTENT: a pending +# ACTION ("will be created/destroyed/updated/replaced") FAILS; an observed-only drift +# ("has changed outside of OpenTofu") WARNS. A gate that is permanently red gets ignored. +if ! command -v tofu >/dev/null 2>&1; then + echo " [warn] tofu not on PATH -- substrate drift NOT evaluated here (run on the vcloud" + echo " substrate host). This is 'could not look', never 'nothing there'." + note 2 "P8 substrate drift" warn +elif [ ! -f opentofu/terraform.tfstate ] && [ ! -d opentofu/.terraform ]; then + echo " [warn] no outer-root state here -- substrate drift NOT evaluated (D-128: the" + echo " outer root lives on vcloud). Not a failure; not a pass either." + note 2 "P8 substrate drift" warn +else + P8_OUT="$(tofu -chdir=opentofu plan -detailed-exitcode -no-color -input=false 2>&1)"; P8_RC=$? + P8_ACT="$(printf '%s\n' "$P8_OUT" | grep -cE 'will be (created|destroyed|updated|replaced)')" + P8_OBS="$(printf '%s\n' "$P8_OUT" | grep -cE 'has changed outside of OpenTofu|Objects have changed')" + case "$P8_RC" in + 0) echo " [ok] substrate matches the outer root -- zero diff" ;; + 2) + if [ "$P8_ACT" -gt 0 ]; then + echo " [FAIL] $P8_ACT pending substrate action(s) -- config and reality have DIVERGED." + printf '%s\n' "$P8_OUT" | grep -E 'will be (created|destroyed|updated|replaced)' | sed 's/^/ /' + echo " Deploying over an unapplied substrate change is how an in-place apply" + echo " mid-deploy bounces a guest (2026-07-20: regenerated 9 node MACs)." + note 1 "P8 substrate drift" fail + elif [ "$P8_OBS" -gt 0 ]; then + echo " [warn] refresh-only drift observed (no pending action). Benign attributes such" + echo " as a libvirt runtime id change on every guest restart; reported so it is" + echo " SEEN, not so it blocks." + note 2 "P8 substrate drift" warn + else + echo " [FAIL] tofu reported changes (exit 2) but no action or drift line parsed --" + echo " refusing to interpret an unrecognised plan shape as a pass." + note 1 "P8 substrate drift" fail + fi ;; + *) + echo " [FAIL] tofu plan exited $P8_RC -- the substrate could not be evaluated." + printf '%s\n' "$P8_OUT" | tail -5 | sed 's/^/ /' + note 1 "P8 substrate drift" fail ;; + esac +fi + echo "================ P6: stage-2 reminders (NOT run here) ================" echo " - after 'juju add-model': bash scripts/juju-spaces-check.sh" echo " - with sudo: bash scripts/osd-blank-check.sh" diff --git a/tests/preflight/run-tests.sh b/tests/preflight/run-tests.sh index c422475..74bc703 100644 --- a/tests/preflight/run-tests.sh +++ b/tests/preflight/run-tests.sh @@ -72,6 +72,23 @@ exit 1 FB chmod +x "$d/fakebin/juju" + # P8 substrate-drift seam (2026-08-01). A fixture with no tofu and no outer-root state + # legitimately WARNS ("could not look"), which would turn every all-green case rc=2 -- + # so the seam supplies BOTH. $7 selects the plan shape: zerodiff | pending | observed | + # unparsed | error. Asserting P8 needs a fake `tofu` because the real one would touch + # the live substrate, which a harness must never do. + mkdir -p "$d/opentofu/.terraform" + cat > "$d/fakebin/tofu" <&2; exit 1 ;; +esac +FB + chmod +x "$d/fakebin/tofu" echo "$d" } run() { # run