diff --git a/docs/archive/session-ledger-rotated-20260806.md b/docs/archive/session-ledger-rotated-20260806.md new file mode 100644 index 0000000..6d09026 --- /dev/null +++ b/docs/archive/session-ledger-rotated-20260806.md @@ -0,0 +1,34 @@ +# Session-ledger rotated summaries -- archived 2026-08-06 (GA-R4 rule 3, 300-line cap) + +Rotated out of docs/session-ledger.md at the 2026-08-06 (part 2) close to keep it under 300 lines. +These are CLOSED-session narratives (history). Status lives ONLY in docs/CURRENT-STATE.md. + +## SESSION CLOSE 2026-08-02 -- dc0 edge destroyed and rebuilt; D-139 steps 1-3 done; runbook fold opened (bounded, GA-R4) + +- Branch `dc-dc-stage5-preconditions`, **24 commits** pushed (`f79c9e8..`). NO stage opened or closed. Scan: 3 open decisions, SEC **28** (SEC-031, -032 opened), D **141** / DOCFIX 207 / BUNDLEFIX 053. +- **D-139 STEPS 1-3 EXECUTED for dc0.** Apex 139 -> 152 prefixes; MAAS 6 GUA + 5 ULA each paired on one vlan; node statics migrated **GUA 54 / ULA 0 with ZERO multi-global NICs**. v4 untouched, which is the ordering step 3 exists to enforce. Step 3.5 done: model created, spaces gate PASS 0 fatal, `apt-mirror` verified. +- **>>> THE dc0 EDGE WAS DESTROYED AND HAS BEEN REBUILT. <<<** Root cause is NOT the update I first claimed -- zero pkg/firmware lines in the whole serial log. It was **UFS soft-update damage from an unclean power cut**: the 08-01 in-place tofu resize BOUNCED the containment VM, hard-cutting every inner guest. fsck salvaged 2533 unreferenced files and `libcrypto`/`libpython` did not survive. +- **dc1's edge took the SAME cut** (76/181 vs dc0's 2533/785) and lost its user DB instead: it **forwards without translating** (tcpdump, both taps, source unchanged) and runs with NO pf ruleset -- an open router serving its GUI, **SEC-031**. Config INTACT; verdict REPAIR not rebuild, blocked on having no credential path. +- **Edge rebuilt by agent**, `dc-egress-check dc0` **PASS 8/8 exit 0**. Plan asserted on `tofu show -json` including the POSITIVE half; `pfctl -s nat` verified rather than assumed. **SEC-032** minted. +- **NEW GATE `dc-egress-check.sh`** (F9): layered route -> edge answers -> traffic leaves -> upstreams, first failure reported as the cause. Proven live on TWO different failure modes. Wired into restart Stage 0 and phase-4 Step 3.9. **Two of its own defects found and fixed the same day.** +- **RUNBOOK FOLD OPENED** (`docs/runbook-fold-register.md`, 12 rows). D-138 and D-139 appear in **no runbook**; the chain as written would rebuild the pre-D-132/D-138/D-139 shape. Both Class-A rows closed -- incl. `SKILL.md`, which every session reads BEFORE any runbook. +- **6 RULINGS (GA-R5, all utterances quoted):** D-139 ordering (carve before deploy); OOB dual-stack; OOB v4 `10.12.40.0/22`/`10.12.88.0/22` superseding `10.12.60.0/22`; VPN deferred to Roosevelt; **D-135 amended** (dc0 converges on the proxy at rebuild); **D-140 PINNED** (tofu manages juju AFTER a hardened, tested dc0 deploy). +- **OWNED:** I diagnosed the edge break as a partial update from the symptom's SHAPE and was wrong; my agent brief carried a **wrong base-image path** where the apply destroys the volume first and no rollback exists; I guessed `/srv/mirror/ubuntu` and a systemd unit name the repo already defines; two harness cases I wrote never ran while the suite said ALL PASS; one assertion passed on its own comment; and I pushed a red lint once by masking the exit code. +- Gauntlet **ALL GREEN (97)**, repo-lint 0 fail / 1 warn, ledger-scan reconciled. **voffice1 1 commit behind** at close (not a loss). +- **NEXT:** re-stage the rack's VIP overlay (sweep F1 -- it is STALE and is the deploy input), settle the mirror's exit-1-with-"All done" (F2), `pg_dump maasdb` (F6), then fold F2-F11 and the dc1 Phase-2 exercise. Sweep: `docs/audit/queued-findings-20260802-stage5-edge-fold.txt` (**6 FIRST SURFACE**). Status ONLY in CURRENT-STATE.md. + +## SESSION CLOSE 2026-08-02 (part 2) -- queued backlog cleared; mirror ROOT-CAUSED; D-139 step 6 EXECUTED (bounded, GA-R4) + +- Branch `dc-dc-stage5-preconditions`, **9 commits** pushed (`1cdd607..56b37f8`). NO stage opened or closed. Scan: 3 open decisions, SEC **28** (none opened this session), D 141 / **DOCFIX 208** / BUNDLEFIX 053 -- DOCFIX moved 207->208, reconciling with the one number assigned. +- **Sweep F1 and F6 CLOSED; F2 diagnosed then ROOT-CAUSED; F3/F4/F5 graduated to platform-traps + script-authoring; DOCFIX-207** corrected preflight P6's plan count (50/97, stale since 2026-07-10, against a measured 56/108). +- **F6 PASS -- the dc0 region DB is proven uncorrupted:** pg_dump read every page of `maasdb` (23,878,796 bytes / 37,199 lines / exit 0 / completion marker). **F6's own stated blocker was WRONG** -- the discriminators are ROLE and TRANSPORT, not snap confinement; over the unix socket the `maas` role needs no credential at all. +- **F2 ROOT CAUSE IS UPSTREAM:** one of NINE `archive.ubuntu.com` backends (`91.189.92.23`) hangs on ONE dep11 object while serving its directory siblings in 0.5s; the resolver rotates, and 11 of 12 fetches succeed. **"Not transient" WITHDRAWN.** apt is unaffected -- it fetches the `.xz`, which is present; `apt-get update` against the mirror returns rc=0. +- **4 rulings, exact utterances:** *"Re-trigger the sync first, decide after"*; *"Root-cause the curl/debmirror anomaly first"*; **"Full step 6 first, then deploy"**; **"Deprecate both, delete nothing"**. Both step-6 rulings were pushed BEFORE the dependent work (`934a1f0`, `53aae78`) and correctly classed OPS, not new D-numbers. +- **>>> D-139 STEP 6 EXECUTED -- the deploy's last stated blocker. <<<** Apex: 26 GUA VIP addresses created, 26 ULA addresses + 9 ULA prefixes deprecated, nothing deleted; idempotent on re-run. The 26 CREATE targets diff EXACTLY against the deploy overlay's 26 GUA VIP legs. +- **MAAS half: 4 of 5 ULA subnets deleted, 1 HELD.** `fd50:840e:74e2:220::/64` carries the juju controller (`::5`) and the MAAS region VM (`::6`), neither with a GUA counterpart -- deleting it would strip the deploy client's only recorded v6. +- **Two tools shipped:** `netbox/d139-step6-vip-rehome.py` (harness 20 cases) and `dc-plane-ipam.sh retire-v6-ula` (harness 25->32). An adversarial review returned **FIX FIRST on four defects, two CRITICAL** (a dc1 orphan-create; a dropped apex-identity guard) -- all fixed and verified live. +- **OWNED -- THREE of my checkers COULD NOT FAIL**, every one written AFTER I landed that exact rule into script-authoring this session: an assertion satisfied by a traceback; a grep covering one file while the tool inherited the other; and a `sid="'$id'"` comparison that returned a clean ZERO, on which four deletes proceeded. **None was caught by re-reading my own work** -- two by an adversarial reviewer, one by the live run. +- **Also owned:** called the four deletes "proven safe twice" when half that proof was inert (the OUTCOME was safe -- measured afterwards, nodes read v4=6 v6=6); wrote `status=active` into a ruling by inference (measured: `reserved`); and inflated the DOCFIX counter with a decoy token TWICE, the second time inside the sentence correcting the first. +- **Durability:** vcloud 0 uncommitted / 0 unpushed; **voffice1 synced** (was 1 behind); dc0 rack `~/repo-stage` all 13 tracked files MATCH the repo. Gates: gauntlet **ALL GREEN (98)**, repo-lint 0 fail / 1 legacy warn. +- **NEXT:** the preflight-P2 / phase4 machines-overlay asymmetry -- P2 validates a merged input the deploy never passes -- then the bundle deploy. The held subnet needs the controller's v6 re-homed to GUA first and is NOT deploy-blocking. +- Sweep: `docs/audit/queued-findings-20260802-step6-queued-items.txt` (**6 FIRST SURFACE**, incl. a broad `Bash(ssh vr1-dc0-maas *)` allow rule, and four destructive MAAS deletes that matched NO ask rule -- the rule-fails-to-MATCH class, now recurring). Body: `docs/changelog-20260802-queued-items.md`. Status ONLY in CURRENT-STATE.md. diff --git a/docs/audit/queued-findings-20260806-postwave-retire-memcached.txt b/docs/audit/queued-findings-20260806-postwave-retire-memcached.txt new file mode 100644 index 0000000..bda1a26 --- /dev/null +++ b/docs/audit/queued-findings-20260806-postwave-retire-memcached.txt @@ -0,0 +1,106 @@ +QUEUED FINDINGS -- session 2026-08-06 "Stage-5 dc0: F4/F8/F9 + F-A + retire dc-ha-scaleup + memcached 1->3 live" +Sweep per savegame Step 3. FIRST SURFACE items lead. Body: docs/changelog-20260806-stage5-dc0-f4-postwave.md +Commits (UNPUSHED, push held by operator): 7555479, 90f15d7, 5c2f335, ef47213, 667252a. Status ONLY in docs/CURRENT-STATE.md. + +================================================================================ +FIRST SURFACE -- originated in this session; recorded by this session's commits/captures +================================================================================ + +F-1. **[OPERATOR-FLAGGED] designate coordination points at ONE memcached unit, not 3.** After + the memcached 1->3 live scale-up, nova-cloud-controller picked up all 3 servers + (memcache_servers = .115,.166,.165:11211) but designate/leader designate.conf still has + backend_url = memcached://10.12.12.115:11211 (memcached/0 only). designate is workload-BLOCKED + (nameservers/Stage-7), so its config is pre-activation and may re-render on unblock -- OR the + designate tooz-memcached driver uses a single URL by design. RE-CHECK AT STAGE-7 designate + activation whether coordination should span all 3 memcached units (a single-URL coordination + backend is a single point of failure for the 3 designate units). RECORDED, not fixed (hard + rule 1). Surface: docs/audit/stage5-dc0-memcached-scaleup-20260806.txt + CURRENT-STATE + (memcached note) + changelog Item 7. + +F-2. **vault ha_enabled MEASURED FALSE on all 3 units (F4).** The vault charm renders storage + "mysql" with no ha_enabled and exposes no such config option (only vip + dns-ha-access-record); + HA is charm/VIP model, not vault-native. CONFIRMS D-121 (v-a) as-built; vault-native/Raft HA + stays owned by D-068. Operator note: 3 unsealed actives on one backend is safe only while the + hacluster VIP is the sole ingress. Surface: CURRENT-STATE (F4 block) + changelog Item 1 + + capture stage5-dc0-juju-status-14of14-20260806.txt. + +F-3. **memcached VALUE drift found + resolved.** The retired dc-ha-scaleup.yaml carried memcached + num_units:3 (operator-directed 2026-07-31) which BUNDLEFIX-053 NEVER folded -- bundle.yaml + + live sat at 1. Operator ruled "3 units (restore intent)"; BUNDLEFIX-055 folded it into the + base bundle + DOCFIX-212 corrected D-121; live scaled to 3. Surface: changelog Items 5/6/7 + + D-121 + CURRENT-STATE. This corrects two in-session misstatements (see OWNED). + +F-4. **provider-bundle-check ENUMERATES per-offender (measured).** An all-13 cluster_count 3->1 + rewrite exits 1 with 13 DECORATIVE lines; a single-sub rewrite exits 1 with 1. So the + re-pointed single-sub mutate() fixtures (T33/T34) are equivalent to the old whole-overlay + rewrite for proving the arity gate FIRES -- not a narrowing. Surface: harness invariant-10 + comment (tests/provider-bundle-check/run-tests.sh). + +F-5. **dc0 rack ~/repo-stage/overlays/vr1-dc0-vips.yaml is stale-by-ONE-COMMENT.** The vips overlay + was re-rendered (comment-only) AFTER F9 staged it (3f403408); functionally identical, no + VIP/data change. bundle.yaml WAS re-staged current (42845edb). Re-sync the vips comment at the + next real deploy per D-138 (sha-verify). Surface: changelog Item 6. + +================================================================================ +ON SURFACE -- verified already recorded (where) +================================================================================ + +S-1. 14/14 HA measurement-backed (12 active/idle + 2 known-blocked); F8 ceph-radosgw resolved: + CURRENT-STATE + changelog Items 1/2 + capture stage5-dc0-juju-status-14of14-20260806.txt. +S-2. F-A/BUNDLEFIX-054 bundle.yaml HA-chain header (13 subs; 12 triple + vault metal-pair): + bundle.yaml:23 + changelog Item 4. +S-3. dc-ha-scaleup.yaml RETIRED + archived; R6 SUPERSEDED (GA-R5, "Retire the redundancy and + archive"): design-decisions.md (R6 note) + docs/archive/dc-ha-scaleup-RETIRED-20260806.yaml + + changelog Items 5/6 + DOCFIX-211. Harness re-pointed (T17 dropped; T17b/T32/T33/T34 onto base). +S-4. memcached 3 LIVE (add-unit, 3/3 active/idle) + re-stage (42845edb): capture + CURRENT-STATE + + changelog Item 7. + +================================================================================ +OWED / NOT DONE (for the next session) +================================================================================ + +O-1. F-1 designate coordination Stage-7 re-check (above). +O-2. 5 UNPUSHED commits (push held by operator). voffice1 clone will lag until pushed; cannot be + synced (savegame Step 1b) until this host pushes. dc0 rack ~/repo-stage: bundle.yaml current, + vips stale-by-comment (F-5). +O-3. Pre-existing carried items unchanged this session: preflight P5 red (ruled-accepted, 6 + findings), the D-142 vault-init QoL (PROPOSED, impl deferred), F8-era ceph/octavia/designate + blocks are Stage-6/7 activation work. + +================================================================================ +ALWAYS-SWEEP-5 (structurally invisible) +================================================================================ + +A-1. GITIGNORED STATE: no .claude/settings.local.json permission rule was added this session -- + the classifier gated the rack scp (F9) and the juju add-unit; both cleared by IN-BAND operator + approval ("Approved" / "Both approved"), not a persisted rule. So NO gitignored permission + state to preserve. The octavia-pki per-DC overlays remain gitignored (untouched this session). +A-2. DANGLING REFERENCES: all paths cited by this session's commits resolve -- the archived overlay, + both new captures (juju-status-14of14, memcached-scaleup), the changelog. Verified. +A-3. RULING FIDELITY (exact utterances recorded): "Retire the redundancy and archive" (R6 + supersession, design-decisions); "3 units (restore intent)" (memcached, changelog Item 7 + + D-121); "Both approved" (live add-unit + re-stage, changelog Item 7). All dated 2026-08-06. +A-4. AS-EXECUTED LOG GAP: run-logged.sh was NOT used (background session). The live mutations + (juju add-unit memcached, the two rack scp re-stages) were run via direct `ssh vr1-dc0-rack` + and are recorded in the changelog + captures rather than an as-executed script(1) log. DECLARED + here as the log gap; the captures are the durable record. +A-5. CONTRADICTION DETECTOR: (a) bundle.yaml memcached=1 contradicted the claimed "=3" -> measured, + folded to 3 (BUNDLEFIX-055). (b) D-121:55 "memcached=1 (optional 3)" contradicted the + 2026-07-31 operator direction -> DOCFIX-212. (c) vault "active:true" x3 vs HA Enabled=false -> + resolved (charm VIP-HA, not vault-native). All reconciled, none left open. + +================================================================================ +OWNED (own-mistakes, corrected in-session) +================================================================================ + +W-1. Told the operator "memcached=3 is in bundle.yaml:1044" -- WRONG, it was num_units:1. Caught + when reading the block directly for the currency fix. Corrected (changelog Items 5/6/7). +W-2. The retirement commit/changelog first claimed the overlay was "wholly redundant / all-keys + no-op deep-merge." A full merge-diff (run only when the memcached question forced it) showed + it diverged on memcached (3 vs 1). Corrected in the changelog; BUNDLEFIX-055 makes it true. + LESSON: run the merge-diff BEFORE asserting redundancy, not after a downstream question forces it. +W-3. (earlier this session) First F4 draft mis-blamed the mysql backend as HA-incapable; the advisor + caught it -- the real cause is the charm renders no ha_enabled option. Corrected before commit. +W-4. Nearly hand-edited the RENDERED vr1-dc0-vips.yaml comment before a render/ grep showed it is + generated from render/values (D-136); edited the source + re-rendered instead (the W-1 render + trap from the 2026-08-05 session, avoided this time by grepping render/ first). diff --git a/docs/session-ledger.md b/docs/session-ledger.md index 8e5a9fe..a10c8f9 100644 --- a/docs/session-ledger.md +++ b/docs/session-ledger.md @@ -207,36 +207,6 @@ `docs/archive/session-ledger-rotated-20260802.md`. The live ledger stood at 283 lines and this close's summary would have breached the 300-line cap. -## SESSION CLOSE 2026-08-02 -- dc0 edge destroyed and rebuilt; D-139 steps 1-3 done; runbook fold opened (bounded, GA-R4) - -- Branch `dc-dc-stage5-preconditions`, **24 commits** pushed (`f79c9e8..`). NO stage opened or closed. Scan: 3 open decisions, SEC **28** (SEC-031, -032 opened), D **141** / DOCFIX 207 / BUNDLEFIX 053. -- **D-139 STEPS 1-3 EXECUTED for dc0.** Apex 139 -> 152 prefixes; MAAS 6 GUA + 5 ULA each paired on one vlan; node statics migrated **GUA 54 / ULA 0 with ZERO multi-global NICs**. v4 untouched, which is the ordering step 3 exists to enforce. Step 3.5 done: model created, spaces gate PASS 0 fatal, `apt-mirror` verified. -- **>>> THE dc0 EDGE WAS DESTROYED AND HAS BEEN REBUILT. <<<** Root cause is NOT the update I first claimed -- zero pkg/firmware lines in the whole serial log. It was **UFS soft-update damage from an unclean power cut**: the 08-01 in-place tofu resize BOUNCED the containment VM, hard-cutting every inner guest. fsck salvaged 2533 unreferenced files and `libcrypto`/`libpython` did not survive. -- **dc1's edge took the SAME cut** (76/181 vs dc0's 2533/785) and lost its user DB instead: it **forwards without translating** (tcpdump, both taps, source unchanged) and runs with NO pf ruleset -- an open router serving its GUI, **SEC-031**. Config INTACT; verdict REPAIR not rebuild, blocked on having no credential path. -- **Edge rebuilt by agent**, `dc-egress-check dc0` **PASS 8/8 exit 0**. Plan asserted on `tofu show -json` including the POSITIVE half; `pfctl -s nat` verified rather than assumed. **SEC-032** minted. -- **NEW GATE `dc-egress-check.sh`** (F9): layered route -> edge answers -> traffic leaves -> upstreams, first failure reported as the cause. Proven live on TWO different failure modes. Wired into restart Stage 0 and phase-4 Step 3.9. **Two of its own defects found and fixed the same day.** -- **RUNBOOK FOLD OPENED** (`docs/runbook-fold-register.md`, 12 rows). D-138 and D-139 appear in **no runbook**; the chain as written would rebuild the pre-D-132/D-138/D-139 shape. Both Class-A rows closed -- incl. `SKILL.md`, which every session reads BEFORE any runbook. -- **6 RULINGS (GA-R5, all utterances quoted):** D-139 ordering (carve before deploy); OOB dual-stack; OOB v4 `10.12.40.0/22`/`10.12.88.0/22` superseding `10.12.60.0/22`; VPN deferred to Roosevelt; **D-135 amended** (dc0 converges on the proxy at rebuild); **D-140 PINNED** (tofu manages juju AFTER a hardened, tested dc0 deploy). -- **OWNED:** I diagnosed the edge break as a partial update from the symptom's SHAPE and was wrong; my agent brief carried a **wrong base-image path** where the apply destroys the volume first and no rollback exists; I guessed `/srv/mirror/ubuntu` and a systemd unit name the repo already defines; two harness cases I wrote never ran while the suite said ALL PASS; one assertion passed on its own comment; and I pushed a red lint once by masking the exit code. -- Gauntlet **ALL GREEN (97)**, repo-lint 0 fail / 1 warn, ledger-scan reconciled. **voffice1 1 commit behind** at close (not a loss). -- **NEXT:** re-stage the rack's VIP overlay (sweep F1 -- it is STALE and is the deploy input), settle the mirror's exit-1-with-"All done" (F2), `pg_dump maasdb` (F6), then fold F2-F11 and the dc1 Phase-2 exercise. Sweep: `docs/audit/queued-findings-20260802-stage5-edge-fold.txt` (**6 FIRST SURFACE**). Status ONLY in CURRENT-STATE.md. - -## SESSION CLOSE 2026-08-02 (part 2) -- queued backlog cleared; mirror ROOT-CAUSED; D-139 step 6 EXECUTED (bounded, GA-R4) - -- Branch `dc-dc-stage5-preconditions`, **9 commits** pushed (`1cdd607..56b37f8`). NO stage opened or closed. Scan: 3 open decisions, SEC **28** (none opened this session), D 141 / **DOCFIX 208** / BUNDLEFIX 053 -- DOCFIX moved 207->208, reconciling with the one number assigned. -- **Sweep F1 and F6 CLOSED; F2 diagnosed then ROOT-CAUSED; F3/F4/F5 graduated to platform-traps + script-authoring; DOCFIX-207** corrected preflight P6's plan count (50/97, stale since 2026-07-10, against a measured 56/108). -- **F6 PASS -- the dc0 region DB is proven uncorrupted:** pg_dump read every page of `maasdb` (23,878,796 bytes / 37,199 lines / exit 0 / completion marker). **F6's own stated blocker was WRONG** -- the discriminators are ROLE and TRANSPORT, not snap confinement; over the unix socket the `maas` role needs no credential at all. -- **F2 ROOT CAUSE IS UPSTREAM:** one of NINE `archive.ubuntu.com` backends (`91.189.92.23`) hangs on ONE dep11 object while serving its directory siblings in 0.5s; the resolver rotates, and 11 of 12 fetches succeed. **"Not transient" WITHDRAWN.** apt is unaffected -- it fetches the `.xz`, which is present; `apt-get update` against the mirror returns rc=0. -- **4 rulings, exact utterances:** *"Re-trigger the sync first, decide after"*; *"Root-cause the curl/debmirror anomaly first"*; **"Full step 6 first, then deploy"**; **"Deprecate both, delete nothing"**. Both step-6 rulings were pushed BEFORE the dependent work (`934a1f0`, `53aae78`) and correctly classed OPS, not new D-numbers. -- **>>> D-139 STEP 6 EXECUTED -- the deploy's last stated blocker. <<<** Apex: 26 GUA VIP addresses created, 26 ULA addresses + 9 ULA prefixes deprecated, nothing deleted; idempotent on re-run. The 26 CREATE targets diff EXACTLY against the deploy overlay's 26 GUA VIP legs. -- **MAAS half: 4 of 5 ULA subnets deleted, 1 HELD.** `fd50:840e:74e2:220::/64` carries the juju controller (`::5`) and the MAAS region VM (`::6`), neither with a GUA counterpart -- deleting it would strip the deploy client's only recorded v6. -- **Two tools shipped:** `netbox/d139-step6-vip-rehome.py` (harness 20 cases) and `dc-plane-ipam.sh retire-v6-ula` (harness 25->32). An adversarial review returned **FIX FIRST on four defects, two CRITICAL** (a dc1 orphan-create; a dropped apex-identity guard) -- all fixed and verified live. -- **OWNED -- THREE of my checkers COULD NOT FAIL**, every one written AFTER I landed that exact rule into script-authoring this session: an assertion satisfied by a traceback; a grep covering one file while the tool inherited the other; and a `sid="'$id'"` comparison that returned a clean ZERO, on which four deletes proceeded. **None was caught by re-reading my own work** -- two by an adversarial reviewer, one by the live run. -- **Also owned:** called the four deletes "proven safe twice" when half that proof was inert (the OUTCOME was safe -- measured afterwards, nodes read v4=6 v6=6); wrote `status=active` into a ruling by inference (measured: `reserved`); and inflated the DOCFIX counter with a decoy token TWICE, the second time inside the sentence correcting the first. -- **Durability:** vcloud 0 uncommitted / 0 unpushed; **voffice1 synced** (was 1 behind); dc0 rack `~/repo-stage` all 13 tracked files MATCH the repo. Gates: gauntlet **ALL GREEN (98)**, repo-lint 0 fail / 1 legacy warn. -- **NEXT:** the preflight-P2 / phase4 machines-overlay asymmetry -- P2 validates a merged input the deploy never passes -- then the bundle deploy. The held subnet needs the controller's v6 re-homed to GUA first and is NOT deploy-blocking. -- Sweep: `docs/audit/queued-findings-20260802-step6-queued-items.txt` (**6 FIRST SURFACE**, incl. a broad `Bash(ssh vr1-dc0-maas *)` allow rule, and four destructive MAAS deletes that matched NO ask rule -- the rule-fails-to-MATCH class, now recurring). Body: `docs/changelog-20260802-queued-items.md`. Status ONLY in CURRENT-STATE.md. - ## SESSION CLOSE 2026-08-03 -- Stage 5 dc0: bundle DEPLOYED, controller rebuilt, vault up; ovn-central cert DEFERRED (bounded, GA-R4) - Branch `dc-dc-stage5-preconditions`, ~23 commits pushed. NO stage opened/closed. Scan: 3 decisions, **SEC 28**, **D 142 / DOCFIX 209 / BUNDLEFIX 053** (D-141 + DOCFIX-208 assigned this session). @@ -297,3 +267,14 @@ - Gates: gauntlet ALL GREEN (99 harnesses); repo-lint 0 fail (1 pre-existing L1 non-ASCII warn). Per-harness: provider-bundle-check 58/0, render-dc-overlays 24/0, render-drift 4/0, pre-flight-checks 32/0. - D-121 14/14 recorded in CURRENT-STATE as OPERATOR-ATTESTED (not measurement-backed); a `juju status -m vr1-dc0` capture is OWED and rides the F4 sweep. - **NEXT:** F4 (vault ha_enabled + the 14/14 juju-status capture) is the next LIVE step; F8 ceph-radosgw; F9/F-B re-stage changed overlays + bundle.yaml to both racks (sha256-verify); Task #1 post-wave review (incl. whether dc-ha-scaleup.yaml is now redundant; F-A bundle.yaml:23 stale "12 charms" -> 13). Sweep: `docs/audit/queued-findings-20260806-task2-vault-metal-only.txt`. Body: `docs/changelog-20260805-task2-vault-metal-only-commit.md`. Status ONLY in CURRENT-STATE.md. + +## SESSION CLOSE 2026-08-06 (part 2) -- F4 measured + dc-ha-scaleup RETIRED (R6 superseded) + memcached 1->3 LIVE (bounded, GA-R4) + +- Branch dc-dc-stage5-preconditions; **5 commits, UNPUSHED (push HELD by operator)**: 7555479 90f15d7 5c2f335 ef47213 667252a. Scan: 4 open decisions, SEC 29, next-free D-143 / DOCFIX-213 / BUNDLEFIX-056 (used DOCFIX-211/212, BUNDLEFIX-054/055). +- **F4:** 14/14 HA now MEASUREMENT-backed (12 active/idle + 2 known-blocked octavia/designate); **vault ha_enabled MEASURED FALSE** x3 -- charm has no ha_enabled option, HA is VIP-model not vault-native, CONFIRMS D-121 (v-a), Raft stays D-068. F8 ceph-radosgw resolved. F-A/BUNDLEFIX-054 bundle HA-chain header (13 subs; 12 triple + vault metal-pair). F9 rack re-stage (operator "Approved"). +- **dc-ha-scaleup.yaml RETIRED + archived; R6 SUPERSEDED** (GA-R5, "Retire the redundancy and archive") -- DOCFIX-211. Harness re-pointed off the retired fixture (T17 dropped; T17b/T32/T33/T34 onto the base HA chain via mutate(); provider-bundle-check enumerates per-offender, MEASURED, so single-sub == all-13); runbook two-phase deploy model retired; vips comment fixed via render SOURCE + re-render. +- **memcached VALUE drift caught + resolved:** overlay carried memcached=3 (2026-07-31 direction) that BUNDLEFIX-053 never folded (bundle+live=1). Operator ruled **"3 units (restore intent)"** -> BUNDLEFIX-055 folds it + DOCFIX-212 (D-121). Operator **"Both approved"** -> scaled LIVE `add-unit -n 2 --to lxd:1,lxd:2` to 3/3 active/idle; nova-cc sees all 3 servers, designate coordination sees 1 (Stage-7 re-check, F-1). Rack bundle.yaml re-staged 42845edb. +- **OWNED:** told operator "memcached=3 in bundle.yaml" -- WRONG (was 1); retirement first claimed "wholly redundant / all-keys no-op" -- overstated, only a merge-diff (run after a downstream question) showed memcached diverged; earlier F4 draft mis-blamed the mysql backend (advisor-caught). All corrected in-record (instrument-currency memory #17). +- Gates: gauntlet **ALL GREEN (99)**, repo-lint 0 fail / 1 legacy warn, ledger-scan reconciled (decisions + SEC unchanged; numbers moved as assigned). Ledger rotated (08-02 x2 -> archive/session-ledger-rotated-20260806.md), 299->under-300. +- **Durability:** vcloud 0 uncommitted / **5 UNPUSHED (operator hold)**; voffice1 LAGS until push (Step 1b pull blocked on push); dc0 rack bundle.yaml current (42845edb), vips STALE-by-comment (re-sync at next deploy). +- **NEXT:** operator PUSH the 5 commits (then sync voffice1); designate coordination Stage-7 re-check (F-1); pre-existing Stage-6/7 activation blocks (octavia/designate/ceph-rbd-mirror). Sweep: `docs/audit/queued-findings-20260806-postwave-retire-memcached.txt` (5 FIRST SURFACE, F-1 leads). Body: `docs/changelog-20260806-stage5-dc0-f4-postwave.md`. Status ONLY in CURRENT-STATE.md.