diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 1b8e0ef..767ac33 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -394,6 +394,44 @@ **REMAINING: the operator-gated merge of `dc-dc-stage4-phase3-maas-deploy` to `main` as a MERGE commit (not squash), then branch retirement (local + remote).** Every substantive in-stage item is closed or split to gate row G17. +- **STAGE-5 GROUNDING AUDIT RUN 2026-07-27 (operator-directed, autonomous, read-only).** + Before opening Stage 5 the operator asked for a full reconciliation and grounding from + a fresh session: current status, the changes required to reach the as-built, the + configuration versus the project goal, and readiness to enter the next stage error-free + -- then a review of the UPCOMING stages to fix problems before the deployment finds them. + A 7-lens read-only committee ran, plus a live measurement sweep. **VERDICT: Stage 5 would + NOT run error-free today and would fail early.** The SUBSTRATE is in excellent shape -- + all three OpenTofu roots plan ZERO DIFF, 18 nodes Ready with shapes exact to D-121 Option + C, all 18 pinned MACs and power addresses matching `lib-hosts.sh`, D-134 statics perfect, + 17 fabrics, zero orphaned interfaces, both artifact paths serving, gauntlet ALL GREEN + (81). What is NOT ready is the layer between the substrate and the deploy. Headline + blockers, each measured: **the Office1 headend clone -- the D-128 Plane-2 host Stage 5 + EXECUTES from -- is 105 commits behind `main` on a branch retired four days ago, with + both dc1 overlays ABSENT and `bundle.yaml` still the VR0 4-node hyperconverged layout**; + the `openstack` client is installed on NEITHER host while ten Stage-5/6/7 scripts invoke + it; the D-104-amendment 10th controller VM is UNAUTHORED (no OpenTofu resource anywhere) + and its MAAS tag does not exist, while dc1 has exactly 9 Ready nodes for a 9-machine + bundle; `ceph-osd` targets `/dev/vdb` and every node has only `vda` (found independently + by two lenses using different methods); per-role MAAS tags are consumed by the machines + block and authored nowhere; there is ZERO IPv6 in the DC substrate while D-101 RULED + dual-stack for both DCs this deployment; and Step 4's "follow phase-01 verbatim" points + at a runbook whose VIP guard ABORTS for dc1 (and will abort for dc0 once the ruled VIP + extraction lands). Several gates that should have caught these CANNOT FAIL -- `repo-lint` + returns `PASS (0 fail, 0 warn)` over ZERO files on a one-character typo of its flag or + root; `provider-bundle-check` passes decorative HA because `cluster_count` is checked + nowhere in `scripts/` or `tests/`; preflight's aggregator ignores any sub-gate exit code + that is not 1 or 2; and P3 verified ZERO of 33 charm-channel pins because `juju` is not + on the host's PATH. **DELIVERABLES (this doc stays the status authority; those are + findings and questions, not status):** ordered precondition checklist + `docs/audit/stage5-readiness-20260727.md` (READ FIRST); verbatim committee record + `docs/audit/stage5-committee-raw-20260727.md`; **11 Stage-5-blocking + 4 standing + questions awaiting GA-R5 rulings, one exchange each, in + `docs/audit/queued-rulings-20260727.md` -- NONE are adopted**; measurements + `docs/audit/stage5-live-measurement-20260727.txt`; charter + `docs/audit/stage5-grounding-audit-scope-20260727.md`. The DOCFIX remediation batch + (21 items, Phase 3 of the readiness doc) is LOGGED NOT EXECUTED -- nearly every runbook + fix interlocks with an unanswered ruling, so landing them now would encode assumptions. + The sole mechanical fix taken this session is the G3 row correction above. - Position inside Stage 3: deploy step A EXECUTED 2026-07-19 (6/0/6 exact; convergence zero -- `docs/audit/outer-plan-20260719-postA-converged.txt`). **Deploy step B @@ -845,7 +883,7 @@ |---|------|----------------|-------|---------------------------| | G1 | Audit Phase 3: fresh-agent grounding test | [V] 3 clean-context probes score the 7-question set against this doc; holes map made | session | CLOSED 2026-07-18: 3 probes, 21/21 PASS, holes H1 (amended into G9) + H2 (no action) -- `docs/audit/phase3-grounding-test-20260718.md` | | G2 | Audit Phase 4: GA-R1..R7 structural rulings + the stage-status vocabulary A/B | [R] ruling-type gate (GA-R6 rule 6): closes when every item carries a GA-R5 Status block | operator | CLOSED 2026-07-18: all seven GA-R + vocabulary (Option A + H1) RATIFIED, utterances quoted (`docs/audit/ga-rulings.md`, through commit `fe4f1c4` + this one) | -| G3 | Audit Phase 5: repair sweep of GA-F01..F15 (incl. memory hygiene GA-F05..F08, skill sweep) | [R] operator-gated fix batches, each commit naming its GA-F | operator + session | Batch 0 OPENED by operator 2026-07-19; items 0.1 (repo-lint L10, GA-R1/C1), 0.2 (SEC repoint, GA-R4/F3), 0.3 (counter hardening, GA-F15), 0.4 (extractor vocab scan, GA-F10/H1) landed; Batch 0 CLOSED (verification passed 2026-07-19); Batches 0-4 CLOSED 2026-07-19 (Batch 4: GA-R4 ledger rotation 1187->131 lines, F1 cap now enforceable; 96 changelogs + 24 history docs consolidated to docs/archive/ with 4 stage records + per-stage manifest commits; top-level docs/ = 16 files < 25; live-surface refs rewritten); Batch 5 CLOSED 2026-07-19 (skill sweep: GA section added reconciled against ratified text, stale phase/UNVALIDATED claims demoted, bookends + stage-close rewritten; checklist docs/audit/skill-sweep-checklist-20260719.md); Batch 6 OPEN: exit runs 1/2/4/5 PASS (adjudication + captures: docs/audit/phase6-exit-runs-20260719.md); exit item 3 PASS 2026-07-19 (fresh trio 21/21, 7/7 all three -- exit record); PENDING only item 6 (operator re-read + re-sign, replaces section 11); Batches 2-6 await gates; FREEZE holds for un-gated surfaces | +| G3 | Audit Phase 5: repair sweep of GA-F01..F15 (incl. memory hygiene GA-F05..F08, skill sweep) | [R] operator-gated fix batches, each commit naming its GA-F | operator + session | Batch 0 OPENED by operator 2026-07-19; items 0.1 (repo-lint L10, GA-R1/C1), 0.2 (SEC repoint, GA-R4/F3), 0.3 (counter hardening, GA-F15), 0.4 (extractor vocab scan, GA-F10/H1) landed; Batch 0 CLOSED (verification passed 2026-07-19); Batches 0-4 CLOSED 2026-07-19 (Batch 4: GA-R4 ledger rotation 1187->131 lines, F1 cap now enforceable; 96 changelogs + 24 history docs consolidated to docs/archive/ with 4 stage records + per-stage manifest commits; top-level docs/ = 16 files < 25; live-surface refs rewritten); Batch 5 CLOSED 2026-07-19 (skill sweep: GA section added reconciled against ratified text, stale phase/UNVALIDATED claims demoted, bookends + stage-close rewritten; checklist docs/audit/skill-sweep-checklist-20260719.md); Batch 6 OPEN: exit runs 1/2/4/5 PASS (adjudication + captures: docs/audit/phase6-exit-runs-20260719.md); exit item 3 PASS 2026-07-19 (fresh trio 21/21, 7/7 all three -- exit record); **CLOSED 2026-07-19 -- CORRECTED 2026-07-27 (Stage-5 grounding audit, finding L1-7).** This cell previously read "Batch 6 OPEN ... PENDING only item 6 (operator re-read + re-sign) ... FREEZE holds for un-gated surfaces". That was contradicted by its OWN cited evidence file: `docs/audit/phase6-exit-runs-20260719.md:84` records "## 6. Operator re-read + re-signature -- SIGNED 2026-07-19 ... PASS" and `:91-92` states "ALL SIX EXIT RUNS PASS. The grounding audit is EXITED; **G3 + G11 CLOSED**". Corroborated by row G11 (CLOSED, re-signed 2026-07-19) and by section 11's recorded utterance "Reviewed, approved, continue." No ruling was required to correct this -- two surfaces already declared it closed. **The consequential half was the FREEZE clause, not the state token: left standing it would have blocked the very DOCFIX remediation batch the 2026-07-27 audit queues.** The freeze was lifted at audit exit 2026-07-19 (section 1) and normal change discipline governs | | G4 | The two D-130 verifications | [V] run them, capture output | session | CLOSED 2026-07-19: v8 suppression CONFIRMED (7/2/7 -> 6/2/6, zero forces-replacement; `docs/audit/outer-plan-20260719-v8-ignorechanges.txt` + `-baseline.txt`); v7 no-bounce under running domain, zero residue (`docs/audit/throwaway-v7-20260719.txt`) | | G5 | D-130 mechanism ruling (seed-volume durable fix) | [R] operator rules in Phase 5, quoting G4's captured output | operator | CLOSED 2026-07-19: D-130 ADOPTED (a) ignore_changes (`docs/design-decisions.md` D-130, GA-R5 utterance quoted); implemented in `modules/cloudinit-vm` + `tests/cloudinit-vm` | | G6 | State reconcile of autostart + seed WITHOUT bouncing guests | [R] gated mechanism, operator-ruled (S3) | operator | CLOSED 2026-07-19: ruled (ii) state surgery (GA-R5); pull -> inject autostart:true on both domains -> push (serial 22->23, backup `terraform.tfstate.pre-G6-20260719`); guests never touched (ids 1/2 unchanged, running) | diff --git a/docs/audit/queued-rulings-20260727.md b/docs/audit/queued-rulings-20260727.md new file mode 100644 index 0000000..051c891 --- /dev/null +++ b/docs/audit/queued-rulings-20260727.md @@ -0,0 +1,427 @@ +# QUEUED RULINGS -- Stage-5 grounding audit (2026-07-27) + +**Nothing here is adopted.** Every item is PRESENTED, never picked (hard rule 1; +GA-R5: PROPOSED means present options). Each question is a SEPARATE exchange -- +GA-R5 rule 1 makes a batch adoption INVALID, so answering "yes to all" rules +NOTHING and the session that reads this must stop and re-ask. + +**How to use this file.** Answer questions ONE AT A TIME. Write your exact words +on the `OPERATOR UTTERANCE:` line. A ruling exists only once its Status block +quotes the question AND your exact utterance, dated, committed and pushed -- +before any dependent work starts. An ambiguous or template answer rules nothing. + +Evidence for every claim below is in `docs/audit/stage5-committee-raw-20260727.md` +(verbatim lens returns) and `docs/audit/stage5-live-measurement-20260727.txt` +(this session's own measurements). Finding IDs are cited so nothing here has to +be taken on trust. + +Ordering note: **R1, R2 and R3 change what gets deployed.** They should be +answered before the rest, because several later questions have different right +answers depending on them. + +--- + +# PART A -- STAGE-5 BLOCKING + +These must be answered before `juju deploy`. Each one, left unanswered, either +stops the deploy or bakes in a state that is expensive to reverse on a live cloud. + +--- + +## R1. The Ceph OSD device does not exist on any node + +**Finding:** L2-1 (measured twice -- `virsh domblklist` on both racks AND MAAS +`physicalblockdevice_set` on all 18 nodes). Verified independently by this session. + +`bundle.yaml:560` sets `osd-devices: /dev/vdb`. Every one of the 18 VR1 node VMs +has exactly ONE block device, `vda`. `opentofu/modules/node-vm/main.tf:104-125` +declares a single disk and neither substrate root has any OSD-disk variable. The +`/dev/vdb` line is a VR0 as-built comment ("libvirt-attached, MAAS-untracked") -- +VR0 hosts had an attached second disk; VR1 nodes never did. + +Consequence if unanswered: `ceph-osd` deploys onto four storage nodes and finds +no device. Ceph never forms, and everything storage-backed behind it stalls. + +**Options:** +- **(a)** Add an OSD volume to `modules/node-vm` and re-apply BOTH substrate + roots. Closest to Roosevelt (real machines have real disks), but it re-opens + the substrate on 18 running-but-powered-off nodes and both roots currently plan + ZERO DIFF -- that property is deliberately being spent. +- **(b)** Re-point `osd-devices` at a directory or partition path on the existing + `vda`. Cheapest, no substrate change, but it rehearses a Ceph layout Roosevelt + will not use, which cuts against MINIMIZE DELTA TO ROOSEVELT. +- **(c)** Re-shape the storage role (e.g. fewer, larger storage nodes with a + dedicated disk each). + +OPERATOR UTTERANCE: + +--- + +## R2. There is no IPv6 anywhere in the DC substrate, but dual-stack is RULED + +**Finding:** L2-3 (measured). D-101's RULING NOTE of 2026-07-25 records your exact +words -- "Dual stack to be used where IPv4 is required, IPv6 where IPv6 only makes +sense" and "Dual stack deployment for DC0 and DC1. This is the deployment when the +dual-stack is added" -- with the stated effect that "the v4-only phasing option ... +is CLOSED, for BOTH DC0 and DC1 in this deployment." + +Measured reality: exactly ONE IPv6 subnet exists cloud-wide +(`2602:f3e2:f01:100::/64` on the Office1 base fabric). None of the 12 DC plane +fabrics carries an IPv6 subnet. Zero IPv6 links across all 18 nodes; every node +reports `default_gateways.ipv6 = NONE`. D-101's own family matrix requires ULA on +data-tenant, storage and replication and a ULA leg on metal-admin/metal-internal -- +none of which has any MAAS v6 presence to bind against. D-101's own "Remaining open +item" is the un-assigned NetBox literals: the org ULA /48 and the per-DC GUA carve. + +This is the single largest fork in the audit. It is BLOCKING because addresses +become as-built Keystone endpoints and Vault-issued cert SANs at deploy time. + +**Options:** +- **(a)** Assign the ULA /48 and per-DC GUA carve in NetBox, carve them into MAAS, + and deploy dual-stack as ruled. Honours D-101; adds real work before Stage 5. +- **(b)** Deploy v4-only now and add v6 legs post-deploy. **This CONTRADICTS a + ruling you already made** -- it would need an explicit amendment, not a silent + choice, and re-addressing endpoints and re-issuing SANs afterwards is the + expensive path. +- **(c)** Deploy dc1 v4-only as a deliberate, recorded rehearsal exception while + dc0 goes dual-stack, making the v4/v6 delta itself the experiment (the D-135 + per-DC-difference pattern applied to address family). + +OPERATOR UTTERANCE: + +--- + +## R3. The MTU budget is in neither of D-101's two permitted states + +**Finding:** L2-5 (measured on both racks and across all 17 MAAS VLANs). + +D-101's Tenant/MTU sub-policy allows exactly two shapes: raise the underlay to +jumbo (9000) end-to-end so tenant MTU stays 1500, OR accept 1500 and pin tenant +MTU to about 1444 consistently across ovn geneve, tenant-network MTU and amphora. +It also says: "The measured underlay MTU is a Phase-0 gate -- do not assume jumbo." + +Measured: the six plane bridges on both racks are MTU **9000**; every MAAS VLAN +record (17/17) says **1500**; the rack transit leg `enp1s0` is **1500**, so +cross-DC replication is not jumbo end-to-end; and `grep -i mtu bundle.yaml +overlays/*.yaml` returns NOTHING. That is neither branch. `scripts/dc-dc-mtu-geneve-budget.sh` +exists and correctly refuses to guess (`FAIL: --underlay-mtu is REQUIRED`), but has +never been run to a recorded verdict. + +Consequence if unanswered: the deploy SUCCEEDS and the failure appears later as +tenant/geneve blackholing -- D-101 names this "the classic nested-OpenStack failure +mode". It is a bundle option that must be set BEFORE deploy. + +**Options:** +- **(a)** Raise the rack transit and the MAAS VLAN records to 9000 so the underlay + is genuinely jumbo end-to-end, and keep tenant MTU at 1500. +- **(b)** Accept a 1500 underlay and pin ~1444 consistently across ovn geneve, + tenant-network MTU and amphora, set explicitly in the bundle. + +OPERATOR UTTERANCE: + +--- + +## R4. D-134's reserved address bands exist in prose only + +**Finding:** L2-2 (measured). + +MAAS holds exactly THREE ipranges cloud-wide and ALL THREE are `type=dynamic`. +There are ZERO `type=reserved` ranges anywhere. `maas admin subnet +unreserved-ip-ranges` reports the `.50-.99` VIP band as allocatable on **12 of 12** +DC plane subnets. Only the node statics and the `.201-.254` dynamic ranges are +protected. The tool that creates these reservations, `phase-00-maas-standup.sh`, +REFUSES to run for any non-VR0 DC (`:136`), so no VR1 path to create them exists. + +Consequence: nothing stops MAAS handing a VIP-band address to a Juju/LXD container +during the deploy, and `phase-04-network-verify.sh:100` already hard-fails if the +FIP pool is not a reserved iprange. + +**Options:** +- **(a)** Create the reserved ipranges on all 12 subnets AND carve the FIP pool + now, as one gated mutation batch before deploy. +- **(b)** Create the VIP-band reservations now; defer the FIP pool to phase-04 + where its own verifier expects it. +- **(c)** Accept the bands as unreserved for the rehearsal and rely on static + assignment, recording the risk explicitly. + +OPERATOR UTTERANCE: + +--- + +## R5. Designate ships in the bundle, so Stage 5 will deploy it -- and that breaks Stage 7's gate + +**Finding:** L6-1, corroborated by L3-2. + +`bundle.yaml` carries live `designate`, `designate-bind`, `designate-mysql-router` +and `designate-hacluster` blocks plus 8 relations (DOCFIX-167 closed that on +2026-07-10). So `juju deploy ./bundle.yaml` deploys Designate at Stage 5. But +Stage 5's own runbook says the bundle "explicitly ships NO designate", and Stage 7 +Step 5's gate requires "the diff shows ONLY the new designate/... applications +being added" -- which can never be true if they are already there. + +**Options:** +- **(a)** Suppress Designate for the Stage-5 deploy (a subtractive overlay, or a + rendered per-stage bundle) and keep Stage 7's incremental-add shape intact. +- **(b)** Accept Designate landing at Stage 5 and rewrite Stage 7 Steps 1/5 into a + "configure, not deploy" shape. **Caveat worth weighing:** D-106's own bootstrap + order puts `os-public-hostname` + FQDN-SAN certificates BEFORE Designate; + option (b) inverts that order. + +OPERATOR UTTERANCE: + +--- + +## R6. Whether to apply the HA scale-up overlay at Stage 5 + +**Finding:** L6-4. + +`overlays/dc-ha-scaleup.yaml` sets `ceph-radosgw` and `designate` to `num_units: 3` +with `cluster_count: 3`. But Stage 6's radosgw multisite procedure and the +DOCFIX-165 script behind it are single-unit-shaped: measured, +`dc-dc-radosgw-multisite.sh --help` exposes `master-init ... --unit U`, one unit, +with no all-units mode, and the runbook says `juju run ceph-radosgw/0 restart`. +Realm/period membership would land on one of three gateways and Stage 6's Step-4 +gate could false-green. + +Scaling 1 -> 3 after the fact is a live-cloud change, which is why this is a +Stage-5-time decision. + +**Options:** +- **(a)** Apply the HA overlay at Stage 5 and fix the Stage-6 radosgw path to be + multi-unit-aware first. +- **(b)** Deploy single-unit at Stage 5 and scale up after Stage 6's multisite + work, accepting a live-cloud scale-out later. + +OPERATOR UTTERANCE: + +--- + +## R7. Do the two DCs share one Octavia CA, or get independent trust domains? + +**Finding:** L7-6. The runbook itself flags this correctly and says it must not +silently default to reuse -- but the call has never been made. + +`overlays/octavia-pki.yaml` is a single unscoped path holding CA private keys plus +a plaintext issuing-CA passphrase. Reusing it across both DCs puts one amphora +control-plane CA private key across two clouds that D-100 defines as independent. +Note the contrast one step later in the same runbook: per D-109 each DC's Vault is +its OWN independent root CA, no regional root-of-trust. + +For a commercial multi-tenant cloud with hard tenant isolation, shared-CA is the +weaker posture -- but it is your call, and the rehearsal cost differs. + +**IMPORTANT -- the runbook offers you a choice that is currently impossible on one +side.** Lens 5 (L5-3) measured the generator: `runbooks/phase-01-bundle-deploy.md:369-373` +reads the octavia VIP out of `bundle.yaml` and hard-gates it with +`grep -qE '^10\.12\.4\.[0-9]{1,3}$' || { echo "FAIL: implausible VIP -- stop"; exit 1; }`. +dc1's octavia VIP is `10.12.64.57` and lives in `overlays/vr1-dc1-vips.yaml:37`, not +in `bundle.yaml` at all. Its CN and SAN are hardcoded +`octavia-controller.omega.dc0.vr0.cloud.neumatrix.local` and the CA subject is +`/CN=VR0 DC0 Omega Cloud Octavia Controller CA`. So option (a) requires a generator +fix first, and option (b) bakes a dc0 CN/SAN into dc1's Octavia trust domain. +Also note `runbooks/phase-01-bundle-deploy.md:144-145` hard-ABORTS the deploy if the +overlay is absent -- which it currently is. + +**Options:** +- **(a)** Regenerate fresh per-DC Octavia PKI -- independent trust domains per DC, + consistent with D-109's per-DC Vault root. **Requires fixing the generator's + dc0-frozen VIP gate and CN/SAN literals first** (a DOCFIX, no choice in it). +- **(b)** Reuse the existing CA across both DCs for the rehearsal, recording both + the divergence from D-109's per-DC posture AND the dc0 CN/SAN in dc1's chain. +- **(c)** Fix the generator now and defer the trust-domain decision until it can + actually be executed either way. + +OPERATOR UTTERANCE: + +--- + +## R8. Octavia `lb-mgmt-net` address family + +**Finding:** L6-9. Register item 13 carries the same open fork. + +Stage 5's runbook states plainly that "Octavia's `lb-mgmt-net` IPv6 support is a +real, open risk, not resolved" and that it must be decided before the +Ceph-over-v6 / geneve-over-v6 gate is declared closed. Stage 6's ENTRY condition +requires that gate to have passed, yet Stage 5's own text sanctions recording +"blocked on Step 6" -- so Stage 5 can close without producing Stage 6's +precondition, which GA-R6 E3 forbids resolving by conditional close. + +This is downstream of R2: if R2 goes v4-only, this question largely dissolves. + +**Options:** +- **(a)** Pin `lb-mgmt-net` to IPv4 for this deployment regardless of R2, and + record it as a scoped exception to the dual-stack ruling. +- **(b)** Attempt IPv6 `lb-mgmt-net` and make it a named Stage-5 verification. +- **(c)** Defer until R2 is answered, then re-present. + +OPERATOR UTTERANCE: + +--- + +## R9. Where do dc1's OpenStack-layer network literals live? + +**Finding:** L1-1 (with an explicit guard), extended by L7-10 and L7-3. + +`scripts/lib-net.sh`'s `vr1-dc1` arm `unset`s `VIP_PREFIX_*`, `FIP_POOL_*`, +`VIP_COUNT_EXPECT` and `KEYSTONE_VIP_DEFAULT` on the stated grounds that they are +"NOT yet ruled/measured". They have since been ruled (D-134 amendment, 2026-07-23) +and built (`overlays/vr1-dc1-vips.yaml`). Any Stage-5 script that correctly calls +the selector now dies under `set -u` on a value that exists. + +**GUARD -- do not let this be "fixed" mechanically.** That arm unsets TWO groups +for TWO different reasons. `METAL_INTERNAL_VID` and `METAL_INTERNAL_IFACE` are +**correctly** unset: D-133 abolished the VLAN-103 / `br-internal` stack for VR1 and +those facts genuinely do not exist. Only the VIP/FIP/keystone group is superseded. + +Compounding context (L7-3, measured): of 27 `lib-net.sh` consumers, only 6 call +`lib_net_select_dc` at all. The rest source it unconditionally and silently get +VR0/dc0's literals -- so Stage 5 Steps 7-9 would write dc0's `10.12.4/8/12` values +against a DC whose planes are `10.12.64-84`. Whichever option you pick, that +consumer sweep is the larger half of the work. + +**Options:** +- **(a)** Populate `lib-net.sh`'s dc1 arm from the D-134 amendment -- `lib-net.sh` + stays the single authority for network literals. +- **(b)** Re-point the `phase-0*` scripts at `overlays/vr1-dc1-vips.yaml` as the + authority, leaving `lib-net.sh` for substrate facts only -- closer to where + D-136 would eventually take this. + +OPERATOR UTTERANCE: + +--- + +## R10. On what basis does Stage 5 start against a red preflight? + +**Finding:** L1-9, with this session's measurement. + +Stage 5's stated entry gate is "`preflight.sh` PASS". Measured today, preflight +exits 1, and the red set is exactly the known one: P4's missing +`overlays/octavia-pki.yaml` (a gitignored secret, absent by design), P4's "MAAS +unreachable from the jumphost" (expected -- the region is on voffice1), and P5's 7 +credential findings. Nothing new has joined. But "PASS" is unreachable as written, +so the gate as stated can never authorise Stage 5. + +Note this interacts with L4-2: P3 currently verifies ZERO of 33 charm-channel pins +because `juju` is not installed on the host preflight runs on. + +**Options:** +- **(a)** Re-express the Stage-5 entry gate as a NAMED subset that must be green + (e.g. P1, P2 and P5-with-known-residuals), with the known-red items listed as + accepted preconditions. +- **(b)** Fix the reds first -- place the octavia overlay, run preflight from the + headend where MAAS is reachable, remediate the 7 credential findings. +- **(c)** Record a dated, explicit exception basis for this stage only. + +OPERATOR UTTERANCE: + +--- + +## R11. Vault and Designate both have HA intent and no VIP + +**Finding:** the vault half was already recorded; L3-2 found the SECOND case. + +`grep -n 'vip' bundle.yaml` returns exactly 11 lines -- none for vault, none for +designate. Both are scaled to 3 with an `hacluster` subordinate related and +`cluster_count: 3`. `provider-bundle-check.py:133-135` skips any app with no `vip` +(`if not vip: continue`), which is why neither has ever been flagged. + +Compounding (L4-3, verified by this session): `cluster_count` is checked NOWHERE in +`scripts/` or `tests/` -- a 3 -> 1 rewrite of all 20 occurrences yields a +byte-identical PASS. So neither before nor after the deploy does anything assert +that HA is real (L4-10: `cloud-assert.sh` reports "Cluster ID uniform across units" +over a SINGLE unit). + +**Options:** +- **(a)** Add per-DC VIPs for both vault and designate in the symmetric overlay + shape, and extend the checker to fail on any app with an hacluster relation and + no VIP, plus any `cluster_count` that does not match its principal's `num_units`. +- **(b)** VIPs for both, checker work deferred to a follow-up. +- **(c)** Deploy them without VIPs deliberately (recording why consumers reaching a + unit address is acceptable here). + +OPERATOR UTTERANCE: + +--- + +# PART B -- STANDING / NOT STAGE-5 BLOCKING + +Real, evidenced, and safe to answer after Stage 5 starts. Kept separate so the +eleven above are not diluted. + +--- + +## R12. G17's scope: does it carry the node time-source check? + +**Finding:** L1-8. `docs/CURRENT-STATE.md:852` lists only the artifact-reachability +commands, while `docs/dc-dc-deployment-workflow.md:206` says "Node-side +reachability **and the node time source** are gate G17" and the phase-4 runbook +requires `chronyc sources` to show the MAAS-served source, not the DC edge +(D-129(iv)). CURRENT-STATE also self-contradicts on whether DoD bullet 6 is STRUCK +or awaiting a DOCFIX. The observation window is one-time -- first boot. + +Separately (L4-7), G17's check as written cannot fail: `curl -sI` exits 0 on +404/500 and the dc0 URL is an autoindex root that answers 200 with nothing behind it. + +**Options:** (a) fold time verification into G17's `[V]` text and fix the check to +assert content with an exit-code predicate; (b) give time verification its own gate +row; (c) confirm it is STRUCK and remove the two conflicting surfaces. + +OPERATOR UTTERANCE: + +--- + +## R13. Credential reproducibility -- convert mint-refs now, or build `creds-mint.sh` first? + +**Finding:** L7-7 (measured: 30 rows across 16 ids carry `mint-ref=operator-terminal`; +`grep -rnI "ssh-keygen" .` returns ZERO hits repo-wide). + +Lens 7's assessment is worth quoting because it changes the shape of the fix: the +sharp edge is VR1-PRESENT, not Roosevelt-future -- SEC-007/-015 make edge SSH the +only management path, so a jumphost rebuild locks you out of both DC edges TODAY. +And the minimum fix needs no new tool: record each mint invocation as a numbered +runbook step and flip those rows' `mint-ref` from `operator-terminal` to +`runbook::`, which the existing S4 check already resolves. +`creds-mint.sh` is orthogonal -- it prevents the NEXT unregistered mint; it does not +make an existing key reproducible. + +**Options:** (a) convert the six edge/service/power key rows to `runbook:` refs +before Stage 5 (they are the unrecoverable-in-place ones); (b) build and rule +`creds-mint.sh` first, since Stage 5 is the largest minting event; (c) both, in +that order. + +OPERATOR UTTERANCE: + +--- + +## R14. The register cannot express a RULED exception + +**Finding:** carried from the 2026-07-27 close and re-measured today -- 3 of the 7 +standing credential findings are S5 power-key asymmetries that SEC-016 RULED to be +correct by design. The register has no way to say "this asymmetry is ruled", so it +reports a permanent red that a reader learns to ignore. That is how a real finding +gets lost. + +**Options:** (a) add a ruled-exception field to the matrix, citing the SEC/D number, +which the checker honours and prints; (b) leave it red and rely on prose; (c) rework +the S5 rule so a ruled per-DC divergence is representable. + +OPERATOR UTTERANCE: + +--- + +## R15. Should the gauntlet and repo-lint pin a floor? + +**Finding:** L4-8 and L4-1, both verified by this session. + +`run-tests-all.sh` counts what it DISCOVERS and compares that count to nothing, so a +renamed or deleted harness is neither run nor failed and the gauntlet still prints +ALL GREEN. `repo-lint` reports `PASS (0 fail, 0 warn)` over ZERO files given a +one-character typo. Both are the gates every stage close cites. The fixes are small +and mechanical, but they change what "green" means, so they are worth your explicit +sign-off rather than my assumption. + +**Options:** (a) add a floor to both (minimum harness count; refuse a non-directory +root and a zero-file scan); (b) floor on the gauntlet only; (c) leave as-is and rely +on the operator noticing a changed count. + +OPERATOR UTTERANCE: diff --git a/docs/audit/stage5-committee-raw-20260727.md b/docs/audit/stage5-committee-raw-20260727.md index 9a4d268..d9cdb80 100644 --- a/docs/audit/stage5-committee-raw-20260727.md +++ b/docs/audit/stage5-committee-raw-20260727.md @@ -510,3 +510,115 @@ - `bash scripts/cloud-assert.sh` -- not run; no OpenStack model exists yet. L4-10's A3 claim is from source plus a reproduction of the identical shell arithmetic. - G17's own check -- not runnable; all 18 nodes powered off. L4-7 is proved against a local HTTP server, not against `10.12.8.4`. - `creds-matrix.py --tier2 --remote --privileged` -- not run by this lens. + +--- + +## LENS 5 -- STAGE-5 PRE-MORTEM (returned seventh; first attempt died on an API error and was relaunched) + +LENS 5's framing note: Stage 5 executes FROM voffice1 (D-128 Plane 2). That clone is 105 commits behind `main` on a branch deleted upstream. **Every "artifact exists" row reads true only against the vcloud clone and is FALSE on the host where the commands actually run, until P1 completes. Tick nothing before P1.** + +**ID: L5-1** +CLAIM: `bundle.yaml`'s `ceph-osd` targets `/dev/vdb`, but every dc1 node has exactly one block device (`vda`) -- the deploy produces ZERO OSDs. +EVIDENCE: `maas admin machine read | jq` for all 9 dc1 nodes -> one device each, e.g. `{"h":"cute-satyr","mem":24576,"cpu":8,"bd":[{"name":"vda","size":590558003200,"type":"physical"}]}`. `bundle.yaml:560` `osd-devices: /dev/vdb`; `:533` `expected-osd-count: 4`. Definition side `opentofu/modules/node-vm/main.tf:26` + `:105-118` (`target.dev = "vda"`). Contrast the VR0 assumption still in the runbook: `runbooks/phase-01-bundle-deploy.md:125-131` globs `${h}-1.qcow2`, a second volume `node-vm` does not create. [INDEPENDENT CONFIRMATION of L2-1 by a different method and a different lens.] +BLOCKS-STAGE-5: yes | DISPOSITION: ruling + +**ID: L5-2** +CLAIM: Step 4 says "follow phase-01 verbatim", but phase-01 is VR0-frozen and **ACTIVELY REFUSES** to deploy dc1. +EVIDENCE: `runbooks/dc-dc-phase4-juju-bundle-per-dc.md:198-200`. In phase-01: hardcoded VR0 system_ids `for SID in 4na83t qdbqd6 h8frng tmsafc` (`:107`) and `select(.system_id|IN("4na83t",...))` (`:118`); a JUMPHOST-LOCAL libvirt loop `for h in openstack0 openstack1 openstack2 openstack3` over `/var/lib/libvirt/images/${h}-1.qcow2` (`:125-131`) -- dc1 nodes live inside `vvr1-dc1` on the dc1 rack, not on any jumphost; a plan gate of "56 apps, 108 relations, **4 machines**, 24 LXD" (`:55-58`, `:152`) against a 9-machine bundle; and the deploy guard `TOT=$(grep -cE '...vip:...10\.12\.4\.' bundle.yaml)` ... `if [ "$TOT" = 11 ] && [ "$HI" = 11 ] && [ "$LO" = 0 ]` (`:170-175`). For dc1 the VIPs are `10.12.64/68/72.x` in an OVERLAY, so the guard evaluates 0/0/0 and takes the `ABORT: VIP guard failed` branch (`:176`). **After the ruling-3 extraction it evaluates 0/0/0 for dc0 too.** +BLOCKS-STAGE-5: yes | DISPOSITION: DOCFIX (phase-4 must stop delegating "verbatim" and carry its own DC-parameterised deploy step, or phase-01 must be DC-parameterised) + +**ID: L5-3** +CLAIM: A dc1-correct `overlays/octavia-pki.yaml` CANNOT be produced by the documented generator, and its absence hard-aborts the deploy. +EVIDENCE: `runbooks/phase-01-bundle-deploy.md:144-145` `ABORT: overlays/octavia-pki.yaml missing (Step 1.0)`. Generator 1.0-GEN.c at `:369-373` reads `bundle.yaml`'s `octavia.options.vip` then hard-gates `grep -qE '^10\.12\.4\.[0-9]{1,3}$' || { echo "FAIL: implausible VIP -- stop"; exit 1; }`. dc1's octavia VIP is `10.12.64.57` and lives in `overlays/vr1-dc1-vips.yaml:37`, ABSENT from `bundle.yaml`. Identity strings are dc0/VR0 literals: `CN = octavia-controller.omega.dc0.vr0.cloud.neumatrix.local` (`:344`), `/CN=VR0 DC0 Omega Cloud Octavia Controller CA` (`:327`). phase-4 frames reuse-vs-regenerate as an operator call (`:204-209`) **without noting that regenerate is currently impossible**. +BLOCKS-STAGE-5: yes | DISPOSITION: ruling (per-DC trust domain vs reuse) + DOCFIX (the generator's dc0-frozen VIP gate and CN/SAN literals) + +**ID: L5-4** +CLAIM: The Stage-5 execution host's clone is 105 commits behind on a branch DELETED upstream, so a plain `git pull` does not recover it. +EVIDENCE: `ssh voffice1 'git ... status -sb; log -1; rev-list --count HEAD..origin/main'` -> `## dc-dc-g12-dc1-substrate...origin/dc-dc-g12-dc1-substrate` / `61c416e` / `105`. `git ls-remote --heads origin` -> only `dc-dc-stage5-grounding-audit` and `main`. +BLOCKS-STAGE-5: yes | DISPOSITION: DOCFIX (add an explicit step `git fetch origin && git switch main` with a HEAD-equals-origin/main assertion; phase-4's prerequisites at `:23-34` name NO repo-state precondition at all) + +**ID: L5-5** +CLAIM: Step 12's Ceph gate uses `juju run` to execute an arbitrary shell command, but on the juju actually installed `juju run` is the **ACTION** runner. +EVIDENCE: `runbooks/dc-dc-phase4-juju-bundle-per-dc.md:350`: `juju run -m ceph-mon/leader 'ceph -s'`. MEASURED: `ssh voffice1 'juju version'` -> `3.6.27-genericlinux-amd64`; `juju help run` -> `Usage: juju run [options] ... ` / "Run an action on a specified unit"; `juju help exec` -> `Usage: juju exec [options] `. As written the gate asks ceph-mon to run an action named `ceph -s`. +BLOCKS-STAGE-5: yes (a named exit-gate command that cannot execute) | DISPOSITION: DOCFIX (`juju exec -m --unit ceph-mon/leader -- ceph -s`, re-verified against 3.6.27 before landing) + +**ID: L5-6** +CLAIM: Step 12's geneve gate greps for an `ovn-central` config key that Step 6 of the SAME runbook establishes does not exist -- the exit gate can never pass. +EVIDENCE: `:352` `juju config -m ovn-central | grep -i encap`. Same file `:270-272`: "**NO entry for OVN** (confirmed no such charm-config option exists -- geneve family follows the bound interface automatically once the plane itself is ULA-only)." Exit gate `:360-363` requires "Ceph-over-v6 and geneve-over-v6 verified (Step 12)". A `grep` with no match exits 1 and prints nothing; no pass criterion is stated either way. +BLOCKS-STAGE-5: yes | DISPOSITION: DOCFIX (DOCFIX-204 class -- the gate must assert the OBSERVABLE, e.g. the geneve encap address family on a live chassis, not a nonexistent charm option) + +**ID: L5-7** +CLAIM: Step 11's controller-backup commands are wrong-shaped for the installed juju. +EVIDENCE: `:334-335` `juju create-backup -m ` / `juju download-backup -m `. MEASURED on 3.6.27: `juju help download-backup` -> `Usage: juju download-backup [options] /full/path/to/backup/on/controller` (a controller-side PATH, not a backup-id); `-m, --model ... Accepts [:]|` (a MODEL, not a controller name). `juju help create-backup` -> `--filename (= "juju-backup--