diff --git a/.claude/skills/openstack-cloud-ops/references/platform-traps.md b/.claude/skills/openstack-cloud-ops/references/platform-traps.md index e9b0cba..8612fdc 100644 --- a/.claude/skills/openstack-cloud-ops/references/platform-traps.md +++ b/.claude/skills/openstack-cloud-ops/references/platform-traps.md @@ -71,8 +71,8 @@ `/boot.config` (262 bytes of serial output) and then triple-faults handing off to `/boot/loader`. `tofu validate` CANNOT catch it (the attribute is optional) -- `scripts/opentofu-validate.sh` check **S1** does. -Evidence: `docs/changelog-20260712-libvirt-memory-unit-rootcause.md`, -`docs/incident-20260712-opnsense-edge-boot-triplefault.md`. +Evidence: `docs/archive/changelogs/changelog-20260712-libvirt-memory-unit-rootcause.md`, +`docs/archive/incident-20260712-opnsense-edge-boot-triplefault.md`. The provider's own current domain examples now carry `memory_unit = "MiB"` (`docs/resources/domain.md`) -- a fresh reader of the docs would get this @@ -101,7 +101,7 @@ - **MAAS drives power off/on via ACPI**, so a MAAS-managed node VM without ACPI can only ever be hard-stopped. That would have been a nasty Stage-3 debug. -Evidence: `docs/changelog-20260712-libvirt-acpi-kernel-panic.md`. +Evidence: `docs/archive/changelogs/changelog-20260712-libvirt-acpi-kernel-panic.md`. Guard: `scripts/opentofu-validate.sh` check **S2**. ### 1d. No `cpu` block => generic CPU => NO `svm` => nested KVM impossible @@ -137,7 +137,7 @@ nested KVM at that depth; and `nova-compute` on the DC nodes needs it one level deeper still. `modules/node-vm` still has the missing-`cpu`-block defect -- logged, NOT fixed, because DC1 is gated behind Office1 -(`docs/changelog-20260713-d114-voffice1-nested-virt.md`). +(`docs/archive/changelogs/changelog-20260713-d114-voffice1-nested-virt.md`). ### 1e. "will be updated in-place" DOES NOT mean "the guest stays up" @@ -238,7 +238,7 @@ (pool parent overridable via `VR1_POOL_PARENT`), wired into `scripts/prereqs/install-all.sh` and reported by `check-prereqs.sh`. It cost a session on 2026-07-12 (DOCFIX-186) because it existed only as hand-applied host -state; `docs/changelog-20260713-apparmor-libvirt-prereq.md`. +state; `docs/archive/changelogs/changelog-20260713-apparmor-libvirt-prereq.md`. ## 3. MAAS + LXD @@ -396,7 +396,7 @@ It already cost a defect: a pre-install snapshot silently no-op'd and the install then ran with no rollback point behind it -(`docs/changelog-20260713-office1-dhcp-apply.md`). Because the step was +(`docs/archive/changelogs/changelog-20260713-office1-dhcp-apply.md`). Because the step was non-fatal, nothing stopped. **Always feed remote commands to `sh -s`:** diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index 0c73a00..a62e838 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -202,7 +202,7 @@ |---|------|----------------|-------|---------------------------| | G1 | Audit Phase 3: fresh-agent grounding test | [V] 3 clean-context probes score the 7-question set against this doc; holes map made | session | CLOSED 2026-07-18: 3 probes, 21/21 PASS, holes H1 (amended into G9) + H2 (no action) -- `docs/audit/phase3-grounding-test-20260718.md` | | G2 | Audit Phase 4: GA-R1..R7 structural rulings + the stage-status vocabulary A/B | [R] ruling-type gate (GA-R6 rule 6): closes when every item carries a GA-R5 Status block | operator | CLOSED 2026-07-18: all seven GA-R + vocabulary (Option A + H1) RATIFIED, utterances quoted (`docs/audit/ga-rulings.md`, through commit `fe4f1c4` + this one) | -| G3 | Audit Phase 5: repair sweep of GA-F01..F15 (incl. memory hygiene GA-F05..F08, skill sweep) | [R] operator-gated fix batches, each commit naming its GA-F | operator + session | Batch 0 OPENED by operator 2026-07-19; items 0.1 (repo-lint L10, GA-R1/C1), 0.2 (SEC repoint, GA-R4/F3), 0.3 (counter hardening, GA-F15), 0.4 (extractor vocab scan, GA-F10/H1) landed; Batch 0 CLOSED (verification passed 2026-07-19); Batches 0-3 CLOSED 2026-07-19 (Batch 3: memory hygiene -- GA-F06/F07 files removed per provenance rulings, index rebuilt, skill line added, G1 review zero violations, new finding GA-F16 registered+fixed; pre-edit copies docs/archive/memory-20260718/); Batches 4-6 await gates; Batches 2-6 await gates; FREEZE holds for un-gated surfaces | +| G3 | Audit Phase 5: repair sweep of GA-F01..F15 (incl. memory hygiene GA-F05..F08, skill sweep) | [R] operator-gated fix batches, each commit naming its GA-F | operator + session | Batch 0 OPENED by operator 2026-07-19; items 0.1 (repo-lint L10, GA-R1/C1), 0.2 (SEC repoint, GA-R4/F3), 0.3 (counter hardening, GA-F15), 0.4 (extractor vocab scan, GA-F10/H1) landed; Batch 0 CLOSED (verification passed 2026-07-19); Batches 0-4 CLOSED 2026-07-19 (Batch 4: GA-R4 ledger rotation 1187->131 lines, F1 cap now enforceable; 96 changelogs + 24 history docs consolidated to docs/archive/ with 4 stage records + per-stage manifest commits; top-level docs/ = 16 files < 25; live-surface refs rewritten); Batches 5-6 await gates; Batches 2-6 await gates; FREEZE holds for un-gated surfaces | | G4 | The two D-130 verifications | [V] run them, capture output | session | CLOSED 2026-07-19: v8 suppression CONFIRMED (7/2/7 -> 6/2/6, zero forces-replacement; `docs/audit/outer-plan-20260719-v8-ignorechanges.txt` + `-baseline.txt`); v7 no-bounce under running domain, zero residue (`docs/audit/throwaway-v7-20260719.txt`) | | G5 | D-130 mechanism ruling (seed-volume durable fix) | [R] operator rules in Phase 5, quoting G4's captured output | operator | CLOSED 2026-07-19: D-130 ADOPTED (a) ignore_changes (`docs/design-decisions.md` D-130, GA-R5 utterance quoted); implemented in `modules/cloudinit-vm` + `tests/cloudinit-vm` | | G6 | State reconcile of autostart + seed WITHOUT bouncing guests | [R] gated mechanism, operator-ruled (S3) | operator | CLOSED 2026-07-19: ruled (ii) state surgery (GA-R5); pull -> inject autostart:true on both domains -> push (serial 22->23, backup `terraform.tfstate.pre-G6-20260719`); guests never touched (ids 1/2 unchanged, running) | diff --git a/docs/D-057-DECIDED-append.md b/docs/D-057-DECIDED-append.md deleted file mode 100644 index 258bfcf..0000000 --- a/docs/D-057-DECIDED-append.md +++ /dev/null @@ -1,37 +0,0 @@ -> SUPERSEDED IN PART BY D-058 (full plane renumber, 2026-06-29): provider-vip is now -> **10.12.8.0/22** (gateway 10.12.8.1), not 10.12.24.0/22. metal-admin->.12, metal-internal->.16, -> data-tenant->.20, oob->.60. The .24/.8/.12 values below are the original D-057 instantiation, -> kept for history; the durable framework (separate tagged routed VIP plane) is unchanged. -> See D-058-renumber.md for the authoritative scheme. - -## D-057: Provider NIC L3-on-OVS-uplink -- split the public API VIP plane onto its own routed VLAN (2026-06-27) - -**Status:** DECIDED. Root cause CONFIRMED live (this session) and direction CONFIRMED by operator. Supersedes the standalone PROPOSED/OPEN note `D-057-provider-uplink-l3-separation.md` and resolves its one open question (the D-003B interaction) in favor of option (a)/(c). Implementation pending: MAAS plane add -> `carve-host-interfaces.sh` delta -> bundle delta -> teardown/redeploy -> D-011 re-validate. grep-before-assign confirmed D-057 free (max prior D-053; D-054/055/056 are DOCFIX). - -**Symptom:** every provider-ext floating IP (pool 10.12.5.0-10.12.7.254) unreachable cloud-wide; the .4.x public API VIPs answered. Blocked phase-06 Step 6.3 (SSH to capi-mgmt-v2 via FIP 10.12.7.107). - -**Root cause (CONFIRMED, not inferred -- measured on all three ovn-chassis hosts this session):** a two-consumer collision on the untagged provider NIC `enp1s0`. (1) ovn-chassis is configured `bridge-interface-mappings: br-ex:` and wants `enp1s0` as an OVS `br-ex` port. (2) ~11 `public: provider-public` API charms are LXD containers; Juju attaches a container to a subnet by bridging the host NIC, which created the Linux bridge `br-enp1s0` and captured `enp1s0`. A single untagged physical NIC cannot be both a Linux-bridge member and an OVS `br-ex` port: the Linux bridge won, `br-ex` was starved (no-carrier on all three), FIPs dark while the containerized .4.x VIPs kept answering. Evidence: `br-ex` operstate=down / carrier=Invalid-argument on openstack1/2/3; `enp1s0 master br-enp1s0` on all three; `br-enp1s0` ports = `enp1s0` + container `veth*` taps (1/3/5 taps); host static (.41/.42/.43) riding `br-enp1s0`; no public VIP on any host (all in containers); host default route via metal-admin `10.12.8.1`, provider `10.12.4.1` only on-link. The host has NO provider-routing role -- only its containers do. - -**Why option (b) was rejected:** "remove the host provider static" addresses consumer (1)'s L3 but NOT consumer (2). `br-enp1s0` exists for container attach; removing the host static cannot release `enp1s0`. Refuted by the measured `veth` taps. - -**Decision (a)/(c):** relocate the container `public` attach off the untagged uplink onto a tagged sub-interface, freeing untagged `enp1s0` for OVS `br-ex`. This is the Canonical shared-NIC pattern and is the exact mirror of the existing metal-internal tagged-secondary stack (`br-metal.103 -> br-internal`). Because a single CIDR cannot span two L2s, the public API VIPs re-IP onto a new subnet on the tagged VLAN -- which removes D-003B's same-L2 property and replaces it with L3 routing (see amendment below). - -**Target interface tree (per host; octet N = .40-.43 by HOST_OCTET):** -- `enp1s0` -- RAW, untagged on fabric 1_provider, **NO L3 / NO subnet link / NO bridge**. ovn-chassis MAC-enslaves it into OVS `br-ex`. (The old host static 10.12.4.N is removed FROM THIS UPLINK -- that removal is the fix.) -- `enp1s0.104` (VLAN, VID 104, on fabric 1_provider) -> `br-prov-api` (standard bridge) -> **STATIC 10.12.24.N**. Carries the new `provider-vip` space. Container `public` endpoints bind here. The host's provider-plane presence MOVES from `enp1s0` to `br-prov-api` (tagged, OVS-free, no `br-ex` competition -> zero D-057-class risk), mirroring metal-internal's `br-internal` host static -- the proven container-attach pattern on this cloud. (L3-less `br-prov-api` considered and rejected: unproven for Juju attach here; not worth gambling the redeploy.) - -**New plane (framework, not final subnetting):** `provider-vip` -- 10.12.24.0/22, VID 104, fabric 1_provider, **routed** (gateway 10.12.24.1), VIP reserve 10.12.24.2-.100 (VIPs .50-.60). The subnet/VID values are this deployment's instantiation; the DURABLE decision is the framework: the public API VIP plane is its own tagged, routed plane distinct from the provider ext_net/FIP plane. The DC-DC byte-aligned plan adopts the same split (its provider VLAN 240 -> 240 ext_net + 241 provider-vip; its already-present NN-11 role is the home). - -**provider-public after the split:** keeps only FIP/ext_net (10.12.5.0-10.12.7.254) + the OVN gateway SNAT + mgmt reserve; untagged enp1s0 -> OVS br-ex; NO container attach, NO host L3. Neutron stays flat physnet1 (no retype) -- the OVN/Neutron layer, confirmed correct this session, is untouched. - -**D-003B AMENDMENT:** D-003B deliberately co-located public API VIPs and FIPs on one provider L2 ("tenant->API by construction") -- a Bobcat-era convenience that this session proved unworkable on the NIC-limited host (the OVS-vs-container collision). It is hereby amended: the public API VIP plane is a separate, routed VLAN. tenant->API is preserved via L3 routing (tenant SNAT egress on provider 10.12.4.0/22 -> gateway -> provider-vip 10.12.24.0/22), and re-validated per D-011 #3. This is also better for the commercial hard-isolation goal (API VIPs out of the FIP broadcast domain; an L3 policy enforcement point) and matches the Roosevelt NIC-limited reality. - -**ROUTED-GATEWAY PREREQUISITE (gating, highest risk):** today only provider-public and metal-admin route (as-built: all other planes `gateway: none`). provider-vip must be routed: gateway 10.12.24.1 must exist and route to/from 10.12.4.0/22 on whatever owns 10.12.4.1 (jumphost/libvirt), with the return path 10.12.24.1 -> 10.12.4.0/22. Without it the redeploy passes D-011 #1/#2 and silently fails #3. Confirm/establish before redeploy -- do not discover at validation. - -**bridge-interface-mappings:** trim openstack0's provider MAC `52:54:00:3d:fd:54` (measured: the only one of the four with no ovn-chassis unit; nova-compute/ovn-chassis run on machines 9/10/11 = openstack1/2/3). Keep `9d:63:77`/`89:7f:ce`/`99:fc:c2`. Settles openstack0 as control+storage role (feeds Roosevelt node-role split). The mapping still targets the untagged `enp1s0` MAC (now free). - -**bundle changes (summary):** on the 11 `public:`-bound API charms, rebind `public: provider-public -> public: provider-vip`, and re-IP each `vip:` first token `10.12.4.X -> 10.12.24.X` (X=50-60; metal-admin .8.X and metal-internal .12.X tokens unchanged). Trim the openstack0 MAC. No Neutron/OVN config change. - -**Roosevelt relevance:** the provider-vip tagged-routed-plane split joins metal-internal VID-103 as a per-host interface-tree region-invariant (maas-as-built reference). The DC-DC byte-aligned plan inherits the split. - -**Related:** root-causes phase-06 6.3; amends D-003B; extends D-052/D-053 plane model with the host-interface-layer rule the rebuild exposed; the committed `docs/maas-as-built-reference.md` carve table needs the follow-on correction (enp1s0 raw/L3-less + the enp1s0.104->br-prov-api stack). diff --git a/docs/D-057-REVIEW-ITEMS.md b/docs/D-057-REVIEW-ITEMS.md deleted file mode 100644 index 51ec889..0000000 --- a/docs/D-057-REVIEW-ITEMS.md +++ /dev/null @@ -1,58 +0,0 @@ -# D-057 -- review items for END OF DEPLOYMENT (do not action mid-deploy) - -These are real findings surfaced while building the D-057 provider-vip split. None -block the deploy; each is deferred to a post-D-011 reconciliation sweep so we don't -re-architect inside a step. Logged here so they are not lost. - --------------------------------------------------------------------------------- -R1. provider-vip-maas-standup.md CREATE blocks are now redundant with the script. - scripts/provider-vip-standup.sh is the tested execution path for creating the - plane. The runbook's Phase-2 create blocks duplicate that logic (two sources of - truth -- this is how the MTU-source bug nearly drifted between them). KEEP the - runbook for its still-unique parts: Phase-1 audit, the virbr1 vlan_filtering gate, - and the deferred jumphost-gateway reads. RECONCILE: trim the create blocks to a - pointer at the script, or annotate them "superseded -- see script". (User: noted - for end-of-deployment review, not trimmed now.) - --------------------------------------------------------------------------------- -R2. scripts/review-bundle.py is STALE (pre-D-052) and is NOT a current gate. - Run against today's committed bundle it reports ~30 FAILs: its PHANTOM_BINDING_KEYS - check forbids exactly the per-endpoint bindings (shared-db/amqp/certificates/cluster/ - ha/internal) that D-052 deliberately added, and its VIP check expects DUAL VIPs at - octets .224-235 in the old 10.12.8 "metal" net -- the bundle has TRIPLE VIPs at - .50-.60 across provider-public/metal-admin/metal-internal. So review-bundle.py - predates D-052 entirely and does not describe the live bundle. - - The "verify_bundle.py / 8/8 harness" referenced in the redeploy notes is NOT at - this repo HEAD (d575a25) -- only review-bundle.py is. - - INTERIM GATE for the D-057 change: scripts/d057-bundle-check.py (focused, fail- - closed, proven FAIL-on-pre / PASS-on-post). It checks only the D-057 invariants. - - RECONCILE: either bring review-bundle.py forward to D-052/053/D-057 (rewrite - PHANTOM check, VIP check to triple + provider-vip 10.12.8, octets 50-60), or - restore/commit the newer verify_bundle.py and retire review-bundle.py. - --------------------------------------------------------------------------------- -R3. Committed bundle uses machines 8/9/10/11; the live cloud ran 0/1/2/3. - Repo-fidelity gap (the committed bundle is not the as-deployed bundle). Does NOT - affect D-057 (the carve is hostname-based; the MAC trim is by MAC). RECONCILE the - bundle machine block + `to:` placements to whatever the redeploy actually uses. - --------------------------------------------------------------------------------- -R4. oob CIDR -- RESOLVED by D-058. oob adopts 10.12.60.0/22 (the design-docs value). - The live virsh power gateway 10.12.64.1 -> 10.12.60.1 (scripts/lib-hosts.sh - VIRSH_POWER_ADDRESS) is part of the D-058 foundation cascade -- see D-058-renumber.md. - --------------------------------------------------------------------------------- -R5. gateway_ip default-route watch-item (post-deploy, not end-of-deploy). - provider-vip subnet carries gateway_ip=10.12.8.1 (post-D-058). This mirrors provider-public's - existing .4.1 (which already coexists with metal-admin .8.1 without hijacking the - node default), so risk is low -- but VERIFY after redeploy that every API unit's - default route is still via metal-admin 10.12.12.1, not .8.1: - juju exec --all -- ip route show default - If any unit defaults via .8.1: blank the subnet gateway_ip (set PVIP_GATEWAY="" in - scripts/provider-vip-standup.sh and re-apply) or pin the node default-gateway subnet - to metal-admin. provider-vip reachability does not depend on its own gateway in v1. - --------------------------------------------------------------------------------- -R6. ovn-chassis-octavia has no bridge-interface-mappings (expected -- it is the - octavia-side chassis). Left unchanged. Noted only so it is not mistaken for an - omission during review. diff --git a/docs/D-058-renumber.md b/docs/D-058-renumber.md deleted file mode 100644 index 590f13a..0000000 --- a/docs/D-058-renumber.md +++ /dev/null @@ -1,77 +0,0 @@ -# D-058: full plane renumber -- clean fabric-grouped /22 scheme (2026-06-29) - -**Status:** DECIDED (operator). Supersedes the D-057 minimal-delta placement of -provider-vip at 10.12.24.0/22, and resolves R4 (oob). grep-before-assign: D-058 free -(max prior D-057; D-054/055/056 are DOCFIX). - -**What:** renumber the v1 plane scheme so CIDRs are contiguous /22 blocks grouped by -fabric and ordered to match the layer model, instead of the historical scatter. This -is a cloud-wide re-IP, intentionally larger than D-057, accepted by the operator for -Roosevelt addressing fidelity. It is executed as a teardown/redeploy (no in-place -re-CIDR), so there is no transient subnet overlap. - -## The map (authoritative) - -| Plane | old | NEW | gateway (was -> now) | -|-----------------|---------------|---------------|-----------------------------| -| provider-public | 10.12.4.0/22 | 10.12.4.0/22 | 10.12.4.1 (unchanged) | -| provider-vip | 10.12.24.0/22 | 10.12.8.0/22 | 10.12.24.1 -> 10.12.8.1 | -| metal-admin | 10.12.8.0/22 | 10.12.12.0/22 | 10.12.8.1 -> 10.12.12.1 | -| metal-internal | 10.12.12.0/22 | 10.12.16.0/22 | none (L2 east-west) | -| data-tenant | 10.12.16.0/22 | 10.12.20.0/22 | none (isolated L2) | -| storage | 10.12.32.0/22 | 10.12.32.0/22 | none (unchanged) | -| replication | 10.12.36.0/22 | 10.12.36.0/22 | none (unchanged) | -| oob | 10.12.64.0/22 | 10.12.60.0/22 | 10.12.64.1 -> 10.12.60.1 | - -Rotate rule (collision-safe): 8->12, 12->16, 16->20, 24->8, 64->60; 4/32/36 fixed. -VLAN IDs unchanged (metal-internal VID 103, provider-vip VID 104). VIP triple becomes -provider-vip .8.5x / metal-admin .12.5x / metal-internal .16.5x (octets 50-60). -Host statics (.40-.43) follow each plane. metal-admin PXE-DHCP band -> 10.12.12.9-.11. - -## JUMPHOST ORDERING TRAP (must respect on the host) - -The jumphost owns three gateways that move. provider-vip's NEW gateway 10.12.8.1 -is metal-admin's OLD address. So on the jumphost, in this order: - 1. move virbr2 (metal-admin) 10.12.8.1 -> 10.12.12.1 - 2. move virbr7 (oob) 10.12.64.1 -> 10.12.60.1 - 3. THEN add virbr1.104 (provider-vip) = 10.12.8.1 <-- only after step 1 frees .8.1 -Adding virbr1.104=.8.1 while virbr2 still holds .8.1 is a same-subnet collision. In a -clean rebuild the bridges are reconfigured as a set, but the free-then-claim order -still applies. (Step 3 is the jumphost-provider-vip-gateway.md runbook.) - -## APEX / NetBox note (IaC discipline) - -NetBox is the apex; this renumber's authoritative home is NetBox. BUT the committed -netbox/ipv4-prefixes-import.py is itself stale (pre-D-052: only Metal/Provider/LBaaS-mgmt, -provider VLAN VID 240, API VIPs at .224-.254 -- none of the 6-plane D-052/053 model). It -must FIRST be brought current to D-052/053, THEN carry the D-058 scheme, before it can be -the source of truth. Until that reconciliation, scripts/lib-net.sh is the working contract -and already carries D-058. Do NOT hand-edit downstream MAAS for these values once NetBox -is current -- regenerate. - -## DONE in this pack (renumbered + re-validated: both suites ALL PASS, d057-check PASS) - -scripts/lib-net.sh (PLANE_CIDRS, PLANE_NAME, PLANE_GW, DATA_PLANE_CIDRS, -METAL_INTERNAL_CIDR, PROVIDER_VIP_CIDR, VIP_PREFIX_* triple), -scripts/carve-host-interfaces.sh, scripts/provider-vip-standup.sh, -scripts/d057-bundle-check.py, bundle.yaml (11 VIP triples), both test suites + -fixtures, provider-vip-maas-standup.md, jumphost-provider-vip-gateway.md, README. - -## COMMITTED-FOUNDATION CASCADE (still on the OLD scheme -- next sweep) - -Apply the same rotate (8->12, 12->16, 16->20, 24->8, 64->60; 4/32/36 fixed). These are -in the committed repo, not this pack, and several are prose runbooks -- sweep with care, -NetBox-anchored: - - netbox/ipv4-prefixes-import.py (APEX -- de-stale to D-052/053 first, then D-058) - - netbox/README.md - - scripts/phase-00-maas-carve.sh (METAL_CIDR default 10.12.8 -> 10.12.12; ranges) - - scripts/lib-hosts.sh (VIRSH_POWER_ADDRESS 10.12.64.1 -> 10.12.60.1) - - scripts/review-bundle.py (stale pre-D-052 already -- R2; fold in with that) - - runbooks/phase-00-teardown-maas-reset.md, phase-01-bundle-deploy.md, - phase-03-core-verify.md, phase-04-network-carve.md, phase-05-octavia-enablement.md, - phase-08-workload-cluster-acceptance.md, appendix-A-troubleshooting.md - - docs/maas-as-built-reference.md, docs/design-decisions.md (append D-058), - docs/v1-redeploy-changelog.md, docs/netbox-vip-queue.md - - tests/phase-00-carve/run-tests.sh, tests/phase-04/make_fixtures.py - - jumphost underlay: virbr2 -> 10.12.12.1, virbr7 -> 10.12.60.1 (see ordering trap) - - host-nginx :81 upstream: Horizon -> 10.12.8.58 diff --git a/docs/D-068-openbao-assessment-DRAFT.md b/docs/D-068-openbao-assessment-DRAFT.md deleted file mode 100644 index 4a6ec5d..0000000 --- a/docs/D-068-openbao-assessment-DRAFT.md +++ /dev/null @@ -1,108 +0,0 @@ -# D-068 item 1 -- OpenBao / off-EOL-vault assessment (RESEARCH DRAFT) - -Status: RESEARCH INPUT ONLY (2026-07-06, jumphost stream, web research). The D-068 -off-EOL path decision remains OPEN and operator-ruled; this draft exists so the B2 -discussion starts from verified facts. Items that could not be confirmed are marked -UNVERIFIED. Companion: docs/D-068-vault-1.8-vs-1.16-analysis.md (the 1.16 ruling). - -## Summary - -- OpenBao is healthy upstream (Linux Foundation, MPL-2.0; latest stable 2.5.5, - 2026-06-17) and API-compatible for our surfaces (kv v1/v2, AppRole, PKI). BUT: - **no Juju charm for OpenBao exists anywhere** (Charmhub API search empty, - 2026-07-06), and **OpenBao REMOVED the MySQL storage backend** (raft/postgres - only; upstream issue #651). An OpenBao path today = bespoke charm development - we would own + a storage migration -- it cannot serve our 18 V0 `certificates` - relations. -- **No plan exists to bring tls-certificates V1 to the legacy/reactive OpenStack - charms.** Canonical's V1 world is Sunbeam ("Canonical OpenStack", 2024.1 LTS, - up-to-12-year support): vault + manual-tls-certificates + traefik on V1. The - charm-guide still documents legacy TLS via the V0 vault relation. Practical - read: V1 will likely NEVER come to the reactive set; "wait for V1" is risk - acceptance without a deadline. (Absence-of-evidence finding -- flagged as such.) -- The operator-lineage vault charm moved on (tracks to 1.19 + a new 2.0 track, - 2026-07-02, tracking IBM Vault 2.0 GA 2026-04-14) -- same disqualifiers as the - ruled-out 1.16 (V1-only, Raft-only, BUSL). The reactive 1.8/stable charm is in - steady-state maintenance and still being rebuilt (Charmhub revision 2026-07-03). -- CVE posture of 1.8.8: no known unauthenticated remote-compromise CVE applies to - an internal-only, AppRole-driven, TLS-issuing deployment. Notable but gated: - CVE-2021-45042 (PKI wildcard issuance to AUTHORIZED users, medium, in-range); - CVE-2024-2048 (cert-auth bypass, 9.8 -- only if the cert auth method is enabled, - ours is not); Aug-2025 "Vault Fault" set incl. CVE-2025-6000 RCE (requires - privileged plugin-catalog access; per-CVE 1.8.8 applicability UNVERIFIED). - Honest characterization: **manageable today, unauditable tomorrow** -- nobody - assesses 1.8.8 anymore; each future disclosure needs our own analysis with no - patch option. Supports deadline-bound acceptance, not indefinite stay. - -## Candidate paths compared - -1. **Wait-for-V1 in legacy charms** -- no published plan; strong signal V1 - investment goes to Sunbeam only. Equivalent to acceptance WITHOUT a deadline. - Reject as a plan; it is the null hypothesis. -2. **OpenBao now** -- right long-term horse upstream, but requires (a) a charm - that does not exist (bespoke, we own it), (b) storage migration off MySQL, - (c) an aggressive support cadence (only latest cycle maintained). Not viable - as a like-for-like swap; keep on the WATCH list (an OpenBao charm appearing, - or any Canonical position on OpenBao, changes this materially). -3. **Risk-acceptance with a hard deadline (research recommendation)** -- stay on - 1.8/stable (Canonical still rebuilds it; commercially supported under - Pro/PCB). Compensating controls: mgmt-plane-only API exposure, audit logging - shipped, minimal policies, confirm userpass/cert auth methods and plugin - registration absent, unseal/root discipline (already the norm -- D-069). - Hard re-decision deadline proposed: 2027-04. Two follow-ups; (a) is now - MEASURED (read-only, 2026-07-06, evidence - ~/openstack-baseline/vault-config-20260706.yaml + vault-snap-tracks-20260706.txt): - the reactive charm EXPOSES a `channel` config (currently 1.8/stable; its own - description warns a change seals ALL vault units via the snap refresh -- - manual 3-of-5 unseal is our norm, so that is a planned-window event, not a - blocker). The Canonical vault snap on the live unit publishes tracks - 1.1-1.12, then 1.15-1.20 and 2.0: **there is NO 1.13 or 1.14 track** -- the - last MPL-licensed payload reachable via the charm is **1.12.11** - (1.12/stable), not 1.14.x. So the in-charm payload-bump option is: - 1.8.8 -> 1.12.11 (stays MPL; still EOL upstream, but 4 minors newer), or - 1.15.6+ (BUSL -- licensing review required). Remaining UNVERIFIED before any - such move: charm-vs-newer-payload API compatibility (does the reactive charm - drive 1.12 correctly -- Canonical test coverage unknown), storage-format - stepping across 4 minors on vault-on-mysql, and whether MySQL backend - behavior changed in 1.9-1.12. This is a rehearse-first D-NNN of its own if - pursued; NOT a routine refresh. (b) track Sunbeam feature parity -- the - durable exit from the V0 certificates world is a PLATFORM decision, not a - vault swap. Also verified live 2026-07-06: the charm's Vault-listener TLS - options exist and are empty (ssl-ca / ssl-cert / ssl-chain / ssl-key) -- - the concrete D-068 item 2 mechanism, no fabricated option names. - -## Decision hooks for the operator (nothing ruled yet) - -- Adopt path 3 with deadline? (Composes with keeping path 2 on watch.) -- Authorize the read-only snap-channel headroom measurement (follow-up a)? -- Fold "Sunbeam parity watch" into the Roosevelt planning inputs? -- D-068 items 2 (listener TLS) and 3 (AppRole TTL audit) are unaffected by this - and proceed on their own track. - -## Sources (accessed 2026-07-05/06; full trail) - -- openbao.org: release-notes 2.6.0-beta (2026-06-22); blog/roadmap-2.0 (LF - governance); docs/configuration/storage; docs/guides/migration -- endoflife.date/openbao (2.5.5 latest stable, 2026-06-17; latest-cycle-only support) -- github.com/openbao/openbao issue #651 (MySQL backend removed/requested) -- api.charmhub.io/v2/charms/find?q=openbao (EMPTY, 2026-07-06); charmhub.io/openbao (404) -- charmhub.io/vault (tracks 1.5-1.19 + 2.0 candidate/beta/edge; 1.8/stable rev 2026-07-03) -- canonical-vault-charms.readthedocs-hosted.com (2.0 upgrade docs; 1.8-vs-1.15 differences) -- opendev.org/openstack/charm-vault commits (steady-state; last commit 2025-10-29) -- docs.openstack.org/charm-guide latest: admin/security/tls (V0 model current) -- canonical-openstack.readthedocs-hosted.com: vault feature, service-endpoint-encryption, - release-cycle (Sunbeam 2024.1 LTS, 12-year support) -- infoq.com 2026/04 vault-2-0-ibm-identity + IBM support-lifecycle addendum - (Vault 2.0 GA 2026-04-14; 1.21->2.0 renumbering) -- nvd.nist.gov CVE-2024-2048 (9.8 NIST / 8.1 vendor; <=1.14.9; cert auth only) -- cyata.ai vault-fault (Aug-2025 nine zero-days incl. CVE-2025-6000 RCE) -- stack.watch + cvedetails (historic ranges; NOTE conflicting aggregator text for - CVE-2021-45042 / CVE-2021-43998 -- re-check NVD before quoting in a decision record) -- stackhpc-kayobe-config readthedocs stackhpc-2025.1 configuration/openbao - (non-Juju OpenStack distro already uses OpenBao for internal PKI) -- duokey.com BUSL-change/OpenBao-migration backgrounder (2023-08-10 BUSL date) - -NOT confirmed despite searching: any Canonical statement on OpenBao; any V1 plan -for legacy charms; exact Vault 1.8.x EOL date/final point release; per-CVE -applicability of the 2025 Cyata set to 1.8.8; reactive-charm snap-channel -headroom (needs on-host measurement, not web research). diff --git a/docs/DOCFIX-064-phase08-changelist.md b/docs/DOCFIX-064-phase08-changelist.md deleted file mode 100644 index d7224b7..0000000 --- a/docs/DOCFIX-064-phase08-changelist.md +++ /dev/null @@ -1,81 +0,0 @@ -# DOCFIX-064 -- phase-08 runbook change-list (DRAFT, 2026-07-01) - -RESERVED number: DOCFIX-064 (per changelog next-free note). This is the accumulated -phase-08 operator-runbook (single-consumer acceptance) sweep. Written as a change-LIST with -exact anchors + evidence so the edit is mechanical when phase-08 is finalized. NOT yet applied -to runbooks/phase-08-workload-cluster-acceptance.md. - -Scope note: these are fixes to the OPERATOR single-consumer acceptance path (capi-test-1 in -capi-mgmt scope). The multi-tenant tenant->cluster flow is a SEPARATE deliverable -(tenant-onboarding-v2-DRAFT.md). Some items overlap (image --public, image-by-UUID, template -ownership scope) because both paths hit them. - --------------------------------------------------------------------------------- -## Items --------------------------------------------------------------------------------- - -1. IMAGE SEED MUST create the image `--public` [Step 8.0, image create] - Evidence: a shared/owner-only kube image causes magnum cluster/template create to fail with - `Cluster type (vm, Unset, kubernetes) not supported` -- a non-owner (or the driver acting in - another project) cannot read `os_distro`, so type-derivation returns Unset. Fix: the seed - `openstack image create ... --public` (and re-verify visibility=public post-create). - -2. SEED HARDENING [Step 8.0] - - curl with retry + connect/max timeout (fail loud on partial/hung download). - - sha512 verify against the published manifest is a hard GATE (already present -- keep). - - poll image to `active` as a hard-gate loop (not a fixed sleep). - - POST-active property re-verify: kube_version, os_distro, visibility=public, disk_format. - -3. IMAGE-ABSENT PRESENCE GUARD [Step 8.0] - Explicitly branch "image present -> verify props" vs "absent -> seed", so a re-run does not - double-seed and a present-but-wrong-visibility image is caught (ties to item 1). - -4. IMAGE BY UUID, not name [Step 8.0 template create; 8.1] - Evidence: a doubled-quoted image NAME resolved to the literal `'name'` (no image) -> Unset - type -> 400. Passing the resolved UUID removes the quoting/resolution surface. Gate the UUID - with `grep -qE '^[0-9a-f-]{36}$'` before use. - -5. TEMPLATE CREATE -- OWNER PROJECT SCOPE [Step 8.0] - Evidence: `coe cluster template create/show` and `cluster create --cluster-template ` - resolve the template within the CALLER'S project (templates are visible by ownership). - A private template created in capi-mgmt is NOT selectable by name from admin scope (create - 404s while `template list` still shows it). Fix: run the template create AND the cluster - create in the SAME project scope that owns the template (capi-mgmt for the operator path). - Add the capi-mgmt scope preamble (resolve `capi-mgmt` --domain capi dynamically; export - OS_PROJECT_ID) before both. - -6. FLAVOR-FLOOR PRE-CHECK [Step 8.0 template create] - Magnum requires master/node flavors >= 2 vcpu and >= 2048 MB. Pre-check the chosen flavors - against the floor and fail loud, rather than surfacing an opaque driver error later. - -7. OCTAVIA PREREQ -- CAPTURE REAL EXIT [Prerequisites / Step 8.0] - The octavia-healthy probe must capture the actual command result and test it, NOT - `... | head || echo` (which masks failure -- head succeeds on empty input). Same - capture-and-test-result discipline applied across the onboarding v2 blocks. - -8. 8.1 PRE-CHECKS -- D-039 role + keypair [Step 8.1] - Before cluster create, assert (a) the trustor holds member + load-balancer_member (+ reader) - on the cluster project (D-039 -- else CAPO 403s at the Octavia LB step), and (b) the keypair - exists in the creating scope. Fail loud pre-create. - -9. POLICYD ZIP PATH UNDER $HOME (snap confinement) [appendix-C section C.3] - Evidence: `juju attach-resource ... /tmp/overrides.zip` failed "no such file or directory" - though the shell saw the file -- the confined juju snap cannot read /tmp. Build the zip under - $HOME. Also: `zip` is absent on the jumphost -- build via python3 zipfile (arcname=top-level). - Fix appendix-C C.3 to use a $HOME path and the python3 zipfile method (currently shows - `zip -j /tmp/overrides.zip`). - --------------------------------------------------------------------------------- -## Cross-doc corrections (already staged in this package) --------------------------------------------------------------------------------- -- appendix-C: manager domain-enumeration is own-domain-only on this cloud (2.5d finding); - the cloud-wide names-only leak does NOT manifest. (Applied in appendix-C-identity-rbac.md here.) -- appendix-D: cluster-create trust model; D.7 status updated (Stages 1-4 validated, Stage 6 - create_trust outstanding). Needs committing (was packaged, not yet in repo). - --------------------------------------------------------------------------------- -## Sequencing --------------------------------------------------------------------------------- -Apply items 1-8 to phase-08 and item 9 to appendix-C only AFTER Stage 6 (create_trust) is -resolved -- if the multi-tenant trust step surfaces a further phase-08-relevant fix (e.g. a -CONF.trust.roles pin), fold it into the same DOCFIX-064 sweep rather than reopening. diff --git a/docs/README-D057-PACK.md b/docs/README-D057-PACK.md deleted file mode 100644 index fde799c..0000000 --- a/docs/README-D057-PACK.md +++ /dev/null @@ -1,108 +0,0 @@ -# D-057 provider-vip split -- install pack - -> RENUMBERED per D-058 (2026-06-29): provider-vip=10.12.8.0/22, metal-admin=10.12.12.0/22, -> metal-internal=10.12.16.0/22, data-tenant=10.12.20.0/22, oob=10.12.60.0/22. See -> docs/D-058-renumber.md for the map, the jumphost ordering trap, and the committed-foundation -> cascade still to sweep. This pack carries the renumbered scheme and is re-validated. - -Latest versions of every file produced for the D-057 remediation (move the public -API VIP plane onto a tagged routed VLAN so the untagged provider NIC is free for OVS -br-ex, restoring floating-IP reachability). Files are laid out in repo-relative -folders -- drop them into the repo at the paths shown. - -LEGEND: [NEW] new file | [CHG] modified existing file | [DOC] documentation - --------------------------------------------------------------------------------- -## Contents and destination paths --------------------------------------------------------------------------------- - bundle.yaml -> bundle.yaml [CHG] - D-057 delta: 11 charms public->provider-vip; 11 VIP provider legs 10.12.4.X-> - 10.12.8.X (admin .8 / internal .12 legs unchanged); openstack0 MAC trimmed from - ovn-chassis bridge-interface-mappings; header comments updated. Nothing else. - - scripts/provider-vip-standup.sh -> scripts/provider-vip-standup.sh [NEW] - Creates the MAAS provider-vip plane (space + VID 104 on the provider fabric + - subnet 10.12.8.0/22 + gateway + reserved band). Dry-run by default; --apply to - execute. Idempotent. MTU mirrors the PROVIDER parent fabric (not metal-internal). - - scripts/carve-host-interfaces.sh -> scripts/carve-host-interfaces.sh [CHG] - Host interface carve: enp1s0 -> raw + L3-less (OVS br-ex uplink); new - enp1s0.104 -> br-prov-api (standard bridge) -> static 10.12.8.N. Dry-run default; - per-host; idempotent. - - scripts/lib-net.sh -> scripts/lib-net.sh [CHG] - Adds the shared contract: PROVIDER_VIP_CIDR=10.12.8.0/22, PROVIDER_VIP_VID=104. - Consumed by both the carve and the stand-up. - - scripts/d057-bundle-check.py -> scripts/d057-bundle-check.py [NEW] - Focused, fail-closed checker of the D-057 bundle invariants. Run as a pre-deploy - gate: `python3 scripts/d057-bundle-check.py bundle.yaml`. (Interim gate -- see - R2 in docs/D-057-REVIEW-ITEMS.md: review-bundle.py is pre-D-052 and not current.) - - tests/provider-vip-standup/ -> tests/provider-vip-standup/ [NEW] - tests/carve-host-interfaces/ -> tests/carve-host-interfaces/ [CHG] - Behavior tests (fake `maas` + real jq). Run: `bash tests//run-tests.sh`. - Harnesses self-`chmod +x` their fakebin at runtime (GitHub Desktop strips exec - bits). The standup FRESH case also guards the MTU source (asserts mtu comes from - the provider parent, not metal-internal). - - runbooks/provider-vip-maas-standup.md -> runbooks/provider-vip-maas-standup.md [DOC] - Gated manual runbook for the MAAS plane. Phase-1 audit + the virbr1 - vlan_filtering gate are uniquely useful; Phase-2 creates are superseded by the - script (noted in-file; R1). - - runbooks/jumphost-provider-vip-gateway.md-> runbooks/jumphost-provider-vip-gateway.md[DOC] - Gated runbook to set the jumphost L3 gateway (virbr1.104 = 10.12.8.1): - audit -> reversible runtime apply -> systemd-oneshot persistence (recommended) - or netplan. Deliberately a runbook, not a script (one-time, non-portable, - libvirt-persistence risk untestable by fixtures). - - docs/D-057-REVIEW-ITEMS.md -> docs/D-057-REVIEW-ITEMS.md [DOC] - End-of-deployment reconciliation log (R1-R6): runbook redundancy, stale - review-bundle.py, bundle machine-id fidelity, oob CIDR, gateway-default-route - watch-item, octavia chassis bim. - - docs/D-058-renumber.md -> docs/D-058-renumber.md [DOC] - The plane renumber: authoritative map, jumphost ordering trap, NetBox-apex note, - and the committed-foundation cascade list. Read this first if CIDRs look unfamiliar. - - docs/D-057-DECIDED-append.md -> APPEND to docs/design-decisions.md [DOC] - The D-057 decision record. Append its body to docs/design-decisions.md (do not - keep as a standalone file long-term). - --------------------------------------------------------------------------------- -## Dependencies (NOT shipped -- already in repo / environment) --------------------------------------------------------------------------------- - scripts/lib-hosts.sh UNCHANGED repo file. Required at runtime by - carve-host-interfaces.sh AND by the carve test harness. - Ensure it is present; this pack does not modify it. - jq on the jumphost (scripts + harnesses). - PyYAML for d057-bundle-check.py: pip install pyyaml --break-system-packages - --------------------------------------------------------------------------------- -## CRITICAL: these changes are ATOMIC -- land them in the SAME redeploy --------------------------------------------------------------------------------- -The carve frees enp1s0 and moves the container `public` attach to br-prov-api. If the -NEW carve/stand-up deploy against the OLD bundle (public still -> provider-public), -Juju rebuilds the Linux bridge br-enp1s0 and REPRODUCES D-057. Land together: - (1) scripts/lib-net.sh + provider-vip-standup.sh + carve-host-interfaces.sh - (2) bundle.yaml - (3) the host-nginx :81 line on the proxy VM 10.12.4.7 (Horizon VIP 10.12.4.58 -> - 10.12.8.58) -- a proxy-VM config change, not a repo file; do it in the same window. - --------------------------------------------------------------------------------- -## Execution order (rehearsal) --------------------------------------------------------------------------------- - 0. GATE on the jumphost: `cat /sys/class/net/virbr1/bridge/vlan_filtering` MUST be 0. - 1. PULL the committed pack to the jumphost (commit from Windows; jumphost pulls). - 2. GATE `python3 scripts/d057-bundle-check.py bundle.yaml` -> must PASS. - 3. MAAS `bash scripts/provider-vip-standup.sh` (review dry-run) then `--apply`. - then `juju reload-spaces` so Juju sees the provider-vip space. - 4. CARVE per host at MAAS-Ready: `bash scripts/carve-host-interfaces.sh ` - (dry-run) then apply. (Exact invocation per the carve's own usage.) - 5. DEPLOY the bundle (atomic partner of steps 3-4) + the host-nginx :81 change. - 6. GW run runbooks/jumphost-provider-vip-gateway.md (set virbr1.104 = 10.12.8.1). - 7. VALIDATE D-011 (FIP reachability; resume phase-06 Step 6.3). - -Tests can be run any time, offline: `bash tests/provider-vip-standup/run-tests.sh` -and `bash tests/carve-host-interfaces/run-tests.sh` (both expect ALL PASS). diff --git a/docs/archive/D-057-DECIDED-append.md b/docs/archive/D-057-DECIDED-append.md new file mode 100644 index 0000000..258bfcf --- /dev/null +++ b/docs/archive/D-057-DECIDED-append.md @@ -0,0 +1,37 @@ +> SUPERSEDED IN PART BY D-058 (full plane renumber, 2026-06-29): provider-vip is now +> **10.12.8.0/22** (gateway 10.12.8.1), not 10.12.24.0/22. metal-admin->.12, metal-internal->.16, +> data-tenant->.20, oob->.60. The .24/.8/.12 values below are the original D-057 instantiation, +> kept for history; the durable framework (separate tagged routed VIP plane) is unchanged. +> See D-058-renumber.md for the authoritative scheme. + +## D-057: Provider NIC L3-on-OVS-uplink -- split the public API VIP plane onto its own routed VLAN (2026-06-27) + +**Status:** DECIDED. Root cause CONFIRMED live (this session) and direction CONFIRMED by operator. Supersedes the standalone PROPOSED/OPEN note `D-057-provider-uplink-l3-separation.md` and resolves its one open question (the D-003B interaction) in favor of option (a)/(c). Implementation pending: MAAS plane add -> `carve-host-interfaces.sh` delta -> bundle delta -> teardown/redeploy -> D-011 re-validate. grep-before-assign confirmed D-057 free (max prior D-053; D-054/055/056 are DOCFIX). + +**Symptom:** every provider-ext floating IP (pool 10.12.5.0-10.12.7.254) unreachable cloud-wide; the .4.x public API VIPs answered. Blocked phase-06 Step 6.3 (SSH to capi-mgmt-v2 via FIP 10.12.7.107). + +**Root cause (CONFIRMED, not inferred -- measured on all three ovn-chassis hosts this session):** a two-consumer collision on the untagged provider NIC `enp1s0`. (1) ovn-chassis is configured `bridge-interface-mappings: br-ex:` and wants `enp1s0` as an OVS `br-ex` port. (2) ~11 `public: provider-public` API charms are LXD containers; Juju attaches a container to a subnet by bridging the host NIC, which created the Linux bridge `br-enp1s0` and captured `enp1s0`. A single untagged physical NIC cannot be both a Linux-bridge member and an OVS `br-ex` port: the Linux bridge won, `br-ex` was starved (no-carrier on all three), FIPs dark while the containerized .4.x VIPs kept answering. Evidence: `br-ex` operstate=down / carrier=Invalid-argument on openstack1/2/3; `enp1s0 master br-enp1s0` on all three; `br-enp1s0` ports = `enp1s0` + container `veth*` taps (1/3/5 taps); host static (.41/.42/.43) riding `br-enp1s0`; no public VIP on any host (all in containers); host default route via metal-admin `10.12.8.1`, provider `10.12.4.1` only on-link. The host has NO provider-routing role -- only its containers do. + +**Why option (b) was rejected:** "remove the host provider static" addresses consumer (1)'s L3 but NOT consumer (2). `br-enp1s0` exists for container attach; removing the host static cannot release `enp1s0`. Refuted by the measured `veth` taps. + +**Decision (a)/(c):** relocate the container `public` attach off the untagged uplink onto a tagged sub-interface, freeing untagged `enp1s0` for OVS `br-ex`. This is the Canonical shared-NIC pattern and is the exact mirror of the existing metal-internal tagged-secondary stack (`br-metal.103 -> br-internal`). Because a single CIDR cannot span two L2s, the public API VIPs re-IP onto a new subnet on the tagged VLAN -- which removes D-003B's same-L2 property and replaces it with L3 routing (see amendment below). + +**Target interface tree (per host; octet N = .40-.43 by HOST_OCTET):** +- `enp1s0` -- RAW, untagged on fabric 1_provider, **NO L3 / NO subnet link / NO bridge**. ovn-chassis MAC-enslaves it into OVS `br-ex`. (The old host static 10.12.4.N is removed FROM THIS UPLINK -- that removal is the fix.) +- `enp1s0.104` (VLAN, VID 104, on fabric 1_provider) -> `br-prov-api` (standard bridge) -> **STATIC 10.12.24.N**. Carries the new `provider-vip` space. Container `public` endpoints bind here. The host's provider-plane presence MOVES from `enp1s0` to `br-prov-api` (tagged, OVS-free, no `br-ex` competition -> zero D-057-class risk), mirroring metal-internal's `br-internal` host static -- the proven container-attach pattern on this cloud. (L3-less `br-prov-api` considered and rejected: unproven for Juju attach here; not worth gambling the redeploy.) + +**New plane (framework, not final subnetting):** `provider-vip` -- 10.12.24.0/22, VID 104, fabric 1_provider, **routed** (gateway 10.12.24.1), VIP reserve 10.12.24.2-.100 (VIPs .50-.60). The subnet/VID values are this deployment's instantiation; the DURABLE decision is the framework: the public API VIP plane is its own tagged, routed plane distinct from the provider ext_net/FIP plane. The DC-DC byte-aligned plan adopts the same split (its provider VLAN 240 -> 240 ext_net + 241 provider-vip; its already-present NN-11 role is the home). + +**provider-public after the split:** keeps only FIP/ext_net (10.12.5.0-10.12.7.254) + the OVN gateway SNAT + mgmt reserve; untagged enp1s0 -> OVS br-ex; NO container attach, NO host L3. Neutron stays flat physnet1 (no retype) -- the OVN/Neutron layer, confirmed correct this session, is untouched. + +**D-003B AMENDMENT:** D-003B deliberately co-located public API VIPs and FIPs on one provider L2 ("tenant->API by construction") -- a Bobcat-era convenience that this session proved unworkable on the NIC-limited host (the OVS-vs-container collision). It is hereby amended: the public API VIP plane is a separate, routed VLAN. tenant->API is preserved via L3 routing (tenant SNAT egress on provider 10.12.4.0/22 -> gateway -> provider-vip 10.12.24.0/22), and re-validated per D-011 #3. This is also better for the commercial hard-isolation goal (API VIPs out of the FIP broadcast domain; an L3 policy enforcement point) and matches the Roosevelt NIC-limited reality. + +**ROUTED-GATEWAY PREREQUISITE (gating, highest risk):** today only provider-public and metal-admin route (as-built: all other planes `gateway: none`). provider-vip must be routed: gateway 10.12.24.1 must exist and route to/from 10.12.4.0/22 on whatever owns 10.12.4.1 (jumphost/libvirt), with the return path 10.12.24.1 -> 10.12.4.0/22. Without it the redeploy passes D-011 #1/#2 and silently fails #3. Confirm/establish before redeploy -- do not discover at validation. + +**bridge-interface-mappings:** trim openstack0's provider MAC `52:54:00:3d:fd:54` (measured: the only one of the four with no ovn-chassis unit; nova-compute/ovn-chassis run on machines 9/10/11 = openstack1/2/3). Keep `9d:63:77`/`89:7f:ce`/`99:fc:c2`. Settles openstack0 as control+storage role (feeds Roosevelt node-role split). The mapping still targets the untagged `enp1s0` MAC (now free). + +**bundle changes (summary):** on the 11 `public:`-bound API charms, rebind `public: provider-public -> public: provider-vip`, and re-IP each `vip:` first token `10.12.4.X -> 10.12.24.X` (X=50-60; metal-admin .8.X and metal-internal .12.X tokens unchanged). Trim the openstack0 MAC. No Neutron/OVN config change. + +**Roosevelt relevance:** the provider-vip tagged-routed-plane split joins metal-internal VID-103 as a per-host interface-tree region-invariant (maas-as-built reference). The DC-DC byte-aligned plan inherits the split. + +**Related:** root-causes phase-06 6.3; amends D-003B; extends D-052/D-053 plane model with the host-interface-layer rule the rebuild exposed; the committed `docs/maas-as-built-reference.md` carve table needs the follow-on correction (enp1s0 raw/L3-less + the enp1s0.104->br-prov-api stack). diff --git a/docs/archive/D-057-REVIEW-ITEMS.md b/docs/archive/D-057-REVIEW-ITEMS.md new file mode 100644 index 0000000..51ec889 --- /dev/null +++ b/docs/archive/D-057-REVIEW-ITEMS.md @@ -0,0 +1,58 @@ +# D-057 -- review items for END OF DEPLOYMENT (do not action mid-deploy) + +These are real findings surfaced while building the D-057 provider-vip split. None +block the deploy; each is deferred to a post-D-011 reconciliation sweep so we don't +re-architect inside a step. Logged here so they are not lost. + +-------------------------------------------------------------------------------- +R1. provider-vip-maas-standup.md CREATE blocks are now redundant with the script. + scripts/provider-vip-standup.sh is the tested execution path for creating the + plane. The runbook's Phase-2 create blocks duplicate that logic (two sources of + truth -- this is how the MTU-source bug nearly drifted between them). KEEP the + runbook for its still-unique parts: Phase-1 audit, the virbr1 vlan_filtering gate, + and the deferred jumphost-gateway reads. RECONCILE: trim the create blocks to a + pointer at the script, or annotate them "superseded -- see script". (User: noted + for end-of-deployment review, not trimmed now.) + +-------------------------------------------------------------------------------- +R2. scripts/review-bundle.py is STALE (pre-D-052) and is NOT a current gate. + Run against today's committed bundle it reports ~30 FAILs: its PHANTOM_BINDING_KEYS + check forbids exactly the per-endpoint bindings (shared-db/amqp/certificates/cluster/ + ha/internal) that D-052 deliberately added, and its VIP check expects DUAL VIPs at + octets .224-235 in the old 10.12.8 "metal" net -- the bundle has TRIPLE VIPs at + .50-.60 across provider-public/metal-admin/metal-internal. So review-bundle.py + predates D-052 entirely and does not describe the live bundle. + - The "verify_bundle.py / 8/8 harness" referenced in the redeploy notes is NOT at + this repo HEAD (d575a25) -- only review-bundle.py is. + - INTERIM GATE for the D-057 change: scripts/d057-bundle-check.py (focused, fail- + closed, proven FAIL-on-pre / PASS-on-post). It checks only the D-057 invariants. + - RECONCILE: either bring review-bundle.py forward to D-052/053/D-057 (rewrite + PHANTOM check, VIP check to triple + provider-vip 10.12.8, octets 50-60), or + restore/commit the newer verify_bundle.py and retire review-bundle.py. + +-------------------------------------------------------------------------------- +R3. Committed bundle uses machines 8/9/10/11; the live cloud ran 0/1/2/3. + Repo-fidelity gap (the committed bundle is not the as-deployed bundle). Does NOT + affect D-057 (the carve is hostname-based; the MAC trim is by MAC). RECONCILE the + bundle machine block + `to:` placements to whatever the redeploy actually uses. + +-------------------------------------------------------------------------------- +R4. oob CIDR -- RESOLVED by D-058. oob adopts 10.12.60.0/22 (the design-docs value). + The live virsh power gateway 10.12.64.1 -> 10.12.60.1 (scripts/lib-hosts.sh + VIRSH_POWER_ADDRESS) is part of the D-058 foundation cascade -- see D-058-renumber.md. + +-------------------------------------------------------------------------------- +R5. gateway_ip default-route watch-item (post-deploy, not end-of-deploy). + provider-vip subnet carries gateway_ip=10.12.8.1 (post-D-058). This mirrors provider-public's + existing .4.1 (which already coexists with metal-admin .8.1 without hijacking the + node default), so risk is low -- but VERIFY after redeploy that every API unit's + default route is still via metal-admin 10.12.12.1, not .8.1: + juju exec --all -- ip route show default + If any unit defaults via .8.1: blank the subnet gateway_ip (set PVIP_GATEWAY="" in + scripts/provider-vip-standup.sh and re-apply) or pin the node default-gateway subnet + to metal-admin. provider-vip reachability does not depend on its own gateway in v1. + +-------------------------------------------------------------------------------- +R6. ovn-chassis-octavia has no bridge-interface-mappings (expected -- it is the + octavia-side chassis). Left unchanged. Noted only so it is not mistaken for an + omission during review. diff --git a/docs/archive/D-058-renumber.md b/docs/archive/D-058-renumber.md new file mode 100644 index 0000000..590f13a --- /dev/null +++ b/docs/archive/D-058-renumber.md @@ -0,0 +1,77 @@ +# D-058: full plane renumber -- clean fabric-grouped /22 scheme (2026-06-29) + +**Status:** DECIDED (operator). Supersedes the D-057 minimal-delta placement of +provider-vip at 10.12.24.0/22, and resolves R4 (oob). grep-before-assign: D-058 free +(max prior D-057; D-054/055/056 are DOCFIX). + +**What:** renumber the v1 plane scheme so CIDRs are contiguous /22 blocks grouped by +fabric and ordered to match the layer model, instead of the historical scatter. This +is a cloud-wide re-IP, intentionally larger than D-057, accepted by the operator for +Roosevelt addressing fidelity. It is executed as a teardown/redeploy (no in-place +re-CIDR), so there is no transient subnet overlap. + +## The map (authoritative) + +| Plane | old | NEW | gateway (was -> now) | +|-----------------|---------------|---------------|-----------------------------| +| provider-public | 10.12.4.0/22 | 10.12.4.0/22 | 10.12.4.1 (unchanged) | +| provider-vip | 10.12.24.0/22 | 10.12.8.0/22 | 10.12.24.1 -> 10.12.8.1 | +| metal-admin | 10.12.8.0/22 | 10.12.12.0/22 | 10.12.8.1 -> 10.12.12.1 | +| metal-internal | 10.12.12.0/22 | 10.12.16.0/22 | none (L2 east-west) | +| data-tenant | 10.12.16.0/22 | 10.12.20.0/22 | none (isolated L2) | +| storage | 10.12.32.0/22 | 10.12.32.0/22 | none (unchanged) | +| replication | 10.12.36.0/22 | 10.12.36.0/22 | none (unchanged) | +| oob | 10.12.64.0/22 | 10.12.60.0/22 | 10.12.64.1 -> 10.12.60.1 | + +Rotate rule (collision-safe): 8->12, 12->16, 16->20, 24->8, 64->60; 4/32/36 fixed. +VLAN IDs unchanged (metal-internal VID 103, provider-vip VID 104). VIP triple becomes +provider-vip .8.5x / metal-admin .12.5x / metal-internal .16.5x (octets 50-60). +Host statics (.40-.43) follow each plane. metal-admin PXE-DHCP band -> 10.12.12.9-.11. + +## JUMPHOST ORDERING TRAP (must respect on the host) + +The jumphost owns three gateways that move. provider-vip's NEW gateway 10.12.8.1 +is metal-admin's OLD address. So on the jumphost, in this order: + 1. move virbr2 (metal-admin) 10.12.8.1 -> 10.12.12.1 + 2. move virbr7 (oob) 10.12.64.1 -> 10.12.60.1 + 3. THEN add virbr1.104 (provider-vip) = 10.12.8.1 <-- only after step 1 frees .8.1 +Adding virbr1.104=.8.1 while virbr2 still holds .8.1 is a same-subnet collision. In a +clean rebuild the bridges are reconfigured as a set, but the free-then-claim order +still applies. (Step 3 is the jumphost-provider-vip-gateway.md runbook.) + +## APEX / NetBox note (IaC discipline) + +NetBox is the apex; this renumber's authoritative home is NetBox. BUT the committed +netbox/ipv4-prefixes-import.py is itself stale (pre-D-052: only Metal/Provider/LBaaS-mgmt, +provider VLAN VID 240, API VIPs at .224-.254 -- none of the 6-plane D-052/053 model). It +must FIRST be brought current to D-052/053, THEN carry the D-058 scheme, before it can be +the source of truth. Until that reconciliation, scripts/lib-net.sh is the working contract +and already carries D-058. Do NOT hand-edit downstream MAAS for these values once NetBox +is current -- regenerate. + +## DONE in this pack (renumbered + re-validated: both suites ALL PASS, d057-check PASS) + +scripts/lib-net.sh (PLANE_CIDRS, PLANE_NAME, PLANE_GW, DATA_PLANE_CIDRS, +METAL_INTERNAL_CIDR, PROVIDER_VIP_CIDR, VIP_PREFIX_* triple), +scripts/carve-host-interfaces.sh, scripts/provider-vip-standup.sh, +scripts/d057-bundle-check.py, bundle.yaml (11 VIP triples), both test suites + +fixtures, provider-vip-maas-standup.md, jumphost-provider-vip-gateway.md, README. + +## COMMITTED-FOUNDATION CASCADE (still on the OLD scheme -- next sweep) + +Apply the same rotate (8->12, 12->16, 16->20, 24->8, 64->60; 4/32/36 fixed). These are +in the committed repo, not this pack, and several are prose runbooks -- sweep with care, +NetBox-anchored: + - netbox/ipv4-prefixes-import.py (APEX -- de-stale to D-052/053 first, then D-058) + - netbox/README.md + - scripts/phase-00-maas-carve.sh (METAL_CIDR default 10.12.8 -> 10.12.12; ranges) + - scripts/lib-hosts.sh (VIRSH_POWER_ADDRESS 10.12.64.1 -> 10.12.60.1) + - scripts/review-bundle.py (stale pre-D-052 already -- R2; fold in with that) + - runbooks/phase-00-teardown-maas-reset.md, phase-01-bundle-deploy.md, + phase-03-core-verify.md, phase-04-network-carve.md, phase-05-octavia-enablement.md, + phase-08-workload-cluster-acceptance.md, appendix-A-troubleshooting.md + - docs/maas-as-built-reference.md, docs/design-decisions.md (append D-058), + docs/v1-redeploy-changelog.md, docs/netbox-vip-queue.md + - tests/phase-00-carve/run-tests.sh, tests/phase-04/make_fixtures.py + - jumphost underlay: virbr2 -> 10.12.12.1, virbr7 -> 10.12.60.1 (see ordering trap) + - host-nginx :81 upstream: Horizon -> 10.12.8.58 diff --git a/docs/archive/D-068-openbao-assessment-DRAFT.md b/docs/archive/D-068-openbao-assessment-DRAFT.md new file mode 100644 index 0000000..4a6ec5d --- /dev/null +++ b/docs/archive/D-068-openbao-assessment-DRAFT.md @@ -0,0 +1,108 @@ +# D-068 item 1 -- OpenBao / off-EOL-vault assessment (RESEARCH DRAFT) + +Status: RESEARCH INPUT ONLY (2026-07-06, jumphost stream, web research). The D-068 +off-EOL path decision remains OPEN and operator-ruled; this draft exists so the B2 +discussion starts from verified facts. Items that could not be confirmed are marked +UNVERIFIED. Companion: docs/D-068-vault-1.8-vs-1.16-analysis.md (the 1.16 ruling). + +## Summary + +- OpenBao is healthy upstream (Linux Foundation, MPL-2.0; latest stable 2.5.5, + 2026-06-17) and API-compatible for our surfaces (kv v1/v2, AppRole, PKI). BUT: + **no Juju charm for OpenBao exists anywhere** (Charmhub API search empty, + 2026-07-06), and **OpenBao REMOVED the MySQL storage backend** (raft/postgres + only; upstream issue #651). An OpenBao path today = bespoke charm development + we would own + a storage migration -- it cannot serve our 18 V0 `certificates` + relations. +- **No plan exists to bring tls-certificates V1 to the legacy/reactive OpenStack + charms.** Canonical's V1 world is Sunbeam ("Canonical OpenStack", 2024.1 LTS, + up-to-12-year support): vault + manual-tls-certificates + traefik on V1. The + charm-guide still documents legacy TLS via the V0 vault relation. Practical + read: V1 will likely NEVER come to the reactive set; "wait for V1" is risk + acceptance without a deadline. (Absence-of-evidence finding -- flagged as such.) +- The operator-lineage vault charm moved on (tracks to 1.19 + a new 2.0 track, + 2026-07-02, tracking IBM Vault 2.0 GA 2026-04-14) -- same disqualifiers as the + ruled-out 1.16 (V1-only, Raft-only, BUSL). The reactive 1.8/stable charm is in + steady-state maintenance and still being rebuilt (Charmhub revision 2026-07-03). +- CVE posture of 1.8.8: no known unauthenticated remote-compromise CVE applies to + an internal-only, AppRole-driven, TLS-issuing deployment. Notable but gated: + CVE-2021-45042 (PKI wildcard issuance to AUTHORIZED users, medium, in-range); + CVE-2024-2048 (cert-auth bypass, 9.8 -- only if the cert auth method is enabled, + ours is not); Aug-2025 "Vault Fault" set incl. CVE-2025-6000 RCE (requires + privileged plugin-catalog access; per-CVE 1.8.8 applicability UNVERIFIED). + Honest characterization: **manageable today, unauditable tomorrow** -- nobody + assesses 1.8.8 anymore; each future disclosure needs our own analysis with no + patch option. Supports deadline-bound acceptance, not indefinite stay. + +## Candidate paths compared + +1. **Wait-for-V1 in legacy charms** -- no published plan; strong signal V1 + investment goes to Sunbeam only. Equivalent to acceptance WITHOUT a deadline. + Reject as a plan; it is the null hypothesis. +2. **OpenBao now** -- right long-term horse upstream, but requires (a) a charm + that does not exist (bespoke, we own it), (b) storage migration off MySQL, + (c) an aggressive support cadence (only latest cycle maintained). Not viable + as a like-for-like swap; keep on the WATCH list (an OpenBao charm appearing, + or any Canonical position on OpenBao, changes this materially). +3. **Risk-acceptance with a hard deadline (research recommendation)** -- stay on + 1.8/stable (Canonical still rebuilds it; commercially supported under + Pro/PCB). Compensating controls: mgmt-plane-only API exposure, audit logging + shipped, minimal policies, confirm userpass/cert auth methods and plugin + registration absent, unseal/root discipline (already the norm -- D-069). + Hard re-decision deadline proposed: 2027-04. Two follow-ups; (a) is now + MEASURED (read-only, 2026-07-06, evidence + ~/openstack-baseline/vault-config-20260706.yaml + vault-snap-tracks-20260706.txt): + the reactive charm EXPOSES a `channel` config (currently 1.8/stable; its own + description warns a change seals ALL vault units via the snap refresh -- + manual 3-of-5 unseal is our norm, so that is a planned-window event, not a + blocker). The Canonical vault snap on the live unit publishes tracks + 1.1-1.12, then 1.15-1.20 and 2.0: **there is NO 1.13 or 1.14 track** -- the + last MPL-licensed payload reachable via the charm is **1.12.11** + (1.12/stable), not 1.14.x. So the in-charm payload-bump option is: + 1.8.8 -> 1.12.11 (stays MPL; still EOL upstream, but 4 minors newer), or + 1.15.6+ (BUSL -- licensing review required). Remaining UNVERIFIED before any + such move: charm-vs-newer-payload API compatibility (does the reactive charm + drive 1.12 correctly -- Canonical test coverage unknown), storage-format + stepping across 4 minors on vault-on-mysql, and whether MySQL backend + behavior changed in 1.9-1.12. This is a rehearse-first D-NNN of its own if + pursued; NOT a routine refresh. (b) track Sunbeam feature parity -- the + durable exit from the V0 certificates world is a PLATFORM decision, not a + vault swap. Also verified live 2026-07-06: the charm's Vault-listener TLS + options exist and are empty (ssl-ca / ssl-cert / ssl-chain / ssl-key) -- + the concrete D-068 item 2 mechanism, no fabricated option names. + +## Decision hooks for the operator (nothing ruled yet) + +- Adopt path 3 with deadline? (Composes with keeping path 2 on watch.) +- Authorize the read-only snap-channel headroom measurement (follow-up a)? +- Fold "Sunbeam parity watch" into the Roosevelt planning inputs? +- D-068 items 2 (listener TLS) and 3 (AppRole TTL audit) are unaffected by this + and proceed on their own track. + +## Sources (accessed 2026-07-05/06; full trail) + +- openbao.org: release-notes 2.6.0-beta (2026-06-22); blog/roadmap-2.0 (LF + governance); docs/configuration/storage; docs/guides/migration +- endoflife.date/openbao (2.5.5 latest stable, 2026-06-17; latest-cycle-only support) +- github.com/openbao/openbao issue #651 (MySQL backend removed/requested) +- api.charmhub.io/v2/charms/find?q=openbao (EMPTY, 2026-07-06); charmhub.io/openbao (404) +- charmhub.io/vault (tracks 1.5-1.19 + 2.0 candidate/beta/edge; 1.8/stable rev 2026-07-03) +- canonical-vault-charms.readthedocs-hosted.com (2.0 upgrade docs; 1.8-vs-1.15 differences) +- opendev.org/openstack/charm-vault commits (steady-state; last commit 2025-10-29) +- docs.openstack.org/charm-guide latest: admin/security/tls (V0 model current) +- canonical-openstack.readthedocs-hosted.com: vault feature, service-endpoint-encryption, + release-cycle (Sunbeam 2024.1 LTS, 12-year support) +- infoq.com 2026/04 vault-2-0-ibm-identity + IBM support-lifecycle addendum + (Vault 2.0 GA 2026-04-14; 1.21->2.0 renumbering) +- nvd.nist.gov CVE-2024-2048 (9.8 NIST / 8.1 vendor; <=1.14.9; cert auth only) +- cyata.ai vault-fault (Aug-2025 nine zero-days incl. CVE-2025-6000 RCE) +- stack.watch + cvedetails (historic ranges; NOTE conflicting aggregator text for + CVE-2021-45042 / CVE-2021-43998 -- re-check NVD before quoting in a decision record) +- stackhpc-kayobe-config readthedocs stackhpc-2025.1 configuration/openbao + (non-Juju OpenStack distro already uses OpenBao for internal PKI) +- duokey.com BUSL-change/OpenBao-migration backgrounder (2023-08-10 BUSL date) + +NOT confirmed despite searching: any Canonical statement on OpenBao; any V1 plan +for legacy charms; exact Vault 1.8.x EOL date/final point release; per-CVE +applicability of the 2025 Cyata set to 1.8.8; reactive-charm snap-channel +headroom (needs on-host measurement, not web research). diff --git a/docs/archive/DOCFIX-064-phase08-changelist.md b/docs/archive/DOCFIX-064-phase08-changelist.md new file mode 100644 index 0000000..d7224b7 --- /dev/null +++ b/docs/archive/DOCFIX-064-phase08-changelist.md @@ -0,0 +1,81 @@ +# DOCFIX-064 -- phase-08 runbook change-list (DRAFT, 2026-07-01) + +RESERVED number: DOCFIX-064 (per changelog next-free note). This is the accumulated +phase-08 operator-runbook (single-consumer acceptance) sweep. Written as a change-LIST with +exact anchors + evidence so the edit is mechanical when phase-08 is finalized. NOT yet applied +to runbooks/phase-08-workload-cluster-acceptance.md. + +Scope note: these are fixes to the OPERATOR single-consumer acceptance path (capi-test-1 in +capi-mgmt scope). The multi-tenant tenant->cluster flow is a SEPARATE deliverable +(tenant-onboarding-v2-DRAFT.md). Some items overlap (image --public, image-by-UUID, template +ownership scope) because both paths hit them. + +-------------------------------------------------------------------------------- +## Items +-------------------------------------------------------------------------------- + +1. IMAGE SEED MUST create the image `--public` [Step 8.0, image create] + Evidence: a shared/owner-only kube image causes magnum cluster/template create to fail with + `Cluster type (vm, Unset, kubernetes) not supported` -- a non-owner (or the driver acting in + another project) cannot read `os_distro`, so type-derivation returns Unset. Fix: the seed + `openstack image create ... --public` (and re-verify visibility=public post-create). + +2. SEED HARDENING [Step 8.0] + - curl with retry + connect/max timeout (fail loud on partial/hung download). + - sha512 verify against the published manifest is a hard GATE (already present -- keep). + - poll image to `active` as a hard-gate loop (not a fixed sleep). + - POST-active property re-verify: kube_version, os_distro, visibility=public, disk_format. + +3. IMAGE-ABSENT PRESENCE GUARD [Step 8.0] + Explicitly branch "image present -> verify props" vs "absent -> seed", so a re-run does not + double-seed and a present-but-wrong-visibility image is caught (ties to item 1). + +4. IMAGE BY UUID, not name [Step 8.0 template create; 8.1] + Evidence: a doubled-quoted image NAME resolved to the literal `'name'` (no image) -> Unset + type -> 400. Passing the resolved UUID removes the quoting/resolution surface. Gate the UUID + with `grep -qE '^[0-9a-f-]{36}$'` before use. + +5. TEMPLATE CREATE -- OWNER PROJECT SCOPE [Step 8.0] + Evidence: `coe cluster template create/show` and `cluster create --cluster-template ` + resolve the template within the CALLER'S project (templates are visible by ownership). + A private template created in capi-mgmt is NOT selectable by name from admin scope (create + 404s while `template list` still shows it). Fix: run the template create AND the cluster + create in the SAME project scope that owns the template (capi-mgmt for the operator path). + Add the capi-mgmt scope preamble (resolve `capi-mgmt` --domain capi dynamically; export + OS_PROJECT_ID) before both. + +6. FLAVOR-FLOOR PRE-CHECK [Step 8.0 template create] + Magnum requires master/node flavors >= 2 vcpu and >= 2048 MB. Pre-check the chosen flavors + against the floor and fail loud, rather than surfacing an opaque driver error later. + +7. OCTAVIA PREREQ -- CAPTURE REAL EXIT [Prerequisites / Step 8.0] + The octavia-healthy probe must capture the actual command result and test it, NOT + `... | head || echo` (which masks failure -- head succeeds on empty input). Same + capture-and-test-result discipline applied across the onboarding v2 blocks. + +8. 8.1 PRE-CHECKS -- D-039 role + keypair [Step 8.1] + Before cluster create, assert (a) the trustor holds member + load-balancer_member (+ reader) + on the cluster project (D-039 -- else CAPO 403s at the Octavia LB step), and (b) the keypair + exists in the creating scope. Fail loud pre-create. + +9. POLICYD ZIP PATH UNDER $HOME (snap confinement) [appendix-C section C.3] + Evidence: `juju attach-resource ... /tmp/overrides.zip` failed "no such file or directory" + though the shell saw the file -- the confined juju snap cannot read /tmp. Build the zip under + $HOME. Also: `zip` is absent on the jumphost -- build via python3 zipfile (arcname=top-level). + Fix appendix-C C.3 to use a $HOME path and the python3 zipfile method (currently shows + `zip -j /tmp/overrides.zip`). + +-------------------------------------------------------------------------------- +## Cross-doc corrections (already staged in this package) +-------------------------------------------------------------------------------- +- appendix-C: manager domain-enumeration is own-domain-only on this cloud (2.5d finding); + the cloud-wide names-only leak does NOT manifest. (Applied in appendix-C-identity-rbac.md here.) +- appendix-D: cluster-create trust model; D.7 status updated (Stages 1-4 validated, Stage 6 + create_trust outstanding). Needs committing (was packaged, not yet in repo). + +-------------------------------------------------------------------------------- +## Sequencing +-------------------------------------------------------------------------------- +Apply items 1-8 to phase-08 and item 9 to appendix-C only AFTER Stage 6 (create_trust) is +resolved -- if the multi-tenant trust step surfaces a further phase-08-relevant fix (e.g. a +CONF.trust.roles pin), fold it into the same DOCFIX-064 sweep rather than reopening. diff --git a/docs/archive/README-D057-PACK.md b/docs/archive/README-D057-PACK.md new file mode 100644 index 0000000..fde799c --- /dev/null +++ b/docs/archive/README-D057-PACK.md @@ -0,0 +1,108 @@ +# D-057 provider-vip split -- install pack + +> RENUMBERED per D-058 (2026-06-29): provider-vip=10.12.8.0/22, metal-admin=10.12.12.0/22, +> metal-internal=10.12.16.0/22, data-tenant=10.12.20.0/22, oob=10.12.60.0/22. See +> docs/D-058-renumber.md for the map, the jumphost ordering trap, and the committed-foundation +> cascade still to sweep. This pack carries the renumbered scheme and is re-validated. + +Latest versions of every file produced for the D-057 remediation (move the public +API VIP plane onto a tagged routed VLAN so the untagged provider NIC is free for OVS +br-ex, restoring floating-IP reachability). Files are laid out in repo-relative +folders -- drop them into the repo at the paths shown. + +LEGEND: [NEW] new file | [CHG] modified existing file | [DOC] documentation + +-------------------------------------------------------------------------------- +## Contents and destination paths +-------------------------------------------------------------------------------- + bundle.yaml -> bundle.yaml [CHG] + D-057 delta: 11 charms public->provider-vip; 11 VIP provider legs 10.12.4.X-> + 10.12.8.X (admin .8 / internal .12 legs unchanged); openstack0 MAC trimmed from + ovn-chassis bridge-interface-mappings; header comments updated. Nothing else. + + scripts/provider-vip-standup.sh -> scripts/provider-vip-standup.sh [NEW] + Creates the MAAS provider-vip plane (space + VID 104 on the provider fabric + + subnet 10.12.8.0/22 + gateway + reserved band). Dry-run by default; --apply to + execute. Idempotent. MTU mirrors the PROVIDER parent fabric (not metal-internal). + + scripts/carve-host-interfaces.sh -> scripts/carve-host-interfaces.sh [CHG] + Host interface carve: enp1s0 -> raw + L3-less (OVS br-ex uplink); new + enp1s0.104 -> br-prov-api (standard bridge) -> static 10.12.8.N. Dry-run default; + per-host; idempotent. + + scripts/lib-net.sh -> scripts/lib-net.sh [CHG] + Adds the shared contract: PROVIDER_VIP_CIDR=10.12.8.0/22, PROVIDER_VIP_VID=104. + Consumed by both the carve and the stand-up. + + scripts/d057-bundle-check.py -> scripts/d057-bundle-check.py [NEW] + Focused, fail-closed checker of the D-057 bundle invariants. Run as a pre-deploy + gate: `python3 scripts/d057-bundle-check.py bundle.yaml`. (Interim gate -- see + R2 in docs/D-057-REVIEW-ITEMS.md: review-bundle.py is pre-D-052 and not current.) + + tests/provider-vip-standup/ -> tests/provider-vip-standup/ [NEW] + tests/carve-host-interfaces/ -> tests/carve-host-interfaces/ [CHG] + Behavior tests (fake `maas` + real jq). Run: `bash tests//run-tests.sh`. + Harnesses self-`chmod +x` their fakebin at runtime (GitHub Desktop strips exec + bits). The standup FRESH case also guards the MTU source (asserts mtu comes from + the provider parent, not metal-internal). + + runbooks/provider-vip-maas-standup.md -> runbooks/provider-vip-maas-standup.md [DOC] + Gated manual runbook for the MAAS plane. Phase-1 audit + the virbr1 + vlan_filtering gate are uniquely useful; Phase-2 creates are superseded by the + script (noted in-file; R1). + + runbooks/jumphost-provider-vip-gateway.md-> runbooks/jumphost-provider-vip-gateway.md[DOC] + Gated runbook to set the jumphost L3 gateway (virbr1.104 = 10.12.8.1): + audit -> reversible runtime apply -> systemd-oneshot persistence (recommended) + or netplan. Deliberately a runbook, not a script (one-time, non-portable, + libvirt-persistence risk untestable by fixtures). + + docs/D-057-REVIEW-ITEMS.md -> docs/D-057-REVIEW-ITEMS.md [DOC] + End-of-deployment reconciliation log (R1-R6): runbook redundancy, stale + review-bundle.py, bundle machine-id fidelity, oob CIDR, gateway-default-route + watch-item, octavia chassis bim. + + docs/D-058-renumber.md -> docs/D-058-renumber.md [DOC] + The plane renumber: authoritative map, jumphost ordering trap, NetBox-apex note, + and the committed-foundation cascade list. Read this first if CIDRs look unfamiliar. + + docs/D-057-DECIDED-append.md -> APPEND to docs/design-decisions.md [DOC] + The D-057 decision record. Append its body to docs/design-decisions.md (do not + keep as a standalone file long-term). + +-------------------------------------------------------------------------------- +## Dependencies (NOT shipped -- already in repo / environment) +-------------------------------------------------------------------------------- + scripts/lib-hosts.sh UNCHANGED repo file. Required at runtime by + carve-host-interfaces.sh AND by the carve test harness. + Ensure it is present; this pack does not modify it. + jq on the jumphost (scripts + harnesses). + PyYAML for d057-bundle-check.py: pip install pyyaml --break-system-packages + +-------------------------------------------------------------------------------- +## CRITICAL: these changes are ATOMIC -- land them in the SAME redeploy +-------------------------------------------------------------------------------- +The carve frees enp1s0 and moves the container `public` attach to br-prov-api. If the +NEW carve/stand-up deploy against the OLD bundle (public still -> provider-public), +Juju rebuilds the Linux bridge br-enp1s0 and REPRODUCES D-057. Land together: + (1) scripts/lib-net.sh + provider-vip-standup.sh + carve-host-interfaces.sh + (2) bundle.yaml + (3) the host-nginx :81 line on the proxy VM 10.12.4.7 (Horizon VIP 10.12.4.58 -> + 10.12.8.58) -- a proxy-VM config change, not a repo file; do it in the same window. + +-------------------------------------------------------------------------------- +## Execution order (rehearsal) +-------------------------------------------------------------------------------- + 0. GATE on the jumphost: `cat /sys/class/net/virbr1/bridge/vlan_filtering` MUST be 0. + 1. PULL the committed pack to the jumphost (commit from Windows; jumphost pulls). + 2. GATE `python3 scripts/d057-bundle-check.py bundle.yaml` -> must PASS. + 3. MAAS `bash scripts/provider-vip-standup.sh` (review dry-run) then `--apply`. + then `juju reload-spaces` so Juju sees the provider-vip space. + 4. CARVE per host at MAAS-Ready: `bash scripts/carve-host-interfaces.sh ` + (dry-run) then apply. (Exact invocation per the carve's own usage.) + 5. DEPLOY the bundle (atomic partner of steps 3-4) + the host-nginx :81 change. + 6. GW run runbooks/jumphost-provider-vip-gateway.md (set virbr1.104 = 10.12.8.1). + 7. VALIDATE D-011 (FIP reachability; resume phase-06 Step 6.3). + +Tests can be run any time, offline: `bash tests/provider-vip-standup/run-tests.sh` +and `bash tests/carve-host-interfaces/run-tests.sh` (both expect ALL PASS). diff --git a/docs/archive/clientdocs-workflow-review-20260708.md b/docs/archive/clientdocs-workflow-review-20260708.md new file mode 100644 index 0000000..b5c1aa2 --- /dev/null +++ b/docs/archive/clientdocs-workflow-review-20260708.md @@ -0,0 +1,190 @@ +# clientdocs workflow review + consolidation proposal (2026-07-08) + +STATUS: PROPOSAL for operator ruling. Read-only analysis; nothing implemented. +Options are presented per finding; the operator rules before any change. No +identifier numbers are consumed here -- they are assigned at implementation. + +## Why this exists + +Operator's stated problem (2026-07-08): the information a client needs is spread +throughout the handover documentation with no clear flow. A client has to switch +between multiple documents to accomplish one task, and some workflows have holes +that cause errors (the flannel network-driver gap being the live example: a +self-created flannel cluster never converges, and devteam lost a night to it). +The goal for the client-facing set: an easy-to-follow, workflow-shaped path that +makes finding and using their data simple, minimizes document-switching, and +carries enough component detail that the client understands what each piece +does. + +DOCFIX-123 (the new jenkins-kubernetes-guide.md) is the first instance of the +target shape -- a single self-contained workflow that cross-links to detail +rather than scattering it. This review proposes applying the same principle +across the rest of the set. + +## Method + +Read-only inventory of all 20 client-facing files under clientdocs/ (the six +top-level guides, the tenant-skill SKILL.md + four references, and the six +starter-kit scripts), mapping every workflow topic to the file(s) that cover it, +plus duplication, cross-file conflicts, and gaps. Evidence is file:line +throughout the source inventory; the summary below carries representative cites. + +## The client journey (what a client actually does, in order) + +1. Pre-onboarding homework (intake-form.md). +2. Receive credentials + first orientation (welcome.md, handover-pack.md). +3. Prove the tenancy works (acceptance-checklist.md). +4. Run day-to-day (self-service-guide.md; tenant-skill references). +5. Connect automation / CI (ci-integration-guide.md). +6. Stand up a Kubernetes + Jenkins workflow (jenkins-kubernetes-guide.md, new). + +The docs exist for every stage, but stages 2-6 repeat each other's content and +none of them states "you are here, read these in this order." + +## Findings + +### F1 -- Heavy duplication (maintenance-drift + reader confusion) + +The same facts are authored independently in many files. Representative: + +- The three-account model table: 4 renderings (handover-pack.md:28-32, + self-service-guide.md:10-14, tenant-skill/SKILL.md:32-36, and prose in + welcome.md). +- The clouds.yaml auth block: 3 byte-identical copies (ci-integration-guide.md, + tenant-skill/SKILL.md, scripts/clouds.yaml.template). +- "Cluster create cannot use an application credential": stated in 8 places. +- The worked CI sequence: duplicated at length in ci-integration-guide.md and + references/ci-automation.md. +- Floating-IPs-count-when-detached: repeated in 5 files. +- (Full list: 15 duplication clusters in the source inventory.) + +Risk: any change (an endpoint, an account name, a rule) must be made in N places +or the copies drift. It also makes each document longer than it needs to be, +which is itself part of the "hard to follow" problem. + +### F2 -- No single entry point or reading order + +No file tells the client which document to open first, or the order for a given +goal. A client wanting "deploy my app to Kubernetes from Jenkins" previously had +to assemble it from kubernetes.md + ci-integration-guide.md + day2-operations.md ++ troubleshooting.md. (DOCFIX-123 now owns that one journey; the other journeys +still lack a spine.) + +### F3 -- Workflow holes that cause errors + +- Network-driver choice: the self-service path (self-service-guide.md:89-99) + never mentioned calico, so a client building a template via the dashboard got + no steer away from flannel. CLOSED 2026-07-08 in kubernetes.md + the Jenkins + guide; self-service-guide.md still only carries the Public/Hidden rule, not + the driver rule. +- Right-sizing to host capacity: every doc presents quota as the only ceiling on + node_count (self-service-guide.md:91, kubernetes.md:49, welcome.md, intake). + None warned that an under-quota-but-oversized request fails with "No valid + host" -- the exact failure devteam hit. CLOSED in the Jenkins guide; NOT yet + in the general docs. + +### F4 -- Partial component detail / no glossary + +Component terms (application credential, the -cluster account, cluster template, +kubeconfig, LoadBalancer Service, floating IP, ingress) are defined inline where +first used but there is no single glossary a client can consult. The operator +specifically asked for "enough detail so they understand what all the components +do." The Jenkins guide added a scoped glossary (section 10); the set has no +shared one. + +### F5 -- Cross-file inconsistencies + +- Jenkinsfile.example calls scripts at clientdocs-starter-kit/... while the kit + is delivered as scripts/ (README.md:27-33). The file admits paths need + adjusting, but it is a stumble. +- Jenkinsfile.example's "Worked sequence" stage runs an infrastructure smoke + build, not an app deploy -- misleading under that name for a k8s reader. + +## Proposed information architecture + +Principle: each fact has ONE owner; every other mention is a one-line pointer, +not a copy. Each guide owns a JOURNEY (a workflow) and cross-links to the owners +for reference detail. This is the shape DOCFIX-123 already demonstrates. + +Proposed owners (single source of truth): +- Identifiers, endpoints, CA bundle, the full account model: handover-pack.md + (already the most complete; make it canonical). +- clouds.yaml / OS_* auth setup: scripts/clouds.yaml.template (the delivery-ready + copy) + one reference section; others point to it. +- Day-2 resource operations (networks, servers, load balancers, secrets): + references/day2-operations.md. +- Troubleshooting signatures: references/troubleshooting.md (the single triage + ladder; other docs link symptoms to it). +- Journeys (own the flow, link the detail): welcome (orientation), + self-service-guide (day-to-day), acceptance-checklist (proof), + ci-integration-guide (automation), jenkins-kubernetes-guide (k8s+Jenkins). + +Add: +- A top-of-set "Start here / reading order" -- either a short new index doc or a + section in welcome.md -- mapping goal -> ordered doc list. +- A shared glossary -- a new small reference, or a section in handover-pack.md -- + that every guide links to (retire the per-guide inline definitions in favor of + one, keeping only a one-line gloss at first use + a link). + +## Consolidation moves (OPTIONS for operator ruling) + +Each move is independent; rule per-move. Effort/risk are rough. + +- M1 -- De-duplicate the account model. Own it in handover-pack.md; replace the + other 3 renderings with a 1-line summary + pointer. + Options: (a) do it; (b) keep the self-service copy (first-read convenience) but + retire the SKILL.md + welcome copies; (c) leave as-is. + Effort: low. Risk: low. Recommend (b) -- one convenience copy, not four. + +- M2 -- De-duplicate clouds.yaml/auth. Own in clouds.yaml.template; the CI guide + and SKILL.md point to it. + Options: (a) do it; (b) leave. Effort: low. Risk: low. Recommend (a). + +- M3 -- Add a "Start here / reading order" entry point. + Options: (a) new short index doc; (b) a section in welcome.md; (c) none. + Effort: low. Risk: low. Recommend (b) -- no new file to package. + +- M4 -- Shared glossary. + Options: (a) new reference file linked from all guides; (b) a section in + handover-pack.md; (c) leave per-guide inline. Effort: medium (touches many + files if links are added). Risk: low. Recommend (a). + +- M5 -- Close the remaining F3 holes in the general docs: add the driver rule and + the host-capacity/right-size caveat to self-service-guide.md (and anywhere + node_count is discussed). + Options: (a) do it; (b) rely on the Jenkins guide + kubernetes.md only. + Effort: low. Risk: low. Recommend (a) -- these are the error-causing holes the + operator called out. + +- M6 -- Fix F5: correct Jenkinsfile.example's script paths to scripts/ and either + rename or re-scope its misleading stage; consider adding a second + Jenkinsfile.example for the kubeconfig-deploy pattern (or point at the Jenkins + guide's inline pipeline). Effort: low-medium. Risk: low (starter asset). + Recommend: fix paths now; the deploy-pipeline example already lives in the + Jenkins guide, so just cross-link. + +- M7 -- Reconcile the quota-vs-capacity expectation cloud-side (separate from + docs): the addendum-40 finding that testcloud tenant quotas over-promise + schedulable capacity. Doc side is M5; the cloud side (trim quotas to reality) + is an operator decision logged in addendum 40. + +## Recommended sequencing + +1. M5 first (closes the active error-causing holes; low effort). +2. M3 + M1(b) (entry point + de-dup the worst offender; makes the set navigable). +3. M2, M6 (mechanical de-dup + the starter-asset fix). +4. M4 (glossary; larger touch, do once the owners above are settled). + +Each move ships under the standard clientdocs discipline (ASCII/LF, no +internal-term leakage, L7 receipt re-record, harness green, changelog + revert) +and, where a delivered client already has the affected file, a package +re-instantiation. Numbers are assigned at implementation, not here. + +## Not in this proposal + +- The broad rewrite into a single mega-document. The operator's goal is + minimize doc-switching PER TASK, which the journey-owns-flow + cross-link model + achieves without collapsing the set into one unnavigable file. If the operator + prefers fewer, larger documents, that is a different architecture to rule on. +- Any client-specific content. This is about the template set; per-client + instantiation is unchanged. diff --git a/docs/archive/dc-dc-ipv6-charm-research.md b/docs/archive/dc-dc-ipv6-charm-research.md new file mode 100644 index 0000000..9c48ba2 --- /dev/null +++ b/docs/archive/dc-dc-ipv6-charm-research.md @@ -0,0 +1,142 @@ +# D-101 IPv6 family-matrix charm research (2026-07-09/10) + +Basis for `overlays/dc-dc-ipv6-family-matrix.yaml` (tooling gap register item +#13). Researched directly against real charm source/config via WebFetch/ +WebSearch -- not inferred from training-data memory of charm behavior, which +this repo's own discipline treats as a real risk class (the OpenTofu module +work earlier this session found a real syntax bug from exactly that mistake). + +## Method + +Fetched actual `config.yaml` files and, where a config option's semantics +were ambiguous from prose alone, actual charm source/commits, from the +`openstack/charm-*` and `openstack-charmers/charm-layer-ovn` repositories +(GitHub mirrors of the OpenDev-hosted canonical source) and +`bugs.launchpad.net`. Every claim below is labeled CONFIRMED (read the +actual file/commit) or INFERRED-BY-PATTERN (not independently checked, +reasoned from a shared framework) -- do not treat the two as equal +confidence. + +## Findings + +### 1. `prefer-ipv6` is a real, shared charms.openstack-family config option + +**CONFIRMED** directly in `charm-keystone/config.yaml`: a boolean, default +`False`, enabling IPv6 support. + +**CONFIRMED** via `charm-nova-cloud-controller`'s actual "Dual Stack VIPs" +commit (`1fa5f7a673006bc6faaf9d377e791448a5ed661b`): setting `prefer-ipv6: +true` causes the charm's HAProxy config template to add `bind +:::{{ports[0]}}` ALONGSIDE the existing `bind *:{{ports[0]}}` -- i.e. this +is a genuinely ADDITIVE dual-stack switch (HAProxy listens on both address +families simultaneously), not an either/or toggle. The commit message +states this explicitly: "HAProxy always listens on both IPv4 and IPv6 +allowing connectivity on either protocol." + +**INFERRED-BY-PATTERN, NOT independently confirmed per-charm:** glance, +neutron-api, placement, cinder, barbican, magnum, openstack-dashboard, +ceph-radosgw all build on the same charms.openstack base classes as +keystone/nova-cloud-controller, so the same `prefer-ipv6` option and +additive-dual-stack HAProxy behavior is EXPECTED -- but each charm's real +`config.yaml` was not individually fetched this session. Confirm via `juju +config ` once any of these charms is actually deployed, before +trusting the overlay is complete for all eleven. + +### 2. The existing `vip:` option takes additional v6 addresses, appended + +**CONFIRMED**: `charm-keystone/config.yaml`'s `vip` option description: +"Virtual IP(s) to use to front API services in HA configuration. If +multiple networks are being used, a VIP should be provided for each +network, separated by spaces." This repo's `bundle.yaml` already uses this +exact mechanism today (D-020's triple-VIP pattern -- e.g. `keystone`'s `vip: +"10.12.4.50 10.12.8.50 10.12.12.50"`, one address per Juju-space binding). +Adding v6/GUA addresses to the SAME space-separated list, once +`prefer-ipv6: true` makes HAProxy listen on both families, is the +mechanism -- confirmed consistent with finding #1's dual-stack commit, +though the EXACT mapping of "which list position pairs with which network" +inside the charm's own context-building code was not traced line-by-line +this session. Treat the overlay's VIP ordering (existing v4 entries first, +then the new v6/GUA entries appended) as the reasonable, but not +byte-for-byte charm-source-confirmed, convention. + +### 3. `ceph-mon` has a real, separate `prefer-ipv6` + network-CIDR mechanism + +**CONFIRMED** directly in `charm-ceph-mon/config.yaml`: +- `prefer-ipv6` (bool, default False) -- per the charm's OWN description, + "If set to False (default) IPv4 is expected," reading as a straight + EITHER/OR switch for THIS charm specifically (unlike the additive + HAProxy behavior in finding #1) -- which is actually the correct shape + for D-101's ULA-ONLY storage/replication planes (no v4 needed at all, + so a clean switch to v6-only is exactly right, not a limitation). +- `ceph-public-network` / `ceph-cluster-network` (string, space-delimited + CIDR list) -- these map directly to D-101's own cited mechanism name, + `ms_bind_ipv6` (the underlying Ceph daemon config the charm sets when + `prefer-ipv6` is enabled). +- A real caveat quoted from the charm's own docs: "these charms do not + currently support IPv6 privacy extension. In order for this charm to + function correctly, the privacy extension must be disabled and a + non-temporary address must be configured/available on your network + interface." This is a real host/OS-level prerequisite for every ceph-mon + unit's storage/replication interface -- not something an overlay file + can set, needs to be part of Phase-0/node-prep discipline. + +### 4. OVN has NO explicit IPv6/encapsulation charm-config option + +**CONFIRMED**: fetched the FULL `charm-layer-ovn/config.yaml` (the shared +base for `ovn-central`/`ovn-chassis`) -- 25 options total, none related to +IPv6, address binding, or encapsulation IP. geneve-over-v6 for the +data-tenant plane is therefore not a charm-config concern at all; it +follows whichever IP family the unit's bound interface on that plane +actually has. Since data-tenant is ULA-only (no v4 present) per D-101, this +needs no overlay entry -- the absence of an entry IS the correct +configuration, not an oversight. + +### 5. Vault's cert-issuance code has no IPv4/IPv6 distinction + +**CONFIRMED**: fetched `charm-vault/src/lib/charm/vault_pki.py`'s actual +`sort_sans()` function -- it splits SANs into "IP SANs" vs. "name SANs" +using `charmhelpers.contrib.network.ip.is_ip()`, with NO further IPv4/IPv6 +differentiation anywhere in the function or its caller +(`generate_certificate()`, which just joins whatever IP SANs it's given +into a comma-separated string). D-109's "Vault issues v4+IPv6-SAN certs" +requirement is technically supported by the charm's own code path, not +merely assumed. + +### 6. A real, open upstream risk for Octavia's lb-mgmt-net over IPv6 + +**CONFIRMED** (real, still-referenced Launchpad bug reports, not +speculation): #1911788 ("IPv6 mgmt network not working, octavia can't talk +to amphora," OpenStack Octavia Charm) and #1913409 ("octavia_amp_network +does not support IPv6," kolla-ansible, describing the same underlying +Octavia limitation from a different deployment tool). #1911788's root cause +traces to an OVN/LXD/MAAS hostname-resolution mismatch (Neutron's +`binding_host_id` carrying a short hostname while OVN's controller expects +the FQDN) -- it was marked a duplicate of #1896630, "managing /etc/hosts +for containers." That is the SAME CLASS of problem this repo's own D-008 +bootstrap order (static /etc/hosts bootstrap -> os-public-hostname -> vault +certs -> Designate, reactivated by D-106 for VR1) already exists to +harden against -- so this is NOT necessarily a fatal blocker here, but it +IS a real, open, documented upstream risk for exactly the kind of network +(Octavia's amphora management network) D-101 wants to put on ULA-only IPv6. +`overlays/dc-dc-ipv6-family-matrix.yaml` deliberately does NOT attempt an +Octavia lb-mgmt-net IPv6 change for this reason -- that needs its own +explicit decision (accept the risk and test for real once DC1 exists, or +keep Octavia's lb-mgmt-net dual-stack/v4 as a deliberate, logged D-101 +exception), presented here rather than silently forced through. + +## What this closes vs. what remains open + +CLOSES (mechanism): the "no overlay file exists, charm option names +unconfirmed" half of gap #13 -- a real, sourced overlay now exists at +`overlays/dc-dc-ipv6-family-matrix.yaml`, covering 9 of Octavia's 11 sibling +dual-stack API charms directly plus ceph-mon's ULA-only switch, with OVN +correctly requiring no entry. + +STAYS OPEN: (a) the 8 INFERRED-BY-PATTERN charms' `prefer-ipv6` option +needs live `juju config` confirmation once any of them is actually +deployed; (b) the exact VIP-list-position-to-network mapping inside each +charm's context-building code was not traced to source; (c) Octavia's +lb-mgmt-net IPv6 question is a real, open decision, not resolved by this +research; (d) none of this has been applied or tested against a live +model -- UNVALIDATED, same posture as every other OpenTofu/overlay artifact +built this session. diff --git a/docs/archive/dc-dc-netem-and-ula-gua-proposal.md b/docs/archive/dc-dc-netem-and-ula-gua-proposal.md new file mode 100644 index 0000000..74b2820 --- /dev/null +++ b/docs/archive/dc-dc-netem-and-ula-gua-proposal.md @@ -0,0 +1,139 @@ +# Proposed netem parameters + ULA/GUA generation guidance (2026-07-10) + +Addresses tooling gap register items #4(d) and #11: the two D-100/D-101 +sub-items still genuinely "leaning," not ruled, even after Stage 0 closed +the big decisions. This document does NOT rule either -- it presents a +concrete recommendation for #4(d) (present options, get a ruling, per this +repo's own discipline for anything touching an ADOPTED-but-not-fully- +specified decision) and, for #11's literal-generation half, exact commands +the OPERATOR runs to produce a real value (this document does not generate +or invent the actual ULA/48 itself -- that is explicitly the operator's/ +NetBox's job per D-101's own text). + +--- + +## Part 1 -- Proposed `tc netem` parameters (D-100 sub-item #4(d)) + +**Current state:** `docs/dc-dc-buildout-design.md` Section 6 states only a +qualitative lean: "same-metro dark fiber (low single-digit ms, jumbo- +capable)." No specific latency/jitter/loss/rate numbers are ruled. +`opentofu/modules/netem-link` (built 2026-07-09) is ready to apply real +parameters the moment they're decided -- this is the one missing input. + +**Proposal (for operator ratification, not silently applied):** + +| Parameter | Proposed value | Rationale | +|---|---|---| +| Latency | 1ms (each direction) | Same-metro dark fiber circuits (single city, no long-haul) typically show sub-millisecond one-way propagation delay for distances under ~50km (fiber propagation is ~5us/km) -- 1ms per direction (2ms RTT) is a deliberately conservative round number that accounts for real-world switching/equipment overhead beyond pure propagation, while staying solidly within the buildout design's own "low single-digit ms" lean. | +| Jitter | 0.2ms | A small, non-zero jitter keeps the simulation honest (a perfectly flat-latency link is unrealistic even on dedicated fiber) without dominating the 1ms base latency -- roughly 20%, a common rule-of-thumb ratio for stable dedicated circuits (as opposed to shared/internet-transit links, which would warrant a much larger jitter fraction). | +| Loss | 0.01% | Dark fiber with modern optical equipment is very low-loss; this is a nonzero-but-negligible value so packet-loss-handling code paths are still technically exercised during the drill, without meaningfully affecting throughput-sensitive tests (Ceph replication, radosgw sync). | +| Rate | Not capped (no `rate` netem param) | The buildout design's own lean explicitly says "jumbo-capable" -- same-metro dark fiber circuits are typically provisioned well above what this test environment's actual traffic will generate (VR1 is a virtual rehearsal on one physical host, not a real multi-site deployment) -- an artificial rate cap would constrain the TEST more than a real Roosevelt link would constrain PRODUCTION, which is backwards for a rehearsal whose value is proving the failover mechanism works, not proving it works under bandwidth starvation. If the operator wants to also rehearse a bandwidth-constrained scenario, that is a SEPARATE, deliberate test configuration, not this default. | + +**Resulting `tc netem` invocation shape** (for reference; `modules/netem- +link`'s `netem_args` variable takes this as a string): +``` +delay 1ms 0.2ms loss 0.01% +``` + +**This is a recommendation, not a ruling.** Per this repo's standing +discipline (present options for anything not yet ADOPTED, never silently +decide), the operator should explicitly ratify this table (accept as-is, +or adjust) before `modules/netem-link` is actually instantiated with real +values in Stage 3's runbook. Once ratified, this should be recorded as a +D-100 amendment (a new "Sub-item RULED" line in `docs/design-decisions.md` +D-100's entry, following the exact pattern D-100's other redline items +already got at Stage 0 ratification) -- not left as a standalone doc. +**Re-tune when a specific Roosevelt inter-DC target is known** (buildout +design's own stated intent) -- this proposal is a VR1 rehearsal default, +not a permanent value. + +--- + +## Part 2 -- ULA/GUA/DC2-supernet generation guidance (D-101, gap #11 + gap #3's data half) + +**What this section does NOT do:** generate the org ULA /48, obtain the +real per-DC GUA carve, or assign DC2's supernet. D-101's own text is +explicit that these are NetBox-authoritative, populated via the extended +import pipeline -- not hardcoded in any decision, and not something this +session (or any Claude session) should invent, guess, or pre-select on the +operator's behalf. What follows is the exact, safe MECHANISM for the +operator to produce each real value themselves, so that step is a two- +minute copy-paste rather than a research task when they're ready to do it. + +### 2.1 -- Org ULA /48 (RFC 4193) + +RFC 4193 requires the 40-bit Global ID be "randomly generated" for +statistical uniqueness across organizations -- a cryptographically random +40 bits satisfies this directly (this is also what every independent +RFC-4193-compliant ULA generator tool checked during tonight's research +actually does under the hood, not a shortcut invented here): + +```bash +printf 'fd%s::/48\n' "$(openssl rand -hex 5 | sed 's/\(..\)\(....\)\(....\)/\1:\2:\3/')" +``` + +Run this ONCE, by the operator, on a real machine -- the output is the +real `ORG_ULA_48` value `netbox/dc-dc-prefixes-import.py` (DOCFIX-152) +requires. Record it in NetBox as the IPAM apex (D-101's own requirement), +not just as an environment variable -- this is a permanent organizational +identifier, not a per-session value. + +> **Correction (DOCFIX-183):** the `sed` above was fixed to split the 10 hex +> digits as 2+4+4 hextets (`fdXX:XXXX:XXXX::/48`); the prior 4+6 split produced +> INVALID IPv6 (a 6-hex-digit hextet -- caught when this command was actually +> run 2026-07-11). **The ratified VR1 value is `fd50:840e:74e2::/48`** (CSPRNG, +> collision-checked vs the Tailscale `/48` in NetBox -- see +> `docs/dc-dc-netbox-buildout-scope.md` section 4c); this command is the +> reference method, not a re-generation step for VR1. + +### 2.2 -- Per-DC GUA carve (from ARIN 2602:f3e2::/32 region-0 /36) + +This is NOT something to generate randomly -- D-101's own text names this +as ALREADY a real, existing ARIN allocation ("GUA (ARIN 2602:f3e2::/32, +region-0 /36)"), meaning the org already holds this block from a real +Internet numbering authority. The per-DC carve within it (which /40 or /44 +goes to DC1 vs. DC2) is an INTERNAL allocation decision within an already- +owned block, not a new external assignment -- this can be decided by +whoever administers that ARIN allocation for this organization, following +whatever internal IPAM convention they already use for carving GUA space +(e.g. sequential /40s: DC1 = `2602:f3e2:1000::/40`, DC2 = +`2602:f3e2:2000::/40`, matching the example addresses already used in +`netbox/dc-dc-prefixes-import.py`'s own docstring and README -- those are +ILLUSTRATIVE examples in this repo, not yet a real assignment; confirm +with whoever holds the ARIN allocation before treating any specific /40 as +final). + +### 2.3 -- DC2's v4 supernet + +This is a NetBox IPAM assignment task, not a generation task -- DC2 needs +a real, non-overlapping v4 supernet (at least a /19, per `netbox/dc-dc- +prefixes-import.py`'s own `DC2_MIN_SUPERNET_PREFIXLEN` check) chosen from +whatever address space this organization has available and not already in +use by DC1, Office1, client VPN ranges (D-074), or any other real +allocation. This is a real IPAM decision requiring visibility into the +FULL address inventory, which this session does not have -- assign it via +NetBox directly, following whatever process this organization already +uses for carving new supernets. + +### 2.4 -- Once all three exist + +Run `netbox/dc-dc-prefixes-import.py --dc dc1` (needs only `ORG_ULA_48` +and `DC_GUA_PREFIX` from 2.1/2.2) and `--dc dc2` (additionally needs +`DC2_V4_SUPERNET` from 2.3) for real, closing the DATA half of tooling gap +#3 -- the mechanism has been ready since DOCFIX-152; only these three real +inputs were ever missing. + +--- + +## Summary: what this closes vs. leaves open + +**CLOSES:** gap #4(d) now has a concrete, reasoned proposal ready for a +one-line operator ratification (not a re-research task); gap #11's ULA- +generation half now has an exact, safe, two-minute command instead of +"go figure out RFC 4193." + +**STAYS OPEN, by design:** the actual netem ratification (a decision only +the operator can make); the actual ULA-48/GUA-carve/DC2-supernet values +(real-world IPAM/ARIN coordination this session has no authority or +visibility to perform) -- this document makes producing them fast and +low-risk once the operator is ready, it does not produce them. diff --git a/docs/archive/dc-dc-replication-DR-seed.md b/docs/archive/dc-dc-replication-DR-seed.md new file mode 100644 index 0000000..948ec6e --- /dev/null +++ b/docs/archive/dc-dc-replication-DR-seed.md @@ -0,0 +1,75 @@ +# SEED: DC-DC cross-DC replication / DR design (Tier A) + +**Status:** SEED / DRAFT for the DC-DC planning chat. Authored by the main-chat stream +during DC-DC planning. This is a settled *decision* to formalize, not a committed record. +The planning chat files the D-entry: **assign next-free via `ledger-scan` at commit -- +note D-076 is already claimed by the PINNED dashboard workstream (draft in a worktree), so +the DR decision is the next free after that; confirm and coordinate before consuming.** + +**Governing constraint:** minimize delta to Roosevelt. + +--- + +## Decision (settled with the operator) + +Two **independent** clouds (separate Keystones), so RBD-mirror is a **data** primitive, not +service failover. DR posture = **Tier A: data recoverability**, **bidirectional** (each DC +protects the other), operator-runbook failover. + +Mechanism, per data class: +- **Cinder volumes -> `cinder-backup`** to a **cross-DC-replicated object store**. This is + the honest cross-independent-cloud path: a restore reconstitutes the volume *with its + Cinder metadata* in the peer cloud, which raw RBD-mirror cannot. Low delta -- `cinder` + already exposes the `backup-backend` endpoint (bound metal-internal) and `ceph-radosgw` + is present as the S3 target. +- **Glance images -> snapshot-based RBD-mirror** of the Glance pool, two-way, with a + re-register step in the failover runbook. +- **Nova ephemeral -> NOT replicated.** Ephemeral root disks are cattle; recovering them + would also need the un-mirrored Nova instance records. Instances are rebuilt from + image + volume on failover. + +Parameters: +- **RPO** = replication/backup interval; start ~15 min, tunable. +- **RTO** = a **measured output** of the failover drill, not a pre-committed SLA. +- **Staging:** prove **one-way** (single mechanism, single direction) + a clean + failover/failback drill FIRST, then enable **two-way**. +- **Carrier:** the **replication plane (`10.12.36.0/22`)** with an explicit **inter-DC + route + bandwidth budget** -- distinct from intra-DC OSD cluster traffic that already + uses this plane. +- **Not replicated:** Neutron state (ports, FIPs, security groups) -- recovered workloads + land on the peer's networks with new addresses; the runbook re-creates this. + +## Options considered + +- **Tier A (CHOSEN):** mirror/backup + operator rehydration runbook. RPO minutes, RTO + hours/measured. Lowest complexity, most robust, delta-minimal. +- **Tier B (folded in for volumes):** `cinder-backup`/restore -- adopted as the Cinder path + because it crosses the independent-cloud boundary cleanly. +- **Tier C (deferred):** orchestrated workload failover via control-plane metadata replay. + Low RTO, high build cost, brittle (no first-class cross-independent-cloud OpenStack DR). + Revisit only if a hard minutes-RTO requirement emerges. + +## Failover / failback runbook -- SKELETON (planning chat to flesh out against the live env) + +Failover (DC1 lost, recover at DC2): +1. Confirm DC1 truly down (avoid split-brain from a transient partition). +2. Glance: promote the mirrored Glance pool at DC2 (force-promote if DC1 unreachable); + re-register images into DC2 Glance. +3. Cinder: restore required volumes into DC2 Cinder from the replicated backups. +4. Rebuild instances from image + restored volume; re-create Neutron ports/FIPs/SGs. +5. Record RTO/RPO actuals. + +Failback (DC1 recovered) -- split-brain-safe ordering for two-way: +1. Do NOT let recovered DC1 resume as primary automatically. +2. Demote DC1's Glance pool; resync from DC2 (now primary). +3. Reconcile Cinder: back up any DC2-side changes; restore into DC1. +4. In a controlled window, optionally flip primary back (demote DC2, promote DC1). +5. Verify both directions healthy; record the drill. + +## Open sub-items for the planning chat +- radosgw **multisite** vs. a mirrored backup bucket as the replicated backup target. +- Exact mirrored pools; **consistency groups** for multi-volume apps. +- Inter-DC replication route + bandwidth budget on the simulated WAN (MTU-aware). +- Interaction with the **D-071** controller-HA decision (backup/restore of per-DC control + planes). +- `ceph-rbd-mirror` charm topology + daemon HA; peer bootstrap on the replication plane. diff --git a/docs/archive/docfix-draft-20260702.md b/docs/archive/docfix-draft-20260702.md new file mode 100644 index 0000000..c5c65ff --- /dev/null +++ b/docs/archive/docfix-draft-20260702.md @@ -0,0 +1,215 @@ +# DOCFIX draft -- redeploy-readiness review (bundle + channels + runbooks) + +STATUS: DRAFT / OPEN -- accreting during the 2026-07-02 review session. +Numbers are PROVISIONAL from next-free DOCFIX-066 (verified against HEAD +690779a: DOCFIX-065 / D-068 / BUNDLEFIX-008 consumed). Renumber-check again +at commit time. ASCII + LF. + +Scope of this review: + 1. bundle.yaml -- YAML validity, structural consistency, known anti-patterns. + 2. Charm channel pins -- staleness review against current Charmhub guidance. + 3. Runbook sweep -- cross-reference integrity, stale values, anything that + would break the next redeploy. + +Severity key: BLOCKER (breaks redeploy) / RISK (may break or mislead) / +NIT (consistency only). + +-------------------------------------------------------------------------------- +## Findings +-------------------------------------------------------------------------------- + +(appended as found) +### DOCFIX-066 (BLOCKER) -- teardown runbook drives the DEPRECATED teardown script +File: runbooks/phase-00-teardown-maas-reset.md (steps 2, plan table, lines 22/30/67/74/85). +The runbook's execution spine is `scripts/phase-00-teardown.sh --apply` with the narrative +"hosts release to MAAS Ready" -- the exact premise DOCFIX-057/D-061 proved WRONG on this +virsh-pod MAAS (destroy-model DECOMPOSES pod-composed machines; observed 3x). The script +itself carries a DO-NOT-USE banner, so the runbook and script now contradict each other; +an operator following the runbook on the next redeploy either hits the deprecation mid- +teardown or, if they push past it, triggers a fourth decompose + full reenroll/recarve. +The D-061 replacements (phase-00-teardown-release.sh --keep-instance + canary; +phase-00-teardown-destroy.sh) exist but are never mentioned in the runbook. +FIX: rewrite the runbook spine as the D-061 fork -- (a) machine-preserving path: +teardown-release.sh with the MANDATORY first-run canary (--apply --canary, verify +openstack0 survives in MAAS, then all-four); (b) from-scratch path: teardown-destroy.sh ++ reenroll + recarve. State which path the standard redeploy uses. Step-5 "hosts Ready" +premise and the OSD-wipe/8_lbaas step ordering must be revalidated per path (release path +leaves hosts Deployed, not Ready -- the wipe/carve preconditions differ). + +### DOCFIX-067 (RISK -- verify live before ruling) -- octavia PKI cert SAN carries the pre-R14 VIP +File: runbooks/phase-01-bundle-deploy.md 1.0-GEN.c (lines ~292, ~318). +The controller-cert CNF sets `IP.1 = 10.12.4.233` -- the OLD octavia VIP. R14 relocated +all VIPs to .50-.60; bundle HEAD has octavia at 10.12.4.57. The octavia-pki overlay is +regenerated every phase-01, so the LIVE cloud's controller cert most likely carries .233 +today; octavia passed phase-05 validation regardless, so the amphora side evidently does +not verify that SAN IP -- functional impact UNPROVEN, inconsistency CERTAIN, and it is a +latent break if SAN verification ever tightens (or when Roosevelt re-uses this block). +The DNS.1/DNS.2 SANs also reference the D-019-dropped FQDN scheme (harmless, same sweep). +FIX: derive the SAN IP dynamically (lib-net VIP_PREFIX_PROVIDER + octavia's octet from +bundle/juju -- rule 3), not a literal. VERIFY-LIVE first (gated CHECK, jumphost): +read the overlay/live cert SAN and confirm what is actually deployed: + openssl x509 -in -noout -text | grep -A2 'Subject Alternative Name' +Rule on severity after the read: if the live cert has .233 and octavia is green, keep RISK +(doc fix + regenerate at next redeploy); do not hot-rotate certs mid-cloud for this. + +### DOCFIX-068 (RISK) -- phase-01 "Constants and env-literals" block is pre-D-052 stale +File: runbooks/phase-01-bundle-deploy.md lines ~23-26. The block states the RETIRED plane +map (2=metal .8, 6=data .12, 7=storage .16, 8=replication .20, 9=lbaas .32 -- wrong +plane->CIDR pairs under D-052/D-053, incl. the retired `lbaas` space), plus hardcoded MAAS +subnet IDs (violates lib-net PATTERN-1: IDs drift, resolve by CIDR) and hardcoded +system_ids (violates DOCFIX-040: re-minted per enrollment; lib-hosts resolves). Mixed +freshness: the "50 apps, 97 relations" expectation in the same block MATCHES bundle HEAD. +Misleading at the worst moment (mid-deploy reference values). +FIX: replace the stale lines with pointers to scripts/lib-net.sh (planes) and +scripts/lib-hosts.sh (host identity); retain only verified-current literals. + +### DOCFIX-069 (RISK) -- zero exec bits + bare script invocations +git index: ALL 37 files under scripts/ are mode 100644 (GitHub Desktop workflow strips ++x). Fresh clone on the jumphost -> every BARE invocation fails "Permission denied". +runbooks/phase-00-teardown-maas-reset.md invokes bare in ~10 places (teardown, carve, +standup); other runbooks appear bash-prefixed (sweep found no other bare hits). +FIX (durable, matches the Windows commit constraint): bash-prefix every script invocation +in runbooks (`bash scripts/x.sh ...`). Optional belt: `git update-index --chmod=+x +scripts/*.sh` -- but Windows-side recommits can strip again, so the bash prefix is the +invariant; do both if desired. + +### DOCFIX-070 (RISK) -- scripts/review-bundle.py is pre-D-052 stale; NOT CLEAN is noise +Against bundle HEAD it reports FAIL=71: expects space `metal` (retired), DUAL VIPs (D-020 +form; D-052 moved to triples), no per-endpoint bindings (D-052 introduced them), vault +1.8 (D-068 pinned 1.16), baselines 51 apps/98 rels (now 50/97: VIP set changed -- vault +dropped its VIP, ceph-radosgw gained one -- blessed by provider-bundle-check.py, which +PASSES clean). Hazard: a pre-deploy NOT CLEAN verdict that must be ignored trains alarm +fatigue and will eventually mask a real defect. +OPTIONS: (a) update review-bundle.py expectations to the D-052/D-060/D-062/D-068 model; +(b) retire it (git rm) and fold any still-unique checks (relation-endpoint syntax, +phantom-key detection reworked for the per-endpoint model) into provider-bundle-check.py; +(c) banner it historical. RECOMMEND (b): one authoritative gate beats two disagreeing +ones -- same reasoning as the D-060 d057-bundle-check retirement. + +### NIT-A -- D-002 channel matrix drift (design-decisions) +The D-002 table still lists `etcd, easyrsa -> latest/stable` (etcd/easyrsa dropped; R3 / +phase-02 record vault-on-mysql), omits memcached (bundle: latest/stable -- apparently the +only track that charm publishes; upstream's "never latest/stable" applies to OpenStack- +project charms, which memcached is not), and its vault row (1.8) is superseded by D-068 +(1.16). Append-only fix: a dated amendment note under D-002, not an edit. +VERIFY-LIVE (gated CHECK, jumphost) before finalizing: + for c in memcached rabbitmq-server vault hacluster; do juju info "$c" 2>/dev/null | sed -n '/channels:/,$p' | head -12; done +Expected: memcached publishes only latest/*; rabbitmq-server tops out at 3.9; vault +carries 1.16/stable; hacluster 2.4/stable. + +### NIT-B -- channel-pin review conclusion (informational; no change) +All bundle pins judged CURRENT for Caracal/jammy: 2024.1/stable core (18), OVN +24.03/stable, ceph squid/stable, mysql 8.0/stable (12), hacluster 2.4/stable (11), +rabbitmq-server 3.9/stable (terminal track for this charm), vault 1.16/stable (D-068), +memcached latest/stable (sole track; see NIT-A). Upstream charm-guide delivery page is +frozen (last updated 2023-12) -- Charmhub/juju info is the only live authority; the +NIT-A verify block doubles as the pre-deploy channel assert. Candidate: fold that assert +into scripts/pre-flight-checks.sh (D-002 claims pre-flight verifies channels -- confirm +it actually does; not yet audited). + +### NIT-C -- ASCII-rule violations in docs/ +docs/v1-pre-deploy-fixes.md (277 non-ASCII bytes), docs/netbox-vip-queue.md (81). The +repo rule is ASCII-only for all committed files (mod_wsgi lesson). Low functional risk +(docs, not conf), but the rule is stated absolute -- sanitize or record a carve-out. + +### NIT-D -- identifier index gaps +DOCFIX-027/028/029/034/037 and BUNDLEFIX-001..006 are defined only at point of use +(runbook/bundle comments) and absent from appendix-A / the changelog index -- appendix-A +claims to be the index "keyed by the same identifiers used inline". Add one-line index +entries (or mark point-of-use-only identifiers as such). + +### NIT-E -- appendix-A lacks a mysql-innodb-cluster recovery entry +D-062 material (blocked 'Instance not yet configured' = single-unit seed; half-join +instanceErrors = mid-life rescan; reboot-cluster-from-complete-outage ONLY on confirmed +outage -- destructive against a healthy cluster) exists in design-decisions + the restart +procedure but has no appendix-A symptom entry. Add one; also consider committing the +restart-procedure doc to the repo (it currently lives outside it). + +-------------------------------------------------------------------------------- +## Verify-live queue (gated CHECKs for the jumphost before findings finalize) +1. Octavia controller cert SAN (DOCFIX-067) -- read the deployed overlay/cert. +2. juju info channel probe (NIT-A/B) -- memcached / rabbitmq-server / vault / hacluster. +3. pre-flight-checks.sh -- confirm whether it performs the D-002 channel assert. + +-------------------------------------------------------------------------------- +## Deployment-flow parity findings (decision vs bundle vs schedule) +-------------------------------------------------------------------------------- + +### DOCFIX-071 (BLOCKER) -- D-064 keystone policy attach is not reachable from the deploy schedule +Evidence: bundle.yaml keystone has use-policyd-override=True but NO resources: stanza; +`attach-resource keystone` appears in NO phase runbook or script -- only appendix-C:73-74. +phase-01:183 knowingly deploys into "PO (broken)" and phase-02:167 re-notes it as +FINDING-1 "not a regression"; no phase ever resolves it. The live cloud got the policy +via a session action (D-064), never folded into the schedule. NEXT REDEPLOY as written: +keystone stays PO (broken), the SCS Domain Manager RBAC (the commercial tenant-isolation +core, D-051) is ABSENT, and tenant onboarding fails at G3. +Compounding defect: the appendix-C block zips to and attaches FROM /tmp -- the +documented snap-confinement trap (attach-resource cannot read /tmp on this jumphost; +use $HOME). The only written procedure is the known-broken form. +FIX OPTIONS (debate): + (a) Bundle-native resource: add to keystone `resources: {policyd-override: + ./policies/overrides.zip}` and commit the zip beside its source yaml. Deploy-time + attach becomes automatic -- the bundle describes the WHOLE desired state, zero + manual step, zero Roosevelt delta. Sync risk (zip vs yaml drift) is closed by a + pre-flight assert: rebuild the zip from policies/, byte-compare against committed, + HOLD on mismatch. + (b) Schedule step: a gated attach block in phase-03 (post-TLS-settle), $HOME-pathed, + gating on `PO:` in juju status. + EITHER WAY: the G3 BEHAVIORAL gate (manager can self-service own domain; admin-grant + and cross-domain DENIED; cloud-admin unaffected) must be a phase step -- D-051's own + warning: the charm validates YAML only, `PO:` proves parse, not policy. RECOMMEND + (a) + G3 gate in phase-03: post-deploy manual steps are exactly the class D-046 + proved unreliable ("reports ready regardless"). + +### DOCFIX-072 (RISK) -- bundle implements the still-PROPOSED D-043 +bundle.yaml nova-compute sets resume-guests-state-on-host-boot: True while D-043 +(tenant-VM auto-resume) remains PROPOSED / decision-pending. The bundle is ahead of the +decision record -- the exact drift the discipline forbids (and the restart-procedure doc +already assumes the option is in force). FIX: rule on D-043 -- RECOMMEND adopting its +option (a) (auto-resume + monitoring; industry norm for tenant VMs; customers at +Roosevelt expect VMs back after host maintenance; D-041's down-is-a-signal stance is +preserved for CONTROL-PLANE services, which auto-resume does not touch) -- and mark the +decision ADOPTED with the bundle line as its implementation. Alternative: strip the +option until ruled; NOT recommended (regresses the validated restart procedure). + +### NIT-F -- D-011.6 text not amended to the phase-08 ruling +design-decisions D-011 item 6 still reads "Vault unseal + auto-unseal-after-reboot +pattern verified"; phase-08 D-011.6 rules MANUAL unseal is the v1 standard (auto-unseal +NOT configured). Append an amendment note to D-011 so the acceptance bar and the +acceptance runbook agree. + +### NIT-G (Roosevelt-forward) -- rabbitmq-server scale-up will race without min-cluster-size +Testcloud: num_units=1, no min-cluster-size -- correct per D-009 (decorative HA). But the +D-009 promise is "Roosevelt scale-up is mechanical: 1 -> 3 and rerun". For rabbitmq that +is NOT sufficient: without `min-cluster-size: 3` the charm accepts client relations +before the cluster forms (same failure CLASS as D-062's mysql formation race; upstream +charm docs call min-cluster-size best practice). Record now as a Roosevelt bundle-delta +note on D-009 so the mechanical scale-up story stays true. + +-------------------------------------------------------------------------------- +## Patchset status (2026-07-02, patchset-20260702-redeploy-readiness) +-------------------------------------------------------------------------------- +IMPLEMENTED in the delivered ZIP (numbers verified next-free at HEAD 690779a; +re-grep at commit): DOCFIX-066 (teardown runbook rewritten around the D-061 fork, +destroy path = validated spine, reenroll step added, all invocations bash-prefixed), +DOCFIX-067 (octavia SAN IP derived from bundle at generation time; verify-live of the +deployed cert still queued), DOCFIX-068 (phase-01 constants -> lib-net/lib-hosts), +DOCFIX-069 (bash-prefix; optional chmod noted in apply-notes), DOCFIX-070 (checks +absorbed into provider-bundle-check.py + 8-case harness; review-bundle.py to git rm), +DOCFIX-071 (bundle-native keystone policy resource + committed zip + drift guard + +phase-03 Step 3.4 two-stage gate + appendix-C /tmp fix + subshell wrap), DOCFIX-072 +(D-043 RESOLVED->ADOPTED(a)). D-doc amendments appended: D-002, D-009, D-011, D-043, +D-051, D-061. STILL OPEN: NIT-C (docs/ ASCII sanitize), NIT-D (identifier index), +NIT-E (appendix-A mysql entry), verify-live queue items 1-3. + + +## Patchset status addendum (2026-07-03, Block 2) +IMPLEMENTED: DOCFIX-073 (preflight + channel assert + phase-01 gate), DOCFIX-074 +(repo-lint + full ASCII sanitize incl. .gitignore/netbox; closed NIT-C), DOCFIX-075 +(cloud-assert + committed ops-restart-procedure; closed the health-check gap), +DOCFIX-076 (as-executed convention + run-logged + index), DOCFIX-077 (appendix-A +mysql entry + identifier index; closed NIT-D/E), DOCFIX-078 (security ledger), +D-069 (vault custody policy), D-070 (supersedes D-012). Verify-live queue item 2 +(channel probe) is now AUTOMATED by preflight P3. Remaining operator inputs: +SEC-003 custodian assignment; capi-mgmt auto-resume exclusion ruling; octavia +deployed-cert SAN read. See docs/changelog-20260703-process-hardening.md. diff --git a/docs/archive/handoff-20260703-open-items.md b/docs/archive/handoff-20260703-open-items.md new file mode 100644 index 0000000..9b95428 --- /dev/null +++ b/docs/archive/handoff-20260703-open-items.md @@ -0,0 +1,164 @@ +# Handoff -- open items register (2026-07-03 session close) + +STATUS SNAPSHOT at handoff: repo HEAD abc144a; gauntlet ALL GREEN (26 +harnesses); repo-lint 0 fail (1 documented legacy WARN); Claude Code live on +the jumphost with CLAUDE.md + permission rules + PreToolUse guard (13/13); +skill single-sourced at .claude/skills/openstack-cloud-ops (v1.3). Two +blockers from the redeploy-readiness sweep (DOCFIX-066 teardown spine, +DOCFIX-071 policy delivery) are FIXED and merged. Full change record: +docs/changelog-20260703-process-hardening.md (items 1-32, per-item reverts). + +Numbering at handoff (re-grep before assigning -- L5 prints it): +next-free D-071, DOCFIX-082, BUNDLEFIX-009. + +Conventions for whoever picks this up: repo wins over any session memory; +grep design-decisions before changing a built surface; every script change +ships with its harness green + a changelog entry with a revert; mutations are +individually human-gated (the permission ask-rules enforce this on the +jumphost -- do not work around them). + +-------------------------------------------------------------------------------- +## 1. Immediate (small, well-defined; good first Claude Code tasks) +-------------------------------------------------------------------------------- + +H-1 (DOCFIX-082 candidate) -- tenant-offboard.sh hardening + The new offboard script is well-built (audit-default, typed gate, protected- + domain blocklist, dependency-ordered) but runs `set -u` only. It MUTATES in + --apply: a silently failed pipeline mid-sweep strands resources while + reporting progress. FIX: `set -uo pipefail` + capture-then-test conversions + per references/script-authoring.md (both SIGPIPE directions); state the + chosen error regime in the header; strengthen is_id to the 32-hex form. + ACCEPT: tests/tenant-offboard green incl. a new failure-injection case; + changelog entry with revert. + +H-2 (DOCFIX-082 or 083) -- vault-kv-inner-probe.sh: secret off argv + The AppRole probe builds the login body with role_id/secret_id and (verify) + passes it via curl argv -- visible in `ps` on the unit. TTL is 60s so + exposure is bounded, but the house rule is absolute: secrets never transit + argv. FIX: feed the body via stdin (`curl --data @-`). ACCEPT: probe + harness green; a grep proves no secret-bearing var appears in a curl argv. + +H-3 -- record-keeping for the two new scripts + tenant-offboard.sh + tenant-offboard tests and vault-kv-health.sh (+ inner + probe, + tests) have NO changelog entries or identifiers. Add entries + (what/why/revert) under the next-free DOCFIX numbers; vault-kv-health cites + D-068 item 3 -- link it there too. + +H-4 -- claude.ai skill copy parity + The chat-side installed skill visible at session close lacked the v1.3 + hermetic-harness scar. Confirm the re-upload landed (a fresh chat session + sees the current mount); if not, re-upload the .skill regenerated from + .claude/skills/openstack-cloud-ops. Standing habit: any edit to the skill + source -> repackage -> re-upload. + +H-5 -- Claude Code guardrail smoke test (if not already done) + Fresh session on the jumphost: `/permissions` shows the three rule sets; + `maas list` is hard-blocked by the hook; a `juju status` runs unprompted; + any `--apply` prompts. One minute; proves settings loaded. + +-------------------------------------------------------------------------------- +## 2. Verify-live queue (read-only CHECKs; jumphost) +-------------------------------------------------------------------------------- + +V-1 (DOCFIX-067 severity ruling) -- octavia deployed-cert SAN read + The runbook now derives the SAN dynamically; the LIVE cloud's controller + cert was generated under the old literal. Read what is actually deployed: + + CHECK (read-only) -- jumphost + ```bash + ( { + CRT=$(juju ssh -m openstack octavia/leader -- \ + 'sudo cat /etc/octavia/certs/controller_cert.pem 2>/dev/null || sudo ls /etc/octavia/certs' &1) + printf '%s\n' "$CRT" | openssl x509 -noout -text 2>/dev/null \ + | grep -A2 'Subject Alternative Name' || printf '%s\n' "$CRT" | head -5 + } ) + ``` + (Path may differ -- if the ls fallback fires, locate the controller cert per + phase-01 1.0-GEN and re-read.) RULE ON RESULT: if SAN carries the old + 10.12.4.233 and octavia is green (expected), record RISK-accepted in the + changelog -- regenerated at next redeploy; do NOT hot-rotate certs for this. + +V-2 (D-069) -- second-person unseal rehearsal + Someone other than the initializing operator performs the manual 3-of-5 + unseal per ops-restart-procedure Stage 3. This is an ACCEPTANCE item, not a + nicety: "keys work" and "a second human can use them" are different facts. + Requires SEC-003 custodian assignment first (section 3). + +V-3 -- channel assert: AUTOMATED (no action) + preflight.sh P3 now runs the Charmhub channel assert live on every pre- + deploy run; the manual `juju info` probe from the old queue is retired. + +-------------------------------------------------------------------------------- +## 3. Operator rulings pending (nothing proceeds without these) +-------------------------------------------------------------------------------- + +R-1 (SEC-003 / D-069) -- vault unseal-key custodian assignment + Policy is ADOPTED (split custody, no individual holds threshold); the + ASSIGNMENT (who, what media) is operator input, recorded off-repo; flip the + ledger row when done. + +R-2 (D-043 caveat) -- capi-mgmt-v2 auto-resume exclusion + Cloud-wide auto-resume means the mgmt VM WILL resume after host reboots; + its manual-start policy now governs deliberate stops only. If a REAL + exclusion is wanted, it needs a mechanism (options to draft on request); + otherwise close the caveat with a one-line D-043 amendment accepting it. + +R-3 (D-063, PROPOSED/OPEN) -- capi-mgmt SG hardening + Options a/b/c recorded in design-decisions; option (a) requires the + MEASURED post-NAT conductor source, never inferred. Phase-07 depends on + 6443 reachability -- rule before the next redeploy or explicitly defer. + +-------------------------------------------------------------------------------- +## 4. The deploy path (the reason for all of the above) +-------------------------------------------------------------------------------- + +D-1 -- Pattern-A full redeploy (VR0 DC0) + Now runnable end-to-end from the docs alone: phase-00 (D-061 DESTROY path: + teardown-destroy -> 8_lbaas -> OSD wipe -> reenroll -> carve x4 -> standup) + -> `bash scripts/preflight.sh` PASS -> phases 01-08 gated (keystone policy + now arrives via the bundle; phase-03 Step 3.4 gates PO: + behavioral G3) + -> `bash scripts/cloud-assert.sh --capture` -> commit the asbuilt/ BOM. + Session discipline: `bash scripts/run-logged.sh phase-NN-` first. + +D-2 -- D-011 acceptance: implement validate.sh + The placeholder's TODO is now aligned to the AMENDED decisions (manual + unseal, no snapshots, no Designate) and points at the building blocks + (cloud-assert / tenant-acceptance / run-tests-all). Still to author: VIP + reachability from jumphost + tenant VM, the LB round-robin/failover pattern + test, timed fresh-tenant Magnum e2e, and V-2 as a checklist item. + +D-3 -- confirm D-042 fix status + Memory says the seven-stage magnum-capi-helm contract-ref runbook was + staged pending execution; the repo record should say whether it ran. + Confirm from the changelog/decisions and either execute or close. + +-------------------------------------------------------------------------------- +## 5. v1-close checklist (execute after D-011 passes) +-------------------------------------------------------------------------------- + +C-1 Consolidate the 10 per-phase do-documents into docs/v1-deploy-runbook.md + (structure preserved in git history via the consolidation commit). +C-2 Consolidate/close docs/docfix-draft-20260702.md -- everything implemented + graduates to the changelog; anything decision-shaped to design-decisions. +C-3 Repo visibility back to PRIVATE (SEC-004; Settings -> Options). +C-4 Rotate the libvirt SSH credential (SEC-001 -- exposed 2026-06-26). +C-5 Ledger review pass (docs/security-ledger.md) -- every row OWNED + dated. + +-------------------------------------------------------------------------------- +## 6. Deferred, with explicit revisit triggers +-------------------------------------------------------------------------------- + +- OS sandboxing (bubblewrap) for Claude Code Bash: revisit at Roosevelt or if + a prompt-injection-shaped incident ever occurs on the jumphost. +- Managed settings (disableBypassPermissionsMode): when a SECOND operator + gets jumphost access. +- Plugin graduation (wrap the skill + slash-commands /preflight /cloud-assert + + hooks): when Claude Code is the daily driver and prompt-count friction is + measurable. The skill remains the source either way. +- Jenkins CI poll job for repo-lint/gauntlet: when the tenant-rehearsal + Jenkins graduates to an ops instance. +- rabbitmq min-cluster-size:3 (D-009 amendment): Roosevelt bundle delta only. +- SSH on GitBucket (port 29418), IPv6 dual-stack, NetBox import bundle: + v2-deferred as previously recorded. + +-- end of register -- diff --git a/docs/archive/handoff-20260705-open-items.md b/docs/archive/handoff-20260705-open-items.md new file mode 100644 index 0000000..6ac9f74 --- /dev/null +++ b/docs/archive/handoff-20260705-open-items.md @@ -0,0 +1,179 @@ +# Handoff -- open items register (2026-07-05 session close) + +**Supersedes** `docs/handoff-20260703-open-items.md` (kept for history). This is the active handoff. +**Repo HEAD at close:** `f28de74` ("clearing queue"). **Next-free:** D-073, DOCFIX-091, BUNDLEFIX-012. +**Start every new session by:** reading this file + `docs/session-ledger.md`, then running +`bash scripts/ledger-scan.sh` AND `bash scripts/ledger-scan.sh --fences`, and reconciling. + +--- + +## 0. Current state (as-built) + +- Multi-tenant buildout is functionally complete through the beta tenant acceptance tests. +- `validate.sh` D-011 acceptance suite is modular and the **gauntlet is ALL GREEN (31 harnesses)**. +- **Vault decision recorded (D-068 amendment):** bundle pins vault `1.8/stable` (reactive charm, + integration-compatible); the `1.16` operator lineage is RULED OUT (certs interface V0->V1 breaks + the 18 `vault:certificates` relations, Raft-only drops the mysql backend, no upgrade path, BUSL, + community-broken with Ceph). "Get off EOL vault 1.8.8" stays OPEN under D-068. Evidence: + `docs/D-068-vault-1.8-vs-1.16-analysis.md`. Live vault is `1.8/stable` rev 714 (in-channel current). +- **Session-ledger is fenced** (delimited-section convention: machine-derived / main-chat / jumphost / + shared), and `ledger-scan.sh --fences` validates it (4 sections, 0 errors). Stay in your lane. +- The last jumphost ops-update window (`ops-update-20260705`) is CLOSED; fleet refreshed to current + in-channel revisions; controllers/agents at juju 3.6.25; appendix-B re-baselined. + +**Minor reconcile item for the next session:** the ledger "State facts" bullet still reads +"ops-update-20260705 in flight" and "D-068 unruled" -- both stale (window closed; D-068 ruled). It's +in the append-only shared section, so append a correction rather than rewrite another stream's line. + +## 1. Immediate (small, well-defined; good first tasks) + +- Batch-3 live validation of the D-011 checks (see the verify-live queue below) -- the main thing + standing between here and a clean D-011 acceptance close. +- `offboard` v2: `--sweep-magnum-orphans` mode (orphan per-cluster trustee in the magnum domain); + stage-3 app-cred idempotency re-run guard. (Both logged-not-actioned.) +- DOCFIX candidate: sweep the repo for other `-f json ... 2>&1` stderr-merge hazards (DOCFIX-085 class). +- `host_href=None` barbican observation (exercised fine in stage 6; probably closeable after a look). + +**Window runbook (2026-07-06):** the whole batch-3 + foil + D-073-apply sequence is +scripted as ONE do-document: `runbooks/d011-batch3-window-DRAFT.md` (DOCFIX-093). +B2/D-068 discussion input ready: `docs/D-068-openbao-assessment-DRAFT.md`. + +## 2. Verify-live queue (read-only/gated CHECKs; run on the real cloud before trusting) + +ALL CLEARED 2026-07-06 (d011-batch3 window, addendum 27): +- **d011-04** headroom parsing: fixed live (DOCFIX-096: --long for compute_id; + display-name hypervisor keys); OK path confirmed, disruptive half PASS under + the A2 conditional pre-approval. OCCM LB-name + agnhost RR assumptions held. +- **d011-05** foil satisfied by the foil1 tenant (onboarded in-window); PASS + incl. P3 isolation. +- tenant-assert `keypair list --user` confirmed under admin; admin-side + app-cred list measured always-403 -> check reworked (DOCFIX-095). +- H3 CANARY_SSH_USER confirmed (ubuntu). +STILL OPEN here: Horizon manual checks (onboarding contract 5.1/5.3) on a +live tenant -- operator, browser; non-blocking. + +## 3. Operator rulings pending (nothing proceeds without these) + +- **D-071** (controller update cadence / patch policy) -- stays **PROPOSED**; ratification is GATED on + the pre-DC-DC controller HA/backup planning session (see section 6). Do NOT ratify without it. + Backups are confirmed to exist (juju 3.6 `create-backup`); the open half is single-controller HA. +- **D-068 item 1** vault modernization -- `1.16` is ruled out; the OPEN question is HOW to get off EOL + vault 1.8.8 for production. Candidate paths (none adopted): wait for the OpenStack service charms to + gain tls-certificates V1 support, evaluate OpenBao (MPL fork), or explicit EOL risk-acceptance for + VR0 with a production remediation deadline. Plus D-068 items 2 (Vault listener TLS for Roosevelt -- + vault_url is cleartext http today) and 3 (AppRole secret_id TTL audit + proactive auth health probe). +- ~~D-050 / list_trusts hardening~~ RESOLVED: D-050 CLOSED via D-051 (ruling + 2026-07-06); the list_trusts hardening was adopted as D-073 and APPLIED live + 2026-07-06 (addendum 27; behavioral verify green). + +## 4. Open security-ledger rows + +- **SEC-001** OPEN -- rotate credentials after the current rebuild completes. +- **SEC-003** OPEN -- assign unseal-key custodians + rehearse the second-person unseal (D-069). +- **SEC-004** OPEN -- flip repo visibility to PRIVATE at v1 close (currently public for web_fetch). + +## 5. The deploy path (the reason for all of the above) + +Execute the Roosevelt bare-metal multi-datacenter deployment using the Omega Cloud v1 runbooks as the +template; the guiding constraint remains minimizing delta to Roosevelt. D-011 (amended per D-019) +acceptance validation is the near-term gate -- confirm all criteria post multi-tenant work, which is +what the batch-3 verify-live queue closes out. + +## 6. v1-close / project-completion checklist (execute after D-011 passes) + +- Consolidate the 10 per-phase do-documents into a single `docs/v1-deploy-runbook.md`; drop the + per-phase files; preserve the 10-doc structure in git history via the consolidation commit. +- Set repo visibility PRIVATE (SEC-004). +- Hold the pre-DC-DC controller HA/backup planning session -> then rule D-071. +- **Onboarding completion package (operator-ruled 2026-07-06; complete BEFORE closing this + deployment test env).** A written, tested, and validated end-to-end onboarding workflow: + scripts + internal docs + drafted client-facing docs (drafts/living, no final polish needed). + Items (H-numbers per docs/tenant-onboarding-contract.md section 6): + - H1 `tenant-assert.sh` post-onboard verifier, with harness (also offboard preflight + + periodic drift sweep). DELIVERED addendum 22; live validation pending (foil onboarding). + - H2 stage-3 app-cred idempotency re-run guard in tenant-onboard.sh. + DELIVERED addendum 23 (also covers the keypair; harness 12/12). + - H3 canary access-proof stage (boot canary + FIP + SSH via tenant keypair + teardown) -- + the ruled meaning of "confirm tenant SSH access into their domain" (A1.4 confirmed). + DELIVERED addendum 23 as stage7 (explicit-only); live validation pending (foil). + - H4 `keystone-policy-drift.sh` (script backlog item 6) wired as a periodic check. + (Only H-item not yet built; also the main-chat backlog owner -- coordinate.) + - H5 no-ad-hoc rule embedded in the onboarding runbook. DONE addendum 25. + - Horizon manager-role GUI identity probe (contract section 5.3) validated live. + - Client-facing draft package (intake form, welcome/engagement doc, self-service guide) + -- RULED: top-level clientdocs/; drafts DELIVERED addendum 22 (living docs, sweep on + contract-doc changes). + - Full workflow validated on a live onboarding end-to-end (the d011-05 foil tenant is the + natural first candidate). + +## 7. Deferred, with explicit revisit triggers + +- **GitBucket SSH** (System Settings -> Integrations; bind 0.0.0.0:29418, public git.baldurkeep.com:29418) + -- verify the deployment topology (Docker port map vs systemd vs firewall) before declaring it works. + v1 uses HTTPS basic auth. Trigger: v2 / DC-DC phase. +- **IPv6 dual-stack** integration + **NetBox restructure/import** -- deferred to DC-DC / v2. +- **Roosevelt hardware sheets** needed to resolve the four-NIC collapse path (D-059, parked). +- **SpryLogin / identity platform** decision (FreeIPA vs Keycloak vs Authentik). + +--- + +## 8. Startup prompt -- NEW MAIN CHAT (claude.ai stream) + +> You are resuming the Omega Cloud v1 / Baldurkeep project (commercial multi-tenant Charmed OpenStack +> Caracal 2024.1). You are the **main claude.ai chat stream**. A parallel **Claude Code (jumphost)** +> stream runs concurrently -- coordinate, do not collide. +> +> **First, orient (do this before anything else):** pull the repo; read +> `docs/handoff-20260705-open-items.md`, `docs/session-ledger.md`, and `docs/design-decisions.md`; +> run `bash scripts/ledger-scan.sh` and `bash scripts/ledger-scan.sh --fences`; reconcile the ledger +> narrative against the scan. The repo is authoritative over your memory of it. +> +> **Your lane / discipline:** you work a sandbox clone via gated copy-paste; the operator (Jesse) runs +> blocks on the jumphost and pastes output back. Deliver multi-file changes as **repo-relative ZIPs** +> (never loose files). Committed files are **pure ASCII + LF**. In the session-ledger, edit only the +> `main-chat` fenced section; the machine-derived block is regenerated by `ledger-scan` (never +> hand-typed); the shared section is append-only. When you consume a D-/DOCFIX-/BUNDLEFIX- number, +> grep for next-free first and coordinate with Code. **Harness-first**: run test harnesses before +> executing live; own mistakes plainly and immediately. +> +> **Current state:** multi-tenant buildout complete through beta acceptance; validate.sh D-011 suite +> modular, gauntlet ALL GREEN (31). Vault is settled on `1.8/stable` (D-068: 1.16 ruled out; EOL is +> the open item). Reconciliation is done and clean. +> +> **Likely next work (confirm with Jesse):** batch-3 live validation of the D-011 checks (verify-live +> queue in the handoff) -> D-011 acceptance close; then project-completion (consolidate do-docs, flip +> repo private). D-068 EOL-vault modernization research is available when Jesse wants it. Do NOT ratify +> D-071 or move vault to 1.16 -- both are operator/decision-gated. +> +> Standing meta-instruction from Jesse: before answering, state what you need to know and any +> assumptions you'd otherwise make. Debate best-practice deviations with sourced rationale. + +## 9. Startup prompt -- CLAUDE CODE (jumphost stream) + +> You are resuming the Omega Cloud v1 / Baldurkeep project as the **Claude Code (jumphost) stream**, +> executing live on the cloud. A parallel **main claude.ai chat** stream runs concurrently -- +> coordinate, do not collide. +> +> **First, orient:** read `docs/handoff-20260705-open-items.md`, `docs/session-ledger.md`, and +> `docs/design-decisions.md`; run `bash scripts/ledger-scan.sh` and `bash scripts/ledger-scan.sh +> --fences`; reconcile. Repo is authoritative over memory. +> +> **Your lane / discipline:** in the session-ledger, edit only the `jumphost` fenced section; the +> shared section is append-only; do not touch the `main-chat` section. When you consume a +> D-/DOCFIX-/BUNDLEFIX- number, record it in its canonical home (design-decisions / changelog) FIRST, +> then regenerate the machine-derived block by running `ledger-scan` and pasting its output -- never +> hand-type numbers. Keep the ``/`` fences intact; pull before you edit. +> Committed files ASCII + LF. Gate destructive/irreversible steps individually; verify before mutate. +> +> **Current state:** last ops-update window (ops-update-20260705) is CLOSED; fleet current in-channel; +> juju 3.6.25. Vault is `1.8/stable` rev 714 (current in-channel). +> +> **Hard constraints:** +> - **Vault stays `1.8/stable`.** Do NOT move it to `1.16` (D-068: fundamentally different, +> incompatible charm). In-channel currency only. +> - **D-071 stays PROPOSED.** Do NOT ratify it -- that is operator authority, gated on the pre-DC-DC +> controller HA/backup planning session. Controller backups exist (create-backup on 3.6); the open +> half is single-controller HA. +> +> **Likely next work (confirm with Jesse):** support batch-3 live validation of the D-011 checks +> (verify-live queue in the handoff); controller update cadence per D-071 ONCE it is adopted (not yet). diff --git a/docs/archive/incident-20260712-opnsense-edge-boot-triplefault.md b/docs/archive/incident-20260712-opnsense-edge-boot-triplefault.md new file mode 100644 index 0000000..48060e6 --- /dev/null +++ b/docs/archive/incident-20260712-opnsense-edge-boot-triplefault.md @@ -0,0 +1,113 @@ +# Incident 2026-07-12: Office1 OPNsense edge triple-faults at the BTX loader + +> ## ROOT CAUSE FOUND 2026-07-12 (DOCFIX-188) -- READ THIS FIRST, THEN STOP +> +> **The guest had 2 MiB of RAM.** Not a CPU, nesting, machine-type, or console problem. +> Everything below this box is the ORIGINAL (wrong) investigation, retained for the record. +> **Do not work the "ranked next steps" -- they all chase the wrong layer.** +> +> `dmacvicar/libvirt` >= 0.9 changed `memory` from MiB (the 0.8-era meaning the modules were +> written against) to **raw libvirt units, defaulting to KiB**. With no `memory_unit`, +> `memory = 2048` rendered `2048` -> QEMU **`-m size=2048k`** = +> **2 MiB**. Measured end-to-end: module input -> tofu state -> domain XML -> `virsh dominfo` +> (`Max memory: 2048 KiB`) -> the live QEMU cmdline. +> +> `boot2` is tiny and fits in 2 MiB, so it echoes `/boot.config` and *then* triple-faults +> handing off to `/boot/loader`, which does not fit. **That is why the fault was +> deterministic at exactly 262 bytes and immune to every CPU/machine/disk/console change +> tried -- none of them touched the cause.** +> +> **Fix:** `memory_unit = "MiB"` on the `libvirt_domain` in all three VM modules +> (`opnsense-edge`, `cloudinit-vm`, `node-vm` -- the same defect was latent in all of them +> and would have broken every future VR1 VM). Guarded against recurrence by +> `scripts/opentofu-validate.sh` S1. See +> `docs/changelog-20260712-libvirt-memory-unit-rootcause.md`. +> +> **Lesson for the next incident:** a bootloader that dies at a fixed byte offset, immune to +> every knob you turn, is a *resource* problem, not a CPU-feature problem. The domain XML and +> the QEMU cmdline are ground truth -- read them before theorising about nested virt. +> +> **Status:** root cause fixed in repo; live boot verification is the remaining gated step. + +**Original status (SUPERSEDED):** OPEN. The Office1 OPNsense edge VM is built and *starts*, but OPNsense +triple-faults ~262 bytes into boot. Blocks the Office1 headend (router+DHCP for +`office1-local`) and therefore the Office1 NetBox VM. Written at a context limit as a +resume artifact -- a fresh session should read this + `docs/session-ledger.md` first. + +## Symptom (confirmed, reproducible) +- Domain `office1-opnsense` (libvirt, `qemu:///system`) enters `running (booted)` then + goes to `paused (unknown)`. `virsh resume` fails with *"cont: Resetting the Virtual + Machine is required"* -> the guest **triple-faulted**. +- Serial capture (`/var/lib/libvirt/vr1/staging/office1-opnsense-serial.log`, `root:600`, + read with `sudo cat`) contains exactly, every boot: + ``` + /boot.config: -S115200 -h -D + ``` + i.e. FreeBSD `boot2` echoes `/boot.config` (serial 115200, `-h` serial console, `-D` + dual console) and then triple-faults **handing off to the BTX `/boot/loader`**. The log + mtime updates on each boot (confirmed fresh, not stale); it is deterministically 262 bytes. + +## Environment +- **Double-nested virt:** OPNsense guest -> `vcloud` host (itself a VM: virtio NIC/disk) + -> outer hypervisor. Host CPU model reported by libvirt: **AMD `Opteron_G3`**. Nested + KVM on (`kvm_amd/parameters/nested = 1`). + - **CORRECTION 2026-07-12 (DOCFIX-189): the `Opteron_G3` label is a RED HERRING.** The + real CPU is an **AMD EPYC 9965 (Zen 5, family 26)** (`/proc/cpuinfo`). libvirt's CPU + model database does not know family 26, so it falls back to the oldest matching model + name. The prior session built a nested-virt/ancient-CPU theory partly on this artifact. + Read `/proc/cpuinfo`, not libvirt's model guess. +- Image: OPNsense **26.1 nano** amd64 (`scripts/opnsense-prep-image.sh 26.1`), prepped to + `/var/lib/libvirt/vr1/office1/opnsense-26.1-nano.qcow2` (11 GiB virtual). +- Module: `opentofu/modules/opnsense-edge`, instantiated as `module "office1_opnsense"` in + `opentofu/main.tf`. LAN=`office1-local` (`vtnet0`), WAN=`office1-wan` (`vtnet1`, NAT + `172.30.1.0/24`). Config ISO (real-ISP-router config, DOCFIX-185) at + `/var/lib/libvirt/vr1/staging/office1-opnsense-config.iso`. + +## What was tried -- ALL applied and verified in the domain XML, NONE resolved it +1. **Serial console added** (module gap -- nano is serial-only). Got the boot *to* the + loader stage (from a no-console early fault to the 262-byte `/boot.config` point). +2. **`machine` q35 -> i440fx** (`pc-i440fx-noble`, verified). No change. (Note: machine + + cpu are create-time; the provider does an in-place "change" that does NOT apply -- + must recreate the domain: `virsh destroy && virsh undefine`, then `tofu apply`.) +3. **Disk: COW overlay -> direct per-VM copy** of the nano (verified 11 GiB / 2.14 GiB + allocated, no backing). No change. +4. **CPU `host-passthrough`** (verified). No change. +5. **Disable AMD `svm`** (`` in XML -- the documented + AMD nested-virt fix, forum: `-cpu host,-svm`). **Still 262 bytes.** NOTE: the provider's + CPU-feature key is `features` (plural); `feature` validates but is silently dropped. + +## Ranked next steps for a fresh session +1. **Full forum CPU flag set**, not just `-svm`: add `+kvm_pv_eoi,+kvm_pv_unhalt` (and try + without `svm` masking too). Likely via `libvirt_domain.qemu_commandline` since the + provider's CPU `features` may not fully translate under `host-passthrough`. Confirm the + masking actually reaches the guest CPUID. +2. **Video device (`-D` dual console).** `/boot.config` requests dual console but the domain + has **no video/graphics device** -- `boot2`'s VGA init may fault. Add a `graphics` + + `video` (VGA/std) device and retry. (Cheap, plausible, not yet tried.) +3. **Memory 2 GB -> 4 GB** (below OPNsense's 3 GB min; guides use 4096). Recreate to apply. + - **CORRECTION (DOCFIX-189): this step was RIGHT FOR THE WRONG REASON, and its premise is + unsourced.** The memory *was* the problem -- but it was 2 MiB, not "2 GB but a bit + small", and no measurement was taken to notice that. The "3 GB min" figure is + UNVERIFIED (no source given); do not propagate it. If 2 GiB proves insufficient after + the real fix, check OPNsense 26.1's documented minimum before picking a number. +4. **UEFI boot (OVMF)** instead of legacy BIOS/BTX -- sidesteps BTX entirely, but the nano + is a BIOS/MBR image, so this needs care (or a different image build). +5. **Outer-hypervisor CPU:** under double-nesting, vcloud's exposed CPU (`Opteron_G3`) may + not provide what FreeBSD's BTX needs. Consider testing a minimal FreeBSD/OPNsense boot + directly on vcloud to isolate whether it's nesting-depth-specific. + +## Reproduce / operate +- Recreate + boot: `cd opentofu && source ~/vr1-stage1.env && virsh -c qemu:///system + destroy office1-opnsense; virsh -c qemu:///system undefine office1-opnsense; tofu apply`. +- Watch: `virsh domstate --reason office1-opnsense`; serial via `sudo cat` the log above. +- Provider CPU/serial schema is attribute-style + nested under `devices`/`cpu`; introspect + with `tofu providers schema -json` (see how `serials`/`features` were found this session). +- As-executed log: `~/as-executed/2026-07-12-dc-dc-phase1-office1.log`. +- Creds (jumphost-only, 0600, `~/vr1-office1-creds/`): SSH key + OPNsense root pw/hash. + **Operator: revoke the pasted NetBox token + these at close.** + +## Cross-refs +- `docs/changelog-20260712-opnsense-edge-boot-fixes.md` (DOCFIX-187, the module changes). +- `docs/changelog-20260712-office1-opnsense-edge-build.md` (DOCFIX-186, the build + the + apparmor/config-iso-staging findings). +- `docs/changelog-20260712-opnsense-edge-real-isp-router.md` (DOCFIX-185, config posture). diff --git a/docs/archive/model-a-fallback-plan.md b/docs/archive/model-a-fallback-plan.md new file mode 100644 index 0000000..fe843dd --- /dev/null +++ b/docs/archive/model-a-fallback-plan.md @@ -0,0 +1,106 @@ +# Model A fallback + revert plan (D-123) + +**Purpose.** The operator ruled **Model B** for D-123 (nodes nested inside `vvr1-dc0`, single-object +`virsh destroy` site-down) -- the heavier, higher-risk path (depth-4 nested virt, supersedes +D-103/D-114, ~416 GiB containment VM). This document preserves **Model A** as a fully-specified, +already-implemented fallback so that if Model B fails to deploy, we revert WITHOUT re-engineering. + +**Revert anchor (git).** Model A is not theoretical -- it is the CURRENTLY COMMITTED substrate. The +last commit before any Model B reshape is tagged **`model-a-fallback`** (re-cut 2026-07-16 onto +`114d392`, the **R-3-compliant** Model A layout with 4 storage nodes/DC -- the prior anchor at +`87a7a8a` was R-3-stale, flagged by the Model B design cross-check). To restore Model A: +`git checkout model-a-fallback -- opentofu/` +(or cherry-pick the substrate files), then re-run `bash scripts/opentofu-validate.sh`. No file needs +to be re-authored -- Model A already validates (`tofu validate` Success; 11/11 modules). + +--- + +## 1. Model A architecture (the as-built shape) + +Nodes are **vcloud-level libvirt siblings** of the headend -- NOT nested inside it. This is the +VR0-proven shape and the ADOPTED D-103/D-114 seam. + +``` +vcloud (host, L0) +|-- vvr1-dc0 MAAS rack headend (cloudinit-vm; D-124: 4 vCPU / 8192 MiB / 80 GiB; +| expose_nested_virt = false; legs = metal-admin + office1<->dc0 transit) +|-- vr1-dc0-control-01..03 node VMs (16/65536/150) \ +|-- vr1-dc0-compute-01..02 node VMs (12/49152/100) > vcloud-level siblings, on vr1_dc0_planes +|-- vr1-dc0-storage-01..04 node VMs (8/24576/550) / (4 storage/DC per R-3) +|-- vr1-dc0-edge opnsense (2/2048, 2-NIC: provider-public LAN + vr1-dc0-wan WAN) +|-- vr1-dc0-* planes 6 isolated-L2 libvirt networks (dc-planes) at vcloud level +|-- vr1-dc0-wan NAT /24 simulated ISP uplink (site-wan) +`-- mesh-vr1-dc0-office1 transit leg (office1 <-> dc0) + +nesting depth = 2 (vcloud -> node VM -> nova KVM guest) <- VR0-PROVEN +site-down = destroy the vr1-dc0-* domain GROUP (scripted, gated) +MAAS model = region on Office1 + rack (vvr1-dc0); maas-vm-host registers VCLOUD's virsh so + MAAS discovers the OpenTofu-created node domains (D-103/D-114 as-built) +``` + +## 2. The committed artifacts that embody Model A (revert targets) + +| Artifact | Model A content | +|---|---| +| `opentofu/main.tf` `module "vr1_dc0_node"` | `for_each = local.vr1_dc0_nodes`; created on the **vcloud** libvirt provider; attached to `module.vr1_dc0_planes` outputs (6 NICs, metal-admin first = PXE). | +| `opentofu/main.tf` `module "vvr1_dc0"` | `cloudinit-vm`, **4/8192/80** (D-124), `expose_nested_virt = false`, two legs (metal-admin + mesh transit). A small rack headend that holds NO nodes. | +| `opentofu/main.tf` `module "vr1_dc0_planes"` / `mesh_*` / `vr1_dc0_wan` | all created at **vcloud** level. | +| Step-9 `maas-vm-host` (deferred, DOCFIX-179) | registers **vcloud's** virsh to the DC's MAAS -> MAAS discovers the vcloud-level node domains. | +| `scripts/site-headend-install.sh --role rack` | installs the rack controller on `vvr1-dc0` (no LXD/compose in rack mode). | +| Site-down | a scripted group-destroy of the `vr1-dc0-*` domains (owned by `dc-dc-teardown-rollback.md`); NOT a single `virsh destroy`. | + +## 3. What Model B changes vs Model A (the delta to undo on revert) + +Reverting = undoing exactly these; nothing else moves. + +1. **Node placement:** B retargets `module "vr1_dc0_node"` (and the 6 planes + `vr1-dc0-wan`) from + vcloud's libvirt to **`vvr1-dc0`'s inner libvirt**. A restores them to vcloud level. +2. **Headend sizing:** B resizes `vvr1-dc0` from 4/8192/80 to ~416 GiB (must hold one DC's full node + fleet) and sets `expose_nested_virt = true`. A restores D-124's 4/8192/80. +3. **maas-vm-host target:** B registers `vvr1-dc0`'s inner virsh; A registers vcloud's virsh. +4. **Governance:** B supersedes D-103/D-114; A keeps them ADOPTED as-is. On revert, the D-103/D-114 + supersession is withdrawn. +5. **Nesting depth:** B = 4 (unproven); A = 2 (VR0-proven). +6. **Site-down primitive:** B = one `virsh destroy vvr1-dc0`; A = scripted group-destroy. +7. **Inner root:** B adds `opentofu/vr1-dc0-substrate/` (a new root dir + its own state) and a + `site-headend-install.sh --host-nodes` node-host bootstrap; A has neither. Revert removes them. +8. **Per-DC ISP egress (D-125 bridge-in):** because B nests `vr1-dc0-wan` inside `vvr1-dc0`, B adds a + vcloud-level ISP NAT (`module "vr1_dc0_uplink"`, `site-wan`, cidr `172.30.2.0/24`), a 2nd IP-less + uplink NIC + the `br-vr1-dc0-wan` netplan bridge on `vvr1-dc0`, the new `modules/wan-bridge`, and the + bootstrap's `--uplink-if`/`--wan-bridge` verify. **The ADDRESSING is identical to Model A** -- same + `172.30.2.0/24` (D-115), OPNsense WAN still `.2`; only the libvirt realization differs (bridge through + `vvr1-dc0` vs a direct vcloud-level NAT). **In Model A the extra plumbing does not exist** -- `vr1-dc0-wan` + is a vcloud-level NAT the vcloud-level edge attaches to directly, no uplink NIC, no bridge, no + `wan-bridge` module. Revert removes that plumbing; NO re-address is needed (the address never changed). + (No HELD gate here: the /24 is a ruled literal, not a tfvar.) + +## 4. Revert procedure (if Model B deployment fails) + +1. STOP -- do not attempt to fix Model B in place if nested-virt (depth-4) is the failure mode; that + is the known risk this fallback exists for. +2. `git checkout model-a-fallback -- opentofu/main.tf opentofu/variables.tf opentofu/modules/` + (restores the Model A substrate verbatim -- this also drops `module "vr1_dc0_uplink"`, since Model A + has no vcloud uplink), then `git rm -r opentofu/vr1-dc0-substrate` (the inner root does not exist in + Model A) and `git rm -r opentofu/modules/wan-bridge` (D-125, also absent in Model A), and revert the + `site-headend-install.sh` node-host mode (incl. the D-125 `--uplink-if`/`--wan-bridge` WAN-bridge + verify). The OPNsense WAN address is UNCHANGED (`.2` on `172.30.2.0/24`) -- nothing to restore. +3. `bash scripts/opentofu-validate.sh` -> expect 11/11 PASS (Model A already validates). +4. Re-instate D-103/D-114 as ADOPTED (they were only annotated superseded, not deleted -- see the + sweep's supersession notes; revert removes those annotations). +5. Restore D-124's rack sizing (4/8192/80) in tfvars/main.tf. +6. Adopt the scripted group-destroy site-down (the `vr1-dc0-*` group op) in place of the single-object + destroy. +7. Re-run the Layer-1 gate (`repo-lint`, `run-tests-all`, `tofu validate`) before proceeding. + +## 5. Failure signals that should trigger the revert + +- nova-compute guests fail to boot or are unusably slow at 3x-nested KVM (the depth-4 risk). +- `vvr1-dc0` cannot be allocated ~416 GiB on the host alongside the other layers. +- MAAS enrolment/commissioning breaks because the node domains are no longer vcloud-visible. +- `expose_nested_virt = true` on `vvr1-dc0` destabilises the headend/rack. + +--- + +*Model A remains the recommended engineering choice on delta/risk grounds; Model B is the +operator-ruled choice for its single-object site-down primitive. This plan makes the choice +reversible at low cost. Kept in sync with the D-123 sweep; if Model B changes, update section 3.* diff --git a/docs/archive/netbox-vip-queue.md b/docs/archive/netbox-vip-queue.md new file mode 100644 index 0000000..5017bd1 --- /dev/null +++ b/docs/archive/netbox-vip-queue.md @@ -0,0 +1,126 @@ +# Post-deployment NetBox VIP imports (queued from workstream 2) + +**Status:** Queued. To be imported after successful cloud deployment + validation, +once `netbox/ipv4-prefixes-import.py` engineer review unblocks the Provider /22 +prefix import. + +**Background:** Per D-010 (NetBox-upstream policy), IPAM entries should exist in +NetBox before being written into IaC. For v1 testcloud, this rule was relaxed +under workstream 2 (2026-05-22) to avoid blocking the rebuild on the engineer +review. VIPs were written into `bundle.yaml` directly. This document captures +the corresponding NetBox writes that need to happen post-deploy. + +**Scope:** v1 only (IPv4). v2 IPv6 VIPs are out of scope. + +--- + +## Provider prefix (parent -- gating) + +Before any IPAddress entries can be created, the parent prefix must exist: + +| Prefix | Site | Role | Status | +|---|---|---|---| +| `10.12.4.0/22` | VR0 DC0 | provider | Active | + +Created by: `netbox/ipv4-prefixes-import.py` (per D-010, gated on engineer review). + +--- + +## VIP IPAddress entries + +All entries under prefix `10.12.4.0/22`, tenant scope = VR0 DC0 Omega Cloud (or +appropriate testcloud tenant convention). + +| IP | Status | DNS name | Description | +|---|---|---|---| +| `10.12.4.224/22` | Active | `barbican.omega.dc0.vr0.cloud.neumatrix.local` | barbican API VIP -- Charmed OpenStack hacluster | +| `10.12.4.225/22` | Reserved | -- | RESERVED for ceph-radosgw HA VIP in v2 (workstream-2 decision; ceph-radosgw HA deferred to v2) | +| `10.12.4.226/22` | Active | `cinder.omega.dc0.vr0.cloud.neumatrix.local` | cinder API VIP -- Charmed OpenStack hacluster | +| `10.12.4.227/22` | Reserved | -- | RESERVED for designate VIP in v2 (per D-019; Designate deferred to v2) | +| `10.12.4.228/22` | Active | `glance.omega.dc0.vr0.cloud.neumatrix.local` | glance API VIP -- Charmed OpenStack hacluster | +| `10.12.4.229/22` | Active | `keystone.omega.dc0.vr0.cloud.neumatrix.local` | keystone API VIP -- Charmed OpenStack hacluster | +| `10.12.4.230/22` | Active | `magnum.omega.dc0.vr0.cloud.neumatrix.local` | magnum API VIP -- Charmed OpenStack hacluster | +| `10.12.4.231/22` | Active | `neutron.omega.dc0.vr0.cloud.neumatrix.local` | neutron-api API VIP -- Charmed OpenStack hacluster | +| `10.12.4.232/22` | Active | `nova.omega.dc0.vr0.cloud.neumatrix.local` | nova-cloud-controller API VIP -- Charmed OpenStack hacluster | +| `10.12.4.233/22` | Active | `octavia.omega.dc0.vr0.cloud.neumatrix.local` | octavia API VIP -- Charmed OpenStack hacluster | +| `10.12.4.234/22` | Active | `horizon.omega.dc0.vr0.cloud.neumatrix.local` | openstack-dashboard (Horizon) VIP -- Charmed OpenStack hacluster | +| `10.12.4.235/22` | Active | `placement.omega.dc0.vr0.cloud.neumatrix.local` | placement API VIP -- Charmed OpenStack hacluster | +| `10.12.4.236/22` | Active | `vault.omega.dc0.vr0.cloud.neumatrix.local` | vault VIP -- Charmed Vault hacluster (D-006) | + +**Notes:** + +- Mask is `/22` (the parent prefix mask), not `/32` -- NetBox convention for + endpoint IP addresses within a prefix. +- The Reserved slots at `.225` and `.227` document v2 intent without consuming + active allocations. When v2 work brings ceph-radosgw HA and Designate online, + those entries' Status flips Reserved -> Active and the bundle's `# v2-deferred:` + markers are uncommented. +- `nova-cloud-controller` charm -> DNS short name `nova` (catalog service name, + not charm name). +- `openstack-dashboard` charm -> DNS short name `horizon` (project name). +- `neutron-api` charm -> DNS short name `neutron`. + +--- + +## FIP pool -- for completeness (not part of workstream 2) + +Per D-003, the Provider /22 also carries the Neutron FIP pool. These are NOT +individual IPAddress entries; they're modeled as an IP Range under the prefix: + +| Range | Purpose | +|---|---| +| `10.12.4.10 - 10.12.4.223` | Neutron FIP pool (created by `ipv4-prefixes-import.py`) | +| `10.12.4.224 - 10.12.4.254` | API VIP pool (the 13 entries above + future) | + +Neutron `allocation_pools` for the provider subnet MUST exclude `.224-.254` -- +this is enforced in `runbooks/06-tenant-setup.md` (or wherever the provider +subnet is created). + +--- + +## Execution path (when unblocked) + +1. Confirm engineer review of `netbox/ipv4-prefixes-import.py` has signed off. +2. Run `netbox/ipv4-prefixes-import.py` -- creates the Provider /22 prefix + FIP + IP Range + API VIP IP Range. +3. Add the 13 IPAddress entries from the table above. Two paths: + - **Web UI:** Per-entry manual creation. Tedious but reviewable. + - **API/script:** Extend `ipv4-prefixes-import.py` with a VIP-addresses + section, OR write a separate `netbox/ipv4-vips-import.py` that reads + this document (or a YAML/CSV companion). Idempotent (skip-if-exists). +4. Sanity check: NetBox prefix view of `10.12.4.0/22` shows all 13 entries. +5. Cross-check: every active VIP in `bundle.yaml` has a matching Active + entry in NetBox; the Reserved entries at `.225` and `.227` have no + corresponding bundle entries (v2-deferred). + +--- + +## Change log + +| Date | Change | Reference | +|---|---|---| +| 2026-05-22 | Document created. 12 active VIP allocations queued + 1 v2-reserved slot. | Workstream 2 -- VIP allocation + hacluster activation | +| 2026-05-27 | Designate VIP at `.227` flipped Active -> Reserved per D-019 (Designate deferred to v2). Active count: 11; Reserved count: 2. | D-019 | + +--- + +## Management-plane reservations (MAAS-side; draft <- live, queued for NetBox) + +**Status:** Queued / draft. Observed live in MAAS during the 2026-06-11 rebuild; recorded here +draft <- live (live MAAS state is truth; this doc is the reconciliation target). NOT yet written +to NetBox -- D-010 (NetBox is unmutated until the IPAM design is confirmed). Purpose-pending: +the ranges were reserved on the mgmt plane but their per-IP role is not yet assigned. + +Two contiguous reservation ranges on the provider + metal segments (cf. D-003 -- the provider +network carries ext_net + API VIPs on one L2; D-010 -- NetBox-upstream policy): + +| Range | Segment | Status | Purpose | +|---|---|---|---| +| `10.12.4.101` - `10.12.4.110` | provider `10.12.4.0/22` | Reserved (MAAS) | mgmt-plane, role TBD -- reconcile before NetBox write | +| `10.12.8.101` - `10.12.8.110` | metal/internal (`10.12.8.x`; prefix TBC) | Reserved (MAAS) | mgmt-plane, role TBD -- reconcile before NetBox write | + +**To reconcile before NetBox import:** confirm each range's intended role (host mgmt, OOB, spare +pool, etc.) and the metal-segment prefix, then create the matching NetBox IPRange/IPAddress +entries with the resolved role + DNS convention. Do NOT mutate NetBox until the role assignment +is confirmed (D-010). The /22-mask convention (not /32) applies, as for the VIP entries above. + diff --git a/docs/archive/phase-00-maas-standup-notes.md b/docs/archive/phase-00-maas-standup-notes.md new file mode 100644 index 0000000..7e89dbb --- /dev/null +++ b/docs/archive/phase-00-maas-standup-notes.md @@ -0,0 +1,52 @@ +# phase-00 MAAS stand-up (D-058) -- notes + +`scripts/phase-00-maas-standup.sh` brings MAAS to the D-058 plane topology +idempotently. **Dry-run is the audit** (default): it resolves live ids BY +CIDR/name (PATTERN-1) and prints the plan, changing nothing. `--apply` executes. + +## Behavior per resource +- present and correct -> **SKIP** +- absent -> **CREATE** (fabric / VLAN / subnet / space / gateway / managed / dns / API-VIP reserve) +- present but bound to the wrong plane or wrong VID -> **DRIFT** (reported, never touched) + +A re-CIDR is destructive (MAAS cannot change a subnet CIDR in place), so it is +**out of scope by design**: the drift scan reports it as MIGRATE-NEEDED and the +script refuses to build onto a CIDR the wrong plane occupies. Verified idempotent +on `--apply` (zero mutations against an already-correct cloud). + +## D-058 target (what it stands up) +provider-public 10.12.4.0/22 (untagged, gw .4.1) | provider-vip 10.12.8.0/22 +(VID 104 on the provider fabric, gw .8.1, VIP band .8.2-.100) | metal-admin +10.12.12.0/22 (untagged, gw .12.1, VIP band .12.2-.100) | metal-internal +10.12.16.0/22 (VID 103 on the metal fabric, VIP band .16.2-.100) | data-tenant +10.12.20.0/22 | storage 10.12.32.0/22 | replication 10.12.36.0/22. Untagged base +planes are created first so their tagged siblings can ride the same fabric +(provider-public->provider-vip, metal-admin->metal-internal). + +## Single MAAS-address authority (D-058 consolidation) +- THIS script owns topology AND every reserved range: the API-VIP bands, the + Neutron FIP pool (10.12.5.0-.7.254 on provider-public), and the mgmt reserves. +- `phase-00-maas-carve.sh` is RETIRED -- its reserves are folded in here, and its + gated stale-range delete is subsumed by the teardown + re-CIDR step. +- It never deletes anything. + +## Relationship to provider-vip-standup.sh +This generalizes that script from one plane to the whole topology; provider-vip is +now just one row of the table. `provider-vip-standup.sh` remains the targeted +single-plane tool (add provider-vip to an already-D-058 cloud). Both source +`lib-net.sh`; no conflict. + +## Tests +`tests/phase-00-maas-standup/` -- fake `maas` + real jq, fixtures generated by +`make_fixtures.py`. Four scenarios, ALL PASS: fresh MAAS (full create plan), +D-058 done (all SKIP / zero WOULD), D-052 current (the three migrating planes +drift + refuse), wrong-VID (vid drift). Run: `bash tests/phase-00-maas-standup/run-tests.sh`. + +## The current-cloud gap (next deliverable) +The live cloud is D-052/053, so this stand-up will report metal-admin (.8), +metal-internal (.12), data-tenant (.16) as DRIFT -- those CIDRs are reassigned by +D-058. The destructive cutover (release/teardown so subnets have no links, delete +the old subnets in collision-safe order, then `--apply` to build the new scheme) +is a **separate gated step**, not this script. Sequence: teardown -> delete old +subnets -> `phase-00-maas-standup.sh --apply` -> `phase-00-maas-carve.sh` +(D-058) -> jumphost bridge re-IP (D-058 ordering trap) -> deploy. diff --git a/docs/archive/repo-lint-nextfree-bug-FINDING.md b/docs/archive/repo-lint-nextfree-bug-FINDING.md new file mode 100644 index 0000000..42723a1 --- /dev/null +++ b/docs/archive/repo-lint-nextfree-bug-FINDING.md @@ -0,0 +1,68 @@ +# FINDING: repo_lint.py L5 "next-free identifiers" is unreliable -- do not use it for numbering + +**Status:** DOCFIX candidate. Read-only analysis by the main-chat stream at repo HEAD +`59a7c73` (2026-07-06). No code changed. Code assigns the DOCFIX number when actioning +(re-run `ledger-scan` for next-free first; main-chat consumed none of these). + +**One-line:** `scripts/repo_lint.py` prints an `[info] L5 next-free identifiers` line that +is wrong on all three counters and disagrees with `ledger-scan`. Numbers must be taken +from `ledger-scan`, never from repo_lint. This is a live collision risk while the jumphost +stream is consuming DOCFIX numbers rapidly. + +## Evidence + +At this HEAD the two tools disagree: + +| counter | repo_lint L5 says | ledger-scan says (authoritative) | ground truth | +|---|---|---|---| +| D | next-free `D-075` | next-free `D-074` | highest header `D-073`; `D-074` exists only as a PROPOSAL mention | +| DOCFIX | next-free `DOCFIX-100` | next-free `DOCFIX-106` | highest consumed `DOCFIX-105`; `DOCFIX-106` is a next-free pointer | +| BUNDLEFIX | next-free `BUNDLEFIX-051` | next-free `BUNDLEFIX-012` | highest consumed `BUNDLEFIX-011`; `051` traces to a test fixture (below) | + +`ledger-scan` is correct on all three. Each repo_lint number is wrong via a DIFFERENT +defect, which is why this is worth fixing rather than tweaking. + +## Root cause -- three independent defects in the L5 next-free block + +The block (`repo_lint.py`, L5 identifier numbering) computes next-free as `max(seen)+1` +over `re.finditer(r"\b(D|DOCFIX|BUNDLEFIX)-(0\d{2})\b", ...)` across `all_text()`: + +1. **Band-limited regex `0\d{2}`.** It only matches identifiers `000`-`099`. It is blind to + `DOCFIX-100..106` (the current live range), so it reports `DOCFIX-100` when true next-free + is `106`. This is the same defect `ledger-scan` already carried and Code already fixed + (its regex is now `[0-9]{3}`); repo_lint still has the old band-limited form. +2. **Mention-counting, no header-authority, no pointer-exclusion.** It counts every + identifier OCCURRENCE, not definition-headers, and does not skip "Next-free:" pointer + lines. So a PROPOSAL reference inflates it -- the main-chat CIDR plan mentions `D-074`, + which is exactly why repo_lint reports `D-075`. `ledger-scan` uses design-decisions + headers for D and excludes `next-free` pointer lines. +3. **Scans the test tree.** `all_text()` includes `tests/`. `tests/ledger-scan/run-tests.sh` + contains a deliberately-high fixture line + (`Next-free: D-071, DOCFIX-099, BUNDLEFIX-050 ... must NOT inflate`) written to PROVE + `ledger-scan` ignores pointer lines. repo_lint ingests that fixture's `BUNDLEFIX-`050 + and reports `051` -- the precise trap the fixture was authored to catch. + +Correction of the record: an earlier main-chat note called `BUNDLEFIX-`051 "spurious." It is +(tokens above de-fanged 2026-07-06, jumphost stream: identifier-shaped tokens above the +real high-water mark in docs/ prose inflate the ledger-scan counters -- the standing rule +this very FINDING is about; the table rows survive because "next-free" lines are excluded.) +not -- it is `max+1` of the fixture token in defect 3. The imprecision is corrected here. + +## Fix options (Code's lane; not implemented here) + +**Recommended -- remove the next-free print from repo_lint entirely.** repo_lint's actual L5 +job is the duplicate-definition-heading collision guard; keep that. Next-free is +`ledger-scan`'s job, and having two tools derive it by different regexes is how they drifted +apart. Single source of truth (the same principle already floated for the machine-derived +ledger cache). Lowest risk, removes the drift surface. + +**Fallback -- if repo_lint must keep printing next-free, make it match `ledger-scan`:** +widen the regex to `[0-9]{3,}`; use design-decisions headers for D; exclude `next-free` +pointer lines; and exclude `tests/` from `all_text()` for identifier scanning. More code, +same answer -- prefer the removal. + +## Interim guidance (both streams, until fixed) + +Take D / DOCFIX / BUNDLEFIX next-free from `bash scripts/ledger-scan.sh` ONLY. Treat +repo_lint's `[info] L5 next-free` line as advisory-and-currently-wrong; it does not gate +(it is `[info]`, 0-fail), so it is safe to ignore for numbering. diff --git a/docs/archive/script-quality-findings-20260707.md b/docs/archive/script-quality-findings-20260707.md new file mode 100644 index 0000000..bcbc347 --- /dev/null +++ b/docs/archive/script-quality-findings-20260707.md @@ -0,0 +1,141 @@ +# Script-quality findings register -- read-only sweep, 2026-07-07 + +**What this is.** The deferred-findings register from the 2026-07-07 read-only +script-quality sweep, recorded so the items survive session compaction. Nothing +in this file has been executed; every fix listed here is a DOCFIX **candidate** +to be delivered later under the standard change-delivery loop (harness green, +repo-lint clean, changelog entry with revert). No identifier numbers are +assigned in this register -- the batch that fixes an item assigns its number +via the next-free rule at delivery time. + +**Sweep coverage.** 110 files reviewed (scripts/, tests/, .claude/hooks/). +Tools available on the review host: bash and python3 only; shellcheck and +pyflakes were absent, so findings come from manual read plus targeted fixture +probes, not static analysis. + +**Batch note.** The sweep's E/R-class items approved for immediate fix were +implemented in the fix batch of this date and are NOT listed here, except the +two previously-known items marked FIXED-THIS-BATCH below. + +## Deferred findings (R-class; fixes are DOCFIX candidates) + +### R5 -- lib-validate.sh: vr_json tempfile leak +`vr_json` (scripts/lib-validate.sh, ~line 52) creates `_VR_ERR="$(mktemp)"` on +every call and never removes it. `vr_err_tail` legitimately needs the file +after return, but nothing cleans it up afterwards: one leaked tempfile per +`vr_json` call, per check, per validation run. Fix shape: a single per-process +stderr file reused across calls, or an EXIT trap in the library's sourcing +contract. + +### R6 -- `| head -1` SIGPIPE / extract-then-check class in phase-05/06 +Live-pipe `openstack ... | head -1` captures that use the fragment without +whole-output validation: +- scripts/phase-05-amphora-pipeline.sh:62 (`image list --tag ... | head -1`) +- scripts/phase-05-amphora-pipeline.sh:66 (`image list --name ... | head -1`) +- scripts/phase-06-mgmt-vm.sh:106 (`port list --server ... | head -1`) +- scripts/phase-06-mgmt-vm.sh:108 (`floating ip list --port ... | head -1`) +Same class the house style forbids (SIGPIPE race on the producer; a partial +failure yields a plausible-looking fragment). Fix shape: capture whole output, +validate shape (uuid regex), then select the first row. + +### R7 -- phase-07-conductor-graft.sh: remote tempdir leak +The helm-install payload run on the conductor (~line 153) does +`D=$(mktemp -d); cd "$D"` and never removes the directory: one leaked tempdir +(plus tarball) on the remote unit per graft run. Fix shape: trap cleanup inside +the remote payload. + +### R9 -- cloud-assert.sh --capture: `--long` + silent empty images.json +Line ~184: `openstack image list --long -f json "$DIR/images.json" +2>/dev/null || true`. Two defects: (a) the deprecated `--long` flag only adds +stderr noise (and `2>/dev/null` then masks REAL errors, against the +stderr-separation rule); (b) a failed list silently commits an empty/missing +images.json into the captured BOM. Fix shape: drop `--long`, stderr to a +tempfile surfaced on failure, fail the capture loudly if the JSON is empty or +unparseable. + +### R11 -- harness tempfile leaks +Harnesses that mktemp without a trap cleanup (leak per run): +- tests/carve-host-interfaces/run-tests.sh +- tests/lib-validate/run-tests.sh (also amplifies R5: every vr_json case leaks) +- tests/phase-00-maas-standup/run-tests.sh +Fix shape: the standard `W=$(mktemp -d); trap 'rm -rf "$W"' EXIT` pattern the +other 34 harnesses already use. + +## Previously-known items -- FIXED-THIS-BATCH + +- **tenant-offboard.sh Phase B counter (FIXED-THIS-BATCH).** Logged in the + changelog (2026-07-06 addendum 28) as FOUND-NOT-FIXED: the residual-trust + sweep incremented SWEEP_FAIL inside a `printf | while` pipeline subshell, so + failed trust deletes could never raise exit 22. Fixed in this date's batch + (here-string loop; regression case in tests/tenant-offboard). +- **tenant-offboard.sh hardcoded auth URL (FIXED-THIS-BATCH).** The known + hardcoded-OS_AUTH_URL class (house rule 3): tenant_env now reads `auth_url=` + from the cred file with the literal demoted to a KEYSTONE_VIP-overridable + fallback. NOTE: the tenant-acceptance.sh foil_env sibling of this class + (addendum 28 FOUND-NOT-FIXED list) remains OPEN and is NOT covered by this + batch. + +## Consolidation proposals (S-class; S1-S8) + +Transferred verbatim from the sweep report at integration (2026-07-07); +effort estimates are the sweep's (S/M/L). Operator has NOT ruled on S1-S5/ +S7/S8 -- proposals only. + +- **S1 (M)** The unset-OS_*/source-env dance is hand-rolled ~15x across + tenant-onboard (8 subshells), tenant-offboard (2), tenant-acceptance (3), + d063-apply (1), tenant-assert (2) -- while scripts/lib-validate.sh already + ships tested vr_scrub_os / vr_admin_env / vr_tenant_env. Proposal: source + lib-validate.sh from the tenant-* family + d063-apply; add one + `vr_appcred_env ` for the app-cred variant. (The offboard tenant_env + auth-url bug was fixed narrowly in this batch; this proposal remains the + wider consolidation.) +- **S2 (S)** run() capture-then-test helper duplicated with eval-counter + variants in tenant-offboard.sh and d063-apply.sh vs the rc-returning + lib-validate run(). Consolidate; count failures at call sites. +- **S3 (S)** is_id() defined in 4 scripts (tenant-onboard, tenant-offboard, + d063-apply, tenant-assert); lib-validate already has vr_is_hex32. Ride + along with S1. +- **S4 (M)** The trust-list stdout-only capture + inline-python filter block + is near-identical 3x (tenant-offboard x2, tenant-acceptance x1) and is + inline python-in-bash beyond one-liner scale (house rule: .py files). One + fixture-tested scripts/trust_filter.py would serve all three. +- **S5 (S)** Kube-image UUID resolution block (json capture + jq kube filter + + tempfile + 36-hex validation) appears 3x inside tenant-onboard.sh. One + local function. +- **S6 (no action)** cap()/oerr() triplicated across clientdocs/scripts/ + {smoke-test,acceptance-run,ci-cleanup-sweep} -- CORRECT as-is: handover + artifacts must be standalone off-jumphost. Worth a header comment noting + the duplication is deliberate. +- **S7 (S)** Dead code: checks/d011-04 CF assigned never used; checks/d011-03 + vr_is_ipv4-or-VHOST first test subsumed by second. (The other two S7 items + -- offboard no-op is_id, validate.sh lost line-rewrite -- were FIXED in + this batch.) +- **S8 (S)** Move the KEYSTONE_VIP default and the D-003 FIP pool bounds into + lib-net.sh as tagged constants (the literal is duplicated in 4 files: + tenant-onboard, phase-03-admin-openrc, phase-06-kubeconfig-gate, + phase-06-k8s-bootstrap; FIP pool literals duplicated across the two + phase-04-network scripts; provider-bundle-check.py re-states plane CIDRs + as a second source of truth -- python cannot source bash, needs a + generated or parsed form). + +## Found during the fix batch (this date; logged, NOT fixed -- out of scope) + +- **tenant-offboard.sh Phase A cluster captures merge stderr.** The + tenant-scope `coe cluster list` captures (inventory ~line 160; delete loop + and wait loop in Phase A) still use `2>&1`: a stderr warning line would + word-split into `coe cluster delete` argv and keep the wait loop from ever + seeing empty (budget exit 21). Same class as the Phase C fix; was not in the + approved batch scope. The offboard harness deliberately excludes coe + commands from its stderr-noise scenario until this is fixed. +- **Merged-stderr `domain show`/`token issue` captures (fail-closed class).** + tenant-offboard.sh (~line 133, sweep-mode preamble) and tenant-onboard.sh + stage subshells validate `$(... 2>&1)` captures by shape/equality. Failure + direction is CLOSED (safe), but a benign stderr warning on a noisy CLI would + abort with a false precondition/auth failure. Cosmetic-risk cleanup + candidate, same stdout-only pattern. +- **tenant-offboard.sh Phase E0 app-cred inventory merges stderr.** Line ~267 + tests rc (failures route correctly), but a success-with-warning would + word-split the warning into the display-only cascade-inventory lines. + Display pollution only; no delete argv exposure. + +ASCII + LF. diff --git a/docs/archive/session-findings-2026-07-02.md b/docs/archive/session-findings-2026-07-02.md new file mode 100644 index 0000000..2805d16 --- /dev/null +++ b/docs/archive/session-findings-2026-07-02.md @@ -0,0 +1,77 @@ +# Session findings -- 2026-07-02 (multi-tenant tenant->cluster buildout) + +## Executive summary +The tenant IDENTITY/TRUST path is DONE and PROVEN. A tenant password identity now creates a Magnum +cluster through create_user (D-064), create_trust (D-065), and into certificate generation. Cluster +COMPLETION is blocked one step later by an OPERATOR-side Barbican/Vault substrate defect (D-067), +independent of the tenant model. The Option-3 tenant account model (D-066) is adopted and to be used +from the first tenant. Next session: fix D-067 (live), then full tenant buildout + tenant-facing tests. + +## The trust-blocker chain (how we got from "cluster 403s" to "done + one substrate bug") +1. D-064 (prior): create_user template fix unblocked trustee-user creation. +2. create_trust then 403'd for EVERY caller (admin included), even trustor==self via direct + `openstack trust create`. Root cause: base policy shipped identity:create_trust with the + non-resolving `user_id:%(trust.trustor_user_id)s` (Caracal populates target.trust.trustor_user_id). + -> D-065: override with the target-prefixed form keystone itself ships. PROVEN by toggling the + override off (still 403 -> base policy owns it) then on. +3. After D-065, create_trust via APP CRED still failed: keystone `_check_application_credential` + blocks trust creation from app-cred tokens "regardless of the unrestricted flag" (this build's + docstring). -> D-066: cluster-create MUST be PASSWORD auth; adopt Option-3 account split. + `allow_insecure_application_credential_trust_escalation` REJECTED (isolation). +4. Password create_trust PASSED. Cluster then failed at cert-gen -> Barbican 500 -> castellan + vault_key_manager -> Vault AppRole login rejected: "source address 10.12.8.176 unauthorized + through CIDR restrictions". -> D-067. + +## D-067 root cause (and a corrected mis-diagnosis) +barbican reaches Vault on the METAL-ADMIN plane (vault_url=10.12.8.190, egress 10.12.8.176). Vault's +barbican-vault AppRole binds the secret_id to the METAL-INTERNAL CIDR (where east-west service traffic +belongs, D-052/D-053). Off-plane source -> rejected. The bundle is CORRECT (vault/barbican/barbican-vault +all bind secrets endpoints to metal-internal, lines 130/667/700); the LIVE env drifted. Fix = live +rebind to metal-internal (gated, next session), NOT CIDR-widen. +CORRECTED: mid-session I hypothesized "secret_id TTL expiry". REFUTED -- `juju run vault/leader +refresh-secrets` rotated the secret_id (barbican.conf re-rendered, service restarted) and the login +STILL failed with the CIDR error. It was never expiry; it is plane/CIDR. + +## What is validated live (tenant acme) +- Manager persona self-service via CLI (create_project/user/grant) -- D-064 G3. PASS. +- Tenant isolation: anti-escalation (admin grant DENIED); cross-domain resource reads DENIED/hidden; + domain enumeration OWN-DOMAIN-ONLY (tighter than appendix-C's SCS worst-case -- appendix-C corrected). +- App-cred + keypair self-mint; tenant L3 (net/subnet/router/ext-gw, SNAT proven) by a non-admin + app-cred identity. +- Cluster template create (image by UUID -- name form has a quoting/derivation hazard). +- Cluster create through create_user + create_trust (password) into cert-gen. + +## Decisions logged +- D-066: Option-3 tenant accounts (domain-admin/cluster/svc); cluster-create requires password auth. +- D-067: barbican-vault -> Vault must use metal-internal (live drift; the cert-gen blocker). ADOPTED, fix pending. +- D-068: PROPOSED -- Vault substrate hardening (1.16 pin [bundle done], TLS, AppRole lifecycle). + +## Probe-discipline lessons (now runbook conventions -- these recurred and cost time) +1. Validate raw output WHOLE, never extract-then-check. A `tr -dc 0-9` MARK guard turned an error + string ("...10.12.8.30:17070...") into MARK=123101283017070 and passed. Use `case "$raw" in + ''|*[!0-9]*) fail;; *) ok;; esac`. +2. Whitelist-print secrets, never blacklist-redact. `approle_secret_id` leaked past a `secret`-keyed + redact (the key is *_secret_id*). Print only an allowlist of safe fields; never pipe secrets. +3. No `exit`/bare-`return` in interactive PASTE blocks (they escape to the login shell and logged the + operator out). Subshell-wrap `( ... )`. NOTE: executed .sh scripts may use exit normally. +4. Privileged reads over `juju ssh` use `sudo cat file | ...`, never `sudo cmd < file` (the redirect + runs UNPRIVILEGED -> Permission denied). +5. Use the deployment's DECLARED endpoint/scheme, not the conventional one (assumed Vault https; it + serves http -- every probe errored on scheme until corrected). +6. A parser that can print NOTHING has a silent third state -- read raw + self-report inputs (field + lengths, raw body) so a malformed-request 400 can't masquerade as an auth failure. + +## Roosevelt hardening backlog (from this session) +- D-067/D-068: metal-internal binding discipline for ALL vault-kv consumers; Vault 1.16 + TLS; + AppRole secret_id lifecycle (TTL/renewal + proactive auth health probe). +- Endpoint/credential "follow the topology" is now a recurring class (with D-057): consider stable + VIP/DNS endpoints for substrate services so leader/re-IP changes don't silently break consumers. + +## Next session plan +1. Repair live env: read-only binding diagnosis (`juju show-application vault barbican barbican-vault`, + spaces<->subnets), then GATED `juju bind` of the barbican<->Vault secrets path to metal-internal; + re-run refresh-secrets if needed; confirm barbican AppRole login HTTP 200 from metal-internal. +2. Re-run tenant cluster-create (acme, ${CLIENT}-cluster password) -> cert-gen clears -> watch to + CREATE_COMPLETE; capture the CAPO child-cred mint identity (confirms D-066). +3. Full tenant buildout via scripts/tenant-onboard.sh; then clean-room `beta` (zero admin fallback). +4. Tenant-facing tests: kubeconfig, nodes/CNI/CCM, a tenant LB, tenant isolation from a second tenant. diff --git a/docs/archive/stage3-adversarial-review-20260716.md b/docs/archive/stage3-adversarial-review-20260716.md new file mode 100644 index 0000000..f136d6a --- /dev/null +++ b/docs/archive/stage3-adversarial-review-20260716.md @@ -0,0 +1,348 @@ +# Stage-3 / vr1-dc0 pre-deployment adversarial review -- 2026-07-16 + +Adjudicated findings from a four-charter adversarial review of the Stage-3 DC-substrate +batch, run before `vr1-dc0` is deployed for the first time. READ-ONLY: nothing was applied, +mutated, or pushed. Every proposed fix is a PROPOSAL for the operator to gate individually; +none were executed. Fixes are staged for a post-acceptance sweep, not applied mid-review. + +Adjudicator: Code (main loop). Method: Layer-1 deterministic gate (facts) -> Code's own +grounding of the decision-coherence register from primary text + the session transcript -> +four independent charters (A1 author's-advocate, A2 prosecutor, A3 Roosevelt-hawk, A4 +drift-archaeologist) in Round-1-independent + Round-2-cross-examination -> Code adjudicates. +I am NOT a consensus engine: surviving dissent is recorded (section 4), not resolved away. + +--- + +## 1. Review surface + +- Repo: `/home/jessea123/openstack-caracal-dc-dc` (live jumphost clone). Branch + `dc-dc-stage3-phase2-dc-substrate`, frozen at HEAD **`87a7a8a`** (a WIP commit that froze the + two uncommitted plane-review ledger files on top of the batch commit **`a48a60f`**). +- Baseline `@{upstream}` = **`80502e9`**. Review range **`80502e9..87a7a8a`**. +- Immutable artifact: **`docs/stage3-review-base.patch`** (3795 lines, 26 files, +3057/-275). + Every finding cites against the repo files (line numbers as they stand at HEAD). +- IN scope: the D-121..D-124 rulings, the OpenTofu Stage-3 substrate (`modules/site-wan`, + `main.tf` vr1_dc0 section, the 8 node VMs, the `vvr1-dc0` rack, `variables.tf`), the two + NetBox importers + harnesses, `overlays/dc-ha-scaleup.yaml`, `scripts/site-headend-install.sh` + rack mode, the phase2 runbook rewrite, and the three changelogs. +- OUT of scope (operator-gated, correctly deferred): the live `tofu apply`, the NetBox + `--commit`, the rack install, and the Stage-4 `maas-vm-host` wiring (DOCFIX-179). +- **Undocumented-intent note (per the prompt's FIRST ACTION item 4):** one design input exists + as a claim but not as a file -- `scratchpad/optc-calc.py`, the whole-host capacity model that + D-121's Option-C sizing rests on, is NOT in the tree (see R3-F06). No other design element was + found to live only in prior-session discussion. +- **A4 charter note:** the A4 (drift-archaeologist) Round-1 pass died on an API stall mid-stream + and was re-run as a standalone backfill; because Round-2 had already executed, A4 did not + participate in cross-examination. Its lens (drift / status-table archaeology) was substantially + covered by Code's own decision-coherence grounding and by A1. [A4 backfill: INTEGRATION PENDING + -- see section 4a; this document is complete without it and will be annotated on its return.] + +--- + +## 2. Layer-1 facts (mechanically proven; agents debate meaning, not these) + +*Verbatim tool output for every row below is captured under +`scratchpad/L1/` (ledger-scan.txt, repo-lint.txt, gauntlet.txt, tofu-validate.txt, tofu-plan.txt, +ceph-optc-500.txt, state-addresses.txt, etc.); this table is the distilled register.* + +A2's external-authority citations (Ceph size=3-wants->=4-hosts guidance; MAAS rack-statelessness; +CIS-style ip_forward hardening) are reproduced as the charter reported them; the underlying technical +claims hold independently, but specific benchmark/control identifiers should be re-verified before +being quoted as authoritative. + +| id | fact | +|---|---| +| L1-01 | diffstat: 26 files, +3057/-275 vs `80502e9`. | +| L1-02 | `ledger-scan`: PROPOSED/OPEN = **D-068, D-071, D-115**. next-free D=125. Fences OK. SEC-010/011 OPEN. | +| L1-03 | `repo-lint`: 0 fail, 1 WARN (legacy non-ASCII in design-decisions.md D-001..018 only). | +| L1-04 | Byte hygiene: the reviewed batch added **0** non-ASCII and **0** CR bytes. Patch has 0 added non-ASCII lines. | +| L1-05 | Gauntlet: **ALL GREEN, 62 harnesses** (the prompt cited 60 -- a stale number, still green). | +| L1-06 | `tofu validate`: Success (OpenTofu v1.12.3, 11/11 modules). | +| L1-07 | `tofu plan`: **BLOCKED** -- 4 required vars unset/no-default (`vr1_dc0_rack_metal_admin_ip`, `_transit_ip`, `_transit_prefix`, `_transit_peer_ip`). Plan cannot render until office1-netbox assigns the rack transit/30 + rack IP. Hard pre-apply gate. | +| L1-08 | state<->repo: `terraform.tfstate` = 15 resources, all Office1/Stage-1/2; every state module still declared (no orphan); the Stage-3 substrate is repo-only = expected "to create". Nothing applied. | +| L1-09 | Ceph re-run for the RULED Option C (3 OSD hosts/DC @ 500Gi): **PASS, margin 6.48 TiB** (roomier than D-121's carried Option-B 5.31 TiB). D-121 validated Option B, not C. | +| L1-10 | Phase-5 drill Step 10.0(b) = `virsh shutdown ` per node (domain GROUP), not a single `virsh destroy`; Step 10 failover invokes **NO MAAS** (Ceph/Glance/Cinder/Neutron workload failover only). | +| L1-11 | Section-9 shim register = node-VM create (D-103), tc netem (D-100), single-unit **Juju** controllers (D-104), single-unit rbd-mirror (D-108). D-121 scales the OpenStack plane and excludes D-104 + keeps rbd-mirror at 1 (D-108). | +| L1-12 | `scratchpad/optc-calc.py` (D-121's whole-host validation basis) is **NOT tracked in git and ABSENT from disk**; the changelog claims "reproducible: python3 scratchpad/optc-calc.py" (false). | +| L1-13 | Transcript verification (full question text + answers read): the only AskUserQuestion in scope had two questions -- **"Node layout"** (operator answered **"Option C (3+2+3)"** -- an explicit design ruling) and **"Proceed"** (a DELIVERY-workflow question; options were "rewrite the runbook / bundle changelog+ledger / repo-lint / hand over"; operator answered **"Do it all now..."**). Neither vault backend (v-a/v-b) nor node containment (Model A/B) was ever a presented option; the operator's messages contain no vault/backend/unseal utterance anywhere. | + +--- + +## 3. Findings register (ranked; severity, evidence, violation, action, contested?) + +### R3-F01 -- BLOCKING -- Two ADOPTED rulings rest on inferred/agent-authored operator rulings (count = 2) + +The batch's single highest-value defect. An operator ruling is a VALUE; inferring it violates +hard-rule-2, and here two of the four decisions carry an inferred/agent-authored ADOPTED status. +**The count is 2, not 1** -- the four charters converged on 1 (D-123 only) because Code's shared +brief pre-asserted "D-121 is grounded via AskUserQuestion"; that biased them past the vault +sub-ruling. Code verified the second case directly against the transcript (L1-13). + +- **(a) D-121 vault-HA backend v-a -- the code-consequential case.** Status line + `design-decisions.md:3264`: "Vault-HA backend sub-ruling: RESOLVED = (v-a) ... operator ruling". + Its OWN body `:3346` says "Operator sub-ruling needed." Transcript (L1-13, full question text + read): vault backend was NEVER a presented choice -- the only AskUserQuestion asked "Node layout" + (answered Option C) and "Proceed" (a delivery-workflow question answered "Do it all now"); neither + is a vault ruling, and the operator uttered nothing about vault anywhere. So v-a is agent-authored, + with no operator value. Worse, it is ENCODED in committed code: + `overlays/dc-ha-scaleup.yaml:67-80` re-declares `vault-hacluster` and re-adds + `[vault:ha, vault-hacluster:ha]`, REVERSING BUNDLEFIX-002 (which de-HA'd vault) -- a real change + driven by an unratified ruling, not a status-quo path. +- **(b) D-123 Model A -- the record case.** Status `design-decisions.md:3454`: "ADOPTED Model A + ... read from 'Yes, fire off those tasks' in response to the A/B question; flag if B was + intended." The real utterance "Do it all now..." was the operator's answer to the AskUserQuestion + "Proceed" question -- a DELIVERY-workflow choice (L1-13), NOT an A/B answer; Model A vs B was never + a presented option. So the "A/B question" D-123 claims to read is one the operator was never asked. Corroborating inconsistency: the phase-2 runbook already labels D-123 **PROPOSED** + (`runbooks/dc-dc-phase2-tofu-dc-substrate.md:507` "D-123 (PROPOSED, recommend...)") while + design-decisions.md marks it ADOPTED -- the record disagrees with itself about D-123's ratification. + Refinement (A1, upheld in Round 2): Model A happens to equal the already-as-built + D-103/D-114 node-placement seam, so -- unlike vault -- **no NEW code rests on the Model-A + inference** (Model B would have been the change). The region+rack MAAS model and the D-124 rack + block (`main.tf:376-439`, "# D-124", explicitly ADOPTED) are independently authorized. +- Legitimately ADOPTED on explicit authority: D-121 node layout = Option C (AskUserQuestion + selection, `:3397`) + the 14-service scale-up ("this is the time we add the additional HA + nodes"); D-122 (the operator's explicit "Ruling 2" + "Routed by fabric switches"); D-124 + ("1. A 2. Confirmed", `:3513`). +- **Violates:** hard-rule-2 (no inferred value); hard-rule-1 (an inferred ruling drove the committed + vault overlay). **Action:** demote BOTH the D-121 vault-HA axis and the D-123 Model-A axis to + PROPOSED; put both one-line questions back to the operator ("vault HA backend: v-a MySQL-backed + vs v-b etcd/Raft?" and "node containment: Model A vcloud-level group vs Model B nested-in-VM?"). + D-121's "ADOPTED IN PART" then means what it says. Region+rack, Option C, and D-124 stand. +- **Contested (severity):** A1 argues MAJOR-not-BLOCKING for D-123 -- a status LABEL is not a + value "entering a command", nothing is applied (L1-08), and apply is independently gated + (L1-07). Adjudication: kept BLOCKING because (i) the prompt frames C2 as BLOCKING by design, + (ii) the vault case DID drive committed code, and (iii) "ADOPTED" is the authority a downstream + agent acts on. The dissent is real and recorded (section 4): as a matter of LIVE mutation risk, + nothing is currently blocked; as a matter of decision-record integrity, both must be re-ruled + before the record is trusted or apply proceeds. + +### R3-F02 -- MAJOR -- SEC-010: the metal-admin DC-LOCAL invariant is a ledger promise, not a committed artifact or a gate + +Upheld under cross-examination (A3-F6 survived A2's "distro default is fine" refutation attempt). +`main.tf` `vvr1_dc0` cloud-init pins static IPs + one route but has NO `net.ipv4.ip_forward=0` +sysctl and NO host firewall on the transit leg. The rack straddles metal-admin (10.12.8.0/22, +DC-local per D-052 `:767` / D-100 `:1946`) and the office1<->dc0 transit (crosses fiber), so the +"never crosses the fiber" invariant is preserved ONLY by Ubuntu's distro default -- which the +deferred MAAS-rack snap install could flip. SEC-010 records this "Close BEFORE tofu apply", but a +grep of `scripts/` and the phase2 runbook finds NO mechanical gate; the only apply-block (L1-07) +is an addressing gate, not security. **Violates:** D-052/D-100 (and general host-hardening guidance +-- CIS-style benchmarks recommend `ip_forward=0` on non-router multi-homed hosts; verify the exact +control ID before quoting it). +**Action:** add `ip_forward=0` (v4+v6) + an nft/ufw transit-leg pin to the committed rack +cloud-init as the ARTIFACT; wire a mechanical pre-apply gate (not a ledger note); carry the same +pin onto `voffice1` when its transit leg is wired (currently single-homed, `main.tf:184`). The pin +is free -- a MAAS rack proxies at the application layer and needs no kernel forwarding. + +### R3-F03 -- MAJOR -- D-107's DR-independence claim is false as written (and false at Roosevelt too) + +D-107 `:2088`: the per-DC mirror exists "so a DC can redeploy independently even if Office1 or the +peer DC is down -- a DR requirement the drill exercises." Under the RULED region-on-Office1 + +rack-per-DC model (D-123), a MAAS rack is stateless and cannot commission/deploy/power without the +region, so an Office1 outage removes reprovisioning from BOTH DCs. And L1-10: the Phase-5 drill +invokes NO MAAS -- it exercises workload failover, not redeploy. Two clauses fail. +- **Nuance (A2, Round 2):** the claim says "Office1 OR the peer DC is down." The peer-DC-down half + is VALID (region up, mirror up -> the DC can redeploy). Scope the defect to the **Office1-outage** + sub-case, not the whole claim. +- **Elevation (A2, Round 2):** buildout-design `:157` mandates single-region PERMANENTLY, so this + is unachievable at ROOSEVELT too -- not stale VR1 wording. A standby-region-per-DC mitigation + would CONTRADICT the buildout's single-region decision. +- **Violates:** D-107 (its DR-independence + "drill exercises" clauses). **Action (option a):** + amend D-107 to scope DR-independence to the in-DC artifact mirror + workload failover; add a + WRITTEN, PERMANENT acceptance that node provisioning is unavailable during an Office1-region + outage (day-1/2 dependency, not runtime -- running DC clouds are unaffected); delete/correct "a + DR requirement the drill exercises". Do NOT reopen the topology. + +### R3-F04 -- MAJOR -- D-122's MAAS-controller bullet and single-object site-down wording were never superseded + +- **C1:** D-122 `:3441` "Each site runs its own MAAS controller (as voffice1 does)" -- a full + region+rack -- was superseded by D-123's region-on-Office1 + rack-per-DC (operator-confirmed; + matches buildout `:157`) but never marked. A downstream agent quoted the stale bullet back as if + live. **Action:** append a verbatim supersession note pointing to D-123. +- **C3:** D-122 `:3417` "virsh destroy against a single object" is void for DCs -- under + Model A the containment VM does not contain the DC nodes (D-123 admits this `:3482`), and the + Phase-5 drill already uses per-node group shutdown (L1-10), so the drill does NOT break. The + stale claim also leaked into `dc-dc-deployment-workflow.md:86`. **Action:** amend D-122 AND + workflow:86 to "destroy the vr1-dc0-* domain GROUP" for DCs (single-object destroy stays literal + for Office1); note `vvr1-dc0` at a DC is a MAAS rack headend, not a D-114 containment VM (rack + mode runs no LXD/compose). +- **Violates:** D-123 (supersedes both). Topology operator-confirmed; record fix only. + +### R3-F05 -- MAJOR -- D-107 vs D-123 "node artifacts" wording (C4): reconcilable, but D-107 must be scoped + +D-107 `:2087-89` ("no node artifacts served from Office1"; "images including amphora ONLY from an +in-DC mirror") reads against D-123 `:3459` (rack "PROXIES OS images from the Office1 region"). The +counter-reading HOLDS: the intra-MAAS region->rack commission/deploy image channel is architecturally +distinct from D-107's supply-chain mirror (apt/snap/charmhub/registry/amphora served at/after juju +deploy). Both stand once D-107 is scoped. **Violates:** NONE (a wording gap in D-107). **Action:** +amend D-107 to scope "node artifacts" to the supply-chain classes, explicitly excluding the +intra-MAAS provisioning-image proxy D-123 routes region->rack. Pairs with R3-F03. + +### R3-F06 -- MAJOR -- D-121's Option-C sizing rests on an uncommitted, now-absent model; the "reproducible" claim is false + +`scratchpad/optc-calc.py` -- the whole-host validation that grounds Option C's 222 vCPU / 790 GiB / +77%-RAM ruling -- is NOT tracked and is ABSENT from disk (L1-12), yet the changelog claims +"Resource model reproducible: python3 scratchpad/optc-calc.py". **Violates:** repo-authoritative +discipline (CLAUDE.md); D-121 record integrity. **Contested (severity, A1 Round 2):** the Option-C +ruling is an explicit operator selection (not gated on the calculator), and fit is independently +checkable -- node literals are committed in `main.tf` locals and the disk dimension re-ran from the +COMMITTED `dc-dc-ceph-disk-budget.sh` (L1-09 PASS). So the ruling is not "unverifiable"; the defect +is reproducibility hygiene. **Action:** promote `optc-calc.py` into a committed, harnessed +calculator (already logged as a follow-up) and re-run for Option C before the sizing is treated as +measured; correct the changelog's "reproducible" claim until it is. + +### R3-F07 -- MAJOR -- D-121 records an Option-B disk validation for the ruled Option-C layout + +D-121 `:3382` records the disk-budget PASS for "Option B 4+4" while the RULED layout is Option C +(3 storage/DC). Code re-ran for Option C (L1-09): PASS, 6.48 TiB margin -- capacity is fine, so this +is a RECORD defect, not a capacity one. **Action:** re-record the Option-C validation (cite L1-09). + +### R3-F08 -- MINOR -- Option C's 3 OSD hosts at size=3 carry zero Ceph rebuild headroom (unrecorded) + +D-121 flagged "size=3 has ZERO rebuild headroom" for the REJECTED Option A `:3358` but not for the +ADOPTED Option C, whose 3 storage hosts have the identical property. **Walked back from MAJOR in +Round 2** (A1/A2/A3 concur): with Charmed default min_size=2 the pool keeps SERVING on one host +loss (the availability drill is clean), ceph-mon=3 sits on the CONTROL nodes (quorum untouched), +and only the RE-REPLICATION sub-case lacks headroom -- an inherent size=3 economy that resolves at +Roosevelt (>=4 storage hosts). **Action:** add an accepted-risk note to the Option-C record (a +storage-node-loss drill will show degraded-not-self-healing; Roosevelt remedy = >=4 storage/DC). + +### R3-F09 -- MINOR -- D-115 reads PROPOSED to `ledger-scan` though it is ADOPTED by amendment (L1) + +D-115's primary Status line retains the substring "Originally PROPOSED/OPEN", which trips +`ledger-scan`'s regex (L1-02), while the decision IS ratified by its 2026-07-13 amendment +(`:2727`). So the office carve import + PR #1 merge rest on a legitimately ADOPTED decision -- NOT +an unratified-executed one. **Action:** move "Originally PROPOSED/OPEN" out of the Status line (or +harden `ledger-scan` to ignore an "originally/was PROPOSED" clause when ADOPTED is present). Sweep +D-121..D-124 Status lines for the same machine-vs-human record divergence. + +### R3-F10 -- MINOR -- SEC-011 (node least-connectivity) is CONTESTED on Roosevelt-fidelity grounds + +Code's own prior SEC-011 (storage nodes get provider-public + data-tenant legs they never bind) is +challenged by A2 (Round 2, REFUTED) and A1 (WEAKENED): (1) in VR1 every plane is isolated L2 with +no gateway/route out, so an unbound vNIC has NO exploitable reachability -- "attack surface" is nil +today; (2) pruning would OVERRIDE the ADOPTED D-122 6-NIC ruling without a superseding decision AND +INCREASE delta to Roosevelt, where D-052's planes are VLAN-trunked on bonded NICs (ALL planes +present at every node; netplan/MAAS decides L3 binding), so uniform 6 isolated-L2 vNICs is the more +faithful model. **Adjudication:** downgrade SEC-011 from "pre-apply hardening" to an OBSERVATION / +operator's call; recommend amending the SEC-011 ledger row to record the Roosevelt-delta caveat. +Surviving dissent (section 4). + +### R3-F11 -- MINOR -- Stale "DC1/DC2 + supernet unassigned" prose in the workflow doc (L3) + +`dc-dc-deployment-workflow.md` Authoring-status cell still reads "DC1-first; DC2 hard-gated (D-101 +supernet unassigned)" -- superseded by D-119 (code is vr1-dc0/vr1-dc1) and D-115 (vr1-dc1 supernet += 10.12.64.0/19 assigned). No gate depends on the stale reading (the second DC is sequenced out +regardless). **Action:** update the cell; drop the "unassigned" clause. + +### R3-F12 -- OBSERVATION -- Transient Octavia N+1 amphora failover headroom is not modeled (C7b) + +The whole-host validation is a steady-state sum; it does not visibly reserve the transient N+1 +amphora placement headroom Octavia STANDALONE failover needs (a hard-won VR0 finding), and the +Phase-5 drill exercises failover on two clouds sharing one host. Mitigated in this single-host sim +(a hard-downed DC frees its RAM for the survivor) and unverifiable until R3-F06's model is committed. +**Action:** when `optc-calc.py` lands, add a line on per-cloud transient amphora headroom (or state +it is a Roosevelt-only concern for this sim). + +### R3-F13 -- OBSERVATION -- D-122 "6 NICs = baremetal-matched" is imprecise; Section-9 could add a cross-ref + +(i) D-122 sizes nodes at "6 NICs, one per plane" but the Roosevelt realization (D-052) is untagged +metal-admin + tagged VLAN subinterfaces trunked on bonded NICs, not 6 discrete physical NICs -- L2 +isolation is behaviorally equivalent; a characterization nuance, no code change. (ii) A3 argued +Section 9 omits the site-wan simulated-ISP and the virsh-destroy DR fault-injection; largely REFUTED +in Round 2 (both have Roosevelt analogs -- a real circuit, a real facility-down drill -- so neither +meets the register's "no analog" bar; the register is a build-step register and virsh-destroy is an +operational step). At most add a clarifying cross-reference for the containment/virsh-destroy sim +vehicle. Not a material incompleteness. + +--- + +## 4. Minority report (surviving dissent -- recorded, not resolved away) + +- **R3-F01 severity (dissenter: A1).** A1 holds D-123's inferred-status is MAJOR ledger-hygiene, not + deploy-BLOCKING: a status label is not a value "entering a command" (hard-rule-2's literal scope), + nothing is applied (L1-08), and apply is independently gated (L1-07). A1 further showed (Round 2, + refuting A3-F1) that Model A = the already-as-built D-103/D-114 path, so no NEW code rests on that + specific inference. Code's adjudication keeps R3-F01 BLOCKING on decision-record + governance + grounds (and because the VAULT half DID drive committed code), but records A1's point: as pure + live-mutation risk, nothing is currently blocked. +- **R3-F10 SEC-011 (dissenter: A2).** A2 refutes the "attack surface" framing outright: isolated-L2 + planes have no exploitable reachability, and pruning increases Roosevelt delta. Code accepts the + Roosevelt-delta point and downgrades SEC-011 accordingly, but records that A2 would go further and + strike the finding as a non-issue; Code retains it as an OBSERVATION worth the operator's note. +- **R3-F08 Ceph severity (dissenters: A1, A2, A3 concur down; A2-R1 held MAJOR).** A2's Round-1 + MAJOR ("opposite of the clean drill C is sold on") did not survive its own and others' Round-2 + scrutiny (min_size=2 keeps serving; mon on control nodes). Recorded as the walk-back it was; + final severity MINOR. + +**Convergence check (mandated).** The three live charters converged on the C-register verdicts -- +expected, because those items are text-provable quotations, not judgment calls. But they did NOT +fully converge: each surfaced a distinct high-value finding (A2: external CIS/Ceph/MAAS citations; +A3: the Section-9 completeness challenge + the baremetal NIC-realization nuance; A1: the uncommitted +`optc-calc.py`), and Round-2 produced genuine dissent (above). So the charters separated adequately. +The one place convergence WAS a failure mode: all three reported inferred-ruling **count = 1** +because Code's shared brief pre-asserted D-121's grounding -- a demonstration that shared priors +create shared blind spots. Code's independent transcript check (L1-13) corrected the count to 2. + +## 4a. A4 backfill integration (returned; corroborates) + +The A4 drift-archaeologist backfill returned and **independently reached inferred-ruling count = 2** +(D-121 vault-`(v-a)` + D-123 Model A), corroborating R3-F01 by the harder path: A4 was tasked to +scrutinize the vault Status-vs-body contradiction directly, and reached count = 2 WITHOUT the +biasing "D-121 is grounded" hint the three Round-1 charters received. This confirms the section-4 +convergence diagnosis -- the count-1 result was a shared-prior artifact, not a real ceiling. A4 did +NOT participate in Round-2 cross-examination (its Round-1 died on an API stall; Round-2 had already +run). Integration -- with one A4 claim verified-and-rejected (I check agent citations, not +rubber-stamp them): +- **A4's count = 2 corroborated.** A4's F1 (D-123 Model A) + F2 (D-121 vault, Status `:3264` vs + body `:3346`, no vault utterance in the batch changelog) independently match R3-F01. +- **A4's F7 REJECTED on direct check.** A4 claimed the phase-2 exit gate cites the stale D-122 + "each site runs its own controller" bullet (`runbooks/dc-dc-phase2-tofu-dc-substrate.md:636`). + Verified: the runbook does NOT -- at `:507` it cites "D-123 (PROPOSED, recommend...)" and the exit + gate (`:630-635`) uses the rack-controller-per-DC (D-123) model. A4 miscited. NOTABLE side effect: + the runbook already labels D-123 **PROPOSED**, which CONTRADICTS design-decisions.md's ADOPTED + status and independently supports R3-F01 (the D-123 record is internally inconsistent about its own + ratification). Added to R3-F01's evidence. +- A4 otherwise agrees across C1-C7 and L1-L5 (it frames C4 as "REFUTED -- both stand", the same + disposition as R3-F05's "reconcilable, scope D-107"), and adds clean negative findings: no false + DONE/as-executed claims, state reconciles with zero orphans, byte/lint hygiene clean. + +--- + +## 5. Leads register + +| Lead | Status | Evidence | +|---|---|---| +| **L1** (D-115 PROPOSED yet executed+merged) | **CONFIRMED as record-staleness, REFUTED as governance breach** | D-115 ADOPTED by amendment `:2727`; `ledger-scan` false-positive on the "Originally PROPOSED/OPEN" substring (R3-F09). The import + PR #1 rest on a ratified decision. | +| **L2** (D-071 controller-HA gates Stage 3) | **REFUTED as a Stage-3 blocker** | D-104 (`:2031-38`, ADOPTED) dispositions the controller-topology question for VR1 (single-unit per DC, HA deferred to Roosevelt) and is the entry the DC-DC phase was gated on; it references D-071 without amending it. D-071 (patch cadence) is Roosevelt-scoped. Annotate the ledger note (R3-F... minor). | +| **L3** (D-117-class naming drift) | **REFUTED** | Code is unambiguous (vr1-dc0/vr1-dc1, D-119); the double-namespace is RECORDED (D-117 + amendments). Only residual is stale prose (R3-F11); no gate depends on the ambiguous reading. doc-"DC1" = vr1-dc0. | +| **L4** (GUA/ULA IPAM reconciliation) | **REFUTED (one line)** | Stage 3 is isolated-L2 substrate + node/edge/rack VMs; it instantiates no tenant L3 addressing. The GUA/ULA reconciliation is a Stage-5 Neutron concern. Stage 3 does NOT depend on it. | +| **L5** (exit gate "conditionally met at best") | **CONFIRMED -- still conditional, honestly HELD (not run through)** | Node sizing ruled (D-121, but via the uncommitted model, R3-F06); edge sizing carried from applied office1_opnsense (measured basis, not a DC-edge boot measurement); netem still commented/unparameterized (unruled D-100 sub-item); the interface-naming boot measurement is a runtime TODO; `tofu plan` BLOCKED on 4 rack vars (L1-07). netem is correctly held, not fabricated. | + +--- + +## 6. Go / No-Go on Stage 3 + +**NO-GO for `tofu apply` as it stands** -- on hard-gate grounds, NOT because the substrate design is +unsound (it is sound: planes are isolated L2, state reconciles with no orphans, gauntlet green, +Option C fits with margin, validate passes). + +Blocking conditions to clear (each operator-gated, applied in a post-acceptance sweep): +1. **Re-rule the two inferred axes (R3-F01):** D-121 vault-HA (v-a vs v-b) and D-123 node + containment (Model A vs B). Both are one-line operator questions. Until then the vault overlay + and the Model-A record are unratified. +2. **Ship the SEC-010 artifact + mechanical pre-apply gate (R3-F02).** DC-LOCAL must be enforced, + not promised. +3. **Assign the rack transit/30 + rack IP in office1-netbox and feed the 4 tfvars (L1-07)** so a + plan can render and be reconciled against state. Wire `voffice1`'s transit leg (or gate the rack + route) so the route peer exists. + +Conditional-GO once 1-3 are met AND the record amendments (R3-F03..F07, F09, F11) are applied and +`optc-calc.py` is committed + re-run for Option C (R3-F06). The record corrections do not block the +substrate build; they block treating the DECISIONS as coherent, which is the operator's stated +concern. The MINOR/OBSERVATION items (F08, F10, F12, F13) are for the operator's judgment and need +not gate the sweep. + +--- + +*Prepared read-only; not committed, not merged, not pushed. Presented for operator review. Fixes +staged for an individually-gated post-acceptance sweep.* diff --git a/docs/archive/tenant-cidr-overlap-correction-PLAN.md b/docs/archive/tenant-cidr-overlap-correction-PLAN.md new file mode 100644 index 0000000..ee6e7a0 --- /dev/null +++ b/docs/archive/tenant-cidr-overlap-correction-PLAN.md @@ -0,0 +1,154 @@ +# Tenant CIDR overlap correction -- work order / change plan + +**Status:** DRAFT for the Code (jumphost) stream + operator ruling. Authored by the +main-chat stream at repo HEAD `59a7c73` (2026-07-06). Read-only analysis; no live +mutation and no committed-surface edit performed by main-chat. + +**Proposes:** decision **D-074** (PROPOSED; amends D-016). D-074 is next-free per +`bash scripts/ledger-scan.sh` at this HEAD. NOTE: D-073 and DOCFIX up to 105 were +consumed the same day by the jumphost stream -- **Code must re-run the scan and +re-confirm next-free before inserting the decision header.** Do not hand-type it. + +--- + +## 1. Finding (what is actually wrong) + +The stage-4 tenant CIDR guard in `scripts/tenant-onboard.sh` **over-enforces relative +to the onboarding contract it is supposed to implement**, and separately has an +overlap-detection hardening gap. + +- **Contract intent** (`docs/tenant-onboarding-contract.md` sec. 3, item 4): a tenant + CIDR must be RFC1918 and non-colliding with **(a) our allocations** and **(b) the + client's own on-prem/VPN ranges they may later interconnect**. The contract does + **not** require non-collision with *other tenants*. +- **Guard behaviour** (`tenant-onboard.sh` stage 4 -- the `grep -qw "$TENANT_CIDR"` + check run under `admin_env` against a cloud-wide `openstack subnet list`): it dies on + **any** exact subnet-string match cloud-wide, which includes other tenants' subnets. + So it forbids tenant-vs-tenant exact overlap that the contract never intended to + forbid. +- **Independent hardening gap:** the guard is exact-string (`grep -qw`), so it MISSES + partial/containing overlaps (e.g. a requested `10.20.20.128/25` against an existing + `10.20.20.0/24`). It is therefore simultaneously *too strict* on exact tenant matches + and *too loose* on partial overlaps against ranges we must actually protect. + +**Origin of the uniqueness behaviour (and a correction of the record):** it is an IPAM +artifact of **D-016**, which models a single `10.20.0.0/16` tenant pool and carves +non-overlapping /24s from it. It is **not** a dataplane requirement. An earlier +main-chat claim that the uniqueness rule was needed for "CAPI/Magnum reachability" was +wrong and is retracted here: per **D-035**, `capi-mgmt-v2` is single-homed on its own +`10.20.0.0/24` and reaches workload clusters via their API-LB floating IPs and OpenStack +via the API VIPs -- it never routes into other tenants' `10.20.x` space, so overlapping +tenant CIDRs cannot touch the management path. + +## 2. Proposed decision -- D-074 (PROPOSED; amends D-016) + +Tenant private (overlay) CIDRs are **tenant-chosen and MAY overlap across tenants.** +Non-collision is enforced **only** against a reserved-ranges registry (operator/infra +allocations, section 3). NetBox (the IPAM apex) tracks the reserved/routed space; tenant +overlay space leaves NetBox uniqueness tracking. + +Rationale: this aligns the implementation with the onboarding contract's already-stated +intent, and with the hard-isolation (SCS Domain Manager) persona. Overlap-allowed tenant +CIDRs is the standard isolated-multi-tenant model -- it is why Neutron ships +`allow_overlapping_ips=true` by default and why public-cloud VPCs let tenants pick +arbitrary RFC1918. Globally-unique-per-tenant is the enterprise-private-cloud pattern +(shared routed fabric, interoperating internal units), which is the opposite of what this +cloud is. + +Forward-only: existing tenants are **not** re-CIDR'd (northwind stays `10.20.20.0/24`; +design-decisions forbids transient-overlap re-CIDR). + +**PROPOSED means the operator has not ruled.** Nothing in sections 3-4 is implemented +until D-074 is ADOPTED. + +## 3. Reserved-ranges registry (the ONLY thing the guard rejects) + +Non-collision is enforced against these, using real subnet-overlap math (containment in +either direction), not string equality: + +| plane / allocation | CIDR (confirm live before wiring) | +|---|---| +| provider-public (incl. API VIPs + `provider-ext` FIP pool) | `10.12.4.0/22` | +| metal-admin | `10.12.8.0/22` | +| metal-internal | `10.12.12.0/22` | +| data-tenant | `10.12.16.0/22` | +| storage | `10.12.32.0/22` | +| replication | `10.12.36.0/22` | +| capi-mgmt tenant network (D-035) | `10.20.0.0/24` -- OPERATOR DECISION: keep reserved? (recommend YES) | +| metadata service | `169.254.169.254/32` (link-local; not RFC1918; low concern) | + +**These values were read from `docs/maas-as-built-reference.md` by main-chat.** Code MUST +re-derive them from the authoritative live/modeled source (pre-flight `PLANE_CIDRS` / +`scripts/lib-net.sh` / NetBox) before wiring the guard -- do not hardcode from this doc +(hard rules 2 and 3: dynamic lookup, no inferred values). Centralize the list in +`lib-net.sh` keyed by plane name, not as scattered literals. + +## 4. Gated correction procedure (harness-first; each phase gates the next) + +### Phase 0 -- read-only verification (MUST pass before any mutation) +- **0.1 Confirm `allow_overlapping_ips`.** Read it live (`juju config neutron-api ...` + or `neutron.conf` on a unit). If it is `false`, **STOP** -- flipping it is a wider + blast radius that needs its own ruling; do not proceed on the assumption it is `true`. +- **0.2 Overlap+Magnum proof on a FOIL tenant.** Onboard a foil tenant with a CIDR that + deliberately overlaps an existing tenant's /24; confirm (a) the subnet creates and (b) + a Magnum cluster reaches ACTIVE with a working API-LB. Capture evidence to + `~/openstack-baseline/`. (This foil doubles as the `d011-05` P3 isolation foil.) + +### Phase 1 -- guard change (Code's lane; Code owns `tenant-onboard.sh`) +- Replace the exact-string cloud-wide guard with subnet-overlap math (Python + `ipaddress`) against the reserved-ranges registry **only**. Drop the tenant-vs-tenant + check. +- Extend `tests/tenant-onboard/` with fixtures covering: exact tenant-vs-tenant overlap + (now **ALLOWED**), partial overlap against a reserved range (**REJECTED**), exact match + against a reserved range (**REJECTED**), and a disjoint range (**ALLOWED**). No script + change ships without its harness. +- `bash scripts/repo-lint.sh` + the harness green before commit. + +### Phase 2 -- doc alignment +- `tenant-onboarding-contract.md` sec. 3.4: make explicit that tenant-vs-tenant overlap + is permitted (the wording is already close; remove any implication of cloud-wide + uniqueness). +- Amend **D-016** with a forward-pointer to D-074. +- `runbooks/tenant-onboarding-runbook.md`: replace the "carve a non-colliding /24 from + `10.20.0.0/16`" guidance with "tenant CIDR is client-choice or the standard default; + only reserved ranges are rejected." +- `clientdocs/` intake field: reword "requested internal IP address" -> + "requested internal network range (CIDR) -- optional; overlap-safe; default per policy" + (also fixes the IP-vs-range wording mismatch). + +### Phase 3 -- create the devops tenant on the corrected cloud +- `bash scripts/tenant-onboard.sh ` with `TENANT_CIDR` per the default + policy ruled in section 6. Verify with `scripts/tenant-assert.sh`. + +## 5. Lose / gain + +**Gain:** implementation matches the contract and the hard-isolation persona; onboarding +drops the "find a free /24" step; the intake field goes optional; clients can align to +their own on-prem/VPN ranges; NetBox's uniqueness job shrinks to the routed space that +genuinely must be unique (cleaner, and no global per-tenant IPAM registry to maintain +across DCs for Roosevelt); the 256-/24 ceiling of one /16 stops being a scaling limit. + +**Lose / cost:** debuggability -- an IP no longer uniquely identifies a tenant, so +captures/triage need net+tenant context (mitigate with disciplined naming); no +tenant-private-IP routing between tenants (an anti-pattern under hard isolation anyway); +IPAM discipline narrows to reserved ranges rather than disappearing; the guard change +touches a file Code is actively editing (coordination cost); D-016's `/16` pool becomes +vestigial (not reclaimed, just no longer enforced forward). + +## 6. Open operator decisions (nothing proceeds until ruled) + +1. **Rule D-074** -- adopt the overlap-allowed model? +2. **Default tenant CIDR policy** -- (a) uniform default (e.g. `10.0.0.0/24`) for all, or + (b) client-choice with a fallback default. Recommend (b). This sets what the devops + tenant receives. +3. **Keep `capi-mgmt` `10.20.0.0/24` reserved?** Recommend YES. +4. **`allow_overlapping_ips` confirmed `true`?** (Phase 0.1 gates everything.) + +## 7. Coordination / lane notes + +- The jumphost stream is actively editing `tenant-onboard.sh` and the contract today + (changelog addenda 21 / 23 / 26). The Phase-1 guard edit and all live steps are the + jumphost stream's lane; main-chat authored this plan only. **Code: pull and read this + before resuming onboard-script work.** +- D-074 is next-free per the scan at HEAD `59a7c73`; re-confirm before inserting. +- This is a new file with no edits to any shared/fenced file, to minimize collision. diff --git a/docs/archive/upstream-bug-draft-dashboard-tls.md b/docs/archive/upstream-bug-draft-dashboard-tls.md new file mode 100644 index 0000000..bd50f7c --- /dev/null +++ b/docs/archive/upstream-bug-draft-dashboard-tls.md @@ -0,0 +1,48 @@ +# Upstream bug draft -- charm-openstack-dashboard TLS backend on vhost-less address + +STATUS: DRAFT for operator submission (Launchpad: charm-openstack-dashboard). +Prepared 2026-07-05 from the ops-update-20260705 RCA (repo changelog addendum 15, +D-072). ASCII-only. No site secrets: addresses below are RFC1918 lab values. + +## Title +Multi-space deployment: haproxy https backend rendered on cluster-binding address, +which never receives an apache SSL vhost; L4 health check masks the dead TLS path + +## Affected +charm-openstack-dashboard 2024.1/stable (observed rev 728 and rev 750; render logic +identical across both -- charmhelpers get_network_addresses + haproxy context). +Reactive charm generation; juju 3.6. + +## Environment +Charmed OpenStack Caracal (jammy), MAAS spaces deployment. Application bindings: +default ('') = space A (admin), cluster = space B (internal), public = space C +(provider). hacluster VIPs on all three planes. vault:certificates TLS. + +## What happens +- ApacheSSLContext.get_network_addresses() derives SSL vhost addresses from the + DEFAULT-binding fallback (private-address) and the PUBLIC binding: vhosts render + for the space-A and space-C unit addresses only. +- The haproxy context renders the https (443) backend server line from the + CLUSTER-binding address (space B). +- Result: haproxy forwards TLS to :433 where apache has NO SSL + vhost; apache serves the connection from the non-SSL main server in PLAINTEXT + (verified: plain-HTTP GET to :433 answers 200). TLS clients fail handshake + (curl exit 35 / 000) on every dashboard VIP. +- haproxy's check is Layer4-only, so the backend shows UP and nothing alerts. +- juju status is fully green throughout; the condition is silent from day one. + +## Expected +Either the SSL vhost set includes every address the charm itself renders as an +HTTPS backend target, or the https backend uses an address from the vhost set +(default-binding address), or the health check is HTTPS-aware so the breakage +is at least visible. + +## Workaround +Bind `cluster` to the same space as the application default binding. Verified: +backends re-render onto a vhost-served address and VIP https serves immediately. + +## Evidence trail (available on request) +charm rev 728 vs 750 full diff (no vhost-logic change); charmhelpers +get_network_addresses trail log (LP:1952414 style) showing identical tuple sets +across 2800+ renders since deploy; apache2ctl -S vhost list vs haproxy.cfg +backend; plaintext-200-on-ssl-port probe; post-rebind verification. diff --git a/docs/archive/v1-pre-deploy-fixes.md b/docs/archive/v1-pre-deploy-fixes.md new file mode 100644 index 0000000..8bd5206 --- /dev/null +++ b/docs/archive/v1-pre-deploy-fixes.md @@ -0,0 +1,832 @@ +# v1 Pre-Deploy Fixes (v2 -- includes Designate deferral) + +**Purpose:** Single-pass repo hygiene before v1 deployment execution begins. Apply these fixes as one logical commit per group (nine commits total) before any execution document runs. + +**Status:** Authoritative for the v1 deploy track. Supersedes the v1 draft of this document. + +**Scope:** Repo-only changes. No cloud state is touched by this document. All changes are reviewed locally, committed to `main`, and pushed before the next do-document runs. + +**What changed in v2 of this change list (2026-05-27):** + +- Added commits 7-9 implementing the Designate-deferral decision (D-019). +- Amended commit 5 (deprecated runbook moves) -- `07-dns-zones.md` is now permanently deprecated per D-019, not "replaced by v1-do-doc-10-dns." +- Amended commit 6 (deprecated README content) -- same. +- Updated commit 4 (README.md refresh) -- adds language reflecting Designate deferral to the v1 scope section. +- Updated sect.10 verification commands to expect 11 VIPs, not 12. + +## Cross-references + +- D-002 (channel matrix) -- Vault row cleanup +- D-005 (Ceph Squid release) +- D-008 (DNS architecture) -- superseded by D-019; v2-scope +- D-011 (validation bar) -- amended by D-019 (Designate criterion dropped) +- D-014 (repo path) -- stale path correction +- D-017 (CAPI bootstrap cluster lifecycle) -- supersedes runbook 00 Phase 5 +- D-018 (teardown strategy) -- supersedes runbook 00 Phase 4 +- **D-019 (NEW) -- Designate deferral to v2; tenant resolvers use public DNS** +- Charmed Ceph `charm-ceph-osd` `config.yaml` and `metadata.yaml` (osd-devices semantics) +- Ceph BlueStore configuration reference (single-device co-located OSD pattern) + +--- + + +## 1. Bundle: remove ceph-osd `storage:` block + +### What + +In `bundle.yaml`, under the `ceph-osd` application, delete the entire `storage:` block. The `options.osd-devices: /dev/vdb` line stays. + +### Why + +The `osd-devices` declared under `storage:` is **additive** to the `options.osd-devices` value per the `charm-ceph-osd` `config.yaml`: "These devices are the range of devices that will be checked for and used across all service units, in addition to any volumes attached via the `--storage` flag during deployment." + +Concretely, with the current bundle: + +- `options.osd-devices: /dev/vdb` -> one OSD per unit using the 512 GB libvirt-attached disk +- `storage.osd-devices: loop,1024M` -> an *additional* 1 GB loopback OSD per unit + +Total: 2 OSDs per unit x 4 units = 8 OSDs, against `expected-osd-count: 4` on `ceph-mon`. The 1 GB OSDs are below practical minimums, asymmetric with the 512 GB primaries (CRUSH-weighting anti-pattern), and provide no operational value. + +The remaining `bluestore-db`, `bluestore-wal`, `cache-devices`, and `osd-journals` loopback entries are also being removed -- not because they break anything, but because: + +1. BlueStore co-locates DB and WAL on the data device when no separate volume is supplied (Ceph Reef BlueStore reference). +2. Loopback files on the same backing storage as `/dev/vdb` are not "faster than the primary device," so the standard rationale for separate DB/WAL devices doesn't apply. +3. `osd-journals` is unused under BlueStore (default since Luminous). + +### Diff + +```yaml +# BEFORE (in bundle.yaml, applications.ceph-osd) + ceph-osd: + charm: ceph-osd + channel: squid/stable + num_units: 4 + to: ["8", "9", "10", "11"] + options: + source: *ceph-source + osd-devices: /dev/vdb # libvirt-attached, MAAS-untracked, wiped 2026-05-22 + bindings: *internal-bindings + constraints: arch=amd64 tags=openstack + storage: # Loop-backed auxiliaries (testcloud has no real SSDs) + bluestore-db: loop,1024M + bluestore-wal: loop,1024M + cache-devices: loop,10240M + osd-devices: loop,1024M + osd-journals: loop,1024M # Legacy storage name still in squid metadata; benign + +# AFTER + ceph-osd: + charm: ceph-osd + channel: squid/stable + num_units: 4 + to: ["8", "9", "10", "11"] + options: + source: *ceph-source + osd-devices: /dev/vdb # libvirt-attached, MAAS-untracked, wiped 2026-05-22 + bindings: *internal-bindings + constraints: arch=amd64 tags=openstack +``` + +### Commit message + +``` +bundle: remove ceph-osd storage block to match expected-osd-count + +The storage: block declared a second osd-devices entry (loop,1024M) which +is additive to options.osd-devices per the charm config. That produced +8 OSDs against expected-osd-count: 4 on ceph-mon, with 1 GB loopback +OSDs as the asymmetric secondaries -- a CRUSH-weighting anti-pattern. + +Real production storage (DB/WAL on actual SSDs) will be declared on +Roosevelt. For the testcloud, BlueStore co-locates DB/WAL on the data +device which is the documented default for single-device setups. + +osd-journals is unused under BlueStore. +``` + +--- + + +## 2. design-decisions.md: D-002 -- remove Vault from OpenStack-core row + +### What + +In `docs/design-decisions.md` under D-002 (channel matrix), the OpenStack-core row currently lists `, vault` as one of the components on `2024.1/stable`. Remove that token. + +### Why + +Vault has its own track per Canonical's charm delivery table -- it runs on `1.8/stable`, not `2024.1/stable`. The D-002 table elsewhere (and the actual bundle.yaml) already reflects this; the OpenStack-core row description was a leftover from earlier drafting. + +### Diff + +```diff +-| OpenStack core API charms (keystone, glance, nova-cloud-controller, neutron-api, cinder, placement, octavia, barbican, magnum, designate, openstack-dashboard, vault) | `2024.1/stable` | ++| OpenStack core API charms (keystone, glance, nova-cloud-controller, neutron-api, cinder, placement, octavia, barbican, magnum, designate, openstack-dashboard) | `2024.1/stable` | +``` + +### Commit message + +``` +docs/design-decisions: D-002 -- drop vault from OpenStack-core channel row + +Vault uses 1.8/stable per the Canonical charm delivery table, not the +2024.1/stable OpenStack-core track. The bundle.yaml already reflects +this; the design-decisions D-002 description had a stale token. +``` + +> **Note:** the `designate` token in the row above is correct as-of this commit (Designate is still on `2024.1/stable` channel). The Designate row is removed entirely by commit 8 (D-019). + +--- + + +## 3. design-decisions.md: D-014 -- update repo path + +### What + +In `docs/design-decisions.md` under D-014 (repo location), the path currently shows the per-user namespace from before the 2026-05-27 move. Update to the OpenStack-group path. + +### Why + +Per the user-memory pinned note: "Caracal rebuild repo (moved to OpenStack group 2026-05-27): https://git.baldurkeep.com/OpenStack/openstack-caracal-ipv4 (web), https://git.baldurkeep.com/git/OpenStack/openstack-caracal-ipv4.git (clone). Old jesse.austin/openstack-caracal-ipv4 path no longer exists; GitBucket does not redirect." + +### Diff + +```diff + ## D-014: Repository location + +-**Decision:** `git.baldurkeep.com/jesse.austin/openstack-caracal-ipv4` for v1. ++**Decision:** `git.baldurkeep.com/OpenStack/openstack-caracal-ipv4` for v1. ++ ++- Web: `https://git.baldurkeep.com/OpenStack/openstack-caracal-ipv4` ++- Clone: `https://git.baldurkeep.com/git/OpenStack/openstack-caracal-ipv4.git` ++- Moved from `jesse.austin/openstack-caracal-ipv4` to the `OpenStack` group on 2026-05-27. GitBucket does not redirect from the old path. + + **Rationale:** Establishes a single repo per cloud lifecycle. v2 path TBD. +``` + +### Commit message + +``` +docs/design-decisions: D-014 -- update repo path after move to OpenStack group + +Repository moved from jesse.austin/openstack-caracal-ipv4 to the +OpenStack group on 2026-05-27. GitBucket does not redirect, so the +prior path is dead. +``` + +--- + + +## 4. README.md: refresh stale references + +### What + +Three hunks in `README.md`: + +1. Update the design-decisions reference range from "D-001 through D-016" to "D-001 through D-019" (post-D-017/D-018/D-019 additions). +2. Replace the inline runbook 00 description (which still mentions backups + capi-mgmt graceful teardown -- both invalidated by D-017/D-018). +3. Replace the 12-step deploy order with a pointer to the do-document set. +4. Add a v1 scope reduction note for Designate (per D-019). + +### Why + +The README's v1 deployment order block reflects the original 12-step runbook plan; the actual deploy path now flows through the v1-do-doc-NN execution documents. Keeping the README in sync prevents new operators from following the stale path. + +The D-019 note in the v1 scope section makes the Designate deferral explicit for anyone reading the README to understand v1 scope. + +### Diff (four hunks) + +**Hunk A -- D-range:** + +```diff +-`--- docs/ +- `--- design-decisions.md # architectural record (D-001 through D-016) ++`--- docs/ ++ |--- design-decisions.md # architectural record (D-001 through D-019) ++ `--- netbox-vip-queue.md # post-deploy NetBox imports (workstream 2) +``` + +**Hunk B -- runbook 00 description (in the layout block):** + +```diff +-| |--- 00-pre-deploy.md # backups, capi-mgmt graceful teardown ++| # (deprecated; see runbooks/deprecated/ -- superseded by D-017 + D-018 + v1-do-doc-NN set) +``` + +**Hunk C -- replace the deploy-order block:** + +```diff +-## v1 deployment order +- +-1. Verify NetBox state -- run NetBox imports if not already applied +- - `netbox/ipv4-prefixes-import.py` -- required +- - `netbox/ipv6-mark-reserved.py` -- required (Q3: tag existing IPv6 entries) +-2. Run pre-flight checks (`scripts/pre-flight-checks.sh`) +-3. Backup current cloud state (`runbooks/00-pre-deploy.md`) +-4. Destroy existing OpenStack model (`runbooks/01-destroy-model.md`) +-5. Deploy new bundle (`runbooks/02-deploy.md`) +-6. Initialize Vault (`runbooks/03-vault-init.md`) +-7. Set up Magnum domain (`runbooks/04-magnum-domain.md`) +-8. Stand up CAPI bootstrap cluster on `capi-mgmt.maas` (`runbooks/04a-capi-bootstrap-cluster.md`) +-9. Install Magnum CAPI Helm driver (`runbooks/05-magnum-capi-driver.md`) +-10. Recreate tenant resources (`runbooks/06-tenant-setup.md`) +-11. Populate DNS zones (`runbooks/07-dns-zones.md`) +-12. Run validation (`runbooks/08-validate.md` + `scripts/validate.sh`) ++## v1 deployment order ++ ++The deploy is executed via the `runbooks/v1-do-doc-NN-*.md` execution documents in numeric order: ++ ++| Doc | Purpose | ++|---|---| ++| `v1-do-doc-01-prep.md` | Pre-flight state check (repo, openrc, MAAS state of 5 VMs) | ++| `v1-do-doc-02-pki.md` | Octavia PKI overlay generation | ++| `v1-do-doc-03-destroy.md` | Conditional model + MAAS teardown (clean state for rebuild) | ++| `v1-do-doc-04-deploy.md` | `juju deploy` + settle wait + on-disk PKI verification | ++| `v1-do-doc-05-vault-init.md` | Vault initialization + cert cascade + admin-openrc regeneration | ++| `v1-do-doc-06-magnum-domain.md` | Magnum Keystone domain setup | ++| `v1-do-doc-07-capi-bootstrap.md` | CAPI bootstrap cluster + workload pivot | ++| `v1-do-doc-08-magnum-driver.md` | Magnum CAPI Helm driver graft | ++| `v1-do-doc-09-tenant.md` | Tenant project/user/openrc + Snapshot 2 | ++| `v1-do-doc-10-validate.md` | D-011 acceptance criteria + Snapshot 3 | ++ ++NetBox imports are run separately (gated on external NetBox engineer review; see `netbox/README.md`). +``` + +**Hunk D -- v1 scope note about Designate deferral:** + +```diff + ## v1-specific design decisions (summary; see docs/design-decisions.md for full record) + + - **D-015 v1/v2 fork** -- IPv4-only v1; IPv6/dual-stack v2 deferred + - **D-016 IPv4 tenant pool hybrid model** -- NetBox owns upstream `/16` pool; Neutron owns per-project subnets within it + - **D-003 Option B network architecture** -- Provider `/22` carries both ext_net FIPs (`10.12.4.10-.223`) and OpenStack public API VIPs (`10.12.4.224-.254`) on the same L2 segment; fixes the tenant->API unreachability that caused Magnum OCCM crashloop on Bobcat testcloud + - **D-005 Ceph Squid** -- matches Caracal default; rehearses Roosevelt + - **D-006 Vault HA backend = etcd + easyrsa** + - **D-007 Magnum from day one** -- charm in bundle + CAPI Helm driver graft +-- **D-008 DNS via Designate from day one** -- static /etc/hosts for bootstrap; Designate handles tenant-level resolution (A records only for v1) ++- **D-019 (supersedes D-008) DNS scope reduction for v1** -- Designate deferred to v2 alongside corporate DNS / NS-delegation work. Tenant subnets use public DNS (`1.1.1.1` / `1.0.0.1`) directly via `--dns-nameserver`. `*.cloud.neumatrix.local` FQDN tree remains internal-only, resolved via static `/etc/hosts` on bootstrap-relevant hosts. + - **D-009 Hacluster relations included at num_units=1** -- decorative on testcloud; documents the relation pattern for Roosevelt scale-up + - **No OVN pinning on testcloud** -- Roosevelt bare-metal will pin via `ovn-source` +``` + +### Commit message + +``` +README: refresh stale runbook references, reflect D-019 scope reduction + +The v1 deployment order block referenced runbooks 00-08 which are +being moved to deprecated/ in favor of v1-do-doc-NN execution documents. +Replaces the order block with a pointer to the do-document set. + +Also reflects D-019: Designate deferred to v2; v1 tenant resolvers use +public DNS. Adds the netbox-vip-queue.md reference. Updates the design- +decisions D-range from D-001-D-016 to D-001-D-019. +``` + +--- + + +## 5. Move superseded runbooks to `runbooks/deprecated/` + +### What + +`git mv` each superseded runbook into a new `runbooks/deprecated/` directory. Add `runbooks/deprecated/README.md` explaining the deprecation in commit 6. + +### Why + +The v1-do-doc-NN set replaces the prior runbook 00-08 work. Keeping the originals in `runbooks/deprecated/` preserves the audit trail without misleading new operators into following the old paths. + +### Files to move + +| From | To | Replacement | +|---|---|---| +| `runbooks/00-pre-deploy.md` | `runbooks/deprecated/00-pre-deploy.md` | superseded by D-017 + D-018 (no per-cycle backups; teardown direct-to-MAAS); v1-do-doc-01 covers prep | +| `runbooks/01a-octavia-pki-generation.md` | `runbooks/deprecated/01a-octavia-pki-generation.md` | `v1-do-doc-02-pki.md` | +| `runbooks/02-deploy.md` | `runbooks/deprecated/02-deploy.md` | `v1-do-doc-04-deploy.md` | +| `runbooks/03-vault-init.md` | `runbooks/deprecated/03-vault-init.md` | `v1-do-doc-05-vault-init.md` | +| `runbooks/04-magnum-domain.md` | `runbooks/deprecated/04-magnum-domain.md` | `v1-do-doc-06-magnum-domain.md` | +| `runbooks/04a-capi-bootstrap-cluster.md` | `runbooks/deprecated/04a-capi-bootstrap-cluster.md` | `v1-do-doc-07-capi-bootstrap.md` | +| `runbooks/05-magnum-capi-driver.md` | `runbooks/deprecated/05-magnum-capi-driver.md` | `v1-do-doc-08-magnum-driver.md` | +| `runbooks/06-tenant-setup.md` | `runbooks/deprecated/06-tenant-setup.md` | `v1-do-doc-09-tenant.md` | +| **`runbooks/07-dns-zones.md`** | **`runbooks/deprecated/07-dns-zones.md`** | **deferred to v2 per D-019** (no v1 replacement) | +| `runbooks/08-validate.md` | `runbooks/deprecated/08-validate.md` | `v1-do-doc-10-validate.md` | + +### Files NOT moved + +| File | Why kept in `runbooks/` | +|---|---| +| `runbooks/01-destroy-model.md` | Referenced by v1-do-doc-03 as a conditional sub-procedure; still active | + +### Git commands + +```bash +cd "$HOME/openstack-caracal-ipv4" + +mkdir -p runbooks/deprecated + +git mv runbooks/00-pre-deploy.md runbooks/deprecated/ +git mv runbooks/01a-octavia-pki-generation.md runbooks/deprecated/ +git mv runbooks/02-deploy.md runbooks/deprecated/ +git mv runbooks/03-vault-init.md runbooks/deprecated/ +git mv runbooks/04-magnum-domain.md runbooks/deprecated/ +git mv runbooks/04a-capi-bootstrap-cluster.md runbooks/deprecated/ +git mv runbooks/05-magnum-capi-driver.md runbooks/deprecated/ +git mv runbooks/06-tenant-setup.md runbooks/deprecated/ +git mv runbooks/07-dns-zones.md runbooks/deprecated/ +git mv runbooks/08-validate.md runbooks/deprecated/ + +git status +# Expect: 10 renames staged, runbooks/01-destroy-model.md untouched +``` + +### Commit message + +``` +runbooks: move superseded files to runbooks/deprecated/ + +These are replaced by v1-do-doc-NN-*.md execution documents (added in +follow-up commits). The 01-destroy-model.md runbook stays in place -- it's +referenced by v1-do-doc-03 as a conditional sub-procedure. + +07-dns-zones.md is deferred to v2 per D-019 with no v1 replacement +(Designate is no longer in v1 scope). + +History preserved via git mv. +``` + +--- + + +## 6. Add a deprecation banner to `runbooks/deprecated/README.md` + +### What + +Create a new file `runbooks/deprecated/README.md` with a banner explaining the deprecation scope and a replacement map. + +### Content + +```markdown +# Deprecated v1 Runbooks + +The runbooks in this directory have been superseded by the +`runbooks/v1-do-doc-NN-*.md` execution documents (or, in the case of +`07-dns-zones.md`, deferred to v2 entirely per D-019). + +They are preserved here so the audit trail from the early v1 drafting +phase remains accessible. **Do not execute them.** The v1 deploy is +gated through the do-document set. + +## Replacement map + +| Deprecated runbook | Replacement | +|---|---| +| `00-pre-deploy.md` | superseded by D-017 + D-018 (no per-cycle backups; direct MAAS teardown); `v1-do-doc-01-prep.md` covers prep | +| `01a-octavia-pki-generation.md` | `v1-do-doc-02-pki.md` | +| `02-deploy.md` | `v1-do-doc-04-deploy.md` | +| `03-vault-init.md` | `v1-do-doc-05-vault-init.md` | +| `04-magnum-domain.md` | `v1-do-doc-06-magnum-domain.md` | +| `04a-capi-bootstrap-cluster.md` | `v1-do-doc-07-capi-bootstrap.md` | +| `05-magnum-capi-driver.md` | `v1-do-doc-08-magnum-driver.md` | +| `06-tenant-setup.md` | `v1-do-doc-09-tenant.md` | +| `07-dns-zones.md` | **deferred to v2 per D-019** (no v1 replacement) | +| `08-validate.md` | `v1-do-doc-10-validate.md` | + +`01-destroy-model.md` is **not** in this directory -- it remains active in +`runbooks/` and is referenced as a conditional sub-procedure by +`v1-do-doc-03-destroy.md`. +``` + +### Commit message + +``` +runbooks/deprecated: add README explaining the deprecation scope + +Includes the deprecated -> replacement mapping for operators who arrive +via git log searches or stale internal references. +``` + +--- + + +## 7. Bundle: remove Designate (per D-019) + +### What + +In `bundle.yaml`, remove four applications, seven relations, and update a header comment to reflect Designate deferral to v2. VIP `10.12.4.227` becomes unused space in the `10.12.4.224-.254` range (same status as `10.12.4.225` reserved for v2 ceph-radosgw HA). + +### Why + +Per D-019 (added by commit 8 of this change list): Designate is deferred to v2 alongside corporate-DNS / NS-delegation work. v1 testcloud topology investigation (2026-05-27 session) confirmed: + +1. **Outside-in DNS** isn't needed for v1 -- corporate clients reach the cloud through the existing `openstack.baldurkeep.com -> 10.17.4.20 -> 10.12.x` HTTPS proxy chain, not via the `*.cloud.neumatrix.local` FQDN tree. The edge nginx (`neumatrix-nginx` at `10.17.8.7`) cannot route to `10.12.x` directly anyway. +2. **Inside-out DNS** doesn't require Designate -- tenant subnets can specify public DNS (`1.1.1.1`, `1.0.0.1`) directly via `--dns-nameserver` at subnet-create time. +3. **FIP DNS auto-registration** (the remaining v1 use case for Designate) is nice-to-have, not load-bearing for any v1 acceptance criterion. + +Removing Designate now removes one charm, four applications (with subordinate routers), seven relations, and one VIP from the v1 deploy surface, reducing first-deploy troubleshooting scope. + +### Diff + +**Remove four applications** (search-and-delete each block in `bundle.yaml`): + +```yaml + # ===================================================================== + # DNS: Designate (NEW for Caracal v1 per D-008) + # ===================================================================== + # Naming convention: .omega.dc0.vr0.cloud.neumatrix.local + + designate: + charm: designate + channel: 2024.1/stable + num_units: 1 + to: [lxd:8] + options: + openstack-origin: *openstack-origin + nameservers: "ns1.omega.dc0.vr0.cloud.neumatrix.local. ns2.omega.dc0.vr0.cloud.neumatrix.local." + vip: 10.12.4.227 + os-public-hostname: designate.omega.dc0.vr0.cloud.neumatrix.local + bindings: *api-bindings + constraints: arch=amd64 + + designate-mysql-router: + charm: mysql-router + channel: 8.0/stable + + designate-bind: + charm: designate-bind + channel: 2024.1/stable + num_units: 1 + to: [lxd:8] + bindings: + "": provider # unit on provider so bind9:53 reachable from tenants (D-003) + cluster: metal # peer traffic stays internal (decorative with num_units=1) + constraints: arch=amd64 +``` + +**Remove the `designate-hacluster:` line** from the hacluster subordinate block: + +```diff + keystone-hacluster: { charm: hacluster, channel: 2.4/stable } + glance-hacluster: { charm: hacluster, channel: 2.4/stable } + neutron-api-hacluster: { charm: hacluster, channel: 2.4/stable } + nova-cloud-controller-hacluster: { charm: hacluster, channel: 2.4/stable } + placement-hacluster: { charm: hacluster, channel: 2.4/stable } + openstack-dashboard-hacluster: { charm: hacluster, channel: 2.4/stable } + cinder-hacluster: { charm: hacluster, channel: 2.4/stable } + octavia-hacluster: { charm: hacluster, channel: 2.4/stable } + barbican-hacluster: { charm: hacluster, channel: 2.4/stable } + magnum-hacluster: { charm: hacluster, channel: 2.4/stable } + vault-hacluster: { charm: hacluster, channel: 2.4/stable } + # v2-deferred: ceph-radosgw-hacluster: { charm: hacluster, channel: 2.4/stable } +- designate-hacluster: { charm: hacluster, channel: 2.4/stable } ++ # v2-deferred (D-019): designate-hacluster: { charm: hacluster, channel: 2.4/stable } +``` + +**Remove seven relations** from the `relations:` block: + +```yaml + # ---- Designate (DNS) -- NEW for Caracal v1 per D-008 + - [designate-mysql-router:db-router, mysql-innodb-cluster:db-router] + - [designate-mysql-router:shared-db, designate:shared-db] + - [designate:identity-service, keystone:identity-service] + - [designate:amqp, rabbitmq-server:amqp] + - [designate:certificates, vault:certificates] + - [designate:dns-backend, designate-bind:dns-backend] + - [designate:ha, designate-hacluster:ha] +``` + +**Update the header comment block** (decision references in the bundle's top-of-file block): + +```diff + D-006 Vault HA via etcd + easyrsa + D-007 Magnum Layer A + Layer B graft +- D-008 Designate day-one ++ D-019 (supersedes D-008) Designate deferred to v2 + D-009 hacluster subordinates (decorative on testcloud) +``` + +**Update the HA subordinate header comment:** + +```diff + # ===================================================================== +- # HA Cluster Subordinates (12 active for v1; ceph-radosgw deferred to v2) ++ # HA Cluster Subordinates (11 active for v1; ceph-radosgw + designate deferred to v2) + # ===================================================================== +``` + +### Verification post-edit + +```bash +cd "$HOME/openstack-caracal-ipv4" + +# 1. No designate application remains +grep -E "^ designate" bundle.yaml \ + && echo "[FAIL] designate-related application still present" \ + || echo "[OK] no designate applications" + +# 2. No designate relations remain +grep -E "designate" bundle.yaml | grep -vE "^[[:space:]]*#" +# Expect: no output (the only remaining 'designate' tokens should be commented) + +# 3. VIP count is now 11 +VIP_COUNT=$(grep -cE "^[[:space:]]+vip: 10\.12\.4\." bundle.yaml) +echo "VIPs: $VIP_COUNT (expect 11)" + +# 4. YAML still parses +python3 -c "import yaml; yaml.safe_load(open('bundle.yaml')); print('[OK] YAML parses')" +``` + +### Commit message + +``` +bundle: remove Designate per D-019 (deferred to v2) + +Removes the designate, designate-bind, designate-mysql-router, and +designate-hacluster applications, plus all seven designate-related +relations. Updates header comments to reflect the deferral. + +Rationale per D-019: v1 testcloud topology investigation confirmed +outside-in DNS is not needed (corporate clients reach the cloud via +the openstack.baldurkeep.com HTTPS proxy chain, not via *.cloud. +neumatrix.local FQDNs). Tenant subnets use public DNS directly via +--dns-nameserver. FIP DNS auto-registration is not load-bearing for +any v1 acceptance criterion. + +VIP 10.12.4.227 becomes unused space in 10.12.4.224-.254 (same +status as 10.12.4.225 reserved for v2 ceph-radosgw HA). + +Reduces v1 deploy surface by one charm + three subordinates + +seven relations + one VIP. +``` + +--- + + +## 8. design-decisions.md: add D-019 + amend D-008 and D-011 + +### What + +Three coordinated edits in `docs/design-decisions.md`: + +1. Add a new D-019 entry with the full deferral rationale and v2 deltas +2. Mark D-008 as "superseded by D-019 -- v2-scope" +3. Amend D-011 (validation bar) to remove the "Designate resolves" criterion + +### Why + +Captures the decision in the authoritative design-decisions record. Preserves the audit trail (D-008 stays with its original content but with the superseded status flag). + +### Diff -- Hunk A: amend D-008 status + +```diff + ## D-008: DNS architecture + +-**Decision:** Layered -- static /etc/hosts for bootstrap + Designate (in bundle from day one) for tenant-level resolution. ++**Status:** Superseded by D-019 (2026-05-27). v2-scope. Original decision text preserved below for audit. ++ ++**Decision (original; superseded):** Layered -- static /etc/hosts for bootstrap + Designate (in bundle from day one) for tenant-level resolution. +``` + +### Diff -- Hunk B: amend D-011 validation bar + +In the D-011 section, the testcloud validation criteria currently list (among others) a Designate resolution check. Remove the relevant bullet. + +```diff + ## D-011: Roosevelt-rehearsal validation bar + + **Decision:** v1 testcloud must pass these criteria before being declared "deploy-equivalent" to Roosevelt: + + - All charms `active/idle` (per `juju status`) + - Tenant subnet -> OpenStack API reachability (Bobcat Magnum OCCM crashloop regression test) + - Octavia LBaaS end-to-end (LB create + member health + failover + recovery) + - Magnum CAPI end-to-end (cluster template + cluster create + CREATE_COMPLETE) + - Vault unseal-after-reboot survives a power cycle +-- Designate resolves API FQDNs via the Designate VIP + - Snapshots 1, 2, 3 captured at the appropriate stages (per D-012) + ++**Amendment (2026-05-27):** Per D-019, the "Designate resolves" criterion is removed for v1. Designate is deferred to v2; tenant subnets resolve via public DNS. v2 will reinstate a DNS-resolution validation criterion calibrated to whatever DNS mechanism is in place (NS delegation from corporate DNS, or otherwise). +``` + +### Diff -- Hunk C: add D-019 (new entry, append to end of design-decisions.md) + +```markdown +--- + + +## D-019: DNS scope reduction for v1 -- Designate deferred to v2 + +**Decision (2026-05-27):** Designate is removed from the v1 testcloud bundle and deferred to v2 alongside corporate DNS / NS delegation work. v1 tenant subnets resolve via public DNS (`1.1.1.1`, `1.0.0.1`) directly via the `--dns-nameserver` option at subnet-create time. + +**Supersedes:** D-008 (DNS architecture). + +**Amends:** D-011 (validation bar -- removes "Designate resolves" criterion). + +### Rationale + +Three findings from the 2026-05-27 testcloud topology investigation: + +1. **Outside-in DNS** (corporate clients resolving `*.cloud.neumatrix.local`) is not needed for v1. Corporate access to the cloud already flows through the existing `openstack.baldurkeep.com -> 10.17.4.20 -> 10.12.x` HTTPS proxy chain (handled by the edge nginx at `10.17.8.7`), which does not depend on corporate-side resolution of cloud-internal FQDNs. + +2. **The edge nginx cannot route to `10.12.x` directly.** Inspection confirmed the edge has only `10.17.8.7/22` plus a tailscale interface; reaching `10.12.4.x` requires the libvirt-host NAT path. Adding DNS to the testcloud would require parallel UDP/53 NAT/proxy plumbing across three hosts (edge nginx, libvirt host, internal nginx) for a feature that has no v1 consumer. + +3. **Inside-out DNS** (tenant VMs resolving external names) is satisfied by tenant subnets pointing `--dns-nameserver` at public DNS (`1.1.1.1`, `1.0.0.1`). Designate is not needed in the inside-out path either, since: + - Tenant VMs don't need to resolve cloud-internal FQDNs (their API access goes through documented IPs / `--cloud` configs in cloud.conf) + - Cross-tenant DNS visibility is not a v1 requirement + +The remaining v1 use case for Designate (FIP DNS auto-registration via the `neutron-api <-> designate` integration) is informational only -- nothing in v1 consumes those records. + +### v1 implementation + +- Tenant subnets created with `--dns-nameserver 1.1.1.1 --dns-nameserver 1.0.0.1` (or via the openrc `OS_DNS_NAMESERVERS` env) +- CAPI workload cluster template variable `OPENSTACK_DNS_NAMESERVERS` set to `1.1.1.1,1.0.0.1` (per `v1-do-doc-07-capi-bootstrap.md` sect.13) +- Cloud-internal `*.cloud.neumatrix.local` FQDN tree resolved via static `/etc/hosts` on bootstrap-relevant hosts (jumphost, openstack0-3, LXD containers per charm bootstrap, capi-mgmt -- staged in `v1-do-doc-05-vault-init.md` sect.11 and `v1-do-doc-07-capi-bootstrap.md` sect.6) +- Charms continue to use FQDN-based `os-public-hostname` (cert SANs depend on it) -- internal resolution via `/etc/hosts` is sufficient + +### v2 plan + +- Re-introduce Designate (charm + designate-bind + relations + hacluster sub) +- NS delegation from corporate DNS to designate-bind on a real (non-NAT) network VIP +- Tenant subnets transitioning to use Designate VIP as their resolver (after corporate DNS delegation lands) +- Designate v2 deploy on a real-network Roosevelt or v2-testcloud topology where the bridging-host complexity from v1 testcloud doesn't apply +- D-011 validation re-introduces a calibrated DNS-resolution criterion (mechanism TBD: NS delegation working end-to-end vs static A records at corporate DNS) + +### v2-residency note + +The IPv6 prefixes already imported into NetBox (and marked Reservation status) include allocations that would be appropriate for Designate's VIPs in a v2 design -- these stay in NetBox as Reservation until v2 work begins. +``` + +### Commit message + +``` +docs/design-decisions: add D-019, amend D-008/D-011 (DNS scope reduction) + +D-019 captures the Designate deferral to v2 with rationale grounded in +the 2026-05-27 testcloud topology investigation: outside-in DNS not +needed (corporate clients use openstack.baldurkeep.com HTTPS chain); +edge nginx can't route to cloud-internal anyway; inside-out is +satisfied by tenant --dns-nameserver pointing at public DNS. + +D-008 (DNS architecture) marked superseded; original text preserved +for audit trail. + +D-011 validation bar amended to remove "Designate resolves" criterion; +v2 will reinstate a calibrated DNS criterion. +``` + +--- + + +## 9. netbox-vip-queue.md: drop the Designate VIP row + +### What + +In `docs/netbox-vip-queue.md`, remove the row for Designate's VIP `10.12.4.227`. The queue goes from 12 entries to 11. + +### Why + +Per D-019, Designate has no VIP in v1. `10.12.4.227` becomes unused space within the `10.12.4.224-.254` range -- same status as `10.12.4.225` (reserved for v2 ceph-radosgw HA). + +### Diff + +```diff + # NetBox VIP Queue (post-deploy) + + The following 12 IPAddress entries should be added to NetBox after the + v1 deploy completes and the engineer-review of `netbox/ipv4-prefixes- + import.py` has landed. + +-| VIP | Service | Notes | +-|---|---|---| +-| 10.12.4.224 | barbican | per D-003 | +-| 10.12.4.226 | cinder | per D-003 | +-| 10.12.4.227 | designate | per D-003 + D-008 | +-| 10.12.4.228 | glance | per D-003 | +-| 10.12.4.229 | keystone | per D-003 | +-| 10.12.4.230 | magnum | per D-003 | +-| 10.12.4.231 | neutron-api | per D-003 | +-| 10.12.4.232 | nova-cloud-controller | per D-003 | +-| 10.12.4.233 | octavia | per D-003 | +-| 10.12.4.234 | openstack-dashboard | per D-003 | +-| 10.12.4.235 | placement | per D-003 | +-| 10.12.4.236 | vault | per D-003 | ++The following 11 IPAddress entries should be added to NetBox after the ++v1 deploy completes and the engineer-review of `netbox/ipv4-prefixes- ++import.py` has landed. ++ ++| VIP | Service | Notes | ++|---|---|---| ++| 10.12.4.224 | barbican | per D-003 | ++| 10.12.4.226 | cinder | per D-003 | ++| 10.12.4.228 | glance | per D-003 | ++| 10.12.4.229 | keystone | per D-003 | ++| 10.12.4.230 | magnum | per D-003 | ++| 10.12.4.231 | neutron-api | per D-003 | ++| 10.12.4.232 | nova-cloud-controller | per D-003 | ++| 10.12.4.233 | octavia | per D-003 | ++| 10.12.4.234 | openstack-dashboard | per D-003 | ++| 10.12.4.235 | placement | per D-003 | ++| 10.12.4.236 | vault | per D-003 | ++ ++**Reserved (unused in v1):** ++ ++- `10.12.4.225` -- reserved for v2 ceph-radosgw HA (workstream-2 decision) ++- `10.12.4.227` -- reserved for v2 designate (D-019 deferral) +``` + +(The exact line counts may differ if the existing file has more preamble -- adjust to the actual structure when editing.) + +### Commit message + +``` +docs/netbox-vip-queue: drop designate row per D-019 + +VIP 10.12.4.227 is no longer in v1; designate deferred to v2. Adds a +"reserved (unused in v1)" section to make the unused slots in the +.224-.254 range explicit for future reference. +``` + +--- + + +## 10. Verification (read-only) after the nine commits land + +```bash +cd "$HOME/openstack-caracal-ipv4" +git pull +git log --oneline -10 + +echo "=== 1. ceph-osd no longer has a storage block ===" +grep -A 12 "^ ceph-osd:" bundle.yaml | grep "storage:" \ + && echo "[FAIL] storage: block still present in ceph-osd" \ + || echo "[OK] no storage: in ceph-osd" + +echo "=== 2. D-002 vault token removed ===" +grep -E "OpenStack core API charms" docs/design-decisions.md | grep -v "vault)" \ + && echo "[OK]" \ + || echo "[FAIL] D-002 still contains vault token" + +echo "=== 3. D-014 reflects new repo path ===" +grep "OpenStack/openstack-caracal-ipv4" docs/design-decisions.md \ + && echo "[OK]" \ + || echo "[FAIL]" + +echo "=== 4. README reflects v1-do-doc set ===" +grep "v1-do-doc-NN" README.md \ + && echo "[OK]" \ + || echo "[FAIL]" + +echo "=== 5. Designate gone from bundle ===" +grep -E "^ designate" bundle.yaml \ + && echo "[FAIL] designate-related app still present" \ + || echo "[OK] no designate apps" + +grep -E "designate" bundle.yaml | grep -vE "^[[:space:]]*#" +# Expect: empty (only commented references remain) + +echo "=== 6. VIP count is 11 ===" +VIP_COUNT=$(grep -cE "^[[:space:]]+vip: 10\.12\.4\." bundle.yaml) +echo "VIPs: $VIP_COUNT (expect 11)" + +echo "=== 7. D-019 added ===" +grep "^## D-019" docs/design-decisions.md \ + && echo "[OK]" \ + || echo "[FAIL]" + +echo "=== 8. D-008 marked superseded ===" +grep -A 2 "^## D-008" docs/design-decisions.md | grep "Superseded by D-019" \ + && echo "[OK]" \ + || echo "[FAIL]" + +echo "=== 9. netbox-vip-queue.md has 11 entries ===" +QUEUE_COUNT=$(grep -cE "^\| 10\.12\.4\." docs/netbox-vip-queue.md) +echo "VIP queue entries: $QUEUE_COUNT (expect 11)" + +echo "=== 10. runbooks/deprecated/ has 10 files + README ===" +ls runbooks/deprecated/ | wc -l +# Expect: 11 (10 deprecated runbooks + README.md) + +echo "=== 11. YAML still parses ===" +python3 -c "import yaml; yaml.safe_load(open('bundle.yaml')); print('[OK] YAML parses')" +``` + +--- + + +## 11. Acceptance criteria + +- [ ] All 9 commits landed and pushed to `main` +- [ ] Verification section 10 returns `[OK]` for all 11 checks +- [ ] `git log --oneline` shows the 9 commits in order +- [ ] `bundle.yaml` parses cleanly via `python3 -c "import yaml; yaml.safe_load(open('bundle.yaml'))"` +- [ ] No untracked changes (`git status` clean except for any local scratch) + +Once all checked, proceed to `v1-do-doc-01-prep.md` execution. + +--- + + +## 12. Change log + +| Date | Change | Reference | +|---|---|---| +| 2026-05-27 | v1 (six commits): ceph-osd storage block; D-002/D-014 cleanups; README refresh; deprecated runbook moves; deprecation README | Initial drafting | +| 2026-05-27 | v2 (nine commits): + Designate deferral commits 7/8/9; amended commit 5/6/10 to reflect 07-dns-zones permanent deprecation; updated commit 4 README hunks; VIP count expectations updated from 12 to 11 | D-019 decision; testcloud topology investigation | diff --git a/docs/changelog-20260719-phase5-sweep.md b/docs/changelog-20260719-phase5-sweep.md index a917614..4b85f91 100644 --- a/docs/changelog-20260719-phase5-sweep.md +++ b/docs/changelog-20260719-phase5-sweep.md @@ -222,3 +222,30 @@ whitespace/punctuation-only and content-faithful (this note is the record of it). No gauntlet run for this batch: zero script/tooling changes (memory + docs + one skill prose line); repo-lint 0-fail + fence check OK. + +## Batch 4 -- consolidation + rotation (OPENED by operator 2026-07-19) + +- **C1 ledger rotation (GA-R4 rule 6, GA-F13)**: 1187 -> 131 lines; bodies + verbatim to docs/archive/session-ledger-rotated-20260719.md; lessons/facts + ROUTED first (format=line trap -> appendix-A; guard discipline -> + operating-discipline; VR0 facts + admin/controller -> maas-as-built; + lessons 1-6 verified already covered). F1 cap now enforceable. +- **C2-C5 changelog consolidation (GA-R2, GA-F09)**: 96 files (95 + the + v1-redeploy running record) -> docs/archive/changelogs/, four stage + records under docs/archive/stage-records/ (v1-era; stage0-1 = 45; stage2 + = 38; stage3-partial = 11), each commit message MANIFESTS its sources + (D2). The live session changelog stays top-level until session close. + NOTE: v1-redeploy-changelog's own "completion consolidation" TODO is NOT + discharged by the move -- it remains in the ledger's project-completion + block. +- **C6 docs disposition (D2 scope)**: 24 non-changelog history docs -> + docs/archive/ (v1-era packs/handoffs/findings/drafts; executed plans; + the R3-F register; Model-A fallback; DR seed; ipv6/netem research; + D-068-openbao DRAFT). RETAINED top-level, 16 files each justified: + status authority, registers (D/SEC), ledger, workflow (identity), design + doc, readiness checklist, two as-builts, tenant contract (lint L7 + auto-covered), log convention, live finding doc (CURRENT-STATE-cited), + D-129 review (open sub-decisions), NetBox scope (live apex reference), + D-068 analysis (open-decision evidence), live session changelog. + Live-surface references to every moved file rewritten to archive paths + (18 files touched); history-internal refs left verbatim. diff --git a/docs/clientdocs-workflow-review-20260708.md b/docs/clientdocs-workflow-review-20260708.md deleted file mode 100644 index b5c1aa2..0000000 --- a/docs/clientdocs-workflow-review-20260708.md +++ /dev/null @@ -1,190 +0,0 @@ -# clientdocs workflow review + consolidation proposal (2026-07-08) - -STATUS: PROPOSAL for operator ruling. Read-only analysis; nothing implemented. -Options are presented per finding; the operator rules before any change. No -identifier numbers are consumed here -- they are assigned at implementation. - -## Why this exists - -Operator's stated problem (2026-07-08): the information a client needs is spread -throughout the handover documentation with no clear flow. A client has to switch -between multiple documents to accomplish one task, and some workflows have holes -that cause errors (the flannel network-driver gap being the live example: a -self-created flannel cluster never converges, and devteam lost a night to it). -The goal for the client-facing set: an easy-to-follow, workflow-shaped path that -makes finding and using their data simple, minimizes document-switching, and -carries enough component detail that the client understands what each piece -does. - -DOCFIX-123 (the new jenkins-kubernetes-guide.md) is the first instance of the -target shape -- a single self-contained workflow that cross-links to detail -rather than scattering it. This review proposes applying the same principle -across the rest of the set. - -## Method - -Read-only inventory of all 20 client-facing files under clientdocs/ (the six -top-level guides, the tenant-skill SKILL.md + four references, and the six -starter-kit scripts), mapping every workflow topic to the file(s) that cover it, -plus duplication, cross-file conflicts, and gaps. Evidence is file:line -throughout the source inventory; the summary below carries representative cites. - -## The client journey (what a client actually does, in order) - -1. Pre-onboarding homework (intake-form.md). -2. Receive credentials + first orientation (welcome.md, handover-pack.md). -3. Prove the tenancy works (acceptance-checklist.md). -4. Run day-to-day (self-service-guide.md; tenant-skill references). -5. Connect automation / CI (ci-integration-guide.md). -6. Stand up a Kubernetes + Jenkins workflow (jenkins-kubernetes-guide.md, new). - -The docs exist for every stage, but stages 2-6 repeat each other's content and -none of them states "you are here, read these in this order." - -## Findings - -### F1 -- Heavy duplication (maintenance-drift + reader confusion) - -The same facts are authored independently in many files. Representative: - -- The three-account model table: 4 renderings (handover-pack.md:28-32, - self-service-guide.md:10-14, tenant-skill/SKILL.md:32-36, and prose in - welcome.md). -- The clouds.yaml auth block: 3 byte-identical copies (ci-integration-guide.md, - tenant-skill/SKILL.md, scripts/clouds.yaml.template). -- "Cluster create cannot use an application credential": stated in 8 places. -- The worked CI sequence: duplicated at length in ci-integration-guide.md and - references/ci-automation.md. -- Floating-IPs-count-when-detached: repeated in 5 files. -- (Full list: 15 duplication clusters in the source inventory.) - -Risk: any change (an endpoint, an account name, a rule) must be made in N places -or the copies drift. It also makes each document longer than it needs to be, -which is itself part of the "hard to follow" problem. - -### F2 -- No single entry point or reading order - -No file tells the client which document to open first, or the order for a given -goal. A client wanting "deploy my app to Kubernetes from Jenkins" previously had -to assemble it from kubernetes.md + ci-integration-guide.md + day2-operations.md -+ troubleshooting.md. (DOCFIX-123 now owns that one journey; the other journeys -still lack a spine.) - -### F3 -- Workflow holes that cause errors - -- Network-driver choice: the self-service path (self-service-guide.md:89-99) - never mentioned calico, so a client building a template via the dashboard got - no steer away from flannel. CLOSED 2026-07-08 in kubernetes.md + the Jenkins - guide; self-service-guide.md still only carries the Public/Hidden rule, not - the driver rule. -- Right-sizing to host capacity: every doc presents quota as the only ceiling on - node_count (self-service-guide.md:91, kubernetes.md:49, welcome.md, intake). - None warned that an under-quota-but-oversized request fails with "No valid - host" -- the exact failure devteam hit. CLOSED in the Jenkins guide; NOT yet - in the general docs. - -### F4 -- Partial component detail / no glossary - -Component terms (application credential, the -cluster account, cluster template, -kubeconfig, LoadBalancer Service, floating IP, ingress) are defined inline where -first used but there is no single glossary a client can consult. The operator -specifically asked for "enough detail so they understand what all the components -do." The Jenkins guide added a scoped glossary (section 10); the set has no -shared one. - -### F5 -- Cross-file inconsistencies - -- Jenkinsfile.example calls scripts at clientdocs-starter-kit/... while the kit - is delivered as scripts/ (README.md:27-33). The file admits paths need - adjusting, but it is a stumble. -- Jenkinsfile.example's "Worked sequence" stage runs an infrastructure smoke - build, not an app deploy -- misleading under that name for a k8s reader. - -## Proposed information architecture - -Principle: each fact has ONE owner; every other mention is a one-line pointer, -not a copy. Each guide owns a JOURNEY (a workflow) and cross-links to the owners -for reference detail. This is the shape DOCFIX-123 already demonstrates. - -Proposed owners (single source of truth): -- Identifiers, endpoints, CA bundle, the full account model: handover-pack.md - (already the most complete; make it canonical). -- clouds.yaml / OS_* auth setup: scripts/clouds.yaml.template (the delivery-ready - copy) + one reference section; others point to it. -- Day-2 resource operations (networks, servers, load balancers, secrets): - references/day2-operations.md. -- Troubleshooting signatures: references/troubleshooting.md (the single triage - ladder; other docs link symptoms to it). -- Journeys (own the flow, link the detail): welcome (orientation), - self-service-guide (day-to-day), acceptance-checklist (proof), - ci-integration-guide (automation), jenkins-kubernetes-guide (k8s+Jenkins). - -Add: -- A top-of-set "Start here / reading order" -- either a short new index doc or a - section in welcome.md -- mapping goal -> ordered doc list. -- A shared glossary -- a new small reference, or a section in handover-pack.md -- - that every guide links to (retire the per-guide inline definitions in favor of - one, keeping only a one-line gloss at first use + a link). - -## Consolidation moves (OPTIONS for operator ruling) - -Each move is independent; rule per-move. Effort/risk are rough. - -- M1 -- De-duplicate the account model. Own it in handover-pack.md; replace the - other 3 renderings with a 1-line summary + pointer. - Options: (a) do it; (b) keep the self-service copy (first-read convenience) but - retire the SKILL.md + welcome copies; (c) leave as-is. - Effort: low. Risk: low. Recommend (b) -- one convenience copy, not four. - -- M2 -- De-duplicate clouds.yaml/auth. Own in clouds.yaml.template; the CI guide - and SKILL.md point to it. - Options: (a) do it; (b) leave. Effort: low. Risk: low. Recommend (a). - -- M3 -- Add a "Start here / reading order" entry point. - Options: (a) new short index doc; (b) a section in welcome.md; (c) none. - Effort: low. Risk: low. Recommend (b) -- no new file to package. - -- M4 -- Shared glossary. - Options: (a) new reference file linked from all guides; (b) a section in - handover-pack.md; (c) leave per-guide inline. Effort: medium (touches many - files if links are added). Risk: low. Recommend (a). - -- M5 -- Close the remaining F3 holes in the general docs: add the driver rule and - the host-capacity/right-size caveat to self-service-guide.md (and anywhere - node_count is discussed). - Options: (a) do it; (b) rely on the Jenkins guide + kubernetes.md only. - Effort: low. Risk: low. Recommend (a) -- these are the error-causing holes the - operator called out. - -- M6 -- Fix F5: correct Jenkinsfile.example's script paths to scripts/ and either - rename or re-scope its misleading stage; consider adding a second - Jenkinsfile.example for the kubeconfig-deploy pattern (or point at the Jenkins - guide's inline pipeline). Effort: low-medium. Risk: low (starter asset). - Recommend: fix paths now; the deploy-pipeline example already lives in the - Jenkins guide, so just cross-link. - -- M7 -- Reconcile the quota-vs-capacity expectation cloud-side (separate from - docs): the addendum-40 finding that testcloud tenant quotas over-promise - schedulable capacity. Doc side is M5; the cloud side (trim quotas to reality) - is an operator decision logged in addendum 40. - -## Recommended sequencing - -1. M5 first (closes the active error-causing holes; low effort). -2. M3 + M1(b) (entry point + de-dup the worst offender; makes the set navigable). -3. M2, M6 (mechanical de-dup + the starter-asset fix). -4. M4 (glossary; larger touch, do once the owners above are settled). - -Each move ships under the standard clientdocs discipline (ASCII/LF, no -internal-term leakage, L7 receipt re-record, harness green, changelog + revert) -and, where a delivered client already has the affected file, a package -re-instantiation. Numbers are assigned at implementation, not here. - -## Not in this proposal - -- The broad rewrite into a single mega-document. The operator's goal is - minimize doc-switching PER TASK, which the journey-owns-flow + cross-link model - achieves without collapsing the set into one unnavigable file. If the operator - prefers fewer, larger documents, that is a different architecture to rule on. -- Any client-specific content. This is about the template set; per-client - instantiation is unchanged. diff --git a/docs/dc-dc-deployment-workflow.md b/docs/dc-dc-deployment-workflow.md index 6e89f96..54ce425 100644 --- a/docs/dc-dc-deployment-workflow.md +++ b/docs/dc-dc-deployment-workflow.md @@ -27,7 +27,7 @@ stated lean, not a rubber-stamp); COS per-DC-only; mirror sync independent per-DC. Verified via `bash scripts/ledger-scan.sh`: none of D-100..D-110 appear under "PROPOSED / OPEN decisions." Full ruling record: -`docs/changelog-20260709-stage0-ratification.md`. +`docs/archive/changelogs/changelog-20260709-stage0-ratification.md`. **State:** cleared 2026-07-09 (history; status authority: docs/CURRENT-STATE.md). @@ -138,8 +138,8 @@ | # | Item | State | |---|---|---| -| C1 | Office1 LAN IPv6 deployed (`2602:f3e2:f01:100::1/64` + RA on the edge LAN). WAN v6 leg DEFERRED: v6 does not egress the lab. | **done 2026-07-14** -- live and PROVEN ON THE WIRE: address on the kernel, radvd advertising the prefix (`mode=unmanaged`), `voffice1` autoconfigured `...:5054:ff:fe6a:87e5` by SLAAC, edge->voffice1 GUA ping 0.0% loss. NOTE the D-113 AMENDMENT: the ADDRESS half is **not** REST-API-doable (`scripts/opnsense-set-interface-v6.sh`); only the RA half is. `docs/changelog-20260714-office1-lan-ipv6-executed.md` | -| C2 | `office1-netbox` (10.10.1.10) is the VR1 IPAM apex -- a working draft (DOCFIX-195) -- and holds the complete validated dataset: D-115 carve + six-plane roles + both DCs' prefixes/aggregates/RIRs, PLUS the D-120 additions -- `10.10.1.0/24`'s child ranges (`ip-ranges`: static .2-.49 / dynamic .100-.200 / node .201-.254) and the two service IP assignments (`ip-addresses`: office1-netbox 10.10.1.10, office1-tailscale 10.10.1.11). `netbox.baldurkeep.com` stays a READ-ONLY v1 reference draft; the write-back to it is DEFERRED to end-of-deployment (DOCFIX-195, operator ruling 2026-07-15) -- NOT a Stage 2 gate. | **done 2026-07-15.** (i) The HARDENED fidelity check was re-run against office1-netbox -> exit 0 (faithful); it also surfaced + fixed a checker scope false-positive (`docs/changelog-20260715-fidelity-check-scope-correctness.md`). (ii) The tooling gap was CLOSED -- `prod-draft-dump.py` + `sandbox-fidelity-check.py` now cover `ip-ranges`/`ip-addresses` -- and the D-120 objects (3 bands + 2 service IPs) were loaded into office1-netbox via `netbox/d120-compose-bands.py --commit` and VERIFIED: re-dump + fidelity check -> exit 0, "EXACTLY the planned delta (D-115/D-117/D-118/D-120)." `docs/changelog-20260715-d120-load-and-iprange-tooling.md`. Gauntlet ALL GREEN (60). | +| C1 | Office1 LAN IPv6 deployed (`2602:f3e2:f01:100::1/64` + RA on the edge LAN). WAN v6 leg DEFERRED: v6 does not egress the lab. | **done 2026-07-14** -- live and PROVEN ON THE WIRE: address on the kernel, radvd advertising the prefix (`mode=unmanaged`), `voffice1` autoconfigured `...:5054:ff:fe6a:87e5` by SLAAC, edge->voffice1 GUA ping 0.0% loss. NOTE the D-113 AMENDMENT: the ADDRESS half is **not** REST-API-doable (`scripts/opnsense-set-interface-v6.sh`); only the RA half is. `docs/archive/changelogs/changelog-20260714-office1-lan-ipv6-executed.md` | +| C2 | `office1-netbox` (10.10.1.10) is the VR1 IPAM apex -- a working draft (DOCFIX-195) -- and holds the complete validated dataset: D-115 carve + six-plane roles + both DCs' prefixes/aggregates/RIRs, PLUS the D-120 additions -- `10.10.1.0/24`'s child ranges (`ip-ranges`: static .2-.49 / dynamic .100-.200 / node .201-.254) and the two service IP assignments (`ip-addresses`: office1-netbox 10.10.1.10, office1-tailscale 10.10.1.11). `netbox.baldurkeep.com` stays a READ-ONLY v1 reference draft; the write-back to it is DEFERRED to end-of-deployment (DOCFIX-195, operator ruling 2026-07-15) -- NOT a Stage 2 gate. | **done 2026-07-15.** (i) The HARDENED fidelity check was re-run against office1-netbox -> exit 0 (faithful); it also surfaced + fixed a checker scope false-positive (`docs/archive/changelogs/changelog-20260715-fidelity-check-scope-correctness.md`). (ii) The tooling gap was CLOSED -- `prod-draft-dump.py` + `sandbox-fidelity-check.py` now cover `ip-ranges`/`ip-addresses` -- and the D-120 objects (3 bands + 2 service IPs) were loaded into office1-netbox via `netbox/d120-compose-bands.py --commit` and VERIFIED: re-dump + fidelity check -> exit 0, "EXACTLY the planned delta (D-115/D-117/D-118/D-120)." `docs/archive/changelogs/changelog-20260715-d120-load-and-iprange-tooling.md`. Gauntlet ALL GREEN (60). | | C3 | The `$DC` shell-selector collision resolved. | **done 2026-07-14 (D-119)** -- selectors are REGION-QUALIFIED (`vr0-dc0`/`vr1-dc0`/`vr1-dc1`) across `lib-net.sh`, `lib-hosts.sh`, `opentofu/`, the importer and every runbook call site. Bare `dcN` is REJECTED. The importer's DC->site map is now an IDENTITY, so the off-by-one is structurally impossible. Gap #19 CLOSED. | | C4 | The NetBox sandbox loop documented. | **done 2026-07-14** -- `docs/dc-dc-netbox-buildout-scope.md` section 8 (rationale + write-guard table + WAF trap) and `runbooks/dc-dc-phase1-office1-standup.md` Step 10b (operational sequence). | | C5 | This doc's Stage 2 row reconciled against reality (it has drifted three times). | **done 2026-07-14** | @@ -172,7 +172,7 @@ **D-125** (bridge-in) and node PLACEMENT by **D-123**; D-122's intent stands. - **D-123 (ADOPTED Model B)** -- DC nodes nest INSIDE `vvr1-dc0`; a single `virsh destroy vvr1-dc0` = site-down. Forces the TWO-ROOT split (a libvirt provider can't target a same-apply VM). Overturns the - earlier inferred Model A (preserved as `docs/model-a-fallback-plan.md` + git tag `model-a-fallback`). + earlier inferred Model A (preserved as `docs/archive/model-a-fallback-plan.md` + git tag `model-a-fallback`). - **D-124 (ADOPTED)** -- rack transit addressing (office1<->dc0 mesh /30, Scheme A); the entry's rack SIZING is VOID under Model B (`vvr1-dc0` now holds the fleet). - **D-125 (ADOPTED)** -- per-DC ISP egress via bridge-in (resolves OBS-3's design gap; egress efficacy is a @@ -234,7 +234,7 @@ | **Build** | Replication-plane peering (ULA); radosgw multisite zonegroup (two DCs as zones); rbd-mirror for the Glance pool; cinder-backup to the replicated object store; consistency groups where needed. Run the failover/failback drill. | | **Gate** | One-way drill clean (RTO/RPO measured); then two-way; controller restore drill (D-104) exercised; RTO recorded as a measured output, not an SLA. | | **Owns** | D-108 (replication), D-104 (control-plane backup/restore in the drill). | -| **Reuse vs new** | Ground truth: `docs/dc-dc-replication-DR-seed.md` (the original Tier-A/bidirectional/cinder-backup+rbd-mirror decision) -- D-108 corrected its carrier from the seed's v4 replication plane to IPv6-only ULA. The failover/failback drill skeleton is already drafted (buildout design Section 8): confirm DC1 down -> promote Glance -> restore Cinder -> rebuild instances -> record RTO/RPO; failback is split-brain-safe (no auto-resume, controlled window to flip primary back). | +| **Reuse vs new** | Ground truth: `docs/archive/dc-dc-replication-DR-seed.md` (the original Tier-A/bidirectional/cinder-backup+rbd-mirror decision) -- D-108 corrected its carrier from the seed's v4 replication plane to IPv6-only ULA. The failover/failback drill skeleton is already drafted (buildout design Section 8): confirm DC1 down -> promote Glance -> restore Cinder -> rebuild instances -> record RTO/RPO; failback is split-brain-safe (no auto-resume, controlled window to flip primary back). | | **Authoring status** | **Runbook WRITTEN 2026-07-09: `runbooks/dc-dc-phase5-dr-failover-drill.md`.** Command-level implementation of Section 8's failover/failback skeleton, staged one-way-then-two-way per D-108's own requirement, plus the D-104 controller backup/restore drill. Found and flagged three real bundle gaps while writing it (grepped `bundle.yaml`): no `cinder-backup` charm deployed, no `ceph-rbd-mirror` charm deployed, and `ceph-mon`'s existing `rbd-mirror` endpoint binding is on `storage` instead of `replication` (inherited from the VR0/DC0 seed, predates D-108) -- all three become this runbook's own first gated mutation (Step 2), not silently routed around. Every realm/zone/pool/endpoint name is an explicit `` placeholder. Includes a negative-control sub-drill (sever only the replication link first, confirm the split-brain check correctly refuses to promote) before the real hard-down drill. NOT YET EXECUTED -- no live Ceph cluster this session (harness for gap #5's scripts is future work, logged, not built). | **State:** runbook written, not yet executed (status authority: docs/CURRENT-STATE.md). @@ -324,8 +324,8 @@ > > The as-built path that actually worked: prep image -> boot on factory defaults -> **D-112(c)** > console bootstrap -> enable SSH + install key -> configure over SSH and the REST API. See -> `docs/changelog-20260713-opnsense-api-proven.md` and -> `docs/changelog-20260713-opnsense-api-write-proven.md`. +> `docs/archive/changelogs/changelog-20260713-opnsense-api-proven.md` and +> `docs/archive/changelogs/changelog-20260713-opnsense-api-write-proven.md`. > > **done (D-113(a2)) -- reconciled 2026-07-16; the repo won over this register's stale "OPEN".** The > config-ISO path and the module's `config_seed`/cdrom wiring are RETIRED: `modules/opnsense-edge` no @@ -349,7 +349,7 @@ gained `lib_hosts_select_dc()` (dc0 no-op; dc1 AND dc2 BOTH fail loud -- no real per-DC host enrollment exists yet for either, an intentional asymmetry vs. the net selector, documented in both files' comments and - `docs/changelog-20260709-dc-selector-convention.md`). Backward compatible + `docs/archive/changelogs/changelog-20260709-dc-selector-convention.md`). Backward compatible by construction -- every existing caller of either library is unaffected. New harness `tests/dc-selector/run-tests.sh`: 21/21 PASS. The mechanism exists; Stage 5's runbook still needs to actually CALL @@ -378,7 +378,7 @@ the gate now validates EVERY module standalone (S3). **`node-vm` builds every DC node and `netem-link` is the Stage 6 failover-drill mechanism** -- they would have detonated mid-deploy. See - `docs/changelog-20260713-docfix194-opentofu-module-validation.md`. + `docs/archive/changelogs/changelog-20260713-docfix194-opentofu-module-validation.md`. - **`modules/cloudinit-vm` IS NOW INSTANTIATED FOR REAL (2026-07-13, D-114).** The long-standing "the mechanism exists but no image source is chosen and no `user_data`/`meta_data`/`network_config` has been designed for it" caveat -- @@ -523,7 +523,7 @@ pipeline). The MECHANISM exists; the DATA half (real DC1/DC2 Ceph cluster, real names/endpoints, the three `bundle.yaml` gaps in the Stage 6 runbook's own "Known gaps" section) is unchanged and still blocks Stage - 6 for real. See `docs/changelog-20260709-ceph-replication-tooling.md`. + 6 for real. See `docs/archive/changelogs/changelog-20260709-ceph-replication-tooling.md`. 6. **Designate reactivation is a bundle change -- CLOSED 2026-07-10 (DOCFIX-167).** `bundle.yaml` now carries `designate`/`designate- bind`/`designate-mysql-router`/`designate-hacluster` + 8 relations, @@ -534,7 +534,7 @@ the shared bundle -- pushed to a proposed per-DC overlay instead, to keep `bundle.yaml` DC-agnostic. `python3 scripts/provider-bundle- check.py` PASS, all 6 existing invariants unaffected. See - `docs/changelog-20260709-designate-cinderbackup-rbdmirror.md`. NOT + `docs/archive/changelogs/changelog-20260709-designate-cinderbackup-rbdmirror.md`. NOT applied to any live model -- repo-source change only. 7. **MTU/geneve budget + Ceph disk-budget calculators -- CLOSED 2026-07-09 (DOCFIX-162).** D-102 and buildout-design Section 3 both describe @@ -560,7 +560,7 @@ run against the real vcloud host (prep-only session, no live infrastructure reachable) -- that first real run is part of executing `runbooks/dc-dc-phase0-vcloud-prep.md` Step 3. Changelog: - `docs/changelog-20260709-mtu-ceph-budget-calculators.md`. + `docs/archive/changelogs/changelog-20260709-mtu-ceph-budget-calculators.md`. 8. **The "developed IPv6 plan" D-101 cites does not exist in this repo.** D-101's context line references "the developed family-follows-reachability IPv6 plan (in-project, not yet committed)" as prior art -- grepped the whole repo, it is @@ -590,7 +590,7 @@ validate.sh` itself (no `tofu` binary this session to validate against). 11. **Decision sub-items still genuinely open** (the big items are RULED per Stage 0; these specific literals/params are not -- but see - `docs/dc-dc-netem-and-ula-gua-proposal.md` (DOCFIX-168) for a concrete + `docs/archive/dc-dc-netem-and-ula-gua-proposal.md` (DOCFIX-168) for a concrete netem proposal and exact, safe ULA-generation guidance, drafted 2026-07-10): exact netem parameters (D-100), exact NetBox literals for DC2/ULA/GUA (D-101). @@ -610,7 +610,7 @@ (`module "office1_network"`) -- unlike most of gap #2's other scaffolding, this needed no unmeasured value beyond `domain_suffix`/`underlay_mtu`, already real required inputs elsewhere in the same file. See - `opentofu/README.md` and `docs/changelog-20260709-office1-network-edge.md` + `opentofu/README.md` and `docs/archive/changelogs/changelog-20260709-office1-network-edge.md` (DOCFIX-163) for the full design rationale. **Follow-on work also done 2026-07-13 (D-114):** the "designing the service VMs' actual `cloudinit-vm` `network_config` against this network" tail of this gap is closed, but not @@ -665,7 +665,7 @@ `carve-host-interfaces.sh`, and `phase-00-maas-standup.sh` now accept an opt-in `$DC` env var and call `lib_net_select_dc`/`lib_hosts_select_dc` immediately after sourcing (unset `$DC` is byte-for-byte unchanged - behavior; see `docs/changelog-20260709-maas-scripts-dc-param.md`). This + behavior; see `docs/archive/changelogs/changelog-20260709-maas-scripts-dc-param.md`). This closes the CLI-invocation half of the gap only -- it does NOT create any new per-DC data (no dc1/dc2 hostnames, octets, boot MACs, or node-count figures were invented), and `phase-00-maas-standup.sh`'s own `PLANES` @@ -850,7 +850,7 @@ (never a baked `virbrN`, PATTERN-1); `apply` idempotent; `check` = session-verify. INSTALLED + `enabled` on vcloud 2026-07-17; the `After=libvirtd`+poll BOOT-RACE stays UNPROVEN until an actual reboot self-heals `ssh voffice1`. See D-126 amendment + - D-128; `docs/changelog-20260717-d126-baseleg-oneshot-d128-operating-model.md`. + D-128; `docs/archive/changelogs/changelog-20260717-d126-baseleg-oneshot-d128-operating-model.md`. **DC generalization (so this gap does NOT reopen per-DC):** the tool is SITE-KEYED with a commented DC template row, and the harness FAILS on any un-commented DC @@ -879,7 +879,7 @@ | MAAS VM-host registration (`modules/maas-vm-host`) | BUILT 2026-07-09, UNVALIDATED -- needs a real MAAS zone/pool + vcloud power_address, see `opentofu/README.md` | | `tc netem` mechanism (`modules/netem-link`) | **WAS BROKEN, FIXED 2026-07-13 (DOCFIX-194)** -- its destroy provisioner referenced `var.*`, which OpenTofu rejects at INIT, so the module could not even initialize. Now validates. Real latency/jitter/loss/rate parameters are STILL an unruled D-100 item. Never applied. | | `opentofu/` (DC2 planes) | not done -- deliberately deferred pending NetBox CIDR assignment, see `opentofu/README.md` | -| Office1-local network (`modules/office1-network`) | BUILT + INSTANTIATED 2026-07-09 (DOCFIX-163, gap #12 CLOSED) -- see `opentofu/README.md` and `docs/changelog-20260709-office1-network-edge.md` | +| Office1-local network (`modules/office1-network`) | BUILT + INSTANTIATED 2026-07-09 (DOCFIX-163, gap #12 CLOSED) -- see `opentofu/README.md` and `docs/archive/changelogs/changelog-20260709-office1-network-edge.md` | | Office1 OPNsense edge ownership | **BUILT AND LIVE 2026-07-12/13.** `module "office1_opnsense"` is instantiated (not commented). **Gap #17 is CLOSED FOR OFFICE1**: `office1-wan` (virbr11, NAT, 172.30.1.0/24) is the dedicated per-site ISP uplink, and the edge's WAN sits on it -- the D-100-honouring answer (NOT a mesh leg). **DC1/DC2 still have NO uplink network** -- replicating the `office1-wan` shape per DC is now mechanical, not a decision. | | Reusable as-is, repo-agnostic (2026-07-09 sweep) | `bundle.yaml`, `phase-01..08` VR0 DC0 runbooks (Stage 5/7 template, once DC-parameterized -- gap #1), `preflight.sh`, `cloud-assert.sh`, `repo-lint.sh`, `run-logged.sh`, `ledger-scan.sh` | | `lib-net.sh` / `lib-hosts.sh` | `$DC` selector mechanism ADDED 2026-07-09 (DOCFIX-151, gap #1 CLOSED) -- `lib_net_select_dc()`/`lib_hosts_select_dc()`, backward compatible, 21/21 tests; Stage 4 and Stage 5's runbooks (both written 2026-07-09) mandate calling it, not yet executed for real | diff --git a/docs/dc-dc-ipv6-charm-research.md b/docs/dc-dc-ipv6-charm-research.md deleted file mode 100644 index 9c48ba2..0000000 --- a/docs/dc-dc-ipv6-charm-research.md +++ /dev/null @@ -1,142 +0,0 @@ -# D-101 IPv6 family-matrix charm research (2026-07-09/10) - -Basis for `overlays/dc-dc-ipv6-family-matrix.yaml` (tooling gap register item -#13). Researched directly against real charm source/config via WebFetch/ -WebSearch -- not inferred from training-data memory of charm behavior, which -this repo's own discipline treats as a real risk class (the OpenTofu module -work earlier this session found a real syntax bug from exactly that mistake). - -## Method - -Fetched actual `config.yaml` files and, where a config option's semantics -were ambiguous from prose alone, actual charm source/commits, from the -`openstack/charm-*` and `openstack-charmers/charm-layer-ovn` repositories -(GitHub mirrors of the OpenDev-hosted canonical source) and -`bugs.launchpad.net`. Every claim below is labeled CONFIRMED (read the -actual file/commit) or INFERRED-BY-PATTERN (not independently checked, -reasoned from a shared framework) -- do not treat the two as equal -confidence. - -## Findings - -### 1. `prefer-ipv6` is a real, shared charms.openstack-family config option - -**CONFIRMED** directly in `charm-keystone/config.yaml`: a boolean, default -`False`, enabling IPv6 support. - -**CONFIRMED** via `charm-nova-cloud-controller`'s actual "Dual Stack VIPs" -commit (`1fa5f7a673006bc6faaf9d377e791448a5ed661b`): setting `prefer-ipv6: -true` causes the charm's HAProxy config template to add `bind -:::{{ports[0]}}` ALONGSIDE the existing `bind *:{{ports[0]}}` -- i.e. this -is a genuinely ADDITIVE dual-stack switch (HAProxy listens on both address -families simultaneously), not an either/or toggle. The commit message -states this explicitly: "HAProxy always listens on both IPv4 and IPv6 -allowing connectivity on either protocol." - -**INFERRED-BY-PATTERN, NOT independently confirmed per-charm:** glance, -neutron-api, placement, cinder, barbican, magnum, openstack-dashboard, -ceph-radosgw all build on the same charms.openstack base classes as -keystone/nova-cloud-controller, so the same `prefer-ipv6` option and -additive-dual-stack HAProxy behavior is EXPECTED -- but each charm's real -`config.yaml` was not individually fetched this session. Confirm via `juju -config ` once any of these charms is actually deployed, before -trusting the overlay is complete for all eleven. - -### 2. The existing `vip:` option takes additional v6 addresses, appended - -**CONFIRMED**: `charm-keystone/config.yaml`'s `vip` option description: -"Virtual IP(s) to use to front API services in HA configuration. If -multiple networks are being used, a VIP should be provided for each -network, separated by spaces." This repo's `bundle.yaml` already uses this -exact mechanism today (D-020's triple-VIP pattern -- e.g. `keystone`'s `vip: -"10.12.4.50 10.12.8.50 10.12.12.50"`, one address per Juju-space binding). -Adding v6/GUA addresses to the SAME space-separated list, once -`prefer-ipv6: true` makes HAProxy listen on both families, is the -mechanism -- confirmed consistent with finding #1's dual-stack commit, -though the EXACT mapping of "which list position pairs with which network" -inside the charm's own context-building code was not traced line-by-line -this session. Treat the overlay's VIP ordering (existing v4 entries first, -then the new v6/GUA entries appended) as the reasonable, but not -byte-for-byte charm-source-confirmed, convention. - -### 3. `ceph-mon` has a real, separate `prefer-ipv6` + network-CIDR mechanism - -**CONFIRMED** directly in `charm-ceph-mon/config.yaml`: -- `prefer-ipv6` (bool, default False) -- per the charm's OWN description, - "If set to False (default) IPv4 is expected," reading as a straight - EITHER/OR switch for THIS charm specifically (unlike the additive - HAProxy behavior in finding #1) -- which is actually the correct shape - for D-101's ULA-ONLY storage/replication planes (no v4 needed at all, - so a clean switch to v6-only is exactly right, not a limitation). -- `ceph-public-network` / `ceph-cluster-network` (string, space-delimited - CIDR list) -- these map directly to D-101's own cited mechanism name, - `ms_bind_ipv6` (the underlying Ceph daemon config the charm sets when - `prefer-ipv6` is enabled). -- A real caveat quoted from the charm's own docs: "these charms do not - currently support IPv6 privacy extension. In order for this charm to - function correctly, the privacy extension must be disabled and a - non-temporary address must be configured/available on your network - interface." This is a real host/OS-level prerequisite for every ceph-mon - unit's storage/replication interface -- not something an overlay file - can set, needs to be part of Phase-0/node-prep discipline. - -### 4. OVN has NO explicit IPv6/encapsulation charm-config option - -**CONFIRMED**: fetched the FULL `charm-layer-ovn/config.yaml` (the shared -base for `ovn-central`/`ovn-chassis`) -- 25 options total, none related to -IPv6, address binding, or encapsulation IP. geneve-over-v6 for the -data-tenant plane is therefore not a charm-config concern at all; it -follows whichever IP family the unit's bound interface on that plane -actually has. Since data-tenant is ULA-only (no v4 present) per D-101, this -needs no overlay entry -- the absence of an entry IS the correct -configuration, not an oversight. - -### 5. Vault's cert-issuance code has no IPv4/IPv6 distinction - -**CONFIRMED**: fetched `charm-vault/src/lib/charm/vault_pki.py`'s actual -`sort_sans()` function -- it splits SANs into "IP SANs" vs. "name SANs" -using `charmhelpers.contrib.network.ip.is_ip()`, with NO further IPv4/IPv6 -differentiation anywhere in the function or its caller -(`generate_certificate()`, which just joins whatever IP SANs it's given -into a comma-separated string). D-109's "Vault issues v4+IPv6-SAN certs" -requirement is technically supported by the charm's own code path, not -merely assumed. - -### 6. A real, open upstream risk for Octavia's lb-mgmt-net over IPv6 - -**CONFIRMED** (real, still-referenced Launchpad bug reports, not -speculation): #1911788 ("IPv6 mgmt network not working, octavia can't talk -to amphora," OpenStack Octavia Charm) and #1913409 ("octavia_amp_network -does not support IPv6," kolla-ansible, describing the same underlying -Octavia limitation from a different deployment tool). #1911788's root cause -traces to an OVN/LXD/MAAS hostname-resolution mismatch (Neutron's -`binding_host_id` carrying a short hostname while OVN's controller expects -the FQDN) -- it was marked a duplicate of #1896630, "managing /etc/hosts -for containers." That is the SAME CLASS of problem this repo's own D-008 -bootstrap order (static /etc/hosts bootstrap -> os-public-hostname -> vault -certs -> Designate, reactivated by D-106 for VR1) already exists to -harden against -- so this is NOT necessarily a fatal blocker here, but it -IS a real, open, documented upstream risk for exactly the kind of network -(Octavia's amphora management network) D-101 wants to put on ULA-only IPv6. -`overlays/dc-dc-ipv6-family-matrix.yaml` deliberately does NOT attempt an -Octavia lb-mgmt-net IPv6 change for this reason -- that needs its own -explicit decision (accept the risk and test for real once DC1 exists, or -keep Octavia's lb-mgmt-net dual-stack/v4 as a deliberate, logged D-101 -exception), presented here rather than silently forced through. - -## What this closes vs. what remains open - -CLOSES (mechanism): the "no overlay file exists, charm option names -unconfirmed" half of gap #13 -- a real, sourced overlay now exists at -`overlays/dc-dc-ipv6-family-matrix.yaml`, covering 9 of Octavia's 11 sibling -dual-stack API charms directly plus ceph-mon's ULA-only switch, with OVN -correctly requiring no entry. - -STAYS OPEN: (a) the 8 INFERRED-BY-PATTERN charms' `prefer-ipv6` option -needs live `juju config` confirmation once any of them is actually -deployed; (b) the exact VIP-list-position-to-network mapping inside each -charm's context-building code was not traced to source; (c) Octavia's -lb-mgmt-net IPv6 question is a real, open decision, not resolved by this -research; (d) none of this has been applied or tested against a live -model -- UNVALIDATED, same posture as every other OpenTofu/overlay artifact -built this session. diff --git a/docs/dc-dc-netbox-buildout-scope.md b/docs/dc-dc-netbox-buildout-scope.md index 0b4e323..0d0a419 100644 --- a/docs/dc-dc-netbox-buildout-scope.md +++ b/docs/dc-dc-netbox-buildout-scope.md @@ -175,7 +175,7 @@ (NetBox-upstream), D-052/D-053 (six-plane), D-016/D-074 (tenant CIDR). - **Tooling:** `netbox/dc-dc-prefixes-import.py` (+ `README.md`), `netbox/ipv6-mark-reserved.py`, `scripts/lib-net.sh` (`SPACES6`/`PLANE_CIDRS`), `docs/dc-dc-buildout-design.md` (Phase 1 gate), - `docs/maas-as-built-reference.md` (VLANs/fabrics), `docs/dc-dc-replication-DR-seed.md`. + `docs/maas-as-built-reference.md` (VLANs/fabrics), `docs/archive/dc-dc-replication-DR-seed.md`. - **Phase:** this is the **Phase 1** NetBox-population gate ("NetBox authoritative and populated: planes, per-DC v4, ULA/GUA") in `dc-dc-buildout-design.md`. @@ -195,7 +195,7 @@ **DOCFIX-195 (operator ruling 2026-07-15) -- READ THIS FIRST; it INVERTS the roles the rest of this section was originally written around.** During VR1 the VR1 IPAM APEX is `office1-netbox` -(`10.10.1.10:8000`, NetBox 4.6.4 on the Office1 VM, `docs/changelog-20260713-office1-netbox-deployed.md`): +(`10.10.1.10:8000`, NetBox 4.6.4 on the Office1 VM, `docs/archive/changelogs/changelog-20260713-office1-netbox-deployed.md`): it takes ALL VR1 writes (the section-4 buildout, the D-120 additions, any further refinement) and every consuming system (OpenTofu, MAAS, overlays, the Juju bundle) derives its NetBox values from it, not from `netbox.baldurkeep.com`. It is a *working draft* -- edited as the buildout proceeds -- but it is the @@ -232,7 +232,7 @@ a SUBSET of the plan (`unexpected = extra - expected`); it NEVER proved the plan was CARRIED OUT. With ZERO of the 36 DC prefixes created, `extra` and `unexpected` were both empty, and it printed "Nothing lost, nothing stray" and exited 0. `EXPECTED_NEW` was an upper bound masquerading as an -assertion. Fixed 2026-07-14 (`docs/changelog-20260714-netbox-write-path-hardening.md`): `EXPECTED_NEW` +assertion. Fixed 2026-07-14 (`docs/archive/changelogs/changelog-20260714-netbox-write-path-hardening.md`): `EXPECTED_NEW` is now BOTH bounds (`EXPECTED-BUT-ABSENT` is a hard failure), and delta prefixes are checked for scope and role. A NEW harness `tests/sandbox-fidelity-check/run-tests.sh` (10/10; T3 IS the false-green scenario) ships with it -- the checker had shipped with NO harness, which is why the false green @@ -263,9 +263,9 @@ Before 2026-07-14 the two pynetbox importers had NO target guard at all; the only thing stopping an accidental upstream `--commit` was the WAF 403 (8.5) -- and fixing the WAF (below) removed that accidental safety, which is exactly why the explicit `--yes-write-upstream` gate was added in the -same change (`docs/changelog-20260714-netbox-importers-waf-useragent.md`). The whole-plan preflights +same change (`docs/archive/changelogs/changelog-20260714-netbox-importers-waf-useragent.md`). The whole-plan preflights are the OTHER layer: both importers had previously died mid-loop and left the production IPAM -half-populated (`docs/changelog-20260714-netbox-write-path-hardening.md`). +half-populated (`docs/archive/changelogs/changelog-20260714-netbox-write-path-hardening.md`). ### 8.5 The upstream WAF trap (FIXED 2026-07-14) @@ -277,7 +277,7 @@ The three stdlib tools always sent an accepted UA (`curl/8.5.0`). The two pynetbox importers did NOT until 2026-07-14 -- meaning they could NEVER have written to the apex, only to the WAF-less sandbox. `get_nb()` in both now sets `nb.http_session.headers["User-Agent"] = "curl/8.5.0"` -(`docs/changelog-20260714-netbox-importers-waf-useragent.md`). NOTE: `pynetbox` is NOT installed on +(`docs/archive/changelogs/changelog-20260714-netbox-importers-waf-useragent.md`). NOTE: `pynetbox` is NOT installed on `vcloud` -- the two pynetbox importers must run on `office1-netbox` (pynetbox 7.0.0), which reaches upstream through the edge. diff --git a/docs/dc-dc-netem-and-ula-gua-proposal.md b/docs/dc-dc-netem-and-ula-gua-proposal.md deleted file mode 100644 index 74b2820..0000000 --- a/docs/dc-dc-netem-and-ula-gua-proposal.md +++ /dev/null @@ -1,139 +0,0 @@ -# Proposed netem parameters + ULA/GUA generation guidance (2026-07-10) - -Addresses tooling gap register items #4(d) and #11: the two D-100/D-101 -sub-items still genuinely "leaning," not ruled, even after Stage 0 closed -the big decisions. This document does NOT rule either -- it presents a -concrete recommendation for #4(d) (present options, get a ruling, per this -repo's own discipline for anything touching an ADOPTED-but-not-fully- -specified decision) and, for #11's literal-generation half, exact commands -the OPERATOR runs to produce a real value (this document does not generate -or invent the actual ULA/48 itself -- that is explicitly the operator's/ -NetBox's job per D-101's own text). - ---- - -## Part 1 -- Proposed `tc netem` parameters (D-100 sub-item #4(d)) - -**Current state:** `docs/dc-dc-buildout-design.md` Section 6 states only a -qualitative lean: "same-metro dark fiber (low single-digit ms, jumbo- -capable)." No specific latency/jitter/loss/rate numbers are ruled. -`opentofu/modules/netem-link` (built 2026-07-09) is ready to apply real -parameters the moment they're decided -- this is the one missing input. - -**Proposal (for operator ratification, not silently applied):** - -| Parameter | Proposed value | Rationale | -|---|---|---| -| Latency | 1ms (each direction) | Same-metro dark fiber circuits (single city, no long-haul) typically show sub-millisecond one-way propagation delay for distances under ~50km (fiber propagation is ~5us/km) -- 1ms per direction (2ms RTT) is a deliberately conservative round number that accounts for real-world switching/equipment overhead beyond pure propagation, while staying solidly within the buildout design's own "low single-digit ms" lean. | -| Jitter | 0.2ms | A small, non-zero jitter keeps the simulation honest (a perfectly flat-latency link is unrealistic even on dedicated fiber) without dominating the 1ms base latency -- roughly 20%, a common rule-of-thumb ratio for stable dedicated circuits (as opposed to shared/internet-transit links, which would warrant a much larger jitter fraction). | -| Loss | 0.01% | Dark fiber with modern optical equipment is very low-loss; this is a nonzero-but-negligible value so packet-loss-handling code paths are still technically exercised during the drill, without meaningfully affecting throughput-sensitive tests (Ceph replication, radosgw sync). | -| Rate | Not capped (no `rate` netem param) | The buildout design's own lean explicitly says "jumbo-capable" -- same-metro dark fiber circuits are typically provisioned well above what this test environment's actual traffic will generate (VR1 is a virtual rehearsal on one physical host, not a real multi-site deployment) -- an artificial rate cap would constrain the TEST more than a real Roosevelt link would constrain PRODUCTION, which is backwards for a rehearsal whose value is proving the failover mechanism works, not proving it works under bandwidth starvation. If the operator wants to also rehearse a bandwidth-constrained scenario, that is a SEPARATE, deliberate test configuration, not this default. | - -**Resulting `tc netem` invocation shape** (for reference; `modules/netem- -link`'s `netem_args` variable takes this as a string): -``` -delay 1ms 0.2ms loss 0.01% -``` - -**This is a recommendation, not a ruling.** Per this repo's standing -discipline (present options for anything not yet ADOPTED, never silently -decide), the operator should explicitly ratify this table (accept as-is, -or adjust) before `modules/netem-link` is actually instantiated with real -values in Stage 3's runbook. Once ratified, this should be recorded as a -D-100 amendment (a new "Sub-item RULED" line in `docs/design-decisions.md` -D-100's entry, following the exact pattern D-100's other redline items -already got at Stage 0 ratification) -- not left as a standalone doc. -**Re-tune when a specific Roosevelt inter-DC target is known** (buildout -design's own stated intent) -- this proposal is a VR1 rehearsal default, -not a permanent value. - ---- - -## Part 2 -- ULA/GUA/DC2-supernet generation guidance (D-101, gap #11 + gap #3's data half) - -**What this section does NOT do:** generate the org ULA /48, obtain the -real per-DC GUA carve, or assign DC2's supernet. D-101's own text is -explicit that these are NetBox-authoritative, populated via the extended -import pipeline -- not hardcoded in any decision, and not something this -session (or any Claude session) should invent, guess, or pre-select on the -operator's behalf. What follows is the exact, safe MECHANISM for the -operator to produce each real value themselves, so that step is a two- -minute copy-paste rather than a research task when they're ready to do it. - -### 2.1 -- Org ULA /48 (RFC 4193) - -RFC 4193 requires the 40-bit Global ID be "randomly generated" for -statistical uniqueness across organizations -- a cryptographically random -40 bits satisfies this directly (this is also what every independent -RFC-4193-compliant ULA generator tool checked during tonight's research -actually does under the hood, not a shortcut invented here): - -```bash -printf 'fd%s::/48\n' "$(openssl rand -hex 5 | sed 's/\(..\)\(....\)\(....\)/\1:\2:\3/')" -``` - -Run this ONCE, by the operator, on a real machine -- the output is the -real `ORG_ULA_48` value `netbox/dc-dc-prefixes-import.py` (DOCFIX-152) -requires. Record it in NetBox as the IPAM apex (D-101's own requirement), -not just as an environment variable -- this is a permanent organizational -identifier, not a per-session value. - -> **Correction (DOCFIX-183):** the `sed` above was fixed to split the 10 hex -> digits as 2+4+4 hextets (`fdXX:XXXX:XXXX::/48`); the prior 4+6 split produced -> INVALID IPv6 (a 6-hex-digit hextet -- caught when this command was actually -> run 2026-07-11). **The ratified VR1 value is `fd50:840e:74e2::/48`** (CSPRNG, -> collision-checked vs the Tailscale `/48` in NetBox -- see -> `docs/dc-dc-netbox-buildout-scope.md` section 4c); this command is the -> reference method, not a re-generation step for VR1. - -### 2.2 -- Per-DC GUA carve (from ARIN 2602:f3e2::/32 region-0 /36) - -This is NOT something to generate randomly -- D-101's own text names this -as ALREADY a real, existing ARIN allocation ("GUA (ARIN 2602:f3e2::/32, -region-0 /36)"), meaning the org already holds this block from a real -Internet numbering authority. The per-DC carve within it (which /40 or /44 -goes to DC1 vs. DC2) is an INTERNAL allocation decision within an already- -owned block, not a new external assignment -- this can be decided by -whoever administers that ARIN allocation for this organization, following -whatever internal IPAM convention they already use for carving GUA space -(e.g. sequential /40s: DC1 = `2602:f3e2:1000::/40`, DC2 = -`2602:f3e2:2000::/40`, matching the example addresses already used in -`netbox/dc-dc-prefixes-import.py`'s own docstring and README -- those are -ILLUSTRATIVE examples in this repo, not yet a real assignment; confirm -with whoever holds the ARIN allocation before treating any specific /40 as -final). - -### 2.3 -- DC2's v4 supernet - -This is a NetBox IPAM assignment task, not a generation task -- DC2 needs -a real, non-overlapping v4 supernet (at least a /19, per `netbox/dc-dc- -prefixes-import.py`'s own `DC2_MIN_SUPERNET_PREFIXLEN` check) chosen from -whatever address space this organization has available and not already in -use by DC1, Office1, client VPN ranges (D-074), or any other real -allocation. This is a real IPAM decision requiring visibility into the -FULL address inventory, which this session does not have -- assign it via -NetBox directly, following whatever process this organization already -uses for carving new supernets. - -### 2.4 -- Once all three exist - -Run `netbox/dc-dc-prefixes-import.py --dc dc1` (needs only `ORG_ULA_48` -and `DC_GUA_PREFIX` from 2.1/2.2) and `--dc dc2` (additionally needs -`DC2_V4_SUPERNET` from 2.3) for real, closing the DATA half of tooling gap -#3 -- the mechanism has been ready since DOCFIX-152; only these three real -inputs were ever missing. - ---- - -## Summary: what this closes vs. leaves open - -**CLOSES:** gap #4(d) now has a concrete, reasoned proposal ready for a -one-line operator ratification (not a re-research task); gap #11's ULA- -generation half now has an exact, safe, two-minute command instead of -"go figure out RFC 4193." - -**STAYS OPEN, by design:** the actual netem ratification (a decision only -the operator can make); the actual ULA-48/GUA-carve/DC2-supernet values -(real-world IPAM/ARIN coordination this session has no authority or -visibility to perform) -- this document makes producing them fast and -low-risk once the operator is ready, it does not produce them. diff --git a/docs/dc-dc-replication-DR-seed.md b/docs/dc-dc-replication-DR-seed.md deleted file mode 100644 index 948ec6e..0000000 --- a/docs/dc-dc-replication-DR-seed.md +++ /dev/null @@ -1,75 +0,0 @@ -# SEED: DC-DC cross-DC replication / DR design (Tier A) - -**Status:** SEED / DRAFT for the DC-DC planning chat. Authored by the main-chat stream -during DC-DC planning. This is a settled *decision* to formalize, not a committed record. -The planning chat files the D-entry: **assign next-free via `ledger-scan` at commit -- -note D-076 is already claimed by the PINNED dashboard workstream (draft in a worktree), so -the DR decision is the next free after that; confirm and coordinate before consuming.** - -**Governing constraint:** minimize delta to Roosevelt. - ---- - -## Decision (settled with the operator) - -Two **independent** clouds (separate Keystones), so RBD-mirror is a **data** primitive, not -service failover. DR posture = **Tier A: data recoverability**, **bidirectional** (each DC -protects the other), operator-runbook failover. - -Mechanism, per data class: -- **Cinder volumes -> `cinder-backup`** to a **cross-DC-replicated object store**. This is - the honest cross-independent-cloud path: a restore reconstitutes the volume *with its - Cinder metadata* in the peer cloud, which raw RBD-mirror cannot. Low delta -- `cinder` - already exposes the `backup-backend` endpoint (bound metal-internal) and `ceph-radosgw` - is present as the S3 target. -- **Glance images -> snapshot-based RBD-mirror** of the Glance pool, two-way, with a - re-register step in the failover runbook. -- **Nova ephemeral -> NOT replicated.** Ephemeral root disks are cattle; recovering them - would also need the un-mirrored Nova instance records. Instances are rebuilt from - image + volume on failover. - -Parameters: -- **RPO** = replication/backup interval; start ~15 min, tunable. -- **RTO** = a **measured output** of the failover drill, not a pre-committed SLA. -- **Staging:** prove **one-way** (single mechanism, single direction) + a clean - failover/failback drill FIRST, then enable **two-way**. -- **Carrier:** the **replication plane (`10.12.36.0/22`)** with an explicit **inter-DC - route + bandwidth budget** -- distinct from intra-DC OSD cluster traffic that already - uses this plane. -- **Not replicated:** Neutron state (ports, FIPs, security groups) -- recovered workloads - land on the peer's networks with new addresses; the runbook re-creates this. - -## Options considered - -- **Tier A (CHOSEN):** mirror/backup + operator rehydration runbook. RPO minutes, RTO - hours/measured. Lowest complexity, most robust, delta-minimal. -- **Tier B (folded in for volumes):** `cinder-backup`/restore -- adopted as the Cinder path - because it crosses the independent-cloud boundary cleanly. -- **Tier C (deferred):** orchestrated workload failover via control-plane metadata replay. - Low RTO, high build cost, brittle (no first-class cross-independent-cloud OpenStack DR). - Revisit only if a hard minutes-RTO requirement emerges. - -## Failover / failback runbook -- SKELETON (planning chat to flesh out against the live env) - -Failover (DC1 lost, recover at DC2): -1. Confirm DC1 truly down (avoid split-brain from a transient partition). -2. Glance: promote the mirrored Glance pool at DC2 (force-promote if DC1 unreachable); - re-register images into DC2 Glance. -3. Cinder: restore required volumes into DC2 Cinder from the replicated backups. -4. Rebuild instances from image + restored volume; re-create Neutron ports/FIPs/SGs. -5. Record RTO/RPO actuals. - -Failback (DC1 recovered) -- split-brain-safe ordering for two-way: -1. Do NOT let recovered DC1 resume as primary automatically. -2. Demote DC1's Glance pool; resync from DC2 (now primary). -3. Reconcile Cinder: back up any DC2-side changes; restore into DC1. -4. In a controlled window, optionally flip primary back (demote DC2, promote DC1). -5. Verify both directions healthy; record the drill. - -## Open sub-items for the planning chat -- radosgw **multisite** vs. a mirrored backup bucket as the replicated backup target. -- Exact mirrored pools; **consistency groups** for multi-volume apps. -- Inter-DC replication route + bandwidth budget on the simulated WAN (MTU-aware). -- Interaction with the **D-071** controller-HA decision (backup/restore of per-DC control - planes). -- `ceph-rbd-mirror` charm topology + daemon HA; peer bootstrap on the replication plane. diff --git a/docs/dc0-deploy-readiness.md b/docs/dc0-deploy-readiness.md index f922b04..95487e7 100644 --- a/docs/dc0-deploy-readiness.md +++ b/docs/dc0-deploy-readiness.md @@ -96,7 +96,7 @@ ### C5. Record-hygiene sweep -- RESOLVED (this queue was stale; R3-F disposition 2026-07-19) 7. The queued amendments LANDED 2026-07-16 (review sweep Phase A/B, - `docs/changelog-20260716-review-sweep-phaseAB.md`): D-122/D-107/D-121/D-115/SEC-011 all + `docs/archive/changelogs/changelog-20260716-review-sweep-phaseAB.md`): D-122/D-107/D-121/D-115/SEC-011 all amended; the absent `optc-calc.py` replaced by committed `scripts/dc-dc-whole-host-budget.py` (13/13). Residual R3-F items dispositioned in sweep Batch 2 item 2.10 (session changelog). diff --git a/docs/design-decisions.md b/docs/design-decisions.md index 7a62e65..8bd911e 100644 --- a/docs/design-decisions.md +++ b/docs/design-decisions.md @@ -829,7 +829,7 @@ ## D-057: provider-vip plane -- separate tagged, routed plane for public API VIPs (2026-06-29) -**Status:** SUPERSEDED by D-060 (Pattern A revert, 2026-06-29) -- the provider-vip plane is abandoned; public API VIPs return to provider-public. [Originally DECIDED.] Full record + D-003B amendment: `docs/D-057-DECIDED-append.md`. +**Status:** SUPERSEDED by D-060 (Pattern A revert, 2026-06-29) -- the provider-vip plane is abandoned; public API VIPs return to provider-public. [Originally DECIDED.] Full record + D-003B amendment: `docs/archive/D-057-DECIDED-append.md`. Root cause of the phase-06 FIP-unreachability blocker: API LXD containers bind `public` to provider-public (untagged enp1s0); Juju bridges enp1s0 into a Linux bridge, starving @@ -846,7 +846,7 @@ **Status:** SUPERSEDED by D-060 (Pattern A revert, 2026-06-29) -- the plane renumber is abandoned; the cloud stays on the D-052/D-053 scheme. [Originally DECIDED (operator).] Full map, jumphost ordering trap, NetBox-apex note, and the committed-foundation cascade: -`docs/D-058-renumber.md`. Supersedes the D-057 minimal-delta placement of provider-vip at +`docs/archive/D-058-renumber.md`. Supersedes the D-057 minimal-delta placement of provider-vip at 10.12.24.0/22; resolves R4 (oob). Cloud-wide re-IP for Roosevelt addressing fidelity (contiguous /22 blocks grouped by @@ -1720,7 +1720,7 @@ Revert: `juju bind openstack-dashboard cluster=metal-internal` + git revert of the bundle/docs commit. -**Upstream:** bug draft at docs/upstream-bug-draft-dashboard-tls.md (charm renders an +**Upstream:** bug draft at docs/archive/upstream-bug-draft-dashboard-tls.md (charm renders an L4-masked dead TLS backend by default in multi-space deployments) -- operator to file. **Roosevelt:** edge TLS/DNS (D-044 trajectory) supersedes this plumbing; carry the invariant "cluster binding space == default binding space for TLS-fronted charms" into @@ -1791,7 +1791,7 @@ ## D-072 -- AMENDMENT (2026-07-06): upstream bug NOT filed (operator ruling) -Operator ruled 2026-07-06: the drafted upstream bug (docs/upstream-bug-draft-dashboard-tls.md) +Operator ruled 2026-07-06: the drafted upstream bug (docs/archive/upstream-bug-draft-dashboard-tls.md) will NOT be submitted -- the root condition is judged a site misconfiguration (cluster binding pointed at a third space) rather than a charm defect to report. The draft stays in-repo for history; the addendum-18 OPERATOR ACTION item is withdrawn. The Roosevelt @@ -1824,7 +1824,7 @@ ## D-074: Tenant CIDRs are tenant-chosen and may overlap (amends D-016) **Status: ADOPTED 2026-07-06** (operator rulings recorded same day; proposed by -the main-chat stream in docs/tenant-cidr-overlap-correction-PLAN.md). +the main-chat stream in docs/archive/tenant-cidr-overlap-correction-PLAN.md). **Decision.** Tenant private (overlay) CIDRs are tenant-chosen and MAY overlap across tenants. Non-collision is enforced ONLY against the reserved-ranges @@ -2130,7 +2130,7 @@ ## D-108: Cross-DC replication mechanism (VR1) -**Status:** ADOPTED 2026-07-09 (operator ruling). Engineers the mechanism for the SETTLED Tier-A DR decision (`docs/dc-dc-replication-DR-seed.md`); does not reopen the posture. +**Status:** ADOPTED 2026-07-09 (operator ruling). Engineers the mechanism for the SETTLED Tier-A DR decision (`docs/archive/dc-dc-replication-DR-seed.md`); does not reopen the posture. **Decision (mechanism):** - Cinder volumes: `cinder-backup` to a cross-DC-replicated object store via `ceph-radosgw` MULTISITE (a zonegroup with the two DCs as zones; bucket-level bidirectional sync) as the replicated backup target. A restore reconstitutes the volume with its Cinder metadata in the peer -- the honest cross-independent-cloud path. Multisite is chosen over a single mirrored bucket for bidirectional, per-object async replication with conflict handling. @@ -2143,7 +2143,7 @@ **Rationale:** matches the settled seed; multisite plus rbd-mirror are the delta-minimal Ceph-native primitives. This answers the seed's open questions (multisite vs mirrored bucket -> multisite; daemon HA -> single-unit VR1; peer bootstrap -> on the replication plane). -**References:** engineers `docs/dc-dc-replication-DR-seed.md`; carrier per D-101; interacts with D-104 (per-DC control-plane backup / restore in the drill); the failover / failback skeleton is fleshed out in the buildout doc. +**References:** engineers `docs/archive/dc-dc-replication-DR-seed.md`; carrier per D-101; interacts with D-104 (per-DC control-plane backup / restore in the drill); the failover / failback skeleton is fleshed out in the buildout doc. --- @@ -2338,8 +2338,8 @@ **Impacts if not (b):** `scripts/opnsense-build-config-iso.sh`, its harness, the module's `config_seed` volume + cdrom disk, and the README research section all assume the ISO mechanism. -**References:** `docs/incident-20260712-opnsense-edge-boot-triplefault.md`, -`docs/changelog-20260712-libvirt-acpi-kernel-panic.md` (DOCFIX-190), upstream +**References:** `docs/archive/incident-20260712-opnsense-edge-boot-triplefault.md`, +`docs/archive/changelogs/changelog-20260712-libvirt-acpi-kernel-panic.md` (DOCFIX-190), upstream `opnsense/core:src/sbin/opnsense-importer` + `src/etc/rc.syshook.d/import/20-importer`. Also gates Stage 3 (the per-DC OPNsense edges use this same module). @@ -2491,9 +2491,9 @@ Low but not zero: Stage 3 builds two more edges from this same module. Ruling BEFORE Stage 3 means the choice is made once; ruling after means migrating three edges instead of one. -**Evidence:** `docs/changelog-20260712-office1-opnsense-edge-build.md`, -`docs/changelog-20260712-opnsense-edge-boot-fixes.md`, -`docs/changelog-20260713-office1-dhcp-apply.md` (the 667-element self-heal, measured), +**Evidence:** `docs/archive/changelogs/changelog-20260712-office1-opnsense-edge-build.md`, +`docs/archive/changelogs/changelog-20260712-opnsense-edge-boot-fixes.md`, +`docs/archive/changelogs/changelog-20260713-office1-dhcp-apply.md` (the 667-element self-heal, measured), D-112 (the (c) ruling and its stated rationale). ### D-113 -- AMENDMENT (2026-07-14): "interfaces: the REST API" IS FALSE. Base-interface addressing has NO API. @@ -3081,7 +3081,7 @@ cost of re-editing the 7 places that already carry this one and risking a stale literal surviving in one of them. Not worth it. -**Corrects a FALSE claim in the repo.** `docs/changelog-20260711-ula-gen-command-fix.md` called this +**Corrects a FALSE claim in the repo.** `docs/archive/changelogs/changelog-20260711-ula-gen-command-fix.md` called this "the **ratified** VR1 value" -- it was NOT ratified; no D-number ever assigned it. That phantom ratification is very likely why the value has been treated as settled for two days while `grep ORG_ULA docs/design-decisions.md` returned nothing. This entry makes the claim true, and the @@ -3643,7 +3643,7 @@ (3 control + 2 compute + 4 storage per the D-121 R-3 amendment = ~384 GiB node RAM + overhead, ~416 GiB VM) and set `expose_nested_virt = true`. D-124's transit-addressing (Scheme A) still stands. 3. **Depth-4 nested virt ACCEPTED** as a rehearsal cost (VR0 proved depth 2). Functional risk to the - Phase-5 workload boot flagged and accepted; `docs/model-a-fallback-plan.md` + git tag + Phase-5 workload boot flagged and accepted; `docs/archive/model-a-fallback-plan.md` + git tag `model-a-fallback` preserve Model A as a low-cost revert if Model B fails to deploy. **Substrate reshape (Phase C of the review sweep, not yet done):** retarget `modules/node-vm`, the 6 @@ -3727,7 +3727,7 @@ `.2-.49` band). These fill the four `vr1_dc0_rack_*` tfvars: `rack_transit_ip=172.31.0.2`, `rack_transit_prefix=30`, `rack_transit_peer_ip=172.31.0.1`, `rack_metal_admin_ip=10.12.8.2`. The NetBox `--commit` runs ON office1-netbox (apex token local there; unreachable from the vcloud jumphost). See -`docs/changelog-20260716-d124-addressing-pin.md`. +`docs/archive/changelogs/changelog-20260716-d124-addressing-pin.md`. ## D-125: VR1 Model B per-DC ISP egress -- bridge-in single-NAT (resolves OBS-3's design gap; egress efficacy is a deploy-time gate) [ARCH] @@ -3804,7 +3804,7 @@ **Related:** D-122 (dedicated per-site L3 ISP -- this is its Model B realization), D-123 (Model B nesting, which created the defect), D-115 (`vr1-dc0-wan` 172.30.2.0/24 -- the old inner subnet, now the bridge), D-113 (edge config over REST -- OPNsense WAN static addr), SEC-010 (transit FORWARD-drop, scoped), -DOCFIX-185 (edge is a real-ISP router, not an egress airgap). **Fallback:** `docs/model-a-fallback-plan.md` +DOCFIX-185 (edge is a real-ISP router, not an egress airgap). **Fallback:** `docs/archive/model-a-fallback-plan.md` section 3 (revert removes the uplink NIC/network/bridge + `wan-bridge` and restores the OPNsense WAN addr). ## D-126: durable, rootless vcloud->site-service-VM access -- SSH local-forward via systemd --user (RULED: Option A) [OPS] diff --git a/docs/docfix-draft-20260702.md b/docs/docfix-draft-20260702.md deleted file mode 100644 index c5c65ff..0000000 --- a/docs/docfix-draft-20260702.md +++ /dev/null @@ -1,215 +0,0 @@ -# DOCFIX draft -- redeploy-readiness review (bundle + channels + runbooks) - -STATUS: DRAFT / OPEN -- accreting during the 2026-07-02 review session. -Numbers are PROVISIONAL from next-free DOCFIX-066 (verified against HEAD -690779a: DOCFIX-065 / D-068 / BUNDLEFIX-008 consumed). Renumber-check again -at commit time. ASCII + LF. - -Scope of this review: - 1. bundle.yaml -- YAML validity, structural consistency, known anti-patterns. - 2. Charm channel pins -- staleness review against current Charmhub guidance. - 3. Runbook sweep -- cross-reference integrity, stale values, anything that - would break the next redeploy. - -Severity key: BLOCKER (breaks redeploy) / RISK (may break or mislead) / -NIT (consistency only). - --------------------------------------------------------------------------------- -## Findings --------------------------------------------------------------------------------- - -(appended as found) -### DOCFIX-066 (BLOCKER) -- teardown runbook drives the DEPRECATED teardown script -File: runbooks/phase-00-teardown-maas-reset.md (steps 2, plan table, lines 22/30/67/74/85). -The runbook's execution spine is `scripts/phase-00-teardown.sh --apply` with the narrative -"hosts release to MAAS Ready" -- the exact premise DOCFIX-057/D-061 proved WRONG on this -virsh-pod MAAS (destroy-model DECOMPOSES pod-composed machines; observed 3x). The script -itself carries a DO-NOT-USE banner, so the runbook and script now contradict each other; -an operator following the runbook on the next redeploy either hits the deprecation mid- -teardown or, if they push past it, triggers a fourth decompose + full reenroll/recarve. -The D-061 replacements (phase-00-teardown-release.sh --keep-instance + canary; -phase-00-teardown-destroy.sh) exist but are never mentioned in the runbook. -FIX: rewrite the runbook spine as the D-061 fork -- (a) machine-preserving path: -teardown-release.sh with the MANDATORY first-run canary (--apply --canary, verify -openstack0 survives in MAAS, then all-four); (b) from-scratch path: teardown-destroy.sh -+ reenroll + recarve. State which path the standard redeploy uses. Step-5 "hosts Ready" -premise and the OSD-wipe/8_lbaas step ordering must be revalidated per path (release path -leaves hosts Deployed, not Ready -- the wipe/carve preconditions differ). - -### DOCFIX-067 (RISK -- verify live before ruling) -- octavia PKI cert SAN carries the pre-R14 VIP -File: runbooks/phase-01-bundle-deploy.md 1.0-GEN.c (lines ~292, ~318). -The controller-cert CNF sets `IP.1 = 10.12.4.233` -- the OLD octavia VIP. R14 relocated -all VIPs to .50-.60; bundle HEAD has octavia at 10.12.4.57. The octavia-pki overlay is -regenerated every phase-01, so the LIVE cloud's controller cert most likely carries .233 -today; octavia passed phase-05 validation regardless, so the amphora side evidently does -not verify that SAN IP -- functional impact UNPROVEN, inconsistency CERTAIN, and it is a -latent break if SAN verification ever tightens (or when Roosevelt re-uses this block). -The DNS.1/DNS.2 SANs also reference the D-019-dropped FQDN scheme (harmless, same sweep). -FIX: derive the SAN IP dynamically (lib-net VIP_PREFIX_PROVIDER + octavia's octet from -bundle/juju -- rule 3), not a literal. VERIFY-LIVE first (gated CHECK, jumphost): -read the overlay/live cert SAN and confirm what is actually deployed: - openssl x509 -in -noout -text | grep -A2 'Subject Alternative Name' -Rule on severity after the read: if the live cert has .233 and octavia is green, keep RISK -(doc fix + regenerate at next redeploy); do not hot-rotate certs mid-cloud for this. - -### DOCFIX-068 (RISK) -- phase-01 "Constants and env-literals" block is pre-D-052 stale -File: runbooks/phase-01-bundle-deploy.md lines ~23-26. The block states the RETIRED plane -map (2=metal .8, 6=data .12, 7=storage .16, 8=replication .20, 9=lbaas .32 -- wrong -plane->CIDR pairs under D-052/D-053, incl. the retired `lbaas` space), plus hardcoded MAAS -subnet IDs (violates lib-net PATTERN-1: IDs drift, resolve by CIDR) and hardcoded -system_ids (violates DOCFIX-040: re-minted per enrollment; lib-hosts resolves). Mixed -freshness: the "50 apps, 97 relations" expectation in the same block MATCHES bundle HEAD. -Misleading at the worst moment (mid-deploy reference values). -FIX: replace the stale lines with pointers to scripts/lib-net.sh (planes) and -scripts/lib-hosts.sh (host identity); retain only verified-current literals. - -### DOCFIX-069 (RISK) -- zero exec bits + bare script invocations -git index: ALL 37 files under scripts/ are mode 100644 (GitHub Desktop workflow strips -+x). Fresh clone on the jumphost -> every BARE invocation fails "Permission denied". -runbooks/phase-00-teardown-maas-reset.md invokes bare in ~10 places (teardown, carve, -standup); other runbooks appear bash-prefixed (sweep found no other bare hits). -FIX (durable, matches the Windows commit constraint): bash-prefix every script invocation -in runbooks (`bash scripts/x.sh ...`). Optional belt: `git update-index --chmod=+x -scripts/*.sh` -- but Windows-side recommits can strip again, so the bash prefix is the -invariant; do both if desired. - -### DOCFIX-070 (RISK) -- scripts/review-bundle.py is pre-D-052 stale; NOT CLEAN is noise -Against bundle HEAD it reports FAIL=71: expects space `metal` (retired), DUAL VIPs (D-020 -form; D-052 moved to triples), no per-endpoint bindings (D-052 introduced them), vault -1.8 (D-068 pinned 1.16), baselines 51 apps/98 rels (now 50/97: VIP set changed -- vault -dropped its VIP, ceph-radosgw gained one -- blessed by provider-bundle-check.py, which -PASSES clean). Hazard: a pre-deploy NOT CLEAN verdict that must be ignored trains alarm -fatigue and will eventually mask a real defect. -OPTIONS: (a) update review-bundle.py expectations to the D-052/D-060/D-062/D-068 model; -(b) retire it (git rm) and fold any still-unique checks (relation-endpoint syntax, -phantom-key detection reworked for the per-endpoint model) into provider-bundle-check.py; -(c) banner it historical. RECOMMEND (b): one authoritative gate beats two disagreeing -ones -- same reasoning as the D-060 d057-bundle-check retirement. - -### NIT-A -- D-002 channel matrix drift (design-decisions) -The D-002 table still lists `etcd, easyrsa -> latest/stable` (etcd/easyrsa dropped; R3 / -phase-02 record vault-on-mysql), omits memcached (bundle: latest/stable -- apparently the -only track that charm publishes; upstream's "never latest/stable" applies to OpenStack- -project charms, which memcached is not), and its vault row (1.8) is superseded by D-068 -(1.16). Append-only fix: a dated amendment note under D-002, not an edit. -VERIFY-LIVE (gated CHECK, jumphost) before finalizing: - for c in memcached rabbitmq-server vault hacluster; do juju info "$c" 2>/dev/null | sed -n '/channels:/,$p' | head -12; done -Expected: memcached publishes only latest/*; rabbitmq-server tops out at 3.9; vault -carries 1.16/stable; hacluster 2.4/stable. - -### NIT-B -- channel-pin review conclusion (informational; no change) -All bundle pins judged CURRENT for Caracal/jammy: 2024.1/stable core (18), OVN -24.03/stable, ceph squid/stable, mysql 8.0/stable (12), hacluster 2.4/stable (11), -rabbitmq-server 3.9/stable (terminal track for this charm), vault 1.16/stable (D-068), -memcached latest/stable (sole track; see NIT-A). Upstream charm-guide delivery page is -frozen (last updated 2023-12) -- Charmhub/juju info is the only live authority; the -NIT-A verify block doubles as the pre-deploy channel assert. Candidate: fold that assert -into scripts/pre-flight-checks.sh (D-002 claims pre-flight verifies channels -- confirm -it actually does; not yet audited). - -### NIT-C -- ASCII-rule violations in docs/ -docs/v1-pre-deploy-fixes.md (277 non-ASCII bytes), docs/netbox-vip-queue.md (81). The -repo rule is ASCII-only for all committed files (mod_wsgi lesson). Low functional risk -(docs, not conf), but the rule is stated absolute -- sanitize or record a carve-out. - -### NIT-D -- identifier index gaps -DOCFIX-027/028/029/034/037 and BUNDLEFIX-001..006 are defined only at point of use -(runbook/bundle comments) and absent from appendix-A / the changelog index -- appendix-A -claims to be the index "keyed by the same identifiers used inline". Add one-line index -entries (or mark point-of-use-only identifiers as such). - -### NIT-E -- appendix-A lacks a mysql-innodb-cluster recovery entry -D-062 material (blocked 'Instance not yet configured' = single-unit seed; half-join -instanceErrors = mid-life rescan; reboot-cluster-from-complete-outage ONLY on confirmed -outage -- destructive against a healthy cluster) exists in design-decisions + the restart -procedure but has no appendix-A symptom entry. Add one; also consider committing the -restart-procedure doc to the repo (it currently lives outside it). - --------------------------------------------------------------------------------- -## Verify-live queue (gated CHECKs for the jumphost before findings finalize) -1. Octavia controller cert SAN (DOCFIX-067) -- read the deployed overlay/cert. -2. juju info channel probe (NIT-A/B) -- memcached / rabbitmq-server / vault / hacluster. -3. pre-flight-checks.sh -- confirm whether it performs the D-002 channel assert. - --------------------------------------------------------------------------------- -## Deployment-flow parity findings (decision vs bundle vs schedule) --------------------------------------------------------------------------------- - -### DOCFIX-071 (BLOCKER) -- D-064 keystone policy attach is not reachable from the deploy schedule -Evidence: bundle.yaml keystone has use-policyd-override=True but NO resources: stanza; -`attach-resource keystone` appears in NO phase runbook or script -- only appendix-C:73-74. -phase-01:183 knowingly deploys into "PO (broken)" and phase-02:167 re-notes it as -FINDING-1 "not a regression"; no phase ever resolves it. The live cloud got the policy -via a session action (D-064), never folded into the schedule. NEXT REDEPLOY as written: -keystone stays PO (broken), the SCS Domain Manager RBAC (the commercial tenant-isolation -core, D-051) is ABSENT, and tenant onboarding fails at G3. -Compounding defect: the appendix-C block zips to and attaches FROM /tmp -- the -documented snap-confinement trap (attach-resource cannot read /tmp on this jumphost; -use $HOME). The only written procedure is the known-broken form. -FIX OPTIONS (debate): - (a) Bundle-native resource: add to keystone `resources: {policyd-override: - ./policies/overrides.zip}` and commit the zip beside its source yaml. Deploy-time - attach becomes automatic -- the bundle describes the WHOLE desired state, zero - manual step, zero Roosevelt delta. Sync risk (zip vs yaml drift) is closed by a - pre-flight assert: rebuild the zip from policies/, byte-compare against committed, - HOLD on mismatch. - (b) Schedule step: a gated attach block in phase-03 (post-TLS-settle), $HOME-pathed, - gating on `PO:` in juju status. - EITHER WAY: the G3 BEHAVIORAL gate (manager can self-service own domain; admin-grant - and cross-domain DENIED; cloud-admin unaffected) must be a phase step -- D-051's own - warning: the charm validates YAML only, `PO:` proves parse, not policy. RECOMMEND - (a) + G3 gate in phase-03: post-deploy manual steps are exactly the class D-046 - proved unreliable ("reports ready regardless"). - -### DOCFIX-072 (RISK) -- bundle implements the still-PROPOSED D-043 -bundle.yaml nova-compute sets resume-guests-state-on-host-boot: True while D-043 -(tenant-VM auto-resume) remains PROPOSED / decision-pending. The bundle is ahead of the -decision record -- the exact drift the discipline forbids (and the restart-procedure doc -already assumes the option is in force). FIX: rule on D-043 -- RECOMMEND adopting its -option (a) (auto-resume + monitoring; industry norm for tenant VMs; customers at -Roosevelt expect VMs back after host maintenance; D-041's down-is-a-signal stance is -preserved for CONTROL-PLANE services, which auto-resume does not touch) -- and mark the -decision ADOPTED with the bundle line as its implementation. Alternative: strip the -option until ruled; NOT recommended (regresses the validated restart procedure). - -### NIT-F -- D-011.6 text not amended to the phase-08 ruling -design-decisions D-011 item 6 still reads "Vault unseal + auto-unseal-after-reboot -pattern verified"; phase-08 D-011.6 rules MANUAL unseal is the v1 standard (auto-unseal -NOT configured). Append an amendment note to D-011 so the acceptance bar and the -acceptance runbook agree. - -### NIT-G (Roosevelt-forward) -- rabbitmq-server scale-up will race without min-cluster-size -Testcloud: num_units=1, no min-cluster-size -- correct per D-009 (decorative HA). But the -D-009 promise is "Roosevelt scale-up is mechanical: 1 -> 3 and rerun". For rabbitmq that -is NOT sufficient: without `min-cluster-size: 3` the charm accepts client relations -before the cluster forms (same failure CLASS as D-062's mysql formation race; upstream -charm docs call min-cluster-size best practice). Record now as a Roosevelt bundle-delta -note on D-009 so the mechanical scale-up story stays true. - --------------------------------------------------------------------------------- -## Patchset status (2026-07-02, patchset-20260702-redeploy-readiness) --------------------------------------------------------------------------------- -IMPLEMENTED in the delivered ZIP (numbers verified next-free at HEAD 690779a; -re-grep at commit): DOCFIX-066 (teardown runbook rewritten around the D-061 fork, -destroy path = validated spine, reenroll step added, all invocations bash-prefixed), -DOCFIX-067 (octavia SAN IP derived from bundle at generation time; verify-live of the -deployed cert still queued), DOCFIX-068 (phase-01 constants -> lib-net/lib-hosts), -DOCFIX-069 (bash-prefix; optional chmod noted in apply-notes), DOCFIX-070 (checks -absorbed into provider-bundle-check.py + 8-case harness; review-bundle.py to git rm), -DOCFIX-071 (bundle-native keystone policy resource + committed zip + drift guard + -phase-03 Step 3.4 two-stage gate + appendix-C /tmp fix + subshell wrap), DOCFIX-072 -(D-043 RESOLVED->ADOPTED(a)). D-doc amendments appended: D-002, D-009, D-011, D-043, -D-051, D-061. STILL OPEN: NIT-C (docs/ ASCII sanitize), NIT-D (identifier index), -NIT-E (appendix-A mysql entry), verify-live queue items 1-3. - - -## Patchset status addendum (2026-07-03, Block 2) -IMPLEMENTED: DOCFIX-073 (preflight + channel assert + phase-01 gate), DOCFIX-074 -(repo-lint + full ASCII sanitize incl. .gitignore/netbox; closed NIT-C), DOCFIX-075 -(cloud-assert + committed ops-restart-procedure; closed the health-check gap), -DOCFIX-076 (as-executed convention + run-logged + index), DOCFIX-077 (appendix-A -mysql entry + identifier index; closed NIT-D/E), DOCFIX-078 (security ledger), -D-069 (vault custody policy), D-070 (supersedes D-012). Verify-live queue item 2 -(channel probe) is now AUTOMATED by preflight P3. Remaining operator inputs: -SEC-003 custodian assignment; capi-mgmt auto-resume exclusion ruling; octavia -deployed-cert SAN read. See docs/changelog-20260703-process-hardening.md. diff --git a/docs/handoff-20260703-open-items.md b/docs/handoff-20260703-open-items.md deleted file mode 100644 index 9b95428..0000000 --- a/docs/handoff-20260703-open-items.md +++ /dev/null @@ -1,164 +0,0 @@ -# Handoff -- open items register (2026-07-03 session close) - -STATUS SNAPSHOT at handoff: repo HEAD abc144a; gauntlet ALL GREEN (26 -harnesses); repo-lint 0 fail (1 documented legacy WARN); Claude Code live on -the jumphost with CLAUDE.md + permission rules + PreToolUse guard (13/13); -skill single-sourced at .claude/skills/openstack-cloud-ops (v1.3). Two -blockers from the redeploy-readiness sweep (DOCFIX-066 teardown spine, -DOCFIX-071 policy delivery) are FIXED and merged. Full change record: -docs/changelog-20260703-process-hardening.md (items 1-32, per-item reverts). - -Numbering at handoff (re-grep before assigning -- L5 prints it): -next-free D-071, DOCFIX-082, BUNDLEFIX-009. - -Conventions for whoever picks this up: repo wins over any session memory; -grep design-decisions before changing a built surface; every script change -ships with its harness green + a changelog entry with a revert; mutations are -individually human-gated (the permission ask-rules enforce this on the -jumphost -- do not work around them). - --------------------------------------------------------------------------------- -## 1. Immediate (small, well-defined; good first Claude Code tasks) --------------------------------------------------------------------------------- - -H-1 (DOCFIX-082 candidate) -- tenant-offboard.sh hardening - The new offboard script is well-built (audit-default, typed gate, protected- - domain blocklist, dependency-ordered) but runs `set -u` only. It MUTATES in - --apply: a silently failed pipeline mid-sweep strands resources while - reporting progress. FIX: `set -uo pipefail` + capture-then-test conversions - per references/script-authoring.md (both SIGPIPE directions); state the - chosen error regime in the header; strengthen is_id to the 32-hex form. - ACCEPT: tests/tenant-offboard green incl. a new failure-injection case; - changelog entry with revert. - -H-2 (DOCFIX-082 or 083) -- vault-kv-inner-probe.sh: secret off argv - The AppRole probe builds the login body with role_id/secret_id and (verify) - passes it via curl argv -- visible in `ps` on the unit. TTL is 60s so - exposure is bounded, but the house rule is absolute: secrets never transit - argv. FIX: feed the body via stdin (`curl --data @-`). ACCEPT: probe - harness green; a grep proves no secret-bearing var appears in a curl argv. - -H-3 -- record-keeping for the two new scripts - tenant-offboard.sh + tenant-offboard tests and vault-kv-health.sh (+ inner - probe, + tests) have NO changelog entries or identifiers. Add entries - (what/why/revert) under the next-free DOCFIX numbers; vault-kv-health cites - D-068 item 3 -- link it there too. - -H-4 -- claude.ai skill copy parity - The chat-side installed skill visible at session close lacked the v1.3 - hermetic-harness scar. Confirm the re-upload landed (a fresh chat session - sees the current mount); if not, re-upload the .skill regenerated from - .claude/skills/openstack-cloud-ops. Standing habit: any edit to the skill - source -> repackage -> re-upload. - -H-5 -- Claude Code guardrail smoke test (if not already done) - Fresh session on the jumphost: `/permissions` shows the three rule sets; - `maas list` is hard-blocked by the hook; a `juju status` runs unprompted; - any `--apply` prompts. One minute; proves settings loaded. - --------------------------------------------------------------------------------- -## 2. Verify-live queue (read-only CHECKs; jumphost) --------------------------------------------------------------------------------- - -V-1 (DOCFIX-067 severity ruling) -- octavia deployed-cert SAN read - The runbook now derives the SAN dynamically; the LIVE cloud's controller - cert was generated under the old literal. Read what is actually deployed: - - CHECK (read-only) -- jumphost - ```bash - ( { - CRT=$(juju ssh -m openstack octavia/leader -- \ - 'sudo cat /etc/octavia/certs/controller_cert.pem 2>/dev/null || sudo ls /etc/octavia/certs' &1) - printf '%s\n' "$CRT" | openssl x509 -noout -text 2>/dev/null \ - | grep -A2 'Subject Alternative Name' || printf '%s\n' "$CRT" | head -5 - } ) - ``` - (Path may differ -- if the ls fallback fires, locate the controller cert per - phase-01 1.0-GEN and re-read.) RULE ON RESULT: if SAN carries the old - 10.12.4.233 and octavia is green (expected), record RISK-accepted in the - changelog -- regenerated at next redeploy; do NOT hot-rotate certs for this. - -V-2 (D-069) -- second-person unseal rehearsal - Someone other than the initializing operator performs the manual 3-of-5 - unseal per ops-restart-procedure Stage 3. This is an ACCEPTANCE item, not a - nicety: "keys work" and "a second human can use them" are different facts. - Requires SEC-003 custodian assignment first (section 3). - -V-3 -- channel assert: AUTOMATED (no action) - preflight.sh P3 now runs the Charmhub channel assert live on every pre- - deploy run; the manual `juju info` probe from the old queue is retired. - --------------------------------------------------------------------------------- -## 3. Operator rulings pending (nothing proceeds without these) --------------------------------------------------------------------------------- - -R-1 (SEC-003 / D-069) -- vault unseal-key custodian assignment - Policy is ADOPTED (split custody, no individual holds threshold); the - ASSIGNMENT (who, what media) is operator input, recorded off-repo; flip the - ledger row when done. - -R-2 (D-043 caveat) -- capi-mgmt-v2 auto-resume exclusion - Cloud-wide auto-resume means the mgmt VM WILL resume after host reboots; - its manual-start policy now governs deliberate stops only. If a REAL - exclusion is wanted, it needs a mechanism (options to draft on request); - otherwise close the caveat with a one-line D-043 amendment accepting it. - -R-3 (D-063, PROPOSED/OPEN) -- capi-mgmt SG hardening - Options a/b/c recorded in design-decisions; option (a) requires the - MEASURED post-NAT conductor source, never inferred. Phase-07 depends on - 6443 reachability -- rule before the next redeploy or explicitly defer. - --------------------------------------------------------------------------------- -## 4. The deploy path (the reason for all of the above) --------------------------------------------------------------------------------- - -D-1 -- Pattern-A full redeploy (VR0 DC0) - Now runnable end-to-end from the docs alone: phase-00 (D-061 DESTROY path: - teardown-destroy -> 8_lbaas -> OSD wipe -> reenroll -> carve x4 -> standup) - -> `bash scripts/preflight.sh` PASS -> phases 01-08 gated (keystone policy - now arrives via the bundle; phase-03 Step 3.4 gates PO: + behavioral G3) - -> `bash scripts/cloud-assert.sh --capture` -> commit the asbuilt/ BOM. - Session discipline: `bash scripts/run-logged.sh phase-NN-` first. - -D-2 -- D-011 acceptance: implement validate.sh - The placeholder's TODO is now aligned to the AMENDED decisions (manual - unseal, no snapshots, no Designate) and points at the building blocks - (cloud-assert / tenant-acceptance / run-tests-all). Still to author: VIP - reachability from jumphost + tenant VM, the LB round-robin/failover pattern - test, timed fresh-tenant Magnum e2e, and V-2 as a checklist item. - -D-3 -- confirm D-042 fix status - Memory says the seven-stage magnum-capi-helm contract-ref runbook was - staged pending execution; the repo record should say whether it ran. - Confirm from the changelog/decisions and either execute or close. - --------------------------------------------------------------------------------- -## 5. v1-close checklist (execute after D-011 passes) --------------------------------------------------------------------------------- - -C-1 Consolidate the 10 per-phase do-documents into docs/v1-deploy-runbook.md - (structure preserved in git history via the consolidation commit). -C-2 Consolidate/close docs/docfix-draft-20260702.md -- everything implemented - graduates to the changelog; anything decision-shaped to design-decisions. -C-3 Repo visibility back to PRIVATE (SEC-004; Settings -> Options). -C-4 Rotate the libvirt SSH credential (SEC-001 -- exposed 2026-06-26). -C-5 Ledger review pass (docs/security-ledger.md) -- every row OWNED + dated. - --------------------------------------------------------------------------------- -## 6. Deferred, with explicit revisit triggers --------------------------------------------------------------------------------- - -- OS sandboxing (bubblewrap) for Claude Code Bash: revisit at Roosevelt or if - a prompt-injection-shaped incident ever occurs on the jumphost. -- Managed settings (disableBypassPermissionsMode): when a SECOND operator - gets jumphost access. -- Plugin graduation (wrap the skill + slash-commands /preflight /cloud-assert - + hooks): when Claude Code is the daily driver and prompt-count friction is - measurable. The skill remains the source either way. -- Jenkins CI poll job for repo-lint/gauntlet: when the tenant-rehearsal - Jenkins graduates to an ops instance. -- rabbitmq min-cluster-size:3 (D-009 amendment): Roosevelt bundle delta only. -- SSH on GitBucket (port 29418), IPv6 dual-stack, NetBox import bundle: - v2-deferred as previously recorded. - --- end of register -- diff --git a/docs/handoff-20260705-open-items.md b/docs/handoff-20260705-open-items.md deleted file mode 100644 index 6ac9f74..0000000 --- a/docs/handoff-20260705-open-items.md +++ /dev/null @@ -1,179 +0,0 @@ -# Handoff -- open items register (2026-07-05 session close) - -**Supersedes** `docs/handoff-20260703-open-items.md` (kept for history). This is the active handoff. -**Repo HEAD at close:** `f28de74` ("clearing queue"). **Next-free:** D-073, DOCFIX-091, BUNDLEFIX-012. -**Start every new session by:** reading this file + `docs/session-ledger.md`, then running -`bash scripts/ledger-scan.sh` AND `bash scripts/ledger-scan.sh --fences`, and reconciling. - ---- - -## 0. Current state (as-built) - -- Multi-tenant buildout is functionally complete through the beta tenant acceptance tests. -- `validate.sh` D-011 acceptance suite is modular and the **gauntlet is ALL GREEN (31 harnesses)**. -- **Vault decision recorded (D-068 amendment):** bundle pins vault `1.8/stable` (reactive charm, - integration-compatible); the `1.16` operator lineage is RULED OUT (certs interface V0->V1 breaks - the 18 `vault:certificates` relations, Raft-only drops the mysql backend, no upgrade path, BUSL, - community-broken with Ceph). "Get off EOL vault 1.8.8" stays OPEN under D-068. Evidence: - `docs/D-068-vault-1.8-vs-1.16-analysis.md`. Live vault is `1.8/stable` rev 714 (in-channel current). -- **Session-ledger is fenced** (delimited-section convention: machine-derived / main-chat / jumphost / - shared), and `ledger-scan.sh --fences` validates it (4 sections, 0 errors). Stay in your lane. -- The last jumphost ops-update window (`ops-update-20260705`) is CLOSED; fleet refreshed to current - in-channel revisions; controllers/agents at juju 3.6.25; appendix-B re-baselined. - -**Minor reconcile item for the next session:** the ledger "State facts" bullet still reads -"ops-update-20260705 in flight" and "D-068 unruled" -- both stale (window closed; D-068 ruled). It's -in the append-only shared section, so append a correction rather than rewrite another stream's line. - -## 1. Immediate (small, well-defined; good first tasks) - -- Batch-3 live validation of the D-011 checks (see the verify-live queue below) -- the main thing - standing between here and a clean D-011 acceptance close. -- `offboard` v2: `--sweep-magnum-orphans` mode (orphan per-cluster trustee in the magnum domain); - stage-3 app-cred idempotency re-run guard. (Both logged-not-actioned.) -- DOCFIX candidate: sweep the repo for other `-f json ... 2>&1` stderr-merge hazards (DOCFIX-085 class). -- `host_href=None` barbican observation (exercised fine in stage 6; probably closeable after a look). - -**Window runbook (2026-07-06):** the whole batch-3 + foil + D-073-apply sequence is -scripted as ONE do-document: `runbooks/d011-batch3-window-DRAFT.md` (DOCFIX-093). -B2/D-068 discussion input ready: `docs/D-068-openbao-assessment-DRAFT.md`. - -## 2. Verify-live queue (read-only/gated CHECKs; run on the real cloud before trusting) - -ALL CLEARED 2026-07-06 (d011-batch3 window, addendum 27): -- **d011-04** headroom parsing: fixed live (DOCFIX-096: --long for compute_id; - display-name hypervisor keys); OK path confirmed, disruptive half PASS under - the A2 conditional pre-approval. OCCM LB-name + agnhost RR assumptions held. -- **d011-05** foil satisfied by the foil1 tenant (onboarded in-window); PASS - incl. P3 isolation. -- tenant-assert `keypair list --user` confirmed under admin; admin-side - app-cred list measured always-403 -> check reworked (DOCFIX-095). -- H3 CANARY_SSH_USER confirmed (ubuntu). -STILL OPEN here: Horizon manual checks (onboarding contract 5.1/5.3) on a -live tenant -- operator, browser; non-blocking. - -## 3. Operator rulings pending (nothing proceeds without these) - -- **D-071** (controller update cadence / patch policy) -- stays **PROPOSED**; ratification is GATED on - the pre-DC-DC controller HA/backup planning session (see section 6). Do NOT ratify without it. - Backups are confirmed to exist (juju 3.6 `create-backup`); the open half is single-controller HA. -- **D-068 item 1** vault modernization -- `1.16` is ruled out; the OPEN question is HOW to get off EOL - vault 1.8.8 for production. Candidate paths (none adopted): wait for the OpenStack service charms to - gain tls-certificates V1 support, evaluate OpenBao (MPL fork), or explicit EOL risk-acceptance for - VR0 with a production remediation deadline. Plus D-068 items 2 (Vault listener TLS for Roosevelt -- - vault_url is cleartext http today) and 3 (AppRole secret_id TTL audit + proactive auth health probe). -- ~~D-050 / list_trusts hardening~~ RESOLVED: D-050 CLOSED via D-051 (ruling - 2026-07-06); the list_trusts hardening was adopted as D-073 and APPLIED live - 2026-07-06 (addendum 27; behavioral verify green). - -## 4. Open security-ledger rows - -- **SEC-001** OPEN -- rotate credentials after the current rebuild completes. -- **SEC-003** OPEN -- assign unseal-key custodians + rehearse the second-person unseal (D-069). -- **SEC-004** OPEN -- flip repo visibility to PRIVATE at v1 close (currently public for web_fetch). - -## 5. The deploy path (the reason for all of the above) - -Execute the Roosevelt bare-metal multi-datacenter deployment using the Omega Cloud v1 runbooks as the -template; the guiding constraint remains minimizing delta to Roosevelt. D-011 (amended per D-019) -acceptance validation is the near-term gate -- confirm all criteria post multi-tenant work, which is -what the batch-3 verify-live queue closes out. - -## 6. v1-close / project-completion checklist (execute after D-011 passes) - -- Consolidate the 10 per-phase do-documents into a single `docs/v1-deploy-runbook.md`; drop the - per-phase files; preserve the 10-doc structure in git history via the consolidation commit. -- Set repo visibility PRIVATE (SEC-004). -- Hold the pre-DC-DC controller HA/backup planning session -> then rule D-071. -- **Onboarding completion package (operator-ruled 2026-07-06; complete BEFORE closing this - deployment test env).** A written, tested, and validated end-to-end onboarding workflow: - scripts + internal docs + drafted client-facing docs (drafts/living, no final polish needed). - Items (H-numbers per docs/tenant-onboarding-contract.md section 6): - - H1 `tenant-assert.sh` post-onboard verifier, with harness (also offboard preflight + - periodic drift sweep). DELIVERED addendum 22; live validation pending (foil onboarding). - - H2 stage-3 app-cred idempotency re-run guard in tenant-onboard.sh. - DELIVERED addendum 23 (also covers the keypair; harness 12/12). - - H3 canary access-proof stage (boot canary + FIP + SSH via tenant keypair + teardown) -- - the ruled meaning of "confirm tenant SSH access into their domain" (A1.4 confirmed). - DELIVERED addendum 23 as stage7 (explicit-only); live validation pending (foil). - - H4 `keystone-policy-drift.sh` (script backlog item 6) wired as a periodic check. - (Only H-item not yet built; also the main-chat backlog owner -- coordinate.) - - H5 no-ad-hoc rule embedded in the onboarding runbook. DONE addendum 25. - - Horizon manager-role GUI identity probe (contract section 5.3) validated live. - - Client-facing draft package (intake form, welcome/engagement doc, self-service guide) - -- RULED: top-level clientdocs/; drafts DELIVERED addendum 22 (living docs, sweep on - contract-doc changes). - - Full workflow validated on a live onboarding end-to-end (the d011-05 foil tenant is the - natural first candidate). - -## 7. Deferred, with explicit revisit triggers - -- **GitBucket SSH** (System Settings -> Integrations; bind 0.0.0.0:29418, public git.baldurkeep.com:29418) - -- verify the deployment topology (Docker port map vs systemd vs firewall) before declaring it works. - v1 uses HTTPS basic auth. Trigger: v2 / DC-DC phase. -- **IPv6 dual-stack** integration + **NetBox restructure/import** -- deferred to DC-DC / v2. -- **Roosevelt hardware sheets** needed to resolve the four-NIC collapse path (D-059, parked). -- **SpryLogin / identity platform** decision (FreeIPA vs Keycloak vs Authentik). - ---- - -## 8. Startup prompt -- NEW MAIN CHAT (claude.ai stream) - -> You are resuming the Omega Cloud v1 / Baldurkeep project (commercial multi-tenant Charmed OpenStack -> Caracal 2024.1). You are the **main claude.ai chat stream**. A parallel **Claude Code (jumphost)** -> stream runs concurrently -- coordinate, do not collide. -> -> **First, orient (do this before anything else):** pull the repo; read -> `docs/handoff-20260705-open-items.md`, `docs/session-ledger.md`, and `docs/design-decisions.md`; -> run `bash scripts/ledger-scan.sh` and `bash scripts/ledger-scan.sh --fences`; reconcile the ledger -> narrative against the scan. The repo is authoritative over your memory of it. -> -> **Your lane / discipline:** you work a sandbox clone via gated copy-paste; the operator (Jesse) runs -> blocks on the jumphost and pastes output back. Deliver multi-file changes as **repo-relative ZIPs** -> (never loose files). Committed files are **pure ASCII + LF**. In the session-ledger, edit only the -> `main-chat` fenced section; the machine-derived block is regenerated by `ledger-scan` (never -> hand-typed); the shared section is append-only. When you consume a D-/DOCFIX-/BUNDLEFIX- number, -> grep for next-free first and coordinate with Code. **Harness-first**: run test harnesses before -> executing live; own mistakes plainly and immediately. -> -> **Current state:** multi-tenant buildout complete through beta acceptance; validate.sh D-011 suite -> modular, gauntlet ALL GREEN (31). Vault is settled on `1.8/stable` (D-068: 1.16 ruled out; EOL is -> the open item). Reconciliation is done and clean. -> -> **Likely next work (confirm with Jesse):** batch-3 live validation of the D-011 checks (verify-live -> queue in the handoff) -> D-011 acceptance close; then project-completion (consolidate do-docs, flip -> repo private). D-068 EOL-vault modernization research is available when Jesse wants it. Do NOT ratify -> D-071 or move vault to 1.16 -- both are operator/decision-gated. -> -> Standing meta-instruction from Jesse: before answering, state what you need to know and any -> assumptions you'd otherwise make. Debate best-practice deviations with sourced rationale. - -## 9. Startup prompt -- CLAUDE CODE (jumphost stream) - -> You are resuming the Omega Cloud v1 / Baldurkeep project as the **Claude Code (jumphost) stream**, -> executing live on the cloud. A parallel **main claude.ai chat** stream runs concurrently -- -> coordinate, do not collide. -> -> **First, orient:** read `docs/handoff-20260705-open-items.md`, `docs/session-ledger.md`, and -> `docs/design-decisions.md`; run `bash scripts/ledger-scan.sh` and `bash scripts/ledger-scan.sh -> --fences`; reconcile. Repo is authoritative over memory. -> -> **Your lane / discipline:** in the session-ledger, edit only the `jumphost` fenced section; the -> shared section is append-only; do not touch the `main-chat` section. When you consume a -> D-/DOCFIX-/BUNDLEFIX- number, record it in its canonical home (design-decisions / changelog) FIRST, -> then regenerate the machine-derived block by running `ledger-scan` and pasting its output -- never -> hand-type numbers. Keep the ``/`` fences intact; pull before you edit. -> Committed files ASCII + LF. Gate destructive/irreversible steps individually; verify before mutate. -> -> **Current state:** last ops-update window (ops-update-20260705) is CLOSED; fleet current in-channel; -> juju 3.6.25. Vault is `1.8/stable` rev 714 (current in-channel). -> -> **Hard constraints:** -> - **Vault stays `1.8/stable`.** Do NOT move it to `1.16` (D-068: fundamentally different, -> incompatible charm). In-channel currency only. -> - **D-071 stays PROPOSED.** Do NOT ratify it -- that is operator authority, gated on the pre-DC-DC -> controller HA/backup planning session. Controller backups exist (create-backup on 3.6); the open -> half is single-controller HA. -> -> **Likely next work (confirm with Jesse):** support batch-3 live validation of the D-011 checks -> (verify-live queue in the handoff); controller update cadence per D-071 ONCE it is adopted (not yet). diff --git a/docs/incident-20260712-opnsense-edge-boot-triplefault.md b/docs/incident-20260712-opnsense-edge-boot-triplefault.md deleted file mode 100644 index 48060e6..0000000 --- a/docs/incident-20260712-opnsense-edge-boot-triplefault.md +++ /dev/null @@ -1,113 +0,0 @@ -# Incident 2026-07-12: Office1 OPNsense edge triple-faults at the BTX loader - -> ## ROOT CAUSE FOUND 2026-07-12 (DOCFIX-188) -- READ THIS FIRST, THEN STOP -> -> **The guest had 2 MiB of RAM.** Not a CPU, nesting, machine-type, or console problem. -> Everything below this box is the ORIGINAL (wrong) investigation, retained for the record. -> **Do not work the "ranked next steps" -- they all chase the wrong layer.** -> -> `dmacvicar/libvirt` >= 0.9 changed `memory` from MiB (the 0.8-era meaning the modules were -> written against) to **raw libvirt units, defaulting to KiB**. With no `memory_unit`, -> `memory = 2048` rendered `2048` -> QEMU **`-m size=2048k`** = -> **2 MiB**. Measured end-to-end: module input -> tofu state -> domain XML -> `virsh dominfo` -> (`Max memory: 2048 KiB`) -> the live QEMU cmdline. -> -> `boot2` is tiny and fits in 2 MiB, so it echoes `/boot.config` and *then* triple-faults -> handing off to `/boot/loader`, which does not fit. **That is why the fault was -> deterministic at exactly 262 bytes and immune to every CPU/machine/disk/console change -> tried -- none of them touched the cause.** -> -> **Fix:** `memory_unit = "MiB"` on the `libvirt_domain` in all three VM modules -> (`opnsense-edge`, `cloudinit-vm`, `node-vm` -- the same defect was latent in all of them -> and would have broken every future VR1 VM). Guarded against recurrence by -> `scripts/opentofu-validate.sh` S1. See -> `docs/changelog-20260712-libvirt-memory-unit-rootcause.md`. -> -> **Lesson for the next incident:** a bootloader that dies at a fixed byte offset, immune to -> every knob you turn, is a *resource* problem, not a CPU-feature problem. The domain XML and -> the QEMU cmdline are ground truth -- read them before theorising about nested virt. -> -> **Status:** root cause fixed in repo; live boot verification is the remaining gated step. - -**Original status (SUPERSEDED):** OPEN. The Office1 OPNsense edge VM is built and *starts*, but OPNsense -triple-faults ~262 bytes into boot. Blocks the Office1 headend (router+DHCP for -`office1-local`) and therefore the Office1 NetBox VM. Written at a context limit as a -resume artifact -- a fresh session should read this + `docs/session-ledger.md` first. - -## Symptom (confirmed, reproducible) -- Domain `office1-opnsense` (libvirt, `qemu:///system`) enters `running (booted)` then - goes to `paused (unknown)`. `virsh resume` fails with *"cont: Resetting the Virtual - Machine is required"* -> the guest **triple-faulted**. -- Serial capture (`/var/lib/libvirt/vr1/staging/office1-opnsense-serial.log`, `root:600`, - read with `sudo cat`) contains exactly, every boot: - ``` - /boot.config: -S115200 -h -D - ``` - i.e. FreeBSD `boot2` echoes `/boot.config` (serial 115200, `-h` serial console, `-D` - dual console) and then triple-faults **handing off to the BTX `/boot/loader`**. The log - mtime updates on each boot (confirmed fresh, not stale); it is deterministically 262 bytes. - -## Environment -- **Double-nested virt:** OPNsense guest -> `vcloud` host (itself a VM: virtio NIC/disk) - -> outer hypervisor. Host CPU model reported by libvirt: **AMD `Opteron_G3`**. Nested - KVM on (`kvm_amd/parameters/nested = 1`). - - **CORRECTION 2026-07-12 (DOCFIX-189): the `Opteron_G3` label is a RED HERRING.** The - real CPU is an **AMD EPYC 9965 (Zen 5, family 26)** (`/proc/cpuinfo`). libvirt's CPU - model database does not know family 26, so it falls back to the oldest matching model - name. The prior session built a nested-virt/ancient-CPU theory partly on this artifact. - Read `/proc/cpuinfo`, not libvirt's model guess. -- Image: OPNsense **26.1 nano** amd64 (`scripts/opnsense-prep-image.sh 26.1`), prepped to - `/var/lib/libvirt/vr1/office1/opnsense-26.1-nano.qcow2` (11 GiB virtual). -- Module: `opentofu/modules/opnsense-edge`, instantiated as `module "office1_opnsense"` in - `opentofu/main.tf`. LAN=`office1-local` (`vtnet0`), WAN=`office1-wan` (`vtnet1`, NAT - `172.30.1.0/24`). Config ISO (real-ISP-router config, DOCFIX-185) at - `/var/lib/libvirt/vr1/staging/office1-opnsense-config.iso`. - -## What was tried -- ALL applied and verified in the domain XML, NONE resolved it -1. **Serial console added** (module gap -- nano is serial-only). Got the boot *to* the - loader stage (from a no-console early fault to the 262-byte `/boot.config` point). -2. **`machine` q35 -> i440fx** (`pc-i440fx-noble`, verified). No change. (Note: machine + - cpu are create-time; the provider does an in-place "change" that does NOT apply -- - must recreate the domain: `virsh destroy && virsh undefine`, then `tofu apply`.) -3. **Disk: COW overlay -> direct per-VM copy** of the nano (verified 11 GiB / 2.14 GiB - allocated, no backing). No change. -4. **CPU `host-passthrough`** (verified). No change. -5. **Disable AMD `svm`** (`` in XML -- the documented - AMD nested-virt fix, forum: `-cpu host,-svm`). **Still 262 bytes.** NOTE: the provider's - CPU-feature key is `features` (plural); `feature` validates but is silently dropped. - -## Ranked next steps for a fresh session -1. **Full forum CPU flag set**, not just `-svm`: add `+kvm_pv_eoi,+kvm_pv_unhalt` (and try - without `svm` masking too). Likely via `libvirt_domain.qemu_commandline` since the - provider's CPU `features` may not fully translate under `host-passthrough`. Confirm the - masking actually reaches the guest CPUID. -2. **Video device (`-D` dual console).** `/boot.config` requests dual console but the domain - has **no video/graphics device** -- `boot2`'s VGA init may fault. Add a `graphics` + - `video` (VGA/std) device and retry. (Cheap, plausible, not yet tried.) -3. **Memory 2 GB -> 4 GB** (below OPNsense's 3 GB min; guides use 4096). Recreate to apply. - - **CORRECTION (DOCFIX-189): this step was RIGHT FOR THE WRONG REASON, and its premise is - unsourced.** The memory *was* the problem -- but it was 2 MiB, not "2 GB but a bit - small", and no measurement was taken to notice that. The "3 GB min" figure is - UNVERIFIED (no source given); do not propagate it. If 2 GiB proves insufficient after - the real fix, check OPNsense 26.1's documented minimum before picking a number. -4. **UEFI boot (OVMF)** instead of legacy BIOS/BTX -- sidesteps BTX entirely, but the nano - is a BIOS/MBR image, so this needs care (or a different image build). -5. **Outer-hypervisor CPU:** under double-nesting, vcloud's exposed CPU (`Opteron_G3`) may - not provide what FreeBSD's BTX needs. Consider testing a minimal FreeBSD/OPNsense boot - directly on vcloud to isolate whether it's nesting-depth-specific. - -## Reproduce / operate -- Recreate + boot: `cd opentofu && source ~/vr1-stage1.env && virsh -c qemu:///system - destroy office1-opnsense; virsh -c qemu:///system undefine office1-opnsense; tofu apply`. -- Watch: `virsh domstate --reason office1-opnsense`; serial via `sudo cat` the log above. -- Provider CPU/serial schema is attribute-style + nested under `devices`/`cpu`; introspect - with `tofu providers schema -json` (see how `serials`/`features` were found this session). -- As-executed log: `~/as-executed/2026-07-12-dc-dc-phase1-office1.log`. -- Creds (jumphost-only, 0600, `~/vr1-office1-creds/`): SSH key + OPNsense root pw/hash. - **Operator: revoke the pasted NetBox token + these at close.** - -## Cross-refs -- `docs/changelog-20260712-opnsense-edge-boot-fixes.md` (DOCFIX-187, the module changes). -- `docs/changelog-20260712-office1-opnsense-edge-build.md` (DOCFIX-186, the build + the - apparmor/config-iso-staging findings). -- `docs/changelog-20260712-opnsense-edge-real-isp-router.md` (DOCFIX-185, config posture). diff --git a/docs/model-a-fallback-plan.md b/docs/model-a-fallback-plan.md deleted file mode 100644 index fe843dd..0000000 --- a/docs/model-a-fallback-plan.md +++ /dev/null @@ -1,106 +0,0 @@ -# Model A fallback + revert plan (D-123) - -**Purpose.** The operator ruled **Model B** for D-123 (nodes nested inside `vvr1-dc0`, single-object -`virsh destroy` site-down) -- the heavier, higher-risk path (depth-4 nested virt, supersedes -D-103/D-114, ~416 GiB containment VM). This document preserves **Model A** as a fully-specified, -already-implemented fallback so that if Model B fails to deploy, we revert WITHOUT re-engineering. - -**Revert anchor (git).** Model A is not theoretical -- it is the CURRENTLY COMMITTED substrate. The -last commit before any Model B reshape is tagged **`model-a-fallback`** (re-cut 2026-07-16 onto -`114d392`, the **R-3-compliant** Model A layout with 4 storage nodes/DC -- the prior anchor at -`87a7a8a` was R-3-stale, flagged by the Model B design cross-check). To restore Model A: -`git checkout model-a-fallback -- opentofu/` -(or cherry-pick the substrate files), then re-run `bash scripts/opentofu-validate.sh`. No file needs -to be re-authored -- Model A already validates (`tofu validate` Success; 11/11 modules). - ---- - -## 1. Model A architecture (the as-built shape) - -Nodes are **vcloud-level libvirt siblings** of the headend -- NOT nested inside it. This is the -VR0-proven shape and the ADOPTED D-103/D-114 seam. - -``` -vcloud (host, L0) -|-- vvr1-dc0 MAAS rack headend (cloudinit-vm; D-124: 4 vCPU / 8192 MiB / 80 GiB; -| expose_nested_virt = false; legs = metal-admin + office1<->dc0 transit) -|-- vr1-dc0-control-01..03 node VMs (16/65536/150) \ -|-- vr1-dc0-compute-01..02 node VMs (12/49152/100) > vcloud-level siblings, on vr1_dc0_planes -|-- vr1-dc0-storage-01..04 node VMs (8/24576/550) / (4 storage/DC per R-3) -|-- vr1-dc0-edge opnsense (2/2048, 2-NIC: provider-public LAN + vr1-dc0-wan WAN) -|-- vr1-dc0-* planes 6 isolated-L2 libvirt networks (dc-planes) at vcloud level -|-- vr1-dc0-wan NAT /24 simulated ISP uplink (site-wan) -`-- mesh-vr1-dc0-office1 transit leg (office1 <-> dc0) - -nesting depth = 2 (vcloud -> node VM -> nova KVM guest) <- VR0-PROVEN -site-down = destroy the vr1-dc0-* domain GROUP (scripted, gated) -MAAS model = region on Office1 + rack (vvr1-dc0); maas-vm-host registers VCLOUD's virsh so - MAAS discovers the OpenTofu-created node domains (D-103/D-114 as-built) -``` - -## 2. The committed artifacts that embody Model A (revert targets) - -| Artifact | Model A content | -|---|---| -| `opentofu/main.tf` `module "vr1_dc0_node"` | `for_each = local.vr1_dc0_nodes`; created on the **vcloud** libvirt provider; attached to `module.vr1_dc0_planes` outputs (6 NICs, metal-admin first = PXE). | -| `opentofu/main.tf` `module "vvr1_dc0"` | `cloudinit-vm`, **4/8192/80** (D-124), `expose_nested_virt = false`, two legs (metal-admin + mesh transit). A small rack headend that holds NO nodes. | -| `opentofu/main.tf` `module "vr1_dc0_planes"` / `mesh_*` / `vr1_dc0_wan` | all created at **vcloud** level. | -| Step-9 `maas-vm-host` (deferred, DOCFIX-179) | registers **vcloud's** virsh to the DC's MAAS -> MAAS discovers the vcloud-level node domains. | -| `scripts/site-headend-install.sh --role rack` | installs the rack controller on `vvr1-dc0` (no LXD/compose in rack mode). | -| Site-down | a scripted group-destroy of the `vr1-dc0-*` domains (owned by `dc-dc-teardown-rollback.md`); NOT a single `virsh destroy`. | - -## 3. What Model B changes vs Model A (the delta to undo on revert) - -Reverting = undoing exactly these; nothing else moves. - -1. **Node placement:** B retargets `module "vr1_dc0_node"` (and the 6 planes + `vr1-dc0-wan`) from - vcloud's libvirt to **`vvr1-dc0`'s inner libvirt**. A restores them to vcloud level. -2. **Headend sizing:** B resizes `vvr1-dc0` from 4/8192/80 to ~416 GiB (must hold one DC's full node - fleet) and sets `expose_nested_virt = true`. A restores D-124's 4/8192/80. -3. **maas-vm-host target:** B registers `vvr1-dc0`'s inner virsh; A registers vcloud's virsh. -4. **Governance:** B supersedes D-103/D-114; A keeps them ADOPTED as-is. On revert, the D-103/D-114 - supersession is withdrawn. -5. **Nesting depth:** B = 4 (unproven); A = 2 (VR0-proven). -6. **Site-down primitive:** B = one `virsh destroy vvr1-dc0`; A = scripted group-destroy. -7. **Inner root:** B adds `opentofu/vr1-dc0-substrate/` (a new root dir + its own state) and a - `site-headend-install.sh --host-nodes` node-host bootstrap; A has neither. Revert removes them. -8. **Per-DC ISP egress (D-125 bridge-in):** because B nests `vr1-dc0-wan` inside `vvr1-dc0`, B adds a - vcloud-level ISP NAT (`module "vr1_dc0_uplink"`, `site-wan`, cidr `172.30.2.0/24`), a 2nd IP-less - uplink NIC + the `br-vr1-dc0-wan` netplan bridge on `vvr1-dc0`, the new `modules/wan-bridge`, and the - bootstrap's `--uplink-if`/`--wan-bridge` verify. **The ADDRESSING is identical to Model A** -- same - `172.30.2.0/24` (D-115), OPNsense WAN still `.2`; only the libvirt realization differs (bridge through - `vvr1-dc0` vs a direct vcloud-level NAT). **In Model A the extra plumbing does not exist** -- `vr1-dc0-wan` - is a vcloud-level NAT the vcloud-level edge attaches to directly, no uplink NIC, no bridge, no - `wan-bridge` module. Revert removes that plumbing; NO re-address is needed (the address never changed). - (No HELD gate here: the /24 is a ruled literal, not a tfvar.) - -## 4. Revert procedure (if Model B deployment fails) - -1. STOP -- do not attempt to fix Model B in place if nested-virt (depth-4) is the failure mode; that - is the known risk this fallback exists for. -2. `git checkout model-a-fallback -- opentofu/main.tf opentofu/variables.tf opentofu/modules/` - (restores the Model A substrate verbatim -- this also drops `module "vr1_dc0_uplink"`, since Model A - has no vcloud uplink), then `git rm -r opentofu/vr1-dc0-substrate` (the inner root does not exist in - Model A) and `git rm -r opentofu/modules/wan-bridge` (D-125, also absent in Model A), and revert the - `site-headend-install.sh` node-host mode (incl. the D-125 `--uplink-if`/`--wan-bridge` WAN-bridge - verify). The OPNsense WAN address is UNCHANGED (`.2` on `172.30.2.0/24`) -- nothing to restore. -3. `bash scripts/opentofu-validate.sh` -> expect 11/11 PASS (Model A already validates). -4. Re-instate D-103/D-114 as ADOPTED (they were only annotated superseded, not deleted -- see the - sweep's supersession notes; revert removes those annotations). -5. Restore D-124's rack sizing (4/8192/80) in tfvars/main.tf. -6. Adopt the scripted group-destroy site-down (the `vr1-dc0-*` group op) in place of the single-object - destroy. -7. Re-run the Layer-1 gate (`repo-lint`, `run-tests-all`, `tofu validate`) before proceeding. - -## 5. Failure signals that should trigger the revert - -- nova-compute guests fail to boot or are unusably slow at 3x-nested KVM (the depth-4 risk). -- `vvr1-dc0` cannot be allocated ~416 GiB on the host alongside the other layers. -- MAAS enrolment/commissioning breaks because the node domains are no longer vcloud-visible. -- `expose_nested_virt = true` on `vvr1-dc0` destabilises the headend/rack. - ---- - -*Model A remains the recommended engineering choice on delta/risk grounds; Model B is the -operator-ruled choice for its single-object site-down primitive. This plan makes the choice -reversible at low cost. Kept in sync with the D-123 sweep; if Model B changes, update section 3.* diff --git a/docs/netbox-vip-queue.md b/docs/netbox-vip-queue.md deleted file mode 100644 index 5017bd1..0000000 --- a/docs/netbox-vip-queue.md +++ /dev/null @@ -1,126 +0,0 @@ -# Post-deployment NetBox VIP imports (queued from workstream 2) - -**Status:** Queued. To be imported after successful cloud deployment + validation, -once `netbox/ipv4-prefixes-import.py` engineer review unblocks the Provider /22 -prefix import. - -**Background:** Per D-010 (NetBox-upstream policy), IPAM entries should exist in -NetBox before being written into IaC. For v1 testcloud, this rule was relaxed -under workstream 2 (2026-05-22) to avoid blocking the rebuild on the engineer -review. VIPs were written into `bundle.yaml` directly. This document captures -the corresponding NetBox writes that need to happen post-deploy. - -**Scope:** v1 only (IPv4). v2 IPv6 VIPs are out of scope. - ---- - -## Provider prefix (parent -- gating) - -Before any IPAddress entries can be created, the parent prefix must exist: - -| Prefix | Site | Role | Status | -|---|---|---|---| -| `10.12.4.0/22` | VR0 DC0 | provider | Active | - -Created by: `netbox/ipv4-prefixes-import.py` (per D-010, gated on engineer review). - ---- - -## VIP IPAddress entries - -All entries under prefix `10.12.4.0/22`, tenant scope = VR0 DC0 Omega Cloud (or -appropriate testcloud tenant convention). - -| IP | Status | DNS name | Description | -|---|---|---|---| -| `10.12.4.224/22` | Active | `barbican.omega.dc0.vr0.cloud.neumatrix.local` | barbican API VIP -- Charmed OpenStack hacluster | -| `10.12.4.225/22` | Reserved | -- | RESERVED for ceph-radosgw HA VIP in v2 (workstream-2 decision; ceph-radosgw HA deferred to v2) | -| `10.12.4.226/22` | Active | `cinder.omega.dc0.vr0.cloud.neumatrix.local` | cinder API VIP -- Charmed OpenStack hacluster | -| `10.12.4.227/22` | Reserved | -- | RESERVED for designate VIP in v2 (per D-019; Designate deferred to v2) | -| `10.12.4.228/22` | Active | `glance.omega.dc0.vr0.cloud.neumatrix.local` | glance API VIP -- Charmed OpenStack hacluster | -| `10.12.4.229/22` | Active | `keystone.omega.dc0.vr0.cloud.neumatrix.local` | keystone API VIP -- Charmed OpenStack hacluster | -| `10.12.4.230/22` | Active | `magnum.omega.dc0.vr0.cloud.neumatrix.local` | magnum API VIP -- Charmed OpenStack hacluster | -| `10.12.4.231/22` | Active | `neutron.omega.dc0.vr0.cloud.neumatrix.local` | neutron-api API VIP -- Charmed OpenStack hacluster | -| `10.12.4.232/22` | Active | `nova.omega.dc0.vr0.cloud.neumatrix.local` | nova-cloud-controller API VIP -- Charmed OpenStack hacluster | -| `10.12.4.233/22` | Active | `octavia.omega.dc0.vr0.cloud.neumatrix.local` | octavia API VIP -- Charmed OpenStack hacluster | -| `10.12.4.234/22` | Active | `horizon.omega.dc0.vr0.cloud.neumatrix.local` | openstack-dashboard (Horizon) VIP -- Charmed OpenStack hacluster | -| `10.12.4.235/22` | Active | `placement.omega.dc0.vr0.cloud.neumatrix.local` | placement API VIP -- Charmed OpenStack hacluster | -| `10.12.4.236/22` | Active | `vault.omega.dc0.vr0.cloud.neumatrix.local` | vault VIP -- Charmed Vault hacluster (D-006) | - -**Notes:** - -- Mask is `/22` (the parent prefix mask), not `/32` -- NetBox convention for - endpoint IP addresses within a prefix. -- The Reserved slots at `.225` and `.227` document v2 intent without consuming - active allocations. When v2 work brings ceph-radosgw HA and Designate online, - those entries' Status flips Reserved -> Active and the bundle's `# v2-deferred:` - markers are uncommented. -- `nova-cloud-controller` charm -> DNS short name `nova` (catalog service name, - not charm name). -- `openstack-dashboard` charm -> DNS short name `horizon` (project name). -- `neutron-api` charm -> DNS short name `neutron`. - ---- - -## FIP pool -- for completeness (not part of workstream 2) - -Per D-003, the Provider /22 also carries the Neutron FIP pool. These are NOT -individual IPAddress entries; they're modeled as an IP Range under the prefix: - -| Range | Purpose | -|---|---| -| `10.12.4.10 - 10.12.4.223` | Neutron FIP pool (created by `ipv4-prefixes-import.py`) | -| `10.12.4.224 - 10.12.4.254` | API VIP pool (the 13 entries above + future) | - -Neutron `allocation_pools` for the provider subnet MUST exclude `.224-.254` -- -this is enforced in `runbooks/06-tenant-setup.md` (or wherever the provider -subnet is created). - ---- - -## Execution path (when unblocked) - -1. Confirm engineer review of `netbox/ipv4-prefixes-import.py` has signed off. -2. Run `netbox/ipv4-prefixes-import.py` -- creates the Provider /22 prefix + FIP - IP Range + API VIP IP Range. -3. Add the 13 IPAddress entries from the table above. Two paths: - - **Web UI:** Per-entry manual creation. Tedious but reviewable. - - **API/script:** Extend `ipv4-prefixes-import.py` with a VIP-addresses - section, OR write a separate `netbox/ipv4-vips-import.py` that reads - this document (or a YAML/CSV companion). Idempotent (skip-if-exists). -4. Sanity check: NetBox prefix view of `10.12.4.0/22` shows all 13 entries. -5. Cross-check: every active VIP in `bundle.yaml` has a matching Active - entry in NetBox; the Reserved entries at `.225` and `.227` have no - corresponding bundle entries (v2-deferred). - ---- - -## Change log - -| Date | Change | Reference | -|---|---|---| -| 2026-05-22 | Document created. 12 active VIP allocations queued + 1 v2-reserved slot. | Workstream 2 -- VIP allocation + hacluster activation | -| 2026-05-27 | Designate VIP at `.227` flipped Active -> Reserved per D-019 (Designate deferred to v2). Active count: 11; Reserved count: 2. | D-019 | - ---- - -## Management-plane reservations (MAAS-side; draft <- live, queued for NetBox) - -**Status:** Queued / draft. Observed live in MAAS during the 2026-06-11 rebuild; recorded here -draft <- live (live MAAS state is truth; this doc is the reconciliation target). NOT yet written -to NetBox -- D-010 (NetBox is unmutated until the IPAM design is confirmed). Purpose-pending: -the ranges were reserved on the mgmt plane but their per-IP role is not yet assigned. - -Two contiguous reservation ranges on the provider + metal segments (cf. D-003 -- the provider -network carries ext_net + API VIPs on one L2; D-010 -- NetBox-upstream policy): - -| Range | Segment | Status | Purpose | -|---|---|---|---| -| `10.12.4.101` - `10.12.4.110` | provider `10.12.4.0/22` | Reserved (MAAS) | mgmt-plane, role TBD -- reconcile before NetBox write | -| `10.12.8.101` - `10.12.8.110` | metal/internal (`10.12.8.x`; prefix TBC) | Reserved (MAAS) | mgmt-plane, role TBD -- reconcile before NetBox write | - -**To reconcile before NetBox import:** confirm each range's intended role (host mgmt, OOB, spare -pool, etc.) and the metal-segment prefix, then create the matching NetBox IPRange/IPAddress -entries with the resolved role + DNS convention. Do NOT mutate NetBox until the role assignment -is confirmed (D-010). The /22-mask convention (not /32) applies, as for the VIP entries above. - diff --git a/docs/phase-00-maas-standup-notes.md b/docs/phase-00-maas-standup-notes.md deleted file mode 100644 index 7e89dbb..0000000 --- a/docs/phase-00-maas-standup-notes.md +++ /dev/null @@ -1,52 +0,0 @@ -# phase-00 MAAS stand-up (D-058) -- notes - -`scripts/phase-00-maas-standup.sh` brings MAAS to the D-058 plane topology -idempotently. **Dry-run is the audit** (default): it resolves live ids BY -CIDR/name (PATTERN-1) and prints the plan, changing nothing. `--apply` executes. - -## Behavior per resource -- present and correct -> **SKIP** -- absent -> **CREATE** (fabric / VLAN / subnet / space / gateway / managed / dns / API-VIP reserve) -- present but bound to the wrong plane or wrong VID -> **DRIFT** (reported, never touched) - -A re-CIDR is destructive (MAAS cannot change a subnet CIDR in place), so it is -**out of scope by design**: the drift scan reports it as MIGRATE-NEEDED and the -script refuses to build onto a CIDR the wrong plane occupies. Verified idempotent -on `--apply` (zero mutations against an already-correct cloud). - -## D-058 target (what it stands up) -provider-public 10.12.4.0/22 (untagged, gw .4.1) | provider-vip 10.12.8.0/22 -(VID 104 on the provider fabric, gw .8.1, VIP band .8.2-.100) | metal-admin -10.12.12.0/22 (untagged, gw .12.1, VIP band .12.2-.100) | metal-internal -10.12.16.0/22 (VID 103 on the metal fabric, VIP band .16.2-.100) | data-tenant -10.12.20.0/22 | storage 10.12.32.0/22 | replication 10.12.36.0/22. Untagged base -planes are created first so their tagged siblings can ride the same fabric -(provider-public->provider-vip, metal-admin->metal-internal). - -## Single MAAS-address authority (D-058 consolidation) -- THIS script owns topology AND every reserved range: the API-VIP bands, the - Neutron FIP pool (10.12.5.0-.7.254 on provider-public), and the mgmt reserves. -- `phase-00-maas-carve.sh` is RETIRED -- its reserves are folded in here, and its - gated stale-range delete is subsumed by the teardown + re-CIDR step. -- It never deletes anything. - -## Relationship to provider-vip-standup.sh -This generalizes that script from one plane to the whole topology; provider-vip is -now just one row of the table. `provider-vip-standup.sh` remains the targeted -single-plane tool (add provider-vip to an already-D-058 cloud). Both source -`lib-net.sh`; no conflict. - -## Tests -`tests/phase-00-maas-standup/` -- fake `maas` + real jq, fixtures generated by -`make_fixtures.py`. Four scenarios, ALL PASS: fresh MAAS (full create plan), -D-058 done (all SKIP / zero WOULD), D-052 current (the three migrating planes -drift + refuse), wrong-VID (vid drift). Run: `bash tests/phase-00-maas-standup/run-tests.sh`. - -## The current-cloud gap (next deliverable) -The live cloud is D-052/053, so this stand-up will report metal-admin (.8), -metal-internal (.12), data-tenant (.16) as DRIFT -- those CIDRs are reassigned by -D-058. The destructive cutover (release/teardown so subnets have no links, delete -the old subnets in collision-safe order, then `--apply` to build the new scheme) -is a **separate gated step**, not this script. Sequence: teardown -> delete old -subnets -> `phase-00-maas-standup.sh --apply` -> `phase-00-maas-carve.sh` -(D-058) -> jumphost bridge re-IP (D-058 ordering trap) -> deploy. diff --git a/docs/repo-lint-nextfree-bug-FINDING.md b/docs/repo-lint-nextfree-bug-FINDING.md deleted file mode 100644 index 42723a1..0000000 --- a/docs/repo-lint-nextfree-bug-FINDING.md +++ /dev/null @@ -1,68 +0,0 @@ -# FINDING: repo_lint.py L5 "next-free identifiers" is unreliable -- do not use it for numbering - -**Status:** DOCFIX candidate. Read-only analysis by the main-chat stream at repo HEAD -`59a7c73` (2026-07-06). No code changed. Code assigns the DOCFIX number when actioning -(re-run `ledger-scan` for next-free first; main-chat consumed none of these). - -**One-line:** `scripts/repo_lint.py` prints an `[info] L5 next-free identifiers` line that -is wrong on all three counters and disagrees with `ledger-scan`. Numbers must be taken -from `ledger-scan`, never from repo_lint. This is a live collision risk while the jumphost -stream is consuming DOCFIX numbers rapidly. - -## Evidence - -At this HEAD the two tools disagree: - -| counter | repo_lint L5 says | ledger-scan says (authoritative) | ground truth | -|---|---|---|---| -| D | next-free `D-075` | next-free `D-074` | highest header `D-073`; `D-074` exists only as a PROPOSAL mention | -| DOCFIX | next-free `DOCFIX-100` | next-free `DOCFIX-106` | highest consumed `DOCFIX-105`; `DOCFIX-106` is a next-free pointer | -| BUNDLEFIX | next-free `BUNDLEFIX-051` | next-free `BUNDLEFIX-012` | highest consumed `BUNDLEFIX-011`; `051` traces to a test fixture (below) | - -`ledger-scan` is correct on all three. Each repo_lint number is wrong via a DIFFERENT -defect, which is why this is worth fixing rather than tweaking. - -## Root cause -- three independent defects in the L5 next-free block - -The block (`repo_lint.py`, L5 identifier numbering) computes next-free as `max(seen)+1` -over `re.finditer(r"\b(D|DOCFIX|BUNDLEFIX)-(0\d{2})\b", ...)` across `all_text()`: - -1. **Band-limited regex `0\d{2}`.** It only matches identifiers `000`-`099`. It is blind to - `DOCFIX-100..106` (the current live range), so it reports `DOCFIX-100` when true next-free - is `106`. This is the same defect `ledger-scan` already carried and Code already fixed - (its regex is now `[0-9]{3}`); repo_lint still has the old band-limited form. -2. **Mention-counting, no header-authority, no pointer-exclusion.** It counts every - identifier OCCURRENCE, not definition-headers, and does not skip "Next-free:" pointer - lines. So a PROPOSAL reference inflates it -- the main-chat CIDR plan mentions `D-074`, - which is exactly why repo_lint reports `D-075`. `ledger-scan` uses design-decisions - headers for D and excludes `next-free` pointer lines. -3. **Scans the test tree.** `all_text()` includes `tests/`. `tests/ledger-scan/run-tests.sh` - contains a deliberately-high fixture line - (`Next-free: D-071, DOCFIX-099, BUNDLEFIX-050 ... must NOT inflate`) written to PROVE - `ledger-scan` ignores pointer lines. repo_lint ingests that fixture's `BUNDLEFIX-`050 - and reports `051` -- the precise trap the fixture was authored to catch. - -Correction of the record: an earlier main-chat note called `BUNDLEFIX-`051 "spurious." It is -(tokens above de-fanged 2026-07-06, jumphost stream: identifier-shaped tokens above the -real high-water mark in docs/ prose inflate the ledger-scan counters -- the standing rule -this very FINDING is about; the table rows survive because "next-free" lines are excluded.) -not -- it is `max+1` of the fixture token in defect 3. The imprecision is corrected here. - -## Fix options (Code's lane; not implemented here) - -**Recommended -- remove the next-free print from repo_lint entirely.** repo_lint's actual L5 -job is the duplicate-definition-heading collision guard; keep that. Next-free is -`ledger-scan`'s job, and having two tools derive it by different regexes is how they drifted -apart. Single source of truth (the same principle already floated for the machine-derived -ledger cache). Lowest risk, removes the drift surface. - -**Fallback -- if repo_lint must keep printing next-free, make it match `ledger-scan`:** -widen the regex to `[0-9]{3,}`; use design-decisions headers for D; exclude `next-free` -pointer lines; and exclude `tests/` from `all_text()` for identifier scanning. More code, -same answer -- prefer the removal. - -## Interim guidance (both streams, until fixed) - -Take D / DOCFIX / BUNDLEFIX next-free from `bash scripts/ledger-scan.sh` ONLY. Treat -repo_lint's `[info] L5 next-free` line as advisory-and-currently-wrong; it does not gate -(it is `[info]`, 0-fail), so it is safe to ignore for numbering. diff --git a/docs/script-quality-findings-20260707.md b/docs/script-quality-findings-20260707.md deleted file mode 100644 index bcbc347..0000000 --- a/docs/script-quality-findings-20260707.md +++ /dev/null @@ -1,141 +0,0 @@ -# Script-quality findings register -- read-only sweep, 2026-07-07 - -**What this is.** The deferred-findings register from the 2026-07-07 read-only -script-quality sweep, recorded so the items survive session compaction. Nothing -in this file has been executed; every fix listed here is a DOCFIX **candidate** -to be delivered later under the standard change-delivery loop (harness green, -repo-lint clean, changelog entry with revert). No identifier numbers are -assigned in this register -- the batch that fixes an item assigns its number -via the next-free rule at delivery time. - -**Sweep coverage.** 110 files reviewed (scripts/, tests/, .claude/hooks/). -Tools available on the review host: bash and python3 only; shellcheck and -pyflakes were absent, so findings come from manual read plus targeted fixture -probes, not static analysis. - -**Batch note.** The sweep's E/R-class items approved for immediate fix were -implemented in the fix batch of this date and are NOT listed here, except the -two previously-known items marked FIXED-THIS-BATCH below. - -## Deferred findings (R-class; fixes are DOCFIX candidates) - -### R5 -- lib-validate.sh: vr_json tempfile leak -`vr_json` (scripts/lib-validate.sh, ~line 52) creates `_VR_ERR="$(mktemp)"` on -every call and never removes it. `vr_err_tail` legitimately needs the file -after return, but nothing cleans it up afterwards: one leaked tempfile per -`vr_json` call, per check, per validation run. Fix shape: a single per-process -stderr file reused across calls, or an EXIT trap in the library's sourcing -contract. - -### R6 -- `| head -1` SIGPIPE / extract-then-check class in phase-05/06 -Live-pipe `openstack ... | head -1` captures that use the fragment without -whole-output validation: -- scripts/phase-05-amphora-pipeline.sh:62 (`image list --tag ... | head -1`) -- scripts/phase-05-amphora-pipeline.sh:66 (`image list --name ... | head -1`) -- scripts/phase-06-mgmt-vm.sh:106 (`port list --server ... | head -1`) -- scripts/phase-06-mgmt-vm.sh:108 (`floating ip list --port ... | head -1`) -Same class the house style forbids (SIGPIPE race on the producer; a partial -failure yields a plausible-looking fragment). Fix shape: capture whole output, -validate shape (uuid regex), then select the first row. - -### R7 -- phase-07-conductor-graft.sh: remote tempdir leak -The helm-install payload run on the conductor (~line 153) does -`D=$(mktemp -d); cd "$D"` and never removes the directory: one leaked tempdir -(plus tarball) on the remote unit per graft run. Fix shape: trap cleanup inside -the remote payload. - -### R9 -- cloud-assert.sh --capture: `--long` + silent empty images.json -Line ~184: `openstack image list --long -f json "$DIR/images.json" -2>/dev/null || true`. Two defects: (a) the deprecated `--long` flag only adds -stderr noise (and `2>/dev/null` then masks REAL errors, against the -stderr-separation rule); (b) a failed list silently commits an empty/missing -images.json into the captured BOM. Fix shape: drop `--long`, stderr to a -tempfile surfaced on failure, fail the capture loudly if the JSON is empty or -unparseable. - -### R11 -- harness tempfile leaks -Harnesses that mktemp without a trap cleanup (leak per run): -- tests/carve-host-interfaces/run-tests.sh -- tests/lib-validate/run-tests.sh (also amplifies R5: every vr_json case leaks) -- tests/phase-00-maas-standup/run-tests.sh -Fix shape: the standard `W=$(mktemp -d); trap 'rm -rf "$W"' EXIT` pattern the -other 34 harnesses already use. - -## Previously-known items -- FIXED-THIS-BATCH - -- **tenant-offboard.sh Phase B counter (FIXED-THIS-BATCH).** Logged in the - changelog (2026-07-06 addendum 28) as FOUND-NOT-FIXED: the residual-trust - sweep incremented SWEEP_FAIL inside a `printf | while` pipeline subshell, so - failed trust deletes could never raise exit 22. Fixed in this date's batch - (here-string loop; regression case in tests/tenant-offboard). -- **tenant-offboard.sh hardcoded auth URL (FIXED-THIS-BATCH).** The known - hardcoded-OS_AUTH_URL class (house rule 3): tenant_env now reads `auth_url=` - from the cred file with the literal demoted to a KEYSTONE_VIP-overridable - fallback. NOTE: the tenant-acceptance.sh foil_env sibling of this class - (addendum 28 FOUND-NOT-FIXED list) remains OPEN and is NOT covered by this - batch. - -## Consolidation proposals (S-class; S1-S8) - -Transferred verbatim from the sweep report at integration (2026-07-07); -effort estimates are the sweep's (S/M/L). Operator has NOT ruled on S1-S5/ -S7/S8 -- proposals only. - -- **S1 (M)** The unset-OS_*/source-env dance is hand-rolled ~15x across - tenant-onboard (8 subshells), tenant-offboard (2), tenant-acceptance (3), - d063-apply (1), tenant-assert (2) -- while scripts/lib-validate.sh already - ships tested vr_scrub_os / vr_admin_env / vr_tenant_env. Proposal: source - lib-validate.sh from the tenant-* family + d063-apply; add one - `vr_appcred_env ` for the app-cred variant. (The offboard tenant_env - auth-url bug was fixed narrowly in this batch; this proposal remains the - wider consolidation.) -- **S2 (S)** run() capture-then-test helper duplicated with eval-counter - variants in tenant-offboard.sh and d063-apply.sh vs the rc-returning - lib-validate run(). Consolidate; count failures at call sites. -- **S3 (S)** is_id() defined in 4 scripts (tenant-onboard, tenant-offboard, - d063-apply, tenant-assert); lib-validate already has vr_is_hex32. Ride - along with S1. -- **S4 (M)** The trust-list stdout-only capture + inline-python filter block - is near-identical 3x (tenant-offboard x2, tenant-acceptance x1) and is - inline python-in-bash beyond one-liner scale (house rule: .py files). One - fixture-tested scripts/trust_filter.py would serve all three. -- **S5 (S)** Kube-image UUID resolution block (json capture + jq kube filter - + tempfile + 36-hex validation) appears 3x inside tenant-onboard.sh. One - local function. -- **S6 (no action)** cap()/oerr() triplicated across clientdocs/scripts/ - {smoke-test,acceptance-run,ci-cleanup-sweep} -- CORRECT as-is: handover - artifacts must be standalone off-jumphost. Worth a header comment noting - the duplication is deliberate. -- **S7 (S)** Dead code: checks/d011-04 CF assigned never used; checks/d011-03 - vr_is_ipv4-or-VHOST first test subsumed by second. (The other two S7 items - -- offboard no-op is_id, validate.sh lost line-rewrite -- were FIXED in - this batch.) -- **S8 (S)** Move the KEYSTONE_VIP default and the D-003 FIP pool bounds into - lib-net.sh as tagged constants (the literal is duplicated in 4 files: - tenant-onboard, phase-03-admin-openrc, phase-06-kubeconfig-gate, - phase-06-k8s-bootstrap; FIP pool literals duplicated across the two - phase-04-network scripts; provider-bundle-check.py re-states plane CIDRs - as a second source of truth -- python cannot source bash, needs a - generated or parsed form). - -## Found during the fix batch (this date; logged, NOT fixed -- out of scope) - -- **tenant-offboard.sh Phase A cluster captures merge stderr.** The - tenant-scope `coe cluster list` captures (inventory ~line 160; delete loop - and wait loop in Phase A) still use `2>&1`: a stderr warning line would - word-split into `coe cluster delete` argv and keep the wait loop from ever - seeing empty (budget exit 21). Same class as the Phase C fix; was not in the - approved batch scope. The offboard harness deliberately excludes coe - commands from its stderr-noise scenario until this is fixed. -- **Merged-stderr `domain show`/`token issue` captures (fail-closed class).** - tenant-offboard.sh (~line 133, sweep-mode preamble) and tenant-onboard.sh - stage subshells validate `$(... 2>&1)` captures by shape/equality. Failure - direction is CLOSED (safe), but a benign stderr warning on a noisy CLI would - abort with a false precondition/auth failure. Cosmetic-risk cleanup - candidate, same stdout-only pattern. -- **tenant-offboard.sh Phase E0 app-cred inventory merges stderr.** Line ~267 - tests rc (failures route correctly), but a success-with-warning would - word-split the warning into the display-only cascade-inventory lines. - Display pollution only; no delete argv exposure. - -ASCII + LF. diff --git a/docs/security-ledger.md b/docs/security-ledger.md index 1aedfe8..4e19767 100644 --- a/docs/security-ledger.md +++ b/docs/security-ledger.md @@ -13,11 +13,11 @@ | SEC-002 | 2026-06-17 | juju action params persist in the operation log -- charm authorization must use short-lived child tokens | DOCFIX-011 | operator | STANDING RULE (verify each vault authorize) | | SEC-003 | 2026-07-03 | Vault unseal-key custody is single-operator (bus factor) | D-069 | operator | OPEN -- assign custodians + rehearse second-person unseal (DEFERRED 2026-07-06 by operator; note: keeps d011-06 MANUAL, gates D-011 full close) | | SEC-004 | 2026-05-27 | repo temporarily PUBLIC for v1 web_fetch workflow | project completion list | operator | OPEN -- flip to private at v1 close (deferral reaffirmed 2026-07-06) | -| SEC-005 | 2026-07-13 | long-lived GitBucket PAT stored in PLAINTEXT on the jumphost (`credential.helper store`, mode 600) so the agent can push unattended; GitBucket PATs are account-wide -- no per-repo scoping exists | docs/changelog-20260713-git-credential-store.md | operator | OPEN -- dedicated token (not a reused personal one); revoke + rotate at v1 close, or immediately if the jumphost is shared/rebuilt. Revoke: GitBucket -> Account Settings -> Applications | -| SEC-006 | 2026-07-12 | NetBox API token PASTED INTO AGENT CHAT CONTEXT during D-111 IPv6 reconciliation -- the token is BURNED (it exists in a transcript that is not under our control) and must be revoked + reissued, not merely rotated on schedule | docs/changelog-20260712-office1-opnsense-edge-build.md; flagged again at 2026-07-12 session close | operator | OPEN -- **DEFERRED by operator ruling 2026-07-13: revoke at completion of this deployment.** Until then the token is live and exposed. Revoke: NetBox -> Admin -> API Tokens | +| SEC-005 | 2026-07-13 | long-lived GitBucket PAT stored in PLAINTEXT on the jumphost (`credential.helper store`, mode 600) so the agent can push unattended; GitBucket PATs are account-wide -- no per-repo scoping exists | docs/archive/changelogs/changelog-20260713-git-credential-store.md | operator | OPEN -- dedicated token (not a reused personal one); revoke + rotate at v1 close, or immediately if the jumphost is shared/rebuilt. Revoke: GitBucket -> Account Settings -> Applications | +| SEC-006 | 2026-07-12 | NetBox API token PASTED INTO AGENT CHAT CONTEXT during D-111 IPv6 reconciliation -- the token is BURNED (it exists in a transcript that is not under our control) and must be revoked + reissued, not merely rotated on schedule | docs/archive/changelogs/changelog-20260712-office1-opnsense-edge-build.md; flagged again at 2026-07-12 session close | operator | OPEN -- **DEFERRED by operator ruling 2026-07-13: revoke at completion of this deployment.** Until then the token is live and exposed. Revoke: NetBox -> Admin -> API Tokens | | SEC-007 | 2026-07-12 | `~/vr1-office1-creds/` on the jumphost holds the Office1 edge root password + its bcrypt hash + the `office1_svc` SSH PRIVATE key + (added 2026-07-13) the OPNsense **API key/secret** (`opnsense-api.txt`, 0600), all in plaintext (dir 0700, files 0600) | 2026-07-12 edge build; custody detail deliberately not recorded here per D-069 | operator | OPEN -- required for edge management (D-112(c) makes SSH the only management path), so this is a ROTATION obligation, not a delete-me. Rotate at v1 close / if the jumphost is rebuilt or shared | | SEC-008 | 2026-07-13 | `~/vr1-office1-creds/tailscale-authkey.txt` (vcloud, 0600) -- the operator's 48-char Tailscale auth key for the SELF-HOSTED control plane `tailscale.baldurkeep.com`. Used to enrol `office1-tailscale`. Consumed BY PATH and never printed into a session; the copy shipped to the node was SHREDDED after `tailscale up` (tailscaled holds its own node key now). | 2026-07-13 Office1 Tailscale enrolment; docs/vr1-office1-as-built.md | operator | OPEN -- ROTATION obligation, not a delete-me: the key is still on vcloud for re-enrolment. Revoke/reissue on the control server at v1 close, or immediately if vcloud is rebuilt or shared. | -| SEC-009 | 2026-07-15 | Credential/env SPRAWL + world-readable exposure on vcloud: three env files sat loose in `~` outside the consolidated `~/vr1-office1-creds/` -- `.vr1-netbox.env` (upstream NetBox token), `vr1-office1.env` (edge root-password hash + SSH key path), and `vr1-stage1.env` which held `TF_VAR_maas_api_key` (a MAAS API secret) at mode **0664 (group/world-readable)**; `tailscale-authkey.txt` in the creds dir was also 0664. | docs/changelog-20260715-creds-consolidation.md | operator | **REMEDIATED 2026-07-15** -- all three moved into `~/vr1-office1-creds/`, every sensitive file `chmod 600`. STANDING CONVENTION established (below). Note: this is a CLOSED in-house test; the rotation obligations of SEC-005/006/007 still apply at v1 close, this row only closes the SPRAWL + PERMS exposure. | +| SEC-009 | 2026-07-15 | Credential/env SPRAWL + world-readable exposure on vcloud: three env files sat loose in `~` outside the consolidated `~/vr1-office1-creds/` -- `.vr1-netbox.env` (upstream NetBox token), `vr1-office1.env` (edge root-password hash + SSH key path), and `vr1-stage1.env` which held `TF_VAR_maas_api_key` (a MAAS API secret) at mode **0664 (group/world-readable)**; `tailscale-authkey.txt` in the creds dir was also 0664. | docs/archive/changelogs/changelog-20260715-creds-consolidation.md | operator | **REMEDIATED 2026-07-15** -- all three moved into `~/vr1-office1-creds/`, every sensitive file `chmod 600`. STANDING CONVENTION established (below). Note: this is a CLOSED in-house test; the rotation obligations of SEC-005/006/007 still apply at v1 close, this row only closes the SPRAWL + PERMS exposure. | | SEC-010 | 2026-07-16 | **metal-admin DC-LOCAL invariant (D-052/D-100) is PRESERVED but NOT ENFORCED in the committed Stage-3 config.** The `vvr1_dc0` rack straddles metal-admin (10.12.8.0/22, DC-local) + the office1<->dc0 transit (crosses fiber). Its committed cloud-init pins static IPs only -- **no `net.ipv4.ip_forward=0` sysctl, no host firewall on the transit leg.** Cross-plane routing (metal-admin <-> the whole Office1 /22, bidirectional) is blocked ONLY by Ubuntu's distro default; the deferred MAAS-rack install could silently flip it. A MAAS rack proxies at the application layer and needs no kernel forwarding, so pinning is free. **MODEL B UPDATE (2026-07-16):** after the D-123 Model B reshape, the OUTER `vvr1-dc0` is single-leg (transit only); metal-admin + the other 5 planes are now INNER bridges on `vvr1-dc0`'s own libvirtd. The forwarding hardening is MORE critical (vvr1-dc0 now bridges ALL 6 inner planes + the transit) and belongs on the inner libvirt host, not just a 2-leg rack -- wire it in the C3 bootstrap (`site-headend-install.sh` node-host mode). | 2026-07-16 plane-segregation review + Model B reshape cross-check; `opentofu/main.tf` `module "vvr1_dc0"` | operator | **OPEN -- artifact COMMITTED (Phase D 2026-07-16), gated on apply+verify.** Enforced via a FORWARD-drop across the transit leg, NOT a global `ip_forward=0`. `scripts/site-headend-install.sh --host-nodes` writes `/etc/nftables-sec010.nft` (drop FORWARD in+out the transit interface, default `mgmt`) + a boot-persistent `sec010-fw.service`; `--host-nodes --check` is the MECHANICAL pre-apply gate (fails if the rule is absent, AND -- hardened 2026-07-16 -- if the keyed transit interface does not EXIST, since an nftables oifname on an absent iface loads clean but matches nothing = fail-open). **D-125 UPDATE (2026-07-16):** the earlier rationale (a global `ip_forward=0` being unusable because the inner `vr1-dc0-wan` NAT forced `ip_forward=1`) NO LONGER applies -- under bridge-in the inner `vr1-dc0-wan` is a BRIDGE (no inner NAT). The scoped FORWARD-drop is RETAINED as-is (correct either way). **br_netfilter CONSTRAINT:** it MUST stay interface-scoped and NEVER be globalized -- with `br_netfilter` loaded, bridged WAN frames traverse the L3 FORWARD chain, so a global forward-drop would silently kill the `br-vr1-dc0-wan` WAN path. `--check` now also verifies the WAN bridge exists with its uplink enslaved (same fail-open class). Nothing routes across the fiber THROUGH vvr1-dc0; the rack proxies MAAS at the app layer (originated/terminated, not forwarded). Same pin follows onto `voffice1` when its transit leg is wired. Region route must target only the rack transit /30, never 10.12.8.0/22. Stays OPEN until applied + verified on vvr1-dc0. | | SEC-011 | 2026-07-16 | **Node least-connectivity gap (not an L2 breach).** Under D-121 Option C role separation, all nodes get a uniform 6-plane NIC set, so a ceph-osd STORAGE node has a leg on provider-public (external/FIP) + data-tenant (tenant geneve) -- planes it never binds per D-052. Planes stay isolated L2 (no crosstalk). | 2026-07-16 plane-segregation review; `opentofu/main.tf` `local.vr1_dc0_node_nics` | operator | **CLOSED 2026-07-16 (operator ruling -- keep uniform 6-NIC).** Review R3-F10: A2's cross-examination refuted the attack-surface concern -- in the isolated-L2 sim the unbound vNICs have no reachability out, and pruning would INCREASE Roosevelt-delta (baremetal trunks all VLANs to every node on bonded NICs, so all planes are present regardless of L3 binding). Uniform 6-NIC is the more Roosevelt-faithful model. Accepted non-issue; no code change. | @@ -40,11 +40,11 @@ + the assembled `nbt_.` token), copied off the VM without being printed into the session. The `/root/netbox-secrets/` copy remains the source-of-record on the VM; this is a working copy on the jumphost, mirroring how `vr1-netbox.env` holds the upstream token. See -`docs/changelog-20260715-fidelity-check-scope-correctness.md`. +`docs/archive/changelogs/changelog-20260715-fidelity-check-scope-correctness.md`. **ENFORCEMENT (2026-07-15) -- so the miss cannot recur, esp. for DC0/DC1.** The convention was never ENFORCED; a VM-minted secret had no forcing function to get consolidated and no check to catch a miss. -Now built (`docs/changelog-20260715-creds-audit.md`): +Now built (`docs/archive/changelogs/changelog-20260715-creds-audit.md`): - **`creds-manifests/<site>.manifest`** -- declares EXACTLY the secrets a site must hold, each with its mode and PROVENANCE (`local`, or `<vm>:<path>` for VM-minted secrets -- the class that was missed). `creds-manifests/vr1-office1.manifest` seeded (11 entries) and the live folder audits CLEAN. diff --git a/docs/session-findings-2026-07-02.md b/docs/session-findings-2026-07-02.md deleted file mode 100644 index 2805d16..0000000 --- a/docs/session-findings-2026-07-02.md +++ /dev/null @@ -1,77 +0,0 @@ -# Session findings -- 2026-07-02 (multi-tenant tenant->cluster buildout) - -## Executive summary -The tenant IDENTITY/TRUST path is DONE and PROVEN. A tenant password identity now creates a Magnum -cluster through create_user (D-064), create_trust (D-065), and into certificate generation. Cluster -COMPLETION is blocked one step later by an OPERATOR-side Barbican/Vault substrate defect (D-067), -independent of the tenant model. The Option-3 tenant account model (D-066) is adopted and to be used -from the first tenant. Next session: fix D-067 (live), then full tenant buildout + tenant-facing tests. - -## The trust-blocker chain (how we got from "cluster 403s" to "done + one substrate bug") -1. D-064 (prior): create_user template fix unblocked trustee-user creation. -2. create_trust then 403'd for EVERY caller (admin included), even trustor==self via direct - `openstack trust create`. Root cause: base policy shipped identity:create_trust with the - non-resolving `user_id:%(trust.trustor_user_id)s` (Caracal populates target.trust.trustor_user_id). - -> D-065: override with the target-prefixed form keystone itself ships. PROVEN by toggling the - override off (still 403 -> base policy owns it) then on. -3. After D-065, create_trust via APP CRED still failed: keystone `_check_application_credential` - blocks trust creation from app-cred tokens "regardless of the unrestricted flag" (this build's - docstring). -> D-066: cluster-create MUST be PASSWORD auth; adopt Option-3 account split. - `allow_insecure_application_credential_trust_escalation` REJECTED (isolation). -4. Password create_trust PASSED. Cluster then failed at cert-gen -> Barbican 500 -> castellan - vault_key_manager -> Vault AppRole login rejected: "source address 10.12.8.176 unauthorized - through CIDR restrictions". -> D-067. - -## D-067 root cause (and a corrected mis-diagnosis) -barbican reaches Vault on the METAL-ADMIN plane (vault_url=10.12.8.190, egress 10.12.8.176). Vault's -barbican-vault AppRole binds the secret_id to the METAL-INTERNAL CIDR (where east-west service traffic -belongs, D-052/D-053). Off-plane source -> rejected. The bundle is CORRECT (vault/barbican/barbican-vault -all bind secrets endpoints to metal-internal, lines 130/667/700); the LIVE env drifted. Fix = live -rebind to metal-internal (gated, next session), NOT CIDR-widen. -CORRECTED: mid-session I hypothesized "secret_id TTL expiry". REFUTED -- `juju run vault/leader -refresh-secrets` rotated the secret_id (barbican.conf re-rendered, service restarted) and the login -STILL failed with the CIDR error. It was never expiry; it is plane/CIDR. - -## What is validated live (tenant acme) -- Manager persona self-service via CLI (create_project/user/grant) -- D-064 G3. PASS. -- Tenant isolation: anti-escalation (admin grant DENIED); cross-domain resource reads DENIED/hidden; - domain enumeration OWN-DOMAIN-ONLY (tighter than appendix-C's SCS worst-case -- appendix-C corrected). -- App-cred + keypair self-mint; tenant L3 (net/subnet/router/ext-gw, SNAT proven) by a non-admin - app-cred identity. -- Cluster template create (image by UUID -- name form has a quoting/derivation hazard). -- Cluster create through create_user + create_trust (password) into cert-gen. - -## Decisions logged -- D-066: Option-3 tenant accounts (domain-admin/cluster/svc); cluster-create requires password auth. -- D-067: barbican-vault -> Vault must use metal-internal (live drift; the cert-gen blocker). ADOPTED, fix pending. -- D-068: PROPOSED -- Vault substrate hardening (1.16 pin [bundle done], TLS, AppRole lifecycle). - -## Probe-discipline lessons (now runbook conventions -- these recurred and cost time) -1. Validate raw output WHOLE, never extract-then-check. A `tr -dc 0-9` MARK guard turned an error - string ("...10.12.8.30:17070...") into MARK=123101283017070 and passed. Use `case "$raw" in - ''|*[!0-9]*) fail;; *) ok;; esac`. -2. Whitelist-print secrets, never blacklist-redact. `approle_secret_id` leaked past a `secret`-keyed - redact (the key is *_secret_id*). Print only an allowlist of safe fields; never pipe secrets. -3. No `exit`/bare-`return` in interactive PASTE blocks (they escape to the login shell and logged the - operator out). Subshell-wrap `( ... )`. NOTE: executed .sh scripts may use exit normally. -4. Privileged reads over `juju ssh` use `sudo cat file | ...`, never `sudo cmd < file` (the redirect - runs UNPRIVILEGED -> Permission denied). -5. Use the deployment's DECLARED endpoint/scheme, not the conventional one (assumed Vault https; it - serves http -- every probe errored on scheme until corrected). -6. A parser that can print NOTHING has a silent third state -- read raw + self-report inputs (field - lengths, raw body) so a malformed-request 400 can't masquerade as an auth failure. - -## Roosevelt hardening backlog (from this session) -- D-067/D-068: metal-internal binding discipline for ALL vault-kv consumers; Vault 1.16 + TLS; - AppRole secret_id lifecycle (TTL/renewal + proactive auth health probe). -- Endpoint/credential "follow the topology" is now a recurring class (with D-057): consider stable - VIP/DNS endpoints for substrate services so leader/re-IP changes don't silently break consumers. - -## Next session plan -1. Repair live env: read-only binding diagnosis (`juju show-application vault barbican barbican-vault`, - spaces<->subnets), then GATED `juju bind` of the barbican<->Vault secrets path to metal-internal; - re-run refresh-secrets if needed; confirm barbican AppRole login HTTP 200 from metal-internal. -2. Re-run tenant cluster-create (acme, ${CLIENT}-cluster password) -> cert-gen clears -> watch to - CREATE_COMPLETE; capture the CAPO child-cred mint identity (confirms D-066). -3. Full tenant buildout via scripts/tenant-onboard.sh; then clean-room `beta` (zero admin fallback). -4. Tenant-facing tests: kubeconfig, nodes/CNI/CCM, a tenant LB, tenant isolation from a second tenant. diff --git a/docs/session-ledger.md b/docs/session-ledger.md index e00a803..e968e17 100644 --- a/docs/session-ledger.md +++ b/docs/session-ledger.md @@ -129,3 +129,8 @@ added; first full G1 review ran (rules 1-4): one further violation found + corrected + registered (GA-F16, creds pointer reduction). Pre-edit copies: docs/archive/memory-20260718/. Zero violations remain (3 files). Next gate: operator opens Batch 4 (consolidation + rotation -- the big one). + +- BATCH 4 COMPLETE (2026-07-19, same session): ledger rotated (this file, 131 lines -- F1 cap live); + 96 changelogs + 24 history docs consolidated to docs/archive/ (4 stage records, manifest commits + C2-C5); top-level docs/ = 16 (< 25); live refs rewritten. Verification in the close commit. +- Next gate: operator opens Batch 5 (skill sweep) then Batch 6 (exit runs). diff --git a/docs/stage3-adversarial-review-20260716.md b/docs/stage3-adversarial-review-20260716.md deleted file mode 100644 index f136d6a..0000000 --- a/docs/stage3-adversarial-review-20260716.md +++ /dev/null @@ -1,348 +0,0 @@ -# Stage-3 / vr1-dc0 pre-deployment adversarial review -- 2026-07-16 - -Adjudicated findings from a four-charter adversarial review of the Stage-3 DC-substrate -batch, run before `vr1-dc0` is deployed for the first time. READ-ONLY: nothing was applied, -mutated, or pushed. Every proposed fix is a PROPOSAL for the operator to gate individually; -none were executed. Fixes are staged for a post-acceptance sweep, not applied mid-review. - -Adjudicator: Code (main loop). Method: Layer-1 deterministic gate (facts) -> Code's own -grounding of the decision-coherence register from primary text + the session transcript -> -four independent charters (A1 author's-advocate, A2 prosecutor, A3 Roosevelt-hawk, A4 -drift-archaeologist) in Round-1-independent + Round-2-cross-examination -> Code adjudicates. -I am NOT a consensus engine: surviving dissent is recorded (section 4), not resolved away. - ---- - -## 1. Review surface - -- Repo: `/home/jessea123/openstack-caracal-dc-dc` (live jumphost clone). Branch - `dc-dc-stage3-phase2-dc-substrate`, frozen at HEAD **`87a7a8a`** (a WIP commit that froze the - two uncommitted plane-review ledger files on top of the batch commit **`a48a60f`**). -- Baseline `@{upstream}` = **`80502e9`**. Review range **`80502e9..87a7a8a`**. -- Immutable artifact: **`docs/stage3-review-base.patch`** (3795 lines, 26 files, +3057/-275). - Every finding cites against the repo files (line numbers as they stand at HEAD). -- IN scope: the D-121..D-124 rulings, the OpenTofu Stage-3 substrate (`modules/site-wan`, - `main.tf` vr1_dc0 section, the 8 node VMs, the `vvr1-dc0` rack, `variables.tf`), the two - NetBox importers + harnesses, `overlays/dc-ha-scaleup.yaml`, `scripts/site-headend-install.sh` - rack mode, the phase2 runbook rewrite, and the three changelogs. -- OUT of scope (operator-gated, correctly deferred): the live `tofu apply`, the NetBox - `--commit`, the rack install, and the Stage-4 `maas-vm-host` wiring (DOCFIX-179). -- **Undocumented-intent note (per the prompt's FIRST ACTION item 4):** one design input exists - as a claim but not as a file -- `scratchpad/optc-calc.py`, the whole-host capacity model that - D-121's Option-C sizing rests on, is NOT in the tree (see R3-F06). No other design element was - found to live only in prior-session discussion. -- **A4 charter note:** the A4 (drift-archaeologist) Round-1 pass died on an API stall mid-stream - and was re-run as a standalone backfill; because Round-2 had already executed, A4 did not - participate in cross-examination. Its lens (drift / status-table archaeology) was substantially - covered by Code's own decision-coherence grounding and by A1. [A4 backfill: INTEGRATION PENDING - -- see section 4a; this document is complete without it and will be annotated on its return.] - ---- - -## 2. Layer-1 facts (mechanically proven; agents debate meaning, not these) - -*Verbatim tool output for every row below is captured under -`scratchpad/L1/` (ledger-scan.txt, repo-lint.txt, gauntlet.txt, tofu-validate.txt, tofu-plan.txt, -ceph-optc-500.txt, state-addresses.txt, etc.); this table is the distilled register.* - -A2's external-authority citations (Ceph size=3-wants->=4-hosts guidance; MAAS rack-statelessness; -CIS-style ip_forward hardening) are reproduced as the charter reported them; the underlying technical -claims hold independently, but specific benchmark/control identifiers should be re-verified before -being quoted as authoritative. - -| id | fact | -|---|---| -| L1-01 | diffstat: 26 files, +3057/-275 vs `80502e9`. | -| L1-02 | `ledger-scan`: PROPOSED/OPEN = **D-068, D-071, D-115**. next-free D=125. Fences OK. SEC-010/011 OPEN. | -| L1-03 | `repo-lint`: 0 fail, 1 WARN (legacy non-ASCII in design-decisions.md D-001..018 only). | -| L1-04 | Byte hygiene: the reviewed batch added **0** non-ASCII and **0** CR bytes. Patch has 0 added non-ASCII lines. | -| L1-05 | Gauntlet: **ALL GREEN, 62 harnesses** (the prompt cited 60 -- a stale number, still green). | -| L1-06 | `tofu validate`: Success (OpenTofu v1.12.3, 11/11 modules). | -| L1-07 | `tofu plan`: **BLOCKED** -- 4 required vars unset/no-default (`vr1_dc0_rack_metal_admin_ip`, `_transit_ip`, `_transit_prefix`, `_transit_peer_ip`). Plan cannot render until office1-netbox assigns the rack transit/30 + rack IP. Hard pre-apply gate. | -| L1-08 | state<->repo: `terraform.tfstate` = 15 resources, all Office1/Stage-1/2; every state module still declared (no orphan); the Stage-3 substrate is repo-only = expected "to create". Nothing applied. | -| L1-09 | Ceph re-run for the RULED Option C (3 OSD hosts/DC @ 500Gi): **PASS, margin 6.48 TiB** (roomier than D-121's carried Option-B 5.31 TiB). D-121 validated Option B, not C. | -| L1-10 | Phase-5 drill Step 10.0(b) = `virsh shutdown <domain>` per node (domain GROUP), not a single `virsh destroy`; Step 10 failover invokes **NO MAAS** (Ceph/Glance/Cinder/Neutron workload failover only). | -| L1-11 | Section-9 shim register = node-VM create (D-103), tc netem (D-100), single-unit **Juju** controllers (D-104), single-unit rbd-mirror (D-108). D-121 scales the OpenStack plane and excludes D-104 + keeps rbd-mirror at 1 (D-108). | -| L1-12 | `scratchpad/optc-calc.py` (D-121's whole-host validation basis) is **NOT tracked in git and ABSENT from disk**; the changelog claims "reproducible: python3 scratchpad/optc-calc.py" (false). | -| L1-13 | Transcript verification (full question text + answers read): the only AskUserQuestion in scope had two questions -- **"Node layout"** (operator answered **"Option C (3+2+3)"** -- an explicit design ruling) and **"Proceed"** (a DELIVERY-workflow question; options were "rewrite the runbook / bundle changelog+ledger / repo-lint / hand over"; operator answered **"Do it all now..."**). Neither vault backend (v-a/v-b) nor node containment (Model A/B) was ever a presented option; the operator's messages contain no vault/backend/unseal utterance anywhere. | - ---- - -## 3. Findings register (ranked; severity, evidence, violation, action, contested?) - -### R3-F01 -- BLOCKING -- Two ADOPTED rulings rest on inferred/agent-authored operator rulings (count = 2) - -The batch's single highest-value defect. An operator ruling is a VALUE; inferring it violates -hard-rule-2, and here two of the four decisions carry an inferred/agent-authored ADOPTED status. -**The count is 2, not 1** -- the four charters converged on 1 (D-123 only) because Code's shared -brief pre-asserted "D-121 is grounded via AskUserQuestion"; that biased them past the vault -sub-ruling. Code verified the second case directly against the transcript (L1-13). - -- **(a) D-121 vault-HA backend v-a -- the code-consequential case.** Status line - `design-decisions.md:3264`: "Vault-HA backend sub-ruling: RESOLVED = (v-a) ... operator ruling". - Its OWN body `:3346` says "Operator sub-ruling needed." Transcript (L1-13, full question text - read): vault backend was NEVER a presented choice -- the only AskUserQuestion asked "Node layout" - (answered Option C) and "Proceed" (a delivery-workflow question answered "Do it all now"); neither - is a vault ruling, and the operator uttered nothing about vault anywhere. So v-a is agent-authored, - with no operator value. Worse, it is ENCODED in committed code: - `overlays/dc-ha-scaleup.yaml:67-80` re-declares `vault-hacluster` and re-adds - `[vault:ha, vault-hacluster:ha]`, REVERSING BUNDLEFIX-002 (which de-HA'd vault) -- a real change - driven by an unratified ruling, not a status-quo path. -- **(b) D-123 Model A -- the record case.** Status `design-decisions.md:3454`: "ADOPTED Model A - ... read from 'Yes, fire off those tasks' in response to the A/B question; flag if B was - intended." The real utterance "Do it all now..." was the operator's answer to the AskUserQuestion - "Proceed" question -- a DELIVERY-workflow choice (L1-13), NOT an A/B answer; Model A vs B was never - a presented option. So the "A/B question" D-123 claims to read is one the operator was never asked. Corroborating inconsistency: the phase-2 runbook already labels D-123 **PROPOSED** - (`runbooks/dc-dc-phase2-tofu-dc-substrate.md:507` "D-123 (PROPOSED, recommend...)") while - design-decisions.md marks it ADOPTED -- the record disagrees with itself about D-123's ratification. - Refinement (A1, upheld in Round 2): Model A happens to equal the already-as-built - D-103/D-114 node-placement seam, so -- unlike vault -- **no NEW code rests on the Model-A - inference** (Model B would have been the change). The region+rack MAAS model and the D-124 rack - block (`main.tf:376-439`, "# D-124", explicitly ADOPTED) are independently authorized. -- Legitimately ADOPTED on explicit authority: D-121 node layout = Option C (AskUserQuestion - selection, `:3397`) + the 14-service scale-up ("this is the time we add the additional HA - nodes"); D-122 (the operator's explicit "Ruling 2" + "Routed by fabric switches"); D-124 - ("1. A 2. Confirmed", `:3513`). -- **Violates:** hard-rule-2 (no inferred value); hard-rule-1 (an inferred ruling drove the committed - vault overlay). **Action:** demote BOTH the D-121 vault-HA axis and the D-123 Model-A axis to - PROPOSED; put both one-line questions back to the operator ("vault HA backend: v-a MySQL-backed - vs v-b etcd/Raft?" and "node containment: Model A vcloud-level group vs Model B nested-in-VM?"). - D-121's "ADOPTED IN PART" then means what it says. Region+rack, Option C, and D-124 stand. -- **Contested (severity):** A1 argues MAJOR-not-BLOCKING for D-123 -- a status LABEL is not a - value "entering a command", nothing is applied (L1-08), and apply is independently gated - (L1-07). Adjudication: kept BLOCKING because (i) the prompt frames C2 as BLOCKING by design, - (ii) the vault case DID drive committed code, and (iii) "ADOPTED" is the authority a downstream - agent acts on. The dissent is real and recorded (section 4): as a matter of LIVE mutation risk, - nothing is currently blocked; as a matter of decision-record integrity, both must be re-ruled - before the record is trusted or apply proceeds. - -### R3-F02 -- MAJOR -- SEC-010: the metal-admin DC-LOCAL invariant is a ledger promise, not a committed artifact or a gate - -Upheld under cross-examination (A3-F6 survived A2's "distro default is fine" refutation attempt). -`main.tf` `vvr1_dc0` cloud-init pins static IPs + one route but has NO `net.ipv4.ip_forward=0` -sysctl and NO host firewall on the transit leg. The rack straddles metal-admin (10.12.8.0/22, -DC-local per D-052 `:767` / D-100 `:1946`) and the office1<->dc0 transit (crosses fiber), so the -"never crosses the fiber" invariant is preserved ONLY by Ubuntu's distro default -- which the -deferred MAAS-rack snap install could flip. SEC-010 records this "Close BEFORE tofu apply", but a -grep of `scripts/` and the phase2 runbook finds NO mechanical gate; the only apply-block (L1-07) -is an addressing gate, not security. **Violates:** D-052/D-100 (and general host-hardening guidance --- CIS-style benchmarks recommend `ip_forward=0` on non-router multi-homed hosts; verify the exact -control ID before quoting it). -**Action:** add `ip_forward=0` (v4+v6) + an nft/ufw transit-leg pin to the committed rack -cloud-init as the ARTIFACT; wire a mechanical pre-apply gate (not a ledger note); carry the same -pin onto `voffice1` when its transit leg is wired (currently single-homed, `main.tf:184`). The pin -is free -- a MAAS rack proxies at the application layer and needs no kernel forwarding. - -### R3-F03 -- MAJOR -- D-107's DR-independence claim is false as written (and false at Roosevelt too) - -D-107 `:2088`: the per-DC mirror exists "so a DC can redeploy independently even if Office1 or the -peer DC is down -- a DR requirement the drill exercises." Under the RULED region-on-Office1 + -rack-per-DC model (D-123), a MAAS rack is stateless and cannot commission/deploy/power without the -region, so an Office1 outage removes reprovisioning from BOTH DCs. And L1-10: the Phase-5 drill -invokes NO MAAS -- it exercises workload failover, not redeploy. Two clauses fail. -- **Nuance (A2, Round 2):** the claim says "Office1 OR the peer DC is down." The peer-DC-down half - is VALID (region up, mirror up -> the DC can redeploy). Scope the defect to the **Office1-outage** - sub-case, not the whole claim. -- **Elevation (A2, Round 2):** buildout-design `:157` mandates single-region PERMANENTLY, so this - is unachievable at ROOSEVELT too -- not stale VR1 wording. A standby-region-per-DC mitigation - would CONTRADICT the buildout's single-region decision. -- **Violates:** D-107 (its DR-independence + "drill exercises" clauses). **Action (option a):** - amend D-107 to scope DR-independence to the in-DC artifact mirror + workload failover; add a - WRITTEN, PERMANENT acceptance that node provisioning is unavailable during an Office1-region - outage (day-1/2 dependency, not runtime -- running DC clouds are unaffected); delete/correct "a - DR requirement the drill exercises". Do NOT reopen the topology. - -### R3-F04 -- MAJOR -- D-122's MAAS-controller bullet and single-object site-down wording were never superseded - -- **C1:** D-122 `:3441` "Each site runs its own MAAS controller (as voffice1 does)" -- a full - region+rack -- was superseded by D-123's region-on-Office1 + rack-per-DC (operator-confirmed; - matches buildout `:157`) but never marked. A downstream agent quoted the stale bullet back as if - live. **Action:** append a verbatim supersession note pointing to D-123. -- **C3:** D-122 `:3417` "virsh destroy <site-vm> against a single object" is void for DCs -- under - Model A the containment VM does not contain the DC nodes (D-123 admits this `:3482`), and the - Phase-5 drill already uses per-node group shutdown (L1-10), so the drill does NOT break. The - stale claim also leaked into `dc-dc-deployment-workflow.md:86`. **Action:** amend D-122 AND - workflow:86 to "destroy the vr1-dc0-* domain GROUP" for DCs (single-object destroy stays literal - for Office1); note `vvr1-dc0` at a DC is a MAAS rack headend, not a D-114 containment VM (rack - mode runs no LXD/compose). -- **Violates:** D-123 (supersedes both). Topology operator-confirmed; record fix only. - -### R3-F05 -- MAJOR -- D-107 vs D-123 "node artifacts" wording (C4): reconcilable, but D-107 must be scoped - -D-107 `:2087-89` ("no node artifacts served from Office1"; "images including amphora ONLY from an -in-DC mirror") reads against D-123 `:3459` (rack "PROXIES OS images from the Office1 region"). The -counter-reading HOLDS: the intra-MAAS region->rack commission/deploy image channel is architecturally -distinct from D-107's supply-chain mirror (apt/snap/charmhub/registry/amphora served at/after juju -deploy). Both stand once D-107 is scoped. **Violates:** NONE (a wording gap in D-107). **Action:** -amend D-107 to scope "node artifacts" to the supply-chain classes, explicitly excluding the -intra-MAAS provisioning-image proxy D-123 routes region->rack. Pairs with R3-F03. - -### R3-F06 -- MAJOR -- D-121's Option-C sizing rests on an uncommitted, now-absent model; the "reproducible" claim is false - -`scratchpad/optc-calc.py` -- the whole-host validation that grounds Option C's 222 vCPU / 790 GiB / -77%-RAM ruling -- is NOT tracked and is ABSENT from disk (L1-12), yet the changelog claims -"Resource model reproducible: python3 scratchpad/optc-calc.py". **Violates:** repo-authoritative -discipline (CLAUDE.md); D-121 record integrity. **Contested (severity, A1 Round 2):** the Option-C -ruling is an explicit operator selection (not gated on the calculator), and fit is independently -checkable -- node literals are committed in `main.tf` locals and the disk dimension re-ran from the -COMMITTED `dc-dc-ceph-disk-budget.sh` (L1-09 PASS). So the ruling is not "unverifiable"; the defect -is reproducibility hygiene. **Action:** promote `optc-calc.py` into a committed, harnessed -calculator (already logged as a follow-up) and re-run for Option C before the sizing is treated as -measured; correct the changelog's "reproducible" claim until it is. - -### R3-F07 -- MAJOR -- D-121 records an Option-B disk validation for the ruled Option-C layout - -D-121 `:3382` records the disk-budget PASS for "Option B 4+4" while the RULED layout is Option C -(3 storage/DC). Code re-ran for Option C (L1-09): PASS, 6.48 TiB margin -- capacity is fine, so this -is a RECORD defect, not a capacity one. **Action:** re-record the Option-C validation (cite L1-09). - -### R3-F08 -- MINOR -- Option C's 3 OSD hosts at size=3 carry zero Ceph rebuild headroom (unrecorded) - -D-121 flagged "size=3 has ZERO rebuild headroom" for the REJECTED Option A `:3358` but not for the -ADOPTED Option C, whose 3 storage hosts have the identical property. **Walked back from MAJOR in -Round 2** (A1/A2/A3 concur): with Charmed default min_size=2 the pool keeps SERVING on one host -loss (the availability drill is clean), ceph-mon=3 sits on the CONTROL nodes (quorum untouched), -and only the RE-REPLICATION sub-case lacks headroom -- an inherent size=3 economy that resolves at -Roosevelt (>=4 storage hosts). **Action:** add an accepted-risk note to the Option-C record (a -storage-node-loss drill will show degraded-not-self-healing; Roosevelt remedy = >=4 storage/DC). - -### R3-F09 -- MINOR -- D-115 reads PROPOSED to `ledger-scan` though it is ADOPTED by amendment (L1) - -D-115's primary Status line retains the substring "Originally PROPOSED/OPEN", which trips -`ledger-scan`'s regex (L1-02), while the decision IS ratified by its 2026-07-13 amendment -(`:2727`). So the office carve import + PR #1 merge rest on a legitimately ADOPTED decision -- NOT -an unratified-executed one. **Action:** move "Originally PROPOSED/OPEN" out of the Status line (or -harden `ledger-scan` to ignore an "originally/was PROPOSED" clause when ADOPTED is present). Sweep -D-121..D-124 Status lines for the same machine-vs-human record divergence. - -### R3-F10 -- MINOR -- SEC-011 (node least-connectivity) is CONTESTED on Roosevelt-fidelity grounds - -Code's own prior SEC-011 (storage nodes get provider-public + data-tenant legs they never bind) is -challenged by A2 (Round 2, REFUTED) and A1 (WEAKENED): (1) in VR1 every plane is isolated L2 with -no gateway/route out, so an unbound vNIC has NO exploitable reachability -- "attack surface" is nil -today; (2) pruning would OVERRIDE the ADOPTED D-122 6-NIC ruling without a superseding decision AND -INCREASE delta to Roosevelt, where D-052's planes are VLAN-trunked on bonded NICs (ALL planes -present at every node; netplan/MAAS decides L3 binding), so uniform 6 isolated-L2 vNICs is the more -faithful model. **Adjudication:** downgrade SEC-011 from "pre-apply hardening" to an OBSERVATION / -operator's call; recommend amending the SEC-011 ledger row to record the Roosevelt-delta caveat. -Surviving dissent (section 4). - -### R3-F11 -- MINOR -- Stale "DC1/DC2 + supernet unassigned" prose in the workflow doc (L3) - -`dc-dc-deployment-workflow.md` Authoring-status cell still reads "DC1-first; DC2 hard-gated (D-101 -supernet unassigned)" -- superseded by D-119 (code is vr1-dc0/vr1-dc1) and D-115 (vr1-dc1 supernet -= 10.12.64.0/19 assigned). No gate depends on the stale reading (the second DC is sequenced out -regardless). **Action:** update the cell; drop the "unassigned" clause. - -### R3-F12 -- OBSERVATION -- Transient Octavia N+1 amphora failover headroom is not modeled (C7b) - -The whole-host validation is a steady-state sum; it does not visibly reserve the transient N+1 -amphora placement headroom Octavia STANDALONE failover needs (a hard-won VR0 finding), and the -Phase-5 drill exercises failover on two clouds sharing one host. Mitigated in this single-host sim -(a hard-downed DC frees its RAM for the survivor) and unverifiable until R3-F06's model is committed. -**Action:** when `optc-calc.py` lands, add a line on per-cloud transient amphora headroom (or state -it is a Roosevelt-only concern for this sim). - -### R3-F13 -- OBSERVATION -- D-122 "6 NICs = baremetal-matched" is imprecise; Section-9 could add a cross-ref - -(i) D-122 sizes nodes at "6 NICs, one per plane" but the Roosevelt realization (D-052) is untagged -metal-admin + tagged VLAN subinterfaces trunked on bonded NICs, not 6 discrete physical NICs -- L2 -isolation is behaviorally equivalent; a characterization nuance, no code change. (ii) A3 argued -Section 9 omits the site-wan simulated-ISP and the virsh-destroy DR fault-injection; largely REFUTED -in Round 2 (both have Roosevelt analogs -- a real circuit, a real facility-down drill -- so neither -meets the register's "no analog" bar; the register is a build-step register and virsh-destroy is an -operational step). At most add a clarifying cross-reference for the containment/virsh-destroy sim -vehicle. Not a material incompleteness. - ---- - -## 4. Minority report (surviving dissent -- recorded, not resolved away) - -- **R3-F01 severity (dissenter: A1).** A1 holds D-123's inferred-status is MAJOR ledger-hygiene, not - deploy-BLOCKING: a status label is not a value "entering a command" (hard-rule-2's literal scope), - nothing is applied (L1-08), and apply is independently gated (L1-07). A1 further showed (Round 2, - refuting A3-F1) that Model A = the already-as-built D-103/D-114 path, so no NEW code rests on that - specific inference. Code's adjudication keeps R3-F01 BLOCKING on decision-record + governance - grounds (and because the VAULT half DID drive committed code), but records A1's point: as pure - live-mutation risk, nothing is currently blocked. -- **R3-F10 SEC-011 (dissenter: A2).** A2 refutes the "attack surface" framing outright: isolated-L2 - planes have no exploitable reachability, and pruning increases Roosevelt delta. Code accepts the - Roosevelt-delta point and downgrades SEC-011 accordingly, but records that A2 would go further and - strike the finding as a non-issue; Code retains it as an OBSERVATION worth the operator's note. -- **R3-F08 Ceph severity (dissenters: A1, A2, A3 concur down; A2-R1 held MAJOR).** A2's Round-1 - MAJOR ("opposite of the clean drill C is sold on") did not survive its own and others' Round-2 - scrutiny (min_size=2 keeps serving; mon on control nodes). Recorded as the walk-back it was; - final severity MINOR. - -**Convergence check (mandated).** The three live charters converged on the C-register verdicts -- -expected, because those items are text-provable quotations, not judgment calls. But they did NOT -fully converge: each surfaced a distinct high-value finding (A2: external CIS/Ceph/MAAS citations; -A3: the Section-9 completeness challenge + the baremetal NIC-realization nuance; A1: the uncommitted -`optc-calc.py`), and Round-2 produced genuine dissent (above). So the charters separated adequately. -The one place convergence WAS a failure mode: all three reported inferred-ruling **count = 1** -because Code's shared brief pre-asserted D-121's grounding -- a demonstration that shared priors -create shared blind spots. Code's independent transcript check (L1-13) corrected the count to 2. - -## 4a. A4 backfill integration (returned; corroborates) - -The A4 drift-archaeologist backfill returned and **independently reached inferred-ruling count = 2** -(D-121 vault-`(v-a)` + D-123 Model A), corroborating R3-F01 by the harder path: A4 was tasked to -scrutinize the vault Status-vs-body contradiction directly, and reached count = 2 WITHOUT the -biasing "D-121 is grounded" hint the three Round-1 charters received. This confirms the section-4 -convergence diagnosis -- the count-1 result was a shared-prior artifact, not a real ceiling. A4 did -NOT participate in Round-2 cross-examination (its Round-1 died on an API stall; Round-2 had already -run). Integration -- with one A4 claim verified-and-rejected (I check agent citations, not -rubber-stamp them): -- **A4's count = 2 corroborated.** A4's F1 (D-123 Model A) + F2 (D-121 vault, Status `:3264` vs - body `:3346`, no vault utterance in the batch changelog) independently match R3-F01. -- **A4's F7 REJECTED on direct check.** A4 claimed the phase-2 exit gate cites the stale D-122 - "each site runs its own controller" bullet (`runbooks/dc-dc-phase2-tofu-dc-substrate.md:636`). - Verified: the runbook does NOT -- at `:507` it cites "D-123 (PROPOSED, recommend...)" and the exit - gate (`:630-635`) uses the rack-controller-per-DC (D-123) model. A4 miscited. NOTABLE side effect: - the runbook already labels D-123 **PROPOSED**, which CONTRADICTS design-decisions.md's ADOPTED - status and independently supports R3-F01 (the D-123 record is internally inconsistent about its own - ratification). Added to R3-F01's evidence. -- A4 otherwise agrees across C1-C7 and L1-L5 (it frames C4 as "REFUTED -- both stand", the same - disposition as R3-F05's "reconcilable, scope D-107"), and adds clean negative findings: no false - DONE/as-executed claims, state reconciles with zero orphans, byte/lint hygiene clean. - ---- - -## 5. Leads register - -| Lead | Status | Evidence | -|---|---|---| -| **L1** (D-115 PROPOSED yet executed+merged) | **CONFIRMED as record-staleness, REFUTED as governance breach** | D-115 ADOPTED by amendment `:2727`; `ledger-scan` false-positive on the "Originally PROPOSED/OPEN" substring (R3-F09). The import + PR #1 rest on a ratified decision. | -| **L2** (D-071 controller-HA gates Stage 3) | **REFUTED as a Stage-3 blocker** | D-104 (`:2031-38`, ADOPTED) dispositions the controller-topology question for VR1 (single-unit per DC, HA deferred to Roosevelt) and is the entry the DC-DC phase was gated on; it references D-071 without amending it. D-071 (patch cadence) is Roosevelt-scoped. Annotate the ledger note (R3-F... minor). | -| **L3** (D-117-class naming drift) | **REFUTED** | Code is unambiguous (vr1-dc0/vr1-dc1, D-119); the double-namespace is RECORDED (D-117 + amendments). Only residual is stale prose (R3-F11); no gate depends on the ambiguous reading. doc-"DC1" = vr1-dc0. | -| **L4** (GUA/ULA IPAM reconciliation) | **REFUTED (one line)** | Stage 3 is isolated-L2 substrate + node/edge/rack VMs; it instantiates no tenant L3 addressing. The GUA/ULA reconciliation is a Stage-5 Neutron concern. Stage 3 does NOT depend on it. | -| **L5** (exit gate "conditionally met at best") | **CONFIRMED -- still conditional, honestly HELD (not run through)** | Node sizing ruled (D-121, but via the uncommitted model, R3-F06); edge sizing carried from applied office1_opnsense (measured basis, not a DC-edge boot measurement); netem still commented/unparameterized (unruled D-100 sub-item); the interface-naming boot measurement is a runtime TODO; `tofu plan` BLOCKED on 4 rack vars (L1-07). netem is correctly held, not fabricated. | - ---- - -## 6. Go / No-Go on Stage 3 - -**NO-GO for `tofu apply` as it stands** -- on hard-gate grounds, NOT because the substrate design is -unsound (it is sound: planes are isolated L2, state reconciles with no orphans, gauntlet green, -Option C fits with margin, validate passes). - -Blocking conditions to clear (each operator-gated, applied in a post-acceptance sweep): -1. **Re-rule the two inferred axes (R3-F01):** D-121 vault-HA (v-a vs v-b) and D-123 node - containment (Model A vs B). Both are one-line operator questions. Until then the vault overlay - and the Model-A record are unratified. -2. **Ship the SEC-010 artifact + mechanical pre-apply gate (R3-F02).** DC-LOCAL must be enforced, - not promised. -3. **Assign the rack transit/30 + rack IP in office1-netbox and feed the 4 tfvars (L1-07)** so a - plan can render and be reconciled against state. Wire `voffice1`'s transit leg (or gate the rack - route) so the route peer exists. - -Conditional-GO once 1-3 are met AND the record amendments (R3-F03..F07, F09, F11) are applied and -`optc-calc.py` is committed + re-run for Option C (R3-F06). The record corrections do not block the -substrate build; they block treating the DECISIONS as coherent, which is the operator's stated -concern. The MINOR/OBSERVATION items (F08, F10, F12, F13) are for the operator's judgment and need -not gate the sweep. - ---- - -*Prepared read-only; not committed, not merged, not pushed. Presented for operator review. Fixes -staged for an individually-gated post-acceptance sweep.* diff --git a/docs/tenant-cidr-overlap-correction-PLAN.md b/docs/tenant-cidr-overlap-correction-PLAN.md deleted file mode 100644 index ee6e7a0..0000000 --- a/docs/tenant-cidr-overlap-correction-PLAN.md +++ /dev/null @@ -1,154 +0,0 @@ -# Tenant CIDR overlap correction -- work order / change plan - -**Status:** DRAFT for the Code (jumphost) stream + operator ruling. Authored by the -main-chat stream at repo HEAD `59a7c73` (2026-07-06). Read-only analysis; no live -mutation and no committed-surface edit performed by main-chat. - -**Proposes:** decision **D-074** (PROPOSED; amends D-016). D-074 is next-free per -`bash scripts/ledger-scan.sh` at this HEAD. NOTE: D-073 and DOCFIX up to 105 were -consumed the same day by the jumphost stream -- **Code must re-run the scan and -re-confirm next-free before inserting the decision header.** Do not hand-type it. - ---- - -## 1. Finding (what is actually wrong) - -The stage-4 tenant CIDR guard in `scripts/tenant-onboard.sh` **over-enforces relative -to the onboarding contract it is supposed to implement**, and separately has an -overlap-detection hardening gap. - -- **Contract intent** (`docs/tenant-onboarding-contract.md` sec. 3, item 4): a tenant - CIDR must be RFC1918 and non-colliding with **(a) our allocations** and **(b) the - client's own on-prem/VPN ranges they may later interconnect**. The contract does - **not** require non-collision with *other tenants*. -- **Guard behaviour** (`tenant-onboard.sh` stage 4 -- the `grep -qw "$TENANT_CIDR"` - check run under `admin_env` against a cloud-wide `openstack subnet list`): it dies on - **any** exact subnet-string match cloud-wide, which includes other tenants' subnets. - So it forbids tenant-vs-tenant exact overlap that the contract never intended to - forbid. -- **Independent hardening gap:** the guard is exact-string (`grep -qw`), so it MISSES - partial/containing overlaps (e.g. a requested `10.20.20.128/25` against an existing - `10.20.20.0/24`). It is therefore simultaneously *too strict* on exact tenant matches - and *too loose* on partial overlaps against ranges we must actually protect. - -**Origin of the uniqueness behaviour (and a correction of the record):** it is an IPAM -artifact of **D-016**, which models a single `10.20.0.0/16` tenant pool and carves -non-overlapping /24s from it. It is **not** a dataplane requirement. An earlier -main-chat claim that the uniqueness rule was needed for "CAPI/Magnum reachability" was -wrong and is retracted here: per **D-035**, `capi-mgmt-v2` is single-homed on its own -`10.20.0.0/24` and reaches workload clusters via their API-LB floating IPs and OpenStack -via the API VIPs -- it never routes into other tenants' `10.20.x` space, so overlapping -tenant CIDRs cannot touch the management path. - -## 2. Proposed decision -- D-074 (PROPOSED; amends D-016) - -Tenant private (overlay) CIDRs are **tenant-chosen and MAY overlap across tenants.** -Non-collision is enforced **only** against a reserved-ranges registry (operator/infra -allocations, section 3). NetBox (the IPAM apex) tracks the reserved/routed space; tenant -overlay space leaves NetBox uniqueness tracking. - -Rationale: this aligns the implementation with the onboarding contract's already-stated -intent, and with the hard-isolation (SCS Domain Manager) persona. Overlap-allowed tenant -CIDRs is the standard isolated-multi-tenant model -- it is why Neutron ships -`allow_overlapping_ips=true` by default and why public-cloud VPCs let tenants pick -arbitrary RFC1918. Globally-unique-per-tenant is the enterprise-private-cloud pattern -(shared routed fabric, interoperating internal units), which is the opposite of what this -cloud is. - -Forward-only: existing tenants are **not** re-CIDR'd (northwind stays `10.20.20.0/24`; -design-decisions forbids transient-overlap re-CIDR). - -**PROPOSED means the operator has not ruled.** Nothing in sections 3-4 is implemented -until D-074 is ADOPTED. - -## 3. Reserved-ranges registry (the ONLY thing the guard rejects) - -Non-collision is enforced against these, using real subnet-overlap math (containment in -either direction), not string equality: - -| plane / allocation | CIDR (confirm live before wiring) | -|---|---| -| provider-public (incl. API VIPs + `provider-ext` FIP pool) | `10.12.4.0/22` | -| metal-admin | `10.12.8.0/22` | -| metal-internal | `10.12.12.0/22` | -| data-tenant | `10.12.16.0/22` | -| storage | `10.12.32.0/22` | -| replication | `10.12.36.0/22` | -| capi-mgmt tenant network (D-035) | `10.20.0.0/24` -- OPERATOR DECISION: keep reserved? (recommend YES) | -| metadata service | `169.254.169.254/32` (link-local; not RFC1918; low concern) | - -**These values were read from `docs/maas-as-built-reference.md` by main-chat.** Code MUST -re-derive them from the authoritative live/modeled source (pre-flight `PLANE_CIDRS` / -`scripts/lib-net.sh` / NetBox) before wiring the guard -- do not hardcode from this doc -(hard rules 2 and 3: dynamic lookup, no inferred values). Centralize the list in -`lib-net.sh` keyed by plane name, not as scattered literals. - -## 4. Gated correction procedure (harness-first; each phase gates the next) - -### Phase 0 -- read-only verification (MUST pass before any mutation) -- **0.1 Confirm `allow_overlapping_ips`.** Read it live (`juju config neutron-api ...` - or `neutron.conf` on a unit). If it is `false`, **STOP** -- flipping it is a wider - blast radius that needs its own ruling; do not proceed on the assumption it is `true`. -- **0.2 Overlap+Magnum proof on a FOIL tenant.** Onboard a foil tenant with a CIDR that - deliberately overlaps an existing tenant's /24; confirm (a) the subnet creates and (b) - a Magnum cluster reaches ACTIVE with a working API-LB. Capture evidence to - `~/openstack-baseline/`. (This foil doubles as the `d011-05` P3 isolation foil.) - -### Phase 1 -- guard change (Code's lane; Code owns `tenant-onboard.sh`) -- Replace the exact-string cloud-wide guard with subnet-overlap math (Python - `ipaddress`) against the reserved-ranges registry **only**. Drop the tenant-vs-tenant - check. -- Extend `tests/tenant-onboard/` with fixtures covering: exact tenant-vs-tenant overlap - (now **ALLOWED**), partial overlap against a reserved range (**REJECTED**), exact match - against a reserved range (**REJECTED**), and a disjoint range (**ALLOWED**). No script - change ships without its harness. -- `bash scripts/repo-lint.sh` + the harness green before commit. - -### Phase 2 -- doc alignment -- `tenant-onboarding-contract.md` sec. 3.4: make explicit that tenant-vs-tenant overlap - is permitted (the wording is already close; remove any implication of cloud-wide - uniqueness). -- Amend **D-016** with a forward-pointer to D-074. -- `runbooks/tenant-onboarding-runbook.md`: replace the "carve a non-colliding /24 from - `10.20.0.0/16`" guidance with "tenant CIDR is client-choice or the standard default; - only reserved ranges are rejected." -- `clientdocs/` intake field: reword "requested internal IP address" -> - "requested internal network range (CIDR) -- optional; overlap-safe; default per policy" - (also fixes the IP-vs-range wording mismatch). - -### Phase 3 -- create the devops tenant on the corrected cloud -- `bash scripts/tenant-onboard.sh <devops-client>` with `TENANT_CIDR` per the default - policy ruled in section 6. Verify with `scripts/tenant-assert.sh`. - -## 5. Lose / gain - -**Gain:** implementation matches the contract and the hard-isolation persona; onboarding -drops the "find a free /24" step; the intake field goes optional; clients can align to -their own on-prem/VPN ranges; NetBox's uniqueness job shrinks to the routed space that -genuinely must be unique (cleaner, and no global per-tenant IPAM registry to maintain -across DCs for Roosevelt); the 256-/24 ceiling of one /16 stops being a scaling limit. - -**Lose / cost:** debuggability -- an IP no longer uniquely identifies a tenant, so -captures/triage need net+tenant context (mitigate with disciplined naming); no -tenant-private-IP routing between tenants (an anti-pattern under hard isolation anyway); -IPAM discipline narrows to reserved ranges rather than disappearing; the guard change -touches a file Code is actively editing (coordination cost); D-016's `/16` pool becomes -vestigial (not reclaimed, just no longer enforced forward). - -## 6. Open operator decisions (nothing proceeds until ruled) - -1. **Rule D-074** -- adopt the overlap-allowed model? -2. **Default tenant CIDR policy** -- (a) uniform default (e.g. `10.0.0.0/24`) for all, or - (b) client-choice with a fallback default. Recommend (b). This sets what the devops - tenant receives. -3. **Keep `capi-mgmt` `10.20.0.0/24` reserved?** Recommend YES. -4. **`allow_overlapping_ips` confirmed `true`?** (Phase 0.1 gates everything.) - -## 7. Coordination / lane notes - -- The jumphost stream is actively editing `tenant-onboard.sh` and the contract today - (changelog addenda 21 / 23 / 26). The Phase-1 guard edit and all live steps are the - jumphost stream's lane; main-chat authored this plan only. **Code: pull and read this - before resuming onboard-script work.** -- D-074 is next-free per the scan at HEAD `59a7c73`; re-confirm before inserting. -- This is a new file with no edits to any shared/fenced file, to minimize collision. diff --git a/docs/upstream-bug-draft-dashboard-tls.md b/docs/upstream-bug-draft-dashboard-tls.md deleted file mode 100644 index bd50f7c..0000000 --- a/docs/upstream-bug-draft-dashboard-tls.md +++ /dev/null @@ -1,48 +0,0 @@ -# Upstream bug draft -- charm-openstack-dashboard TLS backend on vhost-less address - -STATUS: DRAFT for operator submission (Launchpad: charm-openstack-dashboard). -Prepared 2026-07-05 from the ops-update-20260705 RCA (repo changelog addendum 15, -D-072). ASCII-only. No site secrets: addresses below are RFC1918 lab values. - -## Title -Multi-space deployment: haproxy https backend rendered on cluster-binding address, -which never receives an apache SSL vhost; L4 health check masks the dead TLS path - -## Affected -charm-openstack-dashboard 2024.1/stable (observed rev 728 and rev 750; render logic -identical across both -- charmhelpers get_network_addresses + haproxy context). -Reactive charm generation; juju 3.6. - -## Environment -Charmed OpenStack Caracal (jammy), MAAS spaces deployment. Application bindings: -default ('') = space A (admin), cluster = space B (internal), public = space C -(provider). hacluster VIPs on all three planes. vault:certificates TLS. - -## What happens -- ApacheSSLContext.get_network_addresses() derives SSL vhost addresses from the - DEFAULT-binding fallback (private-address) and the PUBLIC binding: vhosts render - for the space-A and space-C unit addresses only. -- The haproxy context renders the https (443) backend server line from the - CLUSTER-binding address (space B). -- Result: haproxy forwards TLS to <space-B-addr>:433 where apache has NO SSL - vhost; apache serves the connection from the non-SSL main server in PLAINTEXT - (verified: plain-HTTP GET to :433 answers 200). TLS clients fail handshake - (curl exit 35 / 000) on every dashboard VIP. -- haproxy's check is Layer4-only, so the backend shows UP and nothing alerts. -- juju status is fully green throughout; the condition is silent from day one. - -## Expected -Either the SSL vhost set includes every address the charm itself renders as an -HTTPS backend target, or the https backend uses an address from the vhost set -(default-binding address), or the health check is HTTPS-aware so the breakage -is at least visible. - -## Workaround -Bind `cluster` to the same space as the application default binding. Verified: -backends re-render onto a vhost-served address and VIP https serves immediately. - -## Evidence trail (available on request) -charm rev 728 vs 750 full diff (no vhost-logic change); charmhelpers -get_network_addresses trail log (LP:1952414 style) showing identical tuple sets -across 2800+ renders since deploy; apache2ctl -S vhost list vs haproxy.cfg -backend; plaintext-200-on-ssl-port probe; post-rebind verification. diff --git a/docs/v1-pre-deploy-fixes.md b/docs/v1-pre-deploy-fixes.md deleted file mode 100644 index 8bd5206..0000000 --- a/docs/v1-pre-deploy-fixes.md +++ /dev/null @@ -1,832 +0,0 @@ -# v1 Pre-Deploy Fixes (v2 -- includes Designate deferral) - -**Purpose:** Single-pass repo hygiene before v1 deployment execution begins. Apply these fixes as one logical commit per group (nine commits total) before any execution document runs. - -**Status:** Authoritative for the v1 deploy track. Supersedes the v1 draft of this document. - -**Scope:** Repo-only changes. No cloud state is touched by this document. All changes are reviewed locally, committed to `main`, and pushed before the next do-document runs. - -**What changed in v2 of this change list (2026-05-27):** - -- Added commits 7-9 implementing the Designate-deferral decision (D-019). -- Amended commit 5 (deprecated runbook moves) -- `07-dns-zones.md` is now permanently deprecated per D-019, not "replaced by v1-do-doc-10-dns." -- Amended commit 6 (deprecated README content) -- same. -- Updated commit 4 (README.md refresh) -- adds language reflecting Designate deferral to the v1 scope section. -- Updated sect.10 verification commands to expect 11 VIPs, not 12. - -## Cross-references - -- D-002 (channel matrix) -- Vault row cleanup -- D-005 (Ceph Squid release) -- D-008 (DNS architecture) -- superseded by D-019; v2-scope -- D-011 (validation bar) -- amended by D-019 (Designate criterion dropped) -- D-014 (repo path) -- stale path correction -- D-017 (CAPI bootstrap cluster lifecycle) -- supersedes runbook 00 Phase 5 -- D-018 (teardown strategy) -- supersedes runbook 00 Phase 4 -- **D-019 (NEW) -- Designate deferral to v2; tenant resolvers use public DNS** -- Charmed Ceph `charm-ceph-osd` `config.yaml` and `metadata.yaml` (osd-devices semantics) -- Ceph BlueStore configuration reference (single-device co-located OSD pattern) - ---- - - -## 1. Bundle: remove ceph-osd `storage:` block - -### What - -In `bundle.yaml`, under the `ceph-osd` application, delete the entire `storage:` block. The `options.osd-devices: /dev/vdb` line stays. - -### Why - -The `osd-devices` declared under `storage:` is **additive** to the `options.osd-devices` value per the `charm-ceph-osd` `config.yaml`: "These devices are the range of devices that will be checked for and used across all service units, in addition to any volumes attached via the `--storage` flag during deployment." - -Concretely, with the current bundle: - -- `options.osd-devices: /dev/vdb` -> one OSD per unit using the 512 GB libvirt-attached disk -- `storage.osd-devices: loop,1024M` -> an *additional* 1 GB loopback OSD per unit - -Total: 2 OSDs per unit x 4 units = 8 OSDs, against `expected-osd-count: 4` on `ceph-mon`. The 1 GB OSDs are below practical minimums, asymmetric with the 512 GB primaries (CRUSH-weighting anti-pattern), and provide no operational value. - -The remaining `bluestore-db`, `bluestore-wal`, `cache-devices`, and `osd-journals` loopback entries are also being removed -- not because they break anything, but because: - -1. BlueStore co-locates DB and WAL on the data device when no separate volume is supplied (Ceph Reef BlueStore reference). -2. Loopback files on the same backing storage as `/dev/vdb` are not "faster than the primary device," so the standard rationale for separate DB/WAL devices doesn't apply. -3. `osd-journals` is unused under BlueStore (default since Luminous). - -### Diff - -```yaml -# BEFORE (in bundle.yaml, applications.ceph-osd) - ceph-osd: - charm: ceph-osd - channel: squid/stable - num_units: 4 - to: ["8", "9", "10", "11"] - options: - source: *ceph-source - osd-devices: /dev/vdb # libvirt-attached, MAAS-untracked, wiped 2026-05-22 - bindings: *internal-bindings - constraints: arch=amd64 tags=openstack - storage: # Loop-backed auxiliaries (testcloud has no real SSDs) - bluestore-db: loop,1024M - bluestore-wal: loop,1024M - cache-devices: loop,10240M - osd-devices: loop,1024M - osd-journals: loop,1024M # Legacy storage name still in squid metadata; benign - -# AFTER - ceph-osd: - charm: ceph-osd - channel: squid/stable - num_units: 4 - to: ["8", "9", "10", "11"] - options: - source: *ceph-source - osd-devices: /dev/vdb # libvirt-attached, MAAS-untracked, wiped 2026-05-22 - bindings: *internal-bindings - constraints: arch=amd64 tags=openstack -``` - -### Commit message - -``` -bundle: remove ceph-osd storage block to match expected-osd-count - -The storage: block declared a second osd-devices entry (loop,1024M) which -is additive to options.osd-devices per the charm config. That produced -8 OSDs against expected-osd-count: 4 on ceph-mon, with 1 GB loopback -OSDs as the asymmetric secondaries -- a CRUSH-weighting anti-pattern. - -Real production storage (DB/WAL on actual SSDs) will be declared on -Roosevelt. For the testcloud, BlueStore co-locates DB/WAL on the data -device which is the documented default for single-device setups. - -osd-journals is unused under BlueStore. -``` - ---- - - -## 2. design-decisions.md: D-002 -- remove Vault from OpenStack-core row - -### What - -In `docs/design-decisions.md` under D-002 (channel matrix), the OpenStack-core row currently lists `, vault` as one of the components on `2024.1/stable`. Remove that token. - -### Why - -Vault has its own track per Canonical's charm delivery table -- it runs on `1.8/stable`, not `2024.1/stable`. The D-002 table elsewhere (and the actual bundle.yaml) already reflects this; the OpenStack-core row description was a leftover from earlier drafting. - -### Diff - -```diff --| OpenStack core API charms (keystone, glance, nova-cloud-controller, neutron-api, cinder, placement, octavia, barbican, magnum, designate, openstack-dashboard, vault) | `2024.1/stable` | -+| OpenStack core API charms (keystone, glance, nova-cloud-controller, neutron-api, cinder, placement, octavia, barbican, magnum, designate, openstack-dashboard) | `2024.1/stable` | -``` - -### Commit message - -``` -docs/design-decisions: D-002 -- drop vault from OpenStack-core channel row - -Vault uses 1.8/stable per the Canonical charm delivery table, not the -2024.1/stable OpenStack-core track. The bundle.yaml already reflects -this; the design-decisions D-002 description had a stale token. -``` - -> **Note:** the `designate` token in the row above is correct as-of this commit (Designate is still on `2024.1/stable` channel). The Designate row is removed entirely by commit 8 (D-019). - ---- - - -## 3. design-decisions.md: D-014 -- update repo path - -### What - -In `docs/design-decisions.md` under D-014 (repo location), the path currently shows the per-user namespace from before the 2026-05-27 move. Update to the OpenStack-group path. - -### Why - -Per the user-memory pinned note: "Caracal rebuild repo (moved to OpenStack group 2026-05-27): https://git.baldurkeep.com/OpenStack/openstack-caracal-ipv4 (web), https://git.baldurkeep.com/git/OpenStack/openstack-caracal-ipv4.git (clone). Old jesse.austin/openstack-caracal-ipv4 path no longer exists; GitBucket does not redirect." - -### Diff - -```diff - ## D-014: Repository location - --**Decision:** `git.baldurkeep.com/jesse.austin/openstack-caracal-ipv4` for v1. -+**Decision:** `git.baldurkeep.com/OpenStack/openstack-caracal-ipv4` for v1. -+ -+- Web: `https://git.baldurkeep.com/OpenStack/openstack-caracal-ipv4` -+- Clone: `https://git.baldurkeep.com/git/OpenStack/openstack-caracal-ipv4.git` -+- Moved from `jesse.austin/openstack-caracal-ipv4` to the `OpenStack` group on 2026-05-27. GitBucket does not redirect from the old path. - - **Rationale:** Establishes a single repo per cloud lifecycle. v2 path TBD. -``` - -### Commit message - -``` -docs/design-decisions: D-014 -- update repo path after move to OpenStack group - -Repository moved from jesse.austin/openstack-caracal-ipv4 to the -OpenStack group on 2026-05-27. GitBucket does not redirect, so the -prior path is dead. -``` - ---- - - -## 4. README.md: refresh stale references - -### What - -Three hunks in `README.md`: - -1. Update the design-decisions reference range from "D-001 through D-016" to "D-001 through D-019" (post-D-017/D-018/D-019 additions). -2. Replace the inline runbook 00 description (which still mentions backups + capi-mgmt graceful teardown -- both invalidated by D-017/D-018). -3. Replace the 12-step deploy order with a pointer to the do-document set. -4. Add a v1 scope reduction note for Designate (per D-019). - -### Why - -The README's v1 deployment order block reflects the original 12-step runbook plan; the actual deploy path now flows through the v1-do-doc-NN execution documents. Keeping the README in sync prevents new operators from following the stale path. - -The D-019 note in the v1 scope section makes the Designate deferral explicit for anyone reading the README to understand v1 scope. - -### Diff (four hunks) - -**Hunk A -- D-range:** - -```diff --`--- docs/ -- `--- design-decisions.md # architectural record (D-001 through D-016) -+`--- docs/ -+ |--- design-decisions.md # architectural record (D-001 through D-019) -+ `--- netbox-vip-queue.md # post-deploy NetBox imports (workstream 2) -``` - -**Hunk B -- runbook 00 description (in the layout block):** - -```diff --| |--- 00-pre-deploy.md # backups, capi-mgmt graceful teardown -+| # (deprecated; see runbooks/deprecated/ -- superseded by D-017 + D-018 + v1-do-doc-NN set) -``` - -**Hunk C -- replace the deploy-order block:** - -```diff --## v1 deployment order -- --1. Verify NetBox state -- run NetBox imports if not already applied -- - `netbox/ipv4-prefixes-import.py` -- required -- - `netbox/ipv6-mark-reserved.py` -- required (Q3: tag existing IPv6 entries) --2. Run pre-flight checks (`scripts/pre-flight-checks.sh`) --3. Backup current cloud state (`runbooks/00-pre-deploy.md`) --4. Destroy existing OpenStack model (`runbooks/01-destroy-model.md`) --5. Deploy new bundle (`runbooks/02-deploy.md`) --6. Initialize Vault (`runbooks/03-vault-init.md`) --7. Set up Magnum domain (`runbooks/04-magnum-domain.md`) --8. Stand up CAPI bootstrap cluster on `capi-mgmt.maas` (`runbooks/04a-capi-bootstrap-cluster.md`) --9. Install Magnum CAPI Helm driver (`runbooks/05-magnum-capi-driver.md`) --10. Recreate tenant resources (`runbooks/06-tenant-setup.md`) --11. Populate DNS zones (`runbooks/07-dns-zones.md`) --12. Run validation (`runbooks/08-validate.md` + `scripts/validate.sh`) -+## v1 deployment order -+ -+The deploy is executed via the `runbooks/v1-do-doc-NN-*.md` execution documents in numeric order: -+ -+| Doc | Purpose | -+|---|---| -+| `v1-do-doc-01-prep.md` | Pre-flight state check (repo, openrc, MAAS state of 5 VMs) | -+| `v1-do-doc-02-pki.md` | Octavia PKI overlay generation | -+| `v1-do-doc-03-destroy.md` | Conditional model + MAAS teardown (clean state for rebuild) | -+| `v1-do-doc-04-deploy.md` | `juju deploy` + settle wait + on-disk PKI verification | -+| `v1-do-doc-05-vault-init.md` | Vault initialization + cert cascade + admin-openrc regeneration | -+| `v1-do-doc-06-magnum-domain.md` | Magnum Keystone domain setup | -+| `v1-do-doc-07-capi-bootstrap.md` | CAPI bootstrap cluster + workload pivot | -+| `v1-do-doc-08-magnum-driver.md` | Magnum CAPI Helm driver graft | -+| `v1-do-doc-09-tenant.md` | Tenant project/user/openrc + Snapshot 2 | -+| `v1-do-doc-10-validate.md` | D-011 acceptance criteria + Snapshot 3 | -+ -+NetBox imports are run separately (gated on external NetBox engineer review; see `netbox/README.md`). -``` - -**Hunk D -- v1 scope note about Designate deferral:** - -```diff - ## v1-specific design decisions (summary; see docs/design-decisions.md for full record) - - - **D-015 v1/v2 fork** -- IPv4-only v1; IPv6/dual-stack v2 deferred - - **D-016 IPv4 tenant pool hybrid model** -- NetBox owns upstream `/16` pool; Neutron owns per-project subnets within it - - **D-003 Option B network architecture** -- Provider `/22` carries both ext_net FIPs (`10.12.4.10-.223`) and OpenStack public API VIPs (`10.12.4.224-.254`) on the same L2 segment; fixes the tenant->API unreachability that caused Magnum OCCM crashloop on Bobcat testcloud - - **D-005 Ceph Squid** -- matches Caracal default; rehearses Roosevelt - - **D-006 Vault HA backend = etcd + easyrsa** - - **D-007 Magnum from day one** -- charm in bundle + CAPI Helm driver graft --- **D-008 DNS via Designate from day one** -- static /etc/hosts for bootstrap; Designate handles tenant-level resolution (A records only for v1) -+- **D-019 (supersedes D-008) DNS scope reduction for v1** -- Designate deferred to v2 alongside corporate DNS / NS-delegation work. Tenant subnets use public DNS (`1.1.1.1` / `1.0.0.1`) directly via `--dns-nameserver`. `*.cloud.neumatrix.local` FQDN tree remains internal-only, resolved via static `/etc/hosts` on bootstrap-relevant hosts. - - **D-009 Hacluster relations included at num_units=1** -- decorative on testcloud; documents the relation pattern for Roosevelt scale-up - - **No OVN pinning on testcloud** -- Roosevelt bare-metal will pin via `ovn-source` -``` - -### Commit message - -``` -README: refresh stale runbook references, reflect D-019 scope reduction - -The v1 deployment order block referenced runbooks 00-08 which are -being moved to deprecated/ in favor of v1-do-doc-NN execution documents. -Replaces the order block with a pointer to the do-document set. - -Also reflects D-019: Designate deferred to v2; v1 tenant resolvers use -public DNS. Adds the netbox-vip-queue.md reference. Updates the design- -decisions D-range from D-001-D-016 to D-001-D-019. -``` - ---- - - -## 5. Move superseded runbooks to `runbooks/deprecated/` - -### What - -`git mv` each superseded runbook into a new `runbooks/deprecated/` directory. Add `runbooks/deprecated/README.md` explaining the deprecation in commit 6. - -### Why - -The v1-do-doc-NN set replaces the prior runbook 00-08 work. Keeping the originals in `runbooks/deprecated/` preserves the audit trail without misleading new operators into following the old paths. - -### Files to move - -| From | To | Replacement | -|---|---|---| -| `runbooks/00-pre-deploy.md` | `runbooks/deprecated/00-pre-deploy.md` | superseded by D-017 + D-018 (no per-cycle backups; teardown direct-to-MAAS); v1-do-doc-01 covers prep | -| `runbooks/01a-octavia-pki-generation.md` | `runbooks/deprecated/01a-octavia-pki-generation.md` | `v1-do-doc-02-pki.md` | -| `runbooks/02-deploy.md` | `runbooks/deprecated/02-deploy.md` | `v1-do-doc-04-deploy.md` | -| `runbooks/03-vault-init.md` | `runbooks/deprecated/03-vault-init.md` | `v1-do-doc-05-vault-init.md` | -| `runbooks/04-magnum-domain.md` | `runbooks/deprecated/04-magnum-domain.md` | `v1-do-doc-06-magnum-domain.md` | -| `runbooks/04a-capi-bootstrap-cluster.md` | `runbooks/deprecated/04a-capi-bootstrap-cluster.md` | `v1-do-doc-07-capi-bootstrap.md` | -| `runbooks/05-magnum-capi-driver.md` | `runbooks/deprecated/05-magnum-capi-driver.md` | `v1-do-doc-08-magnum-driver.md` | -| `runbooks/06-tenant-setup.md` | `runbooks/deprecated/06-tenant-setup.md` | `v1-do-doc-09-tenant.md` | -| **`runbooks/07-dns-zones.md`** | **`runbooks/deprecated/07-dns-zones.md`** | **deferred to v2 per D-019** (no v1 replacement) | -| `runbooks/08-validate.md` | `runbooks/deprecated/08-validate.md` | `v1-do-doc-10-validate.md` | - -### Files NOT moved - -| File | Why kept in `runbooks/` | -|---|---| -| `runbooks/01-destroy-model.md` | Referenced by v1-do-doc-03 as a conditional sub-procedure; still active | - -### Git commands - -```bash -cd "$HOME/openstack-caracal-ipv4" - -mkdir -p runbooks/deprecated - -git mv runbooks/00-pre-deploy.md runbooks/deprecated/ -git mv runbooks/01a-octavia-pki-generation.md runbooks/deprecated/ -git mv runbooks/02-deploy.md runbooks/deprecated/ -git mv runbooks/03-vault-init.md runbooks/deprecated/ -git mv runbooks/04-magnum-domain.md runbooks/deprecated/ -git mv runbooks/04a-capi-bootstrap-cluster.md runbooks/deprecated/ -git mv runbooks/05-magnum-capi-driver.md runbooks/deprecated/ -git mv runbooks/06-tenant-setup.md runbooks/deprecated/ -git mv runbooks/07-dns-zones.md runbooks/deprecated/ -git mv runbooks/08-validate.md runbooks/deprecated/ - -git status -# Expect: 10 renames staged, runbooks/01-destroy-model.md untouched -``` - -### Commit message - -``` -runbooks: move superseded files to runbooks/deprecated/ - -These are replaced by v1-do-doc-NN-*.md execution documents (added in -follow-up commits). The 01-destroy-model.md runbook stays in place -- it's -referenced by v1-do-doc-03 as a conditional sub-procedure. - -07-dns-zones.md is deferred to v2 per D-019 with no v1 replacement -(Designate is no longer in v1 scope). - -History preserved via git mv. -``` - ---- - - -## 6. Add a deprecation banner to `runbooks/deprecated/README.md` - -### What - -Create a new file `runbooks/deprecated/README.md` with a banner explaining the deprecation scope and a replacement map. - -### Content - -```markdown -# Deprecated v1 Runbooks - -The runbooks in this directory have been superseded by the -`runbooks/v1-do-doc-NN-*.md` execution documents (or, in the case of -`07-dns-zones.md`, deferred to v2 entirely per D-019). - -They are preserved here so the audit trail from the early v1 drafting -phase remains accessible. **Do not execute them.** The v1 deploy is -gated through the do-document set. - -## Replacement map - -| Deprecated runbook | Replacement | -|---|---| -| `00-pre-deploy.md` | superseded by D-017 + D-018 (no per-cycle backups; direct MAAS teardown); `v1-do-doc-01-prep.md` covers prep | -| `01a-octavia-pki-generation.md` | `v1-do-doc-02-pki.md` | -| `02-deploy.md` | `v1-do-doc-04-deploy.md` | -| `03-vault-init.md` | `v1-do-doc-05-vault-init.md` | -| `04-magnum-domain.md` | `v1-do-doc-06-magnum-domain.md` | -| `04a-capi-bootstrap-cluster.md` | `v1-do-doc-07-capi-bootstrap.md` | -| `05-magnum-capi-driver.md` | `v1-do-doc-08-magnum-driver.md` | -| `06-tenant-setup.md` | `v1-do-doc-09-tenant.md` | -| `07-dns-zones.md` | **deferred to v2 per D-019** (no v1 replacement) | -| `08-validate.md` | `v1-do-doc-10-validate.md` | - -`01-destroy-model.md` is **not** in this directory -- it remains active in -`runbooks/` and is referenced as a conditional sub-procedure by -`v1-do-doc-03-destroy.md`. -``` - -### Commit message - -``` -runbooks/deprecated: add README explaining the deprecation scope - -Includes the deprecated -> replacement mapping for operators who arrive -via git log searches or stale internal references. -``` - ---- - - -## 7. Bundle: remove Designate (per D-019) - -### What - -In `bundle.yaml`, remove four applications, seven relations, and update a header comment to reflect Designate deferral to v2. VIP `10.12.4.227` becomes unused space in the `10.12.4.224-.254` range (same status as `10.12.4.225` reserved for v2 ceph-radosgw HA). - -### Why - -Per D-019 (added by commit 8 of this change list): Designate is deferred to v2 alongside corporate-DNS / NS-delegation work. v1 testcloud topology investigation (2026-05-27 session) confirmed: - -1. **Outside-in DNS** isn't needed for v1 -- corporate clients reach the cloud through the existing `openstack.baldurkeep.com -> 10.17.4.20 -> 10.12.x` HTTPS proxy chain, not via the `*.cloud.neumatrix.local` FQDN tree. The edge nginx (`neumatrix-nginx` at `10.17.8.7`) cannot route to `10.12.x` directly anyway. -2. **Inside-out DNS** doesn't require Designate -- tenant subnets can specify public DNS (`1.1.1.1`, `1.0.0.1`) directly via `--dns-nameserver` at subnet-create time. -3. **FIP DNS auto-registration** (the remaining v1 use case for Designate) is nice-to-have, not load-bearing for any v1 acceptance criterion. - -Removing Designate now removes one charm, four applications (with subordinate routers), seven relations, and one VIP from the v1 deploy surface, reducing first-deploy troubleshooting scope. - -### Diff - -**Remove four applications** (search-and-delete each block in `bundle.yaml`): - -```yaml - # ===================================================================== - # DNS: Designate (NEW for Caracal v1 per D-008) - # ===================================================================== - # Naming convention: <service>.omega.dc0.vr0.cloud.neumatrix.local - - designate: - charm: designate - channel: 2024.1/stable - num_units: 1 - to: [lxd:8] - options: - openstack-origin: *openstack-origin - nameservers: "ns1.omega.dc0.vr0.cloud.neumatrix.local. ns2.omega.dc0.vr0.cloud.neumatrix.local." - vip: 10.12.4.227 - os-public-hostname: designate.omega.dc0.vr0.cloud.neumatrix.local - bindings: *api-bindings - constraints: arch=amd64 - - designate-mysql-router: - charm: mysql-router - channel: 8.0/stable - - designate-bind: - charm: designate-bind - channel: 2024.1/stable - num_units: 1 - to: [lxd:8] - bindings: - "": provider # unit on provider so bind9:53 reachable from tenants (D-003) - cluster: metal # peer traffic stays internal (decorative with num_units=1) - constraints: arch=amd64 -``` - -**Remove the `designate-hacluster:` line** from the hacluster subordinate block: - -```diff - keystone-hacluster: { charm: hacluster, channel: 2.4/stable } - glance-hacluster: { charm: hacluster, channel: 2.4/stable } - neutron-api-hacluster: { charm: hacluster, channel: 2.4/stable } - nova-cloud-controller-hacluster: { charm: hacluster, channel: 2.4/stable } - placement-hacluster: { charm: hacluster, channel: 2.4/stable } - openstack-dashboard-hacluster: { charm: hacluster, channel: 2.4/stable } - cinder-hacluster: { charm: hacluster, channel: 2.4/stable } - octavia-hacluster: { charm: hacluster, channel: 2.4/stable } - barbican-hacluster: { charm: hacluster, channel: 2.4/stable } - magnum-hacluster: { charm: hacluster, channel: 2.4/stable } - vault-hacluster: { charm: hacluster, channel: 2.4/stable } - # v2-deferred: ceph-radosgw-hacluster: { charm: hacluster, channel: 2.4/stable } -- designate-hacluster: { charm: hacluster, channel: 2.4/stable } -+ # v2-deferred (D-019): designate-hacluster: { charm: hacluster, channel: 2.4/stable } -``` - -**Remove seven relations** from the `relations:` block: - -```yaml - # ---- Designate (DNS) -- NEW for Caracal v1 per D-008 - - [designate-mysql-router:db-router, mysql-innodb-cluster:db-router] - - [designate-mysql-router:shared-db, designate:shared-db] - - [designate:identity-service, keystone:identity-service] - - [designate:amqp, rabbitmq-server:amqp] - - [designate:certificates, vault:certificates] - - [designate:dns-backend, designate-bind:dns-backend] - - [designate:ha, designate-hacluster:ha] -``` - -**Update the header comment block** (decision references in the bundle's top-of-file block): - -```diff - D-006 Vault HA via etcd + easyrsa - D-007 Magnum Layer A + Layer B graft -- D-008 Designate day-one -+ D-019 (supersedes D-008) Designate deferred to v2 - D-009 hacluster subordinates (decorative on testcloud) -``` - -**Update the HA subordinate header comment:** - -```diff - # ===================================================================== -- # HA Cluster Subordinates (12 active for v1; ceph-radosgw deferred to v2) -+ # HA Cluster Subordinates (11 active for v1; ceph-radosgw + designate deferred to v2) - # ===================================================================== -``` - -### Verification post-edit - -```bash -cd "$HOME/openstack-caracal-ipv4" - -# 1. No designate application remains -grep -E "^ designate" bundle.yaml \ - && echo "[FAIL] designate-related application still present" \ - || echo "[OK] no designate applications" - -# 2. No designate relations remain -grep -E "designate" bundle.yaml | grep -vE "^[[:space:]]*#" -# Expect: no output (the only remaining 'designate' tokens should be commented) - -# 3. VIP count is now 11 -VIP_COUNT=$(grep -cE "^[[:space:]]+vip: 10\.12\.4\." bundle.yaml) -echo "VIPs: $VIP_COUNT (expect 11)" - -# 4. YAML still parses -python3 -c "import yaml; yaml.safe_load(open('bundle.yaml')); print('[OK] YAML parses')" -``` - -### Commit message - -``` -bundle: remove Designate per D-019 (deferred to v2) - -Removes the designate, designate-bind, designate-mysql-router, and -designate-hacluster applications, plus all seven designate-related -relations. Updates header comments to reflect the deferral. - -Rationale per D-019: v1 testcloud topology investigation confirmed -outside-in DNS is not needed (corporate clients reach the cloud via -the openstack.baldurkeep.com HTTPS proxy chain, not via *.cloud. -neumatrix.local FQDNs). Tenant subnets use public DNS directly via ---dns-nameserver. FIP DNS auto-registration is not load-bearing for -any v1 acceptance criterion. - -VIP 10.12.4.227 becomes unused space in 10.12.4.224-.254 (same -status as 10.12.4.225 reserved for v2 ceph-radosgw HA). - -Reduces v1 deploy surface by one charm + three subordinates + -seven relations + one VIP. -``` - ---- - - -## 8. design-decisions.md: add D-019 + amend D-008 and D-011 - -### What - -Three coordinated edits in `docs/design-decisions.md`: - -1. Add a new D-019 entry with the full deferral rationale and v2 deltas -2. Mark D-008 as "superseded by D-019 -- v2-scope" -3. Amend D-011 (validation bar) to remove the "Designate resolves" criterion - -### Why - -Captures the decision in the authoritative design-decisions record. Preserves the audit trail (D-008 stays with its original content but with the superseded status flag). - -### Diff -- Hunk A: amend D-008 status - -```diff - ## D-008: DNS architecture - --**Decision:** Layered -- static /etc/hosts for bootstrap + Designate (in bundle from day one) for tenant-level resolution. -+**Status:** Superseded by D-019 (2026-05-27). v2-scope. Original decision text preserved below for audit. -+ -+**Decision (original; superseded):** Layered -- static /etc/hosts for bootstrap + Designate (in bundle from day one) for tenant-level resolution. -``` - -### Diff -- Hunk B: amend D-011 validation bar - -In the D-011 section, the testcloud validation criteria currently list (among others) a Designate resolution check. Remove the relevant bullet. - -```diff - ## D-011: Roosevelt-rehearsal validation bar - - **Decision:** v1 testcloud must pass these criteria before being declared "deploy-equivalent" to Roosevelt: - - - All charms `active/idle` (per `juju status`) - - Tenant subnet -> OpenStack API reachability (Bobcat Magnum OCCM crashloop regression test) - - Octavia LBaaS end-to-end (LB create + member health + failover + recovery) - - Magnum CAPI end-to-end (cluster template + cluster create + CREATE_COMPLETE) - - Vault unseal-after-reboot survives a power cycle --- Designate resolves API FQDNs via the Designate VIP - - Snapshots 1, 2, 3 captured at the appropriate stages (per D-012) - -+**Amendment (2026-05-27):** Per D-019, the "Designate resolves" criterion is removed for v1. Designate is deferred to v2; tenant subnets resolve via public DNS. v2 will reinstate a DNS-resolution validation criterion calibrated to whatever DNS mechanism is in place (NS delegation from corporate DNS, or otherwise). -``` - -### Diff -- Hunk C: add D-019 (new entry, append to end of design-decisions.md) - -```markdown ---- - - -## D-019: DNS scope reduction for v1 -- Designate deferred to v2 - -**Decision (2026-05-27):** Designate is removed from the v1 testcloud bundle and deferred to v2 alongside corporate DNS / NS delegation work. v1 tenant subnets resolve via public DNS (`1.1.1.1`, `1.0.0.1`) directly via the `--dns-nameserver` option at subnet-create time. - -**Supersedes:** D-008 (DNS architecture). - -**Amends:** D-011 (validation bar -- removes "Designate resolves" criterion). - -### Rationale - -Three findings from the 2026-05-27 testcloud topology investigation: - -1. **Outside-in DNS** (corporate clients resolving `*.cloud.neumatrix.local`) is not needed for v1. Corporate access to the cloud already flows through the existing `openstack.baldurkeep.com -> 10.17.4.20 -> 10.12.x` HTTPS proxy chain (handled by the edge nginx at `10.17.8.7`), which does not depend on corporate-side resolution of cloud-internal FQDNs. - -2. **The edge nginx cannot route to `10.12.x` directly.** Inspection confirmed the edge has only `10.17.8.7/22` plus a tailscale interface; reaching `10.12.4.x` requires the libvirt-host NAT path. Adding DNS to the testcloud would require parallel UDP/53 NAT/proxy plumbing across three hosts (edge nginx, libvirt host, internal nginx) for a feature that has no v1 consumer. - -3. **Inside-out DNS** (tenant VMs resolving external names) is satisfied by tenant subnets pointing `--dns-nameserver` at public DNS (`1.1.1.1`, `1.0.0.1`). Designate is not needed in the inside-out path either, since: - - Tenant VMs don't need to resolve cloud-internal FQDNs (their API access goes through documented IPs / `--cloud` configs in cloud.conf) - - Cross-tenant DNS visibility is not a v1 requirement - -The remaining v1 use case for Designate (FIP DNS auto-registration via the `neutron-api <-> designate` integration) is informational only -- nothing in v1 consumes those records. - -### v1 implementation - -- Tenant subnets created with `--dns-nameserver 1.1.1.1 --dns-nameserver 1.0.0.1` (or via the openrc `OS_DNS_NAMESERVERS` env) -- CAPI workload cluster template variable `OPENSTACK_DNS_NAMESERVERS` set to `1.1.1.1,1.0.0.1` (per `v1-do-doc-07-capi-bootstrap.md` sect.13) -- Cloud-internal `*.cloud.neumatrix.local` FQDN tree resolved via static `/etc/hosts` on bootstrap-relevant hosts (jumphost, openstack0-3, LXD containers per charm bootstrap, capi-mgmt -- staged in `v1-do-doc-05-vault-init.md` sect.11 and `v1-do-doc-07-capi-bootstrap.md` sect.6) -- Charms continue to use FQDN-based `os-public-hostname` (cert SANs depend on it) -- internal resolution via `/etc/hosts` is sufficient - -### v2 plan - -- Re-introduce Designate (charm + designate-bind + relations + hacluster sub) -- NS delegation from corporate DNS to designate-bind on a real (non-NAT) network VIP -- Tenant subnets transitioning to use Designate VIP as their resolver (after corporate DNS delegation lands) -- Designate v2 deploy on a real-network Roosevelt or v2-testcloud topology where the bridging-host complexity from v1 testcloud doesn't apply -- D-011 validation re-introduces a calibrated DNS-resolution criterion (mechanism TBD: NS delegation working end-to-end vs static A records at corporate DNS) - -### v2-residency note - -The IPv6 prefixes already imported into NetBox (and marked Reservation status) include allocations that would be appropriate for Designate's VIPs in a v2 design -- these stay in NetBox as Reservation until v2 work begins. -``` - -### Commit message - -``` -docs/design-decisions: add D-019, amend D-008/D-011 (DNS scope reduction) - -D-019 captures the Designate deferral to v2 with rationale grounded in -the 2026-05-27 testcloud topology investigation: outside-in DNS not -needed (corporate clients use openstack.baldurkeep.com HTTPS chain); -edge nginx can't route to cloud-internal anyway; inside-out is -satisfied by tenant --dns-nameserver pointing at public DNS. - -D-008 (DNS architecture) marked superseded; original text preserved -for audit trail. - -D-011 validation bar amended to remove "Designate resolves" criterion; -v2 will reinstate a calibrated DNS criterion. -``` - ---- - - -## 9. netbox-vip-queue.md: drop the Designate VIP row - -### What - -In `docs/netbox-vip-queue.md`, remove the row for Designate's VIP `10.12.4.227`. The queue goes from 12 entries to 11. - -### Why - -Per D-019, Designate has no VIP in v1. `10.12.4.227` becomes unused space within the `10.12.4.224-.254` range -- same status as `10.12.4.225` (reserved for v2 ceph-radosgw HA). - -### Diff - -```diff - # NetBox VIP Queue (post-deploy) - - The following 12 IPAddress entries should be added to NetBox after the - v1 deploy completes and the engineer-review of `netbox/ipv4-prefixes- - import.py` has landed. - --| VIP | Service | Notes | --|---|---|---| --| 10.12.4.224 | barbican | per D-003 | --| 10.12.4.226 | cinder | per D-003 | --| 10.12.4.227 | designate | per D-003 + D-008 | --| 10.12.4.228 | glance | per D-003 | --| 10.12.4.229 | keystone | per D-003 | --| 10.12.4.230 | magnum | per D-003 | --| 10.12.4.231 | neutron-api | per D-003 | --| 10.12.4.232 | nova-cloud-controller | per D-003 | --| 10.12.4.233 | octavia | per D-003 | --| 10.12.4.234 | openstack-dashboard | per D-003 | --| 10.12.4.235 | placement | per D-003 | --| 10.12.4.236 | vault | per D-003 | -+The following 11 IPAddress entries should be added to NetBox after the -+v1 deploy completes and the engineer-review of `netbox/ipv4-prefixes- -+import.py` has landed. -+ -+| VIP | Service | Notes | -+|---|---|---| -+| 10.12.4.224 | barbican | per D-003 | -+| 10.12.4.226 | cinder | per D-003 | -+| 10.12.4.228 | glance | per D-003 | -+| 10.12.4.229 | keystone | per D-003 | -+| 10.12.4.230 | magnum | per D-003 | -+| 10.12.4.231 | neutron-api | per D-003 | -+| 10.12.4.232 | nova-cloud-controller | per D-003 | -+| 10.12.4.233 | octavia | per D-003 | -+| 10.12.4.234 | openstack-dashboard | per D-003 | -+| 10.12.4.235 | placement | per D-003 | -+| 10.12.4.236 | vault | per D-003 | -+ -+**Reserved (unused in v1):** -+ -+- `10.12.4.225` -- reserved for v2 ceph-radosgw HA (workstream-2 decision) -+- `10.12.4.227` -- reserved for v2 designate (D-019 deferral) -``` - -(The exact line counts may differ if the existing file has more preamble -- adjust to the actual structure when editing.) - -### Commit message - -``` -docs/netbox-vip-queue: drop designate row per D-019 - -VIP 10.12.4.227 is no longer in v1; designate deferred to v2. Adds a -"reserved (unused in v1)" section to make the unused slots in the -.224-.254 range explicit for future reference. -``` - ---- - - -## 10. Verification (read-only) after the nine commits land - -```bash -cd "$HOME/openstack-caracal-ipv4" -git pull -git log --oneline -10 - -echo "=== 1. ceph-osd no longer has a storage block ===" -grep -A 12 "^ ceph-osd:" bundle.yaml | grep "storage:" \ - && echo "[FAIL] storage: block still present in ceph-osd" \ - || echo "[OK] no storage: in ceph-osd" - -echo "=== 2. D-002 vault token removed ===" -grep -E "OpenStack core API charms" docs/design-decisions.md | grep -v "vault)" \ - && echo "[OK]" \ - || echo "[FAIL] D-002 still contains vault token" - -echo "=== 3. D-014 reflects new repo path ===" -grep "OpenStack/openstack-caracal-ipv4" docs/design-decisions.md \ - && echo "[OK]" \ - || echo "[FAIL]" - -echo "=== 4. README reflects v1-do-doc set ===" -grep "v1-do-doc-NN" README.md \ - && echo "[OK]" \ - || echo "[FAIL]" - -echo "=== 5. Designate gone from bundle ===" -grep -E "^ designate" bundle.yaml \ - && echo "[FAIL] designate-related app still present" \ - || echo "[OK] no designate apps" - -grep -E "designate" bundle.yaml | grep -vE "^[[:space:]]*#" -# Expect: empty (only commented references remain) - -echo "=== 6. VIP count is 11 ===" -VIP_COUNT=$(grep -cE "^[[:space:]]+vip: 10\.12\.4\." bundle.yaml) -echo "VIPs: $VIP_COUNT (expect 11)" - -echo "=== 7. D-019 added ===" -grep "^## D-019" docs/design-decisions.md \ - && echo "[OK]" \ - || echo "[FAIL]" - -echo "=== 8. D-008 marked superseded ===" -grep -A 2 "^## D-008" docs/design-decisions.md | grep "Superseded by D-019" \ - && echo "[OK]" \ - || echo "[FAIL]" - -echo "=== 9. netbox-vip-queue.md has 11 entries ===" -QUEUE_COUNT=$(grep -cE "^\| 10\.12\.4\." docs/netbox-vip-queue.md) -echo "VIP queue entries: $QUEUE_COUNT (expect 11)" - -echo "=== 10. runbooks/deprecated/ has 10 files + README ===" -ls runbooks/deprecated/ | wc -l -# Expect: 11 (10 deprecated runbooks + README.md) - -echo "=== 11. YAML still parses ===" -python3 -c "import yaml; yaml.safe_load(open('bundle.yaml')); print('[OK] YAML parses')" -``` - ---- - - -## 11. Acceptance criteria - -- [ ] All 9 commits landed and pushed to `main` -- [ ] Verification section 10 returns `[OK]` for all 11 checks -- [ ] `git log --oneline` shows the 9 commits in order -- [ ] `bundle.yaml` parses cleanly via `python3 -c "import yaml; yaml.safe_load(open('bundle.yaml'))"` -- [ ] No untracked changes (`git status` clean except for any local scratch) - -Once all checked, proceed to `v1-do-doc-01-prep.md` execution. - ---- - - -## 12. Change log - -| Date | Change | Reference | -|---|---|---| -| 2026-05-27 | v1 (six commits): ceph-osd storage block; D-002/D-014 cleanups; README refresh; deprecated runbook moves; deprecation README | Initial drafting | -| 2026-05-27 | v2 (nine commits): + Designate deferral commits 7/8/9; amended commit 5/6/10 to reflect 07-dns-zones permanent deprecation; updated commit 4 README hunks; VIP count expectations updated from 12 to 11 | D-019 decision; testcloud topology investigation | diff --git a/docs/vr1-office1-as-built.md b/docs/vr1-office1-as-built.md index 9e40994..e0cb0dc 100644 --- a/docs/vr1-office1-as-built.md +++ b/docs/vr1-office1-as-built.md @@ -140,7 +140,7 @@ **NetBox API token format (4.6):** the wire form is `nbt_<key>.<plaintext>`. The API's `token` field alone is NOT usable -- present it raw and you get `403 Invalid v1 token`. See -`docs/changelog-20260713-office1-netbox-deployed.md`. +`docs/archive/changelogs/changelog-20260713-office1-netbox-deployed.md`. --- diff --git a/runbooks/dc-dc-phase1-office1-standup.md b/runbooks/dc-dc-phase1-office1-standup.md index e45405c..76fdd61 100644 --- a/runbooks/dc-dc-phase1-office1-standup.md +++ b/runbooks/dc-dc-phase1-office1-standup.md @@ -679,7 +679,7 @@ this runbook. Record that as the blocking condition rather than inventing placeholder-looking values to "get past" it (this script is designed to fail loud on exactly that pattern, per its own header). See -`docs/dc-dc-netem-and-ula-gua-proposal.md` for the drafted ULA-generation +`docs/archive/dc-dc-netem-and-ula-gua-proposal.md` for the drafted ULA-generation guidance -- a PROPOSAL, not a ruling. **MUTATION -- once the literals above are real (a later session/step)** diff --git a/runbooks/dc-dc-phase2-tofu-dc-substrate.md b/runbooks/dc-dc-phase2-tofu-dc-substrate.md index d29c823..cce4ca6 100644 --- a/runbooks/dc-dc-phase2-tofu-dc-substrate.md +++ b/runbooks/dc-dc-phase2-tofu-dc-substrate.md @@ -180,11 +180,11 @@ scope, D-129). Node count is R-3 = 3 control + 2 compute + **4 storage = 9/DC** (Step 6's "8" is superseded). The Model -A single-root layout is preserved as a revert (git tag `model-a-fallback` + `docs/model-a-fallback-plan.md`). +A single-root layout is preserved as a revert (git tag `model-a-fallback` + `docs/archive/model-a-fallback-plan.md`). Sizing derived by `scripts/dc-dc-whole-host-budget.py` (870/1024 GiB, FIT). SEC-010 DC-LOCAL is enforced by the bootstrap's transit FORWARD-drop (NOT a global `ip_forward=0`), kept interface-scoped -- **never globalize it** (D-125 br_netfilter constraint: bridged WAN frames traverse L3 FORWARD). See -`docs/changelog-20260716-review-sweep-phaseC.md`. +`docs/archive/changelogs/changelog-20260716-review-sweep-phaseC.md`. **D-125 (bridge-in per-DC ISP egress -- closes OBS-3).** Under Model B the inner `vr1-dc0-wan` was a NAT with no egress (its host `vvr1-dc0` is transit-only). Fix, folded into the roots above: @@ -196,7 +196,7 @@ WAN keeps its static `.2` on that same /24 (edge config, D-113 -- unchanged, no re-address). - **DEPLOY-TIME GATE (Step 12):** before trusting the OPNsense chain, isolation-test egress -- a throwaway guest on `br-vr1-dc0-wan` must get a vcloud-ISP address + ping out. FAIL => revert to the D-125 double-NAT - fallback (not a redesign). See D-125 + `docs/changelog-20260716-review-sweep-phaseC2-D125.md`. + fallback (not a redesign). See D-125 + `docs/archive/changelogs/changelog-20260716-review-sweep-phaseC2-D125.md`. --- @@ -341,9 +341,9 @@ > **What `vr1-dc0`'s edge does (the proven Office1 path):** prep the image (this step) -> Step 5 > wires the domain -> Step 8 boots it -> D-112(c) console bootstrap -> configure over SSH + the > REST API (the REPLACEMENT chain below). See -> `docs/changelog-20260712-office1-opnsense-edge-build.md`, -> `docs/changelog-20260713-opnsense-api-proven.md`, and -> `docs/changelog-20260713-opnsense-api-write-proven.md`. +> `docs/archive/changelogs/changelog-20260712-office1-opnsense-edge-build.md`, +> `docs/archive/changelogs/changelog-20260713-opnsense-api-proven.md`, and +> `docs/archive/changelogs/changelog-20260713-opnsense-api-write-proven.md`. This step produces one plain file on the vcloud host filesystem -- the prepped base image -- no libvirt or MAAS object is created yet. Run from the vcloud diff --git a/runbooks/dc-dc-phase3-maas-enlist-deploy.md b/runbooks/dc-dc-phase3-maas-enlist-deploy.md index 00cbc70..4662c81 100644 --- a/runbooks/dc-dc-phase3-maas-enlist-deploy.md +++ b/runbooks/dc-dc-phase3-maas-enlist-deploy.md @@ -74,7 +74,7 @@ `scripts/phase-00-maas-standup.sh` now accept an opt-in `$DC` env var and call `lib_net_select_dc`/`lib_hosts_select_dc` immediately after sourcing (unset `$DC` is byte-for-byte unchanged VR0/DC0 behavior). See - `docs/changelog-20260709-maas-scripts-dc-param.md`. This means `DC=vr1-dc0 + `docs/archive/changelogs/changelog-20260709-maas-scripts-dc-param.md`. This means `DC=vr1-dc0 bash scripts/reenroll-hosts.sh` (etc.) is now a REAL invocation, not hypothetical -- but it does not by itself unblock anything: calling any of the three with `DC=vr1-dc0` or `DC=vr1-dc1` today fails loud immediately at @@ -530,7 +530,7 @@ - [x] Gap #2's CLI half CLOSED 2026-07-10 (DOCFIX-166): `reenroll-hosts.sh` / `carve-host-interfaces.sh` / `phase-00-maas-standup.sh` now accept `$DC` and call the selectors -- - see `docs/changelog-20260709-maas-scripts-dc-param.md`. Remaining: + see `docs/archive/changelogs/changelog-20260709-maas-scripts-dc-param.md`. Remaining: once `lib-hosts.sh` gets a real `$DC` block (checklist item above), `reenroll-hosts.sh`/`carve-host-interfaces.sh` become directly callable for that DC with no further script change; separately, diff --git a/runbooks/dc-dc-phase4-juju-bundle-per-dc.md b/runbooks/dc-dc-phase4-juju-bundle-per-dc.md index 81428a2..2f99bd7 100644 --- a/runbooks/dc-dc-phase4-juju-bundle-per-dc.md +++ b/runbooks/dc-dc-phase4-juju-bundle-per-dc.md @@ -212,7 +212,7 @@ **`overlays/dc-dc-ipv6-family-matrix.yaml` now exists**, drafted 2026-07-10 against REAL charm source/config (not guessed option names) -- see -`docs/dc-dc-ipv6-charm-research.md` for full sourcing and per-charm +`docs/archive/dc-dc-ipv6-charm-research.md` for full sourcing and per-charm confidence level. Summary: `prefer-ipv6` (confirmed real option, additive dual-stack HAProxy binding, confirmed via `charm-nova-cloud-controller`'s actual "Dual Stack VIPs" commit) + appended v6/GUA entries on the existing @@ -241,7 +241,7 @@ config <app>` confirmation once first deployed. 2. **Octavia's `lb-mgmt-net` IPv6 support is a real, open risk, not resolved.** The overlay deliberately excludes an Octavia entry -- - `docs/dc-dc-ipv6-charm-research.md` section 6 documents two real, + `docs/archive/dc-dc-ipv6-charm-research.md` section 6 documents two real, still-referenced upstream Launchpad bugs (#1911788, #1913409) about IPv6 lb-mgmt-net failures. Decide explicitly before this stage's Ceph-over-v6/geneve-over-v6 gate is declared closed: accept the risk diff --git a/runbooks/dc-dc-phase5-dr-failover-drill.md b/runbooks/dc-dc-phase5-dr-failover-drill.md index 557c11b..6308b52 100644 --- a/runbooks/dc-dc-phase5-dr-failover-drill.md +++ b/runbooks/dc-dc-phase5-dr-failover-drill.md @@ -1,6 +1,6 @@ # DC-DC Phase 5 -- DR wiring and failover drill (Stage 6) -Wire the Tier-A cross-DC replication mechanism (`docs/dc-dc-replication-DR-seed.md`, +Wire the Tier-A cross-DC replication mechanism (`docs/archive/dc-dc-replication-DR-seed.md`, engineered by **D-108**) between DC1 and DC2, stage it one-way, prove it clean, enable two-way, then run the actual failover and failback drills plus the per-DC Juju controller backup/restore drill (**D-104**). This is the first DC-DC runbook that deliberately takes a @@ -59,7 +59,7 @@ Tooling gap register audit): - **No `cinder-backup` charm is deployed in either DC's bundle.** The DR seed's own text - (`docs/dc-dc-replication-DR-seed.md` line 24) assumes cinder already exposes the + (`docs/archive/dc-dc-replication-DR-seed.md` line 24) assumes cinder already exposes the `backup-backend` endpoint -- true for `cinder` itself, but the separate `cinder-backup` subordinate/charm that actually runs the backup driver is NOT in `bundle.yaml` today. Adding it is a bundle change (a built-surface edit, gated the same as any other bundle change) -- @@ -77,7 +77,7 @@ **UPDATE 2026-07-10 (DOCFIX-167):** `bundle.yaml` ITSELF now carries `cinder-backup`, `ceph-rbd-mirror`, and the corrected `rbd-mirror: replication` binding (see -`docs/changelog-20260709-designate-cinderbackup-rbdmirror.md`) -- so any DC deployed via +`docs/archive/changelogs/changelog-20260709-designate-cinderbackup-rbdmirror.md`) -- so any DC deployed via Stage 5's runbook AFTER this date gets all three from the initial `juju deploy`, and Step 1's CHECK below should find them already correct (skip straight to Step 3). Step 2 remains here, unmodified, for the case this runbook is run against a DC deployed from an OLDER bundle.yaml diff --git a/runbooks/dc-dc-phase6-designate-cos-magnum.md b/runbooks/dc-dc-phase6-designate-cos-magnum.md index 813dc32..cd1a364 100644 --- a/runbooks/dc-dc-phase6-designate-cos-magnum.md +++ b/runbooks/dc-dc-phase6-designate-cos-magnum.md @@ -258,7 +258,7 @@ were NOT individually re-fetched for designate -- confirm with `juju deploy --dry-run` before treating this as fully final. -See `docs/changelog-20260709-designate-cinderbackup-rbdmirror.md` for the +See `docs/archive/changelogs/changelog-20260709-designate-cinderbackup-rbdmirror.md` for the full record. What's still genuinely open: the `os-public-hostname` overlay (Step 1.4 above) -- `bundle.yaml` deliberately stays DC-agnostic, so that overlay is still this runbook's own job at real execution time, not diff --git a/runbooks/dc-dc-teardown-rollback.md b/runbooks/dc-dc-teardown-rollback.md index e8527c3..a5903c0 100644 --- a/runbooks/dc-dc-teardown-rollback.md +++ b/runbooks/dc-dc-teardown-rollback.md @@ -43,7 +43,7 @@ > below still applies -- but the `maas-vm-host` record now points at `vvr1-dc0`'s inner virsh, so > remove it BEFORE destroying `vvr1-dc0`. > - **Reverting Model B -> Model A** (a different operation from teardown) is -> `docs/model-a-fallback-plan.md` (git tag `model-a-fallback`), not this runbook. +> `docs/archive/model-a-fallback-plan.md` (git tag `model-a-fallback`), not this runbook. > > Until the body below is rewritten to the two-root shape, treat its DC-substrate steps as > **single-root reference** and apply the mapping above. The Office1 (`voffice1`) and mesh/pool steps diff --git a/runbooks/ops-update-procedure.md b/runbooks/ops-update-procedure.md index c698201..916c661 100644 --- a/runbooks/ops-update-procedure.md +++ b/runbooks/ops-update-procedure.md @@ -383,7 +383,7 @@ (This is exactly the appendix-B "refresh the table on a successful validated state" event.) 2. Commit the post-change `asbuilt/<ts>/` BOM. -3. `docs/v1-redeploy-changelog.md`: as-executed addendum -- what moved +3. `docs/archive/v1-redeploy-changelog.md`: as-executed addendum -- what moved (controller x.y.z -> x.y.z', per-app rev table), why, and the revert table (per-app `--revision` + `--channel` re-track pairs; controller = none in-band, D-070 posture). diff --git a/runbooks/phase-04-network-carve.md b/runbooks/phase-04-network-carve.md index dc897b5..b6a254e 100644 --- a/runbooks/phase-04-network-carve.md +++ b/runbooks/phase-04-network-carve.md @@ -177,7 +177,7 @@ the cloud -- draft <- live): 10.12.4.101-10.12.4.110 (subnet 1, provider) + 10.12.8.101-10.12.8.110 (subnet 2, metal), both "mgmt-plane reserved" (10 IPs each). Both sit OUTSIDE the FIP pool (10.12.5.0-10.12.7.254) and the VIP /26 blocks -> no conflict with - provider-ext-fip. FOLD into docs/netbox-vip-queue.md + D-003 in the docs sub-pass (purpose + provider-ext-fip. FOLD into docs/archive/netbox-vip-queue.md + D-003 in the docs sub-pass (purpose annotation pending operator confirmation; do NOT mutate NetBox until IPAM design is confirmed -- D-010). - Transitional note: MAAS already carried the front-loaded VIP reservations (.2-.63 provider + .8.2-.63 metal; old D-020 .8.224-.254 gone) ahead of the bundle's interim diff --git a/scripts/opentofu-validate.sh b/scripts/opentofu-validate.sh index ecbee79..cf03b4c 100644 --- a/scripts/opentofu-validate.sh +++ b/scripts/opentofu-validate.sh @@ -42,7 +42,7 @@ # interpreted in libvirt's DEFAULT unit (KiB), not MiB (the 0.8-era meaning), so # `memory = 2048` silently yields a 2 MiB guest that triple-faults in its # bootloader. That cost a full session on 2026-07-12 (Office1 OPNsense edge; see -# docs/incident-20260712-opnsense-edge-boot-triplefault.md). This is the guard +# docs/archive/incident-20260712-opnsense-edge-boot-triplefault.md). This is the guard # that stops the next VM module from reintroducing it. # # Block-scan is safe because `tofu fmt -check` (below) enforces canonical layout: diff --git a/scripts/repo_lint.py b/scripts/repo_lint.py index 0e4ea81..1f31e2a 100644 --- a/scripts/repo_lint.py +++ b/scripts/repo_lint.py @@ -20,7 +20,7 @@ (collision guard). Next-free numbering is ledger-scan's job ALONE (DOCFIX-107: the old L5 next-free print disagreed with it via three independent defects; see - docs/repo-lint-nextfree-bug-FINDING.md). + docs/archive/repo-lint-nextfree-bug-FINDING.md). L6 bare invoke runbook lines executing scripts/*.sh without a bash/source prefix (repo carries NO exec bits -- DOCFIX-069; bare form fails "Permission denied" on a fresh clone).