| 2026-08-07 |
substrate dc0: apply tailscale VM (2 added) + pin its MACs pre-enlistment
...
Gated apply on voffice1: vr1-dc0-tailscale-01 created (domain+disk, 0 change/
0 destroy), powered off. MACs captured via virsh domiflist and pinned in
vr1_dc0_node_nics order before enlistment (MAC-regen trap). Post-pin tofu plan
must show No changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
substrate: add the .7 Tailscale subnet-router VM to both DC node maps
...
opentofu/vr1-dc{0,1}-substrate/main.tf gain "vr1-dc{0,1}-tailscale-01"
(2 vCPU / 2 GiB / 25 GiB, macs=[]), the dedicated per-DC Tailscale subnet
router at utility .7 (D-129(iii) amendment 2026-08-07). for_each keyed ->
expect 1 add / 0 change / 0 destroy per DC.
Capacity re-gated FIT (dc-dc-whole-host-budget.py at overhead 34/14 ->
874/1024 = 85% RAM, 150 GiB headroom). tofu validate PASS (both substrate
roots + all modules). repo-lint 0 fail.
NOT YET APPLIED -- the plan/apply run on voffice1 (inner substrate root,
D-128) and each is individually gated (verify 1 add/0/0, then pin MACs
pre-enlistment, MAAS commission/deploy, carve .7 metal-admin+provider-public).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

D-129(iii) amendment: per-DC Tailscale operator-access rulings (a-d), both DCs
...
Record four GA-R5 rulings (2026-08-07) that pull the gap-21 per-DC Tailscale
subnet-router forward to close phase-03 Horizon properly (operator: "pull the
tailscale steps forward"; "plan and push to both DC0 and DC1"):
(a) dedicated VM at utility .7 (10.12.8.7 / 10.12.68.7)
(b) STAR -- operator->DC only (the Headscale ACL / security boundary)
(c) SINGLE router, HA scale-up PINNED
(d) SNAT ON now, source-IP preservation PINNED
These are D-129(iii) implementation sub-decisions (not a new D-number). Also:
correct the D-107 citation defect (D-107 is airgap/mirror/NTP, governs no
Tailscale; D-129(iii) governs); extend D-134's standing octet map to .7;
update the gap-21 register row and CURRENT-STATE (phase-03 does NOT close this
session -- Step 3.3 Horizon splits to its own gate row, gated on the tailnet
build + vault CA on the workstation + a browser login over the tailnet).
Records only -- no code, no cloud change. The build (site-tailscale.sh + the
.7 VM per DC + Headscale star ACL/autoApprovers + SEC row) is next.
repo-lint 0 fail; Decision C reconciliation measured read-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-06 |

phase-03 Step 3.4 (dc0): G3 domain-manager probe -- named check + live PASS
...
Build scripts/g3-domain-manager-probe.sh + tests/g3-domain-manager-probe (harness
12/12, every exit path proven) as the GA-R6 named executable check for phase-03
Step 3.4 stage 2 (gate G3), filling the hard-rule-4 gap (was a manual runbook walk
only) and giving dc1's Step 7 a reusable probe. Grounded in the real policy
(domain-manager-policy.yaml:103 create_grant managed-role guard).
Ran it LIVE from the dc0 rack (operator-approved): G3 PASS, 7 ok / 0 fail, teardown
verified clean (zero g3-* residue). Stage-1 (PO:) verified read-only: policyd-override
attached, all 3 keystone units 'PO: Unit is ready'. Capture
docs/audit/g3-dc0-probe-20260806.txt.
HARNESS-MANIFEST recorded 99->100; gauntlet ALL GREEN (100), repo-lint 0 fail.
CURRENT-STATE: Step 3.4 -> RESOLVED; phase-03 exit gate now turns on the Horizon
reachable/login-works item (D-044/D-075 per-rebuild + VR0-nginx vs VR1-tailnet access
model = Decision C), to be measured + ruled next.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Session close 2026-08-06 (part 3, GA-R4 bookend): phase-03 core-verify F-CV1/F-CV2/F-CV3 RESOLVED + binding conformance
...
GA-R4 bookend: bounded ledger summary + sweep. F-CV2 (openstack CLI on the dc0
rack), F-CV1 (designate binding, BUNDLEFIX-056), F-CV3 (dashboard TLS via D-072
AMENDMENT VR1 / BUNDLEFIX-057) all RESOLVED; BUNDLEFIX-058 designate-stack
conformance -> binding conformance clean cloud-wide. D-134 Roosevelt-delta +
gap-21 access-model context recorded.
Sweep O11 FIRST SURFACE (dc0 MAAS query method); O10 rack repo-stage bundle STALE
(re-stage before redeploy). Memory instrument-currency #18. repo-lint 0-fail,
gauntlet ALL GREEN (99) at ace0e16, ledger 292 lines (no rotation).
NEXT: Step 3.4 (domain-manager policy) -- last phase-03 exit-gate item -> Steps 8-12
-> Stage-5 exit. Body: docs/changelog-20260806-phase03-coreverify.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-058: designate-stack binding conformance (amqp/cluster -> metal-internal) + ceph-rbd-mirror Stage-6 watch
...
Operator-requested binding-conformance sweep after F-CV1/F-CV3. The only remaining
ACTIVE svc-endpoint deviations were the designate stack: designate amqp (->rabbitmq)
+ cluster (peer), designate-bind cluster (peer) -- all on the metal-admin default,
deviating from the generic svc-to-svc rule (14 apps use metal-internal). No ruled
exception -> conformance repair, no new D-number.
Benign (designate served metal-admin so nothing broke, unlike F-CV1/F-CV3) but
off-plane. Repaired to metal-internal so dc1 inherits a fully-conformant bundle.
Proven live (same method): juju bind designate amqp=metal-internal cluster=metal-internal
+ juju bind designate-bind cluster=metal-internal (individually gated). Verified no
regression: bindings moved; designate still Stage-7-blocked on nameservers ONLY;
designate-bind active/idle; designate haproxy 0 DOWN (F-CV1 intact). No harness asserts
these bindings. gauntlet ALL GREEN, repo-lint 0-fail.
RECORDED: ceph-rbd-mirror certificates/cluster on metal-admin are UNBOUND today
(Stage-6 DR not wired) -> sweep O9 STAGE-6 WATCH + CURRENT-STATE, pre-check before
wiring at dc-dc-phase5. No other active deviations found cloud-wide.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-057 / D-072 AMENDMENT (VR1): dashboard cluster -> metal-internal (F-CV3 RESOLVED)
...
VR1 split-metal INVERTS the D-072 cluster placement. The openstack-dashboard charm
declares no admin/internal extra-binding and no os-*-network (metadata + charmhub
docs), so apache serves its SSL vhost on metal-internal while haproxy dialed
cluster=metal-admin -> vhost-less -> plaintext (the D-072 trap, inverted). Option A
(serve metal-admin) unavailable in-deployment (no charm lever).
Retire the VR0 exception: openstack-dashboard cluster -> metal-internal (generic rule
+ the served plane). Proven LIVE before ratifying (operator process directive), then
ratified GA-R5 "Ratified, land the config-of-record".
bundle.yaml cluster=metal-internal; design-decisions D-072 AMENDMENT (VR1);
binding-reference matrix + exception RETIRED + cross-ref; CURRENT-STATE/sweep/changelog
F-CV3 RESOLVED. Verified live: provider + operator metal-admin VIPs both TLS 200
CA-verified (were plaintext); cert covers all 3 VIP IPs (resolves AH01909); units
active/idle. No harness asserts this binding. gauntlet ALL GREEN, repo-lint 0-fail.
dc1 inherits via the shared bundle (no rebind).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Gap-21: capture the operator-access model (tailnet -> metal-admin dashboards) as forward context
...
Operator discussion 2026-08-06: operators reach each DC's metal-admin plane over
the tailnet and access routed dashboards (Horizon) from there. Added as dated
CONTEXT (not a ruling) to workflow gap-register item 21, informing the four
deferred Tailscale sub-decisions when re-raised. Captures the implications that
reach outside the Tailscale build: the operator-facing reverse proxy becomes
unnecessary; D-044/D-075 per-rebuild accommodations sunset for the operator path;
cert SANs on the metal-admin VIP (and FQDN certs, D-106/D-008) become load-bearing;
and it reinforces metal-admin as the dashboard's serving plane (the access-model
argument for F-CV3 option A over option B). Cross-referenced from the F-CV3 sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
D-134: Roosevelt-delta annotation -- LXD container addresses are auto-picked, not carved
...
Logs the operator-requested Roosevelt-delta observation from the phase-03
core-verify thread. D-134 pins node statics (per-role octet bands) and the VIP
band (.50-.99 reserved); the OpenStack control plane runs in Juju LXD containers
whose addresses are NOT pinned -- MEASURED 2026-08-06 as MAAS device interfaces in
mode=static but auto-picked from the .100-.200 pool independently per subnet, so
container octets float across planes and are not redeploy-stable. Harmless in VR1
(clients use VIPs; certs reissue to cover whatever a unit holds). Roosevelt-delta:
a bare-metal build wanting predictable container addressing would add a
per-container static reservation (a "container carve"). GA-R3: observation, no new
D-number -- annotated onto D-134.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-056: designate public/internal bindings -- F-CV1 RESOLVED (fix executed + verified)
...
designate's REST API public+internal endpoints were omitted from the bundle,
defaulting to the '' metal-admin fallback -> orphaned the provider + metal-internal
legs of its ruled .62 VIP triple (D-020 amendment) -> _admin haproxy backend SSL-DOWN
on the unserved metal-internal address. The deployed charm declares public/admin/
internal extra-bindings (metadata verified); dnsaas (D-106) is ADDITIONAL, not a
replacement -- the prior "no public binding" reading was the defect's root.
Config-of-record: bundle.yaml +public:provider-public +internal:metal-internal;
provider-bundle-check EXPECT_PUBLIC_VIP 11->12 (vault stays out, not 13) with T16c/T16d
failing-direction tests (57->59, ALL PASS); network-space-binding-reference row 88 +
note. Gauntlet ALL GREEN (99); repo-lint 0-fail.
Live (operator-approved): juju bind designate public=provider-public
internal=metal-internal (rc=0, dc0 rack). Verified: full haproxy sweep 0 DOWN
cloud-wide; apache https vhosts span all 3 planes; cert reissued for provider-public;
catalog triple correct (public 10.12.4.62 / internal 10.12.12.62 / admin 10.12.8.62).
Governing: D-052 / D-020 amendment. Evidence:
docs/audit/stage5-dc0-phase03-coreverify-20260806.txt. Body:
docs/changelog-20260806-phase03-coreverify.md (Item 6).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F-CV1: research resolves :public -- fix BOTH endpoints (D-020 triple + charm supports it)
...
Operator asked to investigate designate's tenant surface + charm docs before
ruling on :public. Conclusion: fix BOTH :public and :internal.
Evidence: (1) D-020 AMENDMENT (2026-07-27, operator GA-R5) gave designate the
established provider/admin/internal triple; that ruling implies the bindings, so
the current metal-admin-fallback bundle is a conformance defect. (2) The DEPLOYED
charm-designate metadata declares extra-bindings public/admin/internal (read in
full on the unit -- authoritative; a WebFetch on master wrongly said none, a
small-model error). Charm description: "Multi-tenant ... REST API" -> tenant-facing
by design. (3) Every sibling binds public->provider-public + internal->metal-internal;
designate is the lone deviation, no ruled exception.
Feasibility confirmed: all 3 units hold provider-public + metal-internal addresses
(juju bind not refused); cert covers metal-internal, will reissue for provider-public.
Proposed gated fix (D-072/BUNDLEFIX-011 precedent): bundle bindings
+public:provider-public +internal:metal-internal, live juju bind, verify by
haproxy readback + cert SAN re-read + haproxy sweep. Awaiting operator go-ahead.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F-CV1: reconcile after research -- :internal defect CONFIRMED, :public OPEN; D-141->D-052
...
Prior-art + governing-decision research on the designate bind-plane mismatch.
Corrects the earlier commit (aaeee93) which reached "CONFIRMED" and cited D-141
before checking the governing decisions (backwards ordering, memory #17).
D-072/BUNDLEFIX-011 prior art: the dashboard-plaintext class was root-caused in
VR0 (charm renders haproxy 443 backend on cluster-binding addr, apache SSL vhosts
only for default+public). dc0 bundle already carries that fix -> F-CV3 is a NEW
cause, parked for its own triage.
F-CV1 re-scoped: governing surface is D-052 + the generic binding rule, NOT D-141.
`:internal`->metal-internal is a CONFIRMED defect (on the metal-admin fallback;
every sibling binds it to metal-internal; no ruled exception; cert already covers
the metal-internal SANs; all 3 units have metal-internal addrs -> juju bind won't
be refused). `:public`->provider-public is OPEN -- designate uniquely carries
`:dnsaas` on provider-public (D-106 dual-VIP); its REST API may be intentionally
metal-admin-only. The per-app table is generated from bundle.yaml (descriptive of
the defect), not intent.
Fix method = D-072 precedent (bundle + live juju bind + haproxy-readback verify).
designate is Stage-7-blocked -> no urgency. Awaiting operator ruling on :public.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 Step 7: triage F-CV1 (CONFIRMED designate bind-plane mismatch) + F-CV3; sweep
...
Operator-authorized triage of the two plaintext-vs-TLS findings (read-only; not
fixed, hard rule 1). Both apps have vault certs rendered + the certificates
relation, so NOT the ovn CN-issuance class -- charm apache-TLS-frontend layer.
F-CV1 CONFIRMED (two findings, not one): designate/0 apache https frontend binds
only 10.12.8.198:8991 (metal-admin); haproxy's _admin backend dials
10.12.12.110:8991 (metal-internal) where no SSL vhost exists -> check-ssl hits
plaintext -> DOWN. VR1 dual-metal-plane bind mismatch (D-141), structural (not
Stage-7 collateral). F-CV3 (dashboard) is separate: :433 served by Ubuntu
default-ssl.conf, charm https frontend not effective; root cause not nailed.
Owned instrument caveat: earlier "no SSLEngine in sites-enabled" was a grep -r
false negative (does not follow the symlinks); apache2ctl -S corrected it.
Remediation = a focused, gated session. Sweep:
docs/audit/queued-findings-20260806-phase03-coreverify.txt
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 Step 7: phase-03 core-API VERIFIED; F-CV2 openstack-client on rack; F-CV1/F-CV3 logged
...
Core-API layer of phase-03 core-verify (adapted for vr1-dc0, run from the dc0
rack per D-138) VERIFIED read-only: settle walk (falsifiable prediction matched),
haproxy 0-DOWN across 12 activated VIP apps, admin-openrc scoped token, IP-only
endpoints, two-sourced keystone VIP, vault CA TLS OK.
F-CV2 RESOLVED (07-30 queued-F1, hit at Step 7): installed python3-openstackclient
6.6.0-0ubuntu2 on the dc0 rack; CURRENT-STATE section 7 row amended (GA-R1/C1).
Exit gate NOT MET (2 open, GA-R6/E3 no conditional close): F-CV3 dashboard VIP
serves plaintext (apache-SSL-inactive despite certs) -> Horizon exit-gate fails;
Step 3.4 domain-manager policy NOT RUN. F-CV1 designate-api plaintext vs haproxy
check-ssl (same shape); "collateral of block" reading retracted. Logged not fixed.
Evidence: docs/audit/stage5-dc0-phase03-coreverify-20260806.txt
Body: docs/changelog-20260806-phase03-coreverify.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Session close 2026-08-06 (part 2, GA-R4 bookend): F4 + dc-ha-scaleup retired + memcached 1->3 live
...
GA-R4 bookend for the F4/retire/memcached session:
- Bounded ledger summary (SESSION CLOSE 2026-08-06 (part 2)); ledger rotated (08-02 x2 ->
docs/archive/session-ledger-rotated-20260806.md), 299 -> 280 lines (under the 300 cap).
- Sweep docs/audit/queued-findings-20260806-postwave-retire-memcached.txt -- 5 FIRST SURFACE,
led by F-1 (operator-flagged): designate coordination points at ONE memcached unit,
Stage-7 activation re-check. Always-sweep-5 incl. the as-executed-log gap; 4 OWNED items.
- Memory instance #17 (instrument-currency-before-negatives) persisted outside the repo.
Gauntlet ALL GREEN (99); repo-lint 0 fail (1 legacy L1 warn); ledger-scan reconciled
(4 open decisions, SEC 29 unchanged; DOCFIX 213 / BUNDLEFIX 056 next-free as assigned).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: memcached scaled to 3 LIVE (operator-approved) + consumer verify
...
Live scale-up executed on the dc0 rack (operator "Both approved"):
juju add-unit memcached -n 2 --to lxd:1,lxd:2 -m vr1-dc0 (exit 0).
- CONVERGED 3/3 active/idle, one per control node (memcached/0 control-01, /1 control-02,
/2 control-03; all "Unit is ready and clustered", 11211/tcp). Clean install, no F5 apt-wedge.
Capture docs/audit/stage5-dc0-memcached-scaleup-20260806.txt.
- Consumer verify (D-121 verify-at-deploy): nova-cloud-controller sees all 3 servers
(memcache_servers = .115,.166,.165:11211). designate coordination backend_url shows ONE
(memcached/0) -- designate is workload-blocked pre-Stage-7, so re-check at designate
activation; recorded, not a scale-up defect.
- bundle.yaml re-staged to the dc0 rack (42845edb == repo HEAD).
- CURRENT-STATE reconciled: memcached config=3 AND live=3.
Record-only commit (the live mutation itself was the add-unit above). repo-lint 0 fail.
Revert of the LIVE state (if ever needed): juju remove-unit memcached/1 memcached/2 -m vr1-dc0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: memcached 1->3 config-of-record [BUNDLEFIX-055 + DOCFIX-212]
...
Operator ruling 2026-08-06, exact utterance "3 units (restore intent)" -- restoring the
2026-07-31 memcached->3 direction that BUNDLEFIX-053 never folded and the retired
dc-ha-scaleup.yaml had carried unrealized. Found while reconciling a stale comment:
- Full merge-diff (measured) of the archived overlay vs bundle.yaml showed the ONLY divergence
was memcached (overlay num_units:3 / base 1) -- BUNDLEFIX-053 folded the HA chain but not this.
Corrects two earlier misstatements (this session): "memcached=3 in bundle.yaml" (it was 1) and
the retirement's "wholly redundant / all-keys no-op" claim (it diverged on memcached).
- BUNDLEFIX-055: bundle.yaml memcached num_units 1->3, to:[lxd:0]->[lxd:0,1,2] + the
3-independent-caches rationale (no hacluster/VIP; clients hash across the full server list).
This completes the fold, making the archived overlay genuinely, fully redundant.
provider-bundle-check 57/0 (memcached is orthogonal to the arity/VIP gates).
- DOCFIX-212: design-decisions D-121 "Left single (NOT scaled)" list -- memcached removed,
recorded as =3.
- CURRENT-STATE: records the CONFIG=3 / LIVE=1 divergence -- a live add-unit (operator wants it
deployed live) or a teardown+redeploy reconciles it. The live scale-up is a separate gated step.
Gauntlet ALL GREEN (99 harnesses); repo-lint 0 fail (1 pre-existing legacy L1 warn).
Owed: re-stage the functionally-changed bundle.yaml to the dc0 rack (gated); live add-unit to 3.
Revert: bundle.yaml memcached back to num_units:1 + to:[lxd:0]; revert the DOCFIX-212 D-121 line
and the CURRENT-STATE gap note.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: RETIRE + archive dc-ha-scaleup.yaml (operator-ruled; R6 superseded) [DOCFIX-211]
...
Operator ruling 2026-08-06, exact utterance "Retire the redundancy and archive." The
overlay was wholly redundant -- BUNDLEFIX-053 folded its entire content into bundle.yaml,
so applying it was an all-keys no-op deep-merge (Task-#1 post-wave review, measured).
- Archived: git mv overlays/dc-ha-scaleup.yaml -> docs/archive/dc-ha-scaleup-RETIRED-20260806.yaml
+ a do-not-deploy banner. docs/archive/ chosen explicitly (no prior overlay-archive precedent).
- design-decisions.md: GA-R5 SUPERSESSION note on D-121's R6 (original 2026-07-27 record retained
as history); fixed :503 and :4563 "overlay encodes v-a" refs -> bundle.yaml/BUNDLEFIX-053.
- Harness (tests/provider-bundle-check): T17 REMOVED (== T18 once vault-hacluster is in base);
T17b/T32/T33/T34 RE-POINTED onto good.yaml (base HA chain). T33/T34 now use the deterministic
mutate() helper, not sed-on-overlay (no silent no-op fixture; a wrong key path KeyErrors). Each
FAIL case proven to fire against the real checker message. 58->57 cases; provider-bundle-check 57/0.
- Runbook dc-dc-phase4 Step-4 note + Step-12.3(a) gate: the two-phase "base at cluster_count:1 then
apply the overlay to reach 3" model RETIRED -- base deploys 3-unit HA directly, so cluster_count:3
is the expected post-Step-4 value.
- Rendered vips overlays: edited the render SOURCE (render/values/vr1-dc{0,1}-vips.yaml, D-136) and
re-rendered (not hand-edited); overlay diff = the two comment lines only; render-drift 4/0.
Hand-maintained vr1-dc{0,1}-machines.yaml deploy-command comments + provider-bundle-check.py +
cloud-assert.sh comments + dc-dc-deployment-workflow.md item 22 reconciled.
- CURRENT-STATE: no edit needed -- active status already omits the HA overlay from the deploy input;
remaining mentions are dated historical narrative (retained as history, GA-R1).
- Also finalizes Item 3/5 changelog notes (F9 approved+staged; Task #1 recommendation ruled).
Gauntlet ALL GREEN (99 harnesses); repo-lint 0 fail (1 pre-existing legacy L1 warn).
Staging: dc0 rack vr1-dc0-vips.yaml now stale-by-one-comment vs the re-render (functionally
identical); re-syncs + sha-verifies at the next deploy per D-138 (logged, no rack write for a comment).
Revert: git mv the overlay back (strip banner) + git revert this commit + re-render vips from
reverted values.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F-A/BUNDLEFIX-054: correct bundle.yaml HA-chain header + session changelog
...
- bundle.yaml:23 HA-chain header was stale after BUNDLEFIX-053 (vault-hacluster),
the D-020 amendment (vault metal-only VIP) and R11 (designate VIP). Corrected FROM
the measured merged dc0 deploy input: 12 -> 13 hacluster subordinates/:ha relations;
ALL 13 carry a VIP (12 the full provider+metal-admin+metal-internal triple, vault a
METAL-ONLY pair per D-020); retired the stale "designate has no HAProxy VIP" note
(designate gained triple VIP .62 under R11); clarified rabbitmq is the 14th HA app but
native-clustered (not in the 13). COMMENT-ONLY -- no deploy semantics change.
- Session changelog docs/changelog-20260806-stage5-dc0-f4-postwave.md (GA-R2):
F4 (14/14 measurement-backed + vault ha_enabled), F8 (ceph-radosgw resolved),
F9 (staging drift measured, re-stage gated), F-A, and the Task #1 finding that
dc-ha-scaleup.yaml is now redundant with bundle.yaml (recommend retiring; operator
ruling owed since R6/D-121 reference it -- logged, not executed).
Gauntlet ALL GREEN (99); repo-lint 0 fail (1 pre-existing legacy L1 warn).
Revert: git revert this commit (restores the prior 6-line header; changelog is additive).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F4: 14/14 HA now measurement-backed + vault ha_enabled MEASURED
...
Live read-only capture from the dc0 rack (juju status -m vr1-dc0, D-138 path):
docs/audit/stage5-dc0-juju-status-14of14-20260806.txt.
- 14/14 D-121-enumerated HA apps at scale=3 (per changelog-20260805-d121-ha-scaleup.md:24).
HONEST SPLIT (GA-R6 E3, not rounded): 12 active/idle; 2 blocked-at-scale=3 on KNOWN
non-HA items -- octavia (configure-resources) + designate (nameservers). Model-wide
152 active/idle, 7 blocked, 1 unknown (gss, normal). Converts the CURRENT-STATE
14/14 line from OPERATOR-ATTESTED to MEASUREMENT-BACKED (C2: measurement wins).
- F4 vault ha_enabled MEASURED: vault status HA Enabled=FALSE x3, storage=mysql,
unsealed, shared cluster id. CAUSE (measured, corrects a mid-session mis-blame of the
mysql backend): charm renders storage "mysql" with no ha_enabled and exposes no such
option (only vip + dns-ha-access-record) -- HA is charm/VIP model, not vault-native.
CONFIRMS D-121 (v-a) as-built; vault-native/Raft HA stays owned by D-068. No new decision.
- F8 ceph-radosgw RESOLVED: 3 units active/idle "Unit is ready" (prior stale-status converged).
- F9 staging drift MEASURED (re-stage gated, owed at next deploy): dc0 rack bundle.yaml +
vr1-dc0-vips.yaml STALE vs HEAD; dc1 rack has no ~/repo-stage (HELD).
Revert: git revert this commit (record-only; no live-cloud change was made -- all reads).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Session close 2026-08-06 (GA-R4 bookend): Task #2 D-020 vault-metal-only DONE + sweep
...
Bookend for the Task #2 session (deliverable commits 07e2d9d, a08783c already pushed).
- docs/session-ledger.md: bounded <=15-line SESSION CLOSE 2026-08-06 block (ledger now
299 lines, one from the 300 cap -- next close needs rotation).
- docs/audit/queued-findings-20260806-task2-vault-metal-only.txt: the sweep. FIRST SURFACE
items: bundle.yaml:23 stale "12 charms" (F-A); rack ~/repo-stage overlays now stale (F-B);
Task #2 guidance unexecutable as written (F-C); render-drift red at part-2 close but
undocumented (F-D). Owed: F4/F8/F9/Task-1. Owned: W-1/W-2/W-3.
Gauntlet ALL GREEN (99); repo-lint 0 fail. Memory: instrument-currency instance sixteen
(appended locally, not repo-tracked). Status ONLY in docs/CURRENT-STATE.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Task #2: commit D-020 vault-metal-only amendment; teach renderer + gates; harnesses green
...
The RULED-but-uncommitted D-020 amendment (vault METAL-ONLY, 2026-08-05) is now
committed and enforced end-to-end. Discovered mid-task that the overlays are RENDERED
from render/values/*.yaml (D-136) and hand-editing them is forbidden by the render-drift
gate -- which was ALREADY red at the part-2 close (undocumented) from that session's dc0
overlay hand-edit. Resolved by teaching the renderer, not by hand-editing (advisor-reviewed;
the amendment is ruled, so this is OPS implementation).
Vault VIP shape is now the metal PAIR (metal-admin + metal-internal, no provider-public,
no v6) across all four consumers, each with a failing-direction fixture:
- provider-bundle-check.py: per-DC/family-aware vault exception; STILL band- and
octet-uniqueness-checked (advisor caught the first draft's early `continue` disarming
octet_owner for .61 -- proven rc=0 draft / rc=1 fixed). T54/T55/T56.
- render-dc-overlays.py: name-keyed metal-only render branch (drops provider + v6);
overlays re-rendered (diff vs HEAD = exactly the one vault line each). T15b.
- pre-flight-checks.sh CHECK 1: awk name-tracking + vault metal-pair branch. T28b.
- render/values/*.yaml: vault comment (part of rendered bytes).
D-121 status 12/14 -> 14/14 in CURRENT-STATE, with the missing juju-status capture
DECLARED as an owed gap (GA-R1 rule 2) rather than papered over.
Verify: gauntlet ALL GREEN (99); repo-lint 0 fail; provider-bundle-check 58/0,
render-dc-overlays 24/0, render-drift 4/0, pre-flight-checks 32/0.
Body: docs/changelog-20260805-task2-vault-metal-only-commit.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-05 |
Session close 2026-08-05 (part 2): D-121 HA 14/14 + vault metal-only; GA-R4 bookend
...
Bounded ledger summary + rotation (07-31/08-01 -> archive/session-ledger-rotated-20260805.md,
315->288) + sweep queued-findings-20260805-d121-ha-vault.txt (9 first-surface). D-121 executed
live -- all 14 control-plane apps to 3-unit HA; D-020 amended vault->metal-only (ratified, HELD
uncommitted pending the provider-bundle-check harness reconcile, Task #2). Status ONLY in
CURRENT-STATE.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
CURRENT-STATE: D-121 HA scale-up now 12/14 (nova-cc, rabbitmq, barbican DONE)
...
Measured 2026-08-05: nova-cloud-controller + rabbitmq-server (native erlang 3-node cluster)
+ barbican scaled to 3-unit HA. 12 of 14 HA apps at 3-unit HA. Only keystone (cloud-wide
auth VIP blip) and vault (operator unseal) remain -- both operator-gated. vault HA also
resolves the barbican-vault secrets-storage dependency.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-053: bundle.yaml to full 3-unit HA (D-121); reverse BUNDLEFIX-002 vault de-HA
...
Reverses the D-009/BUNDLEFIX-002/003 decorative-single-unit posture to real 3-unit HA per
operator ruling 2026-08-05 ("superseded the HA removal from all items ... stand up all
remaining HA apps ... HA blocks sign off"; "testing as if Roosevelt"). Executes adopted
D-121 incl. its (v-a) vault sub-ruling.
bundle.yaml:
- num_units 1->3 + to:[lxd:0,1,2] on 14 HA apps (keystone glance nova-cc placement
neutron-api cinder ceph-radosgw dashboard octavia barbican magnum designate rabbitmq vault)
- hacluster cluster_count 1->3 x12
- vault-hacluster (cluster_count:3) + [vault:ha, vault-hacluster:ha] uncommented -- the
BUNDLEFIX-002 reversal (D-121 (v-a): vault MySQL-backed HA)
- rabbitmq min-cluster-size:3 (D-009 amendment)
- governing comments reconciled (D-009/BUNDLEFIX-002/003 -> D-121 built)
overlays unchanged (VIPs already set).
Every change verified against charm/juju docs or the deployed charm (not assumed):
hacluster cluster_count default 3; rabbitmq min-cluster-size; vault ha endpoint=metal-internal
like keystone; vault MySQL-backend HA per HashiCorp docs (ha_enabled+lock_table). Validated:
provider-bundle-check.py --dc vr1-dc0 PASS (13 hacluster subs cluster_count==num_units; 13
principals carry a VIP; 109 relations well-formed); repo-lint 0 fail.
Live: Wave 1 (8 apps) scaled to 3-unit HA (see CURRENT-STATE + changelog). Waves 2-4 and the
post-wave bundle/overlay reconciliation (pinned task) remain. Body:
docs/changelog-20260805-d121-ha-scaleup.md. Status in docs/CURRENT-STATE.md ONLY.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
CURRENT-STATE: barbican-vault/0 verify result (C2) -- genuine incomplete, missing vault_url
...
Item (b) verify supersedes the "settling" note in 71c5b97: barbican-vault/0 is NOT settling
and NOT deferred. vault provided per-unit role_id + token but the secrets-storage databag
lacks vault_url (app-data empty) -> charm logs "Requesting access to vault (None)". Not
deploy-blocking (barbican/0 active on software backend). Measured 2026-08-05 17:55Z.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: remediate ceph-mon/2 + ceph-radosgw/0 "allocating" (apt CLOSE-WAIT); durable capture
...
Live-cloud, operator-gated (2 exchanges). Two units stuck `allocating` were root-caused
NOT to a down proxy (both dc0 proxies PASS; security.ubuntu.com 200/0.67s) but to
cloud-init `apt-get update` parked in CLOSE-WAIT to 10.12.8.4:3142 for ~12.5h (apt has no
client read timeout) since the 08-04 redeploy. Rebuilt each from the rack (D-138):
remove-unit --no-prompt -> remove-machine --force --no-prompt (REQUIRED: remove-unit does
NOT cascade to a never-provisioned dead-agent container) -> add-unit --to <bundle
placement>. Both fresh containers' cloud-init finished (~211s); mons bootstrapped 3/3, 4
OSDs active, storage cascade cleared. A 2nd apt-cacher-ng mode (radosgw-hacluster 404 on a
rotated point-release -> exit 100) SELF-HEALED via juju hook retry. Measured after-state
17:34:08Z: census 62 active; remaining non-active all deferred-by-design + gss +
barbican-vault settling.
Durable capture (operator-directed "a then b"):
- docs/audit/stage5-dc0-ceph-remediation-20260805.txt: transcribed capture (NOT script(1));
the before-state (CLOSE-WAIT sockets, ~45000s etimes, both containers) is unrecoverable
and lives only here.
- docs/CURRENT-STATE.md: dated status block; measured 62 active (NOT the unmeasured 40/47,
C2); records the --force fact and the phase-03 Step-3.1 gate defect (expects 1, VR1 has 4
deferred + gss) as durable finding.
- runbooks/appendix-A-troubleshooting.md: NEW entry for the symptom pair (both modes +
gated remove/--force/re-add remediation). Drafted; operator-reviewed.
- docs/changelog-20260805-stage5-dc0-ceph-remediation.md: session changelog (separate
same-day file from the skill-DOCFIX changelog, disclosed in header); F2/F3/F4 owed,
unnumbered.
repo-lint 0 fail / 1 legacy warn; ledger-scan DOCFIX next-free still 210 (no token leaked);
all touched files ASCII-clean. Status lives in docs/CURRENT-STATE.md ONLY.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

DOCFIX-209: correct SKILL.md session-CLOSE bookend + re-seed ledger machine-block
...
SKILL.md's always-loaded close-bookend line conflated two ratified mechanisms:
GA-R4 rule 2 (ledger rotation) and GA-R2/D1 (changelog = session body-of-record).
"the full session body goes to docs/archive/ in the same close commit" is wrong on
both counts -- the per-session body is the changelog (docs/changelog-<date>-<label>.md,
cited by a Body: line in the SESSION CLOSE block), and docs/archive/ holds rotated
ledgers (session-ledger-rotated-*.md) + per-stage consolidated records
(archive/changelogs/, archive/stage-records/), never per-session bodies. Corrected in
SKILL.md and the byte-identical line in the derived consolidated-20260727 snapshot.
operating-discipline.md (Session continuity) was already correct -- no edit.
Disposition operator-ruled 2026-08-05 "Proceed as DOCFIX-209": both same-day rulings,
read correctly, back the practice and contradict the skill line, so no ratified text
changes. Body: docs/changelog-20260805-skill-close-convention-docfix.md.
Also re-seeded session-ledger.md machine-derived block from ledger-scan (prior seed
2026-08-02, stale): 4 open decisions (D-142 added), 29 open SEC (SEC-033 added),
next-free D-143 / DOCFIX-210 / BUNDLEFIX-053. No status-bearing surface touched
(a skill routing line is not a status claim -- no GA-R1 L10 coupling). Stage 5 dc0
untouched. repo-lint 0 fail / 1 legacy warn; skill files ASCII-clean; ledger fences OK.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Close-fix 2026-08-05: add Body changelog, verify root CA (openssl), reconcile D-142 scan-visibility + durability
...
Advisor-caught close gaps against the repo's own convention:
- Body: NEW docs/changelog-20260805-vault-init-ovn-resolved.md (prior closes cite a
changelog Body: line; the 08-05 close block had Sweep: but no Body:). Ledger + CURRENT-STATE
now cite it.
- Root CA: decoded the ACTUAL pasted PEM with openssl on the rack (not a self-decode).
Confirms notBefore Aug 5 02:05:57 2026 / notAfter Aug 2 01:06:27 2036 GMT; adds sha256
75:DF:33:97:...:35:A1. as-exit as-built + CURRENT-STATE updated to measured fact.
- D-142 Status now leads "PROPOSED / OPEN" so ledger-scan surfaces it (was invisible; scan
keys the open list off the Status token) -> reconciles the close block's "4 open decisions".
- Durability line completed: voffice1 PULLED to sync (was 3321c57, 4 behind); dc0 rack
~/repo-stage unaffected (docs-only; preflight sha verified); gauntlet not owed (docs-only).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
SESSION CLOSE 2026-08-05 (GA-R4 bookend): vault init DONE + ovn-central RESOLVED; D-142 saved
...
Bounded 15-line ledger summary + close sweep (3 first-surface). Stage 5 remains
OPEN -- session bookend, not a stage close.
- vault init complete (operator-run one-shot, dc0 rack, -m vr1-dc0); root CA
generated; vault active/idle.
- ovn-central/3,4,5 all active -- OVN NB/SB cluster formed (the redeploy's purpose).
- D-142 vault-init QoL sweep SAVED (approved-in-principle, impl deferred; R2 open).
- Close sweep: ceph-mon/2 + ceph-radosgw/0 apt-wedge cascade (triage next);
/background unavailable over Remote Control; R2 transport gap.
Durability: 0 uncommitted / 0 unpushed; repo-lint 0 fail / 1 legacy warn.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|