| 2026-08-07 |
dc0 tailscale .7: carve applied+verified + deployed jammy (carved-and-ready)
...
LIVE (gated): carved the .7 router's two legs (metal-admin 10.12.8.7 +
provider-public 10.12.4.7 VLAN 5002, no br-ex) via the new aux-carve;
check pass=8/0. MAAS-deployed jammy -> Deployed, sshd live on 10.12.8.7.
State: dc0 .7 = carved-and-ready. Two JOIN prerequisites remain, both
off-session: (a) tagged pre-auth key + Headscale autoApprovers/ACL
(operator key is PLAIN; join NOT attempted); (b) SSH access via
vr1-office1-svc (region injects only that key; operator holds it).
CURRENT-STATE tailscale block + changelog Items 3-5 updated; dc1 ruling
recorded (build vr1-dc1-region first, no migration).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Session close 2026-08-06/07 (GA-R4 bookend): phase-03 Step 3.4 G3 + per-DC Tailscale rulings & .7 VMs
...
Bounded ledger summary + sweep (queued-findings-20260807, O1-O5) + CURRENT-STATE
tailscale-build status + node-vm harness reconcile (11/66 -> 12/72, red gate that
my substrate commits caused by not re-running the gauntlet -- the 2026-07-30 lesson
repeated). Rotated the 2026-08-03 close to archive (ledger 298 < 300).
Session delivered: Step 3.4 G3 domain-manager probe (built + live PASS, closes the
last phase-03 exec item); Decision C (Horizon reconciled to VR1, Step 3.3 splits to
its own tailnet-gated row); D-129(iii) amendment rulings a-d (dedicated .7 VM / STAR
/ single-HA-pinned / SNAT-on, both DCs) + D-134 octet .7 + D-107 citation DOCFIX;
site-tailscale.sh tooling; the .7 subnet-router VMs applied + MACs pinned on both
DCs (dc0 tailscale; dc1 region + tailscale). Headscale-side join deferred (no
control-plane access); MAAS commission/deploy/carve + dc1 region setup owed.
Gates: repo-lint 0 fail; gauntlet ALL GREEN (101). Memory: +ipv6-primary-posture
(drift-prevention, operator-directed) + instrument-currency #19.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

D-129(iii) amendment: per-DC Tailscale operator-access rulings (a-d), both DCs
...
Record four GA-R5 rulings (2026-08-07) that pull the gap-21 per-DC Tailscale
subnet-router forward to close phase-03 Horizon properly (operator: "pull the
tailscale steps forward"; "plan and push to both DC0 and DC1"):
(a) dedicated VM at utility .7 (10.12.8.7 / 10.12.68.7)
(b) STAR -- operator->DC only (the Headscale ACL / security boundary)
(c) SINGLE router, HA scale-up PINNED
(d) SNAT ON now, source-IP preservation PINNED
These are D-129(iii) implementation sub-decisions (not a new D-number). Also:
correct the D-107 citation defect (D-107 is airgap/mirror/NTP, governs no
Tailscale; D-129(iii) governs); extend D-134's standing octet map to .7;
update the gap-21 register row and CURRENT-STATE (phase-03 does NOT close this
session -- Step 3.3 Horizon splits to its own gate row, gated on the tailnet
build + vault CA on the workstation + a browser login over the tailnet).
Records only -- no code, no cloud change. The build (site-tailscale.sh + the
.7 VM per DC + Headscale star ACL/autoApprovers + SEC row) is next.
repo-lint 0 fail; Decision C reconciliation measured read-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-06 |

phase-03 Step 3.4 (dc0): G3 domain-manager probe -- named check + live PASS
...
Build scripts/g3-domain-manager-probe.sh + tests/g3-domain-manager-probe (harness
12/12, every exit path proven) as the GA-R6 named executable check for phase-03
Step 3.4 stage 2 (gate G3), filling the hard-rule-4 gap (was a manual runbook walk
only) and giving dc1's Step 7 a reusable probe. Grounded in the real policy
(domain-manager-policy.yaml:103 create_grant managed-role guard).
Ran it LIVE from the dc0 rack (operator-approved): G3 PASS, 7 ok / 0 fail, teardown
verified clean (zero g3-* residue). Stage-1 (PO:) verified read-only: policyd-override
attached, all 3 keystone units 'PO: Unit is ready'. Capture
docs/audit/g3-dc0-probe-20260806.txt.
HARNESS-MANIFEST recorded 99->100; gauntlet ALL GREEN (100), repo-lint 0 fail.
CURRENT-STATE: Step 3.4 -> RESOLVED; phase-03 exit gate now turns on the Horizon
reachable/login-works item (D-044/D-075 per-rebuild + VR0-nginx vs VR1-tailnet access
model = Decision C), to be measured + ruled next.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Session close 2026-08-06 (part 3, GA-R4 bookend): phase-03 core-verify F-CV1/F-CV2/F-CV3 RESOLVED + binding conformance
...
GA-R4 bookend: bounded ledger summary + sweep. F-CV2 (openstack CLI on the dc0
rack), F-CV1 (designate binding, BUNDLEFIX-056), F-CV3 (dashboard TLS via D-072
AMENDMENT VR1 / BUNDLEFIX-057) all RESOLVED; BUNDLEFIX-058 designate-stack
conformance -> binding conformance clean cloud-wide. D-134 Roosevelt-delta +
gap-21 access-model context recorded.
Sweep O11 FIRST SURFACE (dc0 MAAS query method); O10 rack repo-stage bundle STALE
(re-stage before redeploy). Memory instrument-currency #18. repo-lint 0-fail,
gauntlet ALL GREEN (99) at ace0e16, ledger 292 lines (no rotation).
NEXT: Step 3.4 (domain-manager policy) -- last phase-03 exit-gate item -> Steps 8-12
-> Stage-5 exit. Body: docs/changelog-20260806-phase03-coreverify.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-058: designate-stack binding conformance (amqp/cluster -> metal-internal) + ceph-rbd-mirror Stage-6 watch
...
Operator-requested binding-conformance sweep after F-CV1/F-CV3. The only remaining
ACTIVE svc-endpoint deviations were the designate stack: designate amqp (->rabbitmq)
+ cluster (peer), designate-bind cluster (peer) -- all on the metal-admin default,
deviating from the generic svc-to-svc rule (14 apps use metal-internal). No ruled
exception -> conformance repair, no new D-number.
Benign (designate served metal-admin so nothing broke, unlike F-CV1/F-CV3) but
off-plane. Repaired to metal-internal so dc1 inherits a fully-conformant bundle.
Proven live (same method): juju bind designate amqp=metal-internal cluster=metal-internal
+ juju bind designate-bind cluster=metal-internal (individually gated). Verified no
regression: bindings moved; designate still Stage-7-blocked on nameservers ONLY;
designate-bind active/idle; designate haproxy 0 DOWN (F-CV1 intact). No harness asserts
these bindings. gauntlet ALL GREEN, repo-lint 0-fail.
RECORDED: ceph-rbd-mirror certificates/cluster on metal-admin are UNBOUND today
(Stage-6 DR not wired) -> sweep O9 STAGE-6 WATCH + CURRENT-STATE, pre-check before
wiring at dc-dc-phase5. No other active deviations found cloud-wide.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-057 / D-072 AMENDMENT (VR1): dashboard cluster -> metal-internal (F-CV3 RESOLVED)
...
VR1 split-metal INVERTS the D-072 cluster placement. The openstack-dashboard charm
declares no admin/internal extra-binding and no os-*-network (metadata + charmhub
docs), so apache serves its SSL vhost on metal-internal while haproxy dialed
cluster=metal-admin -> vhost-less -> plaintext (the D-072 trap, inverted). Option A
(serve metal-admin) unavailable in-deployment (no charm lever).
Retire the VR0 exception: openstack-dashboard cluster -> metal-internal (generic rule
+ the served plane). Proven LIVE before ratifying (operator process directive), then
ratified GA-R5 "Ratified, land the config-of-record".
bundle.yaml cluster=metal-internal; design-decisions D-072 AMENDMENT (VR1);
binding-reference matrix + exception RETIRED + cross-ref; CURRENT-STATE/sweep/changelog
F-CV3 RESOLVED. Verified live: provider + operator metal-admin VIPs both TLS 200
CA-verified (were plaintext); cert covers all 3 VIP IPs (resolves AH01909); units
active/idle. No harness asserts this binding. gauntlet ALL GREEN, repo-lint 0-fail.
dc1 inherits via the shared bundle (no rebind).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Gap-21: capture the operator-access model (tailnet -> metal-admin dashboards) as forward context
...
Operator discussion 2026-08-06: operators reach each DC's metal-admin plane over
the tailnet and access routed dashboards (Horizon) from there. Added as dated
CONTEXT (not a ruling) to workflow gap-register item 21, informing the four
deferred Tailscale sub-decisions when re-raised. Captures the implications that
reach outside the Tailscale build: the operator-facing reverse proxy becomes
unnecessary; D-044/D-075 per-rebuild accommodations sunset for the operator path;
cert SANs on the metal-admin VIP (and FQDN certs, D-106/D-008) become load-bearing;
and it reinforces metal-admin as the dashboard's serving plane (the access-model
argument for F-CV3 option A over option B). Cross-referenced from the F-CV3 sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-056: designate public/internal bindings -- F-CV1 RESOLVED (fix executed + verified)
...
designate's REST API public+internal endpoints were omitted from the bundle,
defaulting to the '' metal-admin fallback -> orphaned the provider + metal-internal
legs of its ruled .62 VIP triple (D-020 amendment) -> _admin haproxy backend SSL-DOWN
on the unserved metal-internal address. The deployed charm declares public/admin/
internal extra-bindings (metadata verified); dnsaas (D-106) is ADDITIONAL, not a
replacement -- the prior "no public binding" reading was the defect's root.
Config-of-record: bundle.yaml +public:provider-public +internal:metal-internal;
provider-bundle-check EXPECT_PUBLIC_VIP 11->12 (vault stays out, not 13) with T16c/T16d
failing-direction tests (57->59, ALL PASS); network-space-binding-reference row 88 +
note. Gauntlet ALL GREEN (99); repo-lint 0-fail.
Live (operator-approved): juju bind designate public=provider-public
internal=metal-internal (rc=0, dc0 rack). Verified: full haproxy sweep 0 DOWN
cloud-wide; apache https vhosts span all 3 planes; cert reissued for provider-public;
catalog triple correct (public 10.12.4.62 / internal 10.12.12.62 / admin 10.12.8.62).
Governing: D-052 / D-020 amendment. Evidence:
docs/audit/stage5-dc0-phase03-coreverify-20260806.txt. Body:
docs/changelog-20260806-phase03-coreverify.md (Item 6).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F-CV1: reconcile after research -- :internal defect CONFIRMED, :public OPEN; D-141->D-052
...
Prior-art + governing-decision research on the designate bind-plane mismatch.
Corrects the earlier commit (aaeee93) which reached "CONFIRMED" and cited D-141
before checking the governing decisions (backwards ordering, memory #17).
D-072/BUNDLEFIX-011 prior art: the dashboard-plaintext class was root-caused in
VR0 (charm renders haproxy 443 backend on cluster-binding addr, apache SSL vhosts
only for default+public). dc0 bundle already carries that fix -> F-CV3 is a NEW
cause, parked for its own triage.
F-CV1 re-scoped: governing surface is D-052 + the generic binding rule, NOT D-141.
`:internal`->metal-internal is a CONFIRMED defect (on the metal-admin fallback;
every sibling binds it to metal-internal; no ruled exception; cert already covers
the metal-internal SANs; all 3 units have metal-internal addrs -> juju bind won't
be refused). `:public`->provider-public is OPEN -- designate uniquely carries
`:dnsaas` on provider-public (D-106 dual-VIP); its REST API may be intentionally
metal-admin-only. The per-app table is generated from bundle.yaml (descriptive of
the defect), not intent.
Fix method = D-072 precedent (bundle + live juju bind + haproxy-readback verify).
designate is Stage-7-blocked -> no urgency. Awaiting operator ruling on :public.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 Step 7: triage F-CV1 (CONFIRMED designate bind-plane mismatch) + F-CV3; sweep
...
Operator-authorized triage of the two plaintext-vs-TLS findings (read-only; not
fixed, hard rule 1). Both apps have vault certs rendered + the certificates
relation, so NOT the ovn CN-issuance class -- charm apache-TLS-frontend layer.
F-CV1 CONFIRMED (two findings, not one): designate/0 apache https frontend binds
only 10.12.8.198:8991 (metal-admin); haproxy's _admin backend dials
10.12.12.110:8991 (metal-internal) where no SSL vhost exists -> check-ssl hits
plaintext -> DOWN. VR1 dual-metal-plane bind mismatch (D-141), structural (not
Stage-7 collateral). F-CV3 (dashboard) is separate: :433 served by Ubuntu
default-ssl.conf, charm https frontend not effective; root cause not nailed.
Owned instrument caveat: earlier "no SSLEngine in sites-enabled" was a grep -r
false negative (does not follow the symlinks); apache2ctl -S corrected it.
Remediation = a focused, gated session. Sweep:
docs/audit/queued-findings-20260806-phase03-coreverify.txt
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 Step 7: phase-03 core-API VERIFIED; F-CV2 openstack-client on rack; F-CV1/F-CV3 logged
...
Core-API layer of phase-03 core-verify (adapted for vr1-dc0, run from the dc0
rack per D-138) VERIFIED read-only: settle walk (falsifiable prediction matched),
haproxy 0-DOWN across 12 activated VIP apps, admin-openrc scoped token, IP-only
endpoints, two-sourced keystone VIP, vault CA TLS OK.
F-CV2 RESOLVED (07-30 queued-F1, hit at Step 7): installed python3-openstackclient
6.6.0-0ubuntu2 on the dc0 rack; CURRENT-STATE section 7 row amended (GA-R1/C1).
Exit gate NOT MET (2 open, GA-R6/E3 no conditional close): F-CV3 dashboard VIP
serves plaintext (apache-SSL-inactive despite certs) -> Horizon exit-gate fails;
Step 3.4 domain-manager policy NOT RUN. F-CV1 designate-api plaintext vs haproxy
check-ssl (same shape); "collateral of block" reading retracted. Logged not fixed.
Evidence: docs/audit/stage5-dc0-phase03-coreverify-20260806.txt
Body: docs/changelog-20260806-phase03-coreverify.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: memcached scaled to 3 LIVE (operator-approved) + consumer verify
...
Live scale-up executed on the dc0 rack (operator "Both approved"):
juju add-unit memcached -n 2 --to lxd:1,lxd:2 -m vr1-dc0 (exit 0).
- CONVERGED 3/3 active/idle, one per control node (memcached/0 control-01, /1 control-02,
/2 control-03; all "Unit is ready and clustered", 11211/tcp). Clean install, no F5 apt-wedge.
Capture docs/audit/stage5-dc0-memcached-scaleup-20260806.txt.
- Consumer verify (D-121 verify-at-deploy): nova-cloud-controller sees all 3 servers
(memcache_servers = .115,.166,.165:11211). designate coordination backend_url shows ONE
(memcached/0) -- designate is workload-blocked pre-Stage-7, so re-check at designate
activation; recorded, not a scale-up defect.
- bundle.yaml re-staged to the dc0 rack (42845edb == repo HEAD).
- CURRENT-STATE reconciled: memcached config=3 AND live=3.
Record-only commit (the live mutation itself was the add-unit above). repo-lint 0 fail.
Revert of the LIVE state (if ever needed): juju remove-unit memcached/1 memcached/2 -m vr1-dc0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: memcached 1->3 config-of-record [BUNDLEFIX-055 + DOCFIX-212]
...
Operator ruling 2026-08-06, exact utterance "3 units (restore intent)" -- restoring the
2026-07-31 memcached->3 direction that BUNDLEFIX-053 never folded and the retired
dc-ha-scaleup.yaml had carried unrealized. Found while reconciling a stale comment:
- Full merge-diff (measured) of the archived overlay vs bundle.yaml showed the ONLY divergence
was memcached (overlay num_units:3 / base 1) -- BUNDLEFIX-053 folded the HA chain but not this.
Corrects two earlier misstatements (this session): "memcached=3 in bundle.yaml" (it was 1) and
the retirement's "wholly redundant / all-keys no-op" claim (it diverged on memcached).
- BUNDLEFIX-055: bundle.yaml memcached num_units 1->3, to:[lxd:0]->[lxd:0,1,2] + the
3-independent-caches rationale (no hacluster/VIP; clients hash across the full server list).
This completes the fold, making the archived overlay genuinely, fully redundant.
provider-bundle-check 57/0 (memcached is orthogonal to the arity/VIP gates).
- DOCFIX-212: design-decisions D-121 "Left single (NOT scaled)" list -- memcached removed,
recorded as =3.
- CURRENT-STATE: records the CONFIG=3 / LIVE=1 divergence -- a live add-unit (operator wants it
deployed live) or a teardown+redeploy reconciles it. The live scale-up is a separate gated step.
Gauntlet ALL GREEN (99 harnesses); repo-lint 0 fail (1 pre-existing legacy L1 warn).
Owed: re-stage the functionally-changed bundle.yaml to the dc0 rack (gated); live add-unit to 3.
Revert: bundle.yaml memcached back to num_units:1 + to:[lxd:0]; revert the DOCFIX-212 D-121 line
and the CURRENT-STATE gap note.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F4: 14/14 HA now measurement-backed + vault ha_enabled MEASURED
...
Live read-only capture from the dc0 rack (juju status -m vr1-dc0, D-138 path):
docs/audit/stage5-dc0-juju-status-14of14-20260806.txt.
- 14/14 D-121-enumerated HA apps at scale=3 (per changelog-20260805-d121-ha-scaleup.md:24).
HONEST SPLIT (GA-R6 E3, not rounded): 12 active/idle; 2 blocked-at-scale=3 on KNOWN
non-HA items -- octavia (configure-resources) + designate (nameservers). Model-wide
152 active/idle, 7 blocked, 1 unknown (gss, normal). Converts the CURRENT-STATE
14/14 line from OPERATOR-ATTESTED to MEASUREMENT-BACKED (C2: measurement wins).
- F4 vault ha_enabled MEASURED: vault status HA Enabled=FALSE x3, storage=mysql,
unsealed, shared cluster id. CAUSE (measured, corrects a mid-session mis-blame of the
mysql backend): charm renders storage "mysql" with no ha_enabled and exposes no such
option (only vip + dns-ha-access-record) -- HA is charm/VIP model, not vault-native.
CONFIRMS D-121 (v-a) as-built; vault-native/Raft HA stays owned by D-068. No new decision.
- F8 ceph-radosgw RESOLVED: 3 units active/idle "Unit is ready" (prior stale-status converged).
- F9 staging drift MEASURED (re-stage gated, owed at next deploy): dc0 rack bundle.yaml +
vr1-dc0-vips.yaml STALE vs HEAD; dc1 rack has no ~/repo-stage (HELD).
Revert: git revert this commit (record-only; no live-cloud change was made -- all reads).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Task #2: commit D-020 vault-metal-only amendment; teach renderer + gates; harnesses green
...
The RULED-but-uncommitted D-020 amendment (vault METAL-ONLY, 2026-08-05) is now
committed and enforced end-to-end. Discovered mid-task that the overlays are RENDERED
from render/values/*.yaml (D-136) and hand-editing them is forbidden by the render-drift
gate -- which was ALREADY red at the part-2 close (undocumented) from that session's dc0
overlay hand-edit. Resolved by teaching the renderer, not by hand-editing (advisor-reviewed;
the amendment is ruled, so this is OPS implementation).
Vault VIP shape is now the metal PAIR (metal-admin + metal-internal, no provider-public,
no v6) across all four consumers, each with a failing-direction fixture:
- provider-bundle-check.py: per-DC/family-aware vault exception; STILL band- and
octet-uniqueness-checked (advisor caught the first draft's early `continue` disarming
octet_owner for .61 -- proven rc=0 draft / rc=1 fixed). T54/T55/T56.
- render-dc-overlays.py: name-keyed metal-only render branch (drops provider + v6);
overlays re-rendered (diff vs HEAD = exactly the one vault line each). T15b.
- pre-flight-checks.sh CHECK 1: awk name-tracking + vault metal-pair branch. T28b.
- render/values/*.yaml: vault comment (part of rendered bytes).
D-121 status 12/14 -> 14/14 in CURRENT-STATE, with the missing juju-status capture
DECLARED as an owed gap (GA-R1 rule 2) rather than papered over.
Verify: gauntlet ALL GREEN (99); repo-lint 0 fail; provider-bundle-check 58/0,
render-dc-overlays 24/0, render-drift 4/0, pre-flight-checks 32/0.
Body: docs/changelog-20260805-task2-vault-metal-only-commit.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-05 |
CURRENT-STATE: D-121 HA scale-up now 12/14 (nova-cc, rabbitmq, barbican DONE)
...
Measured 2026-08-05: nova-cloud-controller + rabbitmq-server (native erlang 3-node cluster)
+ barbican scaled to 3-unit HA. 12 of 14 HA apps at 3-unit HA. Only keystone (cloud-wide
auth VIP blip) and vault (operator unseal) remain -- both operator-gated. vault HA also
resolves the barbican-vault secrets-storage dependency.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-053: bundle.yaml to full 3-unit HA (D-121); reverse BUNDLEFIX-002 vault de-HA
...
Reverses the D-009/BUNDLEFIX-002/003 decorative-single-unit posture to real 3-unit HA per
operator ruling 2026-08-05 ("superseded the HA removal from all items ... stand up all
remaining HA apps ... HA blocks sign off"; "testing as if Roosevelt"). Executes adopted
D-121 incl. its (v-a) vault sub-ruling.
bundle.yaml:
- num_units 1->3 + to:[lxd:0,1,2] on 14 HA apps (keystone glance nova-cc placement
neutron-api cinder ceph-radosgw dashboard octavia barbican magnum designate rabbitmq vault)
- hacluster cluster_count 1->3 x12
- vault-hacluster (cluster_count:3) + [vault:ha, vault-hacluster:ha] uncommented -- the
BUNDLEFIX-002 reversal (D-121 (v-a): vault MySQL-backed HA)
- rabbitmq min-cluster-size:3 (D-009 amendment)
- governing comments reconciled (D-009/BUNDLEFIX-002/003 -> D-121 built)
overlays unchanged (VIPs already set).
Every change verified against charm/juju docs or the deployed charm (not assumed):
hacluster cluster_count default 3; rabbitmq min-cluster-size; vault ha endpoint=metal-internal
like keystone; vault MySQL-backend HA per HashiCorp docs (ha_enabled+lock_table). Validated:
provider-bundle-check.py --dc vr1-dc0 PASS (13 hacluster subs cluster_count==num_units; 13
principals carry a VIP; 109 relations well-formed); repo-lint 0 fail.
Live: Wave 1 (8 apps) scaled to 3-unit HA (see CURRENT-STATE + changelog). Waves 2-4 and the
post-wave bundle/overlay reconciliation (pinned task) remain. Body:
docs/changelog-20260805-d121-ha-scaleup.md. Status in docs/CURRENT-STATE.md ONLY.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
CURRENT-STATE: barbican-vault/0 verify result (C2) -- genuine incomplete, missing vault_url
...
Item (b) verify supersedes the "settling" note in 71c5b97: barbican-vault/0 is NOT settling
and NOT deferred. vault provided per-unit role_id + token but the secrets-storage databag
lacks vault_url (app-data empty) -> charm logs "Requesting access to vault (None)". Not
deploy-blocking (barbican/0 active on software backend). Measured 2026-08-05 17:55Z.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: remediate ceph-mon/2 + ceph-radosgw/0 "allocating" (apt CLOSE-WAIT); durable capture
...
Live-cloud, operator-gated (2 exchanges). Two units stuck `allocating` were root-caused
NOT to a down proxy (both dc0 proxies PASS; security.ubuntu.com 200/0.67s) but to
cloud-init `apt-get update` parked in CLOSE-WAIT to 10.12.8.4:3142 for ~12.5h (apt has no
client read timeout) since the 08-04 redeploy. Rebuilt each from the rack (D-138):
remove-unit --no-prompt -> remove-machine --force --no-prompt (REQUIRED: remove-unit does
NOT cascade to a never-provisioned dead-agent container) -> add-unit --to <bundle
placement>. Both fresh containers' cloud-init finished (~211s); mons bootstrapped 3/3, 4
OSDs active, storage cascade cleared. A 2nd apt-cacher-ng mode (radosgw-hacluster 404 on a
rotated point-release -> exit 100) SELF-HEALED via juju hook retry. Measured after-state
17:34:08Z: census 62 active; remaining non-active all deferred-by-design + gss +
barbican-vault settling.
Durable capture (operator-directed "a then b"):
- docs/audit/stage5-dc0-ceph-remediation-20260805.txt: transcribed capture (NOT script(1));
the before-state (CLOSE-WAIT sockets, ~45000s etimes, both containers) is unrecoverable
and lives only here.
- docs/CURRENT-STATE.md: dated status block; measured 62 active (NOT the unmeasured 40/47,
C2); records the --force fact and the phase-03 Step-3.1 gate defect (expects 1, VR1 has 4
deferred + gss) as durable finding.
- runbooks/appendix-A-troubleshooting.md: NEW entry for the symptom pair (both modes +
gated remove/--force/re-add remediation). Drafted; operator-reviewed.
- docs/changelog-20260805-stage5-dc0-ceph-remediation.md: session changelog (separate
same-day file from the skill-DOCFIX changelog, disclosed in header); F2/F3/F4 owed,
unnumbered.
repo-lint 0 fail / 1 legacy warn; ledger-scan DOCFIX next-free still 210 (no token leaked);
all touched files ASCII-clean. Status lives in docs/CURRENT-STATE.md ONLY.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Close-fix 2026-08-05: add Body changelog, verify root CA (openssl), reconcile D-142 scan-visibility + durability
...
Advisor-caught close gaps against the repo's own convention:
- Body: NEW docs/changelog-20260805-vault-init-ovn-resolved.md (prior closes cite a
changelog Body: line; the 08-05 close block had Sweep: but no Body:). Ledger + CURRENT-STATE
now cite it.
- Root CA: decoded the ACTUAL pasted PEM with openssl on the rack (not a self-decode).
Confirms notBefore Aug 5 02:05:57 2026 / notAfter Aug 2 01:06:27 2036 GMT; adds sha256
75:DF:33:97:...:35:A1. as-exit as-built + CURRENT-STATE updated to measured fact.
- D-142 Status now leads "PROPOSED / OPEN" so ledger-scan surfaces it (was invisible; scan
keys the open list off the Status token) -> reconciles the close block's "4 open decisions".
- Durability line completed: voffice1 PULLED to sync (was 3321c57, 4 behind); dc0 rack
~/repo-stage unaffected (docs-only; preflight sha verified); gauntlet not owed (docs-only).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
SESSION CLOSE 2026-08-05 (GA-R4 bookend): vault init DONE + ovn-central RESOLVED; D-142 saved
...
Bounded 15-line ledger summary + close sweep (3 first-surface). Stage 5 remains
OPEN -- session bookend, not a stage close.
- vault init complete (operator-run one-shot, dc0 rack, -m vr1-dc0); root CA
generated; vault active/idle.
- ovn-central/3,4,5 all active -- OVN NB/SB cluster formed (the redeploy's purpose).
- D-142 vault-init QoL sweep SAVED (approved-in-principle, impl deferred; R2 open).
- Close sweep: ceph-mon/2 + ceph-radosgw/0 apt-wedge cascade (triage next);
/background unavailable over Remote Control; R2 transport gap.
Durability: 0 uncommitted / 0 unpushed; repo-lint 0 fail / 1 legacy warn.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0: vault init DONE + ovn-central RESOLVED; D-142 vault-init QoL saved
...
Live (operator-run one-shot, recorded no-secrets): phase-02-vault-bringup Steps
2.1-2.3 on the dc0 rack (-m vr1-dc0), vault active/idle, root CA generated
(valid 2026-08-05 -> 2036-08-02). ovn-central/3,4,5 ALL active -- OVN NB/SB
cluster formed; the multi-session cert saga is closed by the two fixes landed
this cycle (dc-node-etchosts Step 1.2b CN delivery + D-052 '' -> metal-admin).
Engineering saved (PROPOSED, not executed -- operator: run current commands now,
test QoL next opportunity):
- D-142 PROPOSED: vault-init workflow QoL sweep (APPROVED-IN-PRINCIPLE, IMPL
DEFERRED; R2 off-host transport UNRESOLVED). Distinct from D-068 (substrate) /
D-011.6 (manual-unseal bar).
- docs/audit/vault-init-qol-proposal-20260805.md: R1-R5 + full hidden-prompt
safety analysis + tee-write residual + pre-init writability probe + R3 harness
constraints + operator's verbatim safety constraint.
- runbook-fold-register F13: -m openstack -> -m vr1-dc0 / run-location = DC rack
(D-138) on phase-02-vault-bringup (scope stretch stated).
CURRENT-STATE + as-exec updated (L10 coupling). Security hygiene: child token in
operator paste was ttl=10m/expired -> benign, not stored. Remaining (separate):
ceph-mon/2 + ceph-radosgw/0 allocating (apt-wedge) cascade.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

D-052 RE-AMENDMENT: ovn-central '' default metal-internal -> metal-admin + network binding reference
...
- ROOT: the '' default is the MANAGEMENT binding (primary NIC + juju agent<->controller),
universal '' = metal-admin across all 56 apps. The 08-03 amendment ('' -> metal-internal,
never deployed until 08-04) left ovn-central single-legged on the isolated metal-internal
plane -> no route to controller (10.12.8.5:17070 unreachable) -> agent-binary download failed
-> ovn-central never started. Its cert rationale was already withdrawn; cert CN is fixed
independently by the /etc/hosts postruncmd (dc-node-etchosts.sh, Step 1.2b).
- bundle.yaml: ovn-central '' -> metal-admin; functional endpoints (certificates, ovsdb*,
coordinator) STAY metal-internal. provider-bundle-check 55/55, repo-lint 0 fail.
- D-052 RE-AMENDMENT 2026-08-05 recorded (GA-R5, operator utterance quoted). Scope: ovn-central
only; every other binding re-verified correct vs dataflow (ruled exceptions: D-072 dashboard,
D-106 designate:dnsaas).
- NEW docs/network-space-binding-reference.md: 6-plane roles + CIDRs + RHOSP analogs, the ''
management-binding rule, full 56-app placement matrix, dataflow confirmation, ruled exceptions.
- Stage 5 remains OPEN. NEXT: re-stage bundle to rack, re-home ovn-central on the live model,
then vault init (operator-only) -> ovn cert -> converge.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-04 |

Stage 5 dc0 redeploy: Path M teardown (clean) + ovn-central cert DELIVERY bug fixed
...
- Path M teardown of vr1-dc0 complete + clean: graceful drain (36->15) + juju
resolved --no-retry (error-unit wedge, 15->3) + --force --no-wait (agents-stopped
tail, 3->0). NO orphan (controller healthy), M.5 cascade clean (10 machines
UNCHANGED, 9 Ready + 1 Deployed). Both --force and resolved operator-ruled (GA-R5).
- dc-node-etchosts.sh: render runcmd: -> postruncmd: -- juju model-config forbids a
top-level runcmd in cloudinit-userdata. The 08-04 fix proved the CONCEPT live but
never the DELIVERY (harness graded cloud-init YAML, which accepts runcmd, not juju
acceptance). Harness switched to postruncmd + NEW T10 (juju-accepted key, rejects
bare runcmd), MUTATION-PROVEN, 10/10. Step 1.2b now live-passing; re-staged to rack.
- M.6 rebuild to the deploy step: add-model (cred vr1-dc0-cred), spaces PASS 0 fatal,
egress 8/8 PASS, dry-run PASS (9 machines, per-DC tags). Stage 5 remains OPEN.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

SESSION CLOSE 2026-08-04 (GA-R4 bookend): ovn-central cert PROVEN + WIRED; SEC-033 filed
...
Bookend for the session that root-caused ovn-central's cert failure (isolated
metal-internal plane -> empty CN), proved the /etc/hosts fix live, and wired it
via dc-node-etchosts.sh + phase-01 Step 1.2b. Stage 5 remains OPEN.
- docs/session-ledger.md: bounded 2026-08-04 summary; oldest block (07-31 node
carve) rotated to docs/archive/session-ledger-rotated-20260804.md; 295 lines.
- docs/security-ledger.md: SEC-033 (tls-certificates relation databag exposes
vault's global-client private key to any juju model reader; interface-level).
- docs/audit/queued-findings-20260804-ovn-cert-fix.txt: close sweep, 4 FIRST
SURFACE (SEC-033-now-filed, live-model test drift, rack repo-stage prereq,
juju-routing research note).
- docs/CURRENT-STATE.md: SESSION CLOSE 2026-08-04 pointer; NEXT = step 3 redeploy.
Gauntlet ALL GREEN (99), repo-lint 0 fail. voffice1 synced to HEAD.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

ovn-central cert: root-caused, fix PROVEN live, and wired for redeploy
...
Root cause (measured, corrects the committed rdns_mode/binding framing):
charm-ovn-central derives its TLS common_name from get_hostname(its
metal-internal address); that plane is deliberately isolated (D-052) with no
reachable resolver, so the reverse lookup returns None -> empty CN -> vault
issues no server cert -> OVN cluster never forms. rdns_mode=2 and the PTR
exist; only the reverse is unreachable from the isolated plane.
Binding approach REFUTED live (3 configs): the charm uses the metal-internal
address regardless of which endpoint is rebound, and juju default-route
selection is not tied to the default binding. So the app STAYS on
metal-internal (D-052-correct for its OVSDB/certificates data type).
Fix PROVEN end-to-end (controlled single-unit test): an /etc/hosts reverse
entry -> CN populated -> vault issued ovn-central_0.server.cert ->
/etc/ovn/{cert_host,key_host,ovn-central.crt} written; control units without
it stayed broken. OVN imposes no CN-content rule; vault signs any non-empty CN.
Wired (verify-at-provision on the redeploy):
- NEW scripts/dc-node-etchosts.sh + tests/dc-node-etchosts (harness 9/9):
renders a per-DC cloudinit-userdata adding each node's metal-internal
address -> hostname to /etc/hosts at provision; CIDR derived from lib-net
(dc0 10.12.12.0/22, dc1 10.12.72.0/22). Rendered runcmd executed live and
correctly scoped to metal-internal.
- runbooks/phase-01-bundle-deploy.md Step 1.2b: gated pre-deploy step
(after add-model, before deploy), VR1-only.
- A NEW mechanism borrowing D-008's shape, not D-008 itself.
Records: reeval RESOLUTION + CURRENT-STATE RESOLVED+WIRED + changelog; the old
rdns_mode remediation is VOID. Gauntlet ALL GREEN (99), repo-lint 0 fail.
Live tests were reversible; model left at its captured before-state.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

CORRECT ovn-central root cause: NOT rdns_mode -- metal-internal is isolated (no reachable resolver)
...
Autonomous step-1 measurement refuted the committed rdns_mode diagnosis:
- rdns_mode=2 on ALL planes; the PTR EXISTS at the region BIND (10.12.8.6
answers dig -x 10.12.12.122). "enable rdns_mode" is a VOID no-op.
- metal-internal is deliberately isolated (link-scoped routes only; ping -I
eth1 10.12.8.6 -> NO ROUTE). The region controller has no metal-internal
interface, so no resolver is reachable on that plane.
- systemd-resolved scopes the reverse of a container's OWN metal-internal
address to eth1 -> no reachable resolver -> empty cert CN. Discriminating
test (advisor-directed): both links set to the reachable 10.12.8.6 STILL
failed "No route to host" -> NOT a dns_servers change either.
The fix is therefore a DECISION (decouple ovn-central's cert CN from
metal-internal reverse-DNS), a precondition for any redeploy (same MAAS,
verify at provision time). VOID banners + CORRECTION blocks added to the
reeval, remediation-plan, sweep, and CURRENT-STATE. Refutation of LP #2044324
and the vault-issuance rule are unchanged and stand.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-08-03 |
CURRENT-STATE: correct ovn-central root cause (reeval) + record Stage-5 sweep (GA-R1 C1)
...
Supersede the LP #2044324 framing with the measured root cause (metal-internal
reverse-DNS gap -> empty CN -> no server cert), point to the reeval + remediation
+ sweep audit docs, and record sweep findings F1/F2/F3. Satisfies the L10 status-
in-same-commit rule the prior audit-docs commit (eb63339) tripped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

reconcile provider-bundle-check harness to D-141 (sweep F1 CLOSED)
...
The 2026-08-03 deploy session reverted the dc0 VIP overlay to IPv4-only
under D-141 (3e691cd) but shipped it without the companion harness update,
leaving run-tests-all RED at 1/98 (provider-bundle-check, 4 dual-family
cases T19/T21/T25/T45 asserting a shape the deploy input no longer has).
Reconcile by re-pointing, never deleting (the checker's dual-family path is
still live code and D-141 rule-3 promotes v6 later):
- new synthetic dual.yaml fixture on the MEASURED all-GUA legs of the
pre-revert deploy input (3e691cd^), present in the current apex -- NOT the
stale DUAL6 ULA constant D-139 deprecated
- T19 assertion REPLACED with the v4-only invariant (0 dual-family)
- T21/T25/T45 re-pointed to dual.yaml (dual-family PASS + apex-refusal
controls preserved)
- corrected the now-vacuous v4only.yaml comment
scripts/provider-bundle-check.py is UNCHANGED -- it is family-agnostic (a
vip is a v4 triple OR a dual-family sextet); this is a harness reconcile only.
Verified: provider-bundle-check 55/55 ALL PASS (count unchanged 55->55, pure
re-point); new T19 failing-direction proven; full gauntlet GAUNTLET: ALL
GREEN (98 harnesses) (was 1/98 FAIL); repo-lint 0 fail / 1 legacy warn.
CURRENT-STATE updated same commit (GA-R1 C1). Three follow-ons LOGGED to the
close sweep (FN1 stale DUAL6 ULA; FN2 generated overlay header; FN3 dual.yaml
builder can silently no-op at promotion).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|