| 2026-08-09 |

savegame 2026-08-08 (part 3): amphora build FIXED + geneve-over-v6 wrap-gate ROOT-CAUSED
...
GA-R4 bookend for the dc0-activation part-3 session.
- Body: docs/changelog-20260808-amphora-geneve.md (amphora two-layer fix [loop-device
passthrough + retrofit ubuntu-mirror->dc0 mirror; image 775ebeba ACTIVE], o-hm0 MTU PASS,
1-test-LB blocked by geneve-over-v6, delivery).
- Sweep: docs/audit/queued-findings-20260808-amphora-geneve.txt (FIRST SURFACE: live mutations
w/ reverts, test LB/net/amphora LEFT LIVE [teardown owed], retrofit diag traps).
- CURRENT-STATE part-3 progress block + session-closed pointer.
- session-ledger: bounded close summary; machine-derived block RE-SEEDED from ledger-scan
(SEC 29->28 [SEC-031 closed], DOCFIX 210->214, BUNDLEFIX 053->059); rotated 08-06 x2 ->
archive/session-ledger-rotated-20260809.md (ledger 283 lines).
Gates: repo-lint 0-fail (1 legacy warn); gauntlet ALL GREEN (102); ledger-scan reconciled
(4 open decisions, SEC 28, next-free D-143/DOCFIX-214/BUNDLEFIX-059, no new numbers).
Revert: git revert this commit (records only; live mutations + reverts are in the changelog).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-08 |

geneve-over-v6 root-cause: verify on compute-01 (amphora host) + tighten evidence
...
Advisor-flagged corrections to cd1f9cd's root-cause record (claims now measured, not inferred):
- data-tenant confirmed the SAME dual-stack plane on both node types (lib-net.sh PLANE_CIDRS
10.12.16.0/22=data-tenant, in SPACES6); divergence is address-family SELECTION, not a binding
error (both chassis types bind data-tenant).
- tunnel state re-read on vr1-dc0-compute-01 (ovn-chassis/1, the amphora's ACTUAL host): its
v6 local encap has IPv4 remote_ip to the 3 control chassis -> cross-family, same pattern.
- softened the bfd_status claim to AMBIGUOUS (OVN geneve may not populate BFD); decisive
evidence is the encap-family split + 100% ICMPv6 loss. compute<->compute v6 path noted UNTESTED.
- dropped the pre-allocated "D-143" (number assigned at ruling per grep-for-next-free).
Revert: git revert this commit (records only; no live cloud state changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

geneve-over-v6 wrap-gate ROOT-CAUSED (FAIL): OVN encap family split control(v4)/compute(v6)
...
The dc0 "verify-live geneve-over-v6" checkpoint gate FAILS, root-caused this session
via the 1-test-LB smoke test. Measured: OVN geneve encap IPs are split across address
families -- containerized control-node chassis (octavia LXD) have IPv4-only data-plane
addresses (10.12.16.x, D-134 auto-picked) -> IPv4 encap; carved compute metal uses IPv6
(2602:f3e2:f02:30::x). Cross-family geneve tunnels never form (bfd_status empty) -> 100%
cross-node overlay loss -> amphora unreachable from o-hm0 -> LB stuck PENDING_CREATE.
Roosevelt-delta: v6 builds must give containerized OVN chassis a v6 data-tenant address
so encap is family-consistent with metal. Candidate D-143 (PROPOSED, operator ruling owed).
Fix NOT applied (substantial + rebuild-relevant). Full record:
docs/audit/geneve-over-v6-rootcause-20260808.md. CURRENT-STATE progress block updated
(clears repo-lint L10 for the audit file).
Revert: git revert this commit (records only; no live cloud state changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

bookend: land prior part-2 GA-R4 close + dc0-activation part-3 progress (amphora FIXED, LB blocked on v6 overlay)
...
Lands the prior session's savegame residue (GA-R4 part-2 close block in
session-ledger.md, the 2026-08-08 ledger rotation to archive/, and the
dc0-activation-checkpoint queued-findings sweep) which was left uncommitted
pending the operator's push decision (now pushed: a25ed99..a6340e6).
CURRENT-STATE.md gets this session's (part 3) measured status update, which
also clears repo-lint L10 for the docs/audit/ sweep file:
- amphora build RESOLVED (was part-2 retrofit exit-1 INCIDENT): root cause
loop-devices-absent-in-LXD + dib-apt-targets-unreachable-public-archive;
fix = loop passthrough + retrofit ubuntu-mirror -> dc0 mirror. Image
775ebeba... ACTIVE+tagged octavia-amphora.
- G18 OWED#3 (o-hm0 MTU) PASS.
- 1 test LB blocked on geneve-over-v6 overlay (o-hm0->amphora 100% ICMPv6
loss); UNDER INVESTIGATION (== the verify-live geneve-over-v6 wrap gate).
Revert: git revert this commit (records only; no live cloud state changed by
this commit -- the live mutations are logged for the forthcoming changelog).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
G18 RULED (GA-R5, option b): Octavia lb-mgmt recorded out of apex scope + D-139 /64 reserved
...
Gate G18 CLOSED 2026-08-08. Operator ruled option (b): the charm-created Octavia
lb-mgmt-net (IPv6-ULA fc00::/64, charm-generated per R8, regenerates per deploy) is
deliberately charm-owned and OUT of apex scope; the separate D-139 apex GUA lb-mgmt
/64 is kept reserved (distinct MAAS-underlay object, no charm consumer). R8 not
reopened. lb-mgmt is v6-ULA, outside the 10.12->10.13 v4 re-IP. Primary record in
CURRENT-STATE G18 row; annotations on D-101/R8 + D-139. No new D-number. Prep package
docs/audit/g18-lb-mgmt-ipam-ruling-prep-20260808.md. Also: changelog Items 3-5
(live network-create, G18 ruling, Designate real-Stage-7 decision).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
dc0-activation: phase-04 network scripts MAAS_PROFILE-aware (DOCFIX-213, F3/D-138) + re-IP ruling-prep package
...
phase-04-network-create.sh + phase-04-network-verify.sh honour MAAS_PROFILE
(default admin=VR0/office1; VR1 overrides to the DC regional, e.g. vr1-dc0-region)
instead of hardcoding 'maas admin'. Fixes F3: the dc0 rack has openstack+cloud L3
but no maas profile, and the maas two-source gate needs both on one host. Operator
directive: each DC has its site regional maas; racks register up to the DC regional.
Harnesses extended with EXPECT_PROFILE proof (failability verified out-of-band).
Also lands docs/audit/reip-1013-ga-r5-ruling-prep-20260808.md (Task #6 background
analysis) + a CURRENT-STATE pivot pointer to it (L10 coupling).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

savegame 2026-08-08: dc0 .7 tailscale FIXED (advertise-only) + 10.13 re-IP PIVOT + dc0 checkpoint plan
...
GA-R4 bookend + sweep for the session that fixed the dc0 .7 tailscale install
(advertise-only, --authkey=file:, check guards -- committed 02e0b12/faef662) and
surfaced the 10.12->10.13 re-IP pivot.
- CURRENT-STATE: 2026-08-08 pivot callout -- 10.12 collides with the live IPv4
cloud; drive dc0 to full deployment as a CHECKPOINT, then teardown+redeploy on
10.13; the re-IP is a D-115 supersession + D-101 termination, NOT YET RULED.
- Sweep docs/audit/queued-findings-20260808-...reip-pivot.txt (F1-F16): the pivot,
the D-115 conflict, MAAS profiles missing on the racks (fix existing+rebuild),
no service migration (DC0>MAAS-regional>MAAS-rack), dc0 live inventory,
checkpoint scope, + instrument-currency #22 (F3 wrong-store retraction; pkill
self-match).
- 10.13 NetBox subnetting DRAFT (Task #2; octet-preserving proposal, not ruled).
- changelog-20260807 F3 retraction/correction; bounded ledger close summary.
Gauntlet ALL GREEN (102); repo-lint 0 fail. Tasks #1-#4 pinned.
Status ONLY in CURRENT-STATE.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-07 |
savegame 2026-08-07 (part 2): bookend + sweep -- dc1 region standup + SEC-031 edge rebuilt
...
GA-R4 bookend for the dc1-region-sequence session (F NetBox importer; dc1 region
topology/IPAM/DHCP-cutover/power-key; SEC-031 edge rebuild + runbook/tool). Bounded
ledger summary + rotation (2026-08-05 x2 archived, now 283 lines). Sweep
docs/audit/queued-findings-20260807-dc1-region-sequence.txt: 3 FIRST SURFACE
(console-driver harness owed; run-logged gap; jammy source-index trap). CURRENT-STATE
records the standup progress (steps a-b done, D/E remaining). Gates: repo-lint 0-fail,
gauntlet ALL GREEN (102), ledger-scan SEC 29->28.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
SEC-031 CLOSED: dc1 edge rebuilt + procedure captured as runbook + tool
...
Rebuild the damaged dc1 OPNsense edge via the proven dc0 procedure (operator-
directed). dc-egress-check dc1 8/8; .6 region reaches the internet (ping 1.1.1.1,
curl images.maas.io 200) -- unblocks jammy image sync. Edge-only tofu -replace
(machine-asserted 2/0/2, nodes protected) -> console bootstrap -> WAN/LAN
addressing -> automatic outbound NAT.
Fix the root gap the operator flagged: the dc0 rebuild lived only as an audit
capture, forcing dc1 to reconstruct it. Now a first-class runbook
(runbooks/dc-edge-rebuild.md, site-parameterised) + a site-agnostic tool
(scripts/opnsense-console-rebuild.py, replaces per-DC one-off drivers). SEC-031
closed; CURRENT-STATE + changelog Item 4 updated. Also lands the Stage-5 dc1
region-standup captures (topology, DHCP handover).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

NetBox util-host importer + as-built utility hosts landed into office1-netbox
...
Close the tool gap on the operator-flagged NetBox-pending list: the D-134
utility RANGES were in the apex but no importer recorded the individual
utility-HOST addresses. New netbox/dc-util-hosts-import.py (DERIVES plane CIDRs
from lib-net + host octet from lib-hosts; whole-plan preflight; range
precondition; dns-collision guard; SANDBOX + upstream-write gates; dry-by-
default) + harness tests/dc-util-hosts-import/ (19/19).
Landed 8 ip-address objects (ids 187-194; apex ip-addresses 186->194, idempotent
re-run EXISTS/0): dc0 .5 juju-01, .6 maas-01, .7 tailscale-01; dc1 .6 maas-01
(both planes each). dns_name carries the ruled vr1-dc<N>-<role>-NN names -> dc0
renames recorded by construction.
Deferred as findings: the .4 artifact host (metal-admin-only, per-DC divergent,
no repo name -- needs an operator naming ruling) and dc1 .5/.7 (record when
live). Gauntlet ALL GREEN (102); repo-lint 0-fail/1-legacy-warn. Revert = NetBox
DELETE of ids 187-194. Body: changelog-20260807-dc1-region-sequence.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
NetBox-pending: queue this session's new IPAM objects for office1-netbox (operator-flagged)
...
office1-netbox (10.10.1.10) is the live VR1 IPAM apex (DOCFIX-195); this
session's new objects are not yet in it. Enumerated in changelog Item 11
(dc0 .7 + dc1 .6 region VM + new vr1-dc1-region + dc1 .7-when-live +
D-134 utility .4-.9 assignments + dc0 renames) and flagged as a
do-not-miss NetBox item in the next-session NEXT (CURRENT-STATE + ledger)
so it isn't lost before the end-of-deployment write-back.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
vr1-dc1-region profile registered + verified; session close consolidated
...
Registered the vr1-dc1-region profile on voffice1 (dc0-pattern tunnel
-L 5243:10.12.68.6:5240 via the rack; apikey piped .6->`maas login -`
stdin, never exposed). Verified: rack vr1-dc1-maas-01 (qtw8pm), 0 machines
(empty -- the 9 nodes+juju+.7 get rebuilt in). changelog Item 10 + a
handoff block enumerating the remaining config/rebuild chain.
Consolidated this session's ledger block to a bounded GA-R4 summary
(dc0 .7 carved-and-ready + D-134 naming amendment + vr1-dc1-region LIVE);
ledger 296/300, repo-lint 0 fail.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

vr1-dc1-region MAAS is LIVE -- region init done on the .6 (operator-authorised one-shot)
...
Located the shared vr1-office1-svc key ON VCLOUD (~/vr1-office1-creds/
office1_svc_ed25519, fingerprint matches the injected key -- operator's
"what was used on DC0"); reached the .6 from vcloud (key stays local,
ProxyCommand via voffice1->rack). Snap egress works via the snapd proxy
10.12.68.2:8000. Installed maas 3.7.2 + postgresql 16.14. Operator
authorised ("You run it"); ran the file-staged credential one-shot (creds
generated + stored 0600 on the .6, NEVER in my context): DB role+db,
maas init region+rack, createadmin. http://10.12.68.6:5240/MAAS/ -> 301.
OWED: consolidate the .6 creds to ~/vr1-dc1-creds/ (SEC-020/D-137).
NEXT: register vr1-dc1-region profile + config via tested tools + rebuild
9 nodes+juju+.7 fresh into it. changelog Item 9 + CURRENT-STATE + ledger.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
dc1 .6 region VM DONE to Deployed jammy (carved, both legs live) -- ready for MAAS install one-shot
...
Extended aux-carve for -maas-01 (9fea7c1); carved the .6 in Office1
(--profile admin, all 3 racks; metal-admin 10.12.68.6 + provider-public
10.12.64.6, pass=6/0); MAAS-deployed jammy -> Deployed, power on, both
legs ping 0% from the rack (transient power/virsh flakes cleared on
re-query).
NEXT is the MAAS region install/init on the .6 -- OPERATOR ONE-SHOT
(createadmin=SEC-020 + reaching the .6 needs the operator's
vr1-office1-svc key). Then register vr1-dc1-region + config via tested
tools + rebuild 9 nodes+juju+.7 fresh into it. changelog Item 8 +
CURRENT-STATE + ledger NEXT.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

dc1-region workstream START: .6 region VM bootstrapped to Ready + MAAS 3.7 procedure researched
...
Operator ruled dc1 node-handling "Rebuild fresh into dc1-region". Measured:
dc1's full set is in the Office1 region (9 nodes+juju Ready; .6/.7 Failed
commissioning on unset-power). Bootstrapped the .6 region VM (Office1-side,
like dc0 hot-kid): power set, renamed normal-piglet->vr1-dc1-maas-01
(convention), recommissioned -> Ready (a transient virsh-login error
cleared on retry).
Researched + recorded the MAAS 3.7 region+rack install/init (changelog
Item 7; matches Office1 3.7.2 + PostgreSQL 16; external DB required;
createadmin=SEC-020 operator one-shot). Remaining next-session: extend
aux-carve for -maas-01 + carve .6 (10.12.68.6 + 10.12.64.6) before deploy
(hot-kid under-carve lesson); deploy jammy; MAAS init one-shot; register
region + config via tested tools; rebuild 9 nodes+juju+.7 fresh into it.
CURRENT-STATE dc1 block + ledger NEXT updated. repo-lint 0 fail.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
D-134 amendment: MAAS hostname naming convention (standing) + dc1 rebuild-fresh ruling
...
Operator: "Record that as the preferred naming convention going forward."
Recorded the VR1 MAAS hostname convention (set vr1-<dc>-<role>-NN ==
libvirt domain == power_id after commission/deploy; random-by-default is
not leave-it-random; standup DoD = 0 non-vr1-<dc>-* names) as a D-134
AMENDMENT (per-DC identity standard; ARCH, no new number) + lib-hosts
comment.
dc1 node-handling RULED "Rebuild fresh into dc1-region" (build region,
then power/enlist/commission/deploy the 9 nodes+juju FRESH into it, not
delete+re-enlist) -> CURRENT-STATE dc1 block + ledger NEXT.
repo-lint 0 fail; gauntlet ALL GREEN (101).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
dc0 MAAS hostname convention: rename tailscale + juju controller (operator correction)
...
Operator: "You are not following naming conventions." Renamed the two
random-named vr1-dc0-region machines to convention (hostname == libvirt
domain == power_id, lib-hosts:26): known-marten->vr1-dc0-tailscale-01,
subtle-grouse->vr1-dc0-juju-01. All 11 dc0-region machines now vr1-dc0-*.
Cosmetic to tooling (carve/power resolve by boot MAC); load-bearing for
operability. changelog Item 6 + CURRENT-STATE note.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Session bookend 2026-08-07 (dc0 tailscale .7 carved-and-ready): GA-R4 + sweep
...
Bounded ledger summary + rotation (08-04 block -> archive, 296/300); sweep
docs/audit/queued-findings-20260807-dc0-tailscale-provisioning.txt (F1-F4
first-surface); CURRENT-STATE + changelog ping-verification edits (both .7
legs live). Deliverable d36d815..c8ddfb6 already pushed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
dc0 tailscale .7: carve applied+verified + deployed jammy (carved-and-ready)
...
LIVE (gated): carved the .7 router's two legs (metal-admin 10.12.8.7 +
provider-public 10.12.4.7 VLAN 5002, no br-ex) via the new aux-carve;
check pass=8/0. MAAS-deployed jammy -> Deployed, sshd live on 10.12.8.7.
State: dc0 .7 = carved-and-ready. Two JOIN prerequisites remain, both
off-session: (a) tagged pre-auth key + Headscale autoApprovers/ACL
(operator key is PLAIN; join NOT attempted); (b) SSH access via
vr1-office1-svc (region injects only that key; operator holds it).
CURRENT-STATE tailscale block + changelog Items 3-5 updated; dc1 ruling
recorded (build vr1-dc1-region first, no migration).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Session close 2026-08-06/07 (GA-R4 bookend): phase-03 Step 3.4 G3 + per-DC Tailscale rulings & .7 VMs
...
Bounded ledger summary + sweep (queued-findings-20260807, O1-O5) + CURRENT-STATE
tailscale-build status + node-vm harness reconcile (11/66 -> 12/72, red gate that
my substrate commits caused by not re-running the gauntlet -- the 2026-07-30 lesson
repeated). Rotated the 2026-08-03 close to archive (ledger 298 < 300).
Session delivered: Step 3.4 G3 domain-manager probe (built + live PASS, closes the
last phase-03 exec item); Decision C (Horizon reconciled to VR1, Step 3.3 splits to
its own tailnet-gated row); D-129(iii) amendment rulings a-d (dedicated .7 VM / STAR
/ single-HA-pinned / SNAT-on, both DCs) + D-134 octet .7 + D-107 citation DOCFIX;
site-tailscale.sh tooling; the .7 subnet-router VMs applied + MACs pinned on both
DCs (dc0 tailscale; dc1 region + tailscale). Headscale-side join deferred (no
control-plane access); MAAS commission/deploy/carve + dc1 region setup owed.
Gates: repo-lint 0 fail; gauntlet ALL GREEN (101). Memory: +ipv6-primary-posture
(drift-prevention, operator-directed) + instrument-currency #19.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

D-129(iii) amendment: per-DC Tailscale operator-access rulings (a-d), both DCs
...
Record four GA-R5 rulings (2026-08-07) that pull the gap-21 per-DC Tailscale
subnet-router forward to close phase-03 Horizon properly (operator: "pull the
tailscale steps forward"; "plan and push to both DC0 and DC1"):
(a) dedicated VM at utility .7 (10.12.8.7 / 10.12.68.7)
(b) STAR -- operator->DC only (the Headscale ACL / security boundary)
(c) SINGLE router, HA scale-up PINNED
(d) SNAT ON now, source-IP preservation PINNED
These are D-129(iii) implementation sub-decisions (not a new D-number). Also:
correct the D-107 citation defect (D-107 is airgap/mirror/NTP, governs no
Tailscale; D-129(iii) governs); extend D-134's standing octet map to .7;
update the gap-21 register row and CURRENT-STATE (phase-03 does NOT close this
session -- Step 3.3 Horizon splits to its own gate row, gated on the tailnet
build + vault CA on the workstation + a browser login over the tailnet).
Records only -- no code, no cloud change. The build (site-tailscale.sh + the
.7 VM per DC + Headscale star ACL/autoApprovers + SEC row) is next.
repo-lint 0 fail; Decision C reconciliation measured read-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
| 2026-08-06 |

phase-03 Step 3.4 (dc0): G3 domain-manager probe -- named check + live PASS
...
Build scripts/g3-domain-manager-probe.sh + tests/g3-domain-manager-probe (harness
12/12, every exit path proven) as the GA-R6 named executable check for phase-03
Step 3.4 stage 2 (gate G3), filling the hard-rule-4 gap (was a manual runbook walk
only) and giving dc1's Step 7 a reusable probe. Grounded in the real policy
(domain-manager-policy.yaml:103 create_grant managed-role guard).
Ran it LIVE from the dc0 rack (operator-approved): G3 PASS, 7 ok / 0 fail, teardown
verified clean (zero g3-* residue). Stage-1 (PO:) verified read-only: policyd-override
attached, all 3 keystone units 'PO: Unit is ready'. Capture
docs/audit/g3-dc0-probe-20260806.txt.
HARNESS-MANIFEST recorded 99->100; gauntlet ALL GREEN (100), repo-lint 0 fail.
CURRENT-STATE: Step 3.4 -> RESOLVED; phase-03 exit gate now turns on the Horizon
reachable/login-works item (D-044/D-075 per-rebuild + VR0-nginx vs VR1-tailnet access
model = Decision C), to be measured + ruled next.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Session close 2026-08-06 (part 3, GA-R4 bookend): phase-03 core-verify F-CV1/F-CV2/F-CV3 RESOLVED + binding conformance
...
GA-R4 bookend: bounded ledger summary + sweep. F-CV2 (openstack CLI on the dc0
rack), F-CV1 (designate binding, BUNDLEFIX-056), F-CV3 (dashboard TLS via D-072
AMENDMENT VR1 / BUNDLEFIX-057) all RESOLVED; BUNDLEFIX-058 designate-stack
conformance -> binding conformance clean cloud-wide. D-134 Roosevelt-delta +
gap-21 access-model context recorded.
Sweep O11 FIRST SURFACE (dc0 MAAS query method); O10 rack repo-stage bundle STALE
(re-stage before redeploy). Memory instrument-currency #18. repo-lint 0-fail,
gauntlet ALL GREEN (99) at ace0e16, ledger 292 lines (no rotation).
NEXT: Step 3.4 (domain-manager policy) -- last phase-03 exit-gate item -> Steps 8-12
-> Stage-5 exit. Body: docs/changelog-20260806-phase03-coreverify.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-058: designate-stack binding conformance (amqp/cluster -> metal-internal) + ceph-rbd-mirror Stage-6 watch
...
Operator-requested binding-conformance sweep after F-CV1/F-CV3. The only remaining
ACTIVE svc-endpoint deviations were the designate stack: designate amqp (->rabbitmq)
+ cluster (peer), designate-bind cluster (peer) -- all on the metal-admin default,
deviating from the generic svc-to-svc rule (14 apps use metal-internal). No ruled
exception -> conformance repair, no new D-number.
Benign (designate served metal-admin so nothing broke, unlike F-CV1/F-CV3) but
off-plane. Repaired to metal-internal so dc1 inherits a fully-conformant bundle.
Proven live (same method): juju bind designate amqp=metal-internal cluster=metal-internal
+ juju bind designate-bind cluster=metal-internal (individually gated). Verified no
regression: bindings moved; designate still Stage-7-blocked on nameservers ONLY;
designate-bind active/idle; designate haproxy 0 DOWN (F-CV1 intact). No harness asserts
these bindings. gauntlet ALL GREEN, repo-lint 0-fail.
RECORDED: ceph-rbd-mirror certificates/cluster on metal-admin are UNBOUND today
(Stage-6 DR not wired) -> sweep O9 STAGE-6 WATCH + CURRENT-STATE, pre-check before
wiring at dc-dc-phase5. No other active deviations found cloud-wide.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-057 / D-072 AMENDMENT (VR1): dashboard cluster -> metal-internal (F-CV3 RESOLVED)
...
VR1 split-metal INVERTS the D-072 cluster placement. The openstack-dashboard charm
declares no admin/internal extra-binding and no os-*-network (metadata + charmhub
docs), so apache serves its SSL vhost on metal-internal while haproxy dialed
cluster=metal-admin -> vhost-less -> plaintext (the D-072 trap, inverted). Option A
(serve metal-admin) unavailable in-deployment (no charm lever).
Retire the VR0 exception: openstack-dashboard cluster -> metal-internal (generic rule
+ the served plane). Proven LIVE before ratifying (operator process directive), then
ratified GA-R5 "Ratified, land the config-of-record".
bundle.yaml cluster=metal-internal; design-decisions D-072 AMENDMENT (VR1);
binding-reference matrix + exception RETIRED + cross-ref; CURRENT-STATE/sweep/changelog
F-CV3 RESOLVED. Verified live: provider + operator metal-admin VIPs both TLS 200
CA-verified (were plaintext); cert covers all 3 VIP IPs (resolves AH01909); units
active/idle. No harness asserts this binding. gauntlet ALL GREEN, repo-lint 0-fail.
dc1 inherits via the shared bundle (no rebind).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|
Gap-21: capture the operator-access model (tailnet -> metal-admin dashboards) as forward context
...
Operator discussion 2026-08-06: operators reach each DC's metal-admin plane over
the tailnet and access routed dashboards (Horizon) from there. Added as dated
CONTEXT (not a ruling) to workflow gap-register item 21, informing the four
deferred Tailscale sub-decisions when re-raised. Captures the implications that
reach outside the Tailscale build: the operator-facing reverse proxy becomes
unnecessary; D-044/D-075 per-rebuild accommodations sunset for the operator path;
cert SANs on the metal-admin VIP (and FQDN certs, D-106/D-008) become load-bearing;
and it reinforces metal-admin as the dashboard's serving plane (the access-model
argument for F-CV3 option A over option B). Cross-referenced from the F-CV3 sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

BUNDLEFIX-056: designate public/internal bindings -- F-CV1 RESOLVED (fix executed + verified)
...
designate's REST API public+internal endpoints were omitted from the bundle,
defaulting to the '' metal-admin fallback -> orphaned the provider + metal-internal
legs of its ruled .62 VIP triple (D-020 amendment) -> _admin haproxy backend SSL-DOWN
on the unserved metal-internal address. The deployed charm declares public/admin/
internal extra-bindings (metadata verified); dnsaas (D-106) is ADDITIONAL, not a
replacement -- the prior "no public binding" reading was the defect's root.
Config-of-record: bundle.yaml +public:provider-public +internal:metal-internal;
provider-bundle-check EXPECT_PUBLIC_VIP 11->12 (vault stays out, not 13) with T16c/T16d
failing-direction tests (57->59, ALL PASS); network-space-binding-reference row 88 +
note. Gauntlet ALL GREEN (99); repo-lint 0-fail.
Live (operator-approved): juju bind designate public=provider-public
internal=metal-internal (rc=0, dc0 rack). Verified: full haproxy sweep 0 DOWN
cloud-wide; apache https vhosts span all 3 planes; cert reissued for provider-public;
catalog triple correct (public 10.12.4.62 / internal 10.12.12.62 / admin 10.12.8.62).
Governing: D-052 / D-020 amendment. Evidence:
docs/audit/stage5-dc0-phase03-coreverify-20260806.txt. Body:
docs/changelog-20260806-phase03-coreverify.md (Item 6).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 F-CV1: reconcile after research -- :internal defect CONFIRMED, :public OPEN; D-141->D-052
...
Prior-art + governing-decision research on the designate bind-plane mismatch.
Corrects the earlier commit (aaeee93) which reached "CONFIRMED" and cited D-141
before checking the governing decisions (backwards ordering, memory #17).
D-072/BUNDLEFIX-011 prior art: the dashboard-plaintext class was root-caused in
VR0 (charm renders haproxy 443 backend on cluster-binding addr, apache SSL vhosts
only for default+public). dc0 bundle already carries that fix -> F-CV3 is a NEW
cause, parked for its own triage.
F-CV1 re-scoped: governing surface is D-052 + the generic binding rule, NOT D-141.
`:internal`->metal-internal is a CONFIRMED defect (on the metal-admin fallback;
every sibling binds it to metal-internal; no ruled exception; cert already covers
the metal-internal SANs; all 3 units have metal-internal addrs -> juju bind won't
be refused). `:public`->provider-public is OPEN -- designate uniquely carries
`:dnsaas` on provider-public (D-106 dual-VIP); its REST API may be intentionally
metal-admin-only. The per-app table is generated from bundle.yaml (descriptive of
the defect), not intent.
Fix method = D-072 precedent (bundle + live juju bind + haproxy-readback verify).
designate is Stage-7-blocked -> no urgency. Awaiting operator ruling on :public.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 Step 7: triage F-CV1 (CONFIRMED designate bind-plane mismatch) + F-CV3; sweep
...
Operator-authorized triage of the two plaintext-vs-TLS findings (read-only; not
fixed, hard rule 1). Both apps have vault certs rendered + the certificates
relation, so NOT the ovn CN-issuance class -- charm apache-TLS-frontend layer.
F-CV1 CONFIRMED (two findings, not one): designate/0 apache https frontend binds
only 10.12.8.198:8991 (metal-admin); haproxy's _admin backend dials
10.12.12.110:8991 (metal-internal) where no SSL vhost exists -> check-ssl hits
plaintext -> DOWN. VR1 dual-metal-plane bind mismatch (D-141), structural (not
Stage-7 collateral). F-CV3 (dashboard) is separate: :433 served by Ubuntu
default-ssl.conf, charm https frontend not effective; root cause not nailed.
Owned instrument caveat: earlier "no SSLEngine in sites-enabled" was a grep -r
false negative (does not follow the symlinks); apache2ctl -S corrected it.
Remediation = a focused, gated session. Sweep:
docs/audit/queued-findings-20260806-phase03-coreverify.txt
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|

Stage 5 dc0 Step 7: phase-03 core-API VERIFIED; F-CV2 openstack-client on rack; F-CV1/F-CV3 logged
...
Core-API layer of phase-03 core-verify (adapted for vr1-dc0, run from the dc0
rack per D-138) VERIFIED read-only: settle walk (falsifiable prediction matched),
haproxy 0-DOWN across 12 activated VIP apps, admin-openrc scoped token, IP-only
endpoints, two-sourced keystone VIP, vault CA TLS OK.
F-CV2 RESOLVED (07-30 queued-F1, hit at Step 7): installed python3-openstackclient
6.6.0-0ubuntu2 on the dc0 rack; CURRENT-STATE section 7 row amended (GA-R1/C1).
Exit gate NOT MET (2 open, GA-R6/E3 no conditional close): F-CV3 dashboard VIP
serves plaintext (apache-SSL-inactive despite certs) -> Horizon exit-gate fails;
Step 3.4 domain-manager policy NOT RUN. F-CV1 designate-api plaintext vs haproxy
check-ssl (same shape); "collateral of block" reading retracted. Logged not fixed.
Evidence: docs/audit/stage5-dc0-phase03-coreverify-20260806.txt
Body: docs/changelog-20260806-phase03-coreverify.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fg98z7QyzwYUs8fsWCn728
|