| 2026-07-31 |

Deploy blocker: mirror lacks jammy-backports. Queue per-DC Tailscale (item 21)
...
BLOCKER, logged not fixed (D-135 scope is a ruled surface). 22 of 33 units in
error, all hook failed: "install", all one cause:
E: The repository 'http://10.12.8.4/ubuntu jammy-backports Release'
does not have a Release file.
Measured: the mirror serves jammy/jammy-security/jammy-updates 200 and
jammy-backports 404 -- exactly D-135's "jammy triple" scope -- while the node's
sources.list carries jammy-backports, because the Step-3.5 apt-mirror
model-config rewrites EVERY suite to the DC mirror. Machines are all
started/running; this is purely the artifact layer and units retry.
dc-mirror.sh check dc0 PASSES while this is true -- it verifies sync STATUS,
not suite COVERAGE against what the deployed image asks for.
QUEUED BY OPERATOR DIRECTION, not built: a per-DC Tailscale subnet router is
now a standing DC-standup requirement, dc0 first. Home of record is the tooling
gap register item 21 (new), with a per-DC definition-of-done, the
vendor-documented build constraints, and four sub-decisions needing GA-R5
rulings first: the utility-band OCTET (D-134's map is a standing cross-DC
standard), the admin-reachability model (star vs mesh -- the real architecture
question at region-region scale), HA count, SNAT on/off. This is EXECUTION of
the already-ruled D-129(iii) shape, not a new decision.
Measured while scoping: SEC-010 does NOT need relaxing (the forwarding to
permit is on the router VM, not the transit leg); voffice1 holds no route to
10.12.x at all, so the tailnet stops at Office1 by construction; the existing
Office1 node is UNTAGGED and carries a 180-day key-expiry clock on the
operator's only tailnet path; and D-129(iii) cites D-107 as governing Office1
while D-107 rules nothing about Tailscale -- the second instance of the F7
miscitation class, this time inside a ruling.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-30 |

DOCFIX-205 sweep: correct the artifact that filed Q1, and one stale residue
...
Advisor review caught that correcting CURRENT-STATE and the ledger was not
enough: docs/audit/queued-findings-20260730.txt -- the sweep capture the 07-30
bookend points at -- still presented Q1 as an open ruling that 'blocks any FQDN
cert', enforced by an A12 refusal that no longer exists. A next session reading
it would re-derive Q1 exactly as this one did, reopening the loop just closed.
- queued-findings-20260730.txt: APPENDED correction (not a rewrite; the
stage4-mirror-gate-20260727.txt precedent), naming the two claims that are now
false and pointing at the evidence. Q2 explicitly untouched.
- dc-dc-deployment-workflow.md:257: described phase-6 as proposing the retired
dc1/dc2-hostnames.yaml spelling. MEASURED: the runbook itself was already
corrected to the region-qualified form by an earlier session -- only the
workflow doc's description of it lagged. Under D-119 IS vr1-dc0, so the
-hostnames.yaml interpolation used elsewhere is correct as written.
- The DOCFIX-205 capture's section 1 quoted a 3-line span starting mid-sentence
and left a clause orphaned. Repaired to two verbatim quotes, both re-verified
present in design-decisions.md. CURRENT-STATE cites this file as evidence and
this repo's discipline is exact captures.
octavia-pki 23/23; repo-lint 0 fail (L10 gated this commit correctly -- the
docs/audit/ change set required CURRENT-STATE to move with it).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-29 |

Items 3.9 + 3.10: the teardown false clear existed in three places, not one
...
3.9 was the readiness audit's most dangerous item -- the teardown runbook is what an operator
reaches for DURING a failed Stage 5, under time pressure.
The false clear was in THREE places, not the one the audit named. Besides Step 2's vm-host
read -- which can never return, since this repo uses per-machine power_type=virsh and
instantiates no maas_vm_host module -- the same wrong premise sat in the "READ BEFORE ANY DC
TEARDOWN" header telling the operator to remove the maas-vm-host record, and in the
Relationship-to-D-061 claim that no VR1 DC had reached Stage 4. The header instance is the
consequential one: it is read FIRST, so it bypassed any fix confined to Step 2.
Step 2 is now a two-lens MACHINE census run from the headend (maas is measurably absent on
vcloud): lens 1 enumerates what exists and ends in a countable RECORDS REQUIRING ATTRIBUTION,
lens 2 corroborates against lib-hosts pinned boot MACs and exits 1 on any hit. Demonstrated
three ways against a fixture -- records present -> exit 1, zero -> exit 0, empty roster ->
REFUSE -- so it is a gate rather than a formality. The 2026-07-21 pod-cascade precedent is
retained as the reason associations are read before any destroy.
Step 3's six phantom module.dc1_* targets are retired; all 8 targets now resolve, verified
independently against ^module "X" across all three roots. The real insight: scoping a DC is a
ROOT choice, not a -target choice, since each substrate root holds exactly one site.
Step 4's VERIFY moved to qemu+ssh from the headend behind a virsh version REFUSAL, because an
unreachable URI, a stopped VM and a bad key otherwise return the same empty result. The
decision tree gains "no branch reaches a destroy without Step 2 passing" -- it never mentioned
the MAAS gate -- and the virsh destroy vs tofu destroy verb distinction. Also fixed: Step 1
backed up the wrong state file, "Two paths" pointed twice at a nonexistent Step 6, and $REPO
silently meant two different clones.
3.10: item 17 CLOSED with measured evidence, and its own stated fix corrected (it closed by
D-125 bridge-in, not the replication the entry claimed). Item 19 disambiguated 19a/19b rather
than renumbered, because both are cited by number from outside the file. Item 20 MEASURED, not
asserted -- verdict no leg required, with the rule mismatch written in rather than resolved
silently, plus an expiry condition; the voffice1-side reboot durability recorded as UNMEASURED
with the commands that would resolve it.
repo-lint 0 fail / 615 files, both files ASCII+LF.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Octavia: R8a BUILT (record was false), two stale surfaces, R7 generator de-frozen
...
Follows an upstream/vendor research pass answering "can Octavia run full
IPv6" -- yes; nothing inside Octavia requires IPv4. charm-octavia creates
the lb-mgmt subnet with ip_version: 6 and has no IPv4 code path at all; the
health manager is v6-capable both directions; the amphora agent binds '::'
and verifies certs by amphora UUID, not IP; and Nova metadata is not a
dependency (config_drive=True). Surviving v4 pressure is soft or external.
R8a BUILT. phase-05-octavia-verify.sh now compares o-hm0's MTU against
lb-mgmt-net's (resolved BY TAG, never by name), fails in either direction,
and REFUSES when a value cannot be read. This closes the LP#2018998
obligation, live on our exact pin after its 2025-12-31 recurrence.
THE DECISION RECORD HAD CLAIMED THIS SHIPPED. design-decisions.md carried a
present-tense "Delivery: the assertion ships in ..." written on the day of
the ruling, while the file held zero MTU references and its last commit
predated the ruling by a month. Corrected in place. Sharper than the usual
ruled-but-not-built class: the decision doc was the false witness.
TWO OWNED ERRORS:
- I concluded "no octavia harness exists" from a NAME grep. The script's
harness has always been tests/phase-05. Duplicate deleted, R8a cases
folded into the real one.
- My first MTU draft ABORTED the script: a $( ) under set -euo pipefail with
inherit_errexit exits the run rather than reaching the refusal branch. The
pre-existing harness I had just declared nonexistent caught it. Same class
as commit 1's IFS word-splitting trap; both now locked by cases.
Two live surfaces still contradicted R8 -- dc-dc-phase4 and the workflow doc
both called lb-mgmt IPv6 "a real, open risk" and recommended keeping it
v4-only, citing the two refuted bugs. A reader would have re-litigated a
closed ruling in the wrong direction. Both superseded in place.
R7 executed: phase-01 Step 1.0-GEN gains a $DC selector; the baked
/CN=VR0 DC0 CA subjects derive from it, and the VIP gate derives this DC's
provider prefix from the same overlay it read the VIP from -- R7's explicit
"read the MERGED input" caveat. dc1's 10.12.64.57 used to hard-ABORT, so no
dc1 artifact could be produced. Verified both DCs, negative control still
aborts.
Logged not built: the controller cert's CN/DNS SANs still carry dc0.vr0
(outside R7's ruled scope, inert while os-public-hostname is unset), and it
gains no IPv6 IP SAN though the VIP is dual-family -- a gap, not a break,
and adding v6 SANs needs its own ruling.
tests/phase-05 14 cases (was 9); gauntlet ALL GREEN (85) on vcloud;
repo-lint 0 fail / 611 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-28 |

D-136 ADOPTED at option (D); IPv6 host-part convention RULED (both GA-R5)
...
Operator utterance, quoted verbatim in both entries: "Rule D-136 option D,
mirror the v4 octet for both legs but make a note to review at the end of the
project to review for adjustment."
GA-R5 NOTE, recorded rather than glossed: this single exchange carried TWO
rulings, and GA-R5 says one per exchange. Both recorded because each clause is a
specific unambiguous answer to a question already put -- not the template
"yes to all" the rule exists to reject. Flagged to the operator; either can be
re-put separately and re-taken.
D-136: option (D) was WRITTEN INTO the entry before being adopted, because the
directed shape matched none of the recorded (A)/(B)/(C) and adopting an option
the entry did not contain would leave a future reader unable to find what was
decided. (D) populates the apex first, then builds the renderer, then dry-runs it
without gating this deployment -- differing from (C) by closing the apex gap now
rather than hand-maintaining VIPs/bands indefinitely, and from (A) by keeping an
unproven tool off the deploy's critical path. Its validation advantage over both
is the reproduction baseline frozen earlier today. Now ADOPTED; ledger-scan
confirms it dropped off the open-decision list (4 -> 3) and the ledger's
machine-derived block is re-seeded.
IPv6 host part: mirrors the v4 octet on BOTH leg types, closing the mapping R4's
D-134 amendment left explicitly unruled -- the blocker on populating any v6
value, since the apex holds every v6 prefix and zero per-address v6 objects.
Accepted cost recorded as deliberate: the GUA leg's dedicated VIP /64 is used
from ::50 up, leaving ::1-::49 unused, in exchange for one rule per plane instead
of a per-leg special case in the renderer. Review owed at the close of THIS
deployment as forward item F4 -- whose trigger differs from F1-F3, so the section
intro now says each item carries its own trigger.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Render pipeline step 1: carve-outs pinned, reproduction fixtures frozen
...
CARVE-OUTS PINNED at operator request ("in a place they will not be lost as
they will be required items in the future"). New forward items in the workflow
doc: F2 the render-pipeline AUTOMATION half (CI runner / event delivery /
status-back), deferred because two GitBucket requirements are unverified --
reliable webhook emission with retry+replay, and a status-back API without
which the approve-on-rendered-diff gate cannot see a validator verdict -- plus
open Jenkins placement; and F3 the NetBox scope boundary for MACs and VLANs,
permanent rather than pending.
D-136's MAC scope-out was CORRECTED AT SOURCE rather than only contradicted
elsewhere. Its stated reason -- MACs "drift-free across substrate main.tf <->
lib-hosts.sh <-> bundle ovn-chassis <-> discovery" -- is falsified: bundle.yaml
:462 carries 52:54:01:d1:04:02 / :05:02, which lib-hosts:132 identifies as
vr1-dc1's scheme, in the file documented as vr1-dc0's source of truth. The
correct reason is availability (no dcim/devices or dcim/interfaces in the apex
at all), and the consequence changes: ovn-chassis becomes renderer OUTPUT and
must leave the base bundle.
STEP 1 -- tests/render-baseline/ freezes the two known-good artifacts a renderer
can be validated against, BEFORE the R2/R11 reconciliation destroys them. After
that lands there is nothing left to diff against and a renderer's first output
is its own first draft, which is D-136's own objection to option (A).
The harness asserts fixture INTEGRITY and deliberately not agreement with live
-- live is supposed to diverge, and asserting against it would turn the harness
red for the change it exists to support (the creds-matrix T24 trap).
Proven able to fail before being trusted: seeded a silent fixture edit (caught
by T2b) and an unpinned file in fixtures/ where sha256sum -c passed vacuously
and only the coverage assertion T3 caught it.
Harness 9/9; gauntlet ALL GREEN (82, was 81); repo-lint 0 fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|

Q1 RULED (GA-R5): NOC posture lives in the workflow doc as a named forward item
...
Operator utterance: "use the workflow doc as a named forward item". Executed in
this commit rather than left as a ruling-without-artifact -- the audit found two
RULED-BUT-NEVER-BUILT decisions and this one is a doc edit, so there is no reason
for it to become a third.
docs/dc-dc-deployment-workflow.md gains a "Forward items -- DEFERRED BY RULING to
the next deployment" section, opening with why forward items are NOT D-numbers
(GA-R3 rule 1(b) as amended by A1, quoted). F1 records the four-step Office1 ->
DC0-as-NOC -> remote-push-to-DC1 test verbatim, why it was deferred, what already
supports it so the next deployment does not start cold (D-100's management-only
Office1<->DC fiber; D-128 already executing Plane 2 on voffice1, which IS the
"technician at Office1" model -- what is missing is the DC0-first ORDER and a
gate proving the push came from there), and what would make it D-admissible.
No D-number assigned; next-free stays 138.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-27 |

G17 opened by operator ruling; Stage 4 DoD bullets 5-6 repaired (DOCFIX-204)
...
RULING (GA-R5), 2026-07-27. Question as presented: "The node-side half of bullet 5. Nodes are
powered off by the READY-handoff ruling, so no node-side probe can run as things stand. Either a
gated rescue-boot check on one node per DC now (closes it inside Stage 4), or split it into its
own gate row targeted at Stage 5 first boot (GA-R6 E3 explicitly permits this; a conditional
close is not permitted)." Operator answer, exact utterance: "split it into its own gate row".
Recorded in the G17 row of docs/CURRENT-STATE.md (the authority); pushed before dependent work.
G17 "Per-DC artifact source reachable FROM A NODE", [V], OPEN. dc0 -> curl 200 from a booted
node against the D-135 item-1 mirror; dc1 -> apt fetches through the proxy at 10.12.68.4:3142,
since dc1's ruled path is the CACHING PROXY with no node-facing mirror (D-135 amendment) and
checking it as one fails by design. Trigger: Stage 5 first boot. The row states that G17 is NOT
a Stage-5 precondition -- Stage 5's bootstrap needs open edge egress for the juju agent stream
and snaps (D-135 items 2-3 unbuilt), a different path from the apt artifact source.
Stage 4 retains bullet 5's rack-side half: the source answers on its own address with an
attested-current sync. dc1's proxy PASSES; dc0 pends the sync re-run.
DOCFIX-204 -- bullet 6 was UNSATISFIABLE and its stale text had spread to four surfaces. "NTP
from the DC's own OPNsense edge working" is superseded by D-129(iv) (RULED 2026-07-21, "Keep
MAAS hierarchy", no NTP role on the edge): the DoD asked for confirmation the edge is the time
source, which the ruling had already refused, so a correctly-built DC could never meet it.
Repaired in the phase-3 DoD (bullet 6 struck, bullet 5 rewritten per-DC), the phase-3 Step 7
(REWRITTEN -- it branched on gap #5 that D-135 resolved, applied one mirror check to both DCs
which fails on dc1 by design, and asked for checks "from a deployed node" that the READY handoff
makes impossible), the workflow-doc Gate cell (which also still said "Nodes deployed",
contradicting DOCFIX-200), the buildout design, and the phase-4 prerequisites (same "nodes
Deployed" error -- exactly the expectation DOCFIX-200 removed to stop Stage 5 breaking).
phase-4 also gains a block making it the OWNER of G17's captures, warning that first boot is the
observation window and missing it forces a deliberate rescue-boot.
Text and gate-record only; no live state touched. repo-lint 0-fail. DOCFIX next-free -> 205.
Revert: git revert this commit; all changes are documentation and the G17 gate row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HvCyrwvYTTcDYnRErfMsNf
|
| 2026-07-21 |
STAGE 3 CLOSE (operator-ruled): GA-R2 consolidation + skill sweep; gauntlet 76 ALL GREEN
...
- Four execution-session changelogs -> docs/archive/changelogs/;
vr1-stage3-record.md finalized (close summary + session table)
- CURRENT-STATE stage line CLOSED (vr1-dc0 scope; dc1 = HELD G12);
workflow Stage-3 State row + runbooks/README corrected
- Skill sweep: D-131 forwarder invariant + MAAS pod traps + merge
precedent folded; no stale facts found
- Close-out evidence: gauntlet 76 ALL GREEN, repo-lint 0 fail, GA-R7
memory review done (changelog items 11, 17)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013XRtm5sDQgUsZyTJk3k7dz
|
| 2026-07-19 |

Batch 4 C6: docs disposition -- 24 history docs archived, 16 retained, refs rewritten
...
DISPOSITION MANIFEST (archive, one-line justifications): clientdocs-
workflow-review (v1 tenant-docs era) / D-057-DECIDED-append +
D-057-REVIEW-ITEMS + README-D057-PACK + D-058-renumber (executed v1
decision packs) / D-068-openbao-assessment-DRAFT (supplementary; the
cited analysis retained) / DOCFIX-064-phase08-changelist + docfix-draft-
20260702 (executed/draft) / handoff-20260703 + handoff-20260705
(superseded handoffs) / incident-20260712-triplefault (trap routed to
platform-traps 1b/1c long ago) / netbox-vip-queue + phase-00-maas-
standup-notes + script-quality-findings + session-findings-2026-07-02 +
tenant-cidr-overlap-PLAN + v1-pre-deploy-fixes (v1-era executed/dated) /
repo-lint-nextfree-bug-FINDING (absorbed DOCFIX-105/107; repo_lint.py
pointer updated) / upstream-bug-draft-dashboard-tls (draft, dated) /
model-a-fallback-plan (D-123 Model B adopted; fallback preserved in
archive) / stage3-adversarial-review-20260716 (R3-F register,
dispositioned in 2.10) / dc-dc-replication-DR-seed (absorbed per
workflow) / dc-dc-ipv6-charm-research + dc-dc-netem-and-ula-gua-proposal
(stage 4+ research, refs updated). RETAINED 16 top-level, justified in
the session changelog. Live-surface refs to all moved files rewritten to
archive paths (18 files). G3 cell + ledger + changelog updated same
commit (the L10 fire on a Status-line path rewrite is satisfied by the
CURRENT-STATE touch -- the check worked as designed).
Revert: git revert this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VZbUeRE8weySLSGY1odrCi
|

S1/2.10: R3-F residual disposition -- none unaccounted
...
Full register accounting (13 items): F01 CLOSED (R-1/R-2 rulings +
re-attributions, Phase A 2026-07-16); F02 CLOSED (SEC-010 artifact
committed + G10 gate names the pre-apply check; the SEC row itself stays
open on apply); F03/F04/F05 CLOSED (D-107/D-122 amendments, Phase A);
F06/F07 CLOSED (whole-host-budget.py committed + re-validation 838/1024
FIT); F09 CLOSED (D-115 status line, scan-verified); F10 CLOSED (SEC-011
operator ruling). Residuals dispositioned NOW: F08 CLOSED-SUPERSEDED --
the 3-OSD-host premise fell when R-3 added the 4th storage node/DC (one
host of rebuild headroom exists in the ruled 3+2+4); F11 FIXED -- the
workflow authoring cell's "D-101 supernet unassigned" clause dropped
(10.12.64.0/19 assigned by D-115; vr1-dc1 stays HELD on transit/rack
addressing, G12); F12 CLOSED -- the observation's blocker (uncommitted
model) resolved by the committed calculator; transient N+1 amphora churn
sits well inside the measured 186 GiB headroom, and the Stage-6 drill
exercises the real failover; F13 CLOSED as recorded observation (wording
cosmetic; D-122's bullet already amended). Also: readiness C5 queue was
STALE (queued amendments that landed 2026-07-16) -- corrected. Batch 2
item 2.10 (amendment S1).
Revert: git revert this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VZbUeRE8weySLSGY1odrCi
|

GA-F10/H1 + GA-R1/R6: token migration, State-line demotions, E1 type tags (2.5+2.6)
...
Mapping record: docs/audit/h1-token-map-20260719.md -- all distinct
tokens dispositioned, none unmapped. Live-surface migration: workflow
doc (26 flags -> 0; header drops the tracker model, 8 State lines demote
to identity + CURRENT-STATE pointers), runbooks (18 -> 0; close-out
instructions RETARGETED to record status in CURRENT-STATE, not the
workflow tracker), readiness doc (5 -> 0), CURRENT-STATE itself (6 -> 0:
G9 STOPPED -> BLOCKED with named blocker; stale "GA-R1 NOT yet ratified"
header + pre-D-130 divergence rows corrected per C2). Gate rows gain
GA-R6/E1 type tags ([V]/[R]) with legend. CLAUDE.md expectation line
demoted per S5b. History classes (ledger, changelogs, design-decisions)
deliberately untouched -- their disposition is Batch 4/GA-R3 structural.
Extractor re-run: zero flagged occurrences on live non-history surfaces.
Batch 2 items 2.5 + 2.6.
Revert: git revert this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VZbUeRE8weySLSGY1odrCi
|
GA-F03: stale tofu 1.12.3 pins demoted to CURRENT-STATE pointers
...
opentofu/README.md:41 (historical prose keeps its date, drops the pin)
and docs/dc-dc-deployment-workflow.md:47/:360/:875 (the :360 live claim
"binary EXISTS: v1.12.3" was FALSE -- measured 1.12.4) now point at
docs/CURRENT-STATE.md section 7, the only pin authority (GA-R1 rule 3).
The :47 edit also drops a "DONE" status token (vocab overlap with 2.5).
Batch 2 item 2.2.
Revert: git revert this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VZbUeRE8weySLSGY1odrCi
|
| 2026-07-17 |
Gap register #20: durable host->site reachability CLOSED (Office1); skill carries the per-DC discipline
...
- docs/dc-dc-deployment-workflow.md: new tooling-gap item 20 -- the durable
host->isolated-site base leg had no owning tool; CLOSED for Office1 via
scripts/site-baseleg.sh (installed+enabled 2026-07-17; boot-race unproven until
a reboot). Documents the DC generalization so the gap does not reopen per-DC.
- SKILL.md: operating-model invariant extended -- durable host->site reach is
standing per-SITE tooling (DCs included): at each DC standup MEASURE the reach
and, if the DC net is ISOLATED, add a MEASURED site-baseleg.sh leg row + install
as part of standup DoD (never an inferred DC IP; harness enforces it).
Doc-only follow-up to 1958063. Nothing merged to main (stage-close).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMKdXCquMnh1jKH2sWXPMg
|

D-126 base-leg oneshot (site-baseleg.sh) + D-128 operating model; post-restart sweep
...
vcloud post-restart health sweep all-green (host on 6.8.0-136, patch+reboot done;
libvirt/pools/autostart clean; Office1 stack self-recovered end-to-end). Sweep
surfaced a recurring break: office1-local is an ISOLATED libvirt net, so the host's
base L3 leg (10.10.0.10/24 on the bridge + compose route 10.10.1.0/24 via 10.10.0.20)
lives nowhere in the XML and drops on every reboot -- the un-persisted foundation
D-126 assumed but never owned.
- NEW scripts/site-baseleg.sh + tests/site-baseleg/ (24/24): root Type=oneshot unit,
After=libvirtd (+virtnetworkd for modular hosts; ordering MEASURED here = monolithic
libvirtd), re-adds the leg at boot. Bridge DISCOVERED from the stable net name (never
a baked virbrN, PATTERN-1); apply idempotent + root-guarded; check = session-verify.
INSTALLED + enabled this session; boot-race UNPROVEN until an actual reboot.
- docs/design-decisions.md: D-126 amendment (base-leg durability) + NEW D-128 (operating
model -- Claude on the vcloud jumphost; Plane-1 substrate on vcloud / Plane-2 executes
on voffice1; workstation tailnet is the human path; D-107 untouched).
- Skill invariant + routing row (SKILL.md), Cross-cutting bullet (dc-dc-deployment-workflow.md),
session-ledger + changelog.
Gauntlet ALL GREEN (67 harnesses), repo-lint 0-fail. Nothing merged to main (stage-close).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMKdXCquMnh1jKH2sWXPMg
|
| 2026-07-16 |

D-125 correction: uplink /24 is the ruled 172.30.2.0/24, not a HELD tfvar (no new NetBox object)
...
The initial D-125 records wrongly modeled the vcloud ISP /24 as a NEW,
NetBox-assigned HELD value (`var.vr1_dc0_uplink_cidr`) with an OPNsense WAN
re-address. That is factually wrong: bridge-in is single-NAT, so there is
exactly ONE simulated-ISP /24 for vr1-dc0, and D-115 already ruled it
(172.30.2.0/24, role edge, site vr1-dc0). NetBox models the prefix by role+site
and does not care which libvirt host runs the NAT, so the /24's IPAM identity is
UNCHANGED from Model A; only the libvirt realization moved (a vcloud NAT bridged
through vvr1-dc0). OPNsense WAN stays .2. (Caught reconciling the "create netbox
objects" request against dc-edge-wan-import.py, which already registers this /24.)
- opentofu/main.tf: module vr1_dc0_uplink cidr = "172.30.2.0/24" inlined (ruled
literal, matching the inner root's site-wan pattern) instead of the var.
- opentofu/variables.tf: REMOVED variable vr1_dc0_uplink_cidr (replaced with a
note explaining why it is a ruled literal, not a tfvar).
- Records corrected everywhere the false claim propagated: D-125 body + D-122
amendment, changelog-phaseC2-D125, model-a-fallback (B->A delta item 8 +
revert), session-ledger, dc-dc-deployment-workflow, phase-2 runbook.
Net effect: D-125 adds NO new NetBox object and NO new HELD gate. The only
NetBox work left for the DC substrate is D-124's transit /30 + rack IP.
Both roots validate; gauntlet ALL GREEN (63); repo-lint 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ck6xh3jWQi5b3Su8Dx1LEH
|

docs: reconcile workflow-doc Stage-3 section against repo (post D-121..D-125)
...
The dc-dc-deployment-workflow.md Stage-3 (Phase 2) section + gap register had
drifted behind the rulings that landed through 2026-07-16. Repo wins; each
correction grep-verified against the tree (hard rule 2):
- Build/Gate/Owns rows: single-root -> D-123 Model B two-root
(outer->bootstrap->inner->maas-register); vvr1-dc0 IS the rack; D-125
bridge-in per-DC ISP egress; Owns += D-121/122/123/124/125.
- Authoring status: node sizing RULED (D-121/R-3 = 9/DC), config.xml retired
(D-113(a2) REST), both roots instantiated + validate.
- Replaced the stale "two blocking rulings landed 2026-07-15" note (D-121
"3 storage/16 VMs, 790 GiB, vault fork OPEN") with a current rulings summary:
9 nodes/DC (R-3), FIT 870/1024 (Model B), vault RULED (R-2), + D-123/124/125.
- Rack source no longer a missing-module gap: vvr1-dc0 stood up by
site-headend-install.sh --host-nodes.
- Gap register: D-113(a2) OPEN->DONE in both places (banner + line 838)
(template deleted, config_seed/config_iso_path retired). Stage-4
"node count/sizing undecided" -> RULED.
Bottom line now stated honestly: Stage-3 authoring complete; execution is
operator-gated (HELD NetBox addressing + deploy-time gates). Doc-only; no
code/behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ck6xh3jWQi5b3Su8Dx1LEH
|

Stage 3 (vr1-dc0) prep batch: D-121..D-124 rulings + substrate/rack wiring, HA overlay, NetBox importers
...
Prepares the VR1 Stage-3 DC substrate so the live steps reduce to review-plan-and-apply.
Authoring + offline validation only -- NO cloud mutation; every tofu apply, NetBox --commit,
and rack install stays operator-gated. Four decisions ruled this batch, all recorded:
- D-121 (ADOPTED): make the decorative single-unit control plane REAL. 14 services scale 1->3
(D-009's mechanical scale-up pulled into VR1); node layout = Option C (3 control + 2 compute +
3 storage/DC, sizing set) after a whole-host resource validation (256 vCPU/1024 GiB/10 TiB);
vault sub-ruling = v-a (MySQL-backed, charm-verify gated). Placement-balance rule for Stage 5.
- D-122 (ADOPTED): VR1 site shape -- nested-per-site containment, dark fiber East-West only,
dedicated per-site L3 ISP uplink, 2-NIC fabric-routed DC edge (LAN=provider-public, WAN=uplink),
per-site MAAS. Edge NIC count baremetal-matched (nodes 6, edge 2).
- D-123 (ADOPTED Model A): DC nodes vcloud-level; site-down = destroy the vr1-dc0-* group; MAAS
region on Office1 + a rack controller per DC (vvr1-dc0).
- D-124 (ADOPTED Scheme A): the Office1-region<->DC-rack management overlay -- a transit link on
the office1<->dc0 mesh leg; rack sizing 4/8192/80.
Code (all validated, NOT applied):
- OpenTofu: new modules/site-wan (NAT /24 uplink, provider-schema-confirmed); main.tf Stage-3
section -- vr1_dc0_wan, the 2-NIC edge, 8 Option C node VMs (for_each), the vvr1-dc0 rack
(two legs, static IPs from new vr1_dc0_rack_* tfvars -- NetBox-assigned, not invented), netem
held, mesh-link network_name output; stale vr1_dc1_planes comment fixed. Step-9 maas-vm-host
deliberately deferred (DOCFIX-179: its maas provider would force every plan to demand creds).
tofu validate Success; opentofu-validate 11/11.
- NetBox: dc-edge-wan-import.py (58/58) + dc-rack-mgmt-import.py (96/96) -- register the DC-edge
/24s + the rack transit/IP, dry-by-default, whole-plan preflight, operator-supplied literals.
- overlays/dc-ha-scaleup.yaml: the D-121 14-service HA scale-up (Stage-5), placement tokens +
rabbitmq min-cluster-size + hacluster cluster_count + re-declared vault-hacluster.
- scripts/site-headend-install.sh: rack-only mode (enrolls a DC rack to Office1's region).
- runbooks/dc-dc-phase2-tofu-dc-substrate.md: fix-forward (D-119 naming, Steps 4-5 REST-API,
D-121/122 folded in, stale gate corrected).
docs/session-ledger.md + dc-dc-deployment-workflow.md reconciled; three changelogs added.
Gauntlet ALL GREEN; repo-lint 0-fail (1 documented legacy WARN).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ck6xh3jWQi5b3Su8Dx1LEH
|
| 2026-07-15 |
stage-close discipline: add the operating-skill sweep as part of definition-of-done
...
Per operator: the skill-update sweep must itself be part of every stage close, so it
is never skipped. At stage close (alongside the branch->main merge + branch retire),
sweep .claude/skills/openstack-cloud-ops/ for any new INVARIANT the stage established
and reconcile divergence from repo HEAD. Recorded in both the workflow doc's
Cross-cutting discipline and the skill's own VR1 deploy-loop stage-close note.
repo-lint 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WgVrZE2ryY9aXP7JiBuwj
|

Stage 2 CLOSED+MERGED; open Stage 3 branch; standing merge-at-close discipline; drop stray temp/ files
...
Post-merge housekeeping after PR #1 (e468d2e) put VR1 Stage 1+2 on main.
- Workflow doc: the Stage 2 "branch merge point" note now records the merge as DONE
(PR #1, e468d2e) and GENERALIZES it into a standing rule in the Cross-cutting
discipline section -- "branch-per-stage + MERGE AT STAGE CLOSE": every stage
develops on its own branch and merging it to main (merge commit, operator-gated)
is part of that stage's definition-of-done, then the branch is retired. main is
trunk-of-record only for merged stages; the next stage branches off post-merge main.
- session-ledger: Stage 2 CLOSED+MERGED recorded; the PINNED merge ruling graduated
into the standing discipline; NEXT = Stage 3 on dc-dc-stage3-phase2-dc-substrate,
with D-119 selectors + creds-audit (SEC-009) in place.
- Removed temp/ (3 stray screenshot PNGs that rode onto main via 95f0634).
repo-lint 0 fail. No script/harness touched (gauntlet unaffected).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WgVrZE2ryY9aXP7JiBuwj
|

C2 sub-task (ii): extend sandbox loop to ip-ranges/ip-addresses; load + verify D-120 bands
...
Closes the logged tooling gap and the D-120 Step 6 load, verified end to end. With
this, C2 is fully closed and Stage 2 close-out (C1..C5) is COMPLETE.
Tooling:
- prod-draft-dump.py: ENDPOINTS += ipam/ip-ranges, ipam/ip-addresses (ranges get a
synthetic start-end key). Verified live.
- sandbox-fidelity-check.py: KEY + EXPECTED_NEW cover the 3 D-120 ranges + 2 service
IPs; both-bounds gives endpoint/address correctness (pure delta -- the frozen draft
has none). Success banner now D-115/D-117/D-118/D-120. Harness +T11/T12 (12->14).
- netbox/d120-compose-bands.py: stdlib, UA-aware, SANDBOX_HOSTS-guarded, dry-by-default,
whole-plan preflight (parent 10.10.1.0/24 exists, endpoints in-parent, start<=end) so
it cannot half-write. Harness 16/16 pins the band values + guards.
Load (operator-approved --commit into office1-netbox, the VR1 apex; sandbox-local):
- created 3 ip-ranges (.2-.49 / .100-.200 / .201-.254) + 2 ip-addresses (.10 netbox,
.11 tailscale), 0 already present. Re-dump -> 3 ranges, 2 IPs. Fidelity check -> exit 0,
"EXACTLY the planned delta (D-115/D-117/D-118/D-120)."
D-120 Step 6 CLOSED (no apex write -- write-back stays deferred). Gauntlet ALL GREEN
(60); repo-lint 0 fail. Workflow C2 row + session-ledger updated: Stage 2 close-out
complete -> branch->main merge point reached (merge remains operator-gated).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WgVrZE2ryY9aXP7JiBuwj
|

DOCFIX-195: office1-netbox holds the VR1 IPAM apex role; apex write-back deferred; C2 redefined
...
Operator ruling 2026-07-15: office1-netbox (10.10.1.10) HOLDS THE IPAM APEX
ROLE during VR1 -- the authoritative NetBox every consuming system (OpenTofu,
MAAS, overlays, bundle) derives its values from. It is a working draft (edited
as the buildout proceeds), but it is the apex. netbox.baldurkeep.com takes NO
write during VR1; it is the eventual production apex, frozen as a read-only v1
reference draft. Merge-back is deferred to end-of-deployment (a maybe), NOT a
Stage 2 gate.
The repo encoded the opposite roles and defined close-out C2 (the last Stage 2
gate, and the branch->main merge trigger) as an operator-gated write UPSTREAM.
Under the ruling there is no apex write in Stage 2, so C2 is redefined:
office1-netbox holds the COMPLETE validated dataset -- carve + six-plane roles
+ both DCs + the D-120 compose-/24 child ranges + the two service IPs -- and it
is VERIFIED. C2 stays OPEN.
Finding surfaced while drafting (logged, not fixed): the D-120 objects are
ip-ranges/ip-addresses, object classes the sandbox loop does NOT cover
(prod-draft-dump.py ENDPOINTS and sandbox-fidelity-check.py KEY/EXPECTED_NEW
both stop at rirs/roles/regions/sites/aggregates/prefixes). So the hardened
fidelity check cannot verify the D-120 load; verifying it needs a tool
extension (+harness) or a recorded manual read-back. Confirmed no live system
config points at any NetBox URL -- documentation-only change.
Files: dc-dc-deployment-workflow.md (State + C2 row), dc-dc-netbox-buildout-scope.md
(sec 8.1 + 8.6), session-ledger.md (5 narrative spots + new deferred-item bullet;
machine-derived block re-seeded from the 2026-07-15 scan), design-decisions.md
(D-010 note + D-103 pointer), netbox/README.md (target banner), + changelog.
repo-lint 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WgVrZE2ryY9aXP7JiBuwj
|

D-120 ADOPTED + EXECUTED: re-IP Office1 services into the .2-.49 static band
...
office1-netbox .201->.10 and office1-tailscale .202->.11, on the wire, no
wipe. NetBox HTTP 200 at .10; tailscale still advertises 10.10.0.0/22.
Method (DC replication inherits it): MAAS refuses interface edits on a
Deployed machine and the only lift is Release->redeploy=WIPE, so the live
address is set on the guest (netplan + curtin cfg, reboot-durable) over
lxc exec (socket immune to the netplan apply), and MAAS-model drift is
accepted (reconciles at next teardown/redeploy). NetBox ALLOWED_HOSTS=*
so no restart; tailscale re-asserted automatically.
Swept SANDBOX_HOSTS .201->.10 in the three importers + test + the on-box
working copy; updated live-state docs; flipped D-120 to ADOPTED with an
Execution record; resolved the re-IP runbook's PENDING markers to the
measured method. Gauntlet ALL GREEN (58); repo-lint 0 fail / 1 legacy WARN.
Still owed: D-120 Step 6 (register the /24 child ranges + service IPs in
the sandbox, feed upstream) rides with C2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QpGaiX6UvhSReDwN2ywstJ
|
| 2026-07-14 |

C4 DONE: document the NetBox sandbox loop; add the upstream-write guard it exposed
...
C4 (Stage 2 close-out) -- the sandbox loop is now documented in two places, no
rationale duplicated: docs/dc-dc-netbox-buildout-scope.md section 8 (design: why a
sandbox, the fidelity check + its 2026-07-14 hardening, the write-guard table, the
WAF trap, C2 preconditions) and runbooks/dc-dc-phase1-office1-standup.md Step 10b
(the dump->seed->apply->verify->gated-write sequence, two-host/two-token separation).
Drafted by a subagent, corrected against this session's own fixes before placement.
The guard C4 exposed: roles-aggregates-import.py and dc-dc-prefixes-import.py had NO
hostname gate -- until today the WAF 403 was the only thing stopping an accidental
--commit to the apex, and the UA fix removed that accidental safety. Both now gate a
non-sandbox --commit behind --yes-write-upstream (SANDBOX_HOSTS allowlist), matching
d115-office-carve.py -- so all three write-capable importers behave identically. This
is defence-in-depth with the whole-plan preflights (which prevent a half-write).
Tests: prefixes-import 86->91 (refused without the flag + writes nothing; allowed
with it; sandbox host needs no flag). roles-aggregates 20->24. GAUNTLET ALL GREEN (58).
Stage 2: C1/C3/C4 DONE. C2 (production write) is the only item left -- gated, with
preconditions now documented (re-verify sandbox under the hardened check, run from
office1-netbox, --yes-write-upstream, approve each mutation).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|

D-119: region-qualify the VR1 DC namespace -- C3 / gap #19 CLOSED (Stage 3 unblocked)
...
THE BUG: the token `dc0` meant TWO DIFFERENT CLOUDS depending on the file --
VR0's LIVE testcloud in scripts/lib-net.sh, VR1's FIRST DC in the NetBox
importer. One string, two clouds, one of them in production.
THE RULING (operator): the repo adopts the apex's names verbatim.
vr0-dc0 VR0's DC0 the LIVE testcloud (a DIFFERENT region)
vr1-dc0 VR1's FIRST DC GUA 2602:f3e2:f02::/48
vr1-dc1 VR1's SECOND DC GUA 2602:f3e2:f03::/48
Bare dc0/dc1/dc2 are now REJECTED LOUDLY everywhere.
D-119 COMPLETES D-117, it does not reverse it: it executes D-117's own amendment
across the three surfaces D-117 never touched. ZERO NetBox writes -- the apex was
already correct and self-consistent (MEASURED: f02::/48 -> vr1-dc0, f03::/48 ->
vr1-dc1); the REPO was the only surface out of step. Renaming the apex to match
the repo was considered and REJECTED (production IPAM write; VR1 would become the
only 1-indexed region; and vr1-dc1 would mean VR1's FIRST DC while vr0-dc1 means
VR0's SECOND).
THE PRIZE: the importer's DC->site map is now an IDENTITY -- there is no offset
table left to get wrong, and an assert enforces it. The original defect was
precisely a WRONG LOOKUP TABLE. The bug class is deleted, not defended.
TWO REAL BUGS the naming fix ALONE would NOT have closed (found by review sweep):
- BREAK-1: descriptions were built with f"VR1 {dc.upper()} ..." -- under D-119 that
renders "VR1 VR1-DC0 provider-public" on all 36 prefixes. Same class as the
original bug (deriving a label by munging a token), hidden in the description
field where slug-focused review missed it. Now looked up from SITES[name].
- BREAK-2 (the important one): DC_GUA_PREFIX was NEVER cross-checked against --dc.
`--dc vr1-dc0 DC_GUA_PREFIX=<f03>` was ACCEPTED: it writes the SECOND DC's GUA,
carved with the FIRST DC's ULA nibble, scoped to the FIRST DC's site -- a
silently mis-bound datacenter assembled from two disagreeing sources. Identity-
mapping the slug leaves the ADDRESSING free-floating. Now guarded by EXPECTED_GUA.
ALSO: vdc1/vdc2 -> vvr1-dc0/vvr1-dc1 (D-114 amendment). D-114's own DR primitive is
`virsh destroy vdc1`, which read as "destroy VR1 DC1" while MEANING "destroy VR1
DC0". A mislabelled destroy command in a DR drill is not cosmetic. Neither VM is
built, so it is free. voffice1/Office1 KEEP their number (operator ruling).
Env var: DC1_/DC2_V4_SUPERNET -> VR1_DC1_V4_SUPERNET (both old names rejected).
rbd-mirror's --site-name aligned: `--dc vr1-dc0 --site-name dc1` was a re-created
two-namespace collision on one command line.
TESTS: dc-selector 21->30, prefixes-import 40->82 (identity invariant; EXPECTED_GUA
in BOTH mismatch directions; mismatched run writes NOTHING; matching pair still
succeeds; no munged description). GAUNTLET ALL GREEN (57). repo-lint 0 fail.
The old harnesses PINNED THE WRONG MAPPING -- they would have gone green while
enforcing the bug.
THE tofu apply IS NOT DONE -- IT IS GATED. Plan: 11 to add, 0 to change, 11 to
destroy (libvirt object names are ForceNew; moved{} makes the plan reviewable but
cannot suppress a replace). MEASURED SAFE: all 11 are EMPTY -- no guests, no
volumes, Stage 3 hasn't run. Office1's live objects are not in the plan.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|

Office1 LAN IPv6 is LIVE and proven on the wire -- Stage 2 close-out C1 CLOSED
...
The last build item on Office1. Executed against the live edge, verified against
the kernel and the wire, not by {"result":"saved"}.
scripts/opnsense-set-interface-v6.sh --commit 10.10.0.1 lan 2602:f3e2:f01 :1 64
POST /api/radvd/settings/addEntry {"entries":{...,"mode":"unmanaged"}}
POST /api/radvd/service/reconfigure
PROVEN: address on the kernel (ifconfig, not the config file); radvd advertising
2602:f3e2:f01 :/64 AdvAutonomous/AdvOnLink with no M/O flags; voffice1
autoconfigured ...:5054:ff:fe6a:87e5 by SLAAC + an RA-learned route; edge->voffice1
GUA ping 0.0% loss. The neighbour MAC matches voffice1's known Kea lease hwaddr.
No collateral damage: v4 10.10.0.1 intact, Kea DHCPv4 still serving udp/67, the
compose-net route still good.
mode=unmanaged is DELIBERATELY not the model's default. `stateless` sets the RA
Other-config flag, pointing clients at a DHCPv6 server -- and Kea DHCPv6 is
enabled=0 with no daemon running. The default would have advertised a server that
does not exist. Field names, the request key (addBase('entries','entries')), the
mode values and the apply endpoint were all READ OFF THE EDGE, not recalled.
OPERATOR RULING -- the v6 NAT was NOT built. A v4-to-v6 NAT (a second OPNsense, or
nginx) was proposed to simulate the v6 the lab will not forward. Withdrawn: it
contradicts D-101 ("native GUA routing, no NAT66"), D-115's amendment says the
missing egress is the upstream's problem and not ours, and a translation box is
precisely the component Roosevelt will NOT have -- maximum delta for zero proof.
nginx is an L7 proxy, not NAT, and proves nothing about RA/ICMPv6/Ceph-over-v6.
Ruling: configure IPv6 internally now; configure the IPv6 edge AFTER deployment.
TRAP recorded: voffice1 has forwarding=1 + accept_ra=0 and STILL autoconfigured --
netplan/systemd-networkd processes RA in USERSPACE, independent of the sysctl.
WAN v6 leg stays DEFERRED: re-measured today, v6 internet is 100% loss from vcloud
while v4 is fine. Internet GUA reachability remains UNPROVABLE in VR1.
repo-lint 0 fail; opnsense-set-interface-v6 harness 40/40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|

Pin Stage 2 as the branch merge point; reconcile the Stage 2 tracker; log the OpenTofu naming collision (gap #19)
...
Operator ruling 2026-07-14: branch dc-dc-stage1-phase0-first-exec merges to
main at STAGE 2 CLOSE, not at session end. Recorded in the ledger's PINNED
section (a ruling, not an open question) and in the workflow doc's Stage 2
section, which now carries the C1..C5 close-out checklist that defines what
"closed" means. This is NEW policy: main has never taken a feature merge, so
all 52 commits of VR1 work live only on the branch.
The Stage 2 status block was STALE -- it still listed MAAS-region, LXD, the
LXD-KVM-host registration and both composed service machines as NOT DONE. All
of that landed 2026-07-13, and the route enable + the D-115 carve import have
landed since. Corrected against measured state; the doc has now drifted three
times and says so. Stage 2 is in CLOSE-OUT, not build.
GAP #19 (new, BLOCKING Stage 3): "dc1" now names two different datacenters.
D-117 moved the repo to the apex's 0-indexed convention on the IMPORT PATH
ONLY. opentofu/main.tf still says dc1_planes/dc2_storage for VR1's FIRST and
SECOND DCs, and lib-net.sh's dc0 means VR0's DC0 -- a different region, and a
RUNNING cloud. So the next DC we deploy is simultaneously apex vr1-dc0,
OpenTofu dc1_*, and a $DC value that collides with the live testcloud. A
mechanical sed would hand VR1's first DC VR0's deployed literals. Nothing in
opentofu/ gains a new DC-keyed module until this is closed.
Docs only -- no script, no tooling, no cloud state touched.
repo-lint: 0 fail (1 documented legacy warn).
Revert: git revert HEAD
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZBxzAbmqwLW7jEmMG23Xw
|
| 2026-07-13 |

D-116 ADOPTED: no Office1-local GitBucket -- keep using git.baldurkeep.com
...
Operator ruling: "we've decided not to build a separate one on the test cloud. We will continue to
use the current setup and create new repos for deployment tasks as needed."
SUPERSEDES the GitBucket half of D-103 (which gave Office1 three service VMs) and removes GitBucket
from Stage 2's build list entirely. Office1's MAAS-composed service machines are now exactly TWO --
office1-netbox and office1-tailscale -- and both are LIVE.
WHAT THIS ACTUALLY CLOSES, and why it is a simplification rather than a deferral: the Stage 2 runbook
carried an OPEN QUESTION about the Office1-local GitBucket -- "what it mirrors, and under what repo
path, is not decided" -- because a SECOND GitBucket alongside the pre-existing git.baldurkeep.com
raised a real authority/mirroring question nobody had answered. That question is now MOOT. There is
one git service, it is the existing one, and it is authoritative.
Recorded explicitly in D-116 so nobody misreads the scope: D-107's per-DC ARTIFACT mirror (OS/package
artifacts for node provisioning) is a SEPARATE concern and is NOT affected. And if Roosevelt later
wants a site-local git service, that is a fresh decision against real requirements -- not something
this ruling forecloses.
Applied everywhere it was an instruction, not just recorded:
- runbooks/dc-dc-phase1-office1-standup.md: Step 11 replaced with a TOMBSTONE (kept rather than
deleted -- step numbering is referenced elsewhere, and someone WILL come looking after reading
D-103); removed from the sequence, the compose steps, and the open-questions list.
- docs/dc-dc-deployment-workflow.md: dropped from Stage 2's Build and, importantly, from its GATE --
"GitBucket serving" is no longer a gate condition.
- docs/design-decisions.md: D-103 annotated in place with the supersession, so grepping for GitBucket
lands on the amendment rather than the stale instruction.
- docs/vr1-office1-as-built.md + session ledger updated.
repo-lint 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|

D-114 follow-through: fix node-vm's nested-virt defect NOW; voffice1 pinned to a reserved address
...
Operator ruling: fixes land as we find them -- only DC1 DEPLOYMENT is gated, not DC1 fixes.
1. modules/node-vm: UNCONDITIONAL `cpu = { mode = "host-passthrough" }`. Same defect just found in
cloudinit-vm: with no cpu block libvirt hands the guest a generic emulated CPU with NO svm. A DC
node runs nova-compute, so it MUST run KVM guests -- there is no correct value other than this,
hence no knob that could only ever be set wrong (contrast cloudinit-vm, where it IS a real per-VM
decision, and opnsense-edge, which disables svm as correct hardening for a router). Not yet
instantiated, so no live call breaks. Without it the DC nodes would have come up looking healthy
and been unable to launch a single instance -- surfacing in Stage 5 as an inscrutable scheduler error.
2. voffice1 moved off the dynamic pool onto a RESERVED address. It first took 10.10.0.100 -- the
first address of Kea's pool (.100-.199). A MAAS region controller must not hold an address that
can move. Added a Kea host reservation via the D-113(a2) REST API (a second live exercise of that
write path): 52:54:00:6a:87:e5 -> 10.10.0.20, deliberately OUTSIDE the pool. add_reservation ->
"saved", service/reconfigure -> "ok". Rebooted; voffice1 came back on 10.10.0.20, .100 released,
/dev/kvm still present. Done now because nothing depends on the address yet.
3. runbooks/dc-dc-phase1-office1-standup.md rewritten to the D-114 model (was: three sibling service
VMs + an unruled Option A/B fork -- it would have actively misdirected whoever ran it). New
sequence: voffice1 via tofu -> first contact (lease/SSH/nested-KVM) -> MAAS on voffice1 -> LXD
(5.21 LTS) -> register LXD as a MAAS VM host -> MAAS-COMPOSE NetBox/GitBucket/Tailscale. The
MAAS/LXD command lines carry explicit PENDING VERIFICATION markers rather than fabricated flags.
docs/dc-dc-deployment-workflow.md Stage 2 + gap register updated to match.
repo-lint 0 fail. opentofu-validate PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|

Correct the workflow doc's status: it was materially FALSE and would misdirect work
...
docs/dc-dc-deployment-workflow.md is the first thing a fresh session reads. Its status claims
dated from 2026-07-09 and had drifted into being wrong. Corrected against MEASURED state.
Docs only; no code, no live system touched.
gap #2 "UNVALIDATED -- no tofu binary" -> OpenTofu v1.12.3 IS installed; tree validates
(root + 10/10 modules) and is partly APPLIED
(11 resources in state). But 2 modules were
BROKEN and had never been parsed (DOCFIX-194).
gap #17 "no per-site ISP-uplink network -> CLOSED FOR OFFICE1: office1-wan is live
exists anywhere, for ANY site" (virbr11, NAT, 172.30.1.0/24); the edge's WAN
sits on it at 172.30.1.2, egress verified. The
gap's own proposed answer (a dedicated per-site
uplink, NOT a mesh leg) is what got built.
DC1/DC2 STILL have none.
opnsense-edge "BUILT, UNVALIDATED" -> BUILT, INSTANTIATED, RUNNING (routing, NAT,
egress, console, SSH, Kea DHCP on udp/67).
config.xml tmpl "BUILT, fully tested" -> SUPERSEDED by D-113(a2); a full push would now
CLOBBER live API-managed DHCP.
netem-link "BUILT, UNVALIDATED" -> WAS BROKEN; fixed 2026-07-13 (DOCFIX-194).
Stage 2 "NOT YET EXECUTED" -> PARTIALLY EXECUTED.
THE TRAP THIS REMOVES: Stage 2's gate lists "NetBox authoritative" and "GitBucket serving" -- and
NetBox and GitBucket DO answer, at baldurkeep.com. But those are PRE-EXISTING services, NOT
Office1 headend VMs. Measured: office1-opnsense is the ONLY VM on the vcloud host. It would be
very easy to tick that gate as met and walk into Stage 3 with no headend at all. The corrected
State section now says this explicitly.
What actually remains of Stage 2 (the real critical path): the three headend service VMs
(MAAS-region, NetBox, GitBucket) do not exist. Blocked on a genuine decision, not on typing:
Option A (modules/cloudinit-vm) needs user_data/meta_data/network_config designed and an image
source chosen; Option B is manual virt-install, logged as debt. Both cloudinit-vm and base-image
now at least VALIDATE (DOCFIX-194) -- they never had before.
Revert: git revert this commit (restores the stale, false status claims).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B6Fzre1CxCY8tzwFCudu57
|