Session changelog (GA-R2/D1: ONE per session). Logged window: dc0-deploy (same-day append, per the A/B-boundary handoff). Status lives ONLY in docs/CURRENT-STATE.md; this file is narrative + reverts.
module "voffice1" in the outer root; apply in a planned voffice1 restart window (graceful shutdown first; region/NetBox/Tailscale return via proven autostart).docs/design-decisions.md:3727); no new D-number (GA-R3, doubt resolves DOWN).<forward>/<ip>) with exactly ONE attached interface (vvr1-dc0 vnet3); voffice1 + office1-opnsense have NO mesh NIC; nothing holds 172.31.0.1; vvr1-dc0's applied transit leg = 172.31.0.2/30 with only route 10.10.0.0/22 via 172.31.0.1 (d124-rack.auto.tfvars, consumed by step A). Item-20 disposition: with voffice1 holding .1, vvr1-dc0 is reached by -J voffice1 (D-126 shape) -- NO vcloud host leg / site-baseleg DC row needed for this net.opentofu/main.tf module "voffice1": network_names gains module.mesh_vr1_dc0_office1.network_name SECOND (office1-local stays first -- PCI order preserves enp1s0); network_config rewritten from the single-NIC "en*" glob to per-NIC name matches (lan=enp1s0 dhcp4; transit=enp2s0 static ${vr1_dc0_rack_transit_peer_ip}/${prefix} -- same D-124 tfvars the rack consumes, no new literal). Naming-trap + no-cloud-init- re-run caveats documented in-block: the seed netplan is REBUILD correctness; the live transit config is an in-guest netplan drop-in at the gated attach step.network_names = [module.office1_network.network_name] and the previous "en*" glob network_config (git revert of this commit); if already applied, a follow-up plan shows the interface removal and the in-guest drop-in /etc/netplan/60-transit.yaml is deleted.docs/audit/outer-apply-20260719-voffice1-transit.txt) -> domain up with both NICs, office1-local index 0.52:54:00:6a:87:e5 -> new 52:54:00:89:e1:19). The Kea reservation (MAC-keyed) missed and voffice1 leased pool address .108 (measured via edge ARP). FIX: re-keyed reservation uuid b12621e6-63e9-4714-9726-2e130c84e069 to the new MAC via the D-113 API (kea/dhcpv4/set_reservation -> saved; kea/service/reconfigure -> ok), then a detached networkctl reconfigure enp1s0 -> voffice1 back on 10.10.0.20.modules/cloudinit-vm does not pin MACs, so ANY interface-list change re-rolls them and breaks MAC-keyed DHCP reservations. Roosevelt-relevant; candidate module mac var or a documented ops step. Also: live transit config landed as in-guest netplan drop-in /etc/netplan/60-transit.yaml (narrows the lan glob to enp1s0 + transit enp2s0 static) -- the drop-in must ALSO narrow lan, else the 50-cloud-init en* glob claims enp2s0 (measured: first drop-in attempt without it left enp2s0 addressless).kea/dhcpv4/set_reservation back to the old MAC (senseless unless the NIC change is also reverted); delete the drop-in.172.31.0.2 unreachable; MEASURED: both rack NICs emitted ZERO frames ever (host FDB virbr7/virbr4, voffice1 ARP FAILED); guest OS alive (ACPI-responsive; cloud-init completed per offline logs).modules/cloudinit-vm domains have NO serial console, NO qga channel, NO graphics; qga is NOT in the noble base image and its cloud-init install failed (no egress) -- so no console, no agent, no network. LOGGED FINDING: add opt-in serial console (+ the D-129-style qga channel var) to modules/cloudinit-vm -- observability gap bites every future DC VM. NOT executed mid-step.~/vvr1-dc0-offline-inspect{,2}.txt): NIC names enp1s0/enp2s0 CORRECT (naming trap closed); root cause = netplan match.name + set-name: boot 1 renames don't apply (.link written after udev add) so the rendered config matched nothing; after removing set-name, boot 2 STILL failed because the early-boot netplan generator ran against the OLD config and renamed the devices to mgmt/uplink before cloud-init wrote the new file (one-boot lag, measured in syslog). Boot 3 (current config, no renames) is the verification boot.network_config drops set-name, keeps MEASURED kernel names; instance-id bumped vvr1-dc0-d123 -> -d123-r2 (side effects noted in-block); seed replaced via explicit -replace=module.vvr1_dc0.libvirt_volume.seed (the sanctioned override of D-130 ignore_changes for INTENDED seed changes) -- plan/apply 2/0/2 exact (docs/audit/outer-{plan,apply}-20260719-vvr1dc0-netplan-fix.txt).--transit-if enp1s0 --uplink-if enp2s0 (script defaults mgmt/uplink no longer match reality); runbook prose naming "mgmt" needs a DOCFIX at stage close.-replace apply + reboot.enp1s0 UP 172.31.0.2/30, region route present, br-vr1-dc0-wan up.-i ~/vr1-dc0-creds/vr1-dc0_svc_ed25519 -J voffice1 jessea123@172.31.0.2 works (D-126 -J shape, per-env key); rack -> region ping 10.10.0.20 0% loss (the MAAS-enrollment path for the G10 bootstrap).docs/audit/outer-plan-20260720-postfix-converged.txt); CURRENT-STATE section 5 updated in the same commit.install -m600 /dev/stdin), verified by byte-count only (32), never in context.docs/audit/stepB-bootstrap-20260720.txt): maas snap 3.7/stable installed, rack ENROLLED to http://10.10.0.20:5240/MAAS; FAILED at qemu apt install on stale image package lists (404s via proxy). Fix: apt-get update (proxied), re-run.docs/audit/stepB-bootstrap-20260720-run2.txt): EXIT 0 -- libvirt/qemu + nested KVM on, inner pool + AppArmor, SEC-010 FORWARD-drop loaded on enp1s0 + sec010-fw.service enabled (boot-persistent), br-vr1-dc0-wan verified with enp2s0 enslaved.--check (docs/audit/stepB-check-20260720.txt): all items ok EXCEPT opnsense base (the remaining step-B item); rackd running, enrolled.maas ... rack-controllers delete; SEC-010: disable sec010-fw.service + rm /etc/nftables-sec010.nft; packages: apt/snap remove; snap proxy: snap unset system proxy.http proxy.https.opnsense-26.7-nano.qcow2; vr1-dc0-substrate opnsense_base_path default updated (no inner state exists; pre-apply edit); step-B --check passes --opnsense-base for the 26.7 path. LOGGED follow-items: site-headend-install.sh's 26.1 default + hint strings (script change + harness, stage close); the D-112/D-113 serial/API boot path is REVALIDATED on 26.7 at the step-C boot gate (office1 precedent: 26.1-proven, upgraded in place to 26.7).https://mirror.ams1.nl.leaseweb.net/opnsense -- MEASURED from the vendor's live download page this session (not memory), HEAD-verified serving the 26.7 nano (~468 MiB). LOGGED FINDING: opnsense-prep-image.sh performs no checksum verification of the downloaded image (same property as the original 26.1 prep; candidate hardening).docs/audit/stepB-opnsense-prep-20260720.txt). Inner-root opnsense_base_path default -> 26.7 (validate PASS).--check EXIT 0, ALL items ok (docs/audit/stepB-check-20260720-final.txt).systemctl disable sec010-fw on voffice1; nft delete table inet sec010; ledger row reopens citing this item.29cf7bf (branch); dc0 private key staged to ~/vr1-dc0-creds/ (0700/0600, piped -- SEC-009 convention now spans hosts); ssh config entry for 172.31.0.2; known_hosts seeded via ssh-keyscan with the ED25519 fingerprint cross-checked against the accept-new-trusted entry (MATCH).keyfile+sshauth=privkey URI params -- its Go ssh ignores ~/.ssh/config AND default identities; single-type known_hosts fails as "key mismatch" (multi-type scan required); (ii) modules/wan-bridge: <mtu> is ILLEGAL in bridge-mode networks -- removed (MTU belongs to the host bridge; DOCFIX-194 class: parsed, never applied); (iii) base-image location premise WRONG under remote provider -- content UPLOADS from the executing host; 26.7 qcow2 copied rack->voffice1 (~/vr1-dc0-images/), tfvar override + variables.tf description corrected; (iv) edge serial-log dir /var/lib/libvirt/vr1/staging/ is a HARDCODED vcloud literal in modules/opnsense-edge (:223) -- dir created on the rack (755); LOGGED findings: parameterize the path; fold staging-dir creation into site-headend-install --host-nodes; opentofu-validate.sh does not cover the inner root.docs/audit/inner-{plan,apply}-20260720-stepC*.txt): run 1 = 25/28 (nodes+planes+pool; wan-bridge mtu abort); run 2 = wan bridge (edge disk stat abort); run 3 = edge disk (staging-dir abort; provider self-cleaned the failed domain -- operator undefine found nothing); run 4 = edge domain, EXIT 0. CONVERGED: inner plan zero diff (docs/audit/inner-converge-20260720-stepC.txt).~/openstack-caracal-dc-dc/opentofu/vr1-dc0-substrate/terraform.tfstate) -- back it up with the site; follow-item for the backup set.tofu destroy in the inner root from voffice1 (site-down alternative: virsh destroy vvr1-dc0, D-122); module fixes revert by commit.~/vr1-dc0-creds/vr1-dc0-edge_ed25519, 0600/0644, SEC-009 convention), staged to the rack the same way.10.12.4.2/22 on the provider-public inner bridge (virbr5, MEASURED) + the D-124-ruled 10.12.8.2/22 on metal-admin (virbr2). INTERIM: ip addr add, not persistent. LOGGED FINDING (queued): site-headend-install.sh never implemented the rack legs despite main.tf's comment claiming it consumes vr1_dc0_rack_metal_admin_ip -- add rack-leg support + harness before step D closes. Also interim: 192.168.1.2/24 on virbr5 to reach the edge's FACTORY LAN; REMOVE once the edge is re-addressed.~/d112c-console.log on the rack). Driver evolved v1->v6 against MEASURED failures, each real and worth keeping for the runbook: (i) serial console wraps long lines -> ship payload as <=160-char base64 chunks (openssl base64 -d), never one long line; (ii) set-name in the edge netplan is irrelevant here but the same class bit the rack (item 4); (iii) tcsh on the edge (documented trap) breaks nested-quoted sed -- "Unmatched '"'"'" -- so NEVER patch files via quoted one-liners; re-ship the whole payload instead; (iv) write_config() needs an include beyond config.inc -- MEASURED (grep -l shell_safe /usr/local/etc/inc/*.inc -> auth/console/filter/ interfaces/rrd/system/util.inc), not guessed; (v) writing config.xml does NOT materialize /root/.ssh/authorized_keys -- must call the vendor's local_user_set() (auth.inc). This was the actual reason key auth failed after a "successful" config write.ssh -i vr1-dc0-edge_ed25519 root@192.168.1.1 -> key-only auth OK, uname -r = 15.1-RELEASE-p1, ifconfig -l = vtnet0/vtnet1. INTERFACE MAPPING MEASURED: vtnet0 = LAN = provider-public plane (it answered on virbr5), vtnet1 = WAN -- same order as the office1 edge.opnsense-bootstrap-apikey.sh
opnsense-mint-apikey.php, run ON the rack): key/secret 80 chars -> ~/vr1-dc0-creds/opnsense-api.txt (0600, secret never printed). SMOKE TEST PASS: GET core/firmware/status -> HTTP 200, CORE_ABI 26.7. This is the first proof the D-113(a2) API path works on 26.7 (the 26.7 revalidation the operator's base-version ruling called for). Script note: the bootstrap script requires the minter BESIDE it (its own dir), not in scripts/ -- worth a usage-line DOCFIX.POST auth/user/del_api_key/<id>, the proven negative-tested path) + creds file; ip addr del the three interim addresses; the console bootstrap reverts by re-imaging the edge (inner tofu destroy/re-apply of the edge module).br-vr1-dc0-wan, serial-log probe, then destroyed (leftover-domains=0 asserted). Script: scratchpad d125-egress-test.sh; capture docs/audit/d125-egress-gate-20260720.txt.vr1-dc0-uplink carries an <ip> but NO <dhcp> block -- there is no DHCP server on that /24 by design (the edge WAN takes a STATIC .2, D-113/D-125). DOCFIX QUEUED: the gate's wording "must GET a vcloud-ISP address" implies DHCP; it must say STATIC, or the /24 needs a range. Re-run used static 172.30.2.50/24.ping 1.1.1.1 FAILED (NET-PING-RC=1).curl http://1.1.1.1 -> 000) both fail past the gateway, so it is not ICMP-specific. Controls: vcloud egress 0% loss; the office1 edge egresses through the SAME-SHAPE office1-wan NAT 0% loss. The two network definitions are byte-equivalent in substance (net-dumpxml: both forward mode=nat + <nat><port 1024-65535>; only netmask-vs- prefix notation differs) -- so the delta is in vcloud's LIVE firewall rules for virbr4, NOT in the design or the config.sudo nft list ruleset | grep -nE "172\.30\.(1|2)\.0/24" (or sudo iptables -t nat -S | grep 172.30) comparing the WORKING virbr11 rules against virbr4's./tmp/noble.img on the rack (kept deliberately: it is the base for any future throwaway probe).awk toupper(cell) ~ /OPEN|PENDING/) matched the substring OPEN inside "fail-open" in the closed row's disposition -- the row counted OPEN forever (same class as the GA-F15 wrap trap). Fix: strip FAIL-OPEN from the cell before matching (scripts/ledger-scan.sh); harness gains a CLOSED-row-with-fail-open fixture + nochk (45/45 PASS). Count re-verified: 7 open, SEC-010 absent.sudo apt-get install -y libguestfs-tools (1:1.52.0-5ubuntu3) -- enables offline guest inspection (virt-cat/virt-ls), used for items 4's captures. Standard KVM-host tooling; Roosevelt-transferable.sudo apt-get remove libguestfs-tools.