diff --git a/docs/CURRENT-STATE.md b/docs/CURRENT-STATE.md index ec25b04..46d50e1 100644 --- a/docs/CURRENT-STATE.md +++ b/docs/CURRENT-STATE.md @@ -2729,6 +2729,50 @@ measured plan is **56 apps / 108 relations**; `docs/dc-dc-buildout-design.md:12` records the repo was seeded at `8813efc` specifically so `bundle.yaml` + `phase-01..08` could be reused. **Provider capabilities must be READ (`references/opentofu-provider-docs.md`), never recalled.** + **>>> dc1 EDGE ASSESSED (read-only agent, 2026-08-02): IT FORWARDS BUT DOES NOT TRANSLATE, + ITS CONFIG IS INTACT, AND THE VERDICT IS REPAIR NOT REBUILD. <<<** Capture + `docs/audit/dc1-edge-assessment-20260802.txt` (643 lines). **DECISIVE MEASUREMENT:** + simultaneous tcpdump on BOTH edge taps while the rack pinged `1.1.1.1` -- the same packet + appears on the LAN tap and 0.8 ms later on the WAN tap with **source still `10.12.64.2`**. It + lands on vcloud's `virbr1`, whose libvirt NAT masquerades only `172.30.3.0/24`, so it is never + translated and no reply can return. **Routing works and pf is not blocking; there is simply no + outbound NAT.** My hypothesis (config lost) was half right and the half that mattered was + wrong. + **CONFIG IS INTACT, ON POSITIVE EVIDENCE rather than "it booted":** fsck NAMES every inode it + deletes, all 50 were enumerated, and **`/conf/config.xml` is not among them** -- the whole + `/conf` tree lost one transient event file. The edge came up on its ruled as-built addresses + (LAN `10.12.64.1/22`, WAN `172.30.3.2/24`), not the factory `192.168.1.1/24` the same log + shows pre-bootstrap. QUALIFIER STATED BY THE AGENT: 92 further casualties are reported by + inode only, so this is strong, not absolute; one `xmllint` settles it. + **WHAT fsck ACTUALLY DESTROYED: the FreeBSD BASE-SYSTEM USER DATABASE** -- `/etc/master.passwd` + (`pwd_mkdb: line #1 too long`) and `/etc/group` (`pw: group 'wheel' does not exist`, 17x + `chown: wheel: illegal group name`). Neither is part of config.xml. The before/after sits in + ONE file: the same serial log reads `Configuring firewall.......done.` on the factory boot and + `Configuring firewall.....failed.` plus `Configuring OpenSSH...failed.` on every boot since. + **SO NO pf RULESET IS IN FORCE -- SEC-031 OPENED.** Confirmed against a CONTROL, which is what + makes it a finding: from vcloud the healthy office1 edge (same image, same build path) answers + WAN ICMP 100% loss / GUI `000`; dc1's answers WAN ICMP **0% loss** / GUI **200 in 0.014s**. + Bounded by lab topology -- the "ISP" is a libvirt NAT net on vcloud -- and D-125 is unaffected. + **REPAIR, NOT REBUILD:** dc0 lost `libcrypto`/`libpython` and cannot reach multi-user; dc1's + PHP runs, serves pages, and its config is correct. Rebuilding would discard a good config and + re-incur D-112(c) + D-113 + re-addressing. **THE CAVEAT THAT SHAPES IT: dc1 has NO working + credential path** -- sshd down (tcp/22 dead), authenticated API GETs hang while unauth answers + in 20 ms, console login likely fails on the same unparseable passwd DB -- and **fsck ALSO + deleted `/var/backups/master.passwd.bak` and `/var/backups/group.bak`**, so the repair must + REGENERATE from config.xml (the proven D-112(c) `local_user_set()` path) or donate from the + pristine `opnsense-26.7-nano.qcow2` in the same pool. That needs `boot -s`, a mutation, so the + read-only agent stopped there. Neither edge disk has a backingStore: **no rollback point + exists on either edge today.** + **>>> TWO FINDINGS THAT OUTLIVE THIS INCIDENT. <<<** **(1) `dc-egress-check.sh`'s A2 is + STRUCTURALLY BLIND TO THIS FAILURE MODE** -- it asserts the edge ANSWERS, and dc1's edge does; + the gate therefore passes A2 and fails A3/A4, which cannot distinguish "no NAT" from "upstream + down". **A positive "the edge TRANSLATES" assertion is required.** That is a defect in a gate + built earlier the same day, found by an agent rather than by me. **(2) NEITHER EDGE HAD EVER + BEEN REBOOT-TESTED.** dc1's serial log holds exactly TWO boots, and 2026-08-01 was the FIRST + time its as-built config ever booted. Every edge in this deployment was built, configured + live, and never power-cycled until the resize did it involuntarily -- which is why a + first-reboot defect surfaced at the worst moment. Roosevelt analog: an edge that has never + survived a reboot is not a built edge. **>>> RUNBOOK FOLD OPENED 2026-08-02 (operator: "1 fold first", "3 start at 2", "4 yes, all DC come up with and in their own region"). Register: `docs/runbook-fold-register.md`. <<<** **MEASURED GAP:** `D-138` and `D-139` appear in **NO runbook at all**; nothing in the chain diff --git a/docs/audit/dc1-edge-assessment-20260802.txt b/docs/audit/dc1-edge-assessment-20260802.txt new file mode 100644 index 0000000..ea2ed8b --- /dev/null +++ b/docs/audit/dc1-edge-assessment-20260802.txt @@ -0,0 +1,643 @@ +dc1-edge-assessment-20260802.txt +================================================================================ +READ-ONLY assessment of the vr1-dc1 OPNsense edge after the 2026-08-01 05:45:30 +containment-VM power cut. NOTHING was mutated: no virsh start/destroy/define, no +config write, no service restart, no tofu. Every API call issued was a GET. +Companion to docs/audit/dc0-edge-egress-outage-20260802.txt (the dc0 half). + +Every line below is tagged MEASURED (observed this session, or quoted from a +captured artifact) or INFERRED (reasoning over measurements). Per this repo's +discipline a finding is an observation, not a conclusion, until it is measured. + +================================================================================ +VERDICT +================================================================================ + Q1 Does the dc1 edge forward/NAT? + It FORWARDS. It does NOT NAT. And it is enforcing NO EFFECTIVE FIREWALL + RULESET -- measured: no outbound NAT on the wire, and no WAN block, in a + controlled comparison against a healthy sibling edge. (Whether pf is + disabled or loaded-but-permissive is NOT separable read-only; `pfctl -s + info` in section 8 settles it, and nothing in the recommendation turns on + which it is.) Not a D-125 isolation posture -- D-125 is the decision that + RESTORES per-DC egress, and dc1 rack egress was MEASURED WORKING on + 2026-07-30. This is a regression, not a design. + + Q2 Is the config INTACT? + YES, on positive evidence rather than on "it booted." fsck names every + inode it deletes; all 50 named casualties are listed in section 5a and + /conf/config.xml IS NOT AMONG THEM -- the whole /conf tree lost exactly + one transient event file. The edge came up on its exact ruled as-built + addresses (LAN 10.12.64.1/22, WAN 172.30.3.2/24) with routes, Unbound, + Dnsmasq, NTP and the gateway monitor all "done." What fsck destroyed is + the FreeBSD BASE SYSTEM user database -- /etc/master.passwd and + /etc/group -- which is NOT part of config.xml, plus BOTH of their + /var/backups/*.bak copies. Everything broken is downstream of that. + QUALIFIER: strong but not absolute -- fsck's other 92 casualties are + reported by inode, not by name (see 5a(i)). One `xmllint --noout + /conf/config.xml` settles it and needs a shell (section 8). + + Q3 Rebuild, repair, or nothing? + >>> REPAIR, not rebuild. <<< dc0 lost its shared libraries and cannot + reach multi-user; dc1 lost its passwd/group database and reaches + multi-user with its whole configuration loaded. Different severity AND a + different repair. See "RECOMMENDATION" for the caveat that decides how + the repair has to be delivered. + + Q4 Other damage in dc1? + NONE FOUND. The edge was the ONLY inner guest running at the cut (the + other ten dc1 domains are shut off with autostart=disable). The rack host + vvr1-dc1 (ext4) is clean: no ext4/IO errors, zero failed units, all four + dc1 units active. + +================================================================================ +1. THE BASELINE THAT FRAMES EVERYTHING -- this was working two days ago +================================================================================ +MEASURED (quoted from docs/CURRENT-STATE.md:115-121, a captured 2026-07-30 +measurement, not prose): + + "ARTIFACT SOURCES AND EGRESS RE-MEASURED 2026-07-30, both DCs. ... dc1 proxy + PASS (apt-cacher-ng active on 10.12.68.4:3142, serves archive AND UCA 200). + Egress re-probed from BOTH rack hosts with `--noproxy '*'` so a cache hit + could not fake it: juju agent stream `streams.canonical.com/juju/tools/` 200, + snap store `api.snapcraft.io` answering, `archive.ubuntu.com` jammy Release + 200, 1.1.1.1 0% loss, default routes via 10.12.4.1 / 10.12.64.1." + +This closes three of the four alternatives in the task's hypothesis list before +any live probe is taken: + - "a deliberate D-125 isolation posture I am misreading" -- REFUTED. D-125 is + the ruling that GIVES each DC a working ISP uplink (bridge-in, single NAT); + its whole purpose is egress. And egress was measured working. + - "no NAT rule was ever configured" -- REFUTED. Outbound NAT was translating + on 2026-07-30. + - "interfaces mis-assigned" -- REFUTED, and re-refuted live in section 3. +The only event recorded between that measurement and now is the 2026-08-01 +05:45:30 hard cut. INFERRED: this is a regression caused by that cut. Sections +2-4 measure the mechanism. + +================================================================================ +2. THE SERIAL LOG -- a controlled before/after inside ONE file +================================================================================ +Serial log path taken from the live domain XML, not assumed: + MEASURED `virsh dumpxml vr1-dc1-opnsense` -> + + + +INSTRUMENT CURRENCY (stated before any negative is drawn from this file): + MEASURED size=79424 mtime=2026-08-01 07:31:36.982889689 +0000 + MEASURED the file contains exactly TWO boots, and `append='on'` means it was + never rotated: + boot 1 ends line 259, banner dated "Wed Jul 22 21:51:39 UTC 2026", + LAN (vtnet0) -> v4: 192.168.1.1/24 (FACTORY addressing; + the D-112(c) console bootstrap session follows it inline) + boot 2 ends line 2670, banner dated "Sat Aug 1 07:31:36 UTC 2026", + LAN 10.12.64.1/22, WAN 172.30.3.2/24 (AS-BUILT addressing) + MEASURED fsck damage counts in the file: UNREF FILE 76, UNEXPECTED SOFT + UPDATE INCONSISTENCY 181, SALVAGE 46, ld-elf errors 0. + (Matches the dc0 report's dc1 control column exactly.) + + NOTE, and it is load-bearing: there is NO boot between 2026-07-22 and + 2026-08-01. The as-built config was applied LIVE over the D-113 API on + 2026-07-23 and the edge then ran continuously. The 2026-08-01 boot is + therefore the FIRST TIME this edge has ever booted on its as-built config. + See section 7, finding (c). + +--- 2a. The pre-cut boot (lines 225-245, VERBATIM) ------------------------- + Setting hostname: OPNsense.internal + Generating /etc/resolv.conf...done. + Generating /etc/hosts...done. + Configuring system logging...done. + Configuring firewall.......done. + Configuring hardware interfaces...done. + Configuring loopback interface...done. + Configuring LAN interface...done. + Configuring WAN interface...done. + Generating /etc/resolv.conf...done. + Generating /etc/hosts...done. + Configuring firewall.......done. + Starting web GUI...done. + Setting up routes...done. + Starting Dnsmasq...done. + Starting Unbound DNS...done. + Configuring firewall.......done. + +--- 2b. fsck, then the first casualty (lines 2548-2565, VERBATIM) ---------- + SUMMARY INFORMATION BAD + SALVAGE? yes + + BLK(S) MISSING IN BIT MAPS + SALVAGE? yes + + 77557 files, 534225 used, 2165014 free (86 frags, 270616 blocks, 0.0% frag.) + + ***** FILE SYSTEM MARKED CLEAN ***** + + ***** FILE SYSTEM WAS MODIFIED ***** + pwd_mkdb: line #1 too long + pwd_mkdb: /etc/master.passwd: Inappropriate file type or format + pwd_mkdb: line #1 too long + pwd_mkdb: /etc/master.passwd: Inappropriate file type or format + + This is the dc1 equivalent of dc0's `ld-elf.so.1: libcrypto.so.17 not found`: + the FIRST line after "FILE SYSTEM WAS MODIFIED" names the casualty. On dc0 it + was the shared libraries. On dc1 it is /etc/master.passwd. + +--- 2c. The post-cut boot (lines 2615-2650, VERBATIM) ---------------------- + Setting hostname: OPNsense.internal + Generating /etc/resolv.conf...done. + Generating /etc/hosts...done. + Configuring system logging...done. + Configuring firewall.....failed. <<<< + Configuring hardware interfaces...done. + Configuring loopback interface...done. + Configuring LAN interface...done. + Configuring WAN interface...done. + Generating /etc/resolv.conf...done. + Generating /etc/hosts...done. + Configuring firewall.....failed. <<<< + Configuring OpenSSH...failed. <<<< + Starting web GUI...pwd_mkdb: line #1 too long + pwd_mkdb: /etc/master.passwd: Inappropriate file type or format + pw: passwd file update: No error: 0 + chown: wheel: illegal group name <<<< x9 + ... + done. + Setting up routes...done. + Starting Dnsmasq...done. + Starting Unbound DNS...done. + Configuring firewall.....failed. <<<< + Setting up gateway monitor...done. + Syncing OpenVPN settings...done. + Starting NTP service...done. + Starting Unbound DNS...done. + + and further down, same boot: + Configuring login behaviour...pw: group `wheel' does not exist + Starting Cron: error in configd communication, see syslog for details + error in configd communication, see syslog for details (openvpn) + chown: hostd: illegal user name / illegal group name + /usr/local/etc/rc.d/hostwatch: WARNING: failed to setup hostwatch + +MEASURED, exhaustive grep over the whole file: + "Configuring firewall" appears at lines 229, 236, 241 (all "......done.") + and 2619, 2626, 2644 (all "....failed."). + "Configuring OpenSSH" appears ONCE, at line 2627, and it FAILED. + +HONEST QUALIFIER on 2a: the three "done." lines are from the FACTORY-config +boot, so they prove the firewall subsystem works on THIS IMAGE, not that the +as-built config booted cleanly (it had never booted). The evidence that the +as-built config's NAT worked is the 2026-07-30 measurement in section 1, which +is stronger anyway because it is behavioral. + +MEASURED qemu started the edge at 2026-08-01 05:45 (mtime of + /var/log/libvirt/qemu/vr1-dc1-opnsense.log); the boot banner is + stamped 07:31:36. ~1h46m elapsed. +INFERRED that gap is the fsck pass over a damaged UFS root (the ~2,000 lines + of salvage output between the two). Not independently timed. + +================================================================================ +3. LIVE STATE -- what the edge is doing right now +================================================================================ +All probes run 2026-08-02 ~06:00-06:20 UTC. Rack identity confirmed first: + MEASURED `ssh vr1-dc1-rack 'hostname'` -> vvr1-dc1 ; uptime "up 1 day, 10 min" + (consistent with the 2026-08-01 05:45 boot) + MEASURED ip -4: enp1s0 172.31.0.6/30, virbr4 10.12.64.2/22, + virbr6 10.12.68.2/.3/.4/22, virbr0 192.168.122.1/24 + MEASURED ip route: "default via 10.12.64.1 dev virbr4" + MEASURED ip neigh dev virbr4: 10.12.64.1 lladdr 52:54:00:95:65:64 STALE + -- and 52:54:00:95:65:64 is exactly the net0 MAC in the edge's + domain XML. The gateway answering IS the edge VM, not a squatter. + +Domain XML (MEASURED, `virsh dumpxml vr1-dc1-opnsense`): + net0 mac 52:54:00:95:65:64 network vr1-dc1-provider-public bridge virbr4 + target vnet0 (LAN leg) + net1 mac 52:54:00:87:3b:74 network vr1-dc1-wan bridge br-vr1-dc1-wan + target vnet1 (WAN leg, D-125 bridge-in) + disk pool vr1-dc1-inner-pool vol vr1-dc1-opnsense-disk.qcow2, + EMPTY -- standalone qcow2, no copy-on-write base to roll back to. + Identical to dc0. (qemu-img could not be run on it: the file is locked + by the running guest. The empty is the evidence.) + Interfaces are CORRECTLY assigned. "mis-assigned interfaces" is refuted. + +Reachability (MEASURED): + from the rack (vvr1-dc1): + ping 10.12.64.1 (edge LAN) -> 2/2, 0% loss + ping 172.30.3.2 (edge WAN) -> 2/2, 0% loss + ping 1.1.1.1 -> 0/2, 100% loss + from vcloud: + ping 172.30.3.2 -> 2/2, 0% loss + ip neigh -> 172.30.3.2 dev virbr1 lladdr 52:54:00:87:3b:74 REACHABLE + (matches the net1 MAC: the D-125 bridge-in L2 path is intact end to end, + inner tap -> br-vr1-dc1-wan -> vcloud virbr1) + virsh net-dumpxml vr1-dc1-uplink -> forward mode='nat', ip 172.30.3.1/24, + bridge virbr1 -- the simulated ISP is up and is a NAT for 172.30.3.0/24. + +Service ports on the edge (MEASURED): + LAN 10.12.64.1 WAN 172.30.3.2 (from vcloud) + tcp/22 (sshd) closed / no answer closed / no answer + tcp/80 301 301 + tcp/443 200 200 + GET / title "Login | OPNsense" + sshd is DOWN -- consistent with "Configuring OpenSSH...failed." The web GUI is + UP and serving PHP. + +API (D-113 key, MEASURED; the key was never printed, never copied to another +host -- opnsense-api.sh passes the secret to curl via --config on stdin): + unauthenticated GET /api/core/menu/search -> 302 in 0.022s + unauthenticated GET /api/core/firmware/status -> 302 in 0.017s + AUTHENTICATED GET core/menu/search -> NO RESPONSE, killed at 40s + AUTHENTICATED GET diagnostics/interface/getInterfaceNames -> killed at 30s + AUTHENTICATED GET diagnostics/interface/getRoutes -> killed at 30s + The web stack answers unauthenticated requests in ~20ms and hangs forever the + moment authentication is attempted -- including on core/menu/search, which does + not use configd. INFERRED: the hang is in the authentication path, which is the + one thing that must consult a user database that no longer parses. Not proven + further; the API is simply not a usable instrument on this box. + +================================================================================ +4. THE DECISIVE TEST -- forwards yes, translates no +================================================================================ +tcpdump on BOTH of the edge's tap devices simultaneously, on the rack (the taps +belong to the rack's libvirt), while the rack pinged 1.1.1.1. Read-only. + + MEASURED vnet0 (edge LAN tap): + 06:04:16.702929 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 1 + 06:04:17.757057 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 2 + 06:04:18.781116 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 3 + 06:04:19.805049 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 4 + + MEASURED vnet1 (edge WAN tap), SAME four packets, ~0.8 ms later: + 06:04:16.703747 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 1 + 06:04:17.757724 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 2 + 06:04:18.781871 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 3 + 06:04:19.805778 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 4 + +This settles Q1 completely: + * The edge HAS a working default route out its WAN. Packets are routed LAN->WAN. + * pf is NOT blocking them -- they are passed, not dropped. + * The source address is STILL 10.12.64.2 on the WAN side. NO OUTBOUND NAT IS + APPLIED. The packet lands on vcloud's virbr1, whose libvirt NAT masquerades + only 172.30.3.0/24, so a 10.12.64.2 source is never translated and no reply + can ever come back. That is exactly the observed "0% loss to the gateway, + 100% loss to the internet." + + So the failure is precisely: pf has no outbound-NAT ruleset loaded. + It is NOT a routing failure, NOT an interface failure, NOT an L2 failure, + NOT an upstream failure. + +--- 4a. CONTROL: pf is enforcing NO ruleset at all ------------------------- +An OPNsense edge with any ruleset loaded blocks unsolicited inbound on WAN -- +there is no default WAN pass rule, and the anti-lockout rule is LAN-only. So +"can the ISP side reach the GUI?" is a direct test for a loaded ruleset. + + MEASURED, from vcloud, against the HEALTHY office1 edge (same OPNsense 26.7 + image, same D-112/D-113 build path -- the control): + 172.30.1.2 found via `ip neigh show dev virbr11` (not guessed) + ping 172.30.1.2 -> 0/2, 100% loss (ICMP blocked on WAN) + https://172.30.1.2/ -> 000 after 12.006s (no answer) + http://172.30.1.2/ -> 000 after 12.006s (no answer) + tcp/22 -> closed + + MEASURED, from vcloud, against the SUBJECT dc1 edge: + ping 172.30.3.2 -> 2/2, 0% loss (ICMP answered on WAN) + https://172.30.3.2/ -> 200 in 0.014s (GUI served to the ISP) + http://172.30.3.2/ -> 301 + tcp/22 -> closed (sshd is down, not firewalled) + + The dc1 edge is answering ICMP and serving its management GUI to the simulated + ISP. A healthy sibling on the identical image does neither. Conclusion, stated + to exactly the strength of the evidence: the dc1 edge has NO EFFECTIVE pf + RULESET -- MEASURED, there is no outbound NAT (section 4) and no WAN block + (this control). + WHAT THIS DOES NOT DISTINGUISH: pf disabled / ruleset empty, versus pf loaded + with a partial or permit-all ruleset. Both produce identical external + behaviour and both produce no NAT, and neither can be separated without + `pfctl -s info` on the guest (section 8). The recommendation is the same + either way, so this is recorded as an open sub-question, not chased. + +>>> SECURITY CONSEQUENCE, flagged not actioned (read-only session): MEASURED, +>>> there is NO INBOUND FILTERING IN FORCE ON THE dc1 EDGE'S WAN -- its +>>> management GUI answers 200 on its WAN address from the simulated ISP +>>> segment, and it answers ICMP there, neither of which the healthy control +>>> does. It is functioning as an open router between 172.30.3.0/24 and the dc1 +>>> provider-public plane. The exposure is bounded by the lab's own topology +>>> (172.30.3.0/24 is a vcloud-local libvirt NAT, not routable inbound from the +>>> real internet), which is why this is a finding and not an incident. It +>>> belongs in docs/security-ledger.md as a new SEC row. I did not open one: +>>> numbering and ledger rows are the operator's to assign. + +================================================================================ +5. IS THE CONFIG INTACT? (Q2) +================================================================================ +EVIDENCE THAT config.xml SURVIVED -- and none of it is "it booted": + + MEASURED The post-cut boot configured both interfaces onto the EXACT ruled + as-built values, which exist ONLY in config.xml: + LAN (vtnet0) -> v4: 10.12.64.1/22 + WAN (vtnet1) -> v4: 172.30.3.2/24 + These match changelog-20260723-g12-dc1-edge.md Item 5 and + lib-net.sh's provider-public gateway exactly. A damaged or reverted + config would have come up on the factory 192.168.1.1/24 (which is + what boot 1 in the very same log shows it doing). + MEASURED "Configuring LAN interface...done." / "Configuring WAN + interface...done." / "Setting up routes...done." -- the config + parsed and drove the interface + routing layer. + MEASURED Dnsmasq, Unbound (twice), NTP, the gateway monitor and the web GUI + all started "done." on the post-cut boot. + MEASURED Section 4 proves a default route exists and works -- i.e. the + data is present and usable. + MEASURED The web GUI serves its login page (200, correct ), so the + PHP stack that READS config.xml is functional. dc0's is not. + +WHAT WAS ACTUALLY DESTROYED -- and it is not config: + MEASURED /etc/master.passwd is unparseable: pwd_mkdb reports "line #1 too + long" and "Inappropriate file type or format" on every one of four + attempts across the boot. So /etc/pwd.db and /etc/spwd.db were never + rebuilt. + MEASURED /etc/group is damaged: `pw: group 'wheel' does not exist`, and + 17 x `chown: wheel: illegal group name` plus 4 x `chown: hostd: + illegal user name` / `illegal group name` (exact grep counts). + Both are FreeBSD BASE SYSTEM files. Neither is part of config.xml. They are + fsck casualties in the same class as dc0's libcrypto.so.17 -- the file was + reconnected as garbage or truncated, not removed. + +--- 5a. THE NAMED-CASUALTY INVENTORY -- the strongest evidence in this file --- +fsck names every inode it DELETES. MEASURED, exhaustive over the serial log: + REMOVE? yes 50 UNALLOCATED 50 NAME= 50 + RECONNECT? yes 92 +All 50 named casualties, deduped and complete: + + /conf/event_config_changed.json <-- the ONLY /conf casualty + /root/.history + /usr/local/etc/bogons + /usr/local/etc/bogonsv6 + /usr/local/etc/unbound.opnsense.d/access_lists.conf + /usr/local/etc/unbound.opnsense.d/domainoverrides.conf + /usr/local/etc/unbound.opnsense.d/dot.conf + /usr/local/etc/unbound.opnsense.d/safesearch.conf + /usr/local/etc/unbound/unbound-blocklists.conf + /usr/local/etc/unbound_dhcpd.conf + /usr/local/lib/perl5/5.42/perl/man/mandoc.db + /usr/local/lib/perl5/site_perl/man/mandoc.db + /usr/local/lib/python3.13/site-packages/__pycache__/unboundmodule.cpython-313.pyc + /usr/local/share/man/mandoc.db + /usr/share/openssl/man/mandoc.db + /var/backups/gpart.vtbd0.bak + /var/backups/group.bak <<<< THE BACKUP OF A CORRUPT FILE + /var/backups/kern.geom.conftxt.bak + /var/backups/master.passwd.bak <<<< THE BACKUP OF A CORRUPT FILE + /var/backups/pkg.sql.xz.1 .2 .3 .4 .7 + /var/cron/tabs/root + /var/db/dhcp6c_duid + /var/db/entropy/saved-entropy.8 + /var/db/locate.database + /var/db/pkg/vuln.xml + /var/run/cron.pid dhcp6c.pid dpinger_WAN_GW.pid dpinger_WAN_GW.sock + /var/run/gateway_watcher.pid rtsold.pid sshd.pid unbound.pid utx.active + /var/unbound/advanced.conf cache.dump.gz host_entries.conf + /var/unbound/private_domains.conf resolv.conf.root root.hints root.key + /var/unbound/unbound.conf unbound_control.key unbound_control.pem + /var/unbound/unbound_server.key unbound_server.pem + +TWO CONCLUSIONS, both directly measured: + + (i) /conf/config.xml IS NOT IN THE LIST. The entire /conf tree lost exactly + one file, and it is `event_config_changed.json` -- a transient event + marker, not configuration. Combined with the fact that the config + demonstrably drove the boot onto its exact as-built addresses, this is + positive evidence of survival rather than an inference from "it booted." + QUALIFIER, because absence of evidence is not evidence of absence: the + 92 RECONNECT (UNREF FILE) entries are reported by INODE, not by name, so + fsck could in principle have unlinked a file into lost+found without + printing its path. Absence from this list is strong but not absolute. + The check that settles it is `xmllint --noout /conf/config.xml` -- see + section 8. + + (ii) BOTH FreeBSD-NATIVE BACKUPS OF THE CORRUPT FILES WERE ALSO DESTROYED. + /var/backups/master.passwd.bak and /var/backups/group.bak are on the + REMOVE list. So the obvious repair -- copy the .bak back over the corrupt + original -- IS NOT AVAILABLE. This corrects an earlier draft of this + assessment which read the master.passwd.bak line as evidence the backup + still existed; it appears in the log as + `UNALLOCATED I=77473 OWNER=0 MODE=0 SIZE=0 ... REMOVE? yes`, + i.e. as a deletion, not a survival. Repair must regenerate the user + database from config.xml instead (section 7). + + A third, minor observation: nearly every /var/unbound file was destroyed, + including unbound's TLS keypairs -- yet "Starting Unbound DNS...done." + appears twice on the post-cut boot. INFERRED: OPNsense regenerates the whole + /var/unbound tree from config.xml at boot, which is another datapoint that + config.xml is intact and that the config->runtime generators work. The one + generator that did NOT work is the firewall's. + +CAUSAL CHAIN + MEASURED the damage and the failures are in the same boot, in this order: + fsck salvage -> master.passwd/group unparseable -> firewall + configure failed x3 -> OpenSSH configure failed -> configd + communication errors -> API auth hangs -> no pf ruleset -> no NAT. + INFERRED that the passwd/group corruption is the CAUSE of the firewall + configure failure. The adjacency is measured; the specific internal + dependency (OPNsense's filter_configure resolving users/groups, or + configd failing to start for the same reason) is NOT proven, because + proving it needs a shell on the edge, which I do not have read-only. + ALTERNATIVE NOT FULLY EXCLUDED, stated honestly: because 2026-08-01 was the + FIRST boot this edge ever took on its as-built config, a latent + boot-time config defect cannot be excluded by this evidence alone. + It is unlikely -- the fsck damage to base files is measured and + sufficient, and outbound NAT was working live on 2026-07-30 -- but + it is not zero, and the repair path in section 7 distinguishes them + for free (if repairing the user DB restores NAT, the alternative is + dead). + +WHAT I COULD NOT READ, and why (this is a gap, not a pass): + I have NOT read config.xml itself, nor /conf/backup/, nor the live pf ruleset + (`pfctl -sr` / `pfctl -sn`). All three need a shell on the edge, and all three + read-only routes to one are closed: + - SSH: tcp/22 does not answer (MEASURED). sshd failed to configure. + - REST API: authenticated GETs hang indefinitely (MEASURED). + - Serial console: the guest is sitting at `login:`. Typing at it is an + interaction with the guest, which I read as outside "READ-ONLY, make no + mutations," so I did not attempt it. INFERRED, and worth the operator's + attention before anyone tries: a console login will most likely FAIL + anyway, because authentication needs the same /etc/master.passwd that + pwd_mkdb cannot parse. + If the operator wants config.xml and the pf ruleset read directly, the + commands are listed in section 8. I am stopping rather than guessing. + +================================================================================ +6. ANY OTHER dc1 DAMAGE FROM THE SAME CUT? (Q4) +================================================================================ +INNER GUESTS -- MEASURED, `virsh list --all` + per-domain `dominfo`: + vr1-dc1-opnsense running Autostart=enable + vr1-dc1-control-01 shut off Autostart=disable + vr1-dc1-control-02 shut off Autostart=disable + vr1-dc1-control-03 shut off Autostart=disable + vr1-dc1-compute-01 shut off Autostart=disable + vr1-dc1-compute-02 shut off Autostart=disable + vr1-dc1-storage-01 shut off Autostart=disable + vr1-dc1-storage-02 shut off Autostart=disable + vr1-dc1-storage-03 shut off Autostart=disable + vr1-dc1-storage-04 shut off Autostart=disable + vr1-dc1-juju-01 shut off Autostart=disable + Eleven domains, one running. The edge is the ONLY guest that was hard-cut. + This is D-127 behaving exactly as written and it BOUNDS the blast radius: + there is no other inner filesystem to inspect. (Consistent with the task's + note that dc1 has no per-DC MAAS region VM -- MEASURED: no such domain exists.) + +THE RACK HOST vvr1-dc1 (ext4) -- MEASURED: + df/findmnt /dev/vda1 ext4 2.9T, 1% used, rw,relatime,discard, + errors=remount-ro -- NOT remounted read-only + journal Aug 01 05:45:41 EXT4-fs (vda1): mounted ... ro, ordered data + Aug 01 05:45:41 EXT4-fs (vda1): re-mounted ... r/w + systemd-fsck-root skipped (unmet condition) -- normal + BOOT: recovering journal; BOOT: clean, 604/58496 files + UEFI (vfat): "differences between boot sector and its backup + ... mostly harmless" + journal -p err for the whole boot: ONE kernel line ("RDSEED32 is broken. + Disabling the corresponding CPUID bit." -- a CPU-feature + notice, unrelated) plus a repeating libvirtd line, below. + systemctl --failed: "0 loaded units listed." + dc1 units: apt-cacher-ng.service active running + dc1-cache-proxy-net.service active exited + dc1-node-dns.service active running + dc1-rack-legs.service active exited + libvirtd.service active running + ext4 came through the cut cleanly. A journalling filesystem with an intact + journal is the expected outcome; UFS-with-soft-updates is not, which is the + whole reason the two edges are damaged and neither rack is. + + INSTRUMENT-CURRENCY NOTE, because a negative is being reported here: + `dmesg` is NOT a valid instrument for the boot window on this host. It holds + 4834 lines whose OLDEST timestamp is [75228.4] (~20.9h uptime) against an + uptime of ~24.5h -- the ring buffer HAS WRAPPED and the boot messages are + gone from it. Its "zero ext4 errors" covers only the last ~3.6 hours. The + boot-window evidence above is from `journalctl -b`, which does cover + 2026-08-01 05:45:41 onward. The dmesg ring is dominated by a flood of + `apparmor="DENIED" ... profile="snap.maas.pebble" ... comm="ps"` lines -- + that is the MAAS RACK controller snap on this host reading /proc/N/environ; + pre-existing noise, not damage, and not related to this incident. + + PRE-EXISTING NOISE, explicitly NOT attributed to the cut -- MEASURED: + `libvirtd[2531]: End of file while reading data: Input/output error` + recurs in ~10-minute clusters from Aug 01 05:51:57 onward; 1406 occurrences + before this session began. INFERRED: a periodic libvirt client + disconnecting (power-state polling shape). It predates and is unaffected by + my probes. Logged so nobody later mistakes it for new damage. + +================================================================================ +7. RECOMMENDATION +================================================================================ +>>> REPAIR the dc1 edge. Do NOT rebuild it. <<< + +WHY REPAIR IS THE RIGHT CALL (all measured): + a) The configuration survived. The edge is running its exact as-built + identity -- LAN 10.12.64.1/22, WAN 172.30.3.2/24, working default route, + working forwarding, DNS/NTP/GUI up. A rebuild would throw all of that away + and re-incur the D-112(c) console bootstrap, the D-113 API-key mint, the + re-addressing, and a fresh SEC row for the transient key exposure -- to + recreate a config that is already correct. + b) The break is a bounded, named, base-system defect: two corrupt files + (/etc/master.passwd, /etc/group) and the services that depend on them + (pf ruleset generation, sshd, configd, API auth). + c) dc0 is genuinely different and genuinely needs the rebuild: it lost + libcrypto.so.17 and libpython3.13.so.1.0, so PHP and python cannot even + load and it never reaches multi-user. dc1's PHP runs and serves pages. + Same incident, same damage class, materially different remedy. The task's + hypothesis -- "same incident, milder outcome" -- is CONFIRMED, and the + "milder" is load-bearing: it changes the remedy. + +THE CAVEAT THAT DECIDES HOW, and it is the important part of this section: + MEASURED, dc1 currently has NO working credential path into the guest: + - sshd is down (tcp/22 does not answer) + - authenticated API calls hang + - INFERRED: console login will fail too, since it needs the same unparseable + /etc/master.passwd + dc0, paradoxically, is EASIER to get a shell on -- it dropped to a single-user + prompt, which is a root shell. dc1 is sitting at a `login:` that probably + cannot authenticate anyone. + So the repair almost certainly requires BOOTING THE dc1 EDGE INTO SINGLE-USER + MODE (interrupt the loader, `boot -s`) to get a shell, then rebuilding the + user database. + AND THE OBVIOUS RESTORE PATH IS CLOSED -- MEASURED, section 5a(ii): fsck + DELETED both /var/backups/master.passwd.bak and /var/backups/group.bak, so + there is no native backup to copy back. The repair therefore has to + REGENERATE the user/group database rather than restore it. The two candidate + sources, in preference order: + 1. OPNsense's own generators, driven from the surviving config.xml -- the + same `local_user_set()` path the D-112(c) console bootstrap already used + on this exact edge (see changelog-20260723-g12-dc1-edge.md Item 2, which + records the working PHP shape including the `require_once("util.inc")` + fix). This is the repo-proven route and it reuses tooling that exists. + 2. The pristine base image in the same pool -- MEASURED, `virsh vol-list + vr1-dc1-inner-pool` holds `opnsense-26.7-nano.qcow2` -- as a donor for + the stock /etc/master.passwd and /etc/group, which for an appliance are + essentially image constants. Then re-run (1) to re-apply config users. + THAT IS A REBOOT. It is a mutation. I did not do it and I am not proposing to + do it inside this assessment -- it needs its own gated step, and it should be + preceded by a disk snapshot/copy of vr1-dc1-opnsense-disk.qcow2 (the domain + XML shows an EMPTY <backingStore/>: there is no rollback point today, on + either edge). Copying that disk requires the guest to be shut down or the + copy to be taken with the guest quiesced -- also a mutation, also gated. + +FALLBACK IF REPAIR FAILS: the rebuild material is present and identical to +dc0's -- MEASURED, `virsh vol-list vr1-dc1-inner-pool` contains the pristine +`opnsense-26.7-nano.qcow2` alongside `vr1-dc1-opnsense-disk.qcow2`. So a rebuild +remains available at no additional risk; it is simply the more expensive answer +and the evidence does not call for it. + +DURABLE FINDINGS THIS RAISES (logged, not actioned -- hard rule 1): + (a) The dc0 report's procedure gap is confirmed by a second, independent + datapoint: an "in-place" tofu change to a containment VM is a power cut to + every inner guest, and BOTH edges took filesystem damage from it. The + remedy the dc0 report proposes -- shut inner guests down first -- would + have prevented both. + (b) dc-egress-check.sh's A2 assertion cannot see this failure mode. On dc1 the + edge ANSWERS at 10.12.64.1 (A2 OK) and forwards, and the DC still has no + egress. A2 is doing its job for the dc0 shape and is simply not the + instrument for the dc1 shape. A candidate A2b: "the edge translates" -- + and it should be written in the POSITIVE form: assert that traffic + observed leaving the edge's WAN side carries the WAN SOURCE ADDRESS (the + section 4 tcpdump), which is a real positive assertion and can fail + loudly. The tempting cheap version -- "assert the edge does NOT serve its + GUI on its WAN address", i.e. the section 4a control -- is an INVERTED + assertion that reads PASS whenever a curl times out for any reason, which + is exactly the "a checker that cannot fail is not a gate" / "could not + look is never nothing there" shape this repo has already been bitten by + twice. Use the control as corroboration, never as the gate itself. + Worth raising as a gate improvement; I have NOT touched the script. + (c) NEITHER DC EDGE HAD EVER BEEN REBOOT-TESTED. MEASURED for dc1: the serial + log contains exactly two boots, 2026-07-22 (factory) and 2026-08-01 + (as-built, first ever). The as-built config was applied live over the API + on 2026-07-23 and never proven to survive a restart. Both edges' first + real reboot was an unclean one, which is the worst possible way to find + out. Roosevelt delta: an edge's config is not "built" until it has been + proven across a clean reboot. That reads like D-NNN material rather than a + changelog line, but assigning the number is the operator's call. + (d) SECURITY: section 4a's open-router finding needs a SEC row. + +================================================================================ +8. WHAT I COULD NOT ANSWER READ-ONLY, and the exact commands that would +================================================================================ +Stated rather than guessed, per the ground rules. + + Q: Is config.xml byte-intact, and are /conf/backup/ revisions present? + Q: What exactly is in the live pf ruleset -- empty, or absent? + Q: Is it really master.passwd/group that breaks filter_configure? + + All three need a shell on the edge. Every read-only route to one is closed + (section 5). The minimal commands, once a shell exists, are: + ls -la /conf/config.xml /conf/backup/ | tail -20 + xmllint --noout /conf/config.xml && echo XML-WELL-FORMED + grep -c '<rule>' /conf/config.xml ; grep -o '<outbound>' /conf/config.xml + pfctl -s info ; pfctl -sr | head -40 ; pfctl -sn | head -40 + wc -c /etc/master.passwd /etc/group ; head -3 /etc/group + ls -la /var/backups/ # EXPECT master.passwd.bak and group.bak to + # be ABSENT -- fsck removed both (5a(ii)) + ls -la /lost+found | head # the 92 RECONNECTed inodes land here + Getting that shell requires `boot -s` at the loader, i.e. a REBOOT of the + edge guest. That is a mutation and it is where I stop. + +================================================================================ +END. Nothing in this session was mutated. The repo tree is left dirty and +uncommitted by instruction. +================================================================================ diff --git a/docs/security-ledger.md b/docs/security-ledger.md index a820f7a..d604311 100644 --- a/docs/security-ledger.md +++ b/docs/security-ledger.md @@ -81,3 +81,4 @@ | SEC-028 | 2026-07-30 | **Per-DC JUJU SERVICE credential, and the RACK becomes a credential-bearing host for the first time.** D-132 q1 gives each DC its own MAAS region with its own database, so the Office1 `juju-vr1-dc0` user and its API key do NOT authenticate against dc0's region -- a fact that reads as a network fault at bootstrap rather than as an auth failure. dc0's `juju-vr1-dc0` was therefore MINTED IN THE NEW REGION 2026-07-30 (superuser, mirroring the measured Office1 shape, which MAAS 3.7 requires for machine allocate/deploy) together with its 3-part API key and account password. Minted ON the region VM, never printed, moved host-to-host with all three sha256 digests compared, staging shredded. Consolidated to `~/vr1-dc0-creds/` on the jumphost (0600) and distributed to the dc0 RACK (0600 in a 0700 folder) because **D-138 moved the cloud-facing Juju client INTO the DC** -- so `rack` is now a credential-bearing host-role, which it had never been. dc1's is NOT YET MINTED (its region VM is authored, not applied). | 2026-07-30 Stage-5 Juju re-point; D-132 q1 + D-138; captures in `docs/audit/dc0-node-carve-20260730.txt` and this session's changelog | operator ("Mint juju-vr1-dc0 on the new region (Recommended)") | **OPEN -- rotation obligation + three notes.** (1) ROTATE at v1 close, or immediately if the rack or region VM is rebuilt or shared. (2) **BLAST RADIUS IS DC-LOCAL BY CONSTRUCTION and this is strictly BETTER than the alternative considered:** the rejected option was pointing juju at the region ADMIN key already consolidated on vcloud, which would have placed an admin-scoped credential on the rack -- the residency SEC-026 exists to constrain. This key administers dc0's region ONLY; no cross-DC reuse, per SEC-012/-016/-018/-019. (3) **`vm-secret-locations` gained its FIRST `rack` rows** (four, both DCs) -- definition-of-done for a new mint site, since an unlisted location is not audited (how SEC-022 happened). It also now covers the rack's `~/.local/share/juju/*` credential store, where juju keeps its own runtime copy. (4) dc1's four matrix rows are the FORWARD REGISTER and FAIL S2 by design until dc1's region exists -- that failure is D-137 working, and the rows must NOT be deleted to go green. **These two NEW S2 findings are NOT covered by the operator's 2026-07-30 P5 acceptance, which enumerated six and said so explicitly.** | | SEC-029 | 2026-07-30 | **Octavia PKI overlay resident on the DC RACK.** The overlay carries the Octavia controller CERT and its PRIVATE KEY, and `juju deploy` consumes it at deploy time. D-138 moved the cloud-facing juju client INTO the DC, and the deploy CANNOT run from `voffice1` at all -- there is no L3 path from the headend to the controller API on the provider plane. dc0's overlay was therefore copied to the dc0 rack 2026-07-30 (`~/repo-stage/overlays/vr1-dc0-octavia-pki.yaml`, `0600`, sha256 compared end to end and equal). dc1's is NOT yet copied. | 2026-07-30 Stage-5 bundle deploy; D-138; this session's changelog | operator ("Copy dc0's PKI overlay to the dc0 rack (Recommended)") | **OPEN -- residency, not a new key.** (1) This is a COPY of material that already exists on the headend (SEC-004 `n-overlay-in-clone`), not a new mint, so the rotation obligation remains the existing `octavia-pki.sh reissue` path -- **and that path must now update BOTH copies, or the rack will deploy a superseded cert.** (2) Per-DC isolation holds BY CONSTRUCTION: each rack receives ONLY its own DC's overlay, never the pair; a cross-DC copy would destroy the independence that `octavia-pki.sh verify`'s A11 assertion exists to prove. (3) `~/repo-stage/` on the rack is a STAGING directory created because the rack has NO REPO CLONE -- the same D-138 tooling gap as queued finding F1. If a clone is later placed on DC client hosts, this path moves and the `vm-secret-locations` row must move with it. | | SEC-030 | 2026-07-31 | **Local permission rules bypass the committed ask-gating -- ACCEPTED, KNOWN posture, recorded so it is explicit rather than implicit.** `.claude/settings.json` (committed team policy) deliberately puts `juju deploy`, `maas admin machine delete` and `maas admin * update/create/release` in **ask**, and the controller-destruction command in **deny**. `.claude/settings.local.json` (GITIGNORED, 286 allow rules accumulated across sessions, largely from interactive don't-ask-again approvals) contains `Bash(ssh *)`, `Bash(ssh voffice1 *)`, `Bash(ssh vr1-dc0-rack *)`, `Bash(ssh vr1-dc0-maas *)` and `Bash(python3 *)`. Because an `ssh <host> '<cmd>'` invocation matches on the OUTER ssh command, a destructive `maas admin machine delete` issued over ssh never reaches the ask rule; `python3 *` permits arbitrary execution outright. **CLAUDE.md hard rule 3 states the ask rules enforce per-mutation approval and are not to be worked around, so this is a real divergence between stated and enforced posture.** | 2026-07-31 permission carry-over review; continuation of `queued-findings-20260730-dc0-region-migration` F1 | operator ("Leave them as they are") | **OPEN -- ACCEPTED BY RULING, NOT AN OVERSIGHT.** (1) The operator was shown the five rules verbatim and the gating they bypass, and ruled to keep them: the mutation gate in practice is the operator's conversational approval of each presented command, not the rule engine. That has held all session. (2) **The compensating control is the PRESENTATION discipline** -- every mutation is still stated, justified as minimal, and approved before running. If that discipline lapses, nothing else catches it. (3) **NOT promoted to committed team policy** (ruling: "Promote a curated safe subset"): 36 read-only rules were promoted to `settings.json`; these five stay local-only so the bypass is NOT baked into team policy or carried to Roosevelt. (4) **`Read(//home/jessea123/**)` was deliberately EXCLUDED from promotion** -- it grants read of `~/vr1-*-creds/`, `~/vault-init/`, `~/as-executed/` and `~/tenant-*`, every one of which CLAUDE.md forbids reading into context. Promoting it would have put a secrets-read permission into team policy. (5) Four local wildcards were NARROWED before promotion because they permit mutation: `ip route *`, `ip neigh *`, `bridge fdb *`, `tc qdisc *` all allow add/del on the host running every DC node; only the `show`/`get` verbs were promoted. (6) A rebuild still loses the ~250 unpromoted local rules -- that half of F1 remains open by choice. | +| SEC-031 | 2026-08-02 | **The vr1-dc1 OPNsense edge is currently an OPEN ROUTER serving its management GUI to the simulated ISP.** The 2026-08-01 hard power cut (in-place tofu resize bouncing the containment VM) destroyed the edge's FreeBSD base-system user database -- `/etc/master.passwd` (`pwd_mkdb: line #1 too long`) and `/etc/group` (`pw: group 'wheel' does not exist`) -- so `Configuring firewall.....failed.` on every boot since and NO pf ruleset is in force. MEASURED against a control, which is what makes this a finding rather than a worry: from vcloud, the healthy office1 edge (same 26.7 image, same build path) answers WAN ICMP with 100% loss and its GUI with `000`; the dc1 edge answers WAN ICMP at **0% loss** and its GUI **200 in 0.014s**. It also forwards without translating -- simultaneous tcpdump on both taps shows the same packet leaving the WAN leg with source still `10.12.64.2`. Bounded by lab topology: the "ISP" is the `vr1-dc1-uplink` libvirt NAT net on vcloud, not the internet, and D-125 egress isolation is unaffected. Discovered by the read-only dc1 edge assessment, `docs/audit/dc1-edge-assessment-20260802.txt`. | docs/audit/dc1-edge-assessment-20260802.txt | operator | OPEN -- closes when the edge's pf ruleset is restored (the repair regenerates the user DB from config.xml via the proven D-112(c) `local_user_set()` path, or donates it from the pristine `opnsense-26.7-nano.qcow2` in the same pool). Config.xml itself is INTACT -- fsck names every inode it deletes and `/conf/config.xml` is not among them. NOT rebuild-blocking; the exposure ends with the repair. |