dc1-edge-assessment-20260802.txt
================================================================================
READ-ONLY assessment of the vr1-dc1 OPNsense edge after the 2026-08-01 05:45:30
containment-VM power cut. NOTHING was mutated: no virsh start/destroy/define, no
config write, no service restart, no tofu. Every API call issued was a GET.
Companion to docs/audit/dc0-edge-egress-outage-20260802.txt (the dc0 half).

Every line below is tagged MEASURED (observed this session, or quoted from a
captured artifact) or INFERRED (reasoning over measurements). Per this repo's
discipline a finding is an observation, not a conclusion, until it is measured.

================================================================================
VERDICT
================================================================================
  Q1  Does the dc1 edge forward/NAT?
      It FORWARDS. It does NOT NAT. And it is enforcing NO EFFECTIVE FIREWALL
      RULESET -- measured: no outbound NAT on the wire, and no WAN block, in a
      controlled comparison against a healthy sibling edge. (Whether pf is
      disabled or loaded-but-permissive is NOT separable read-only; `pfctl -s
      info` in section 8 settles it, and nothing in the recommendation turns on
      which it is.) Not a D-125 isolation posture -- D-125 is the decision that
      RESTORES per-DC egress, and dc1 rack egress was MEASURED WORKING on
      2026-07-30. This is a regression, not a design.

  Q2  Is the config INTACT?
      YES, on positive evidence rather than on "it booted." fsck names every
      inode it deletes; all 50 named casualties are listed in section 5a and
      /conf/config.xml IS NOT AMONG THEM -- the whole /conf tree lost exactly
      one transient event file. The edge came up on its exact ruled as-built
      addresses (LAN 10.12.64.1/22, WAN 172.30.3.2/24) with routes, Unbound,
      Dnsmasq, NTP and the gateway monitor all "done." What fsck destroyed is
      the FreeBSD BASE SYSTEM user database -- /etc/master.passwd and
      /etc/group -- which is NOT part of config.xml, plus BOTH of their
      /var/backups/*.bak copies. Everything broken is downstream of that.
      QUALIFIER: strong but not absolute -- fsck's other 92 casualties are
      reported by inode, not by name (see 5a(i)). One `xmllint --noout
      /conf/config.xml` settles it and needs a shell (section 8).

  Q3  Rebuild, repair, or nothing?
      >>> REPAIR, not rebuild. <<<  dc0 lost its shared libraries and cannot
      reach multi-user; dc1 lost its passwd/group database and reaches
      multi-user with its whole configuration loaded. Different severity AND a
      different repair. See "RECOMMENDATION" for the caveat that decides how
      the repair has to be delivered.

  Q4  Other damage in dc1?
      NONE FOUND. The edge was the ONLY inner guest running at the cut (the
      other ten dc1 domains are shut off with autostart=disable). The rack host
      vvr1-dc1 (ext4) is clean: no ext4/IO errors, zero failed units, all four
      dc1 units active.

================================================================================
1. THE BASELINE THAT FRAMES EVERYTHING -- this was working two days ago
================================================================================
MEASURED (quoted from docs/CURRENT-STATE.md:115-121, a captured 2026-07-30
measurement, not prose):

  "ARTIFACT SOURCES AND EGRESS RE-MEASURED 2026-07-30, both DCs. ... dc1 proxy
   PASS (apt-cacher-ng active on 10.12.68.4:3142, serves archive AND UCA 200).
   Egress re-probed from BOTH rack hosts with `--noproxy '*'` so a cache hit
   could not fake it: juju agent stream `streams.canonical.com/juju/tools/` 200,
   snap store `api.snapcraft.io` answering, `archive.ubuntu.com` jammy Release
   200, 1.1.1.1 0% loss, default routes via 10.12.4.1 / 10.12.64.1."

This closes three of the four alternatives in the task's hypothesis list before
any live probe is taken:
  - "a deliberate D-125 isolation posture I am misreading" -- REFUTED. D-125 is
    the ruling that GIVES each DC a working ISP uplink (bridge-in, single NAT);
    its whole purpose is egress. And egress was measured working.
  - "no NAT rule was ever configured" -- REFUTED. Outbound NAT was translating
    on 2026-07-30.
  - "interfaces mis-assigned" -- REFUTED, and re-refuted live in section 3.
The only event recorded between that measurement and now is the 2026-08-01
05:45:30 hard cut. INFERRED: this is a regression caused by that cut. Sections
2-4 measure the mechanism.

================================================================================
2. THE SERIAL LOG -- a controlled before/after inside ONE file
================================================================================
Serial log path taken from the live domain XML, not assumed:
  MEASURED  `virsh dumpxml vr1-dc1-opnsense` ->
    <serial type='pty'>
      <log file='/var/lib/libvirt/vr1/staging/vr1-dc1-opnsense-serial.log'
           append='on'/>

INSTRUMENT CURRENCY (stated before any negative is drawn from this file):
  MEASURED  size=79424  mtime=2026-08-01 07:31:36.982889689 +0000
  MEASURED  the file contains exactly TWO boots, and `append='on'` means it was
            never rotated:
              boot 1 ends line 259, banner dated "Wed Jul 22 21:51:39 UTC 2026",
                     LAN (vtnet0) -> v4: 192.168.1.1/24  (FACTORY addressing;
                     the D-112(c) console bootstrap session follows it inline)
              boot 2 ends line 2670, banner dated "Sat Aug  1 07:31:36 UTC 2026",
                     LAN 10.12.64.1/22, WAN 172.30.3.2/24  (AS-BUILT addressing)
  MEASURED  fsck damage counts in the file: UNREF FILE 76, UNEXPECTED SOFT
            UPDATE INCONSISTENCY 181, SALVAGE 46, ld-elf errors 0.
            (Matches the dc0 report's dc1 control column exactly.)

  NOTE, and it is load-bearing: there is NO boot between 2026-07-22 and
  2026-08-01. The as-built config was applied LIVE over the D-113 API on
  2026-07-23 and the edge then ran continuously. The 2026-08-01 boot is
  therefore the FIRST TIME this edge has ever booted on its as-built config.
  See section 7, finding (c).

--- 2a. The pre-cut boot (lines 225-245, VERBATIM) -------------------------
  Setting hostname: OPNsense.internal
  Generating /etc/resolv.conf...done.
  Generating /etc/hosts...done.
  Configuring system logging...done.
  Configuring firewall.......done.
  Configuring hardware interfaces...done.
  Configuring loopback interface...done.
  Configuring LAN interface...done.
  Configuring WAN interface...done.
  Generating /etc/resolv.conf...done.
  Generating /etc/hosts...done.
  Configuring firewall.......done.
  Starting web GUI...done.
  Setting up routes...done.
  Starting Dnsmasq...done.
  Starting Unbound DNS...done.
  Configuring firewall.......done.

--- 2b. fsck, then the first casualty (lines 2548-2565, VERBATIM) ----------
  SUMMARY INFORMATION BAD
  SALVAGE? yes

  BLK(S) MISSING IN BIT MAPS
  SALVAGE? yes

  77557 files, 534225 used, 2165014 free (86 frags, 270616 blocks, 0.0% frag.)

  ***** FILE SYSTEM MARKED CLEAN *****

  ***** FILE SYSTEM WAS MODIFIED *****
  pwd_mkdb: line #1 too long
  pwd_mkdb: /etc/master.passwd: Inappropriate file type or format
  pwd_mkdb: line #1 too long
  pwd_mkdb: /etc/master.passwd: Inappropriate file type or format

  This is the dc1 equivalent of dc0's `ld-elf.so.1: libcrypto.so.17 not found`:
  the FIRST line after "FILE SYSTEM WAS MODIFIED" names the casualty. On dc0 it
  was the shared libraries. On dc1 it is /etc/master.passwd.

--- 2c. The post-cut boot (lines 2615-2650, VERBATIM) ----------------------
  Setting hostname: OPNsense.internal
  Generating /etc/resolv.conf...done.
  Generating /etc/hosts...done.
  Configuring system logging...done.
  Configuring firewall.....failed.               <<<<
  Configuring hardware interfaces...done.
  Configuring loopback interface...done.
  Configuring LAN interface...done.
  Configuring WAN interface...done.
  Generating /etc/resolv.conf...done.
  Generating /etc/hosts...done.
  Configuring firewall.....failed.               <<<<
  Configuring OpenSSH...failed.                  <<<<
  Starting web GUI...pwd_mkdb: line #1 too long
  pwd_mkdb: /etc/master.passwd: Inappropriate file type or format
  pw: passwd file update: No error: 0
  chown: wheel: illegal group name               <<<< x9
  ...
  done.
  Setting up routes...done.
  Starting Dnsmasq...done.
  Starting Unbound DNS...done.
  Configuring firewall.....failed.               <<<<
  Setting up gateway monitor...done.
  Syncing OpenVPN settings...done.
  Starting NTP service...done.
  Starting Unbound DNS...done.

  and further down, same boot:
    Configuring login behaviour...pw: group `wheel' does not exist
    Starting Cron: error in configd communication, see syslog for details
    error in configd communication, see syslog for details       (openvpn)
    chown: hostd: illegal user name / illegal group name
    /usr/local/etc/rc.d/hostwatch: WARNING: failed to setup hostwatch

MEASURED, exhaustive grep over the whole file:
  "Configuring firewall" appears at lines 229, 236, 241 (all "......done.")
  and 2619, 2626, 2644 (all "....failed.").
  "Configuring OpenSSH" appears ONCE, at line 2627, and it FAILED.

HONEST QUALIFIER on 2a: the three "done." lines are from the FACTORY-config
boot, so they prove the firewall subsystem works on THIS IMAGE, not that the
as-built config booted cleanly (it had never booted). The evidence that the
as-built config's NAT worked is the 2026-07-30 measurement in section 1, which
is stronger anyway because it is behavioral.

MEASURED  qemu started the edge at 2026-08-01 05:45 (mtime of
          /var/log/libvirt/qemu/vr1-dc1-opnsense.log); the boot banner is
          stamped 07:31:36. ~1h46m elapsed.
INFERRED  that gap is the fsck pass over a damaged UFS root (the ~2,000 lines
          of salvage output between the two). Not independently timed.

================================================================================
3. LIVE STATE -- what the edge is doing right now
================================================================================
All probes run 2026-08-02 ~06:00-06:20 UTC. Rack identity confirmed first:
  MEASURED  `ssh vr1-dc1-rack 'hostname'` -> vvr1-dc1 ; uptime "up 1 day, 10 min"
            (consistent with the 2026-08-01 05:45 boot)
  MEASURED  ip -4: enp1s0 172.31.0.6/30, virbr4 10.12.64.2/22,
            virbr6 10.12.68.2/.3/.4/22, virbr0 192.168.122.1/24
  MEASURED  ip route: "default via 10.12.64.1 dev virbr4"
  MEASURED  ip neigh dev virbr4: 10.12.64.1 lladdr 52:54:00:95:65:64 STALE
            -- and 52:54:00:95:65:64 is exactly the net0 MAC in the edge's
            domain XML. The gateway answering IS the edge VM, not a squatter.

Domain XML (MEASURED, `virsh dumpxml vr1-dc1-opnsense`):
  net0  mac 52:54:00:95:65:64  network vr1-dc1-provider-public  bridge virbr4
        target vnet0                                    (LAN leg)
  net1  mac 52:54:00:87:3b:74  network vr1-dc1-wan  bridge br-vr1-dc1-wan
        target vnet1                                    (WAN leg, D-125 bridge-in)
  disk  pool vr1-dc1-inner-pool vol vr1-dc1-opnsense-disk.qcow2, <backingStore/>
        EMPTY -- standalone qcow2, no copy-on-write base to roll back to.
        Identical to dc0. (qemu-img could not be run on it: the file is locked
        by the running guest. The empty <backingStore/> is the evidence.)
  Interfaces are CORRECTLY assigned. "mis-assigned interfaces" is refuted.

Reachability (MEASURED):
  from the rack (vvr1-dc1):
    ping 10.12.64.1   (edge LAN)  -> 2/2, 0% loss
    ping 172.30.3.2   (edge WAN)  -> 2/2, 0% loss
    ping 1.1.1.1                  -> 0/2, 100% loss
  from vcloud:
    ping 172.30.3.2               -> 2/2, 0% loss
    ip neigh -> 172.30.3.2 dev virbr1 lladdr 52:54:00:87:3b:74 REACHABLE
       (matches the net1 MAC: the D-125 bridge-in L2 path is intact end to end,
        inner tap -> br-vr1-dc1-wan -> vcloud virbr1)
    virsh net-dumpxml vr1-dc1-uplink -> forward mode='nat', ip 172.30.3.1/24,
       bridge virbr1 -- the simulated ISP is up and is a NAT for 172.30.3.0/24.

Service ports on the edge (MEASURED):
                       LAN 10.12.64.1        WAN 172.30.3.2 (from vcloud)
    tcp/22   (sshd)     closed / no answer    closed / no answer
    tcp/80              301                   301
    tcp/443             200                   200
    GET /  title        "<title>Login | OPNsense</title>"
  sshd is DOWN -- consistent with "Configuring OpenSSH...failed." The web GUI is
  UP and serving PHP.

API (D-113 key, MEASURED; the key was never printed, never copied to another
host -- opnsense-api.sh passes the secret to curl via --config on stdin):
    unauthenticated  GET /api/core/menu/search      -> 302 in 0.022s
    unauthenticated  GET /api/core/firmware/status  -> 302 in 0.017s
    AUTHENTICATED    GET core/menu/search           -> NO RESPONSE, killed at 40s
    AUTHENTICATED    GET diagnostics/interface/getInterfaceNames -> killed at 30s
    AUTHENTICATED    GET diagnostics/interface/getRoutes         -> killed at 30s
  The web stack answers unauthenticated requests in ~20ms and hangs forever the
  moment authentication is attempted -- including on core/menu/search, which does
  not use configd. INFERRED: the hang is in the authentication path, which is the
  one thing that must consult a user database that no longer parses. Not proven
  further; the API is simply not a usable instrument on this box.

================================================================================
4. THE DECISIVE TEST -- forwards yes, translates no
================================================================================
tcpdump on BOTH of the edge's tap devices simultaneously, on the rack (the taps
belong to the rack's libvirt), while the rack pinged 1.1.1.1. Read-only.

  MEASURED  vnet0 (edge LAN tap):
    06:04:16.702929 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 1
    06:04:17.757057 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 2
    06:04:18.781116 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 3
    06:04:19.805049 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 4

  MEASURED  vnet1 (edge WAN tap), SAME four packets, ~0.8 ms later:
    06:04:16.703747 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 1
    06:04:17.757724 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 2
    06:04:18.781871 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 3
    06:04:19.805778 IP 10.12.64.2 > 1.1.1.1: ICMP echo request, id 34238, seq 4

This settles Q1 completely:
  * The edge HAS a working default route out its WAN. Packets are routed LAN->WAN.
  * pf is NOT blocking them -- they are passed, not dropped.
  * The source address is STILL 10.12.64.2 on the WAN side. NO OUTBOUND NAT IS
    APPLIED. The packet lands on vcloud's virbr1, whose libvirt NAT masquerades
    only 172.30.3.0/24, so a 10.12.64.2 source is never translated and no reply
    can ever come back. That is exactly the observed "0% loss to the gateway,
    100% loss to the internet."

  So the failure is precisely: pf has no outbound-NAT ruleset loaded.
  It is NOT a routing failure, NOT an interface failure, NOT an L2 failure,
  NOT an upstream failure.

--- 4a. CONTROL: pf is enforcing NO ruleset at all -------------------------
An OPNsense edge with any ruleset loaded blocks unsolicited inbound on WAN --
there is no default WAN pass rule, and the anti-lockout rule is LAN-only. So
"can the ISP side reach the GUI?" is a direct test for a loaded ruleset.

  MEASURED, from vcloud, against the HEALTHY office1 edge (same OPNsense 26.7
  image, same D-112/D-113 build path -- the control):
    172.30.1.2 found via `ip neigh show dev virbr11` (not guessed)
    ping 172.30.1.2          -> 0/2, 100% loss     (ICMP blocked on WAN)
    https://172.30.1.2/      -> 000 after 12.006s  (no answer)
    http://172.30.1.2/       -> 000 after 12.006s  (no answer)
    tcp/22                   -> closed

  MEASURED, from vcloud, against the SUBJECT dc1 edge:
    ping 172.30.3.2          -> 2/2, 0% loss       (ICMP answered on WAN)
    https://172.30.3.2/      -> 200 in 0.014s      (GUI served to the ISP)
    http://172.30.3.2/       -> 301
    tcp/22                   -> closed (sshd is down, not firewalled)

  The dc1 edge is answering ICMP and serving its management GUI to the simulated
  ISP. A healthy sibling on the identical image does neither. Conclusion, stated
  to exactly the strength of the evidence: the dc1 edge has NO EFFECTIVE pf
  RULESET -- MEASURED, there is no outbound NAT (section 4) and no WAN block
  (this control).
  WHAT THIS DOES NOT DISTINGUISH: pf disabled / ruleset empty, versus pf loaded
  with a partial or permit-all ruleset. Both produce identical external
  behaviour and both produce no NAT, and neither can be separated without
  `pfctl -s info` on the guest (section 8). The recommendation is the same
  either way, so this is recorded as an open sub-question, not chased.

>>> SECURITY CONSEQUENCE, flagged not actioned (read-only session): MEASURED,
>>> there is NO INBOUND FILTERING IN FORCE ON THE dc1 EDGE'S WAN -- its
>>> management GUI answers 200 on its WAN address from the simulated ISP
>>> segment, and it answers ICMP there, neither of which the healthy control
>>> does. It is functioning as an open router between 172.30.3.0/24 and the dc1
>>> provider-public plane. The exposure is bounded by the lab's own topology
>>> (172.30.3.0/24 is a vcloud-local libvirt NAT, not routable inbound from the
>>> real internet), which is why this is a finding and not an incident. It
>>> belongs in docs/security-ledger.md as a new SEC row. I did not open one:
>>> numbering and ledger rows are the operator's to assign.

================================================================================
5. IS THE CONFIG INTACT? (Q2)
================================================================================
EVIDENCE THAT config.xml SURVIVED -- and none of it is "it booted":

  MEASURED  The post-cut boot configured both interfaces onto the EXACT ruled
            as-built values, which exist ONLY in config.xml:
              LAN (vtnet0) -> v4: 10.12.64.1/22
              WAN (vtnet1) -> v4: 172.30.3.2/24
            These match changelog-20260723-g12-dc1-edge.md Item 5 and
            lib-net.sh's provider-public gateway exactly. A damaged or reverted
            config would have come up on the factory 192.168.1.1/24 (which is
            what boot 1 in the very same log shows it doing).
  MEASURED  "Configuring LAN interface...done." / "Configuring WAN
            interface...done." / "Setting up routes...done." -- the config
            parsed and drove the interface + routing layer.
  MEASURED  Dnsmasq, Unbound (twice), NTP, the gateway monitor and the web GUI
            all started "done." on the post-cut boot.
  MEASURED  Section 4 proves a default route exists and works -- i.e. the
            <gateways> data is present and usable.
  MEASURED  The web GUI serves its login page (200, correct <title>), so the
            PHP stack that READS config.xml is functional. dc0's is not.

WHAT WAS ACTUALLY DESTROYED -- and it is not config:
  MEASURED  /etc/master.passwd is unparseable: pwd_mkdb reports "line #1 too
            long" and "Inappropriate file type or format" on every one of four
            attempts across the boot. So /etc/pwd.db and /etc/spwd.db were never
            rebuilt.
  MEASURED  /etc/group is damaged: `pw: group 'wheel' does not exist`, and
            17 x `chown: wheel: illegal group name` plus 4 x `chown: hostd:
            illegal user name` / `illegal group name` (exact grep counts).
  Both are FreeBSD BASE SYSTEM files. Neither is part of config.xml. They are
  fsck casualties in the same class as dc0's libcrypto.so.17 -- the file was
  reconnected as garbage or truncated, not removed.

--- 5a. THE NAMED-CASUALTY INVENTORY -- the strongest evidence in this file ---
fsck names every inode it DELETES. MEASURED, exhaustive over the serial log:
    REMOVE? yes        50        UNALLOCATED  50        NAME=  50
    RECONNECT? yes     92
All 50 named casualties, deduped and complete:

    /conf/event_config_changed.json          <-- the ONLY /conf casualty
    /root/.history
    /usr/local/etc/bogons
    /usr/local/etc/bogonsv6
    /usr/local/etc/unbound.opnsense.d/access_lists.conf
    /usr/local/etc/unbound.opnsense.d/domainoverrides.conf
    /usr/local/etc/unbound.opnsense.d/dot.conf
    /usr/local/etc/unbound.opnsense.d/safesearch.conf
    /usr/local/etc/unbound/unbound-blocklists.conf
    /usr/local/etc/unbound_dhcpd.conf
    /usr/local/lib/perl5/5.42/perl/man/mandoc.db
    /usr/local/lib/perl5/site_perl/man/mandoc.db
    /usr/local/lib/python3.13/site-packages/__pycache__/unboundmodule.cpython-313.pyc
    /usr/local/share/man/mandoc.db
    /usr/share/openssl/man/mandoc.db
    /var/backups/gpart.vtbd0.bak
    /var/backups/group.bak                   <<<< THE BACKUP OF A CORRUPT FILE
    /var/backups/kern.geom.conftxt.bak
    /var/backups/master.passwd.bak           <<<< THE BACKUP OF A CORRUPT FILE
    /var/backups/pkg.sql.xz.1 .2 .3 .4 .7
    /var/cron/tabs/root
    /var/db/dhcp6c_duid
    /var/db/entropy/saved-entropy.8
    /var/db/locate.database
    /var/db/pkg/vuln.xml
    /var/run/cron.pid  dhcp6c.pid  dpinger_WAN_GW.pid  dpinger_WAN_GW.sock
    /var/run/gateway_watcher.pid  rtsold.pid  sshd.pid  unbound.pid  utx.active
    /var/unbound/advanced.conf  cache.dump.gz  host_entries.conf
    /var/unbound/private_domains.conf  resolv.conf.root  root.hints  root.key
    /var/unbound/unbound.conf  unbound_control.key  unbound_control.pem
    /var/unbound/unbound_server.key  unbound_server.pem

TWO CONCLUSIONS, both directly measured:

  (i) /conf/config.xml IS NOT IN THE LIST. The entire /conf tree lost exactly
      one file, and it is `event_config_changed.json` -- a transient event
      marker, not configuration. Combined with the fact that the config
      demonstrably drove the boot onto its exact as-built addresses, this is
      positive evidence of survival rather than an inference from "it booted."
      QUALIFIER, because absence of evidence is not evidence of absence: the
      92 RECONNECT (UNREF FILE) entries are reported by INODE, not by name, so
      fsck could in principle have unlinked a file into lost+found without
      printing its path. Absence from this list is strong but not absolute.
      The check that settles it is `xmllint --noout /conf/config.xml` -- see
      section 8.

  (ii) BOTH FreeBSD-NATIVE BACKUPS OF THE CORRUPT FILES WERE ALSO DESTROYED.
      /var/backups/master.passwd.bak and /var/backups/group.bak are on the
      REMOVE list. So the obvious repair -- copy the .bak back over the corrupt
      original -- IS NOT AVAILABLE. This corrects an earlier draft of this
      assessment which read the master.passwd.bak line as evidence the backup
      still existed; it appears in the log as
      `UNALLOCATED I=77473 OWNER=0 MODE=0 SIZE=0 ... REMOVE? yes`,
      i.e. as a deletion, not a survival. Repair must regenerate the user
      database from config.xml instead (section 7).

  A third, minor observation: nearly every /var/unbound file was destroyed,
  including unbound's TLS keypairs -- yet "Starting Unbound DNS...done."
  appears twice on the post-cut boot. INFERRED: OPNsense regenerates the whole
  /var/unbound tree from config.xml at boot, which is another datapoint that
  config.xml is intact and that the config->runtime generators work. The one
  generator that did NOT work is the firewall's.

CAUSAL CHAIN
  MEASURED  the damage and the failures are in the same boot, in this order:
            fsck salvage -> master.passwd/group unparseable -> firewall
            configure failed x3 -> OpenSSH configure failed -> configd
            communication errors -> API auth hangs -> no pf ruleset -> no NAT.
  INFERRED  that the passwd/group corruption is the CAUSE of the firewall
            configure failure. The adjacency is measured; the specific internal
            dependency (OPNsense's filter_configure resolving users/groups, or
            configd failing to start for the same reason) is NOT proven, because
            proving it needs a shell on the edge, which I do not have read-only.
  ALTERNATIVE NOT FULLY EXCLUDED, stated honestly: because 2026-08-01 was the
            FIRST boot this edge ever took on its as-built config, a latent
            boot-time config defect cannot be excluded by this evidence alone.
            It is unlikely -- the fsck damage to base files is measured and
            sufficient, and outbound NAT was working live on 2026-07-30 -- but
            it is not zero, and the repair path in section 7 distinguishes them
            for free (if repairing the user DB restores NAT, the alternative is
            dead).

WHAT I COULD NOT READ, and why (this is a gap, not a pass):
  I have NOT read config.xml itself, nor /conf/backup/, nor the live pf ruleset
  (`pfctl -sr` / `pfctl -sn`). All three need a shell on the edge, and all three
  read-only routes to one are closed:
    - SSH: tcp/22 does not answer (MEASURED). sshd failed to configure.
    - REST API: authenticated GETs hang indefinitely (MEASURED).
    - Serial console: the guest is sitting at `login:`. Typing at it is an
      interaction with the guest, which I read as outside "READ-ONLY, make no
      mutations," so I did not attempt it. INFERRED, and worth the operator's
      attention before anyone tries: a console login will most likely FAIL
      anyway, because authentication needs the same /etc/master.passwd that
      pwd_mkdb cannot parse.
  If the operator wants config.xml and the pf ruleset read directly, the
  commands are listed in section 8. I am stopping rather than guessing.

================================================================================
6. ANY OTHER dc1 DAMAGE FROM THE SAME CUT? (Q4)
================================================================================
INNER GUESTS -- MEASURED, `virsh list --all` + per-domain `dominfo`:
    vr1-dc1-opnsense     running    Autostart=enable
    vr1-dc1-control-01   shut off   Autostart=disable
    vr1-dc1-control-02   shut off   Autostart=disable
    vr1-dc1-control-03   shut off   Autostart=disable
    vr1-dc1-compute-01   shut off   Autostart=disable
    vr1-dc1-compute-02   shut off   Autostart=disable
    vr1-dc1-storage-01   shut off   Autostart=disable
    vr1-dc1-storage-02   shut off   Autostart=disable
    vr1-dc1-storage-03   shut off   Autostart=disable
    vr1-dc1-storage-04   shut off   Autostart=disable
    vr1-dc1-juju-01      shut off   Autostart=disable
  Eleven domains, one running. The edge is the ONLY guest that was hard-cut.
  This is D-127 behaving exactly as written and it BOUNDS the blast radius:
  there is no other inner filesystem to inspect. (Consistent with the task's
  note that dc1 has no per-DC MAAS region VM -- MEASURED: no such domain exists.)

THE RACK HOST vvr1-dc1 (ext4) -- MEASURED:
    df/findmnt   /dev/vda1 ext4 2.9T, 1% used, rw,relatime,discard,
                 errors=remount-ro -- NOT remounted read-only
    journal      Aug 01 05:45:41 EXT4-fs (vda1): mounted ... ro, ordered data
                 Aug 01 05:45:41 EXT4-fs (vda1): re-mounted ... r/w
                 systemd-fsck-root skipped (unmet condition) -- normal
                 BOOT: recovering journal; BOOT: clean, 604/58496 files
                 UEFI (vfat): "differences between boot sector and its backup
                 ... mostly harmless"
    journal -p err for the whole boot: ONE kernel line ("RDSEED32 is broken.
                 Disabling the corresponding CPUID bit." -- a CPU-feature
                 notice, unrelated) plus a repeating libvirtd line, below.
    systemctl --failed:  "0 loaded units listed."
    dc1 units:   apt-cacher-ng.service          active running
                 dc1-cache-proxy-net.service    active exited
                 dc1-node-dns.service           active running
                 dc1-rack-legs.service          active exited
                 libvirtd.service               active running
  ext4 came through the cut cleanly. A journalling filesystem with an intact
  journal is the expected outcome; UFS-with-soft-updates is not, which is the
  whole reason the two edges are damaged and neither rack is.

  INSTRUMENT-CURRENCY NOTE, because a negative is being reported here:
    `dmesg` is NOT a valid instrument for the boot window on this host. It holds
    4834 lines whose OLDEST timestamp is [75228.4] (~20.9h uptime) against an
    uptime of ~24.5h -- the ring buffer HAS WRAPPED and the boot messages are
    gone from it. Its "zero ext4 errors" covers only the last ~3.6 hours. The
    boot-window evidence above is from `journalctl -b`, which does cover
    2026-08-01 05:45:41 onward. The dmesg ring is dominated by a flood of
    `apparmor="DENIED" ... profile="snap.maas.pebble" ... comm="ps"` lines --
    that is the MAAS RACK controller snap on this host reading /proc/N/environ;
    pre-existing noise, not damage, and not related to this incident.

  PRE-EXISTING NOISE, explicitly NOT attributed to the cut -- MEASURED:
    `libvirtd[2531]: End of file while reading data: Input/output error`
    recurs in ~10-minute clusters from Aug 01 05:51:57 onward; 1406 occurrences
    before this session began. INFERRED: a periodic libvirt client
    disconnecting (power-state polling shape). It predates and is unaffected by
    my probes. Logged so nobody later mistakes it for new damage.

================================================================================
7. RECOMMENDATION
================================================================================
>>> REPAIR the dc1 edge. Do NOT rebuild it. <<<

WHY REPAIR IS THE RIGHT CALL (all measured):
  a) The configuration survived. The edge is running its exact as-built
     identity -- LAN 10.12.64.1/22, WAN 172.30.3.2/24, working default route,
     working forwarding, DNS/NTP/GUI up. A rebuild would throw all of that away
     and re-incur the D-112(c) console bootstrap, the D-113 API-key mint, the
     re-addressing, and a fresh SEC row for the transient key exposure -- to
     recreate a config that is already correct.
  b) The break is a bounded, named, base-system defect: two corrupt files
     (/etc/master.passwd, /etc/group) and the services that depend on them
     (pf ruleset generation, sshd, configd, API auth).
  c) dc0 is genuinely different and genuinely needs the rebuild: it lost
     libcrypto.so.17 and libpython3.13.so.1.0, so PHP and python cannot even
     load and it never reaches multi-user. dc1's PHP runs and serves pages.
     Same incident, same damage class, materially different remedy. The task's
     hypothesis -- "same incident, milder outcome" -- is CONFIRMED, and the
     "milder" is load-bearing: it changes the remedy.

THE CAVEAT THAT DECIDES HOW, and it is the important part of this section:
  MEASURED, dc1 currently has NO working credential path into the guest:
    - sshd is down (tcp/22 does not answer)
    - authenticated API calls hang
    - INFERRED: console login will fail too, since it needs the same unparseable
      /etc/master.passwd
  dc0, paradoxically, is EASIER to get a shell on -- it dropped to a single-user
  prompt, which is a root shell. dc1 is sitting at a `login:` that probably
  cannot authenticate anyone.
  So the repair almost certainly requires BOOTING THE dc1 EDGE INTO SINGLE-USER
  MODE (interrupt the loader, `boot -s`) to get a shell, then rebuilding the
  user database.
  AND THE OBVIOUS RESTORE PATH IS CLOSED -- MEASURED, section 5a(ii): fsck
  DELETED both /var/backups/master.passwd.bak and /var/backups/group.bak, so
  there is no native backup to copy back. The repair therefore has to
  REGENERATE the user/group database rather than restore it. The two candidate
  sources, in preference order:
    1. OPNsense's own generators, driven from the surviving config.xml -- the
       same `local_user_set()` path the D-112(c) console bootstrap already used
       on this exact edge (see changelog-20260723-g12-dc1-edge.md Item 2, which
       records the working PHP shape including the `require_once("util.inc")`
       fix). This is the repo-proven route and it reuses tooling that exists.
    2. The pristine base image in the same pool -- MEASURED, `virsh vol-list
       vr1-dc1-inner-pool` holds `opnsense-26.7-nano.qcow2` -- as a donor for
       the stock /etc/master.passwd and /etc/group, which for an appliance are
       essentially image constants. Then re-run (1) to re-apply config users.
  THAT IS A REBOOT. It is a mutation. I did not do it and I am not proposing to
  do it inside this assessment -- it needs its own gated step, and it should be
  preceded by a disk snapshot/copy of vr1-dc1-opnsense-disk.qcow2 (the domain
  XML shows an EMPTY <backingStore/>: there is no rollback point today, on
  either edge). Copying that disk requires the guest to be shut down or the
  copy to be taken with the guest quiesced -- also a mutation, also gated.

FALLBACK IF REPAIR FAILS: the rebuild material is present and identical to
dc0's -- MEASURED, `virsh vol-list vr1-dc1-inner-pool` contains the pristine
`opnsense-26.7-nano.qcow2` alongside `vr1-dc1-opnsense-disk.qcow2`. So a rebuild
remains available at no additional risk; it is simply the more expensive answer
and the evidence does not call for it.

DURABLE FINDINGS THIS RAISES (logged, not actioned -- hard rule 1):
  (a) The dc0 report's procedure gap is confirmed by a second, independent
      datapoint: an "in-place" tofu change to a containment VM is a power cut to
      every inner guest, and BOTH edges took filesystem damage from it. The
      remedy the dc0 report proposes -- shut inner guests down first -- would
      have prevented both.
  (b) dc-egress-check.sh's A2 assertion cannot see this failure mode. On dc1 the
      edge ANSWERS at 10.12.64.1 (A2 OK) and forwards, and the DC still has no
      egress. A2 is doing its job for the dc0 shape and is simply not the
      instrument for the dc1 shape. A candidate A2b: "the edge translates" --
      and it should be written in the POSITIVE form: assert that traffic
      observed leaving the edge's WAN side carries the WAN SOURCE ADDRESS (the
      section 4 tcpdump), which is a real positive assertion and can fail
      loudly. The tempting cheap version -- "assert the edge does NOT serve its
      GUI on its WAN address", i.e. the section 4a control -- is an INVERTED
      assertion that reads PASS whenever a curl times out for any reason, which
      is exactly the "a checker that cannot fail is not a gate" / "could not
      look is never nothing there" shape this repo has already been bitten by
      twice. Use the control as corroboration, never as the gate itself.
      Worth raising as a gate improvement; I have NOT touched the script.
  (c) NEITHER DC EDGE HAD EVER BEEN REBOOT-TESTED. MEASURED for dc1: the serial
      log contains exactly two boots, 2026-07-22 (factory) and 2026-08-01
      (as-built, first ever). The as-built config was applied live over the API
      on 2026-07-23 and never proven to survive a restart. Both edges' first
      real reboot was an unclean one, which is the worst possible way to find
      out. Roosevelt delta: an edge's config is not "built" until it has been
      proven across a clean reboot. That reads like D-NNN material rather than a
      changelog line, but assigning the number is the operator's call.
  (d) SECURITY: section 4a's open-router finding needs a SEC row.

================================================================================
8. WHAT I COULD NOT ANSWER READ-ONLY, and the exact commands that would
================================================================================
Stated rather than guessed, per the ground rules.

  Q: Is config.xml byte-intact, and are /conf/backup/ revisions present?
  Q: What exactly is in the live pf ruleset -- empty, or absent?
  Q: Is it really master.passwd/group that breaks filter_configure?

  All three need a shell on the edge. Every read-only route to one is closed
  (section 5). The minimal commands, once a shell exists, are:
      ls -la /conf/config.xml /conf/backup/ | tail -20
      xmllint --noout /conf/config.xml && echo XML-WELL-FORMED
      grep -c '<rule>' /conf/config.xml ; grep -o '<outbound>' /conf/config.xml
      pfctl -s info ; pfctl -sr | head -40 ; pfctl -sn | head -40
      wc -c /etc/master.passwd /etc/group ; head -3 /etc/group
      ls -la /var/backups/          # EXPECT master.passwd.bak and group.bak to
                                    # be ABSENT -- fsck removed both (5a(ii))
      ls -la /lost+found | head     # the 92 RECONNECTed inodes land here
  Getting that shell requires `boot -s` at the loader, i.e. a REBOOT of the
  edge guest. That is a mutation and it is where I stop.

================================================================================
END. Nothing in this session was mutated. The repo tree is left dirty and
uncommitted by instruction.
================================================================================
