STAGE 5 dc0 -- ceph apt-wedge remediation (ceph-mon/2 + ceph-radosgw/0)
========================================================================
Date: 2026-08-05. Model: vr1-dc0. Operator-gated live-cloud mutation session.
PROVENANCE / HOW TO READ THIS FILE
----------------------------------
This is TRANSCRIBED FROM THE CLAUDE-CODE SESSION TRANSCRIPT, not a script(1)
`-aqe` capture -- no as-executed wrapper ran (the harness classifier refuses the
wrapped form; see memory as-executed-log-vs-classifier). Every command below was
run from the vcloud jumphost via `ssh vr1-dc0-rack '<juju ...>'` (D-138: the juju
client lives on the dc0 rack). The command lines and their output blocks are
VERBATIM as observed this session. GA-R1 rule 2 basis: quoted command + its output.
>>> THE BEFORE-STATE IS UNRECOVERABLE. <<< The CLOSE-WAIT sockets, the ~12.5h
apt etimes, and the two hung containers (2/lxd/0, 0/lxd/2) were DESTROYED by the
remediation. The only evidence they ever existed is this transcript. That is why
this section is captured verbatim: without it the root cause would be re-derived
from scratch (this is the 3rd recurrence of the apt-cacher-ng class).
== BEFORE (measured this session, pre-remediation) ==============================
(1) juju status -- ceph layer BLOCKED, two units allocating on PENDING containers:
ceph-mon blocked 2/3 Insufficient peer units to bootstrap cluster (require 3)
ceph-osd waiting 4 Incomplete relation: monitor
ceph-radosgw waiting 0/1 waiting for machine
ceph-rbd-mirror blocked 1 'ceph-local' incomplete, 'ceph-remote' missing [deferred: dc1 peer]
cinder-backup waiting 1 Incomplete relations: ceph
glance waiting 1 Incomplete relations: storage-backend
nova-compute waiting 2 Incomplete relations: storage-backend
ceph-mon/2 waiting allocating 2/lxd/0 waiting for machine
ceph-radosgw/0 waiting allocating 0/lxd/2 waiting for machine
Machine 2/lxd/0 pending juju-d4f45a-2-lxd-0 Container started
Machine 0/lxd/2 pending juju-d4f45a-0-lxd-2 Container started
(2) The PROXY IS HEALTHY -- refutes the carried "apt-cacher-ng is down" conclusion:
dc-cache-proxy check (dc0): PASS (apt-cacher-ng active, :3142, archive+UCA 200)
dc-snap-proxy check (dc0): PASS (squid :3129, store API 200)
(3) Container level -- both RUNNING with addresses, but cloud-init NEVER finished,
no jujud agent installed:
juju-d4f45a-0-lxd-2 RUNNING 10.12.9.4/10.12.4.147/10.12.32.103/10.12.12.118
juju-d4f45a-2-lxd-0 RUNNING 10.12.8.193/10.12.36.1/10.12.32.1
cloud-init status --long -> status: running (both)
systemctl is-active "jujud-machine-*" -> (empty; agent never installed)
(4) THE SMOKING GUN -- cloud-init modules:final hung in apt-get update for ~12.5h:
(ps -eo pid,etimes,stat,args ; etimes = elapsed SECONDS)
0/lxd/2: 1201 45241 S /usr/bin/apt-get ... --quiet update
1212 45241 S /usr/lib/apt/methods/http
1213 45241 S /usr/lib/apt/methods/http
2/lxd/0: 1329 45297 S /usr/bin/apt-get ... --quiet update
1340 45297 S /usr/lib/apt/methods/http
cloud-init-output.log tail: a run of "Get:.. archive.ubuntu.com .." lines then
STOPS mid-fetch. Started Tue 04 Aug 2026 20:24-20:25 = the 08-04 redeploy window.
(5) ROOT CAUSE -- apt http methods parked in CLOSE-WAIT to the proxy, no timeout:
(ss -tnp)
0/lxd/2: CLOSE-WAIT 10.12.9.4:40210 10.12.8.4:3142 users:(("http",pid=1212))
CLOSE-WAIT 10.12.9.4:40242 10.12.8.4:3142 users:(("http",pid=1213))
2/lxd/0: CLOSE-WAIT 10.12.8.193:40236 10.12.8.4:3142 users:(("http",pid=1340))
apt proxy config (correct): Acquire::http::Proxy "http://10.12.8.4:3142";
Proxy serves the upstream the hang stalled on:
curl -x http://10.12.8.4:3142 http://security.ubuntu.com/.../jammy-security/InRelease
-> HTTP 200 (0.668s)
CONCLUSION: proxy closed its side ~12.5h ago (concurrent-load blip during the
08-04 redeploy); apt never closed its side and has NO read timeout -> parked
forever. NOT a live proxy fault, config fault, or routing fault. Stale half-closed
connections apt never recovered from. Same class remediated 08-03/08-04 (mysql).
== REMEDIATION (operator-gated mutations, run from vr1-dc0-rack) =================
Bundle placement (source of truth, from bundle.yaml -- NO inferred values):
ceph-mon to: [lxd:0, lxd:1, lxd:2]
ceph-radosgw to: [lxd:0]
--- ceph-mon/2 (operator-approved "Yes, try it") ---
juju remove-unit ceph-mon/2 -m vr1-dc0
-> ERROR unit removal: aborted (interactive "Continue [y/N]?" with no tty)
juju remove-unit ceph-mon/2 -m vr1-dc0 --no-prompt -> will remove unit ceph-mon/2 (rc=0)
>>> KEY OPERATIONAL FACT: remove-unit did NOT cascade to the never-provisioned
container. `life` stayed alive; 2/lxd/0 lingered RUNNING/pending. A graceful
`remove-machine 2/lxd/0` marked it dying but the reaper BLOCKED waiting to
stop an agent that never came up. Cleanup required --force. <<<
juju remove-machine 2/lxd/0 -m vr1-dc0 --force --no-prompt (operator-approved)
-> 2/lxd/0 GONE from juju + CONFIRMED-GONE-ON-HOST
juju add-unit ceph-mon -m vr1-dc0 -n 1 --to lxd:2 -> ceph-mon/3 on fresh 2/lxd/5
RESULT: fresh container cloud-init FINISHED in ~211s (no apt hang; jujud active).
ceph-mon 3/3 -> "Monitor bootstrapped" -> all 3 mons active "Unit is ready and
clustered"; 4 ceph-osd -> active "Unit is ready (1 OSD)". Storage cascade cleared:
cinder-ceph, cinder-backup, glance, nova-compute -> active.
--- ceph-radosgw/0 (operator-approved "Approved") ---
juju remove-unit ceph-radosgw/0 -m vr1-dc0 --no-prompt
juju remove-machine 0/lxd/2 -m vr1-dc0 --force --no-prompt -> GONE (juju + host)
juju add-unit ceph-radosgw -m vr1-dc0 -n 1 --to lxd:0 -> ceph-radosgw/1 on 0/lxd/20
RESULT: container provisioned cleanly. SECOND apt-cacher-ng failure MODE observed
here (distinct from CLOSE-WAIT): ceph-radosgw-hacluster/0 install hook failed 4x,
apt exit status 100:
E: Failed to fetch http://archive.ubuntu.com/.../linux-libc-dev_5.15.0-187.197_amd64.deb
404 Not Found [IP: 10.12.8.4 3142]
= STALE apt-cacher-ng cached index pointing at a point-release .deb upstream had
rotated out of the pool. SELF-HEALED: juju auto-retried; once the cached index
refreshed, `apt-get update` in-container returned rc=0 (all Hits) and the retry
installed pacemaker/corosync. hacluster -> active "Unit is ready and clustered";
rgw + apache2 services active (running); ceph-radosgw/1 -> active (stale
"Services not running" status cleared on the next update-status).
== AFTER (freshly measured 2026-08-05 17:34:08Z) ===============================
CENSUS: {active: 62, waiting: 1, blocked: 3, unknown: 1}
Non-active, all accounted for:
barbican-vault/0 waiting 'secrets-storage' incomplete [settling -- verify]
ceph-rbd-mirror/0 blocked 'ceph-remote' missing [DEFERRED: dc1 peer, Stage 6 / dc-dc-phase5]
designate/0 blocked nameservers must be set [DEFERRED: DNS activation, Stage 7 / dc-dc-phase6]
glance-simplestreams-sync/0 unknown [EXPECTED: gss one-shot resting state]
octavia/0 blocked Awaiting configure-resources [DEFERRED: D-021, Stage 5 / phase-05]
Census progression actually measured this session: 59 (mid-settle, both re-adds
done) -> 61 -> 62. NOTE: no pre-remediation TOTAL was measured; do not cite a
"40->" or "47->" transition (those are prose figures from the 08-05 close block,
not a census). The measured transition is qualitative (ceph layer blocked -> the
storage layer fully converged) plus 59->62 on the settle.
== FINDINGS (owed; logged here + in the session changelog, NOT executed) ========
F1. appendix-A GAP -> entry drafted this session. TWO apt-cacher-ng failure modes:
(a) apt http method CLOSE-WAIT to :3142, no client timeout -> cloud-init hangs
forever -> unit stuck `allocating`, container `pending`/"Container started".
(b) stale cached index -> 404 on a rotated point-release .deb -> charm apt-get
install fails (exit 100); self-heals on index refresh + juju hook retry.
Remediation both: confirm proxy healthy -> remove-unit --no-prompt ->
remove-machine <c> --force --no-prompt (dead-agent container does NOT cascade)
-> add-unit <app> --to <bundle placement>.
F2. phase-03-core-verify.md Step 3.1 gate is STALE for VR1. It asserts expected
non-active/idle = 1 (octavia). MEASURED here: 4 deferred-by-design (octavia,
designate, ceph-rbd-mirror) + gss unknown. A DOCFIX is owed (unnumbered until
implemented). Also owed: -m openstack -> -m vr1-dc0, and run-from-dc0-rack
(D-138), throughout that runbook.
F3. dc-cache-proxy.sh `check` does not test security.ubuntu.com (the exact upstream
the CLOSE-WAIT hang stalled on; it serves fine, 200/0.67s, but the checker is
blind to it). A DOCFIX is owed (unnumbered).
F4. RECURRENCE (3rd: 08-03 mysql, 08-04 mysql, 08-05 ceph x2). remove/re-add is a
WORKAROUND. Preventive fix is DESIGN material for an operator ruling: inject
Acquire::http::Timeout / Acquire::Retries into node apt config at provision so
apt cannot hang (mode a) or hard-fail without retry (mode b). D-NNN candidate.