fix: bound monitoring subprocesses and collect fresh system readings

This commit is contained in:
archipelago
2026-10-05 21:31:32 -04:00
parent 3d0c67eb9b
commit e97f958f45
4 changed files with 95 additions and 19 deletions
+23
View File
@@ -38,3 +38,26 @@ mobile/desktop browser layout; upgraded sender and receiver comparing local
Monitoring to received Fleet values; restart/reconnect, offline ageing and real
transport evidence. Legacy senders still advertising zeros need the sender fix;
the receiver cannot reliably distinguish a legacy fabricated zero from idle CPU.
## Dev deployment and collector follow-up
Dev qualification deployment uses backend041f1fa2 (SHA256
11b62f697a71b762bf8638063d1a858a68eea6d7268e300f6320c99068efa78a)
and frontend3d0c67eb. Health and unchanged app-container identities/start times
pass. Live federation CPU/memory/disk values exactly match local Monitoring;
Web5 mobile390px/desktop1440px checks pass. Rollback retained on the node at
/var/lib/archipelago/support/followup-20261005-2120. Yaya remains on the prior
backend; no receiver/fleet-wide acceptance is claimed.
Live testing caught an unbounded podman stats subprocess delaying the first
snapshot, and a300-second collection interval conflicting with180-second Fleet
freshness. The next candidate starts after5seconds and collects once per minute,
with3-second df and8-second podman deadlines and kill-on-drop cleanup. Failed
system reads no longer become invented zero samples. Container-stat failure
leaves container readings unavailable while retaining valid system readings.
The actual subprocess timeout/reaping regression passes; all1,686 backend tests
pass (four ignored). This collector correction is not deployed yet.
Framework SSH and backend health work, but its stored dashboard session returns
401. Authenticated Framework Monitoring acceptance remains open. No authentication
boundary was bypassed to produce an apparent pass.