fix(container): reap ghost containers so an app can't be locked out of itself
Demo images / Build & push demo images (push) Successful in 3m24s
Demo images / Build & push demo images (push) Successful in 3m24s
A ghost is a container whose process tree is still running while podman
has no record of it: the exit-command's `cleanup --rm` deletes the record,
conmon and the payload survive. It keeps owning exactly what the app needs
— the published host port and the file locks in its data dir — so the
replacement container either fails to bind ("address already in use") or
starts and dies on the lock, and Restart=always loops it there forever.
Nothing in the stack could see it: every podman-level stop/rm/recreate
misses a container podman lost.
Seen twice now: 752 restarts on a fleet node (2026-08-10) and again on the
dev box today, where Gitea flapped until it fell out of My Apps. Both were
cleared by hand; container-doctor.sh has the same logic but is an
out-of-band script the daemon never calls.
- New container::ghost_reaper: finds conmon processes whose 64-hex
container id is absent from `podman ps -a --no-trunc -q`, then kills the
payload's children and conmon (TERM, 5s grace, then KILL — the Gitea
ghost ignored TERM). Id-based, never name-based: killing by name would
hit the live managed container. A failed `podman ps` reaps nothing
rather than treating every container as a ghost.
- Hooked at repair_before_package_start (covers package.start,
package.restart and the orchestrator start path) and in the boot
reconciler's 30s tick, so ghosts are cleared before an app is asked to
start and swept for every app continuously.
Restart feedback: the lifecycle RPCs return {"status":"restarting"} in
milliseconds and work in the background, so "Restarting..." flashed for a
few frames and the buttons went idle while the app was still down — the
click read as a no-op. The hero buttons now show a spinner and hold it off
the node's own state (starting/stopping/restarting/updating, plus running
+ health=starting), and the just-clicked action is held until the backend
confirms it picked the work up, with a 12s cap so an unresponsive node
still releases the controls.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
b113fafee4
commit
9ccc325a4d
@@ -312,10 +312,32 @@ function launchApp() {
|
||||
}
|
||||
|
||||
|
||||
/**
|
||||
* Hold the just-clicked action until the node's own state confirms it picked
|
||||
* the work up (or we give up waiting).
|
||||
*
|
||||
* The lifecycle RPCs return in milliseconds and do the real work in the
|
||||
* background, so clearing `pendingAction` on the promise left the buttons idle
|
||||
* while the app was still down — the operator sees a flash and assumes the
|
||||
* click did nothing. The hero section keeps the spinner running off the
|
||||
* backend's transitional state; this only has to bridge the gap until that
|
||||
* first state push lands. The timeout means a node that never reports back
|
||||
* still releases the controls instead of wedging them.
|
||||
*/
|
||||
async function holdUntilBackendPicksUp(timeoutMs = 12000) {
|
||||
const started = Date.now()
|
||||
while (Date.now() - started < timeoutMs) {
|
||||
const s = pkg.value?.state
|
||||
if (s === 'starting' || s === 'stopping' || s === 'restarting' || s === 'updating') return
|
||||
await new Promise((r) => setTimeout(r, 250))
|
||||
}
|
||||
}
|
||||
|
||||
async function startApp() {
|
||||
pendingAction.value = 'start'
|
||||
try {
|
||||
await store.startPackage(appId.value)
|
||||
await holdUntilBackendPicksUp()
|
||||
} catch (err) {
|
||||
showActionError(`Failed to start: ${err instanceof Error ? err.message : 'Unknown error'}`)
|
||||
} finally {
|
||||
@@ -330,6 +352,7 @@ async function stopApp() {
|
||||
// Stopping the app can take its admin credentials offline — invalidate
|
||||
// rather than show a stale "healthy" credentials card (T-02-12).
|
||||
credentialsResource.invalidate()
|
||||
await holdUntilBackendPicksUp()
|
||||
} catch (err) {
|
||||
showActionError(`Failed to stop: ${err instanceof Error ? err.message : 'Unknown error'}`)
|
||||
} finally {
|
||||
@@ -344,6 +367,7 @@ async function restartApp() {
|
||||
// A restart can rotate credentials/admin URLs — invalidate so the next
|
||||
// read is fresh rather than the pre-restart cache (T-02-12).
|
||||
credentialsResource.invalidate()
|
||||
await holdUntilBackendPicksUp()
|
||||
} catch (err) {
|
||||
showActionError(`Failed to restart: ${err instanceof Error ? err.message : 'Unknown error'}`)
|
||||
} finally {
|
||||
|
||||
Reference in New Issue
Block a user