fix(boot): signal systemd READY before heavy boot recovery + Restart=always
Root cause of 'server starting up' forever / crash-on-install (framework-pt, v1.7.114->115, 2026-07-26): on a node with many stacks, the synchronous boot recovery (recover + start_stopped_containers) runs BEFORE sd_notify(Ready), so the unit sits in 'activating' for minutes. Anything touching the service in that window — a superseding start/restart, an install-time reconcile churn — killed a half-started instance; it exits 0 on SIGTERM and Restart=on-failure then never restarts it. Node dead behind 'server starting up'. Fixes: - signal READY (+ start the watchdog keepalive) BEFORE boot recovery, so the unit reaches 'active' in seconds; recovery/reconcile/listener continue after. No more minutes-long activating window. - Restart=always (was on-failure): a clean-exit SIGTERM must still bring the daemon back. Manual is still honored. - OTA restart via a PID1-owned transient timer (systemd-run --on-active=2) instead of a tokio-sleep child of the process being stopped, whose start-half was being lost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1868,13 +1868,38 @@ pub async fn apply_update(data_dir: &Path) -> Result<()> {
|
||||
// UI before systemd kills us. --no-block makes sure systemctl doesn't
|
||||
// try to wait for the current service (us) to exit cleanly before
|
||||
// starting the new process — it would deadlock otherwise.
|
||||
tokio::spawn(async {
|
||||
tokio::time::sleep(std::time::Duration::from_secs(2)).await;
|
||||
// systemctl talks to PID 1 over D-Bus — doesn't need the host
|
||||
// mount namespace, but routing through host_sudo keeps the
|
||||
// apply flow's sudo calls uniform.
|
||||
let _ = host_sudo(&["systemctl", "--no-block", "restart", "archipelago"]).await;
|
||||
});
|
||||
// PID1-owned timer transient: submit NOW (synchronously, while we are
|
||||
// definitely alive), fire in 2s from systemd itself. The old approach —
|
||||
// tokio sleep + `systemd-run --wait -- systemctl --no-block restart` —
|
||||
// ran as a child of the process being stopped; on v1.7.114->115 the
|
||||
// stop landed but the start never fired and the node sat dead all
|
||||
// night. A timer unit owned by PID1 cannot be killed by our own death,
|
||||
// and Restart=always on the unit is the second net.
|
||||
let submitted = tokio::process::Command::new("sudo")
|
||||
.args([
|
||||
"systemd-run",
|
||||
"--collect",
|
||||
"--on-active=2",
|
||||
"--timer-property=AccuracySec=100ms",
|
||||
"--",
|
||||
"systemctl",
|
||||
"restart",
|
||||
"archipelago",
|
||||
])
|
||||
.status()
|
||||
.await;
|
||||
match submitted {
|
||||
Ok(st) if st.success() => {}
|
||||
other => {
|
||||
tracing::warn!(
|
||||
"detached restart submission failed ({other:?}) — falling back to in-process restart"
|
||||
);
|
||||
tokio::spawn(async {
|
||||
tokio::time::sleep(std::time::Duration::from_secs(2)).await;
|
||||
let _ = host_sudo(&["systemctl", "--no-block", "restart", "archipelago"]).await;
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user