test(lifecycle): refuse to start the gate on a loaded host
A gate run on a box at load ~14 failed five times over, and every failure read "could not create a container" — searxng:start, package.start btcpay-server, 3x electrumx — never a lifecycle fault. That sent two separate sessions hunting a phantom host-wide cgroup failure. It was contention. Measured on the 4-core node: at load ~14 podman runs 9-16 processes deep and healthchecks time out 3-8/min; at load ~3.7, podman ~1 and zero timeouts. Restoring four containers whose HealthTimeout equals their HealthInterval, unchanged, made the box *better* once load fell — so the config is a latent hazard, not the cause here. Preflight now checks, once, before iteration 1: - aardvark-dns is singular (duplicates desync name resolution) - load1 is under nproc+1, waiting up to 15 min for a spike to pass - podman can actually create a container, 3/3 Deliberately NOT checked: the count of "Failed to create container" in the journal. Those lines come from healthcheck exec churn and post-boot settling, never reach 0 on a busy node, and gating on them would block the gate forever. The probe proves creation positively instead. Escape hatches: ARCHY_PREFLIGHT=0, ARCHY_MAX_LOAD, ARCHY_PREFLIGHT_SECS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
cfa6c6cb0d
commit
b05222a342
@@ -11,6 +11,10 @@
|
||||
# ARCHY_GATE_CASCADE=1 after the 5× loop, run ONE cascade pass
|
||||
# (uninstall→no-ghost→reinstall a throwaway
|
||||
# app); requires ARCHY_ALLOW_DESTRUCTIVE=1
|
||||
# ARCHY_PREFLIGHT=0 skip the host-readiness preflight
|
||||
# ARCHY_MAX_LOAD load1 ceiling (default: nproc + 1)
|
||||
# ARCHY_PREFLIGHT_SECS how long to wait for load to fall
|
||||
# (default: 900)
|
||||
# plus everything run.sh / lib/rpc.bash respects
|
||||
# (ARCHY_PASSWORD, ARCHY_HOST, ARCHY_SCHEME, ARCHY_ALLOW_DESTRUCTIVE,
|
||||
# ARCHY_ALLOW_CASCADE_DESTRUCTIVE, ARCHY_ALLOW_NOAUTH)
|
||||
@@ -77,6 +81,87 @@ settle_stack() {
|
||||
echo " (stack settle deadline reached — proceeding anyway)"
|
||||
}
|
||||
|
||||
# Host readiness, checked ONCE before iteration 1.
|
||||
#
|
||||
# Why this exists: on 2026-08-08 a gate run on a box at load ~14 failed five
|
||||
# times over, every failure reading "could not create a container" (searxng:start,
|
||||
# package.start btcpay-server, 3× electrumx) and never a lifecycle fault. That
|
||||
# sent two separate sessions hunting a phantom host-wide cgroup failure. It was
|
||||
# load. Measured on that 4-core box: at load ~14 podman runs 9-16 processes deep
|
||||
# and healthchecks time out 3-8/min; at load ~3.7, podman ~1 and zero timeouts.
|
||||
#
|
||||
# So: refuse to start on a loaded host rather than emit misleading failures.
|
||||
# Note what is deliberately NOT checked — the count of "Failed to create
|
||||
# container" in the journal. Those lines are emitted by healthcheck exec churn
|
||||
# and by settling after a boot; they never reach 0 on a busy node, and gating on
|
||||
# them blocks the gate forever. Prove container creation POSITIVELY instead.
|
||||
preflight_host() {
|
||||
[[ "${ARCHY_PREFLIGHT:-1}" == "1" ]] || return 0
|
||||
command -v podman >/dev/null 2>&1 || return 0 # remote/off-node run
|
||||
|
||||
local cores max_load
|
||||
cores=$(nproc 2>/dev/null || echo 4)
|
||||
max_load="${ARCHY_MAX_LOAD:-$((cores + 1))}"
|
||||
|
||||
echo "── preflight: host readiness ──"
|
||||
|
||||
# 1. aardvark-dns must be singular. Two of them serve divergent state and make
|
||||
# container-name resolution flaky, which then looks like a lifecycle bug.
|
||||
local dns
|
||||
dns=$(pgrep -c aardvark-dns 2>/dev/null || echo 0)
|
||||
if (( dns > 1 )); then
|
||||
echo " FAIL: $dns aardvark-dns processes running (expected 1)." >&2
|
||||
echo " Duplicate DNS servers desync container-name resolution." >&2
|
||||
return 1
|
||||
fi
|
||||
echo " aardvark-dns: $dns"
|
||||
|
||||
# 2. Wait for load to fall. It oscillates on nodes doing IBD or media
|
||||
# indexing, so a brief spike is not fatal — a sustained one is.
|
||||
local deadline=$(( $(date +%s) + ${ARCHY_PREFLIGHT_SECS:-900} ))
|
||||
local load1
|
||||
while :; do
|
||||
load1=$(awk '{print $1}' /proc/loadavg)
|
||||
awk -v l="$load1" -v m="$max_load" 'BEGIN { exit !(l < m) }' && break
|
||||
if (( $(date +%s) >= deadline )); then
|
||||
echo " FAIL: load1 $load1 still above $max_load after ${ARCHY_PREFLIGHT_SECS:-900}s." >&2
|
||||
echo " Quiesce the node (bitcoind IBD, electrumx indexing, CI runners)" >&2
|
||||
echo " or override with ARCHY_MAX_LOAD=. Running now yields failures" >&2
|
||||
echo " that look like lifecycle bugs but are contention." >&2
|
||||
return 1
|
||||
fi
|
||||
echo " load1 $load1 > $max_load — waiting…"
|
||||
sleep 20
|
||||
done
|
||||
echo " load1: $load1 (ceiling $max_load, $cores cores)"
|
||||
|
||||
# 3. Prove the host can actually create a container, 3× — the positive test
|
||||
# that the journal grep only ever approximated.
|
||||
local img
|
||||
img=$(podman images --format '{{.Repository}}:{{.Tag}}' 2>/dev/null \
|
||||
| grep -v '<none>' | head -1)
|
||||
if [[ -z "$img" ]]; then
|
||||
echo " (no local image — skipping container-create probe)"
|
||||
else
|
||||
local n
|
||||
for n in 1 2 3; do
|
||||
if ! timeout 60 podman run --rm "$img" /bin/true >/dev/null 2>&1; then
|
||||
echo " FAIL: container-create probe $n/3 failed using $img." >&2
|
||||
echo " The host genuinely cannot create containers; fix that first." >&2
|
||||
return 1
|
||||
fi
|
||||
done
|
||||
echo " container-create probe: 3/3 via $img"
|
||||
fi
|
||||
|
||||
echo "── preflight: OK ──"
|
||||
}
|
||||
|
||||
if ! preflight_host; then
|
||||
echo "Preflight failed — refusing to start the gate. (ARCHY_PREFLIGHT=0 to skip.)" >&2
|
||||
exit 3
|
||||
fi
|
||||
|
||||
# One initial teardown so a previous run's cookies don't poison iteration 1.
|
||||
./setup-teardown.sh
|
||||
|
||||
|
||||
Reference in New Issue
Block a user