test(lifecycle): refuse to start the gate on a loaded host

A gate run on a box at load ~14 failed five times over, and every failure
read "could not create a container" — searxng:start, package.start
btcpay-server, 3x electrumx — never a lifecycle fault. That sent two
separate sessions hunting a phantom host-wide cgroup failure. It was
contention.

Measured on the 4-core node: at load ~14 podman runs 9-16 processes deep
and healthchecks time out 3-8/min; at load ~3.7, podman ~1 and zero
timeouts. Restoring four containers whose HealthTimeout equals their
HealthInterval, unchanged, made the box *better* once load fell — so the
config is a latent hazard, not the cause here.

Preflight now checks, once, before iteration 1:
  - aardvark-dns is singular (duplicates desync name resolution)
  - load1 is under nproc+1, waiting up to 15 min for a spike to pass
  - podman can actually create a container, 3/3

Deliberately NOT checked: the count of "Failed to create container" in
the journal. Those lines come from healthcheck exec churn and post-boot
settling, never reach 0 on a busy node, and gating on them would block
the gate forever. The probe proves creation positively instead.

Escape hatches: ARCHY_PREFLIGHT=0, ARCHY_MAX_LOAD, ARCHY_PREFLIGHT_SECS.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
archipelago
2026-08-08 20:01:35 -04:00
co-authored by Claude Opus 5
parent cfa6c6cb0d
commit b05222a342
+85
View File
@@ -11,6 +11,10 @@
# ARCHY_GATE_CASCADE=1 after the 5× loop, run ONE cascade pass
# (uninstall→no-ghost→reinstall a throwaway
# app); requires ARCHY_ALLOW_DESTRUCTIVE=1
# ARCHY_PREFLIGHT=0 skip the host-readiness preflight
# ARCHY_MAX_LOAD load1 ceiling (default: nproc + 1)
# ARCHY_PREFLIGHT_SECS how long to wait for load to fall
# (default: 900)
# plus everything run.sh / lib/rpc.bash respects
# (ARCHY_PASSWORD, ARCHY_HOST, ARCHY_SCHEME, ARCHY_ALLOW_DESTRUCTIVE,
# ARCHY_ALLOW_CASCADE_DESTRUCTIVE, ARCHY_ALLOW_NOAUTH)
@@ -77,6 +81,87 @@ settle_stack() {
echo " (stack settle deadline reached — proceeding anyway)"
}
# Host readiness, checked ONCE before iteration 1.
#
# Why this exists: on 2026-08-08 a gate run on a box at load ~14 failed five
# times over, every failure reading "could not create a container" (searxng:start,
# package.start btcpay-server, 3× electrumx) and never a lifecycle fault. That
# sent two separate sessions hunting a phantom host-wide cgroup failure. It was
# load. Measured on that 4-core box: at load ~14 podman runs 9-16 processes deep
# and healthchecks time out 3-8/min; at load ~3.7, podman ~1 and zero timeouts.
#
# So: refuse to start on a loaded host rather than emit misleading failures.
# Note what is deliberately NOT checked — the count of "Failed to create
# container" in the journal. Those lines are emitted by healthcheck exec churn
# and by settling after a boot; they never reach 0 on a busy node, and gating on
# them blocks the gate forever. Prove container creation POSITIVELY instead.
preflight_host() {
[[ "${ARCHY_PREFLIGHT:-1}" == "1" ]] || return 0
command -v podman >/dev/null 2>&1 || return 0 # remote/off-node run
local cores max_load
cores=$(nproc 2>/dev/null || echo 4)
max_load="${ARCHY_MAX_LOAD:-$((cores + 1))}"
echo "── preflight: host readiness ──"
# 1. aardvark-dns must be singular. Two of them serve divergent state and make
# container-name resolution flaky, which then looks like a lifecycle bug.
local dns
dns=$(pgrep -c aardvark-dns 2>/dev/null || echo 0)
if (( dns > 1 )); then
echo " FAIL: $dns aardvark-dns processes running (expected 1)." >&2
echo " Duplicate DNS servers desync container-name resolution." >&2
return 1
fi
echo " aardvark-dns: $dns"
# 2. Wait for load to fall. It oscillates on nodes doing IBD or media
# indexing, so a brief spike is not fatal — a sustained one is.
local deadline=$(( $(date +%s) + ${ARCHY_PREFLIGHT_SECS:-900} ))
local load1
while :; do
load1=$(awk '{print $1}' /proc/loadavg)
awk -v l="$load1" -v m="$max_load" 'BEGIN { exit !(l < m) }' && break
if (( $(date +%s) >= deadline )); then
echo " FAIL: load1 $load1 still above $max_load after ${ARCHY_PREFLIGHT_SECS:-900}s." >&2
echo " Quiesce the node (bitcoind IBD, electrumx indexing, CI runners)" >&2
echo " or override with ARCHY_MAX_LOAD=. Running now yields failures" >&2
echo " that look like lifecycle bugs but are contention." >&2
return 1
fi
echo " load1 $load1 > $max_load — waiting…"
sleep 20
done
echo " load1: $load1 (ceiling $max_load, $cores cores)"
# 3. Prove the host can actually create a container, 3× — the positive test
# that the journal grep only ever approximated.
local img
img=$(podman images --format '{{.Repository}}:{{.Tag}}' 2>/dev/null \
| grep -v '<none>' | head -1)
if [[ -z "$img" ]]; then
echo " (no local image — skipping container-create probe)"
else
local n
for n in 1 2 3; do
if ! timeout 60 podman run --rm "$img" /bin/true >/dev/null 2>&1; then
echo " FAIL: container-create probe $n/3 failed using $img." >&2
echo " The host genuinely cannot create containers; fix that first." >&2
return 1
fi
done
echo " container-create probe: 3/3 via $img"
fi
echo "── preflight: OK ──"
}
if ! preflight_host; then
echo "Preflight failed — refusing to start the gate. (ARCHY_PREFLIGHT=0 to skip.)" >&2
exit 3
fi
# One initial teardown so a previous run's cookies don't poison iteration 1.
./setup-teardown.sh