Compare commits

...
Author SHA1 Message Date
archipelago 51a5473e22 docs(release): v1.8.5-alpha changelog section + What's New sync
Demo images / Build & push demo images (push) Failing after 42s
Curated release notes for the pending v1.8.5-alpha: Cuprate (with the
two review catches), kdump/rasdaemon + the host-fixup OTA channel, the
uninstall-abort fix, federation inline-picture routing, honest disk
usage, the three lying-screens fixes (#143/#127/#129), durable mesh
notifications + router recovery (#57/#103), and upstream-release tracking
with the first-sweep safe bumps.

What's New modal synced via scripts/sync-whats-new.py (--check passes;
89 versions, all present). Per docs/RELEASE_NOTES_BACKLOG.md the
v1.7.44-alpha -> current section audit remains the open item before the
tag.
2026-08-31 07:23:53 -04:00
archipelago 1872fc20ee feat(image): bake kdump + rasdaemon into fresh installs (#144)
The ISO's Dockerfile.rootfs gains kdump-tools/kexec-tools/rasdaemon with
USE_KDUMP=1, dumps to /var/crash and a compressed core collector, the
hang/panic sysctl drop-in, and rasdaemon + kdump-tools enabled — and the
installed target's GRUB cmdline gains crashkernel=256M next to the
existing quiet/splash line.

Source of truth note: the edit lands in
image-recipe/_archived/build-auto-installer-iso.sh — the builder that
generates the (git-ignored) image-recipe/build/auto-installer/ workspace,
which a cache-hit can reuse. The workspace copy was updated to match so
even a cached build ships the same state. Host fixups (previous commit)
converge already-deployed nodes to exactly this end state, so fresh and
old installs agree.

bash -n clean on the builder.
2026-08-31 07:23:53 -04:00
archipelago cbd463e980 feat(host): crash/hardware-error capture, delivered by a new host-fixup OTA channel (#144)
kdump + rasdaemon on every node, per docs/kdump-rasdaemon-design.md with
the approved decisions: hang capture ON (a wedged kiosk dumps and reboots
itself instead of sitting dead), crashkernel=256M, backfill ships with
this release, phase-2 UI surfacing deferred.

Host fixups (docs/system-level-ota-design.md) are the general answer to
'deliver system-level updates OTA': curated OS packages, sysctl drop-ins,
service enablement and the GRUB crashkernel line, carried by the signed
binary and applied idempotently at startup — non-fatal by construction
(offline/locked-dpkg nodes converge on a later boot), skipped on dev
boxes and non-Debian hosts. This formalizes the polkit/audio repair
precedents into a channel with a stated policy: pinned packages and
parameter intent only, never dist-upgrade automation; the ISO bakes the
identical end state into fresh installs (next commit).

The one runtime limitation is honest: crashkernel memory can only be
reserved at boot, so the fixup writes GRUB, runs update-grub, and logs
that it takes effect on the next reboot.

tests/lifecycle/os-audit.sh gains section D — a graded baseline check:
FAIL if capture never landed, WARN if written but awaiting reboot, PASS
when reserved, policy live and rasdaemon recording. Section D runs
independently of RPC health: a wedged backend must not mask that the
node also stopped capturing evidence.

Verification: host_fixups unit tests 4/4; cargo fmt clean; full suite
runs in the release gate (create-release) and the archi-dev-box
lifecycle gate before the tag.
2026-08-31 07:23:44 -04:00
archipelago 9df580bf2b docs: peering trust terminology — names for the four concepts (#134)
Gives stable names to what issue #134 showed gets conflated: Trusted peer
(invite-verified, operator decision), Discovered peer (learned from a
Trusted peer's advertisement, hard-capped at Observer — TRUST IS NOT
TRANSITIVE), Routing hint (what a Discovered peer actually contributes:
reachability, not trust), and Peer advertisement (the mechanism itself,
a feature not a leak).

Records the two rules that make the model sound (trust requires a
traceable operator decision; discovery is transitive, trust is not), why
advertisement exists (one invite makes a node reachable to the trusted
set without granting anything), and the deferred open questions: the
'don't advertise my peers' privacy toggle and UI tier vocabulary.
2026-08-31 07:23:44 -04:00
archipelago aee7ecaac1 docs: index the kdump/rasdaemon design 2026-08-31 06:11:15 -04:00
archipelago e51ceaa250 docs: draft kdump + rasdaemon troubleshooting design (#144)
Design for capturing post-mortem and hardware-error evidence on fleet
nodes: kdump (crashkernel=256M, dump to /var/crash on the unencrypted
root — never the LUKS data partition, so the crash kernel never handles
key material; makedumpfile-compressed, keep-2 retention) and rasdaemon
(EDAC/ECC events into sqlite on the same root).

Deliberately phased: phase 1 = capture on the image + bootstrap backfill
for existing nodes (kernel cmdline can't travel by OTA; takes effect on
next reboot); phase 2 = a read-only system.diagnostics surface in the
UI, only after a fleet node has produced a real dump.

Four decisions flagged in the doc: hang-capture on/off (recommended ON
— a wedged kiosk is useless anyway, and this turns every freeze into
evidence + self-reboot), crashkernel size, backfill timing, and phase-2
scope. Implementation touchpoints listed (Dockerfile.rootfs,
auto-install.sh:1810 cmdline, kdump-tools config, bootstrap, lifecycle
gate assertions).
2026-08-31 06:10:55 -04:00
archipelago 7c9559aa57 chore(catalog): sign the catalog — Cuprate ships, safe pin bumps land
Signed by the release root (ceremony verify passed locally before push).
Contents of this catalog over the previous one:

  NEW   cuprate           0.1.0-preview-18-g618ff14 — alternative Monero
                        node (Rust); image verified present in the mirror
                        registry; manifest embedded; store entry curated
                        (money / optional)
  BUMP  strfry            1.1.1 -> 1.1.2
  BUMP  btcpay-server     2.4.2 -> 2.4.3
  BUMP  netbird (nginx)   1.31.3-alpine -> 1.31.4-alpine
  BUMP  pine   (nginx)    1.31.3-alpine -> 1.31.4-alpine

All bump targets verified pullable from their public registries before
editing. The three mirror-backed bumps (vaultwarden 1.37.2-alpine,
archy-nbxplorer 2.6.11, home-assistant 2026.8.3) remain parked on
app-bumps-mirror-pending until a live registry-push token exists for the
lfg2025 namespace.

Drift gate clean: check-app-catalog-drift.py --release --strict
(31 store entries, 0 drift, 0 missing). 69 catalog entries total.

Nodes pick this up on their next hourly catalog refresh (or at startup)
— signature verified against the release-root key before application.
2026-08-31 05:56:00 -04:00
11 changed files with 845 additions and 15 deletions
+18
View File
@@ -1,5 +1,23 @@
# Changelog
## v1.8.5-alpha (2026-08-30)
- **Cuprate — an independent Monero node — is now an app.** Monero consensus validated by a second, unrelated codebase (Rust), the same layer of security-in-depth Bitcoin gets from Knots. Review caught two problems before anything shipped: the unrestricted RPC that can move funds stayed bound to the container's loopback (never published to the node, let alone the LAN — anything on the node could previously have reached it), and its restricted RPC moved off port 18089 to avoid colliding with Penpot. Honest caveat: upstream has cut no stable release yet, so the pin tracks an exact preview build (0.1.0-preview-18-g618ff14) and moves to their first tagged release when there is one.
- **A frozen node now explains itself — and comes back on its own.** The host now captures a memory dump into /var/crash when the kernel panics *or* wedges (a hung kiosk used to sit dead until someone power-cycled it; now it dumps, reboots itself, and leaves the evidence behind), and records failing-memory signals (ECC errors) into a database as they happen. This is the first change delivered by a new host-update channel: the node's own updater now carries OS-level packages and settings to already-deployed machines — the crash-kernel's memory reservation is the one part that waits for a reboot, and the node says so rather than pretending.
- **Uninstalling an app can no longer report success when it failed.** The declarative path used to swallow every teardown error and report the app uninstalled, leaving the tile behind and the truth in the logs. A failed uninstall now stops and shows the real per-app errors, so "still there" is never presented as "gone".
- **Pictures to internet-only mesh contacts work now.** Sending an attachment inline always took the radio path and failed with "Peer is federation-only (no radio twin)" for contacts reachable only over the internet — and the size-adviser kept recommending a radio transfer those peers can't receive. Both fixed: inline sends route over the federation when that's the only way to reach the peer, and the advice no longer offers radio-only transfers to radio-unreachable contacts.
- **Disk cleanup finally has honest numbers.** Space "free" on a drive was counted including the slice the filesystem keeps reserved for root — roughly 5% of the disk, 92 GB on one dev box — so the automatic cleanup that's supposed to kick in at 90% never triggered and stale container images piled up unnoticed. Reserved space now counts as used, which is what the threshold was always meant to measure.
- **Three small screens that were lying to you, fixed.** The "Bitcoin is synced — fund your wallet" toast no longer appears on a node where the wallet it means (LND) isn't installed — it points at installing LND instead. The seed-reveal screen hides its third prompt unless the password actually fails to decrypt (the backup passphrase only exists if you set one). And multi-version store cards stop quoting a version number you'll be asked to choose on the next screen anyway.
- **Mesh notifications survive a refresh, and a stale router no longer hides the fix.** Radio message unread counts are now remembered per contact instead of guessed from session state (the "one new message showed 11 unread" bug), cover Meshtastic, MeshCore and Reticulum alike, and deep-link to the right conversation; a single new message announces itself once. Separately, when the cached router address goes stale, the error card gains a "Reconfigure router" action instead of a Retry loop that can never succeed.
- **The app updater now knows what upstream shipped.** Every app's manifest records where it comes from — including the odd corners (GitLab-only projects, ghcr-only images) — and a checker sweeps all of them against upstream releases, so a pin that quietly rots for months is now visible instead of invisible. The first full sweep found 27 pins behind; the safe patch-level ones shipped with this release (strfry, BTCPay Server 2.4.3, the two nginx frontends), and the major jumps that may carry data migrations are deliberately held for their own careful passes.
## v1.8.4-alpha (2026-08-20)
- **Apps with their own login can now skip the node's login screen — Gitea and BTCPay Server do so out of the box.** Some apps bring a complete account system of their own, and putting the node's password page in front of them broke real workflows: git clients can't answer a browser login, and a BTCPay checkout link handed to a customer must open for that customer. These apps are now served directly on their own login, while the node still fronts the connection for everything else it does (embedding fixes, the "app is restarting" page, Tor). Every app gets a new **Settings → app → Access control** switch, so you can put the node login back in front of any app — or take it away from one — with one click, effective immediately. App developers declare the default in their manifest (`auth: open`), documented in the developer guide.
+344
View File
@@ -0,0 +1,344 @@
//! Host-level fixups: OS packages, kernel parameters and system services the
//! node needs, delivered by the same signed-binary OTA that ships everything
//! else (docs/system-level-ota-design.md).
//!
//! Scope and posture — read before adding anything here:
//!
//! * **Idempotent + non-fatal.** Every step is a no-op when the host already
//! has the desired state, and a failure (offline box, locked dpkg, missing
//! package in the release's Debian suite) logs a warning and moves on. A
//! host fixup must never be able to stop the node from starting.
//! * **Curated, pinned intent — not dist-upgrade automation.** We deliver the
//! specific packages and settings a release deliberately adds (crash
//! capture, hardware-error logging, later: unattended-upgrades posture, host
//! firewall). Regular Debian upgrades stay with the operator; this channel
//! never silently swaps a kernel or a libc.
//! * **Fresh installs converge too.** The ISO bakes the same end state in
//! (Dockerfile.rootfs, auto-install.sh cmdline), so the fixup is a no-op on
//! new machines and only does real work on already-deployed nodes.
//! * **Kernel cmdline can't move at runtime.** `crashkernel=` reserves memory
//! at boot; the fixup writes GRUB and update-grub so the change lands on the
//! next reboot, and says so in the log. Everything else (packages, sysctls,
//! services) applies immediately.
//!
//! First payload (#144, docs/kdump-rasdaemon-design.md): kdump + rasdaemon —
//! post-mortem and hardware-error capture:
//! * kdump-tools/kexec-tools/rasdaemon installed
//! * /etc/default/kdump-tools: USE_KDUMP=1, dumps to /var/crash, compressed
//! core collector
//! * /etc/sysctl.d/99-archipelago-kdump.conf: a wedged node dumps and
//! reboots rather than sitting dead until power-cycled
//! * crashkernel=256M appended to the installed GRUB cmdline (next reboot)
//! * /var/crash pruned to the two newest dumps
//!
//! The module is skipped on dev boxes (same guard bootstrap::run uses) and on
//! hosts without dpkg.
use anyhow::{Context, Result};
use tracing::{debug, info, warn};
use crate::update::host_sudo;
/// Packages the node's host must have. Keep this list short and justified —
/// every entry is state we now own on the fleet's OS images.
const HOST_PACKAGES: &[&str] = &["kdump-tools", "kexec-tools", "rasdaemon"];
/// Crash-kernel reservation. 256M covers the capture kernel plus makedumpfile
/// on the fleet's 16–64GB amd64 machines (~1–2% of RAM, permanently reserved).
/// The arm image (RPi) is out of scope for phase 1 — see the design doc.
const CRASHKERNEL_PARAM: &str = "crashkernel=256M";
const KDUMP_SYSDROPIN_PATH: &str = "/etc/sysctl.d/99-archipelago-kdump.conf";
const KDUMP_SYSDROPIN: &str = "\
# Archipelago kdump policy (#144). A wedged kiosk is useless until someone
# power-cycles it — capture the evidence, then reboot by itself. Dumps land in
# /var/crash (see docs/kdump-rasdaemon-design.md); keep-2 pruning is done by
# the host fixup pass, not a timer.
kernel.panic = 10
kernel.panic_on_oops = 1
kernel.hung_task_panic = 1
kernel.hardlockup_panic = 1
";
/// How many dumps to keep in /var/crash. Two ≈ 4 GiB worst case on the 30 GiB
/// unencrypted root — the partition usage itself is tracked by disk_monitor.
const KEEP_DUMPS: usize = 2;
/// Entry point, spawned from main.rs at startup like the other ensure_* heals.
pub async fn ensure_host_fixups() {
// Dev-box guard (same rationale as bootstrap::run): on contributor
// machines /home/archipelago/archy is a symlink into a git checkout and
// the host is the contributor's own OS — never touch it.
let home_archy = std::path::Path::new("/home/archipelago/archy");
if tokio::fs::symlink_metadata(home_archy)
.await
.map(|m| m.file_type().is_symlink())
.unwrap_or(false)
{
debug!("/home/archipelago/archy is a symlink — skipping host fixups (dev box)");
return;
}
// Non-Debian hosts: nothing we manage here applies.
if tokio::fs::symlink_metadata("/usr/bin/dpkg").await.is_err() {
debug!("no dpkg on this host — skipping host fixups");
return;
}
if let Err(e) = run_host_fixups().await {
warn!("host fixups failed (non-fatal): {:#}", e);
}
}
async fn run_host_fixups() -> Result<()> {
// 1. Packages — install only what's missing; a locked/offline apt must
// never block anything downstream (steps below degrade to no-ops).
match ensure_packages().await {
Ok(true) => info!("host fixups: installed missing packages"),
Ok(false) => debug!("host fixups: all packages present"),
Err(e) => warn!("host fixups: package install failed (non-fatal): {:#}", e),
}
// 2. kdump config + sysctl drop-in + GRUB cmdline + services. One helper
// per concern so a failure in one logs and leaves the others running.
if let Err(e) = ensure_kdump_sysdropin().await {
warn!(
"host fixups: kdump sysctl drop-in failed (non-fatal): {:#}",
e
);
}
if let Err(e) = ensure_kdump_defaults().await {
warn!(
"host fixups: kdump-tools config failed (non-fatal): {:#}",
e
);
}
match ensure_crashkernel_cmdline().await? {
true => {
warn!("host fixups: crashkernel= written to GRUB — takes effect on the NEXT reboot")
}
false => debug!("host fixups: crashkernel already in GRUB cmdline"),
}
if let Err(e) = ensure_rasdaemon_enabled().await {
warn!("host fixups: rasdaemon enable failed (non-fatal): {:#}", e);
}
if let Err(e) = prune_crash_dumps().await {
debug!("host fixups: /var/crash prune skipped: {:#}", e);
}
Ok(())
}
/// True if any package was installed. Mirrors the polkit repair's apt posture:
/// install without `apt-get update` first; only if that fails (fresh suite,
/// stale index), update once and retry. Both under timeout, both non-fatal.
async fn ensure_packages() -> Result<bool> {
let wanted = HOST_PACKAGES
.iter()
.map(|p| format!("'{p}'"))
.collect::<Vec<_>>()
.join(" ");
let script = format!(
r#"
set -u
WANTED="{wanted}"
MISSING=""
for p in $WANTED; do
dpkg-query -W -f='${{Status}}' "$p" 2>/dev/null | grep -q 'install ok installed' || MISSING="$MISSING $p"
done
[ -z "$MISSING" ] && exit 0
timeout 240 apt-get install -y --no-install-recommends $MISSING >/dev/null 2>&1 \
|| timeout 240 sh -c 'apt-get update >/dev/null 2>&1 && apt-get install -y --no-install-recommends $MISSING >/dev/null 2>&1' \
|| exit 3
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("install host packages")?;
match status.code() {
Some(0) => Ok(false),
Some(2) => Ok(true),
code => anyhow::bail!("host package install exited with {code:?}"),
}
}
/// Write the sysctl drop-in and apply it live (these four keys are all
/// runtime-settable, so the hang/panic policy takes effect without a reboot).
async fn ensure_kdump_sysdropin() -> Result<()> {
let script = format!(
r#"
set -u
PATH_FILE='{KDUMP_SYSDROPIN_PATH}'
CONTENT_FILE=/tmp/archy-kdump-sysctl.$$.tmp
cat > "$CONTENT_FILE" <<'SYSEOF'
{KDUMP_SYSDROPIN}SYSEOF
if [ -f "$PATH_FILE" ] && cmp -s "$CONTENT_FILE" "$PATH_FILE"; then
rm -f "$CONTENT_FILE"
exit 0
fi
mv "$CONTENT_FILE" "$PATH_FILE"
chmod 644 "$PATH_FILE"
sysctl --system >/dev/null 2>&1 || true
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("write kdump sysctl drop-in")?;
match status.code() {
Some(0) => Ok(()),
Some(2) => {
info!("host fixups: installed {KDUMP_SYSDROPIN_PATH} (hang/panic policy)");
Ok(())
}
code => anyhow::bail!("kdump sysctl drop-in exited with {code:?}"),
}
}
/// Point kdump-tools at /var/crash with a compressed core collector. Works on
/// the package's shipped defaults file (USE_KDUMP=0, commented KDUMP_COREDIR)
/// and on any state we already wrote — pure line surgery, idempotent.
async fn ensure_kdump_defaults() -> Result<()> {
let script = r#"
set -u
CONF=/etc/default/kdump-tools
[ -f "$CONF" ] || exit 3
CHANGED=0
set_kv() {
# set_kv KEY VALUE — replace any (possibly commented) KEY= line with
# KEY='VALUE', appending at the end when absent.
KEY="$1"; VAL="$2"
if grep -qE "^${KEY}=" "$CONF" 2>/dev/null; then
if ! grep -qE "^${KEY}='?${VAL}'?$" "$CONF"; then
sed -i "s|^${KEY}=.*|${KEY}=\"${VAL}\"|" "$CONF"
CHANGED=1
fi
else
printf '\n%s="%s"\n' "$KEY" "$VAL" >> "$CONF"
CHANGED=1
fi
}
set_kv USE_KDUMP 1
set_kv KDUMP_COREDIR /var/crash
set_kv CORE_COLLECTOR 'makedumpfile -l --message-level 1 -d 31'
[ "$CHANGED" -eq 1 ] || exit 0
systemctl enable kdump-tools >/dev/null 2>&1 || true
exit 2
"#;
let status = host_sudo(&["sh", "-lc", script])
.await
.context("configure kdump-tools")?;
match status.code() {
Some(0) => Ok(()),
Some(2) => {
info!("host fixups: kdump-tools configured (USE_KDUMP=1, /var/crash)");
Ok(())
}
code => anyhow::bail!("kdump-tools config exited with {code:?}"),
}
}
/// Append `crashkernel=` to the installed GRUB cmdline and run update-grub.
/// The reservation itself only exists after the next reboot — memory cannot
/// be set aside at runtime — so the caller must log the reboot caveat.
/// Returns true if the cmdline changed.
async fn ensure_crashkernel_cmdline() -> Result<bool> {
let script = format!(
r#"
set -u
GRUB=/etc/default/grub
PARAM='{CRASHKERNEL_PARAM}'
[ -f "$GRUB" ] || exit 3
LINE=$(grep -E '^GRUB_CMDLINE_LINUX_DEFAULT=' "$GRUB" | head -1)
[ -n "$LINE" ] || exit 3
case "$LINE" in
*"$PARAM"*) exit 0 ;;
esac
NEWLINE=$(printf '%s' "$LINE" | sed "s/\"$/ $PARAM\"/")
sed -i "s|^GRUB_CMDLINE_LINUX_DEFAULT=.*|$NEWLINE|" "$GRUB"
timeout 120 update-grub >/dev/null 2>&1 || true
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("set crashkernel= in GRUB")?;
match status.code() {
Some(0) => Ok(false),
Some(2) => Ok(true),
code => anyhow::bail!("crashkernel cmdline fixup exited with {code:?}"),
}
}
async fn ensure_rasdaemon_enabled() -> Result<()> {
let status = host_sudo(&["systemctl", "enable", "--now", "rasdaemon"])
.await
.context("enable rasdaemon")?;
if status.success() {
Ok(())
} else {
anyhow::bail!("systemctl enable --now rasdaemon exited with {status}")
}
}
/// Keep only the newest [`KEEP_DUMPS`] dumps in /var/crash. Called on every
/// fixup pass rather than by a timer: the pass runs at every startup, which is
/// exactly the cadence at which new dumps appear (a dump ends in a reboot).
async fn prune_crash_dumps() -> Result<()> {
let script = format!(
r#"
set -u
DIR=/var/crash
[ -d "$DIR" ] || exit 0
KEEP={KEEP_DUMPS}
COUNT=$(ls -1 "$DIR" 2>/dev/null | wc -l)
[ "$COUNT" -gt "$KEEP" ] || exit 0
ls -1dt "$DIR"/* 2>/dev/null | tail -n +"$((KEEP + 1))" | while IFS= read -r victim; do
rm -rf -- "$victim"
done
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("prune /var/crash")?;
match status.code() {
Some(0) => Ok(()),
Some(2) => {
info!("host fixups: pruned old dumps in /var/crash (keep {KEEP_DUMPS})");
Ok(())
}
code => anyhow::bail!("/var/crash prune exited with {code:?}"),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn sysctl_dropin_carries_the_full_hang_capture_policy() {
for key in [
"kernel.panic = 10",
"kernel.panic_on_oops = 1",
"kernel.hung_task_panic = 1",
"kernel.hardlockup_panic = 1",
] {
assert!(KDUMP_SYSDROPIN.contains(key), "drop-in missing {key}");
}
}
#[test]
fn package_list_is_exactly_the_kdump_rasdaemon_set() {
assert_eq!(HOST_PACKAGES, &["kdump-tools", "kexec-tools", "rasdaemon"]);
}
#[test]
fn crashkernel_param_is_sized_and_unprefixed() {
assert_eq!(CRASHKERNEL_PARAM, "crashkernel=256M");
}
#[test]
fn keep_dumps_is_two() {
assert_eq!(KEEP_DUMPS, 2);
}
}
+7
View File
@@ -55,6 +55,7 @@ mod entropy;
mod federation;
mod fips;
mod health_monitor;
mod host_fixups;
mod host_ip;
mod identity;
mod identity_manager;
@@ -435,6 +436,12 @@ async fn main() -> Result<()> {
// iframe on kiosk nodes (docs/tv-input-iframe-apps.md).
tokio::spawn(bootstrap::ensure_gamepad_keys());
// Host-level fixups (#144 + docs/system-level-ota-design.md): kdump +
// rasdaemon — crash/hardware-error capture delivered to already-deployed
// nodes over the signed binary OTA. Idempotent, non-fatal, background;
// the crashkernel= GRUB edit lands on the next reboot.
tokio::spawn(host_fixups::ensure_host_fixups());
// Mesh access: mirror IPv4-published app ports onto [::] so direct-port
// app URLs (http://[<fips0 ULA>]:<port>) work from the companion.
tokio::spawn(mesh_ports::run_mesh_port_mirror());
+3
View File
@@ -54,6 +54,9 @@ step-by-step guides, and some predate the current implementation.
- [Dual Ecash](dual-ecash-design.md)
- [Hardware Signer](hardware-signer-design.md)
- [Manifest Hooks](manifest-hooks-design.md)
- [Peering & Federation Trust](peering-trust-model.md) — naming/semantics of trust levels vs discovery (#134)
- [kdump + rasdaemon Troubleshooting](kdump-rasdaemon-design.md) — post-mortem and hardware-error capture on nodes (#144)
- [System-Level OTA](system-level-ota-design.md) — how host-level packages/config reach already-deployed nodes
- [Meshroller Integration](meshroller-integration-design.md)
- [Nostr Git Source Hosting](nostr-git-source-hosting.md)
- [Nostr Identity Import](nostr-identity-import-plan.md) · [Nostr Signer Login (research)](nostr-signer-login-research.md)
+131
View File
@@ -0,0 +1,131 @@
# kdump + rasdaemon — post-mortem and hardware-error capture (#144)
Status: IMPLEMENTED (phase 1) — decisions approved 2026-08-30: hang capture ON,
crashkernel=256M, ship the backfill with this release, phase-2 UI deferred.
Delivery: image-recipe (Dockerfile.rootfs, auto-install.sh cmdline) +
`core/archipelago/src/host_fixups.rs` (existing nodes, see
docs/system-level-ota-design.md) + `tests/lifecycle/os-audit.sh` section D.
Owner: node image (image-recipe) + lifecycle gate
Issue: #144 — "Configure kdump and rasdaemon for troubleshooting"
## The problem
When a fleet node hard-locks or a memory stick starts failing, today we get
nothing: a frozen kiosk is power-cycled and the evidence is gone; a DIMM
throwing correctable ECC errors for weeks is invisible until it starts
corrupting things. Two standard kernel mechanisms capture this evidence:
- **kdump** — reserves a small crash kernel at boot; on a kernel panic (or,
configured so, a hang) the running kernel hands the machine over to the
crash kernel, which writes a compressed dump of memory to disk and
reboots. The node comes back by itself *and* leaves a post-mortem.
- **rasdaemon** — a userspace daemon that records hardware error events
(correctable/uncorrectable ECC per DIMM, PCIe AER) from EDAC/sysfs into a
sqlite database: persistent evidence of degrading hardware with no crash
required.
## Facts the design rests on
- Installed-disk layout (auto-install.sh): BIOS boot 1MiB · EFI 512MiB ·
**root ext4 30GiB, unencrypted** · data (rest, LUKS).
- The data partition is LUKS and unlocked late by the node itself — the
crash kernel must never be asked to handle key material.
- The installed system's kernel command line is written by
auto-install.sh:1810 (`GRUB_CMDLINE_LINUX_DEFAULT="quiet splash …"`).
- Packages land via `Dockerfile.rootfs` (trixie) with `systemctl enable`
in the same RUN block (nginx/tor/avahi pattern).
- Kernel cmdline cannot be changed by OTA — it lives in GRUB. Existing
nodes need a backfill step (bootstrap) plus a deliberate reboot.
## Design
### kdump
- **Packages:** `kdump-tools kexec-tools` added to Dockerfile.rootfs.
- **Command line:** append `crashkernel=256M` to
`GRUB_CMDLINE_LINUX_DEFAULT` in auto-install.sh. 256M covers the capture
kernel plus makedumpfile on the fleet's 16–64GB amd64 machines (~1–2% of
RAM reserved, permanently). The arm image (RPi, config.txt boot) is out
of scope for phase 1.
- **Dump target:** `local filesystem /var/crash` — on the unencrypted 30GiB
root, deliberately *not* the encrypted data partition. No key handling
in the crash initramfs, no dependency on the node's own unlock logic.
- **Core collector:** `makedumpfile -l --message-level 1 -d 31`
(compressed, zero/free pages excluded) — a dump lands at roughly 5–15%
of RAM, i.e. ~1–2 GiB on a 16 GiB machine.
- **Retention:** keep the **2 newest** dumps only. A small systemd timer
(or kdump-tools' `KDUMP_POST_SCRIPT`) prunes older vmcores; a full root
partition is already caught by disk_monitor's usage tracking. Two dumps
≈ 4 GiB worst case on 30 GiB root — safe.
- **When to dump — the deliberate trade-off (decision needed):**
- Baseline: dump on real panics (`kernel.panic` path) — no behavioral
change to a wedged node.
- Recommended for this fleet: also enable hang capture
(`kernel.hung_task_panic=1`, hardlockup via NMI watchdog). A kiosk
that hard-locks is useless until power-cycled anyway; converting the
hang into "dump + automatic reboot" turns every freeze into evidence
*and* self-heals the node. Cost: a genuinely-busy-but-alive machine
that trips the watchdog reboots — the threshold is kernel-default
conservative (40s), so this should be rare.
### rasdaemon
- **Packages:** `rasdaemon`; `systemctl enable rasdaemon` in the
Dockerfile.rootfs enable block (same pattern as nginx).
- **Storage:** its default sqlite DB at
`/var/lib/rasdaemon/ras-mc_event.db` on the unencrypted root.
- **Human access today:** `ras-mc-ctl --summary` / `--errors` over SSH.
No UI in phase 1.
### Surfacing (phase 2 — separate follow-up, not in this cut)
A small read-only `system.diagnostics` surface: last-crash timestamp and
vmcore sizes from `/var/crash`, plus ECC error totals per DIMM from the
rasdaemon DB — shown in Settings → System. Deliberately deferred: capture
first, UI once there is something to show and a node in the fleet has
actually produced a dump.
### Existing nodes (phase 1.5 backfill)
The OTA cannot change the bootloader. Bootstrap (which already delivers
fixes to existing nodes) appends `crashkernel=256M` (and the chosen
panic/hang params) to `/etc/default/grub` on machines that don't have it,
and enables `rasdaemon` via the node's package install path. **Takes
effect on the next reboot** — the operator reboots nodes when applying the
release; no special ceremony needed beyond that.
## Testing
- Image: the new packages appear in the ISO; QEMU boot smoke
(build-iso-release.sh stage 5) still green.
- Lifecycle gate additions (bats, archi-dev-box first): `kdump-config show`
reports a loaded crash kernel reservation; `systemctl is-active
rasdaemon`; `/etc/default/grub` carries `crashkernel=`.
- Live drill (once, on archi-dev-box, not in the gate): trigger
`sysrq c` → vmcore appears in `/var/crash`, node reboots itself,
second boot is clean. Keep this manual — it reboots the box.
## Implementation touchpoints
1. `image-recipe/build/auto-installer/Dockerfile.rootfs` — packages +
`systemctl enable rasdaemon`.
2. `image-recipe/build/auto-installer/installer-iso/archipelago/auto-install.sh:1810`
— append `crashkernel=256M` (+ hang params if approved) to
`GRUB_CMDLINE_LINUX_DEFAULT`.
3. `kdump-tools` config: `/etc/default/kdump-tools` (dump target
`/var/crash`, core_collector line, `KDUMP_POST_SCRIPT` or timer for
retention).
4. Bootstrap backfill for existing nodes.
5. `tests/lifecycle` — presence assertions (crash kernel reserved,
rasdaemon active).
## Decisions needed before implementation
1. **Hang capture on or off?** Recommended ON (`hung_task_panic=1` +
NMI watchdog): every hard lockup becomes a dump + self-reboot. OFF
means dumps only on true panics; wedged nodes still need the button.
2. **crashkernel=256M vs 320M** — 256M is the common default for
16–64GB machines; 320M if we expect large io-heavy kernels.
3. **Backfill now or new-installs-only?** Recommended: ship the backfill
with the next release so the whole fleet gains capture on reboot.
4. Phase-2 UI surfacing scope — confirm "later" so phase 1 stays small.
+47
View File
@@ -0,0 +1,47 @@
# Peering & Federation Trust — naming and semantics
Status: TERMINOLOGY SET — records what the code does today (#134).
Deferred: the "don't advertise my peers" opt-out (see §Open questions).
The code is the authority; this doc gives names to the four concepts that
issue #134 showed get conflated in conversation. Where a name changed in
user-facing discussion, the term below is the one to use everywhere
(UI copy, docs, issues, reviews).
## The four concepts
| Term (use this) | What it is | Where it lives |
|---|---|---|
| **Trusted peer** | A node THIS operator invited and verified: bilateral DID challenge over an out-of-band invite code (`federation::sync`, ADR-007). The only level that grants full access. | `TrustLevel::Trusted`, set via `TrustSource::Invite` or `Manual` |
| **Discovered peer** | A peer we learned about from a Trusted peer's advertised list — the transitive merge. Never better than **Observer**: `TRUST IS NOT TRANSITIVE` (sync.rs guard). | `TrustLevel::Observer`, `TrustSource::TransitiveMerge` |
| **Routing hint** | What a Discovered peer actually contributes: an address that lets us route directly over FIPS without a second invite hop. Reachability, not trust. | Observer-level sync + FIPS endpoint records |
| **Peer advertisement** | The act of a Trusted peer sharing its own peer list during sync. This is the *mechanism* #134 observed — a feature, not a leak. | sync.rs merge path |
## The two rules that make it sound
1. **Trust requires an operator decision, always traceable.** Every trust
level carries a `TrustSource`. Only a minted invite (or an explicit
operator change) can produce `Trusted`; uninvited joins and transitive
merges are hard-capped at `Observer` — a peer can never expand our
trusted set on its own authority.
2. **Discovery is transitive; trust is not.** Seeing more nodes through a
Trusted peer is expected and useful (routing). Granting those nodes
anything is an operator action, never automatic.
## Why a Trusted peer advertising its list is by design
Without advertisement, every new node needs a direct invite from every node
that wants to reach it — the invite graph becomes the routing bottleneck
AdDR-007 set out to remove. With it, one invite makes a node *reachable* to
the trusted set (routing hints), while *authorization* still requires each
operator's own invite. Reachability ≠ access.
## Open questions (deferred, tracked in #134)
- **"Don't advertise my peers"** — an operator privacy toggle suppressing
peer advertisement during sync. Small code change, real design questions:
it hides peers who may WANT discovery, and it degrades the routing benefit
for every node trusting you. Needs a product decision, not just code.
- **Tier vocabulary in the UI** — whether to surface "Observer" as such or
a friendlier term ("Connected"/"Visible") — part of the TODO.md peering
trust-model item.
+82
View File
@@ -0,0 +1,82 @@
# System-Level OTA — host fixups
Status: Implemented (first payload shipped alongside this doc)
Owner: `core/archipelago/src/host_fixups.rs`
Related: docs/kdump-rasdaemon-design.md (first payload), CLAUDE.md invariants
## The problem
The binary OTA updates the node's own software, and the signed app catalog
updates apps. But the **host OS** — Debian packages, kernel parameters,
system services — previously moved only through ISO re-installs. A node
deployed a year ago can be running today's node software on a host that
never gained anything the image learned since. Issue #99 (missing polkit
rule on old nodes) and the audio-stack heal were each hand-carved
one-off bootstrap repairs; there was no general channel and no stated
policy for touching the host from the node.
## The mechanism
`host_fixups::ensure_host_fixups()` — spawned from `main.rs` at startup
alongside the other `ensure_*` heals, in the background, best-effort:
1. **Dev-box guard** — skip when `/home/archipelago/archy` is a symlink
(contributor checkout) and when there's no dpkg (non-Debian host).
2. **Packages** — install only what's missing, from a curated, in-code
list (`HOST_PACKAGES`), `apt-get install` first, one `apt-get update`
retry, both under timeout, never fatal (offline/locked-dpkg nodes
converge on a later boot).
3. **Configuration** — idempotent per-concern helpers writing root-owned
config (via the existing `host_sudo` path): sysctl drop-ins, service
defaults, GRUB cmdline, service enablement.
4. **Reporting** — every step logs what it did; failures log warnings and
move on. A host fixup must never be able to stop the node from starting.
### Why embedded-in-the-binary rather than fetched
Same reasoning as the tor-helper (`bootstrap.rs`): the signed binary OTA
is the only authenticated delivery channel every node already trusts and
pulls on schedule. Fixups compiled into the binary travel with a version,
are reviewable in git, and can't be served to a subset of the fleet.
## Policy — what may travel this channel
| May | May not |
|---|---|
| Specific, pinned packages the node needs (kdump-tools, rasdaemon, …) | `dist-upgrade` or silent kernel/libc swaps — regular Debian upgrades stay with the operator |
| Kernel *parameters* via GRUB/sysctl — with the next-reboot caveat logged loudly | Anything requiring a secret, or touching LUKS key material |
| Service enablement + config the image also bakes in | Divergence: the ISO must converge to the SAME end state so fresh installs are a no-op |
| Small, reviewable, per-concern Rust functions with tests | Shell-script-of-things payloads beyond a single concern |
The rule: **the ISO and the fixup must express the same intent twice,
in reviewable places** — Dockerfile.rootfs/auto-install.sh for fresh
installs, `host_fixups.rs` for the deployed fleet. A change that lands in
one and not the other is a bug.
## Kernel cmdline caveat
`crashkernel=` (and any future `hugepages=`-style reservation) only takes
effect at boot: the fixup writes `/etc/default/grub` + `update-grub` and
logs `takes effect on the NEXT reboot`. Operators reboot nodes when
applying releases; no special ceremony is required beyond that, but the
lifecycle gate grades this state honestly (WARN for written-but-not-yet-
rebooted, FAIL for never-written — see `tests/lifecycle/os-audit.sh`
section D).
## Verification story
- Unit tests pin the policy constants and script shapes
(`host_fixups` tests in `core/archipelago`).
- `tests/lifecycle/os-audit.sh` section D asserts the end state on a real
node (config present, crashkernel reserved or pending reboot, hang
policy live, rasdaemon active).
- The lifecycle gate runs on archi-dev-box per release; the QEMU ISO
smoke covers fresh installs.
## Future payloads (candidates, not commitments)
- `unattended-upgrades` posture + a default-deny host nftables ruleset
(the §F hardening-plan item — needs its own design first).
- Host firewall rules for mesh/WG ports.
- Chronic: anything the image learns post-deploy that old nodes must
converge on (the polkit and audio precedents, formalized).
@@ -567,6 +567,33 @@ RUN mkdir -p /etc/polkit-1/rules.d && \
> /etc/polkit-1/rules.d/49-archipelago-networkmanager.rules && \
chmod 644 /etc/polkit-1/rules.d/49-archipelago-networkmanager.rules
# kdump + rasdaemon (#144, docs/kdump-rasdaemon-design.md): crash dumps and
# hardware-error capture on the host. Packages + config are baked in for fresh
# installs; the binary's host_fixups module delivers the identical end state to
# already-deployed nodes over OTA (idempotent no-op here once applied).
RUN set -eu; \
apt-get update; \
apt-get install -y --no-install-recommends kdump-tools kexec-tools rasdaemon; \
apt-get clean; rm -rf /var/lib/apt/lists/*; \
CONF=/etc/default/kdump-tools; \
sed -i 's|^#\?USE_KDUMP=.*|USE_KDUMP="1"|' "$CONF"; \
grep -q '^KDUMP_COREDIR=' "$CONF" \
&& sed -i 's|^KDUMP_COREDIR=.*|KDUMP_COREDIR="/var/crash"|' "$CONF" \
|| printf '\nKDUMP_COREDIR="/var/crash"\n' >> "$CONF"; \
grep -q '^CORE_COLLECTOR=' "$CONF" \
&& sed -i 's|^CORE_COLLECTOR=.*|CORE_COLLECTOR="makedumpfile -l --message-level 1 -d 31"|' "$CONF" \
|| printf '\nCORE_COLLECTOR="makedumpfile -l --message-level 1 -d 31"\n' >> "$CONF"; \
printf '%s\n' \
'# Archipelago kdump policy (#144). A wedged kiosk is useless until someone' \
'# power-cycles it — capture the evidence, then reboot by itself. Dumps land in' \
'# /var/crash (see docs/kdump-rasdaemon-design.md); keep-2 pruning is done by' \
'# the host fixup pass, not a timer.' \
'kernel.panic = 10' \
'kernel.panic_on_oops = 1' \
'kernel.hung_task_panic = 1' \
'kernel.hardlockup_panic = 1' \
> /etc/sysctl.d/99-archipelago-kdump.conf
# Enable services
RUN systemctl enable NetworkManager || true && \
systemctl enable polkit || systemctl enable polkit.service || true && \
@@ -580,7 +607,9 @@ RUN systemctl enable NetworkManager || true && \
systemctl enable archipelago-update.timer || true && \
systemctl enable archipelago-doctor.timer || true && \
systemctl enable archipelago-tor-helper.path || true && \
systemctl enable nostr-relay || true
systemctl enable nostr-relay || true && \
systemctl enable rasdaemon || true && \
systemctl enable kdump-tools || true
# archipelago-fips.service + archipelago-wg.service + archipelago-wg-address.service
# stay installed and enabled. They all use `ConditionPathExists=` on their
# respective seed-derived key files, so on a fresh pre-onboarding boot
@@ -3715,7 +3744,7 @@ if [ -d "$BOOT_MEDIA/archipelago/plymouth-theme" ]; then
ln -sf /usr/share/plymouth/themes/archipelago/archipelago.plymouth \
/mnt/target/etc/alternatives/default.plymouth 2>/dev/null || true
# Configure clean boot: splash, suppress kernel noise, hide cursor
sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="quiet splash loglevel=0 rd.systemd.show_status=false vt.global_cursor_default=0 acpi=force"/' \
sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="quiet splash loglevel=0 rd.systemd.show_status=false vt.global_cursor_default=0 acpi=force crashkernel=256M"/' \
/mnt/target/etc/default/grub 2>/dev/null || true
echo " Installed Archipelago Plymouth theme on target"
fi
@@ -362,6 +362,23 @@ init()
</button>
</div>
<div class="overflow-y-auto flex-1 min-h-0 space-y-6 pr-1">
<!-- v1.8.5-alpha -->
<div>
<div class="flex items-center gap-2 mb-3">
<span class="text-xs font-mono px-2 py-0.5 rounded bg-orange-500/20 text-orange-300">v1.8.5-alpha</span>
<span class="text-xs text-white/40">August 30, 2026</span>
</div>
<div class="space-y-3 text-sm text-white/80 pl-3 border-l border-white/10">
<p>**Cuprate — an independent Monero node — is now an app.** Monero consensus validated by a second, unrelated codebase (Rust), the same layer of security-in-depth Bitcoin gets from Knots. Review caught two problems before anything shipped: the unrestricted RPC that can move funds stayed bound to the container's loopback (never published to the node, let alone the LAN — anything on the node could previously have reached it), and its restricted RPC moved off port 18089 to avoid colliding with Penpot. Honest caveat: upstream has cut no stable release yet, so the pin tracks an exact preview build (0.1.0-preview-18-g618ff14) and moves to their first tagged release when there is one.</p>
<p>**A frozen node now explains itself — and comes back on its own.** The host now captures a memory dump into /var/crash when the kernel panics *or* wedges (a hung kiosk used to sit dead until someone power-cycled it; now it dumps, reboots itself, and leaves the evidence behind), and records failing-memory signals (ECC errors) into a database as they happen. This is the first change delivered by a new host-update channel: the node's own updater now carries OS-level packages and settings to already-deployed machines — the crash-kernel's memory reservation is the one part that waits for a reboot, and the node says so rather than pretending.</p>
<p>**Uninstalling an app can no longer report success when it failed.** The declarative path used to swallow every teardown error and report the app uninstalled, leaving the tile behind and the truth in the logs. A failed uninstall now stops and shows the real per-app errors, so "still there" is never presented as "gone".</p>
<p>**Pictures to internet-only mesh contacts work now.** Sending an attachment inline always took the radio path and failed with "Peer is federation-only (no radio twin)" for contacts reachable only over the internet — and the size-adviser kept recommending a radio transfer those peers can't receive. Both fixed: inline sends route over the federation when that's the only way to reach the peer, and the advice no longer offers radio-only transfers to radio-unreachable contacts.</p>
<p>**Disk cleanup finally has honest numbers.** Space "free" on a drive was counted including the slice the filesystem keeps reserved for root — roughly 5% of the disk, 92 GB on one dev box — so the automatic cleanup that's supposed to kick in at 90% never triggered and stale container images piled up unnoticed. Reserved space now counts as used, which is what the threshold was always meant to measure.</p>
<p>**Three small screens that were lying to you, fixed.** The "Bitcoin is synced — fund your wallet" toast no longer appears on a node where the wallet it means (LND) isn't installed — it points at installing LND instead. The seed-reveal screen hides its third prompt unless the password actually fails to decrypt (the backup passphrase only exists if you set one). And multi-version store cards stop quoting a version number you'll be asked to choose on the next screen anyway.</p>
<p>**Mesh notifications survive a refresh, and a stale router no longer hides the fix.** Radio message unread counts are now remembered per contact instead of guessed from session state (the "one new message showed 11 unread" bug), cover Meshtastic, MeshCore and Reticulum alike, and deep-link to the right conversation; a single new message announces itself once. Separately, when the cached router address goes stale, the error card gains a "Reconfigure router" action instead of a Retry loop that can never succeed.</p>
<p>**The app updater now knows what upstream shipped.** Every app's manifest records where it comes from — including the odd corners (GitLab-only projects, ghcr-only images) — and a checker sweeps all of them against upstream releases, so a pin that quietly rots for months is now visible instead of invisible. The first full sweep found 27 pins behind; the safe patch-level ones shipped with this release (strfry, BTCPay Server 2.4.3, the two nginx frontends), and the major jumps that may carry data migrations are deliberately held for their own careful passes.</p>
</div>
</div>
<!-- v1.8.4-alpha -->
<div>
<div class="flex items-center gap-2 mb-3">
+123 -13
View File
@@ -505,6 +505,10 @@
"network_policy": "bridge",
"readonly_root": true
},
"upstream": {
"kind": "gitlab",
"repo": "ark-bitcoin/bark"
},
"version": "0.3.0",
"volumes": [
{
@@ -978,13 +982,13 @@
"version": "1.2.11"
},
"btcpay": {
"image": "docker.io/btcpayserver/btcpayserver:2.4.2",
"image": "docker.io/btcpayserver/btcpayserver:2.4.3",
"images": {
"archy-btcpay-db": "source.archipelago-foundation.org/lfg2025/postgres:15.17",
"archy-nbxplorer": "source.archipelago-foundation.org/lfg2025/nbxplorer:2.6.0",
"btcpay-server": "docker.io/btcpayserver/btcpayserver:2.4.2"
"btcpay-server": "docker.io/btcpayserver/btcpayserver:2.4.3"
},
"version": "2.4.2"
"version": "2.4.3"
},
"btcpay-server": {
"manifest": {
@@ -1000,7 +1004,7 @@
"template": "{{HOST_IP}}:23000"
}
],
"image": "docker.io/btcpayserver/btcpayserver:2.4.2",
"image": "docker.io/btcpayserver/btcpayserver:2.4.3",
"network": "archy-net",
"pull_policy": "if-not-present",
"secret_env": [
@@ -1097,7 +1101,7 @@
"kind": "github",
"repo": "btcpayserver/btcpayserver"
},
"version": "2.4.2",
"version": "2.4.3",
"volumes": [
{
"options": [
@@ -1110,7 +1114,7 @@
]
}
},
"version": "2.4.2"
"version": "2.4.3"
},
"core-lightning": {
"manifest": {
@@ -1205,6 +1209,96 @@
"image": "source.archipelago-foundation.org/lfg2025/cryptpad:2024.12.0",
"version": "2024.12.0"
},
"cuprate": {
"manifest": {
"app": {
"category": "money",
"container": {
"custom_args": [
"--config-file",
"/home/cuprate/Cuprated.toml"
],
"data_uid": "1000:1000",
"image": "source.archipelago-foundation.org/lfg2025/cuprate:0.1.0-preview-18-g618ff14",
"network": "archy-net",
"pull_policy": "if-not-present"
},
"dependencies": [
{
"storage": "300Gi"
}
],
"description": "Alternative Monero node implementation in Rust. Independently validates Monero consensus rules, providing a layer of security and redundancy for the network.",
"files": [
{
"content": "network = \"Mainnet\"\ntarget_max_memory = 3000000000\n\n[rpc.restricted]\nenable = true\n",
"overwrite": false,
"path": "/var/lib/archipelago/cuprate/Cuprated.toml"
}
],
"health_check": {
"endpoint": "localhost:18090",
"interval": "30s",
"retries": 3,
"start_period": "5m",
"timeout": "5s",
"type": "tcp"
},
"id": "cuprate",
"metadata": {
"author": "Cuprate",
"category": "money",
"icon": "/assets/img/app-icons/cuprate.svg",
"repo": "https://github.com/Cuprate/cuprate",
"tier": "optional"
},
"name": "Cuprate",
"ports": [
{
"auth": "none",
"auth_rationale": "Monero p2p gossip. Peers are anonymous by design and speak the Monero wire protocol, not HTTP.",
"container": 18080,
"host": 18183,
"protocol": "tcp"
},
{
"auth": "none",
"auth_rationale": "Monero restricted RPC — the subset upstream considers safe for public/remote-node use. Wallets (Feather, monero-wallet-rpc, GUI) connect directly over plain HTTP JSON-RPC and cannot hold a dashboard session cookie.",
"container": 18089,
"host": 18090,
"protocol": "tcp"
}
],
"resources": {
"cpu_limit": 0,
"disk_limit": "300Gi",
"memory_limit": "4Gi"
},
"security": {
"capabilities": [],
"network_policy": "isolated",
"no_new_privileges": true,
"readonly_root": true
},
"upstream": {
"kind": "github",
"repo": "Cuprate/cuprate"
},
"version": "0.1.0-preview",
"volumes": [
{
"options": [
"rw"
],
"source": "/var/lib/archipelago/cuprate",
"target": "/home/cuprate",
"type": "bind"
}
]
}
},
"version": "0.1.0-preview"
},
"did-wallet": {
"manifest": {
"app": {
@@ -2375,6 +2469,10 @@
"network_policy": "isolated",
"readonly_root": false
},
"upstream": {
"kind": "ghcr",
"repo": "immich-app/postgres"
},
"version": "14-vectorchord0.4.3-pgvectors0.2.0",
"volumes": [
{
@@ -2788,6 +2886,10 @@
"network_policy": "isolated",
"readonly_root": false
},
"upstream": {
"kind": "github",
"repo": "minio/minio"
},
"version": "RELEASE.2024-11-07T00-52-20Z",
"volumes": [
{
@@ -3175,6 +3277,10 @@
"seccomp_profile": "default",
"user": 1000
},
"upstream": {
"kind": "manual",
"url": "no public listing for lightninglabs/lightning-stack — verify by hand"
},
"version": "0.12.0",
"volumes": [
{
@@ -3625,7 +3731,7 @@
"key": "/var/lib/archipelago/netbird/tls.key"
}
],
"image": "docker.io/library/nginx:1.31.3-alpine",
"image": "docker.io/library/nginx:1.31.4-alpine",
"network": "netbird-net",
"pull_policy": "if-not-present"
},
@@ -4315,7 +4421,7 @@
"key": "/var/lib/archipelago/pine/tls.key"
}
],
"image": "docker.io/library/nginx:1.31.3-alpine",
"image": "docker.io/library/nginx:1.31.4-alpine",
"network": "archy-net",
"network_aliases": [
"pine"
@@ -4701,6 +4807,10 @@
"no_new_privileges": true,
"readonly_root": false
},
"upstream": {
"kind": "dockerhub",
"repo": "rhasspy/wyoming-whisper"
},
"version": "3.4.2",
"volumes": [
{
@@ -4998,7 +5108,7 @@
"manifest": {
"app": {
"container": {
"image": "dockurr/strfry:1.1.1",
"image": "dockurr/strfry:1.1.2",
"image_signature": "cosign://...",
"pull_policy": "verify-signature"
},
@@ -5055,7 +5165,7 @@
"kind": "github",
"repo": "hoytech/strfry"
},
"version": "1.1.1",
"version": "1.1.2",
"volumes": [
{
"options": [
@@ -5076,7 +5186,7 @@
]
}
},
"version": "1.1.1"
"version": "1.1.2"
},
"tailscale": {
"image": "source.archipelago-foundation.org/lfg2025/tailscale:stable",
@@ -5256,7 +5366,7 @@
}
},
"schema": 1,
"signature": "97628de24e3ffa17f639c663e19881cf6dea8c79aab272fe9c5442a4e951b3f0d257fee21aa9ce6158e3824b56a337acb71805f6fd34245e48345b86b46ec007",
"signature": "da5b6b183ac46c062945c27abdc06affb558e805e1ccf67ac0ee17e5e3dd85cc05a0656dd83bdacb1e1d237445145d00995f55e77209e1cbb2b6d8ce47084e0a",
"signed_by": "did:key:z6Mkfu5LT8d4DjETtrkATvHh9Dvcbnr7zBCUwfau8Sw7DLWT",
"updated": "2026-08-19"
"updated": "2026-08-30"
}
+42
View File
@@ -9,6 +9,8 @@
# C. FM-guards — the concrete failure modes that have bitten the
# fleet: port-drift (FM8), secret-completeness (FM2),
# orphaned container states (FM9), OTA wedge (FM12)
# D. Host capture (#144) — kdump + rasdaemon baseline: crash dumps configured
# and reserved, hang policy live, ECC recording running
#
# Everything here is READ-ONLY: no install/stop/start/uninstall, no service bounce.
# Safe to run against a live production node. It is the per-boot building block the
@@ -226,6 +228,43 @@ section_c() {
fi
}
# ══ Section D — host capture (#144): kdump + rasdaemon ═══════════════════════
section_d() {
echo
echo "== D. Host capture — crash + hardware-error evidence (#144) =="
if [[ "$ARCHY_LOCAL" != "1" ]]; then
record WARN "kdump + rasdaemon baseline" "remote node — host checks skipped"
return
fi
# D1. kdump enabled in config (image bakes it in; OTA host fixups converge)
if grep -qE '^USE_KDUMP=.?1' /etc/default/kdump-tools 2>/dev/null; then
record PASS "kdump-tools configured" "USE_KDUMP=1, dumps to /var/crash"
else
record FAIL "kdump-tools configured" "/etc/default/kdump-tools missing USE_KDUMP=1 — host fixup didn't land"
fi
# D2. crashkernel reservation — memory is reserved at BOOT, so a node that
# took the OTA fixup but hasn't rebooted yet is WARN, not FAIL.
if grep -q 'crashkernel=' /proc/cmdline 2>/dev/null; then
record PASS "crashkernel reserved" "$(grep -oE 'crashkernel=[^ ]+' /proc/cmdline | head -1)"
elif grep -q 'crashkernel=' /etc/default/grub 2>/dev/null; then
record WARN "crashkernel reserved" "written to GRUB — applies on next reboot"
else
record FAIL "crashkernel reserved" "absent from /proc/cmdline AND /etc/default/grub"
fi
# D3. hang/panic capture policy — runtime-settable, expected immediately
if [[ "$(cat /proc/sys/kernel/hung_task_panic 2>/dev/null)" == "1" ]]; then
record PASS "hang-capture policy live" "kernel.hung_task_panic=1"
else
record FAIL "hang-capture policy live" "kernel.hung_task_panic!=1 — sysctl drop-in not applied"
fi
# D4. rasdaemon recording hardware errors (ECC/AER events → sqlite)
if systemctl is-active --quiet rasdaemon 2>/dev/null; then
record PASS "rasdaemon active" "hardware-error events recorded to /var/lib/rasdaemon"
else
record FAIL "rasdaemon active" "service not running — package missing or host fixup failed"
fi
}
# ── run ────────────────────────────────────────────────────────────────────────
echo "=============================================================="
echo " OS-wide audit — ${BASE_URL} ($(date '+%Y-%m-%d %H:%M:%S'))"
@@ -237,6 +276,9 @@ if (( FAIL == 0 )) || [[ -n "$SESSION" ]]; then
section_b
section_c
fi
# Host-capture baseline is independent of RPC health: a wedged backend must
# not mask that the node also stopped capturing evidence.
section_d
echo
echo "=============================================================="