Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
51a5473e22 | ||
|
|
1872fc20ee | ||
|
|
cbd463e980 | ||
|
|
9df580bf2b | ||
|
|
aee7ecaac1 | ||
|
|
e51ceaa250 | ||
|
|
7c9559aa57 |
@@ -1,5 +1,23 @@
|
||||
# Changelog
|
||||
|
||||
## v1.8.5-alpha (2026-08-30)
|
||||
|
||||
- **Cuprate — an independent Monero node — is now an app.** Monero consensus validated by a second, unrelated codebase (Rust), the same layer of security-in-depth Bitcoin gets from Knots. Review caught two problems before anything shipped: the unrestricted RPC that can move funds stayed bound to the container's loopback (never published to the node, let alone the LAN — anything on the node could previously have reached it), and its restricted RPC moved off port 18089 to avoid colliding with Penpot. Honest caveat: upstream has cut no stable release yet, so the pin tracks an exact preview build (0.1.0-preview-18-g618ff14) and moves to their first tagged release when there is one.
|
||||
|
||||
- **A frozen node now explains itself — and comes back on its own.** The host now captures a memory dump into /var/crash when the kernel panics *or* wedges (a hung kiosk used to sit dead until someone power-cycled it; now it dumps, reboots itself, and leaves the evidence behind), and records failing-memory signals (ECC errors) into a database as they happen. This is the first change delivered by a new host-update channel: the node's own updater now carries OS-level packages and settings to already-deployed machines — the crash-kernel's memory reservation is the one part that waits for a reboot, and the node says so rather than pretending.
|
||||
|
||||
- **Uninstalling an app can no longer report success when it failed.** The declarative path used to swallow every teardown error and report the app uninstalled, leaving the tile behind and the truth in the logs. A failed uninstall now stops and shows the real per-app errors, so "still there" is never presented as "gone".
|
||||
|
||||
- **Pictures to internet-only mesh contacts work now.** Sending an attachment inline always took the radio path and failed with "Peer is federation-only (no radio twin)" for contacts reachable only over the internet — and the size-adviser kept recommending a radio transfer those peers can't receive. Both fixed: inline sends route over the federation when that's the only way to reach the peer, and the advice no longer offers radio-only transfers to radio-unreachable contacts.
|
||||
|
||||
- **Disk cleanup finally has honest numbers.** Space "free" on a drive was counted including the slice the filesystem keeps reserved for root — roughly 5% of the disk, 92 GB on one dev box — so the automatic cleanup that's supposed to kick in at 90% never triggered and stale container images piled up unnoticed. Reserved space now counts as used, which is what the threshold was always meant to measure.
|
||||
|
||||
- **Three small screens that were lying to you, fixed.** The "Bitcoin is synced — fund your wallet" toast no longer appears on a node where the wallet it means (LND) isn't installed — it points at installing LND instead. The seed-reveal screen hides its third prompt unless the password actually fails to decrypt (the backup passphrase only exists if you set one). And multi-version store cards stop quoting a version number you'll be asked to choose on the next screen anyway.
|
||||
|
||||
- **Mesh notifications survive a refresh, and a stale router no longer hides the fix.** Radio message unread counts are now remembered per contact instead of guessed from session state (the "one new message showed 11 unread" bug), cover Meshtastic, MeshCore and Reticulum alike, and deep-link to the right conversation; a single new message announces itself once. Separately, when the cached router address goes stale, the error card gains a "Reconfigure router" action instead of a Retry loop that can never succeed.
|
||||
|
||||
- **The app updater now knows what upstream shipped.** Every app's manifest records where it comes from — including the odd corners (GitLab-only projects, ghcr-only images) — and a checker sweeps all of them against upstream releases, so a pin that quietly rots for months is now visible instead of invisible. The first full sweep found 27 pins behind; the safe patch-level ones shipped with this release (strfry, BTCPay Server 2.4.3, the two nginx frontends), and the major jumps that may carry data migrations are deliberately held for their own careful passes.
|
||||
|
||||
## v1.8.4-alpha (2026-08-20)
|
||||
|
||||
- **Apps with their own login can now skip the node's login screen — Gitea and BTCPay Server do so out of the box.** Some apps bring a complete account system of their own, and putting the node's password page in front of them broke real workflows: git clients can't answer a browser login, and a BTCPay checkout link handed to a customer must open for that customer. These apps are now served directly on their own login, while the node still fronts the connection for everything else it does (embedding fixes, the "app is restarting" page, Tor). Every app gets a new **Settings → app → Access control** switch, so you can put the node login back in front of any app — or take it away from one — with one click, effective immediately. App developers declare the default in their manifest (`auth: open`), documented in the developer guide.
|
||||
|
||||
@@ -0,0 +1,344 @@
|
||||
//! Host-level fixups: OS packages, kernel parameters and system services the
|
||||
//! node needs, delivered by the same signed-binary OTA that ships everything
|
||||
//! else (docs/system-level-ota-design.md).
|
||||
//!
|
||||
//! Scope and posture — read before adding anything here:
|
||||
//!
|
||||
//! * **Idempotent + non-fatal.** Every step is a no-op when the host already
|
||||
//! has the desired state, and a failure (offline box, locked dpkg, missing
|
||||
//! package in the release's Debian suite) logs a warning and moves on. A
|
||||
//! host fixup must never be able to stop the node from starting.
|
||||
//! * **Curated, pinned intent — not dist-upgrade automation.** We deliver the
|
||||
//! specific packages and settings a release deliberately adds (crash
|
||||
//! capture, hardware-error logging, later: unattended-upgrades posture, host
|
||||
//! firewall). Regular Debian upgrades stay with the operator; this channel
|
||||
//! never silently swaps a kernel or a libc.
|
||||
//! * **Fresh installs converge too.** The ISO bakes the same end state in
|
||||
//! (Dockerfile.rootfs, auto-install.sh cmdline), so the fixup is a no-op on
|
||||
//! new machines and only does real work on already-deployed nodes.
|
||||
//! * **Kernel cmdline can't move at runtime.** `crashkernel=` reserves memory
|
||||
//! at boot; the fixup writes GRUB and update-grub so the change lands on the
|
||||
//! next reboot, and says so in the log. Everything else (packages, sysctls,
|
||||
//! services) applies immediately.
|
||||
//!
|
||||
//! First payload (#144, docs/kdump-rasdaemon-design.md): kdump + rasdaemon —
|
||||
//! post-mortem and hardware-error capture:
|
||||
//! * kdump-tools/kexec-tools/rasdaemon installed
|
||||
//! * /etc/default/kdump-tools: USE_KDUMP=1, dumps to /var/crash, compressed
|
||||
//! core collector
|
||||
//! * /etc/sysctl.d/99-archipelago-kdump.conf: a wedged node dumps and
|
||||
//! reboots rather than sitting dead until power-cycled
|
||||
//! * crashkernel=256M appended to the installed GRUB cmdline (next reboot)
|
||||
//! * /var/crash pruned to the two newest dumps
|
||||
//!
|
||||
//! The module is skipped on dev boxes (same guard bootstrap::run uses) and on
|
||||
//! hosts without dpkg.
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use tracing::{debug, info, warn};
|
||||
|
||||
use crate::update::host_sudo;
|
||||
|
||||
/// Packages the node's host must have. Keep this list short and justified —
|
||||
/// every entry is state we now own on the fleet's OS images.
|
||||
const HOST_PACKAGES: &[&str] = &["kdump-tools", "kexec-tools", "rasdaemon"];
|
||||
|
||||
/// Crash-kernel reservation. 256M covers the capture kernel plus makedumpfile
|
||||
/// on the fleet's 16–64GB amd64 machines (~1–2% of RAM, permanently reserved).
|
||||
/// The arm image (RPi) is out of scope for phase 1 — see the design doc.
|
||||
const CRASHKERNEL_PARAM: &str = "crashkernel=256M";
|
||||
|
||||
const KDUMP_SYSDROPIN_PATH: &str = "/etc/sysctl.d/99-archipelago-kdump.conf";
|
||||
const KDUMP_SYSDROPIN: &str = "\
|
||||
# Archipelago kdump policy (#144). A wedged kiosk is useless until someone
|
||||
# power-cycles it — capture the evidence, then reboot by itself. Dumps land in
|
||||
# /var/crash (see docs/kdump-rasdaemon-design.md); keep-2 pruning is done by
|
||||
# the host fixup pass, not a timer.
|
||||
kernel.panic = 10
|
||||
kernel.panic_on_oops = 1
|
||||
kernel.hung_task_panic = 1
|
||||
kernel.hardlockup_panic = 1
|
||||
";
|
||||
|
||||
/// How many dumps to keep in /var/crash. Two ≈ 4 GiB worst case on the 30 GiB
|
||||
/// unencrypted root — the partition usage itself is tracked by disk_monitor.
|
||||
const KEEP_DUMPS: usize = 2;
|
||||
|
||||
/// Entry point, spawned from main.rs at startup like the other ensure_* heals.
|
||||
pub async fn ensure_host_fixups() {
|
||||
// Dev-box guard (same rationale as bootstrap::run): on contributor
|
||||
// machines /home/archipelago/archy is a symlink into a git checkout and
|
||||
// the host is the contributor's own OS — never touch it.
|
||||
let home_archy = std::path::Path::new("/home/archipelago/archy");
|
||||
if tokio::fs::symlink_metadata(home_archy)
|
||||
.await
|
||||
.map(|m| m.file_type().is_symlink())
|
||||
.unwrap_or(false)
|
||||
{
|
||||
debug!("/home/archipelago/archy is a symlink — skipping host fixups (dev box)");
|
||||
return;
|
||||
}
|
||||
// Non-Debian hosts: nothing we manage here applies.
|
||||
if tokio::fs::symlink_metadata("/usr/bin/dpkg").await.is_err() {
|
||||
debug!("no dpkg on this host — skipping host fixups");
|
||||
return;
|
||||
}
|
||||
|
||||
if let Err(e) = run_host_fixups().await {
|
||||
warn!("host fixups failed (non-fatal): {:#}", e);
|
||||
}
|
||||
}
|
||||
|
||||
async fn run_host_fixups() -> Result<()> {
|
||||
// 1. Packages — install only what's missing; a locked/offline apt must
|
||||
// never block anything downstream (steps below degrade to no-ops).
|
||||
match ensure_packages().await {
|
||||
Ok(true) => info!("host fixups: installed missing packages"),
|
||||
Ok(false) => debug!("host fixups: all packages present"),
|
||||
Err(e) => warn!("host fixups: package install failed (non-fatal): {:#}", e),
|
||||
}
|
||||
|
||||
// 2. kdump config + sysctl drop-in + GRUB cmdline + services. One helper
|
||||
// per concern so a failure in one logs and leaves the others running.
|
||||
if let Err(e) = ensure_kdump_sysdropin().await {
|
||||
warn!(
|
||||
"host fixups: kdump sysctl drop-in failed (non-fatal): {:#}",
|
||||
e
|
||||
);
|
||||
}
|
||||
if let Err(e) = ensure_kdump_defaults().await {
|
||||
warn!(
|
||||
"host fixups: kdump-tools config failed (non-fatal): {:#}",
|
||||
e
|
||||
);
|
||||
}
|
||||
match ensure_crashkernel_cmdline().await? {
|
||||
true => {
|
||||
warn!("host fixups: crashkernel= written to GRUB — takes effect on the NEXT reboot")
|
||||
}
|
||||
false => debug!("host fixups: crashkernel already in GRUB cmdline"),
|
||||
}
|
||||
if let Err(e) = ensure_rasdaemon_enabled().await {
|
||||
warn!("host fixups: rasdaemon enable failed (non-fatal): {:#}", e);
|
||||
}
|
||||
if let Err(e) = prune_crash_dumps().await {
|
||||
debug!("host fixups: /var/crash prune skipped: {:#}", e);
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// True if any package was installed. Mirrors the polkit repair's apt posture:
|
||||
/// install without `apt-get update` first; only if that fails (fresh suite,
|
||||
/// stale index), update once and retry. Both under timeout, both non-fatal.
|
||||
async fn ensure_packages() -> Result<bool> {
|
||||
let wanted = HOST_PACKAGES
|
||||
.iter()
|
||||
.map(|p| format!("'{p}'"))
|
||||
.collect::<Vec<_>>()
|
||||
.join(" ");
|
||||
let script = format!(
|
||||
r#"
|
||||
set -u
|
||||
WANTED="{wanted}"
|
||||
MISSING=""
|
||||
for p in $WANTED; do
|
||||
dpkg-query -W -f='${{Status}}' "$p" 2>/dev/null | grep -q 'install ok installed' || MISSING="$MISSING $p"
|
||||
done
|
||||
[ -z "$MISSING" ] && exit 0
|
||||
timeout 240 apt-get install -y --no-install-recommends $MISSING >/dev/null 2>&1 \
|
||||
|| timeout 240 sh -c 'apt-get update >/dev/null 2>&1 && apt-get install -y --no-install-recommends $MISSING >/dev/null 2>&1' \
|
||||
|| exit 3
|
||||
exit 2
|
||||
"#
|
||||
);
|
||||
let status = host_sudo(&["sh", "-lc", &script])
|
||||
.await
|
||||
.context("install host packages")?;
|
||||
match status.code() {
|
||||
Some(0) => Ok(false),
|
||||
Some(2) => Ok(true),
|
||||
code => anyhow::bail!("host package install exited with {code:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
/// Write the sysctl drop-in and apply it live (these four keys are all
|
||||
/// runtime-settable, so the hang/panic policy takes effect without a reboot).
|
||||
async fn ensure_kdump_sysdropin() -> Result<()> {
|
||||
let script = format!(
|
||||
r#"
|
||||
set -u
|
||||
PATH_FILE='{KDUMP_SYSDROPIN_PATH}'
|
||||
CONTENT_FILE=/tmp/archy-kdump-sysctl.$$.tmp
|
||||
cat > "$CONTENT_FILE" <<'SYSEOF'
|
||||
{KDUMP_SYSDROPIN}SYSEOF
|
||||
if [ -f "$PATH_FILE" ] && cmp -s "$CONTENT_FILE" "$PATH_FILE"; then
|
||||
rm -f "$CONTENT_FILE"
|
||||
exit 0
|
||||
fi
|
||||
mv "$CONTENT_FILE" "$PATH_FILE"
|
||||
chmod 644 "$PATH_FILE"
|
||||
sysctl --system >/dev/null 2>&1 || true
|
||||
exit 2
|
||||
"#
|
||||
);
|
||||
let status = host_sudo(&["sh", "-lc", &script])
|
||||
.await
|
||||
.context("write kdump sysctl drop-in")?;
|
||||
match status.code() {
|
||||
Some(0) => Ok(()),
|
||||
Some(2) => {
|
||||
info!("host fixups: installed {KDUMP_SYSDROPIN_PATH} (hang/panic policy)");
|
||||
Ok(())
|
||||
}
|
||||
code => anyhow::bail!("kdump sysctl drop-in exited with {code:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
/// Point kdump-tools at /var/crash with a compressed core collector. Works on
|
||||
/// the package's shipped defaults file (USE_KDUMP=0, commented KDUMP_COREDIR)
|
||||
/// and on any state we already wrote — pure line surgery, idempotent.
|
||||
async fn ensure_kdump_defaults() -> Result<()> {
|
||||
let script = r#"
|
||||
set -u
|
||||
CONF=/etc/default/kdump-tools
|
||||
[ -f "$CONF" ] || exit 3
|
||||
CHANGED=0
|
||||
set_kv() {
|
||||
# set_kv KEY VALUE — replace any (possibly commented) KEY= line with
|
||||
# KEY='VALUE', appending at the end when absent.
|
||||
KEY="$1"; VAL="$2"
|
||||
if grep -qE "^${KEY}=" "$CONF" 2>/dev/null; then
|
||||
if ! grep -qE "^${KEY}='?${VAL}'?$" "$CONF"; then
|
||||
sed -i "s|^${KEY}=.*|${KEY}=\"${VAL}\"|" "$CONF"
|
||||
CHANGED=1
|
||||
fi
|
||||
else
|
||||
printf '\n%s="%s"\n' "$KEY" "$VAL" >> "$CONF"
|
||||
CHANGED=1
|
||||
fi
|
||||
}
|
||||
set_kv USE_KDUMP 1
|
||||
set_kv KDUMP_COREDIR /var/crash
|
||||
set_kv CORE_COLLECTOR 'makedumpfile -l --message-level 1 -d 31'
|
||||
[ "$CHANGED" -eq 1 ] || exit 0
|
||||
systemctl enable kdump-tools >/dev/null 2>&1 || true
|
||||
exit 2
|
||||
"#;
|
||||
let status = host_sudo(&["sh", "-lc", script])
|
||||
.await
|
||||
.context("configure kdump-tools")?;
|
||||
match status.code() {
|
||||
Some(0) => Ok(()),
|
||||
Some(2) => {
|
||||
info!("host fixups: kdump-tools configured (USE_KDUMP=1, /var/crash)");
|
||||
Ok(())
|
||||
}
|
||||
code => anyhow::bail!("kdump-tools config exited with {code:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
/// Append `crashkernel=` to the installed GRUB cmdline and run update-grub.
|
||||
/// The reservation itself only exists after the next reboot — memory cannot
|
||||
/// be set aside at runtime — so the caller must log the reboot caveat.
|
||||
/// Returns true if the cmdline changed.
|
||||
async fn ensure_crashkernel_cmdline() -> Result<bool> {
|
||||
let script = format!(
|
||||
r#"
|
||||
set -u
|
||||
GRUB=/etc/default/grub
|
||||
PARAM='{CRASHKERNEL_PARAM}'
|
||||
[ -f "$GRUB" ] || exit 3
|
||||
LINE=$(grep -E '^GRUB_CMDLINE_LINUX_DEFAULT=' "$GRUB" | head -1)
|
||||
[ -n "$LINE" ] || exit 3
|
||||
case "$LINE" in
|
||||
*"$PARAM"*) exit 0 ;;
|
||||
esac
|
||||
NEWLINE=$(printf '%s' "$LINE" | sed "s/\"$/ $PARAM\"/")
|
||||
sed -i "s|^GRUB_CMDLINE_LINUX_DEFAULT=.*|$NEWLINE|" "$GRUB"
|
||||
timeout 120 update-grub >/dev/null 2>&1 || true
|
||||
exit 2
|
||||
"#
|
||||
);
|
||||
let status = host_sudo(&["sh", "-lc", &script])
|
||||
.await
|
||||
.context("set crashkernel= in GRUB")?;
|
||||
match status.code() {
|
||||
Some(0) => Ok(false),
|
||||
Some(2) => Ok(true),
|
||||
code => anyhow::bail!("crashkernel cmdline fixup exited with {code:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
async fn ensure_rasdaemon_enabled() -> Result<()> {
|
||||
let status = host_sudo(&["systemctl", "enable", "--now", "rasdaemon"])
|
||||
.await
|
||||
.context("enable rasdaemon")?;
|
||||
if status.success() {
|
||||
Ok(())
|
||||
} else {
|
||||
anyhow::bail!("systemctl enable --now rasdaemon exited with {status}")
|
||||
}
|
||||
}
|
||||
|
||||
/// Keep only the newest [`KEEP_DUMPS`] dumps in /var/crash. Called on every
|
||||
/// fixup pass rather than by a timer: the pass runs at every startup, which is
|
||||
/// exactly the cadence at which new dumps appear (a dump ends in a reboot).
|
||||
async fn prune_crash_dumps() -> Result<()> {
|
||||
let script = format!(
|
||||
r#"
|
||||
set -u
|
||||
DIR=/var/crash
|
||||
[ -d "$DIR" ] || exit 0
|
||||
KEEP={KEEP_DUMPS}
|
||||
COUNT=$(ls -1 "$DIR" 2>/dev/null | wc -l)
|
||||
[ "$COUNT" -gt "$KEEP" ] || exit 0
|
||||
ls -1dt "$DIR"/* 2>/dev/null | tail -n +"$((KEEP + 1))" | while IFS= read -r victim; do
|
||||
rm -rf -- "$victim"
|
||||
done
|
||||
exit 2
|
||||
"#
|
||||
);
|
||||
let status = host_sudo(&["sh", "-lc", &script])
|
||||
.await
|
||||
.context("prune /var/crash")?;
|
||||
match status.code() {
|
||||
Some(0) => Ok(()),
|
||||
Some(2) => {
|
||||
info!("host fixups: pruned old dumps in /var/crash (keep {KEEP_DUMPS})");
|
||||
Ok(())
|
||||
}
|
||||
code => anyhow::bail!("/var/crash prune exited with {code:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn sysctl_dropin_carries_the_full_hang_capture_policy() {
|
||||
for key in [
|
||||
"kernel.panic = 10",
|
||||
"kernel.panic_on_oops = 1",
|
||||
"kernel.hung_task_panic = 1",
|
||||
"kernel.hardlockup_panic = 1",
|
||||
] {
|
||||
assert!(KDUMP_SYSDROPIN.contains(key), "drop-in missing {key}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn package_list_is_exactly_the_kdump_rasdaemon_set() {
|
||||
assert_eq!(HOST_PACKAGES, &["kdump-tools", "kexec-tools", "rasdaemon"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn crashkernel_param_is_sized_and_unprefixed() {
|
||||
assert_eq!(CRASHKERNEL_PARAM, "crashkernel=256M");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn keep_dumps_is_two() {
|
||||
assert_eq!(KEEP_DUMPS, 2);
|
||||
}
|
||||
}
|
||||
@@ -55,6 +55,7 @@ mod entropy;
|
||||
mod federation;
|
||||
mod fips;
|
||||
mod health_monitor;
|
||||
mod host_fixups;
|
||||
mod host_ip;
|
||||
mod identity;
|
||||
mod identity_manager;
|
||||
@@ -435,6 +436,12 @@ async fn main() -> Result<()> {
|
||||
// iframe on kiosk nodes (docs/tv-input-iframe-apps.md).
|
||||
tokio::spawn(bootstrap::ensure_gamepad_keys());
|
||||
|
||||
// Host-level fixups (#144 + docs/system-level-ota-design.md): kdump +
|
||||
// rasdaemon — crash/hardware-error capture delivered to already-deployed
|
||||
// nodes over the signed binary OTA. Idempotent, non-fatal, background;
|
||||
// the crashkernel= GRUB edit lands on the next reboot.
|
||||
tokio::spawn(host_fixups::ensure_host_fixups());
|
||||
|
||||
// Mesh access: mirror IPv4-published app ports onto [::] so direct-port
|
||||
// app URLs (http://[<fips0 ULA>]:<port>) work from the companion.
|
||||
tokio::spawn(mesh_ports::run_mesh_port_mirror());
|
||||
|
||||
@@ -54,6 +54,9 @@ step-by-step guides, and some predate the current implementation.
|
||||
- [Dual Ecash](dual-ecash-design.md)
|
||||
- [Hardware Signer](hardware-signer-design.md)
|
||||
- [Manifest Hooks](manifest-hooks-design.md)
|
||||
- [Peering & Federation Trust](peering-trust-model.md) — naming/semantics of trust levels vs discovery (#134)
|
||||
- [kdump + rasdaemon Troubleshooting](kdump-rasdaemon-design.md) — post-mortem and hardware-error capture on nodes (#144)
|
||||
- [System-Level OTA](system-level-ota-design.md) — how host-level packages/config reach already-deployed nodes
|
||||
- [Meshroller Integration](meshroller-integration-design.md)
|
||||
- [Nostr Git Source Hosting](nostr-git-source-hosting.md)
|
||||
- [Nostr Identity Import](nostr-identity-import-plan.md) · [Nostr Signer Login (research)](nostr-signer-login-research.md)
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
# kdump + rasdaemon — post-mortem and hardware-error capture (#144)
|
||||
|
||||
Status: IMPLEMENTED (phase 1) — decisions approved 2026-08-30: hang capture ON,
|
||||
crashkernel=256M, ship the backfill with this release, phase-2 UI deferred.
|
||||
Delivery: image-recipe (Dockerfile.rootfs, auto-install.sh cmdline) +
|
||||
`core/archipelago/src/host_fixups.rs` (existing nodes, see
|
||||
docs/system-level-ota-design.md) + `tests/lifecycle/os-audit.sh` section D.
|
||||
Owner: node image (image-recipe) + lifecycle gate
|
||||
Issue: #144 — "Configure kdump and rasdaemon for troubleshooting"
|
||||
|
||||
## The problem
|
||||
|
||||
When a fleet node hard-locks or a memory stick starts failing, today we get
|
||||
nothing: a frozen kiosk is power-cycled and the evidence is gone; a DIMM
|
||||
throwing correctable ECC errors for weeks is invisible until it starts
|
||||
corrupting things. Two standard kernel mechanisms capture this evidence:
|
||||
|
||||
- **kdump** — reserves a small crash kernel at boot; on a kernel panic (or,
|
||||
configured so, a hang) the running kernel hands the machine over to the
|
||||
crash kernel, which writes a compressed dump of memory to disk and
|
||||
reboots. The node comes back by itself *and* leaves a post-mortem.
|
||||
- **rasdaemon** — a userspace daemon that records hardware error events
|
||||
(correctable/uncorrectable ECC per DIMM, PCIe AER) from EDAC/sysfs into a
|
||||
sqlite database: persistent evidence of degrading hardware with no crash
|
||||
required.
|
||||
|
||||
## Facts the design rests on
|
||||
|
||||
- Installed-disk layout (auto-install.sh): BIOS boot 1MiB · EFI 512MiB ·
|
||||
**root ext4 30GiB, unencrypted** · data (rest, LUKS).
|
||||
- The data partition is LUKS and unlocked late by the node itself — the
|
||||
crash kernel must never be asked to handle key material.
|
||||
- The installed system's kernel command line is written by
|
||||
auto-install.sh:1810 (`GRUB_CMDLINE_LINUX_DEFAULT="quiet splash …"`).
|
||||
- Packages land via `Dockerfile.rootfs` (trixie) with `systemctl enable`
|
||||
in the same RUN block (nginx/tor/avahi pattern).
|
||||
- Kernel cmdline cannot be changed by OTA — it lives in GRUB. Existing
|
||||
nodes need a backfill step (bootstrap) plus a deliberate reboot.
|
||||
|
||||
## Design
|
||||
|
||||
### kdump
|
||||
|
||||
- **Packages:** `kdump-tools kexec-tools` added to Dockerfile.rootfs.
|
||||
- **Command line:** append `crashkernel=256M` to
|
||||
`GRUB_CMDLINE_LINUX_DEFAULT` in auto-install.sh. 256M covers the capture
|
||||
kernel plus makedumpfile on the fleet's 16–64GB amd64 machines (~1–2% of
|
||||
RAM reserved, permanently). The arm image (RPi, config.txt boot) is out
|
||||
of scope for phase 1.
|
||||
- **Dump target:** `local filesystem /var/crash` — on the unencrypted 30GiB
|
||||
root, deliberately *not* the encrypted data partition. No key handling
|
||||
in the crash initramfs, no dependency on the node's own unlock logic.
|
||||
- **Core collector:** `makedumpfile -l --message-level 1 -d 31`
|
||||
(compressed, zero/free pages excluded) — a dump lands at roughly 5–15%
|
||||
of RAM, i.e. ~1–2 GiB on a 16 GiB machine.
|
||||
- **Retention:** keep the **2 newest** dumps only. A small systemd timer
|
||||
(or kdump-tools' `KDUMP_POST_SCRIPT`) prunes older vmcores; a full root
|
||||
partition is already caught by disk_monitor's usage tracking. Two dumps
|
||||
≈ 4 GiB worst case on 30 GiB root — safe.
|
||||
- **When to dump — the deliberate trade-off (decision needed):**
|
||||
- Baseline: dump on real panics (`kernel.panic` path) — no behavioral
|
||||
change to a wedged node.
|
||||
- Recommended for this fleet: also enable hang capture
|
||||
(`kernel.hung_task_panic=1`, hardlockup via NMI watchdog). A kiosk
|
||||
that hard-locks is useless until power-cycled anyway; converting the
|
||||
hang into "dump + automatic reboot" turns every freeze into evidence
|
||||
*and* self-heals the node. Cost: a genuinely-busy-but-alive machine
|
||||
that trips the watchdog reboots — the threshold is kernel-default
|
||||
conservative (40s), so this should be rare.
|
||||
|
||||
### rasdaemon
|
||||
|
||||
- **Packages:** `rasdaemon`; `systemctl enable rasdaemon` in the
|
||||
Dockerfile.rootfs enable block (same pattern as nginx).
|
||||
- **Storage:** its default sqlite DB at
|
||||
`/var/lib/rasdaemon/ras-mc_event.db` on the unencrypted root.
|
||||
- **Human access today:** `ras-mc-ctl --summary` / `--errors` over SSH.
|
||||
No UI in phase 1.
|
||||
|
||||
### Surfacing (phase 2 — separate follow-up, not in this cut)
|
||||
|
||||
A small read-only `system.diagnostics` surface: last-crash timestamp and
|
||||
vmcore sizes from `/var/crash`, plus ECC error totals per DIMM from the
|
||||
rasdaemon DB — shown in Settings → System. Deliberately deferred: capture
|
||||
first, UI once there is something to show and a node in the fleet has
|
||||
actually produced a dump.
|
||||
|
||||
### Existing nodes (phase 1.5 backfill)
|
||||
|
||||
The OTA cannot change the bootloader. Bootstrap (which already delivers
|
||||
fixes to existing nodes) appends `crashkernel=256M` (and the chosen
|
||||
panic/hang params) to `/etc/default/grub` on machines that don't have it,
|
||||
and enables `rasdaemon` via the node's package install path. **Takes
|
||||
effect on the next reboot** — the operator reboots nodes when applying the
|
||||
release; no special ceremony needed beyond that.
|
||||
|
||||
## Testing
|
||||
|
||||
- Image: the new packages appear in the ISO; QEMU boot smoke
|
||||
(build-iso-release.sh stage 5) still green.
|
||||
- Lifecycle gate additions (bats, archi-dev-box first): `kdump-config show`
|
||||
reports a loaded crash kernel reservation; `systemctl is-active
|
||||
rasdaemon`; `/etc/default/grub` carries `crashkernel=`.
|
||||
- Live drill (once, on archi-dev-box, not in the gate): trigger
|
||||
`sysrq c` → vmcore appears in `/var/crash`, node reboots itself,
|
||||
second boot is clean. Keep this manual — it reboots the box.
|
||||
|
||||
## Implementation touchpoints
|
||||
|
||||
1. `image-recipe/build/auto-installer/Dockerfile.rootfs` — packages +
|
||||
`systemctl enable rasdaemon`.
|
||||
2. `image-recipe/build/auto-installer/installer-iso/archipelago/auto-install.sh:1810`
|
||||
— append `crashkernel=256M` (+ hang params if approved) to
|
||||
`GRUB_CMDLINE_LINUX_DEFAULT`.
|
||||
3. `kdump-tools` config: `/etc/default/kdump-tools` (dump target
|
||||
`/var/crash`, core_collector line, `KDUMP_POST_SCRIPT` or timer for
|
||||
retention).
|
||||
4. Bootstrap backfill for existing nodes.
|
||||
5. `tests/lifecycle` — presence assertions (crash kernel reserved,
|
||||
rasdaemon active).
|
||||
|
||||
## Decisions needed before implementation
|
||||
|
||||
1. **Hang capture on or off?** Recommended ON (`hung_task_panic=1` +
|
||||
NMI watchdog): every hard lockup becomes a dump + self-reboot. OFF
|
||||
means dumps only on true panics; wedged nodes still need the button.
|
||||
2. **crashkernel=256M vs 320M** — 256M is the common default for
|
||||
16–64GB machines; 320M if we expect large io-heavy kernels.
|
||||
3. **Backfill now or new-installs-only?** Recommended: ship the backfill
|
||||
with the next release so the whole fleet gains capture on reboot.
|
||||
4. Phase-2 UI surfacing scope — confirm "later" so phase 1 stays small.
|
||||
@@ -0,0 +1,47 @@
|
||||
# Peering & Federation Trust — naming and semantics
|
||||
|
||||
Status: TERMINOLOGY SET — records what the code does today (#134).
|
||||
Deferred: the "don't advertise my peers" opt-out (see §Open questions).
|
||||
|
||||
The code is the authority; this doc gives names to the four concepts that
|
||||
issue #134 showed get conflated in conversation. Where a name changed in
|
||||
user-facing discussion, the term below is the one to use everywhere
|
||||
(UI copy, docs, issues, reviews).
|
||||
|
||||
## The four concepts
|
||||
|
||||
| Term (use this) | What it is | Where it lives |
|
||||
|---|---|---|
|
||||
| **Trusted peer** | A node THIS operator invited and verified: bilateral DID challenge over an out-of-band invite code (`federation::sync`, ADR-007). The only level that grants full access. | `TrustLevel::Trusted`, set via `TrustSource::Invite` or `Manual` |
|
||||
| **Discovered peer** | A peer we learned about from a Trusted peer's advertised list — the transitive merge. Never better than **Observer**: `TRUST IS NOT TRANSITIVE` (sync.rs guard). | `TrustLevel::Observer`, `TrustSource::TransitiveMerge` |
|
||||
| **Routing hint** | What a Discovered peer actually contributes: an address that lets us route directly over FIPS without a second invite hop. Reachability, not trust. | Observer-level sync + FIPS endpoint records |
|
||||
| **Peer advertisement** | The act of a Trusted peer sharing its own peer list during sync. This is the *mechanism* #134 observed — a feature, not a leak. | sync.rs merge path |
|
||||
|
||||
## The two rules that make it sound
|
||||
|
||||
1. **Trust requires an operator decision, always traceable.** Every trust
|
||||
level carries a `TrustSource`. Only a minted invite (or an explicit
|
||||
operator change) can produce `Trusted`; uninvited joins and transitive
|
||||
merges are hard-capped at `Observer` — a peer can never expand our
|
||||
trusted set on its own authority.
|
||||
2. **Discovery is transitive; trust is not.** Seeing more nodes through a
|
||||
Trusted peer is expected and useful (routing). Granting those nodes
|
||||
anything is an operator action, never automatic.
|
||||
|
||||
## Why a Trusted peer advertising its list is by design
|
||||
|
||||
Without advertisement, every new node needs a direct invite from every node
|
||||
that wants to reach it — the invite graph becomes the routing bottleneck
|
||||
AdDR-007 set out to remove. With it, one invite makes a node *reachable* to
|
||||
the trusted set (routing hints), while *authorization* still requires each
|
||||
operator's own invite. Reachability ≠ access.
|
||||
|
||||
## Open questions (deferred, tracked in #134)
|
||||
|
||||
- **"Don't advertise my peers"** — an operator privacy toggle suppressing
|
||||
peer advertisement during sync. Small code change, real design questions:
|
||||
it hides peers who may WANT discovery, and it degrades the routing benefit
|
||||
for every node trusting you. Needs a product decision, not just code.
|
||||
- **Tier vocabulary in the UI** — whether to surface "Observer" as such or
|
||||
a friendlier term ("Connected"/"Visible") — part of the TODO.md peering
|
||||
trust-model item.
|
||||
@@ -0,0 +1,82 @@
|
||||
# System-Level OTA — host fixups
|
||||
|
||||
Status: Implemented (first payload shipped alongside this doc)
|
||||
Owner: `core/archipelago/src/host_fixups.rs`
|
||||
Related: docs/kdump-rasdaemon-design.md (first payload), CLAUDE.md invariants
|
||||
|
||||
## The problem
|
||||
|
||||
The binary OTA updates the node's own software, and the signed app catalog
|
||||
updates apps. But the **host OS** — Debian packages, kernel parameters,
|
||||
system services — previously moved only through ISO re-installs. A node
|
||||
deployed a year ago can be running today's node software on a host that
|
||||
never gained anything the image learned since. Issue #99 (missing polkit
|
||||
rule on old nodes) and the audio-stack heal were each hand-carved
|
||||
one-off bootstrap repairs; there was no general channel and no stated
|
||||
policy for touching the host from the node.
|
||||
|
||||
## The mechanism
|
||||
|
||||
`host_fixups::ensure_host_fixups()` — spawned from `main.rs` at startup
|
||||
alongside the other `ensure_*` heals, in the background, best-effort:
|
||||
|
||||
1. **Dev-box guard** — skip when `/home/archipelago/archy` is a symlink
|
||||
(contributor checkout) and when there's no dpkg (non-Debian host).
|
||||
2. **Packages** — install only what's missing, from a curated, in-code
|
||||
list (`HOST_PACKAGES`), `apt-get install` first, one `apt-get update`
|
||||
retry, both under timeout, never fatal (offline/locked-dpkg nodes
|
||||
converge on a later boot).
|
||||
3. **Configuration** — idempotent per-concern helpers writing root-owned
|
||||
config (via the existing `host_sudo` path): sysctl drop-ins, service
|
||||
defaults, GRUB cmdline, service enablement.
|
||||
4. **Reporting** — every step logs what it did; failures log warnings and
|
||||
move on. A host fixup must never be able to stop the node from starting.
|
||||
|
||||
### Why embedded-in-the-binary rather than fetched
|
||||
|
||||
Same reasoning as the tor-helper (`bootstrap.rs`): the signed binary OTA
|
||||
is the only authenticated delivery channel every node already trusts and
|
||||
pulls on schedule. Fixups compiled into the binary travel with a version,
|
||||
are reviewable in git, and can't be served to a subset of the fleet.
|
||||
|
||||
## Policy — what may travel this channel
|
||||
|
||||
| May | May not |
|
||||
|---|---|
|
||||
| Specific, pinned packages the node needs (kdump-tools, rasdaemon, …) | `dist-upgrade` or silent kernel/libc swaps — regular Debian upgrades stay with the operator |
|
||||
| Kernel *parameters* via GRUB/sysctl — with the next-reboot caveat logged loudly | Anything requiring a secret, or touching LUKS key material |
|
||||
| Service enablement + config the image also bakes in | Divergence: the ISO must converge to the SAME end state so fresh installs are a no-op |
|
||||
| Small, reviewable, per-concern Rust functions with tests | Shell-script-of-things payloads beyond a single concern |
|
||||
|
||||
The rule: **the ISO and the fixup must express the same intent twice,
|
||||
in reviewable places** — Dockerfile.rootfs/auto-install.sh for fresh
|
||||
installs, `host_fixups.rs` for the deployed fleet. A change that lands in
|
||||
one and not the other is a bug.
|
||||
|
||||
## Kernel cmdline caveat
|
||||
|
||||
`crashkernel=` (and any future `hugepages=`-style reservation) only takes
|
||||
effect at boot: the fixup writes `/etc/default/grub` + `update-grub` and
|
||||
logs `takes effect on the NEXT reboot`. Operators reboot nodes when
|
||||
applying releases; no special ceremony is required beyond that, but the
|
||||
lifecycle gate grades this state honestly (WARN for written-but-not-yet-
|
||||
rebooted, FAIL for never-written — see `tests/lifecycle/os-audit.sh`
|
||||
section D).
|
||||
|
||||
## Verification story
|
||||
|
||||
- Unit tests pin the policy constants and script shapes
|
||||
(`host_fixups` tests in `core/archipelago`).
|
||||
- `tests/lifecycle/os-audit.sh` section D asserts the end state on a real
|
||||
node (config present, crashkernel reserved or pending reboot, hang
|
||||
policy live, rasdaemon active).
|
||||
- The lifecycle gate runs on archi-dev-box per release; the QEMU ISO
|
||||
smoke covers fresh installs.
|
||||
|
||||
## Future payloads (candidates, not commitments)
|
||||
|
||||
- `unattended-upgrades` posture + a default-deny host nftables ruleset
|
||||
(the §F hardening-plan item — needs its own design first).
|
||||
- Host firewall rules for mesh/WG ports.
|
||||
- Chronic: anything the image learns post-deploy that old nodes must
|
||||
converge on (the polkit and audio precedents, formalized).
|
||||
@@ -567,6 +567,33 @@ RUN mkdir -p /etc/polkit-1/rules.d && \
|
||||
> /etc/polkit-1/rules.d/49-archipelago-networkmanager.rules && \
|
||||
chmod 644 /etc/polkit-1/rules.d/49-archipelago-networkmanager.rules
|
||||
|
||||
# kdump + rasdaemon (#144, docs/kdump-rasdaemon-design.md): crash dumps and
|
||||
# hardware-error capture on the host. Packages + config are baked in for fresh
|
||||
# installs; the binary's host_fixups module delivers the identical end state to
|
||||
# already-deployed nodes over OTA (idempotent no-op here once applied).
|
||||
RUN set -eu; \
|
||||
apt-get update; \
|
||||
apt-get install -y --no-install-recommends kdump-tools kexec-tools rasdaemon; \
|
||||
apt-get clean; rm -rf /var/lib/apt/lists/*; \
|
||||
CONF=/etc/default/kdump-tools; \
|
||||
sed -i 's|^#\?USE_KDUMP=.*|USE_KDUMP="1"|' "$CONF"; \
|
||||
grep -q '^KDUMP_COREDIR=' "$CONF" \
|
||||
&& sed -i 's|^KDUMP_COREDIR=.*|KDUMP_COREDIR="/var/crash"|' "$CONF" \
|
||||
|| printf '\nKDUMP_COREDIR="/var/crash"\n' >> "$CONF"; \
|
||||
grep -q '^CORE_COLLECTOR=' "$CONF" \
|
||||
&& sed -i 's|^CORE_COLLECTOR=.*|CORE_COLLECTOR="makedumpfile -l --message-level 1 -d 31"|' "$CONF" \
|
||||
|| printf '\nCORE_COLLECTOR="makedumpfile -l --message-level 1 -d 31"\n' >> "$CONF"; \
|
||||
printf '%s\n' \
|
||||
'# Archipelago kdump policy (#144). A wedged kiosk is useless until someone' \
|
||||
'# power-cycles it — capture the evidence, then reboot by itself. Dumps land in' \
|
||||
'# /var/crash (see docs/kdump-rasdaemon-design.md); keep-2 pruning is done by' \
|
||||
'# the host fixup pass, not a timer.' \
|
||||
'kernel.panic = 10' \
|
||||
'kernel.panic_on_oops = 1' \
|
||||
'kernel.hung_task_panic = 1' \
|
||||
'kernel.hardlockup_panic = 1' \
|
||||
> /etc/sysctl.d/99-archipelago-kdump.conf
|
||||
|
||||
# Enable services
|
||||
RUN systemctl enable NetworkManager || true && \
|
||||
systemctl enable polkit || systemctl enable polkit.service || true && \
|
||||
@@ -580,7 +607,9 @@ RUN systemctl enable NetworkManager || true && \
|
||||
systemctl enable archipelago-update.timer || true && \
|
||||
systemctl enable archipelago-doctor.timer || true && \
|
||||
systemctl enable archipelago-tor-helper.path || true && \
|
||||
systemctl enable nostr-relay || true
|
||||
systemctl enable nostr-relay || true && \
|
||||
systemctl enable rasdaemon || true && \
|
||||
systemctl enable kdump-tools || true
|
||||
# archipelago-fips.service + archipelago-wg.service + archipelago-wg-address.service
|
||||
# stay installed and enabled. They all use `ConditionPathExists=` on their
|
||||
# respective seed-derived key files, so on a fresh pre-onboarding boot
|
||||
@@ -3715,7 +3744,7 @@ if [ -d "$BOOT_MEDIA/archipelago/plymouth-theme" ]; then
|
||||
ln -sf /usr/share/plymouth/themes/archipelago/archipelago.plymouth \
|
||||
/mnt/target/etc/alternatives/default.plymouth 2>/dev/null || true
|
||||
# Configure clean boot: splash, suppress kernel noise, hide cursor
|
||||
sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="quiet splash loglevel=0 rd.systemd.show_status=false vt.global_cursor_default=0 acpi=force"/' \
|
||||
sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="quiet splash loglevel=0 rd.systemd.show_status=false vt.global_cursor_default=0 acpi=force crashkernel=256M"/' \
|
||||
/mnt/target/etc/default/grub 2>/dev/null || true
|
||||
echo " Installed Archipelago Plymouth theme on target"
|
||||
fi
|
||||
|
||||
@@ -362,6 +362,23 @@ init()
|
||||
</button>
|
||||
</div>
|
||||
<div class="overflow-y-auto flex-1 min-h-0 space-y-6 pr-1">
|
||||
<!-- v1.8.5-alpha -->
|
||||
<div>
|
||||
<div class="flex items-center gap-2 mb-3">
|
||||
<span class="text-xs font-mono px-2 py-0.5 rounded bg-orange-500/20 text-orange-300">v1.8.5-alpha</span>
|
||||
<span class="text-xs text-white/40">August 30, 2026</span>
|
||||
</div>
|
||||
<div class="space-y-3 text-sm text-white/80 pl-3 border-l border-white/10">
|
||||
<p>**Cuprate — an independent Monero node — is now an app.** Monero consensus validated by a second, unrelated codebase (Rust), the same layer of security-in-depth Bitcoin gets from Knots. Review caught two problems before anything shipped: the unrestricted RPC that can move funds stayed bound to the container's loopback (never published to the node, let alone the LAN — anything on the node could previously have reached it), and its restricted RPC moved off port 18089 to avoid colliding with Penpot. Honest caveat: upstream has cut no stable release yet, so the pin tracks an exact preview build (0.1.0-preview-18-g618ff14) and moves to their first tagged release when there is one.</p>
|
||||
<p>**A frozen node now explains itself — and comes back on its own.** The host now captures a memory dump into /var/crash when the kernel panics *or* wedges (a hung kiosk used to sit dead until someone power-cycled it; now it dumps, reboots itself, and leaves the evidence behind), and records failing-memory signals (ECC errors) into a database as they happen. This is the first change delivered by a new host-update channel: the node's own updater now carries OS-level packages and settings to already-deployed machines — the crash-kernel's memory reservation is the one part that waits for a reboot, and the node says so rather than pretending.</p>
|
||||
<p>**Uninstalling an app can no longer report success when it failed.** The declarative path used to swallow every teardown error and report the app uninstalled, leaving the tile behind and the truth in the logs. A failed uninstall now stops and shows the real per-app errors, so "still there" is never presented as "gone".</p>
|
||||
<p>**Pictures to internet-only mesh contacts work now.** Sending an attachment inline always took the radio path and failed with "Peer is federation-only (no radio twin)" for contacts reachable only over the internet — and the size-adviser kept recommending a radio transfer those peers can't receive. Both fixed: inline sends route over the federation when that's the only way to reach the peer, and the advice no longer offers radio-only transfers to radio-unreachable contacts.</p>
|
||||
<p>**Disk cleanup finally has honest numbers.** Space "free" on a drive was counted including the slice the filesystem keeps reserved for root — roughly 5% of the disk, 92 GB on one dev box — so the automatic cleanup that's supposed to kick in at 90% never triggered and stale container images piled up unnoticed. Reserved space now counts as used, which is what the threshold was always meant to measure.</p>
|
||||
<p>**Three small screens that were lying to you, fixed.** The "Bitcoin is synced — fund your wallet" toast no longer appears on a node where the wallet it means (LND) isn't installed — it points at installing LND instead. The seed-reveal screen hides its third prompt unless the password actually fails to decrypt (the backup passphrase only exists if you set one). And multi-version store cards stop quoting a version number you'll be asked to choose on the next screen anyway.</p>
|
||||
<p>**Mesh notifications survive a refresh, and a stale router no longer hides the fix.** Radio message unread counts are now remembered per contact instead of guessed from session state (the "one new message showed 11 unread" bug), cover Meshtastic, MeshCore and Reticulum alike, and deep-link to the right conversation; a single new message announces itself once. Separately, when the cached router address goes stale, the error card gains a "Reconfigure router" action instead of a Retry loop that can never succeed.</p>
|
||||
<p>**The app updater now knows what upstream shipped.** Every app's manifest records where it comes from — including the odd corners (GitLab-only projects, ghcr-only images) — and a checker sweeps all of them against upstream releases, so a pin that quietly rots for months is now visible instead of invisible. The first full sweep found 27 pins behind; the safe patch-level ones shipped with this release (strfry, BTCPay Server 2.4.3, the two nginx frontends), and the major jumps that may carry data migrations are deliberately held for their own careful passes.</p>
|
||||
</div>
|
||||
</div>
|
||||
<!-- v1.8.4-alpha -->
|
||||
<div>
|
||||
<div class="flex items-center gap-2 mb-3">
|
||||
|
||||
+123
-13
@@ -505,6 +505,10 @@
|
||||
"network_policy": "bridge",
|
||||
"readonly_root": true
|
||||
},
|
||||
"upstream": {
|
||||
"kind": "gitlab",
|
||||
"repo": "ark-bitcoin/bark"
|
||||
},
|
||||
"version": "0.3.0",
|
||||
"volumes": [
|
||||
{
|
||||
@@ -978,13 +982,13 @@
|
||||
"version": "1.2.11"
|
||||
},
|
||||
"btcpay": {
|
||||
"image": "docker.io/btcpayserver/btcpayserver:2.4.2",
|
||||
"image": "docker.io/btcpayserver/btcpayserver:2.4.3",
|
||||
"images": {
|
||||
"archy-btcpay-db": "source.archipelago-foundation.org/lfg2025/postgres:15.17",
|
||||
"archy-nbxplorer": "source.archipelago-foundation.org/lfg2025/nbxplorer:2.6.0",
|
||||
"btcpay-server": "docker.io/btcpayserver/btcpayserver:2.4.2"
|
||||
"btcpay-server": "docker.io/btcpayserver/btcpayserver:2.4.3"
|
||||
},
|
||||
"version": "2.4.2"
|
||||
"version": "2.4.3"
|
||||
},
|
||||
"btcpay-server": {
|
||||
"manifest": {
|
||||
@@ -1000,7 +1004,7 @@
|
||||
"template": "{{HOST_IP}}:23000"
|
||||
}
|
||||
],
|
||||
"image": "docker.io/btcpayserver/btcpayserver:2.4.2",
|
||||
"image": "docker.io/btcpayserver/btcpayserver:2.4.3",
|
||||
"network": "archy-net",
|
||||
"pull_policy": "if-not-present",
|
||||
"secret_env": [
|
||||
@@ -1097,7 +1101,7 @@
|
||||
"kind": "github",
|
||||
"repo": "btcpayserver/btcpayserver"
|
||||
},
|
||||
"version": "2.4.2",
|
||||
"version": "2.4.3",
|
||||
"volumes": [
|
||||
{
|
||||
"options": [
|
||||
@@ -1110,7 +1114,7 @@
|
||||
]
|
||||
}
|
||||
},
|
||||
"version": "2.4.2"
|
||||
"version": "2.4.3"
|
||||
},
|
||||
"core-lightning": {
|
||||
"manifest": {
|
||||
@@ -1205,6 +1209,96 @@
|
||||
"image": "source.archipelago-foundation.org/lfg2025/cryptpad:2024.12.0",
|
||||
"version": "2024.12.0"
|
||||
},
|
||||
"cuprate": {
|
||||
"manifest": {
|
||||
"app": {
|
||||
"category": "money",
|
||||
"container": {
|
||||
"custom_args": [
|
||||
"--config-file",
|
||||
"/home/cuprate/Cuprated.toml"
|
||||
],
|
||||
"data_uid": "1000:1000",
|
||||
"image": "source.archipelago-foundation.org/lfg2025/cuprate:0.1.0-preview-18-g618ff14",
|
||||
"network": "archy-net",
|
||||
"pull_policy": "if-not-present"
|
||||
},
|
||||
"dependencies": [
|
||||
{
|
||||
"storage": "300Gi"
|
||||
}
|
||||
],
|
||||
"description": "Alternative Monero node implementation in Rust. Independently validates Monero consensus rules, providing a layer of security and redundancy for the network.",
|
||||
"files": [
|
||||
{
|
||||
"content": "network = \"Mainnet\"\ntarget_max_memory = 3000000000\n\n[rpc.restricted]\nenable = true\n",
|
||||
"overwrite": false,
|
||||
"path": "/var/lib/archipelago/cuprate/Cuprated.toml"
|
||||
}
|
||||
],
|
||||
"health_check": {
|
||||
"endpoint": "localhost:18090",
|
||||
"interval": "30s",
|
||||
"retries": 3,
|
||||
"start_period": "5m",
|
||||
"timeout": "5s",
|
||||
"type": "tcp"
|
||||
},
|
||||
"id": "cuprate",
|
||||
"metadata": {
|
||||
"author": "Cuprate",
|
||||
"category": "money",
|
||||
"icon": "/assets/img/app-icons/cuprate.svg",
|
||||
"repo": "https://github.com/Cuprate/cuprate",
|
||||
"tier": "optional"
|
||||
},
|
||||
"name": "Cuprate",
|
||||
"ports": [
|
||||
{
|
||||
"auth": "none",
|
||||
"auth_rationale": "Monero p2p gossip. Peers are anonymous by design and speak the Monero wire protocol, not HTTP.",
|
||||
"container": 18080,
|
||||
"host": 18183,
|
||||
"protocol": "tcp"
|
||||
},
|
||||
{
|
||||
"auth": "none",
|
||||
"auth_rationale": "Monero restricted RPC — the subset upstream considers safe for public/remote-node use. Wallets (Feather, monero-wallet-rpc, GUI) connect directly over plain HTTP JSON-RPC and cannot hold a dashboard session cookie.",
|
||||
"container": 18089,
|
||||
"host": 18090,
|
||||
"protocol": "tcp"
|
||||
}
|
||||
],
|
||||
"resources": {
|
||||
"cpu_limit": 0,
|
||||
"disk_limit": "300Gi",
|
||||
"memory_limit": "4Gi"
|
||||
},
|
||||
"security": {
|
||||
"capabilities": [],
|
||||
"network_policy": "isolated",
|
||||
"no_new_privileges": true,
|
||||
"readonly_root": true
|
||||
},
|
||||
"upstream": {
|
||||
"kind": "github",
|
||||
"repo": "Cuprate/cuprate"
|
||||
},
|
||||
"version": "0.1.0-preview",
|
||||
"volumes": [
|
||||
{
|
||||
"options": [
|
||||
"rw"
|
||||
],
|
||||
"source": "/var/lib/archipelago/cuprate",
|
||||
"target": "/home/cuprate",
|
||||
"type": "bind"
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"version": "0.1.0-preview"
|
||||
},
|
||||
"did-wallet": {
|
||||
"manifest": {
|
||||
"app": {
|
||||
@@ -2375,6 +2469,10 @@
|
||||
"network_policy": "isolated",
|
||||
"readonly_root": false
|
||||
},
|
||||
"upstream": {
|
||||
"kind": "ghcr",
|
||||
"repo": "immich-app/postgres"
|
||||
},
|
||||
"version": "14-vectorchord0.4.3-pgvectors0.2.0",
|
||||
"volumes": [
|
||||
{
|
||||
@@ -2788,6 +2886,10 @@
|
||||
"network_policy": "isolated",
|
||||
"readonly_root": false
|
||||
},
|
||||
"upstream": {
|
||||
"kind": "github",
|
||||
"repo": "minio/minio"
|
||||
},
|
||||
"version": "RELEASE.2024-11-07T00-52-20Z",
|
||||
"volumes": [
|
||||
{
|
||||
@@ -3175,6 +3277,10 @@
|
||||
"seccomp_profile": "default",
|
||||
"user": 1000
|
||||
},
|
||||
"upstream": {
|
||||
"kind": "manual",
|
||||
"url": "no public listing for lightninglabs/lightning-stack — verify by hand"
|
||||
},
|
||||
"version": "0.12.0",
|
||||
"volumes": [
|
||||
{
|
||||
@@ -3625,7 +3731,7 @@
|
||||
"key": "/var/lib/archipelago/netbird/tls.key"
|
||||
}
|
||||
],
|
||||
"image": "docker.io/library/nginx:1.31.3-alpine",
|
||||
"image": "docker.io/library/nginx:1.31.4-alpine",
|
||||
"network": "netbird-net",
|
||||
"pull_policy": "if-not-present"
|
||||
},
|
||||
@@ -4315,7 +4421,7 @@
|
||||
"key": "/var/lib/archipelago/pine/tls.key"
|
||||
}
|
||||
],
|
||||
"image": "docker.io/library/nginx:1.31.3-alpine",
|
||||
"image": "docker.io/library/nginx:1.31.4-alpine",
|
||||
"network": "archy-net",
|
||||
"network_aliases": [
|
||||
"pine"
|
||||
@@ -4701,6 +4807,10 @@
|
||||
"no_new_privileges": true,
|
||||
"readonly_root": false
|
||||
},
|
||||
"upstream": {
|
||||
"kind": "dockerhub",
|
||||
"repo": "rhasspy/wyoming-whisper"
|
||||
},
|
||||
"version": "3.4.2",
|
||||
"volumes": [
|
||||
{
|
||||
@@ -4998,7 +5108,7 @@
|
||||
"manifest": {
|
||||
"app": {
|
||||
"container": {
|
||||
"image": "dockurr/strfry:1.1.1",
|
||||
"image": "dockurr/strfry:1.1.2",
|
||||
"image_signature": "cosign://...",
|
||||
"pull_policy": "verify-signature"
|
||||
},
|
||||
@@ -5055,7 +5165,7 @@
|
||||
"kind": "github",
|
||||
"repo": "hoytech/strfry"
|
||||
},
|
||||
"version": "1.1.1",
|
||||
"version": "1.1.2",
|
||||
"volumes": [
|
||||
{
|
||||
"options": [
|
||||
@@ -5076,7 +5186,7 @@
|
||||
]
|
||||
}
|
||||
},
|
||||
"version": "1.1.1"
|
||||
"version": "1.1.2"
|
||||
},
|
||||
"tailscale": {
|
||||
"image": "source.archipelago-foundation.org/lfg2025/tailscale:stable",
|
||||
@@ -5256,7 +5366,7 @@
|
||||
}
|
||||
},
|
||||
"schema": 1,
|
||||
"signature": "97628de24e3ffa17f639c663e19881cf6dea8c79aab272fe9c5442a4e951b3f0d257fee21aa9ce6158e3824b56a337acb71805f6fd34245e48345b86b46ec007",
|
||||
"signature": "da5b6b183ac46c062945c27abdc06affb558e805e1ccf67ac0ee17e5e3dd85cc05a0656dd83bdacb1e1d237445145d00995f55e77209e1cbb2b6d8ce47084e0a",
|
||||
"signed_by": "did:key:z6Mkfu5LT8d4DjETtrkATvHh9Dvcbnr7zBCUwfau8Sw7DLWT",
|
||||
"updated": "2026-08-19"
|
||||
"updated": "2026-08-30"
|
||||
}
|
||||
|
||||
@@ -9,6 +9,8 @@
|
||||
# C. FM-guards — the concrete failure modes that have bitten the
|
||||
# fleet: port-drift (FM8), secret-completeness (FM2),
|
||||
# orphaned container states (FM9), OTA wedge (FM12)
|
||||
# D. Host capture (#144) — kdump + rasdaemon baseline: crash dumps configured
|
||||
# and reserved, hang policy live, ECC recording running
|
||||
#
|
||||
# Everything here is READ-ONLY: no install/stop/start/uninstall, no service bounce.
|
||||
# Safe to run against a live production node. It is the per-boot building block the
|
||||
@@ -226,6 +228,43 @@ section_c() {
|
||||
fi
|
||||
}
|
||||
|
||||
# ══ Section D — host capture (#144): kdump + rasdaemon ═══════════════════════
|
||||
section_d() {
|
||||
echo
|
||||
echo "== D. Host capture — crash + hardware-error evidence (#144) =="
|
||||
if [[ "$ARCHY_LOCAL" != "1" ]]; then
|
||||
record WARN "kdump + rasdaemon baseline" "remote node — host checks skipped"
|
||||
return
|
||||
fi
|
||||
# D1. kdump enabled in config (image bakes it in; OTA host fixups converge)
|
||||
if grep -qE '^USE_KDUMP=.?1' /etc/default/kdump-tools 2>/dev/null; then
|
||||
record PASS "kdump-tools configured" "USE_KDUMP=1, dumps to /var/crash"
|
||||
else
|
||||
record FAIL "kdump-tools configured" "/etc/default/kdump-tools missing USE_KDUMP=1 — host fixup didn't land"
|
||||
fi
|
||||
# D2. crashkernel reservation — memory is reserved at BOOT, so a node that
|
||||
# took the OTA fixup but hasn't rebooted yet is WARN, not FAIL.
|
||||
if grep -q 'crashkernel=' /proc/cmdline 2>/dev/null; then
|
||||
record PASS "crashkernel reserved" "$(grep -oE 'crashkernel=[^ ]+' /proc/cmdline | head -1)"
|
||||
elif grep -q 'crashkernel=' /etc/default/grub 2>/dev/null; then
|
||||
record WARN "crashkernel reserved" "written to GRUB — applies on next reboot"
|
||||
else
|
||||
record FAIL "crashkernel reserved" "absent from /proc/cmdline AND /etc/default/grub"
|
||||
fi
|
||||
# D3. hang/panic capture policy — runtime-settable, expected immediately
|
||||
if [[ "$(cat /proc/sys/kernel/hung_task_panic 2>/dev/null)" == "1" ]]; then
|
||||
record PASS "hang-capture policy live" "kernel.hung_task_panic=1"
|
||||
else
|
||||
record FAIL "hang-capture policy live" "kernel.hung_task_panic!=1 — sysctl drop-in not applied"
|
||||
fi
|
||||
# D4. rasdaemon recording hardware errors (ECC/AER events → sqlite)
|
||||
if systemctl is-active --quiet rasdaemon 2>/dev/null; then
|
||||
record PASS "rasdaemon active" "hardware-error events recorded to /var/lib/rasdaemon"
|
||||
else
|
||||
record FAIL "rasdaemon active" "service not running — package missing or host fixup failed"
|
||||
fi
|
||||
}
|
||||
|
||||
# ── run ────────────────────────────────────────────────────────────────────────
|
||||
echo "=============================================================="
|
||||
echo " OS-wide audit — ${BASE_URL} ($(date '+%Y-%m-%d %H:%M:%S'))"
|
||||
@@ -237,6 +276,9 @@ if (( FAIL == 0 )) || [[ -n "$SESSION" ]]; then
|
||||
section_b
|
||||
section_c
|
||||
fi
|
||||
# Host-capture baseline is independent of RPC health: a wedged backend must
|
||||
# not mask that the node also stopped capturing evidence.
|
||||
section_d
|
||||
|
||||
echo
|
||||
echo "=============================================================="
|
||||
|
||||
Reference in New Issue
Block a user