Compare commits

..
25 Commits
Author SHA1 Message Date
archipelago 51a5473e22 docs(release): v1.8.5-alpha changelog section + What's New sync
Demo images / Build & push demo images (push) Failing after 42s
Curated release notes for the pending v1.8.5-alpha: Cuprate (with the
two review catches), kdump/rasdaemon + the host-fixup OTA channel, the
uninstall-abort fix, federation inline-picture routing, honest disk
usage, the three lying-screens fixes (#143/#127/#129), durable mesh
notifications + router recovery (#57/#103), and upstream-release tracking
with the first-sweep safe bumps.

What's New modal synced via scripts/sync-whats-new.py (--check passes;
89 versions, all present). Per docs/RELEASE_NOTES_BACKLOG.md the
v1.7.44-alpha -> current section audit remains the open item before the
tag.
2026-08-31 07:23:53 -04:00
archipelago 1872fc20ee feat(image): bake kdump + rasdaemon into fresh installs (#144)
The ISO's Dockerfile.rootfs gains kdump-tools/kexec-tools/rasdaemon with
USE_KDUMP=1, dumps to /var/crash and a compressed core collector, the
hang/panic sysctl drop-in, and rasdaemon + kdump-tools enabled — and the
installed target's GRUB cmdline gains crashkernel=256M next to the
existing quiet/splash line.

Source of truth note: the edit lands in
image-recipe/_archived/build-auto-installer-iso.sh — the builder that
generates the (git-ignored) image-recipe/build/auto-installer/ workspace,
which a cache-hit can reuse. The workspace copy was updated to match so
even a cached build ships the same state. Host fixups (previous commit)
converge already-deployed nodes to exactly this end state, so fresh and
old installs agree.

bash -n clean on the builder.
2026-08-31 07:23:53 -04:00
archipelago cbd463e980 feat(host): crash/hardware-error capture, delivered by a new host-fixup OTA channel (#144)
kdump + rasdaemon on every node, per docs/kdump-rasdaemon-design.md with
the approved decisions: hang capture ON (a wedged kiosk dumps and reboots
itself instead of sitting dead), crashkernel=256M, backfill ships with
this release, phase-2 UI surfacing deferred.

Host fixups (docs/system-level-ota-design.md) are the general answer to
'deliver system-level updates OTA': curated OS packages, sysctl drop-ins,
service enablement and the GRUB crashkernel line, carried by the signed
binary and applied idempotently at startup — non-fatal by construction
(offline/locked-dpkg nodes converge on a later boot), skipped on dev
boxes and non-Debian hosts. This formalizes the polkit/audio repair
precedents into a channel with a stated policy: pinned packages and
parameter intent only, never dist-upgrade automation; the ISO bakes the
identical end state into fresh installs (next commit).

The one runtime limitation is honest: crashkernel memory can only be
reserved at boot, so the fixup writes GRUB, runs update-grub, and logs
that it takes effect on the next reboot.

tests/lifecycle/os-audit.sh gains section D — a graded baseline check:
FAIL if capture never landed, WARN if written but awaiting reboot, PASS
when reserved, policy live and rasdaemon recording. Section D runs
independently of RPC health: a wedged backend must not mask that the
node also stopped capturing evidence.

Verification: host_fixups unit tests 4/4; cargo fmt clean; full suite
runs in the release gate (create-release) and the archi-dev-box
lifecycle gate before the tag.
2026-08-31 07:23:44 -04:00
archipelago 9df580bf2b docs: peering trust terminology — names for the four concepts (#134)
Gives stable names to what issue #134 showed gets conflated: Trusted peer
(invite-verified, operator decision), Discovered peer (learned from a
Trusted peer's advertisement, hard-capped at Observer — TRUST IS NOT
TRANSITIVE), Routing hint (what a Discovered peer actually contributes:
reachability, not trust), and Peer advertisement (the mechanism itself,
a feature not a leak).

Records the two rules that make the model sound (trust requires a
traceable operator decision; discovery is transitive, trust is not), why
advertisement exists (one invite makes a node reachable to the trusted
set without granting anything), and the deferred open questions: the
'don't advertise my peers' privacy toggle and UI tier vocabulary.
2026-08-31 07:23:44 -04:00
archipelago aee7ecaac1 docs: index the kdump/rasdaemon design 2026-08-31 06:11:15 -04:00
archipelago e51ceaa250 docs: draft kdump + rasdaemon troubleshooting design (#144)
Design for capturing post-mortem and hardware-error evidence on fleet
nodes: kdump (crashkernel=256M, dump to /var/crash on the unencrypted
root — never the LUKS data partition, so the crash kernel never handles
key material; makedumpfile-compressed, keep-2 retention) and rasdaemon
(EDAC/ECC events into sqlite on the same root).

Deliberately phased: phase 1 = capture on the image + bootstrap backfill
for existing nodes (kernel cmdline can't travel by OTA; takes effect on
next reboot); phase 2 = a read-only system.diagnostics surface in the
UI, only after a fleet node has produced a real dump.

Four decisions flagged in the doc: hang-capture on/off (recommended ON
— a wedged kiosk is useless anyway, and this turns every freeze into
evidence + self-reboot), crashkernel size, backfill timing, and phase-2
scope. Implementation touchpoints listed (Dockerfile.rootfs,
auto-install.sh:1810 cmdline, kdump-tools config, bootstrap, lifecycle
gate assertions).
2026-08-31 06:10:55 -04:00
archipelago 7c9559aa57 chore(catalog): sign the catalog — Cuprate ships, safe pin bumps land
Signed by the release root (ceremony verify passed locally before push).
Contents of this catalog over the previous one:

  NEW   cuprate           0.1.0-preview-18-g618ff14 — alternative Monero
                        node (Rust); image verified present in the mirror
                        registry; manifest embedded; store entry curated
                        (money / optional)
  BUMP  strfry            1.1.1 -> 1.1.2
  BUMP  btcpay-server     2.4.2 -> 2.4.3
  BUMP  netbird (nginx)   1.31.3-alpine -> 1.31.4-alpine
  BUMP  pine   (nginx)    1.31.3-alpine -> 1.31.4-alpine

All bump targets verified pullable from their public registries before
editing. The three mirror-backed bumps (vaultwarden 1.37.2-alpine,
archy-nbxplorer 2.6.11, home-assistant 2026.8.3) remain parked on
app-bumps-mirror-pending until a live registry-push token exists for the
lfg2025 namespace.

Drift gate clean: check-app-catalog-drift.py --release --strict
(31 store entries, 0 drift, 0 missing). 69 catalog entries total.

Nodes pick this up on their next hourly catalog refresh (or at startup)
— signature verified against the release-root key before application.
2026-08-31 05:56:00 -04:00
archipelago 7b88ba59b2 chore(apps): bump the pins that need no mirroring; curate Cuprate's store entry
Demo images / Build & push demo images (push) Failing after 40s
Pin bumps (all verified pullable from their public registries before
editing, so none can become an image-not-found on a node):

  strfry           1.1.1 -> 1.1.2              (dockurr/strfry, direct pull)
  btcpay-server    2.4.2 -> 2.4.3             (docker.io/btcpayserver, direct pull)
  netbird (nginx)  1.31.3-alpine -> 1.31.4-alpine
  pine   (nginx)   1.31.3-alpine -> 1.31.4-alpine

image-versions.sh moved in lockstep for BTCPAY_IMAGE — it is the baseline
the update badge compares against. Held back deliberately, per the risk
policy from the Aug-17 pass: gitea (four minors of DB migrations),
portainer (six minors), filebrowser (2.27 -> 2.63), fedimint/gateway
(0.8 -> 0.12, real migrations), lnd (money-critical), netbird-server/
netbird-dashboard (0.x, must move in lockstep), and everything with a
major jump or a data migration.

Cuprate also gets its curated store entry (category money, tier optional,
icon, repo) — same shape as the Alby Hub / phoenixd entries — synced
through generate-app-catalog.py into both store catalogs and the
app-session config. The fips launch-port list is unchanged (Cuprate has
no UI port; the generated file round-trips to the committed bytes after
cargo fmt).

Three further bumps are prepared and parked on the
app-bumps-mirror-pending branch, blocked only on a registry-push token:
vaultwarden 1.37.2-alpine, archy-nbxplorer 2.6.11, home-assistant
2026.8.3 — all mirror-backed, and the push credential on record for the
lfg2025 namespace is dead.

Drift gate: check-app-catalog-drift.py --release --strict clean
(31 store entries, 0 drift, 0 missing). appSessionConfig tests 7/7.
2026-08-30 16:22:26 -04:00
archipelago b12d1d3826 feat(apps): track the last untracked apps' upstreams
Five apps had no app.upstream block, so nothing could ever tell us
when their pins fell behind upstream:

  barkd           gitlab ark-bitcoin/bark   (GitLab-only project)
  immich-postgres ghcr  immich-app/postgres (image exists only on ghcr.io)
  indeedhub-minio github minio/minio
  pine-whisper    dockerhub rhasspy/wyoming-whisper
  lightning-stack manual — no public listing exists for
                   lightninglabs/lightning-stack anywhere (docker.io,
                   ghcr.io, github.com all checked), so it is tracked by hand

This adds two fetchers to scripts/check-upstream-releases.py to reach the
first two: latest_gitlab (GitLab releases API; strips the project-name
tag prefix, e.g. bark-0.6.2 -> 0.6.2) and latest_ghcr (anonymous pull
token + tags/list, the same handshake a docker pull performs).

Live-verified after the change:
  barkd            0.3.0 -> 0.6.2   (bump gated on ark_client.rs REST compat)
  immich-postgres  14-vectorchord0.4.3-pgvectors0.2.0 -> 17-vectorchord0.4.3-pgvector0.8.0
  indeedhub-minio  RELEASE.2024-11-07T00-52-20Z -> latest (date-opaque: UNCOMPARABLE, shown for hand comparison)
  pine-whisper     3.4.1 -> 3.6.0   (tuned-args revision needs re-basing, not just a pin move)

Offline coverage check: 59 apps, 0 untracked.
2026-08-30 16:22:11 -04:00
archipelago 698e915df2 Merge PR #141: package Cuprate, an alternative Monero node
Demo images / Build & push demo images (push) Failing after 37s
2026-08-30 14:18:42 -04:00
ssmithxandarchipelago a179df66d8 docs: add app update strategy, SSH access, and app wishlist to TODO
Flags the app update policy already noted as unresolved in
app-developer-guide.md, adds a section for SSH access strategy, and
starts an app wishlist (Cashu wallet, phoenixd) for packaging.
2026-08-30 14:01:20 -04:00
ssmithxandarchipelago 771ff0d28b docs: add TODO.md backlog and link from docs index
Captures unscoped forward-looking items (peering/federation model,
distributed git & OTA, nostr integration, platform/OS, app testing,
observability, and the dev/build process) so they're tracked outside
of ROADMAP.md's curated public summary.
2026-08-30 14:01:20 -04:00
92111385b7 fix(mesh): don't offer radio-only resource transfer to radio-unreachable peers
The federation fallback in the plain content-inline path wasn't enough —
mesh.transport-advice recommended the "resource-mesh" tier purely from our
own device being Reticulum-capable, without checking that THIS peer
actually has a radio route. For a federation-only contact (no radio twin)
that steered the frontend into send-content-inline's Reticulum
resource-transfer path, which has no dest_prefix to send to and fails with
"Peer is federation-only (no radio twin)" — reproduced after deploying the
first fix on a live node.

Adds MeshService::has_radio_route(contact_id), and gates both the
"resource-mesh" tier in mesh.transport-advice and the resource-transfer
branch in mesh.send-content-inline on it. Federation-only peers now fall
through to the has_tor branches, which route the frontend to
mesh.send-content (already correctly federation-aware) instead.

Landed from PR #133 (re-committed to drop private host details from the
original message; content identical).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-30 13:18:08 -04:00
a9e52fa310 fix(mesh): route send-content-inline over federation for radio-less peers
mesh.send-content-inline always called send_typed_wire (the LoRa/radio
path), which fails with "Peer is federation-only (no radio twin)" for
any contact reachable only via Tor federation — reproduced sending a
picture from the companion app to a federation-only peer. mesh.send-content
already resolves the peer's federation onion and falls back to
send_typed_wire_via_federation; mirror that same lookup here.

Landed from PR #133 (re-committed to drop private host details from the
original message; content identical).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-30 13:18:08 -04:00
archipelago b4714f1773 fix(store): defer multi-version app version choice (#129)
Demo images / Build & push demo images (push) Failing after 39s
2026-08-30 10:23:58 -04:00
archipelago d79ca54019 fix(wallet): disclose backup passphrase only when needed (#127) 2026-08-30 10:23:58 -04:00
archipelago 758332d63d fix(openwrt): make stale router config recoverable (#103) 2026-08-30 10:23:58 -04:00
archipelago ee5123af68 test(ui): satisfy strict build indexing
Demo images / Build & push demo images (push) Failing after 41s
2026-08-30 10:18:02 -04:00
archipelago a624d11b6a fix(mesh): make radio message notifications durable (#57) 2026-08-30 10:16:33 -04:00
archipelagoandClaude Opus 5 2c984fbd49 fix(ui): the IBD-finished toast no longer tells a node without LND to fund its wallet
Demo images / Build & push demo images (push) Failing after 52s
When Bitcoin's IBD completed mid-Lightning-goal, the watcher toasted
"you can now fund your wallet" — but the on-chain wallet lives in LND,
not Bitcoin Core. The watcher only checked that the goal had pending
manual steps, never that the install-LND step had completed, so a user
whose LND wasn't installed yet was pointed at a flow that could not
work: the fund modal's address comes from lnd.newaddress and does not
exist until LND is installed (issue #143).

The toast now checks LND's install state at fire time. With LND
installed the message is unchanged; without it, the toast says the
actual next step — install Lightning (LND) — and the Finish setup
button lands on the goal wizard, whose active step is the pending
install-LND one (the wizard itself was already correctly sequenced).

The watcher had no tests; added four pinning its contract: the two
message branches, silence with no in-progress goal, and silence when
the chain was already synced at page load.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 09:24:01 -04:00
archipelago c188d9de78 fix(lifecycle): abort unsafe declarative uninstall 2026-08-23 07:59:40 -04:00
archipelago 37a82fd2f9 fix(cuprate): avoid Penpot RPC port collision 2026-08-23 01:43:09 -04:00
archipelagoandClaude Opus 5 a9a30406df fix(disk): count reserved blocks as used, not free
Disk usage was computed as used/size, where size is the raw device size.
ext4 reserves 5% of the filesystem for root — 92.4 GiB of this node's
1.8 TiB — which size includes but nothing can allocate. Two consequences,
both live on archi-dev-box today:

The dashboard advertised 251 GiB free when only 159 GiB could actually be
written, and reported 86.2% usage against df's 90.8%.

Worse, disk_monitor triggers automatic cleanup (podman image prune) at
90%. The disk has been genuinely above that threshold while this returned
86.2%, so the cleanup never once fired — which is exactly how ~72 GB of
dangling images accumulated unnoticed, and why deleting apps appeared to
free nothing.

Both call sites now ask df for avail and use used/(used+avail): the same
figure df itself prints, and the space an operator can actually spend.
Callers deriving free as total - used now get avail.

Note this shifts disk_total_bytes in the analytics series down by the
reserve; historical samples are not comparable across this change.

Tests updated for the three-column output, plus a regression test built
from this box's real numbers asserting the corrected math crosses the 90%
threshold the old math missed. 15/15 disk_monitor tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 05:01:18 -04:00
archipelagoandClaude Opus 5 f1b5d2d267 fix(cuprate): stop publishing the unauthenticated unrestricted RPC
The manifest bound cuprated's unrestricted RPC (full node control) to
0.0.0.0 inside the container with
i_know_what_im_doing_allow_public_unrestricted_rpc = true, relying on
ports[].bind: 127.0.0.1 to keep it private. That only restricts the HOST
side. Verified live on archi-dev-box 2026-08-22: a peer container got a
valid unauthenticated get_info off container port 18081 — and still did
after cuprate was moved to its own network, because podman bridges route
to each other unless created with --opt isolate=true, which the
orchestrator's auto-create does not pass. Every app on the node could
therefore drive full node control with no credential.

The PR justified this as the pattern bitcoin-knots already uses, but
knots writes rpcuser/rpcpassword from generated secrets, so a 0.0.0.0
bind there still is not control without credentials. cuprated has no RPC
authentication at all, so the two are not equivalent.

Unrestricted RPC is now left at cuprated's own default — container
loopback only, published nowhere, reachable by nothing — which is what
upstream intends by refusing a non-local bind without an explicit
override. Restricted RPC (the safe-for-public subset wallets use) and p2p
are unchanged, and health_check moves to 18089 since 18184 is gone.

Re-verified after the change: peer container gets connection refused on
18081 (exit 7), restricted RPC and the health endpoint still answer, the
node still syncs, validator APPROVED, 76/76 container tests pass
including the unauthenticated-port canary (still 28 — an auth: local
port was removed, not an auth: none one).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 03:10:53 -04:00
ssmithxandClaude Sonnet 5 d6b48ce095 feat(apps): package Cuprate, an alternative Monero node
Full-node daemon: P2P + Monero's own restricted RPC (the safe-for-public
subset wallets use as a "remote node") are auth:none like bitcoin/electrumx's
equivalents; unrestricted RPC (full node control) stays gated auth:local.
readonly_root works cleanly since the upstream image is FROM scratch with
ownership fixed at build time — no runtime chown/setuid needed, unlike
bitcoin-knots/core.

Verified locally end-to-end before committing: built the upstream Dockerfile,
confirmed the generated Cuprated.toml against `cuprated --generate-config`/
`--dry-run`, and ran the real image with the manifest's exact ports/volumes —
including discovering that cuprated's own 127.0.0.1-default RPC bind is
unreachable through a published host port and needs to bind 0.0.0.0
internally with ports[].bind:127.0.0.1 doing the actual restriction, the
same pattern bitcoin-knots' RPC port already uses in this repo.

Bumps the unauthenticated_ports_are_all_accounted_for canary (26 -> 28) for
cuprate's two auth:none ports, per that test's own review-before-updating
contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 13:58:06 +00:00
49 changed files with 1862 additions and 94 deletions
+18
View File
@@ -1,5 +1,23 @@
# Changelog
## v1.8.5-alpha (2026-08-30)
- **Cuprate — an independent Monero node — is now an app.** Monero consensus validated by a second, unrelated codebase (Rust), the same layer of security-in-depth Bitcoin gets from Knots. Review caught two problems before anything shipped: the unrestricted RPC that can move funds stayed bound to the container's loopback (never published to the node, let alone the LAN — anything on the node could previously have reached it), and its restricted RPC moved off port 18089 to avoid colliding with Penpot. Honest caveat: upstream has cut no stable release yet, so the pin tracks an exact preview build (0.1.0-preview-18-g618ff14) and moves to their first tagged release when there is one.
- **A frozen node now explains itself — and comes back on its own.** The host now captures a memory dump into /var/crash when the kernel panics *or* wedges (a hung kiosk used to sit dead until someone power-cycled it; now it dumps, reboots itself, and leaves the evidence behind), and records failing-memory signals (ECC errors) into a database as they happen. This is the first change delivered by a new host-update channel: the node's own updater now carries OS-level packages and settings to already-deployed machines — the crash-kernel's memory reservation is the one part that waits for a reboot, and the node says so rather than pretending.
- **Uninstalling an app can no longer report success when it failed.** The declarative path used to swallow every teardown error and report the app uninstalled, leaving the tile behind and the truth in the logs. A failed uninstall now stops and shows the real per-app errors, so "still there" is never presented as "gone".
- **Pictures to internet-only mesh contacts work now.** Sending an attachment inline always took the radio path and failed with "Peer is federation-only (no radio twin)" for contacts reachable only over the internet — and the size-adviser kept recommending a radio transfer those peers can't receive. Both fixed: inline sends route over the federation when that's the only way to reach the peer, and the advice no longer offers radio-only transfers to radio-unreachable contacts.
- **Disk cleanup finally has honest numbers.** Space "free" on a drive was counted including the slice the filesystem keeps reserved for root — roughly 5% of the disk, 92 GB on one dev box — so the automatic cleanup that's supposed to kick in at 90% never triggered and stale container images piled up unnoticed. Reserved space now counts as used, which is what the threshold was always meant to measure.
- **Three small screens that were lying to you, fixed.** The "Bitcoin is synced — fund your wallet" toast no longer appears on a node where the wallet it means (LND) isn't installed — it points at installing LND instead. The seed-reveal screen hides its third prompt unless the password actually fails to decrypt (the backup passphrase only exists if you set one). And multi-version store cards stop quoting a version number you'll be asked to choose on the next screen anyway.
- **Mesh notifications survive a refresh, and a stale router no longer hides the fix.** Radio message unread counts are now remembered per contact instead of guessed from session state (the "one new message showed 11 unread" bug), cover Meshtastic, MeshCore and Reticulum alike, and deep-link to the right conversation; a single new message announces itself once. Separately, when the cached router address goes stale, the error card gains a "Reconfigure router" action instead of a Retry loop that can never succeed.
- **The app updater now knows what upstream shipped.** Every app's manifest records where it comes from — including the odd corners (GitLab-only projects, ghcr-only images) — and a checker sweeps all of them against upstream releases, so a pin that quietly rots for months is now visible instead of invisible. The first full sweep found 27 pins behind; the safe patch-level ones shipped with this release (strfry, BTCPay Server 2.4.3, the two nginx frontends), and the major jumps that may carry data migrations are deliberately held for their own careful passes.
## v1.8.4-alpha (2026-08-20)
- **Apps with their own login can now skip the node's login screen — Gitea and BTCPay Server do so out of the box.** Some apps bring a complete account system of their own, and putting the node's password page in front of them broke real workflows: git clients can't answer a browser login, and a BTCPay checkout link handed to a customer must open for that customer. These apps are now served directly on their own login, while the node still fronts the connection for everything else it does (embedding fixes, the "app is restarting" page, Tor). Every app gets a new **Settings → app → Access control** switch, so you can put the node login back in front of any app — or take it away from one — with one click, effective immediately. App developers declare the default in their manifest (`auth: open`), documented in the developer guide.
+16 -4
View File
@@ -52,13 +52,13 @@
{
"id": "btcpay-server",
"title": "BTCPay Server",
"version": "2.4.2",
"version": "2.4.3",
"description": "Self-hosted Bitcoin payment processor. Accept Bitcoin payments without intermediaries.",
"icon": "/assets/img/app-icons/btcpay-server.png",
"author": "BTCPay Server Foundation",
"category": "commerce",
"tier": "core",
"dockerImage": "docker.io/btcpayserver/btcpayserver:2.4.2",
"dockerImage": "docker.io/btcpayserver/btcpayserver:2.4.3",
"repoUrl": "https://github.com/btcpayserver/btcpayserver",
"requires": [
"bitcoin-knots"
@@ -378,7 +378,7 @@
"icon": "/assets/img/app-icons/pine.svg",
"author": "Archipelago",
"category": "home",
"dockerImage": "docker.io/library/nginx:1.31.3-alpine",
"dockerImage": "docker.io/library/nginx:1.31.4-alpine",
"repoUrl": "https://github.com/rhasspy/wyoming"
},
{
@@ -464,7 +464,7 @@
"author": "NetBird",
"category": "networking",
"tier": "recommended",
"dockerImage": "docker.io/library/nginx:1.31.3-alpine",
"dockerImage": "docker.io/library/nginx:1.31.4-alpine",
"repoUrl": "https://github.com/netbirdio/netbird",
"containerConfig": {
"ports": [
@@ -571,6 +571,18 @@
"tier": "optional",
"dockerImage": "source.archipelago-foundation.org/lfg2025/phoenixd:0.9.0",
"repoUrl": "https://github.com/ACINQ/phoenixd"
},
{
"id": "cuprate",
"title": "Cuprate",
"version": "0.1.0-preview",
"description": "Alternative Monero node implementation in Rust. Independently validates Monero consensus rules, providing a layer of security and redundancy for the network.",
"icon": "/assets/img/app-icons/cuprate.svg",
"author": "Cuprate contributors",
"category": "money",
"tier": "optional",
"dockerImage": "source.archipelago-foundation.org/lfg2025/cuprate:0.1.0-preview-18-g618ff14",
"repoUrl": "https://github.com/Cuprate/cuprate"
}
]
}
+8
View File
@@ -2,6 +2,14 @@ app:
id: barkd
name: Ark Wallet
version: 0.3.0
# Where this app comes from, so scripts/check-upstream-releases.py can
# tell us when the pin below has fallen behind. bark ships on GitLab only
# (no GitHub mirror), so the gitlab fetcher is the one that can see it.
# NOTE: a version bump is code work, not a pin move — the REST shapes are
# coded in core/archipelago/src/wallet/ark_client.rs (see Dockerfile note).
upstream:
kind: gitlab
repo: ark-bitcoin/bark
description: Ark protocol wallet daemon (barkd). Lets the node hold self-custodial off-chain bitcoin via an Ark server; the wallet talks to it over a local REST API. Signet by default while Ark matures.
container:
+2 -2
View File
@@ -1,7 +1,7 @@
app:
id: btcpay-server
name: BTCPay Server
version: 2.4.2
version: 2.4.3
# Where this app comes from, so scripts/check-upstream-releases.py can
# tell us when the pin below has fallen behind. Without it nothing can:
# container.image names our mirror, not the project it was mirrored from.
@@ -11,7 +11,7 @@ app:
description: Self-hosted Bitcoin payment processor. Accept Bitcoin payments without intermediaries.
container:
image: docker.io/btcpayserver/btcpayserver:2.4.2
image: docker.io/btcpayserver/btcpayserver:2.4.3
pull_policy: if-not-present
network: archy-net
secret_env:
+157
View File
@@ -0,0 +1,157 @@
app:
id: cuprate
name: Cuprate
# Matches the crate's own Cargo.toml version (binaries/cuprated/Cargo.toml).
# Cuprate has no stable release yet — this is explicitly work-in-progress
# software (see upstream README). The image tag below pins the exact
# commit built, since "0.1.0-preview" alone is not reproducible.
version: 0.1.0-preview
# Where this app comes from, so scripts/check-upstream-releases.py can
# tell us when the pin below has fallen behind. Without it nothing can:
# container.image names our mirror, not the project it was mirrored from.
upstream:
kind: github
repo: Cuprate/cuprate
description: Alternative Monero node implementation in Rust. Independently validates Monero consensus rules, providing a layer of security and redundancy for the network.
category: money
metadata:
icon: /assets/img/app-icons/cuprate.svg
repo: https://github.com/Cuprate/cuprate
tier: optional
container:
# Built from the upstream Dockerfile at the tip of main, 18 commits past
# the cuprated-0.1.0-preview tag (commit 618ff14, 2026-08-19) — there is
# no newer tagged release as of this writing. Re-pin to a tagged release
# once upstream cuts one.
image: source.archipelago-foundation.org/lfg2025/cuprate:0.1.0-preview-18-g618ff14
pull_policy: if-not-present
network: archy-net
# The image's own ENTRYPOINT is ["/usr/local/bin/cuprated"]; these are
# appended as its argv, matching the project's own systemd unit
# (cuprated.service) invocation exactly.
custom_args: ["--config-file", "/home/cuprate/Cuprated.toml"]
# The image (FROM scratch) creates uid:gid 1000:1000 for the `cuprate`
# user at build time and runs as it unconditionally (USER 1000:1000,
# no shell to switch users at runtime) — same pattern as
# apps/phoenixd, apps/electrumx, apps/nostr-rs-relay, apps/portainer,
# apps/barkd. The bind-mounted data dir must be owned by that literal
# uid or cuprated dies on a permission error the first time it writes.
data_uid: "1000:1000"
dependencies:
# Monero mainnet is ~250GiB unpruned as of 2026 and growing a few GB a
# month; cuprated's pruning support is not confirmed stable yet (the
# `pruning` crate exists in the workspace but nothing in this config
# surface toggles it), so this sizes for a full unpruned chain plus
# headroom rather than assuming pruning is available.
- storage: 300Gi
resources:
cpu_limit: 0
memory_limit: 4Gi
disk_limit: 300Gi
security:
# FROM scratch, no package manager/shell, ownership fixed at build time
# — unlike bitcoin-knots this needs no runtime chown/setuid dance, so it
# can run fully read-only with an empty capability set.
capabilities: []
readonly_root: true
no_new_privileges: true
network_policy: isolated
ports:
# P2P. Cuprate's own default listen address is already 0.0.0.0
# (p2p.clear_net.listen_on), so no config override is needed — only the
# host-side port differs from Monero's canonical 18080 because that
# number is already taken on this fleet by lnd's REST port.
- host: 18183
container: 18080
protocol: tcp
auth: none
auth_rationale: >-
Monero p2p gossip. Peers are anonymous by design and speak the Monero wire protocol, not HTTP.
# Unrestricted RPC (full node control) is deliberately NOT published.
# cuprated has no RPC authentication, and for a published port to reach
# it the service would have to bind 0.0.0.0 inside the container — at
# which point every other app can reach it directly on 18081, since
# ports[].bind only restricts the HOST side and podman bridges route to
# each other (verified live 2026-08-22: a peer container on archy-net
# got an unauthenticated get_info, from a *different* network). That is
# unlike bitcoin-knots, whose 0.0.0.0 RPC still demands the rpcuser /
# rpcpassword it writes from generated secrets. So unrestricted RPC is
# left at cuprated's own default — container loopback only, reachable by
# nothing — which is also what upstream intends by refusing a non-local
# bind without an explicit i_know_what_im_doing override.
# Restricted RPC: Monero's own purpose-built safe-for-public subset —
# what wallets use when connecting to a "remote node". Disabled by
# cuprated's own default; enabled via files[] below. A dashboard login
# would break wallet clients connecting programmatically, same
# reasoning as electrumx's port. The daemon still uses its canonical
# container port 18089, but Penpot already owns host port 18089, so this
# maps the public host port to the free 18090 instead.
- host: 18090
container: 18089
protocol: tcp
auth: none
auth_rationale: >-
Monero restricted RPC — the subset upstream considers safe for public/remote-node use. Wallets (Feather, monero-wallet-rpc, GUI) connect directly over plain HTTP JSON-RPC and cannot hold a dashboard session cookie.
volumes:
- type: bind
source: /var/lib/archipelago/cuprate
target: /home/cuprate
options: [rw]
# Settings that need to differ from cuprated's own documented defaults
# (verified against `cuprated --generate-config` and `--dry-run` locally,
# 2026-08-21):
# - target_max_memory: cuprated's own default auto-detects total *host*
# RAM via sysinfo, which inside a memory-limited container would let
# it size caches far past what resources.memory_limit above actually
# grants — same class of problem bitcoin-knots' -dbcache sizing
# comment addresses. Set explicitly, comfortably under the 4Gi limit.
# - rpc.restricted.enable: cuprated ships this off by default; flip on
# so the auth:none host port above actually serves something instead
# of refusing every connection. port stays at its documented default
# (canonical 18089), and advertise stays false — this node is not
# opting in to being listed as a public remote node over the p2p
# network, just reachable if someone points a wallet at it directly.
# - rpc.unrestricted.address + the allow-public flag: cuprated's own
# default (127.0.0.1) looks like the obviously-correct choice for a
# port meant to stay loopback-only, but verified live (2026-08-21)
# that a service bound literally to 127.0.0.1 *inside* the container
# is unreachable through the host's published port — connections
# reset regardless of how long the daemon has been up. Binding
# 0.0.0.0 inside and letting ports[].bind: 127.0.0.1 below be the
# actual restriction is the same pattern apps/bitcoin-knots already
# uses for its own RPC port (-rpcbind=0.0.0.0:8332 internally, gate
# restricts it externally) — not a new risk, the same one already
# reviewed and accepted for Bitcoin's RPC.
files:
- path: /var/lib/archipelago/cuprate/Cuprated.toml
content: |
network = "Mainnet"
target_max_memory = 3000000000
[rpc.restricted]
enable = true
overwrite: false
health_check:
type: tcp
# Restricted RPC — the only RPC surface published now.
endpoint: localhost:18090
interval: 30s
timeout: 5s
retries: 3
start_period: 5m
metadata:
icon: /assets/img/app-icons/cuprate.svg
category: money
tier: optional
author: Cuprate
repo: https://github.com/Cuprate/cuprate
+6
View File
@@ -2,6 +2,12 @@ app:
id: immich-postgres
name: Immich Postgres
version: "14-vectorchord0.4.3-pgvectors0.2.0"
# Upstream is the Immich-built Postgres image, published only on ghcr.io
# (no GitHub release tags, no Docker Hub repo) — the ghcr fetcher in
# scripts/check-upstream-releases.py is the only one that can see it.
upstream:
kind: ghcr
repo: immich-app/postgres
description: Postgres (pgvecto.rs / vectorchord) backend for Immich.
# Container named immich_postgres (underscore) to match the runtime's existing
+6
View File
@@ -2,6 +2,12 @@ app:
id: indeedhub-minio
name: IndeedHub MinIO
version: "RELEASE.2024-11-07T00-52-20Z"
# MinIO's release tags are date-opaque (RELEASE.YYYY-MM-DD…), so the
# checker reports them as UNCOMPARABLE rather than ordering them — the
# latest tag is still shown for hand comparison, which is the point.
upstream:
kind: github
repo: minio/minio
description: MinIO S3-compatible object storage for IndeedHub media.
category: community
+6
View File
@@ -2,6 +2,12 @@ app:
id: lightning-stack
name: Lightning Stack
version: 0.12.0
# No public listing exists for lightninglabs/lightning-stack (checked
# docker.io, ghcr.io and github.com) — nothing can be queried automatically,
# so this one is tracked by hand.
upstream:
kind: manual
url: no public listing for lightninglabs/lightning-stack — verify by hand
description: Complete Lightning Network implementation. Includes LND, CLN, and management tools.
container:
+1 -1
View File
@@ -18,7 +18,7 @@ app:
container_name: netbird
container:
image: docker.io/library/nginx:1.31.3-alpine
image: docker.io/library/nginx:1.31.4-alpine
pull_policy: if-not-present
network: netbird-net
# Self-signed TLS cert materialised before create — the dashboard needs a
+8
View File
@@ -6,6 +6,14 @@ app:
# pick up the args change; the pre-release form "3.4.1-1" would compare
# LOWER than 3.4.1 under semver and never roll out.
version: "3.4.2"
# Tracks the rhasspy/wyoming-whisper image we pin (Docker Hub — the
# project's GitHub tags are not the image tags). NOTE: this manifest
# deliberately ships an args-tuned revision AHEAD of the image tag (see
# comment above) — BEHIND here means the image tag moved and the tuned
# revision needs re-basing onto it, not just a pin bump.
upstream:
kind: dockerhub
repo: rhasspy/wyoming-whisper
description: Wyoming-protocol faster-whisper speech-to-text engine. Internal Pine voice-assistant stack member — turns speech captured by a PineVoice satellite into text for Home Assistant Assist.
category: home
+1 -1
View File
@@ -19,7 +19,7 @@ app:
container_name: pine
container:
image: docker.io/library/nginx:1.31.3-alpine
image: docker.io/library/nginx:1.31.4-alpine
pull_policy: if-not-present
network: archy-net
network_aliases: [pine]
+2 -2
View File
@@ -1,7 +1,7 @@
app:
id: strfry
name: Strfry Nostr Relay
version: 1.1.1
version: 1.1.2
# Where this app comes from, so scripts/check-upstream-releases.py can
# tell us when the pin below has fallen behind. Without it nothing can:
# container.image names our mirror, not the project it was mirrored from.
@@ -11,7 +11,7 @@ app:
description: Lightweight Nostr relay written in C++. Alternative to nostr-rs-relay with lower resource usage.
container:
image: dockurr/strfry:1.1.1
image: dockurr/strfry:1.1.2
image_signature: cosign://...
pull_policy: verify-signature
@@ -405,9 +405,17 @@ impl RpcHandler {
.as_ref()
.ok_or_else(|| anyhow::anyhow!("Mesh service not running"))?;
let device_type = svc.shared_state().status.read().await.device_type;
// Resource transfer is a native RNS transfer over LoRa — it needs an
// actual radio route to this contact, not just a Reticulum device on
// our end. A federation-only peer with no radio twin fits the size
// and device-type checks but has no dest_prefix to send to; without
// this check the send falls into send_content_resource and fails
// with "Peer is federation-only (no radio twin)" (picture-send,
// 2026-08-07) instead of falling back to the federation path below.
let use_resource_transfer = bytes.len() > INLINE_HARD_MAX
&& device_type == crate::mesh::types::DeviceType::Reticulum
&& bytes.len() <= RETICULUM_RESOURCE_MAX;
&& bytes.len() <= RETICULUM_RESOURCE_MAX
&& svc.has_radio_route(contact_id).await;
if bytes.len() > INLINE_HARD_MAX && !use_resource_transfer {
anyhow::bail!(
@@ -492,15 +500,58 @@ impl RpcHandler {
)
.await?
} else {
svc.send_typed_wire(
contact_id,
wire,
"content_ref",
&display,
Some(typed_json),
seq,
)
.await?
// Federation-only peers have no radio twin for
// send_typed_wire's LoRa dest-prefix resolution — route over
// Tor federation instead, mirroring mesh.send-content's onion
// lookup, or the send fails with "Peer is federation-only (no
// radio twin)" (picture-send from a federation-only contact,
// 2026-08-07).
let federation_onion = {
let state = svc.shared_state();
let peers = state.peers.read().await;
peers
.get(&contact_id)
.map(|p| (p.pubkey_hex.clone(), p.did.clone()))
};
let federation_onion = match federation_onion {
Some((Some(pubkey_hex), did)) => {
let nodes = crate::federation::load_nodes(&self.config.data_dir)
.await
.unwrap_or_default();
nodes
.iter()
.find(|n| n.pubkey == pubkey_hex)
.map(|n| n.onion.clone())
.or_else(|| {
did.as_ref().and_then(|d| {
nodes.iter().find(|n| &n.did == d).map(|n| n.onion.clone())
})
})
}
_ => None,
};
if let Some(onion) = federation_onion {
svc.send_typed_wire_via_federation(
contact_id,
&onion,
wire,
"content_ref",
&display,
Some(typed_json),
seq,
)
.await?
} else {
svc.send_typed_wire(
contact_id,
wire,
"content_ref",
&display,
Some(typed_json),
seq,
)
.await?
}
}
};
@@ -590,6 +641,16 @@ impl RpcHandler {
let est_seconds = (size.saturating_add(lora_bytes_per_sec - 1) / lora_bytes_per_sec).max(1);
let is_reticulum = device_type == crate::mesh::types::DeviceType::Reticulum;
// A Reticulum device on our end doesn't mean THIS peer is radio
// reachable — a federation-only contact (no radio twin) has no dest
// prefix for a resource transfer, even though it's small enough and
// our device type qualifies. Without this check the frontend was
// steered into mesh.send-content-inline's resource-transfer path,
// which fails with "Peer is federation-only (no radio twin)"
// (picture-send, 2026-08-07); the tier below now defers to the
// has_tor branches for such peers, which route via mesh.send-content
// (federation) instead.
let has_radio_route = is_reticulum && svc.has_radio_route(contact_id).await;
let (tier, reason) = if size <= MESH_AUTO_MAX {
("auto-mesh", "Small enough to send inline over mesh")
} else if size <= MESH_HARD_MAX {
@@ -598,7 +659,7 @@ impl RpcHandler {
} else {
("auto-mesh", "No Tor path — sending inline over mesh")
}
} else if is_reticulum && size <= RETICULUM_RESOURCE_MAX {
} else if has_radio_route && size <= RETICULUM_RESOURCE_MAX {
(
"resource-mesh",
"Sending directly over LoRa via a Reticulum resource transfer",
@@ -365,8 +365,18 @@ impl RpcHandler {
// after uninstall. The reconciler owns a manifest map independent of
// podman state, so a raw `podman rm` alone is not enough.
if let Some(orchestrator) = &self.orchestrator {
let mut teardown_errors = Vec::new();
for app_id in orchestrator_uninstall_app_ids(package_id) {
let _ = orchestrator.remove(&app_id, preserve_data).await;
if let Err(err) = orchestrator.remove(&app_id, preserve_data).await {
teardown_errors.push(format!("{app_id}: {err:#}"));
}
}
if !teardown_errors.is_empty() {
return Err(anyhow::anyhow!(
"Uninstall {} aborted: failed to remove declarative app unit(s): {}",
package_id,
teardown_errors.join("; ")
));
}
}
@@ -2182,6 +2192,11 @@ mod tests {
assert!(!is_missing_container_error("Error: OCI runtime error"));
}
#[test]
fn single_app_uninstall_targets_its_declarative_unit() {
assert_eq!(orchestrator_uninstall_app_ids("cuprate"), vec!["cuprate"]);
}
#[test]
fn runtime_host_ports_are_manifest_derived_for_public_apps() {
assert_eq!(runtime_host_ports("photoprism"), vec![2342]);
+15 -4
View File
@@ -168,7 +168,7 @@ pub(super) async fn read_disk_usage() -> Result<(u64, u64)> {
/// Read disk usage via `df` for a given path.
pub(super) async fn read_disk_usage_path(path: &str) -> Result<(u64, u64)> {
let output = tokio::process::Command::new("df")
.args(["--block-size=1", "--output=used,size", path])
.args(["--block-size=1", "--output=used,size,avail", path])
.output()
.await
.context("Failed to run df")?;
@@ -189,11 +189,22 @@ pub(super) async fn read_disk_usage_path(path: &str) -> Result<(u64, u64)> {
.ok_or_else(|| anyhow::anyhow!("Missing used"))?
.parse()
.context("parse df used")?;
let total: u64 = parts
// Raw `size` includes the filesystem's root-reserved blocks (5% by default
// on ext4 — 92 GiB of this node's 1.8 TiB), which nothing can allocate.
// Reporting it as capacity told the dashboard there were 251 GiB free when
// only 159 GiB were writable. Callers derive free as total - used, so total
// must mean "what can actually be used".
let _size: u64 = parts
.next()
.ok_or_else(|| anyhow::anyhow!("Missing total"))?
.ok_or_else(|| anyhow::anyhow!("Missing size"))?
.parse()
.context("parse df total")?;
.context("parse df size")?;
let avail: u64 = parts
.next()
.ok_or_else(|| anyhow::anyhow!("Missing avail"))?
.parse()
.context("parse df avail")?;
let total = used.saturating_add(avail);
Ok((used, total))
}
+58 -19
View File
@@ -4,9 +4,19 @@
use anyhow::{Context, Result};
use tracing::{info, warn};
/// Parse df output into (used_bytes, total_bytes, used_percent).
/// Expects output from `df --block-size=1 --output=used,size /` which has a header line
/// followed by a data line with two whitespace-separated numbers.
/// Parse df output into (used_bytes, usable_total_bytes, used_percent).
/// Expects `df --block-size=1 --output=used,size,avail <path>`: a header line
/// followed by used, size and avail.
///
/// `size` is deliberately NOT the denominator. ext4 reserves 5% of the
/// filesystem for root — 92 GiB on archi-dev-box's 1.8 TiB disk — which `size`
/// counts but no ordinary process can ever allocate. Dividing by `size`
/// under-reports usage by about five points: on 2026-08-22 that disk was
/// genuinely 90.8% full (159 GiB usable left) while this returned 86.2%, so the
/// 90% auto-cleanup below had never once fired and ~72 GB of dangling images
/// had accumulated. It also meant the dashboard advertised 251 GiB free when
/// only 159 GiB could actually be written. used/(used+avail) is what `df`
/// itself prints and what the operator can actually spend.
fn parse_df_output(stdout: &str) -> Result<(u64, u64, f64)> {
let data_line = stdout
.lines()
@@ -18,11 +28,19 @@ fn parse_df_output(stdout: &str) -> Result<(u64, u64, f64)> {
.ok_or_else(|| anyhow::anyhow!("Missing used"))?
.parse()
.context("parse df used")?;
let total: u64 = parts
// Parsed to keep the column contract explicit, then intentionally unused —
// see the note above on why raw size is the wrong denominator.
let _size: u64 = parts
.next()
.ok_or_else(|| anyhow::anyhow!("Missing total"))?
.ok_or_else(|| anyhow::anyhow!("Missing size"))?
.parse()
.context("parse df total")?;
.context("parse df size")?;
let avail: u64 = parts
.next()
.ok_or_else(|| anyhow::anyhow!("Missing avail"))?
.parse()
.context("parse df avail")?;
let total = used.saturating_add(avail);
let percent = if total > 0 {
(used as f64 / total as f64) * 100.0
@@ -44,7 +62,7 @@ pub async fn check_disk_usage() -> Result<(u64, u64, f64)> {
"/"
};
let output = tokio::process::Command::new("df")
.args(["--block-size=1", "--output=used,size", data_path])
.args(["--block-size=1", "--output=used,size,avail", data_path])
.output()
.await
.context("Failed to run df")?;
@@ -257,8 +275,8 @@ mod tests {
#[test]
fn test_parse_df_output_normal() {
// Simulates typical df --block-size=1 --output=used,size / output
let output = " Used Size\n 500000000000 1000000000000\n";
// df --block-size=1 --output=used,size,avail : used, size, avail
let output = " Used Size Avail\n 500000000000 1000000000000 500000000000\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 500_000_000_000);
assert_eq!(total, 1_000_000_000_000);
@@ -267,16 +285,35 @@ mod tests {
#[test]
fn test_parse_df_output_high_usage() {
let output = " Used Size\n 900000000000 1000000000000\n";
let output = " Used Size Avail\n 900000000000 1000000000000 100000000000\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 900_000_000_000);
assert_eq!(total, 1_000_000_000_000);
assert!((percent - 90.0).abs() < 0.01);
}
/// The bug this function existed to hide: reserved blocks are counted by
/// `size` but are not available to anyone. Real numbers from archi-dev-box,
/// 2026-08-22 — 1.8 TiB disk, ext4 5% reserve, genuinely 90.8% full. The old
/// used/size math returned 86.2%, so the 90% auto-cleanup never triggered.
#[test]
fn reserved_blocks_are_not_counted_as_free() {
let output = "Used Size Avail\n1681459122176 1951249276928 170581372928\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 1_681_459_122_176);
// Total is what can actually be written, not the raw device size.
assert_eq!(total, 1_852_040_495_104);
assert!(
total < 1_951_249_276_928,
"raw size must not be the denominator"
);
assert!((percent - 90.8).abs() < 0.1, "got {percent}");
assert!(percent >= 90.0, "must cross the auto-cleanup threshold");
}
#[test]
fn test_parse_df_output_almost_full() {
let output = "Used Size\n999 1000\n";
let output = "Used Size Avail\n999 1000 1\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 999);
assert_eq!(total, 1000);
@@ -285,7 +322,7 @@ mod tests {
#[test]
fn test_parse_df_output_empty_disk() {
let output = "Used Size\n0 1000000000000\n";
let output = "Used Size Avail\n0 1000000000000 1000000000000\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 0);
assert_eq!(total, 1_000_000_000_000);
@@ -295,7 +332,7 @@ mod tests {
#[test]
fn test_parse_df_output_zero_total() {
// Edge case: total is 0 (should not happen but should not panic/divide-by-zero)
let output = "Used Size\n0 0\n";
let output = "Used Size Avail\n0 0 0\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 0);
assert_eq!(total, 0);
@@ -338,21 +375,23 @@ mod tests {
#[test]
fn test_parse_df_output_extra_whitespace() {
let output = " Used Size \n 123456 7890000 \n";
let output = " Used Size Avail \n 123456 7890000 7766544 \n";
let (used, total, _) = parse_df_output(output).unwrap();
assert_eq!(used, 123456);
assert_eq!(total, 7890000);
assert_eq!(total, 7_890_000);
}
#[test]
fn test_parse_df_output_real_world_format() {
// Closer to real df output with header padding
let output = " Used Size\n 328000000000 1800000000000\n";
// Real df output carries a reserved-block gap: size here is 1.8 TB but
// only 1.382 TB is available, so usable total is used + avail.
let output = " Used Size Avail\n 328000000000 1800000000000 1382000000000\n";
let (used, total, percent) = parse_df_output(output).unwrap();
assert_eq!(used, 328_000_000_000);
assert_eq!(total, 1_800_000_000_000);
// ~18.2%
assert!(percent > 18.0 && percent < 19.0);
assert_eq!(total, 1_710_000_000_000);
// ~19.2% against usable space, not 18.2% against the raw device.
assert!(percent > 19.0 && percent < 20.0, "got {percent}");
}
#[tokio::test]
+344
View File
@@ -0,0 +1,344 @@
//! Host-level fixups: OS packages, kernel parameters and system services the
//! node needs, delivered by the same signed-binary OTA that ships everything
//! else (docs/system-level-ota-design.md).
//!
//! Scope and posture — read before adding anything here:
//!
//! * **Idempotent + non-fatal.** Every step is a no-op when the host already
//! has the desired state, and a failure (offline box, locked dpkg, missing
//! package in the release's Debian suite) logs a warning and moves on. A
//! host fixup must never be able to stop the node from starting.
//! * **Curated, pinned intent — not dist-upgrade automation.** We deliver the
//! specific packages and settings a release deliberately adds (crash
//! capture, hardware-error logging, later: unattended-upgrades posture, host
//! firewall). Regular Debian upgrades stay with the operator; this channel
//! never silently swaps a kernel or a libc.
//! * **Fresh installs converge too.** The ISO bakes the same end state in
//! (Dockerfile.rootfs, auto-install.sh cmdline), so the fixup is a no-op on
//! new machines and only does real work on already-deployed nodes.
//! * **Kernel cmdline can't move at runtime.** `crashkernel=` reserves memory
//! at boot; the fixup writes GRUB and update-grub so the change lands on the
//! next reboot, and says so in the log. Everything else (packages, sysctls,
//! services) applies immediately.
//!
//! First payload (#144, docs/kdump-rasdaemon-design.md): kdump + rasdaemon —
//! post-mortem and hardware-error capture:
//! * kdump-tools/kexec-tools/rasdaemon installed
//! * /etc/default/kdump-tools: USE_KDUMP=1, dumps to /var/crash, compressed
//! core collector
//! * /etc/sysctl.d/99-archipelago-kdump.conf: a wedged node dumps and
//! reboots rather than sitting dead until power-cycled
//! * crashkernel=256M appended to the installed GRUB cmdline (next reboot)
//! * /var/crash pruned to the two newest dumps
//!
//! The module is skipped on dev boxes (same guard bootstrap::run uses) and on
//! hosts without dpkg.
use anyhow::{Context, Result};
use tracing::{debug, info, warn};
use crate::update::host_sudo;
/// Packages the node's host must have. Keep this list short and justified —
/// every entry is state we now own on the fleet's OS images.
const HOST_PACKAGES: &[&str] = &["kdump-tools", "kexec-tools", "rasdaemon"];
/// Crash-kernel reservation. 256M covers the capture kernel plus makedumpfile
/// on the fleet's 16–64GB amd64 machines (~1–2% of RAM, permanently reserved).
/// The arm image (RPi) is out of scope for phase 1 — see the design doc.
const CRASHKERNEL_PARAM: &str = "crashkernel=256M";
const KDUMP_SYSDROPIN_PATH: &str = "/etc/sysctl.d/99-archipelago-kdump.conf";
const KDUMP_SYSDROPIN: &str = "\
# Archipelago kdump policy (#144). A wedged kiosk is useless until someone
# power-cycles it — capture the evidence, then reboot by itself. Dumps land in
# /var/crash (see docs/kdump-rasdaemon-design.md); keep-2 pruning is done by
# the host fixup pass, not a timer.
kernel.panic = 10
kernel.panic_on_oops = 1
kernel.hung_task_panic = 1
kernel.hardlockup_panic = 1
";
/// How many dumps to keep in /var/crash. Two ≈ 4 GiB worst case on the 30 GiB
/// unencrypted root — the partition usage itself is tracked by disk_monitor.
const KEEP_DUMPS: usize = 2;
/// Entry point, spawned from main.rs at startup like the other ensure_* heals.
pub async fn ensure_host_fixups() {
// Dev-box guard (same rationale as bootstrap::run): on contributor
// machines /home/archipelago/archy is a symlink into a git checkout and
// the host is the contributor's own OS — never touch it.
let home_archy = std::path::Path::new("/home/archipelago/archy");
if tokio::fs::symlink_metadata(home_archy)
.await
.map(|m| m.file_type().is_symlink())
.unwrap_or(false)
{
debug!("/home/archipelago/archy is a symlink — skipping host fixups (dev box)");
return;
}
// Non-Debian hosts: nothing we manage here applies.
if tokio::fs::symlink_metadata("/usr/bin/dpkg").await.is_err() {
debug!("no dpkg on this host — skipping host fixups");
return;
}
if let Err(e) = run_host_fixups().await {
warn!("host fixups failed (non-fatal): {:#}", e);
}
}
async fn run_host_fixups() -> Result<()> {
// 1. Packages — install only what's missing; a locked/offline apt must
// never block anything downstream (steps below degrade to no-ops).
match ensure_packages().await {
Ok(true) => info!("host fixups: installed missing packages"),
Ok(false) => debug!("host fixups: all packages present"),
Err(e) => warn!("host fixups: package install failed (non-fatal): {:#}", e),
}
// 2. kdump config + sysctl drop-in + GRUB cmdline + services. One helper
// per concern so a failure in one logs and leaves the others running.
if let Err(e) = ensure_kdump_sysdropin().await {
warn!(
"host fixups: kdump sysctl drop-in failed (non-fatal): {:#}",
e
);
}
if let Err(e) = ensure_kdump_defaults().await {
warn!(
"host fixups: kdump-tools config failed (non-fatal): {:#}",
e
);
}
match ensure_crashkernel_cmdline().await? {
true => {
warn!("host fixups: crashkernel= written to GRUB — takes effect on the NEXT reboot")
}
false => debug!("host fixups: crashkernel already in GRUB cmdline"),
}
if let Err(e) = ensure_rasdaemon_enabled().await {
warn!("host fixups: rasdaemon enable failed (non-fatal): {:#}", e);
}
if let Err(e) = prune_crash_dumps().await {
debug!("host fixups: /var/crash prune skipped: {:#}", e);
}
Ok(())
}
/// True if any package was installed. Mirrors the polkit repair's apt posture:
/// install without `apt-get update` first; only if that fails (fresh suite,
/// stale index), update once and retry. Both under timeout, both non-fatal.
async fn ensure_packages() -> Result<bool> {
let wanted = HOST_PACKAGES
.iter()
.map(|p| format!("'{p}'"))
.collect::<Vec<_>>()
.join(" ");
let script = format!(
r#"
set -u
WANTED="{wanted}"
MISSING=""
for p in $WANTED; do
dpkg-query -W -f='${{Status}}' "$p" 2>/dev/null | grep -q 'install ok installed' || MISSING="$MISSING $p"
done
[ -z "$MISSING" ] && exit 0
timeout 240 apt-get install -y --no-install-recommends $MISSING >/dev/null 2>&1 \
|| timeout 240 sh -c 'apt-get update >/dev/null 2>&1 && apt-get install -y --no-install-recommends $MISSING >/dev/null 2>&1' \
|| exit 3
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("install host packages")?;
match status.code() {
Some(0) => Ok(false),
Some(2) => Ok(true),
code => anyhow::bail!("host package install exited with {code:?}"),
}
}
/// Write the sysctl drop-in and apply it live (these four keys are all
/// runtime-settable, so the hang/panic policy takes effect without a reboot).
async fn ensure_kdump_sysdropin() -> Result<()> {
let script = format!(
r#"
set -u
PATH_FILE='{KDUMP_SYSDROPIN_PATH}'
CONTENT_FILE=/tmp/archy-kdump-sysctl.$$.tmp
cat > "$CONTENT_FILE" <<'SYSEOF'
{KDUMP_SYSDROPIN}SYSEOF
if [ -f "$PATH_FILE" ] && cmp -s "$CONTENT_FILE" "$PATH_FILE"; then
rm -f "$CONTENT_FILE"
exit 0
fi
mv "$CONTENT_FILE" "$PATH_FILE"
chmod 644 "$PATH_FILE"
sysctl --system >/dev/null 2>&1 || true
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("write kdump sysctl drop-in")?;
match status.code() {
Some(0) => Ok(()),
Some(2) => {
info!("host fixups: installed {KDUMP_SYSDROPIN_PATH} (hang/panic policy)");
Ok(())
}
code => anyhow::bail!("kdump sysctl drop-in exited with {code:?}"),
}
}
/// Point kdump-tools at /var/crash with a compressed core collector. Works on
/// the package's shipped defaults file (USE_KDUMP=0, commented KDUMP_COREDIR)
/// and on any state we already wrote — pure line surgery, idempotent.
async fn ensure_kdump_defaults() -> Result<()> {
let script = r#"
set -u
CONF=/etc/default/kdump-tools
[ -f "$CONF" ] || exit 3
CHANGED=0
set_kv() {
# set_kv KEY VALUE — replace any (possibly commented) KEY= line with
# KEY='VALUE', appending at the end when absent.
KEY="$1"; VAL="$2"
if grep -qE "^${KEY}=" "$CONF" 2>/dev/null; then
if ! grep -qE "^${KEY}='?${VAL}'?$" "$CONF"; then
sed -i "s|^${KEY}=.*|${KEY}=\"${VAL}\"|" "$CONF"
CHANGED=1
fi
else
printf '\n%s="%s"\n' "$KEY" "$VAL" >> "$CONF"
CHANGED=1
fi
}
set_kv USE_KDUMP 1
set_kv KDUMP_COREDIR /var/crash
set_kv CORE_COLLECTOR 'makedumpfile -l --message-level 1 -d 31'
[ "$CHANGED" -eq 1 ] || exit 0
systemctl enable kdump-tools >/dev/null 2>&1 || true
exit 2
"#;
let status = host_sudo(&["sh", "-lc", script])
.await
.context("configure kdump-tools")?;
match status.code() {
Some(0) => Ok(()),
Some(2) => {
info!("host fixups: kdump-tools configured (USE_KDUMP=1, /var/crash)");
Ok(())
}
code => anyhow::bail!("kdump-tools config exited with {code:?}"),
}
}
/// Append `crashkernel=` to the installed GRUB cmdline and run update-grub.
/// The reservation itself only exists after the next reboot — memory cannot
/// be set aside at runtime — so the caller must log the reboot caveat.
/// Returns true if the cmdline changed.
async fn ensure_crashkernel_cmdline() -> Result<bool> {
let script = format!(
r#"
set -u
GRUB=/etc/default/grub
PARAM='{CRASHKERNEL_PARAM}'
[ -f "$GRUB" ] || exit 3
LINE=$(grep -E '^GRUB_CMDLINE_LINUX_DEFAULT=' "$GRUB" | head -1)
[ -n "$LINE" ] || exit 3
case "$LINE" in
*"$PARAM"*) exit 0 ;;
esac
NEWLINE=$(printf '%s' "$LINE" | sed "s/\"$/ $PARAM\"/")
sed -i "s|^GRUB_CMDLINE_LINUX_DEFAULT=.*|$NEWLINE|" "$GRUB"
timeout 120 update-grub >/dev/null 2>&1 || true
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("set crashkernel= in GRUB")?;
match status.code() {
Some(0) => Ok(false),
Some(2) => Ok(true),
code => anyhow::bail!("crashkernel cmdline fixup exited with {code:?}"),
}
}
async fn ensure_rasdaemon_enabled() -> Result<()> {
let status = host_sudo(&["systemctl", "enable", "--now", "rasdaemon"])
.await
.context("enable rasdaemon")?;
if status.success() {
Ok(())
} else {
anyhow::bail!("systemctl enable --now rasdaemon exited with {status}")
}
}
/// Keep only the newest [`KEEP_DUMPS`] dumps in /var/crash. Called on every
/// fixup pass rather than by a timer: the pass runs at every startup, which is
/// exactly the cadence at which new dumps appear (a dump ends in a reboot).
async fn prune_crash_dumps() -> Result<()> {
let script = format!(
r#"
set -u
DIR=/var/crash
[ -d "$DIR" ] || exit 0
KEEP={KEEP_DUMPS}
COUNT=$(ls -1 "$DIR" 2>/dev/null | wc -l)
[ "$COUNT" -gt "$KEEP" ] || exit 0
ls -1dt "$DIR"/* 2>/dev/null | tail -n +"$((KEEP + 1))" | while IFS= read -r victim; do
rm -rf -- "$victim"
done
exit 2
"#
);
let status = host_sudo(&["sh", "-lc", &script])
.await
.context("prune /var/crash")?;
match status.code() {
Some(0) => Ok(()),
Some(2) => {
info!("host fixups: pruned old dumps in /var/crash (keep {KEEP_DUMPS})");
Ok(())
}
code => anyhow::bail!("/var/crash prune exited with {code:?}"),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn sysctl_dropin_carries_the_full_hang_capture_policy() {
for key in [
"kernel.panic = 10",
"kernel.panic_on_oops = 1",
"kernel.hung_task_panic = 1",
"kernel.hardlockup_panic = 1",
] {
assert!(KDUMP_SYSDROPIN.contains(key), "drop-in missing {key}");
}
}
#[test]
fn package_list_is_exactly_the_kdump_rasdaemon_set() {
assert_eq!(HOST_PACKAGES, &["kdump-tools", "kexec-tools", "rasdaemon"]);
}
#[test]
fn crashkernel_param_is_sized_and_unprefixed() {
assert_eq!(CRASHKERNEL_PARAM, "crashkernel=256M");
}
#[test]
fn keep_dumps_is_two() {
assert_eq!(KEEP_DUMPS, 2);
}
}
+7
View File
@@ -55,6 +55,7 @@ mod entropy;
mod federation;
mod fips;
mod health_monitor;
mod host_fixups;
mod host_ip;
mod identity;
mod identity_manager;
@@ -435,6 +436,12 @@ async fn main() -> Result<()> {
// iframe on kiosk nodes (docs/tv-input-iframe-apps.md).
tokio::spawn(bootstrap::ensure_gamepad_keys());
// Host-level fixups (#144 + docs/system-level-ota-design.md): kdump +
// rasdaemon — crash/hardware-error capture delivered to already-deployed
// nodes over the signed binary OTA. Idempotent, non-fatal, background;
// the crashkernel= GRUB edit lands on the next reboot.
tokio::spawn(host_fixups::ensure_host_fixups());
// Mesh access: mirror IPv4-published app ports onto [::] so direct-port
// app URLs (http://[<fips0 ULA>]:<port>) work from the companion.
tokio::spawn(mesh_ports::run_mesh_port_mirror());
+13
View File
@@ -1222,6 +1222,19 @@ impl MeshService {
Ok(dest_prefix)
}
/// True if `contact_id` is reachable over the mesh radio right now — the
/// same peer/twin resolution `peer_dest_prefix` performs, exposed as a
/// cheap bool so RPC handlers can gate radio-only transports (LXMF
/// native image, Reticulum resource transfer) without duplicating the
/// twin-resolution logic. A federation-only contact_id with no matching
/// radio twin returns false here — offering "resource-mesh" or native
/// image to such a peer sends it straight into `peer_dest_prefix`'s
/// "federation-only (no radio twin)" error (picture-send from a
/// federation-only contact, 2026-08-07).
pub async fn has_radio_route(&self, contact_id: u32) -> bool {
self.peer_dest_prefix(contact_id).await.is_ok()
}
/// Split an oversized wire payload into MC-framed base64 chunks and send
/// each via the mesh device. Matches the receive-side reassembly in
/// `mesh/listener/decode.rs::handle_chunked_frame` (header `MCIIXXTT`,
+10 -1
View File
@@ -1746,6 +1746,15 @@ app:
}
}
exempt.sort();
// 28 as of 2026-08-23: the 26 below plus cuprate's two exemptions —
// 18183 (Monero p2p gossip, same reasoning as bitcoin's 8333) and
// 18090 (host mapping for Monero's canonical 18089 restricted RPC,
// upstream's own safe-for-public
// subset that wallets connect to directly as a "remote node" over
// plain HTTP JSON-RPC — same reasoning as electrumx's 50001).
// cuprate's unrestricted RPC (full node control) stays loopback-only
// (auth: local), not in this set.
//
// 26 as of 2026-08-16: the 25 below plus phoenixd 9740, a
// loopback-only JSON API whose own generated http password
// authenticates every request (added with the phoenixd onboarding,
@@ -1762,7 +1771,7 @@ app:
// stage timed out that cycle, so the count here lagged at 17.
assert_eq!(
exempt.len(),
26,
28,
"unauthenticated port set changed — review before updating this count: {exempt:?}"
);
}
+4
View File
@@ -54,6 +54,9 @@ step-by-step guides, and some predate the current implementation.
- [Dual Ecash](dual-ecash-design.md)
- [Hardware Signer](hardware-signer-design.md)
- [Manifest Hooks](manifest-hooks-design.md)
- [Peering & Federation Trust](peering-trust-model.md) — naming/semantics of trust levels vs discovery (#134)
- [kdump + rasdaemon Troubleshooting](kdump-rasdaemon-design.md) — post-mortem and hardware-error capture on nodes (#144)
- [System-Level OTA](system-level-ota-design.md) — how host-level packages/config reach already-deployed nodes
- [Meshroller Integration](meshroller-integration-design.md)
- [Nostr Git Source Hosting](nostr-git-source-hosting.md)
- [Nostr Identity Import](nostr-identity-import-plan.md) · [Nostr Signer Login (research)](nostr-signer-login-research.md)
@@ -86,4 +89,5 @@ file.
## Roadmap & history
- [Roadmap](ROADMAP.md) — where the project is going
- [TODO](TODO.md) — working backlog of unscoped forward-looking items
- [archive/](archive/README.md) — superseded design and status documents, kept for provenance
+52
View File
@@ -0,0 +1,52 @@
# TODO
Working backlog of forward-looking items not yet scoped into a dedicated plan
doc. See [`ROADMAP.md`](ROADMAP.md) for the curated, public-facing direction.
## Dev & build process (priority)
- Formalize the contributor workflow: releases, CI, maintainers, automated
builds, PR/issue flow, branch naming, and reproducible builds.
## Federation & peering
- Peering trust model — define tiers (trusted / public / private / peered)
on top of the existing federation DID trust levels.
- Federation architecture built on the above peering model.
## Distributed git & OTA
- Nostr-hosted git for the alpha (see
[`nostr-git-source-hosting.md`](nostr-git-source-hosting.md)).
- Distributed git beyond the nostr-hosting case.
- Distributed OTA / app delivery.
## Nostr integration
- Nostr signer integration.
## Platform / OS
- Source-availability ISO — define the build/distribution story.
- HW/OS update pipeline.
- Deeper OpenWRT integration.
- GrapheneOS integration — backups, attestation, profiles.
## App ecosystem
- Full pass testing every app in the catalog; expect issues across the board.
- App update strategy — finalize the update policy referenced in
[`app-developer-guide.md`](app-developer-guide.md) (pinned vs. mutable
tags, catalog-vs-disk precedence, rollout/rollback).
- App wishlist — candidates not yet packaged: Cashu wallet, phoenixd.
(CLN is already shipped as `apps/core-lightning`.)
## Access & security
- SSH access strategy — define the access model (keys, rotation, recovery
path, remote-support access).
## Observability
- Capture error logs to troubleshoot customer issues.
- Stats & visualization for traffic, blocked attacks, VPNs, routing.
+131
View File
@@ -0,0 +1,131 @@
# kdump + rasdaemon — post-mortem and hardware-error capture (#144)
Status: IMPLEMENTED (phase 1) — decisions approved 2026-08-30: hang capture ON,
crashkernel=256M, ship the backfill with this release, phase-2 UI deferred.
Delivery: image-recipe (Dockerfile.rootfs, auto-install.sh cmdline) +
`core/archipelago/src/host_fixups.rs` (existing nodes, see
docs/system-level-ota-design.md) + `tests/lifecycle/os-audit.sh` section D.
Owner: node image (image-recipe) + lifecycle gate
Issue: #144 — "Configure kdump and rasdaemon for troubleshooting"
## The problem
When a fleet node hard-locks or a memory stick starts failing, today we get
nothing: a frozen kiosk is power-cycled and the evidence is gone; a DIMM
throwing correctable ECC errors for weeks is invisible until it starts
corrupting things. Two standard kernel mechanisms capture this evidence:
- **kdump** — reserves a small crash kernel at boot; on a kernel panic (or,
configured so, a hang) the running kernel hands the machine over to the
crash kernel, which writes a compressed dump of memory to disk and
reboots. The node comes back by itself *and* leaves a post-mortem.
- **rasdaemon** — a userspace daemon that records hardware error events
(correctable/uncorrectable ECC per DIMM, PCIe AER) from EDAC/sysfs into a
sqlite database: persistent evidence of degrading hardware with no crash
required.
## Facts the design rests on
- Installed-disk layout (auto-install.sh): BIOS boot 1MiB · EFI 512MiB ·
**root ext4 30GiB, unencrypted** · data (rest, LUKS).
- The data partition is LUKS and unlocked late by the node itself — the
crash kernel must never be asked to handle key material.
- The installed system's kernel command line is written by
auto-install.sh:1810 (`GRUB_CMDLINE_LINUX_DEFAULT="quiet splash …"`).
- Packages land via `Dockerfile.rootfs` (trixie) with `systemctl enable`
in the same RUN block (nginx/tor/avahi pattern).
- Kernel cmdline cannot be changed by OTA — it lives in GRUB. Existing
nodes need a backfill step (bootstrap) plus a deliberate reboot.
## Design
### kdump
- **Packages:** `kdump-tools kexec-tools` added to Dockerfile.rootfs.
- **Command line:** append `crashkernel=256M` to
`GRUB_CMDLINE_LINUX_DEFAULT` in auto-install.sh. 256M covers the capture
kernel plus makedumpfile on the fleet's 16–64GB amd64 machines (~1–2% of
RAM reserved, permanently). The arm image (RPi, config.txt boot) is out
of scope for phase 1.
- **Dump target:** `local filesystem /var/crash` — on the unencrypted 30GiB
root, deliberately *not* the encrypted data partition. No key handling
in the crash initramfs, no dependency on the node's own unlock logic.
- **Core collector:** `makedumpfile -l --message-level 1 -d 31`
(compressed, zero/free pages excluded) — a dump lands at roughly 5–15%
of RAM, i.e. ~1–2 GiB on a 16 GiB machine.
- **Retention:** keep the **2 newest** dumps only. A small systemd timer
(or kdump-tools' `KDUMP_POST_SCRIPT`) prunes older vmcores; a full root
partition is already caught by disk_monitor's usage tracking. Two dumps
≈ 4 GiB worst case on 30 GiB root — safe.
- **When to dump — the deliberate trade-off (decision needed):**
- Baseline: dump on real panics (`kernel.panic` path) — no behavioral
change to a wedged node.
- Recommended for this fleet: also enable hang capture
(`kernel.hung_task_panic=1`, hardlockup via NMI watchdog). A kiosk
that hard-locks is useless until power-cycled anyway; converting the
hang into "dump + automatic reboot" turns every freeze into evidence
*and* self-heals the node. Cost: a genuinely-busy-but-alive machine
that trips the watchdog reboots — the threshold is kernel-default
conservative (40s), so this should be rare.
### rasdaemon
- **Packages:** `rasdaemon`; `systemctl enable rasdaemon` in the
Dockerfile.rootfs enable block (same pattern as nginx).
- **Storage:** its default sqlite DB at
`/var/lib/rasdaemon/ras-mc_event.db` on the unencrypted root.
- **Human access today:** `ras-mc-ctl --summary` / `--errors` over SSH.
No UI in phase 1.
### Surfacing (phase 2 — separate follow-up, not in this cut)
A small read-only `system.diagnostics` surface: last-crash timestamp and
vmcore sizes from `/var/crash`, plus ECC error totals per DIMM from the
rasdaemon DB — shown in Settings → System. Deliberately deferred: capture
first, UI once there is something to show and a node in the fleet has
actually produced a dump.
### Existing nodes (phase 1.5 backfill)
The OTA cannot change the bootloader. Bootstrap (which already delivers
fixes to existing nodes) appends `crashkernel=256M` (and the chosen
panic/hang params) to `/etc/default/grub` on machines that don't have it,
and enables `rasdaemon` via the node's package install path. **Takes
effect on the next reboot** — the operator reboots nodes when applying the
release; no special ceremony needed beyond that.
## Testing
- Image: the new packages appear in the ISO; QEMU boot smoke
(build-iso-release.sh stage 5) still green.
- Lifecycle gate additions (bats, archi-dev-box first): `kdump-config show`
reports a loaded crash kernel reservation; `systemctl is-active
rasdaemon`; `/etc/default/grub` carries `crashkernel=`.
- Live drill (once, on archi-dev-box, not in the gate): trigger
`sysrq c` → vmcore appears in `/var/crash`, node reboots itself,
second boot is clean. Keep this manual — it reboots the box.
## Implementation touchpoints
1. `image-recipe/build/auto-installer/Dockerfile.rootfs` — packages +
`systemctl enable rasdaemon`.
2. `image-recipe/build/auto-installer/installer-iso/archipelago/auto-install.sh:1810`
— append `crashkernel=256M` (+ hang params if approved) to
`GRUB_CMDLINE_LINUX_DEFAULT`.
3. `kdump-tools` config: `/etc/default/kdump-tools` (dump target
`/var/crash`, core_collector line, `KDUMP_POST_SCRIPT` or timer for
retention).
4. Bootstrap backfill for existing nodes.
5. `tests/lifecycle` — presence assertions (crash kernel reserved,
rasdaemon active).
## Decisions needed before implementation
1. **Hang capture on or off?** Recommended ON (`hung_task_panic=1` +
NMI watchdog): every hard lockup becomes a dump + self-reboot. OFF
means dumps only on true panics; wedged nodes still need the button.
2. **crashkernel=256M vs 320M** — 256M is the common default for
16–64GB machines; 320M if we expect large io-heavy kernels.
3. **Backfill now or new-installs-only?** Recommended: ship the backfill
with the next release so the whole fleet gains capture on reboot.
4. Phase-2 UI surfacing scope — confirm "later" so phase 1 stays small.
+47
View File
@@ -0,0 +1,47 @@
# Peering & Federation Trust — naming and semantics
Status: TERMINOLOGY SET — records what the code does today (#134).
Deferred: the "don't advertise my peers" opt-out (see §Open questions).
The code is the authority; this doc gives names to the four concepts that
issue #134 showed get conflated in conversation. Where a name changed in
user-facing discussion, the term below is the one to use everywhere
(UI copy, docs, issues, reviews).
## The four concepts
| Term (use this) | What it is | Where it lives |
|---|---|---|
| **Trusted peer** | A node THIS operator invited and verified: bilateral DID challenge over an out-of-band invite code (`federation::sync`, ADR-007). The only level that grants full access. | `TrustLevel::Trusted`, set via `TrustSource::Invite` or `Manual` |
| **Discovered peer** | A peer we learned about from a Trusted peer's advertised list — the transitive merge. Never better than **Observer**: `TRUST IS NOT TRANSITIVE` (sync.rs guard). | `TrustLevel::Observer`, `TrustSource::TransitiveMerge` |
| **Routing hint** | What a Discovered peer actually contributes: an address that lets us route directly over FIPS without a second invite hop. Reachability, not trust. | Observer-level sync + FIPS endpoint records |
| **Peer advertisement** | The act of a Trusted peer sharing its own peer list during sync. This is the *mechanism* #134 observed — a feature, not a leak. | sync.rs merge path |
## The two rules that make it sound
1. **Trust requires an operator decision, always traceable.** Every trust
level carries a `TrustSource`. Only a minted invite (or an explicit
operator change) can produce `Trusted`; uninvited joins and transitive
merges are hard-capped at `Observer` — a peer can never expand our
trusted set on its own authority.
2. **Discovery is transitive; trust is not.** Seeing more nodes through a
Trusted peer is expected and useful (routing). Granting those nodes
anything is an operator action, never automatic.
## Why a Trusted peer advertising its list is by design
Without advertisement, every new node needs a direct invite from every node
that wants to reach it — the invite graph becomes the routing bottleneck
AdDR-007 set out to remove. With it, one invite makes a node *reachable* to
the trusted set (routing hints), while *authorization* still requires each
operator's own invite. Reachability ≠ access.
## Open questions (deferred, tracked in #134)
- **"Don't advertise my peers"** — an operator privacy toggle suppressing
peer advertisement during sync. Small code change, real design questions:
it hides peers who may WANT discovery, and it degrades the routing benefit
for every node trusting you. Needs a product decision, not just code.
- **Tier vocabulary in the UI** — whether to surface "Observer" as such or
a friendlier term ("Connected"/"Visible") — part of the TODO.md peering
trust-model item.
+82
View File
@@ -0,0 +1,82 @@
# System-Level OTA — host fixups
Status: Implemented (first payload shipped alongside this doc)
Owner: `core/archipelago/src/host_fixups.rs`
Related: docs/kdump-rasdaemon-design.md (first payload), CLAUDE.md invariants
## The problem
The binary OTA updates the node's own software, and the signed app catalog
updates apps. But the **host OS** — Debian packages, kernel parameters,
system services — previously moved only through ISO re-installs. A node
deployed a year ago can be running today's node software on a host that
never gained anything the image learned since. Issue #99 (missing polkit
rule on old nodes) and the audio-stack heal were each hand-carved
one-off bootstrap repairs; there was no general channel and no stated
policy for touching the host from the node.
## The mechanism
`host_fixups::ensure_host_fixups()` — spawned from `main.rs` at startup
alongside the other `ensure_*` heals, in the background, best-effort:
1. **Dev-box guard** — skip when `/home/archipelago/archy` is a symlink
(contributor checkout) and when there's no dpkg (non-Debian host).
2. **Packages** — install only what's missing, from a curated, in-code
list (`HOST_PACKAGES`), `apt-get install` first, one `apt-get update`
retry, both under timeout, never fatal (offline/locked-dpkg nodes
converge on a later boot).
3. **Configuration** — idempotent per-concern helpers writing root-owned
config (via the existing `host_sudo` path): sysctl drop-ins, service
defaults, GRUB cmdline, service enablement.
4. **Reporting** — every step logs what it did; failures log warnings and
move on. A host fixup must never be able to stop the node from starting.
### Why embedded-in-the-binary rather than fetched
Same reasoning as the tor-helper (`bootstrap.rs`): the signed binary OTA
is the only authenticated delivery channel every node already trusts and
pulls on schedule. Fixups compiled into the binary travel with a version,
are reviewable in git, and can't be served to a subset of the fleet.
## Policy — what may travel this channel
| May | May not |
|---|---|
| Specific, pinned packages the node needs (kdump-tools, rasdaemon, …) | `dist-upgrade` or silent kernel/libc swaps — regular Debian upgrades stay with the operator |
| Kernel *parameters* via GRUB/sysctl — with the next-reboot caveat logged loudly | Anything requiring a secret, or touching LUKS key material |
| Service enablement + config the image also bakes in | Divergence: the ISO must converge to the SAME end state so fresh installs are a no-op |
| Small, reviewable, per-concern Rust functions with tests | Shell-script-of-things payloads beyond a single concern |
The rule: **the ISO and the fixup must express the same intent twice,
in reviewable places** — Dockerfile.rootfs/auto-install.sh for fresh
installs, `host_fixups.rs` for the deployed fleet. A change that lands in
one and not the other is a bug.
## Kernel cmdline caveat
`crashkernel=` (and any future `hugepages=`-style reservation) only takes
effect at boot: the fixup writes `/etc/default/grub` + `update-grub` and
logs `takes effect on the NEXT reboot`. Operators reboot nodes when
applying releases; no special ceremony is required beyond that, but the
lifecycle gate grades this state honestly (WARN for written-but-not-yet-
rebooted, FAIL for never-written — see `tests/lifecycle/os-audit.sh`
section D).
## Verification story
- Unit tests pin the policy constants and script shapes
(`host_fixups` tests in `core/archipelago`).
- `tests/lifecycle/os-audit.sh` section D asserts the end state on a real
node (config present, crashkernel reserved or pending reboot, hang
policy live, rasdaemon active).
- The lifecycle gate runs on archi-dev-box per release; the QEMU ISO
smoke covers fresh installs.
## Future payloads (candidates, not commitments)
- `unattended-upgrades` posture + a default-deny host nftables ruleset
(the §F hardening-plan item — needs its own design first).
- Host firewall rules for mesh/WG ports.
- Chronic: anything the image learns post-deploy that old nodes must
converge on (the polkit and audio precedents, formalized).
@@ -567,6 +567,33 @@ RUN mkdir -p /etc/polkit-1/rules.d && \
> /etc/polkit-1/rules.d/49-archipelago-networkmanager.rules && \
chmod 644 /etc/polkit-1/rules.d/49-archipelago-networkmanager.rules
# kdump + rasdaemon (#144, docs/kdump-rasdaemon-design.md): crash dumps and
# hardware-error capture on the host. Packages + config are baked in for fresh
# installs; the binary's host_fixups module delivers the identical end state to
# already-deployed nodes over OTA (idempotent no-op here once applied).
RUN set -eu; \
apt-get update; \
apt-get install -y --no-install-recommends kdump-tools kexec-tools rasdaemon; \
apt-get clean; rm -rf /var/lib/apt/lists/*; \
CONF=/etc/default/kdump-tools; \
sed -i 's|^#\?USE_KDUMP=.*|USE_KDUMP="1"|' "$CONF"; \
grep -q '^KDUMP_COREDIR=' "$CONF" \
&& sed -i 's|^KDUMP_COREDIR=.*|KDUMP_COREDIR="/var/crash"|' "$CONF" \
|| printf '\nKDUMP_COREDIR="/var/crash"\n' >> "$CONF"; \
grep -q '^CORE_COLLECTOR=' "$CONF" \
&& sed -i 's|^CORE_COLLECTOR=.*|CORE_COLLECTOR="makedumpfile -l --message-level 1 -d 31"|' "$CONF" \
|| printf '\nCORE_COLLECTOR="makedumpfile -l --message-level 1 -d 31"\n' >> "$CONF"; \
printf '%s\n' \
'# Archipelago kdump policy (#144). A wedged kiosk is useless until someone' \
'# power-cycles it — capture the evidence, then reboot by itself. Dumps land in' \
'# /var/crash (see docs/kdump-rasdaemon-design.md); keep-2 pruning is done by' \
'# the host fixup pass, not a timer.' \
'kernel.panic = 10' \
'kernel.panic_on_oops = 1' \
'kernel.hung_task_panic = 1' \
'kernel.hardlockup_panic = 1' \
> /etc/sysctl.d/99-archipelago-kdump.conf
# Enable services
RUN systemctl enable NetworkManager || true && \
systemctl enable polkit || systemctl enable polkit.service || true && \
@@ -580,7 +607,9 @@ RUN systemctl enable NetworkManager || true && \
systemctl enable archipelago-update.timer || true && \
systemctl enable archipelago-doctor.timer || true && \
systemctl enable archipelago-tor-helper.path || true && \
systemctl enable nostr-relay || true
systemctl enable nostr-relay || true && \
systemctl enable rasdaemon || true && \
systemctl enable kdump-tools || true
# archipelago-fips.service + archipelago-wg.service + archipelago-wg-address.service
# stay installed and enabled. They all use `ConditionPathExists=` on their
# respective seed-derived key files, so on a fresh pre-onboarding boot
@@ -3715,7 +3744,7 @@ if [ -d "$BOOT_MEDIA/archipelago/plymouth-theme" ]; then
ln -sf /usr/share/plymouth/themes/archipelago/archipelago.plymouth \
/mnt/target/etc/alternatives/default.plymouth 2>/dev/null || true
# Configure clean boot: splash, suppress kernel noise, hide cursor
sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="quiet splash loglevel=0 rd.systemd.show_status=false vt.global_cursor_default=0 acpi=force"/' \
sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="quiet splash loglevel=0 rd.systemd.show_status=false vt.global_cursor_default=0 acpi=force crashkernel=256M"/' \
/mnt/target/etc/default/grub 2>/dev/null || true
echo " Installed Archipelago Plymouth theme on target"
fi
File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 6.2 KiB

+4 -4
View File
@@ -52,13 +52,13 @@
{
"id": "btcpay-server",
"title": "BTCPay Server",
"version": "2.4.2",
"version": "2.4.3",
"description": "Self-hosted Bitcoin payment processor. Accept Bitcoin payments without intermediaries.",
"icon": "/assets/img/app-icons/btcpay-server.png",
"author": "BTCPay Server Foundation",
"category": "commerce",
"tier": "core",
"dockerImage": "docker.io/btcpayserver/btcpayserver:2.4.2",
"dockerImage": "docker.io/btcpayserver/btcpayserver:2.4.3",
"repoUrl": "https://github.com/btcpayserver/btcpayserver",
"requires": [
"bitcoin-knots"
@@ -378,7 +378,7 @@
"icon": "/assets/img/app-icons/pine.svg",
"author": "Archipelago",
"category": "home",
"dockerImage": "docker.io/library/nginx:1.31.3-alpine",
"dockerImage": "docker.io/library/nginx:1.31.4-alpine",
"repoUrl": "https://github.com/rhasspy/wyoming"
},
{
@@ -464,7 +464,7 @@
"author": "NetBird",
"category": "networking",
"tier": "recommended",
"dockerImage": "docker.io/library/nginx:1.31.3-alpine",
"dockerImage": "docker.io/library/nginx:1.31.4-alpine",
"repoUrl": "https://github.com/netbirdio/netbird",
"containerConfig": {
"ports": [
+1 -1
View File
@@ -181,7 +181,7 @@ watch(() => appStore.isAuthenticated, (authenticated) => {
startRemoteRelay()
} else {
messageToast.stopPolling()
toastMessage.value = { show: false, text: '', fromPubkey: '' }
toastMessage.value = { show: false, text: '', fromPubkey: '', contactId: null }
screensaverStore.clearInactivityTimer()
screensaverStore.deactivate()
stopRemoteRelay()
+17 -4
View File
@@ -50,6 +50,7 @@ const showRevealModal = ref(false)
const revealPassword = ref('')
const revealCode = ref('')
const revealPassphrase = ref('')
const showRevealPassphrase = ref(false)
const revealing = ref(false)
const revealError = ref('')
const revealedWords = ref<string[]>([])
@@ -60,6 +61,7 @@ function openReveal() {
revealPassword.value = ''
revealCode.value = ''
revealPassphrase.value = ''
showRevealPassphrase.value = false
revealError.value = ''
revealedWords.value = []
showRevealModal.value = true
@@ -83,7 +85,17 @@ async function submitReveal() {
// to set up a backup that now exists.
void loadStatus()
} catch (e: unknown) {
revealError.value = e instanceof Error ? e.message : 'Failed to reveal the ecash phrase'
const message = e instanceof Error ? e.message : 'Failed to reveal the ecash phrase'
// Most operators used their login password as the backup passphrase. Do
// not confront everyone with an unexplained third credential up front;
// disclose it only when the authenticated password could not decrypt the
// node seed and a distinct setup-time passphrase may actually exist.
if (!status.value?.active && /could not decrypt the saved seed/i.test(message)) {
showRevealPassphrase.value = true
revealError.value = 'Your login password did not unlock the saved seed. Enter the separate backup passphrase you chose during setup.'
} else {
revealError.value = message
}
} finally {
revealing.value = false
}
@@ -95,6 +107,7 @@ function closeReveal() {
revealPassword.value = ''
revealCode.value = ''
revealPassphrase.value = ''
showRevealPassphrase.value = false
}
async function copyRevealedWords() {
@@ -376,9 +389,9 @@ async function restoreFromPhrase() {
<label class="block text-xs text-white/60 mb-1">2FA code <span class="text-white/30">(if enabled)</span></label>
<input v-model="revealCode" inputmode="numeric" autocomplete="one-time-code" class="w-full px-3 py-2 rounded-lg bg-white/5 border border-white/10 text-white text-sm font-mono tracking-widest focus:outline-none focus:border-white/30" placeholder="123456" />
</div>
<div v-if="!status?.active">
<label class="block text-xs text-white/60 mb-1">Backup passphrase <span class="text-white/30">(only if different from password)</span></label>
<input v-model="revealPassphrase" type="password" class="w-full px-3 py-2 rounded-lg bg-white/5 border border-white/10 text-white text-sm focus:outline-none focus:border-white/30" placeholder="Leave blank to use password" />
<div v-if="showRevealPassphrase">
<label class="block text-xs text-white/60 mb-1">Separate backup passphrase</label>
<input v-model="revealPassphrase" type="password" autocomplete="off" autofocus class="w-full px-3 py-2 rounded-lg bg-white/5 border border-white/10 text-white text-sm focus:outline-none focus:border-white/30" placeholder="Passphrase chosen during setup" />
</div>
<p v-if="revealError" class="text-xs text-red-300 bg-red-500/10 border border-red-400/20 rounded-lg px-3 py-2">{{ revealError }}</p>
<div class="flex gap-2 pt-1">
@@ -0,0 +1,52 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { flushPromises, mount, type VueWrapper } from '@vue/test-utils'
vi.mock('@/api/rpc-client', () => ({
rpcClient: { call: vi.fn() },
}))
import { rpcClient } from '@/api/rpc-client'
import EcashSeedBackup from '../EcashSeedBackup.vue'
let wrapper: VueWrapper | null = null
describe('EcashSeedBackup reveal credentials (#127)', () => {
beforeEach(() => {
document.body.innerHTML = ''
vi.clearAllMocks()
})
afterEach(() => {
wrapper?.unmount()
wrapper = null
document.body.innerHTML = ''
})
it('asks for a separate backup passphrase only after password decryption fails', async () => {
vi.mocked(rpcClient.call)
.mockResolvedValueOnce({
active: false,
source: null,
can_activate: true,
derivable_from_node_seed: true,
})
.mockRejectedValueOnce(new Error(
'Could not decrypt the saved seed. If you set a separate backup passphrase during setup, enter that passphrase.',
))
wrapper = mount(EcashSeedBackup, { attachTo: document.body })
await flushPromises()
await wrapper.get('button').trigger('click')
expect(document.body.textContent).not.toContain('Separate backup passphrase')
const password = document.body.querySelector<HTMLInputElement>('input[autocomplete="current-password"]')!
password.value = 'login-password'
password.dispatchEvent(new Event('input', { bubbles: true }))
document.body.querySelector('form')!.dispatchEvent(new Event('submit', { bubbles: true, cancelable: true }))
await flushPromises()
expect(document.body.textContent).toContain('Separate backup passphrase')
expect(document.body.textContent).toContain('Your login password did not unlock the saved seed')
expect(document.body.querySelector('input[placeholder="Passphrase chosen during setup"]')).not.toBeNull()
})
})
@@ -0,0 +1,150 @@
import { describe, it, expect, beforeEach, vi } from 'vitest'
import { mount } from '@vue/test-utils'
import { defineComponent, nextTick } from 'vue'
// Controllable doubles shared between the hoisted block and the mock
// factories. Plain holders — each test writes to them before importing the
// composable under a fresh module registry (the watcher keeps a module-level
// `firedThisSession` session guard, so every case needs its own module).
const state = vi.hoisted(() => ({
packages: {} as Record<string, unknown>,
goalStatus: 'in-progress',
goalProgress: {} as Record<string, { completedSteps: string[] }>,
toastAction: vi.fn(),
routerPush: vi.fn(),
// Re-bound every time the useBitcoinSync factory is (re)evaluated; holds the
// exact refs the freshly imported composable watches.
syncRefs: null as null | { synced: { value: boolean }; loaded: { value: boolean } },
}))
vi.mock('@/composables/useBitcoinSync', async () => {
const { ref } = await import('vue')
const synced = ref(false)
const loaded = ref(false)
state.syncRefs = { synced, loaded }
return {
bitcoinSynced: synced,
bitcoinSyncLoaded: loaded,
acquireBitcoinSync: () => () => {},
}
})
vi.mock('@/stores/goals', () => ({
useGoalStore: () => ({
getGoalStatus: () => state.goalStatus,
progress: state.goalProgress,
}),
}))
vi.mock('@/stores/app', () => ({
useAppStore: () => ({
get packages() {
return state.packages
},
}),
}))
vi.mock('@/composables/useToast', () => ({
useToast: () => ({ action: state.toastAction }),
}))
vi.mock('vue-router', () => ({
useRouter: () => ({ push: state.routerPush }),
}))
/**
* Fresh module registry → fresh `firedThisSession`, then mount the composable
* inside a real component so its watchers live in a proper effect scope.
*/
async function mountWatcher() {
vi.resetModules()
const { useIbdFinishWatcher } = await import('../useIbdFinishWatcher')
const Host = defineComponent({
setup() {
useIbdFinishWatcher()
return () => null
},
})
return mount(Host)
}
/** Drive a real unsynced→synced transition through the mocked sync refs. */
async function completeSync() {
const refs = state.syncRefs!
refs.loaded.value = true
refs.synced.value = false // the watcher must observe unsynced at least once
await nextTick()
refs.synced.value = true
await nextTick()
await nextTick()
}
describe('useIbdFinishWatcher', () => {
beforeEach(() => {
state.packages = {}
state.goalStatus = 'in-progress'
state.goalProgress = {}
state.toastAction.mockClear()
state.routerPush.mockClear()
})
it('says to install LND next when Lightning is not installed yet (#143)', async () => {
// Bitcoin synced mid-goal, but the goal's install-LND step is still
// pending: the on-chain wallet lives in LND, so "fund your wallet" would
// promise a flow that cannot work yet.
state.packages = { 'bitcoin-knots': { state: 'running' } }
const wrapper = await mountWatcher()
await completeSync()
expect(state.toastAction).toHaveBeenCalledTimes(1)
const [message, opts] = state.toastAction.mock.calls[0]!
expect(message).toBe(
"Bitcoin is fully synced — next, install Lightning (LND) to get your node's on-chain wallet.",
)
expect(opts.label).toBe('Finish setup')
opts.onClick()
// "Finish setup" lands on the goal wizard, whose active step is the
// pending install-LND one — the correct next action.
expect(state.routerPush).toHaveBeenCalledWith('/dashboard/goals/open-a-shop')
wrapper.unmount()
})
it('says to fund the wallet when LND is already installed', async () => {
state.packages = { 'bitcoin-knots': { state: 'running' }, lnd: { state: 'running' } }
const wrapper = await mountWatcher()
await completeSync()
expect(state.toastAction).toHaveBeenCalledTimes(1)
const [message, opts] = state.toastAction.mock.calls[0]!
expect(message).toBe(
'Bitcoin is fully synced — you can now fund your wallet and open your Lightning channel.',
)
expect(opts.label).toBe('Finish setup')
opts.onClick()
expect(state.routerPush).toHaveBeenCalledWith('/dashboard/goals/open-a-shop')
wrapper.unmount()
})
it('stays silent when no Lightning goal is in progress', async () => {
state.goalStatus = 'not-started'
const wrapper = await mountWatcher()
await completeSync()
expect(state.toastAction).not.toHaveBeenCalled()
wrapper.unmount()
})
it('stays silent when the chain was already synced at page load', async () => {
// A node that's already synced never shows unsynced this session, so the
// toast must not fire (it only marks real IBD-completion transitions).
const wrapper = await mountWatcher()
const refs = state.syncRefs!
refs.loaded.value = true
refs.synced.value = true
await nextTick()
await nextTick()
expect(state.toastAction).not.toHaveBeenCalled()
wrapper.unmount()
})
})
@@ -1,4 +1,5 @@
import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest'
import { createPinia, setActivePinia } from 'pinia'
const mockPush = vi.fn()
@@ -9,6 +10,7 @@ vi.mock('vue-router', () => ({
vi.mock('@/api/rpc-client', () => ({
rpcClient: {
getReceivedMessages: vi.fn(),
call: vi.fn(),
},
}))
@@ -21,13 +23,16 @@ describe('useMessageToast', () => {
beforeEach(() => {
vi.clearAllMocks()
vi.useFakeTimers()
localStorage.clear()
setActivePinia(createPinia())
vi.mocked(rpcClient.call).mockResolvedValue({ messages: [], count: 0 })
// Reset shared singleton state
const toast = useMessageToast()
toast.stopPolling()
toast.receivedMessages.value = []
toast.lastMessageCount.value = 0
toast.loadingMessages.value = false
toast.toastMessage.value = { show: false, text: '', fromPubkey: '' }
toast.toastMessage.value = { show: false, text: '', fromPubkey: '', contactId: null }
})
afterEach(() => {
@@ -143,9 +148,43 @@ describe('useMessageToast', () => {
expect(toast.unreadCount.value).toBe(0)
})
it('shows a radio-mesh toast and deep-links to its contact', async () => {
const toast = useMessageToast()
mockedRpc.getReceivedMessages.mockResolvedValue({ messages: [] })
// Initialize an empty node, then deliver its first Meshtastic message.
await toast.loadReceivedMessages()
vi.mocked(rpcClient.call).mockResolvedValueOnce({
messages: [{
id: 1,
direction: 'received',
peer_contact_id: 42,
peer_name: 'Alice',
plaintext: 'Over LoRa',
timestamp: '2026-01-01',
delivered: true,
encrypted: true,
transport: 'meshtastic',
}],
count: 1,
})
await toast.loadReceivedMessages()
expect(toast.toastMessage.value).toMatchObject({
show: true,
text: 'Over LoRa',
contactId: 42,
})
toast.dismissToastAndOpenMessages()
expect(mockPush).toHaveBeenCalledWith({
path: '/dashboard/mesh',
query: { contact: '42' },
})
})
it('dismissToastAndOpenMessages clears toast and navigates', () => {
const toast = useMessageToast()
toast.toastMessage.value = { show: true, text: 'New message', fromPubkey: '' }
toast.toastMessage.value = { show: true, text: 'New message', fromPubkey: '', contactId: null }
toast.dismissToastAndOpenMessages()
expect(toast.toastMessage.value.show).toBe(false)
@@ -2,6 +2,7 @@ import { computed, watch, watchEffect, onUnmounted } from 'vue'
import { useRouter } from 'vue-router'
import { GOALS } from '@/data/goals'
import { useGoalStore } from '@/stores/goals'
import { useAppStore } from '@/stores/app'
import { useToast } from '@/composables/useToast'
import {
acquireBitcoinSync,
@@ -20,6 +21,7 @@ let firedThisSession = false
*/
export function useIbdFinishWatcher() {
const goalStore = useGoalStore()
const appStore = useAppStore()
const router = useRouter()
const toast = useToast()
@@ -68,8 +70,16 @@ export function useIbdFinishWatcher() {
const goalId = pendingLightningGoalId.value
if (!goalId) return
firedThisSession = true
// The on-chain wallet lives in LND, not Bitcoin Core — the address the
// fund flow shows comes from `lnd.newaddress`. While the goal's
// install-LND step is still pending, "fund your wallet" would point at
// something that doesn't exist yet, so the toast names the actual next
// step instead (issue #143).
const lndInstalled = Object.keys(appStore.packages).includes('lnd')
toast.action(
'Bitcoin is fully synced — you can now fund your wallet and open your Lightning channel.',
lndInstalled
? 'Bitcoin is fully synced — you can now fund your wallet and open your Lightning channel.'
: "Bitcoin is fully synced — next, install Lightning (LND) to get your node's on-chain wallet.",
{
label: 'Finish setup',
onClick: () => { router.push(`/dashboard/goals/${goalId}`) },
+41 -6
View File
@@ -1,6 +1,7 @@
import { ref, computed } from 'vue'
import { useRouter } from 'vue-router'
import { rpcClient } from '@/api/rpc-client'
import { useMeshStore } from '@/stores/mesh'
export interface ReceivedMessage {
from_pubkey: string
@@ -14,11 +15,19 @@ const MESSAGE_POLL_INTERVAL = 30000 // 30s
const receivedMessages = ref<ReceivedMessage[]>([])
const lastMessageCount = ref(0)
const loadingMessages = ref(false)
const toastMessage = ref<{ show: boolean; text: string; fromPubkey: string }>({ show: false, text: '', fromPubkey: '' })
type MessageToast = {
show: boolean
text: string
fromPubkey: string
contactId: number | null
}
const emptyToast = (): MessageToast => ({ show: false, text: '', fromPubkey: '', contactId: null })
const toastMessage = ref<MessageToast>(emptyToast())
let pollTimer: ReturnType<typeof setInterval> | null = null
export function useMessageToast() {
const router = useRouter()
const mesh = useMeshStore()
const unreadCount = computed(() =>
Math.max(0, receivedMessages.value.length - lastMessageCount.value)
@@ -40,6 +49,7 @@ export function useMessageToast() {
// Only deep-link to a specific chat when it's a single new message
// from one sender; otherwise open the mesh list.
fromPubkey: newCount === 1 ? (latest?.from_pubkey ?? '') : '',
contactId: null,
}
lastMessageCount.value = msgs.length
} else {
@@ -55,6 +65,26 @@ export function useMessageToast() {
} finally {
loadingMessages.value = false
}
// Federation messages and radio-mesh messages use separate backend
// queues. Poll the mesh store too so Meshtastic/MeshCore/Reticulum
// arrivals produce the same app-wide toast. fetchMessages returns only
// the newly-unread batch computed from its durable per-contact watermark.
const newMeshMessages = await mesh.fetchMessages()
if (newMeshMessages.length > 0) {
const latest = newMeshMessages[newMeshMessages.length - 1]!
const oneConversation = newMeshMessages.every(
msg => msg.peer_contact_id === latest.peer_contact_id
)
toastMessage.value = {
show: true,
text: newMeshMessages.length === 1
? latest.plaintext
: `${newMeshMessages.length} new messages`,
fromPubkey: '',
contactId: oneConversation ? latest.peer_contact_id : null,
}
}
}
function isAuthenticated(): boolean {
@@ -86,16 +116,21 @@ export function useMessageToast() {
}
function dismissToastAndOpenMessages() {
const peer = toastMessage.value.fromPubkey
toastMessage.value = { show: false, text: '', fromPubkey: '' }
const { fromPubkey: peer, contactId } = toastMessage.value
toastMessage.value = emptyToast()
markAsRead()
// Open the specific conversation when we know the sender; else the mesh list.
router.push(peer ? { path: '/dashboard/mesh', query: { peer } } : '/dashboard/mesh')
// Open the exact radio conversation by contact id, or the federation
// conversation by pubkey. Multiple conversations fall back to the list.
if (contactId !== null) {
router.push({ path: '/dashboard/mesh', query: { contact: String(contactId) } })
} else {
router.push(peer ? { path: '/dashboard/mesh', query: { peer } } : '/dashboard/mesh')
}
}
// Dismiss the toast without navigating (the close icon).
function closeToast() {
toastMessage.value = { show: false, text: '', fromPubkey: '' }
toastMessage.value = emptyToast()
}
return {
@@ -0,0 +1,66 @@
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { createPinia, setActivePinia } from 'pinia'
vi.mock('@/api/rpc-client', () => ({
rpcClient: { call: vi.fn() },
}))
import { rpcClient } from '@/api/rpc-client'
import { useMeshStore, type MeshMessage } from '../mesh'
const message = (id: number, contact = 7): MeshMessage => ({
id,
direction: 'received',
peer_contact_id: contact,
peer_name: 'Alice',
plaintext: `message ${id}`,
timestamp: `2026-01-${String(id).padStart(2, '0')}`,
delivered: true,
encrypted: true,
transport: 'meshtastic',
})
function reply(messages: MeshMessage[]) {
vi.mocked(rpcClient.call).mockResolvedValueOnce({ messages, count: messages.length })
}
describe('mesh unread persistence', () => {
beforeEach(() => {
localStorage.clear()
setActivePinia(createPinia())
vi.clearAllMocks()
})
it('does not swallow the first message after initializing with empty history', async () => {
const store = useMeshStore()
reply([])
expect(await store.fetchMessages()).toEqual([])
expect(localStorage.getItem('archipelago.mesh.last-seen.v1')).toBe('{}')
reply([message(1)])
expect(await store.fetchMessages()).toEqual([message(1)])
expect(store.unreadCounts[7]).toBe(1)
})
it('keeps read messages read across a page refresh', async () => {
const firstPage = useMeshStore()
reply([message(1), message(2)])
await firstPage.fetchMessages() // migration seeds existing history as read
firstPage.markChatRead(7)
setActivePinia(createPinia()) // simulate a full page/store reload
const refreshedPage = useMeshStore()
reply([message(1), message(2), message(3)])
const newlyUnread = await refreshedPage.fetchMessages()
expect(newlyUnread.map(m => m.id)).toEqual([3])
expect(refreshedPage.unreadCounts[7]).toBe(1)
refreshedPage.markChatRead(7)
setActivePinia(createPinia())
const readAgain = useMeshStore()
reply([message(1), message(2), message(3)])
expect(await readAgain.fetchMessages()).toEqual([])
expect(readAgain.totalUnread).toBe(0)
})
})
+30 -7
View File
@@ -288,12 +288,27 @@ export const useMeshStore = defineStore('mesh', () => {
// are safe watermarks: the backend allocates them monotonically and
// restores the counter as max(persisted)+1 across restarts.
const LAST_SEEN_KEY = 'archipelago.mesh.last-seen.v1'
const lastSeenId = ref<Record<number, number>>(
JSON.parse(localStorage.getItem(LAST_SEEN_KEY) || '{}') as Record<number, number>
)
let storedLastSeen = localStorage.getItem(LAST_SEEN_KEY)
function parseLastSeen(raw: string | null): Record<number, number> {
if (!raw) return {}
try {
const parsed = JSON.parse(raw)
if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) {
return parsed as Record<number, number>
}
} catch {
// Treat corrupt browser state like a first run and safely reseed it.
}
storedLastSeen = null
return {}
}
const lastSeenId = ref<Record<number, number>>(parseLastSeen(storedLastSeen))
// First run after this feature ships: treat existing history as seen so
// nobody gets a wall of phantom badges for months-old messages.
let seedLastSeenFromHistory = localStorage.getItem(LAST_SEEN_KEY) === null
// nobody gets a wall of phantom badges for months-old messages. Complete
// this initialization even when history is empty; otherwise the first real
// message to arrive on a brand-new node is mistaken for old history and its
// notification is silently swallowed.
let seedLastSeenFromHistory = storedLastSeen === null
function persistLastSeen() {
localStorage.setItem(LAST_SEEN_KEY, JSON.stringify(lastSeenId.value))
}
@@ -481,17 +496,18 @@ export const useMeshStore = defineStore('mesh', () => {
}
}
async function fetchMessages(limit?: number) {
async function fetchMessages(limit?: number): Promise<MeshMessage[]> {
try {
const res = await rpcClient.call<{ messages: MeshMessage[]; count: number }>({
method: 'mesh.messages',
params: limit ? { limit } : {},
dedup: true,
})
if (seedLastSeenFromHistory && res.messages.length > 0) {
if (seedLastSeenFromHistory) {
for (const m of res.messages) {
if (m.direction === 'received') advanceLastSeen(m.peer_contact_id, m.id)
}
// Persist even an empty object as the initialization sentinel.
persistLastSeen()
seedLastSeenFromHistory = false
}
@@ -520,8 +536,15 @@ export const useMeshStore = defineStore('mesh', () => {
messages.value = res.messages
// Extract node positions from coordinate messages
updateNodePositionsFromMessages(res.messages)
// The app-wide notification poll uses this exact batch, rather than a
// session message-count delta, so one arrival can never resurrect old
// messages as "11 unread" after a refresh.
return newMsgs.filter(msg => !(
viewingChatIds.value.includes(msg.peer_contact_id) && viewingAtBottom.value
))
} catch (err: unknown) {
error.value = err instanceof Error ? err.message : 'Failed to fetch mesh messages'
return []
}
}
+7 -1
View File
@@ -565,7 +565,13 @@ function armMeshLive() {
// match an entry in mesh.peers, so without this fallback the deep-link
// silently failed and just landed on the bare mesh page every time.
const targetPeer = typeof route.query.peer === 'string' ? route.query.peer : ''
if (targetPeer) {
const targetContact = typeof route.query.contact === 'string'
? Number(route.query.contact)
: NaN
if (Number.isInteger(targetContact)) {
const match = mesh.peers.find(p => p.contact_id === targetContact)
if (match) openChat(match)
} else if (targetPeer) {
const match = mesh.peers.find(
(p) => p.pubkey_hex === targetPeer || p.did === targetPeer
)
@@ -51,6 +51,7 @@ export const GENERATED_APP_TITLES: Record<string, string> = {
"botfights": "BotFights",
"btcpay-server": "BTCPay Server",
"core-lightning": "Core Lightning (CLN)",
"cuprate": "Cuprate",
"did-wallet": "Web5 DID Wallet",
"electrs-ui": "Electrs UI",
"electrumx": "ElectrumX",
@@ -26,7 +26,8 @@
:class="tierLabel === 'core' ? 'tier-badge-core' : 'tier-badge-recommended'"
>{{ tierLabel }}</span>
</h3>
<p class="text-sm text-white/60">{{ app.version ? $ver(app.version) : 'latest' }}</p>
<p v-if="!isMultiVersion" class="text-sm text-white/60">{{ app.version ? $ver(app.version) : 'latest' }}</p>
<p v-else class="text-sm text-white/60">Choose version when installing</p>
<p v-if="app.author" class="text-xs text-white/50 mt-1">by {{ app.author }}</p>
</div>
</div>
@@ -175,7 +176,7 @@
<script setup lang="ts">
import { computed } from 'vue'
import { useI18n } from 'vue-i18n'
import type { MarketplaceApp, InstallProgress } from './marketplaceData'
import { MULTI_VERSION_APP_IDS, type MarketplaceApp, type InstallProgress } from './marketplaceData'
import { DEFAULT_APP_ICON } from '@/views/apps/appsConfig'
const { t } = useI18n()
@@ -200,6 +201,8 @@ defineEmits<{
launch: [app: MarketplaceApp]
}>()
const isMultiVersion = computed(() => MULTI_VERSION_APP_IDS.has(props.app.id))
const signatureLabel = computed(() => {
switch (props.app.signature?.status) {
case 'valid': return 'signed'
@@ -15,7 +15,7 @@ const app: MarketplaceApp = {
source: 'community',
}
function mountCard(installed: boolean, installBlockedReason?: string) {
function mountCard(installed: boolean, installBlockedReason?: string, appOverride: MarketplaceApp = app) {
const i18n = createI18n({
legacy: false,
locale: 'en',
@@ -29,7 +29,7 @@ function mountCard(installed: boolean, installBlockedReason?: string) {
return mount(MarketplaceAppCard, {
props: {
app,
app: appOverride,
index: 0,
stagger: false,
installed,
@@ -65,4 +65,16 @@ describe('MarketplaceAppCard', () => {
expect(wrapper.text()).toContain('Requires a full archive Bitcoin node before install.')
expect(wrapper.text()).toContain('Bitcoin Pruned')
})
it('does not present one catalog version as definitive for multi-version apps (#129)', () => {
const wrapper = mountCard(false, undefined, {
...app,
id: 'bitcoin-core',
title: 'Bitcoin Core',
version: '28.4.0',
})
expect(wrapper.text()).toContain('Choose version when installing')
expect(wrapper.text()).not.toContain('28.4')
})
})
@@ -57,6 +57,12 @@ export interface InstallProgress {
attempt: number
}
/** Apps that ask for their concrete version in InstallVersionModal. Their
* store tiles deliberately omit a single catalog version: showing “v28.4”
* there implies that is the only version immediately before asking the user
* to choose a different one. */
export const MULTI_VERSION_APP_IDS = new Set(['bitcoin-knots', 'bitcoin-core'])
/** Archipelago app registry — all app images are mirrored here */
const REGISTRY = 'source.archipelago-foundation.org/lfg2025'
@@ -0,0 +1,37 @@
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { createPinia } from 'pinia'
import { flushPromises, mount } from '@vue/test-utils'
vi.mock('vue-router', () => ({
useRouter: () => ({ push: vi.fn() }),
}))
vi.mock('@/api/rpc-client', () => ({
rpcClient: { call: vi.fn() },
}))
import { rpcClient } from '@/api/rpc-client'
import OpenWrtGateway from './OpenWrtGateway.vue'
describe('OpenWrtGateway stale cached router recovery (#103)', () => {
beforeEach(() => {
vi.clearAllMocks()
sessionStorage.clear()
})
it('offers reconfiguration when the saved router can no longer connect', async () => {
vi.mocked(rpcClient.call).mockRejectedValue(new Error('Connection timed out'))
const wrapper = mount(OpenWrtGateway, {
global: { plugins: [createPinia()] },
})
await flushPromises()
const reconfigure = wrapper.findAll('button').find(button => button.text() === 'Reconfigure router')
expect(reconfigure).toBeDefined()
await reconfigure!.trigger('click')
expect(wrapper.text()).toContain('Connect to Router')
expect(wrapper.text()).not.toContain('Connection timed out')
wrapper.unmount()
})
})
+5 -1
View File
@@ -260,6 +260,7 @@ function pickDetectedRouter(ip: string) {
function disconnectRouter() {
host.value = status.value?.host ?? host.value
connectedParams.value = null
error.value = ''
detectError.value = ''
detectedCandidates.value = []
showConnectForm.value = true
@@ -533,7 +534,10 @@ onMounted(() => {
<!-- Error state -->
<div v-else-if="error" class="glass-card p-6 mb-4">
<p class="text-sm text-red-300">{{ error }}</p>
<button class="mt-3 text-xs text-white/50 hover:text-white transition-colors underline" @click="load()">Retry</button>
<div class="mt-3 flex items-center gap-4">
<button class="text-xs text-white/50 hover:text-white transition-colors underline" @click="load()">Retry</button>
<button class="text-xs text-orange-300/80 hover:text-orange-200 transition-colors underline" @click="disconnectRouter">Reconfigure router</button>
</div>
</div>
<!-- Status panels -->
@@ -362,6 +362,23 @@ init()
</button>
</div>
<div class="overflow-y-auto flex-1 min-h-0 space-y-6 pr-1">
<!-- v1.8.5-alpha -->
<div>
<div class="flex items-center gap-2 mb-3">
<span class="text-xs font-mono px-2 py-0.5 rounded bg-orange-500/20 text-orange-300">v1.8.5-alpha</span>
<span class="text-xs text-white/40">August 30, 2026</span>
</div>
<div class="space-y-3 text-sm text-white/80 pl-3 border-l border-white/10">
<p>**Cuprate — an independent Monero node — is now an app.** Monero consensus validated by a second, unrelated codebase (Rust), the same layer of security-in-depth Bitcoin gets from Knots. Review caught two problems before anything shipped: the unrestricted RPC that can move funds stayed bound to the container's loopback (never published to the node, let alone the LAN — anything on the node could previously have reached it), and its restricted RPC moved off port 18089 to avoid colliding with Penpot. Honest caveat: upstream has cut no stable release yet, so the pin tracks an exact preview build (0.1.0-preview-18-g618ff14) and moves to their first tagged release when there is one.</p>
<p>**A frozen node now explains itself — and comes back on its own.** The host now captures a memory dump into /var/crash when the kernel panics *or* wedges (a hung kiosk used to sit dead until someone power-cycled it; now it dumps, reboots itself, and leaves the evidence behind), and records failing-memory signals (ECC errors) into a database as they happen. This is the first change delivered by a new host-update channel: the node's own updater now carries OS-level packages and settings to already-deployed machines — the crash-kernel's memory reservation is the one part that waits for a reboot, and the node says so rather than pretending.</p>
<p>**Uninstalling an app can no longer report success when it failed.** The declarative path used to swallow every teardown error and report the app uninstalled, leaving the tile behind and the truth in the logs. A failed uninstall now stops and shows the real per-app errors, so "still there" is never presented as "gone".</p>
<p>**Pictures to internet-only mesh contacts work now.** Sending an attachment inline always took the radio path and failed with "Peer is federation-only (no radio twin)" for contacts reachable only over the internet — and the size-adviser kept recommending a radio transfer those peers can't receive. Both fixed: inline sends route over the federation when that's the only way to reach the peer, and the advice no longer offers radio-only transfers to radio-unreachable contacts.</p>
<p>**Disk cleanup finally has honest numbers.** Space "free" on a drive was counted including the slice the filesystem keeps reserved for root — roughly 5% of the disk, 92 GB on one dev box — so the automatic cleanup that's supposed to kick in at 90% never triggered and stale container images piled up unnoticed. Reserved space now counts as used, which is what the threshold was always meant to measure.</p>
<p>**Three small screens that were lying to you, fixed.** The "Bitcoin is synced — fund your wallet" toast no longer appears on a node where the wallet it means (LND) isn't installed — it points at installing LND instead. The seed-reveal screen hides its third prompt unless the password actually fails to decrypt (the backup passphrase only exists if you set one). And multi-version store cards stop quoting a version number you'll be asked to choose on the next screen anyway.</p>
<p>**Mesh notifications survive a refresh, and a stale router no longer hides the fix.** Radio message unread counts are now remembered per contact instead of guessed from session state (the "one new message showed 11 unread" bug), cover Meshtastic, MeshCore and Reticulum alike, and deep-link to the right conversation; a single new message announces itself once. Separately, when the cached router address goes stale, the error card gains a "Reconfigure router" action instead of a Retry loop that can never succeed.</p>
<p>**The app updater now knows what upstream shipped.** Every app's manifest records where it comes from — including the odd corners (GitLab-only projects, ghcr-only images) — and a checker sweeps all of them against upstream releases, so a pin that quietly rots for months is now visible instead of invisible. The first full sweep found 27 pins behind; the safe patch-level ones shipped with this release (strfry, BTCPay Server 2.4.3, the two nginx frontends), and the major jumps that may carry data migrations are deliberately held for their own careful passes.</p>
</div>
</div>
<!-- v1.8.4-alpha -->
<div>
<div class="flex items-center gap-2 mb-3">
+123 -13
View File
@@ -505,6 +505,10 @@
"network_policy": "bridge",
"readonly_root": true
},
"upstream": {
"kind": "gitlab",
"repo": "ark-bitcoin/bark"
},
"version": "0.3.0",
"volumes": [
{
@@ -978,13 +982,13 @@
"version": "1.2.11"
},
"btcpay": {
"image": "docker.io/btcpayserver/btcpayserver:2.4.2",
"image": "docker.io/btcpayserver/btcpayserver:2.4.3",
"images": {
"archy-btcpay-db": "source.archipelago-foundation.org/lfg2025/postgres:15.17",
"archy-nbxplorer": "source.archipelago-foundation.org/lfg2025/nbxplorer:2.6.0",
"btcpay-server": "docker.io/btcpayserver/btcpayserver:2.4.2"
"btcpay-server": "docker.io/btcpayserver/btcpayserver:2.4.3"
},
"version": "2.4.2"
"version": "2.4.3"
},
"btcpay-server": {
"manifest": {
@@ -1000,7 +1004,7 @@
"template": "{{HOST_IP}}:23000"
}
],
"image": "docker.io/btcpayserver/btcpayserver:2.4.2",
"image": "docker.io/btcpayserver/btcpayserver:2.4.3",
"network": "archy-net",
"pull_policy": "if-not-present",
"secret_env": [
@@ -1097,7 +1101,7 @@
"kind": "github",
"repo": "btcpayserver/btcpayserver"
},
"version": "2.4.2",
"version": "2.4.3",
"volumes": [
{
"options": [
@@ -1110,7 +1114,7 @@
]
}
},
"version": "2.4.2"
"version": "2.4.3"
},
"core-lightning": {
"manifest": {
@@ -1205,6 +1209,96 @@
"image": "source.archipelago-foundation.org/lfg2025/cryptpad:2024.12.0",
"version": "2024.12.0"
},
"cuprate": {
"manifest": {
"app": {
"category": "money",
"container": {
"custom_args": [
"--config-file",
"/home/cuprate/Cuprated.toml"
],
"data_uid": "1000:1000",
"image": "source.archipelago-foundation.org/lfg2025/cuprate:0.1.0-preview-18-g618ff14",
"network": "archy-net",
"pull_policy": "if-not-present"
},
"dependencies": [
{
"storage": "300Gi"
}
],
"description": "Alternative Monero node implementation in Rust. Independently validates Monero consensus rules, providing a layer of security and redundancy for the network.",
"files": [
{
"content": "network = \"Mainnet\"\ntarget_max_memory = 3000000000\n\n[rpc.restricted]\nenable = true\n",
"overwrite": false,
"path": "/var/lib/archipelago/cuprate/Cuprated.toml"
}
],
"health_check": {
"endpoint": "localhost:18090",
"interval": "30s",
"retries": 3,
"start_period": "5m",
"timeout": "5s",
"type": "tcp"
},
"id": "cuprate",
"metadata": {
"author": "Cuprate",
"category": "money",
"icon": "/assets/img/app-icons/cuprate.svg",
"repo": "https://github.com/Cuprate/cuprate",
"tier": "optional"
},
"name": "Cuprate",
"ports": [
{
"auth": "none",
"auth_rationale": "Monero p2p gossip. Peers are anonymous by design and speak the Monero wire protocol, not HTTP.",
"container": 18080,
"host": 18183,
"protocol": "tcp"
},
{
"auth": "none",
"auth_rationale": "Monero restricted RPC — the subset upstream considers safe for public/remote-node use. Wallets (Feather, monero-wallet-rpc, GUI) connect directly over plain HTTP JSON-RPC and cannot hold a dashboard session cookie.",
"container": 18089,
"host": 18090,
"protocol": "tcp"
}
],
"resources": {
"cpu_limit": 0,
"disk_limit": "300Gi",
"memory_limit": "4Gi"
},
"security": {
"capabilities": [],
"network_policy": "isolated",
"no_new_privileges": true,
"readonly_root": true
},
"upstream": {
"kind": "github",
"repo": "Cuprate/cuprate"
},
"version": "0.1.0-preview",
"volumes": [
{
"options": [
"rw"
],
"source": "/var/lib/archipelago/cuprate",
"target": "/home/cuprate",
"type": "bind"
}
]
}
},
"version": "0.1.0-preview"
},
"did-wallet": {
"manifest": {
"app": {
@@ -2375,6 +2469,10 @@
"network_policy": "isolated",
"readonly_root": false
},
"upstream": {
"kind": "ghcr",
"repo": "immich-app/postgres"
},
"version": "14-vectorchord0.4.3-pgvectors0.2.0",
"volumes": [
{
@@ -2788,6 +2886,10 @@
"network_policy": "isolated",
"readonly_root": false
},
"upstream": {
"kind": "github",
"repo": "minio/minio"
},
"version": "RELEASE.2024-11-07T00-52-20Z",
"volumes": [
{
@@ -3175,6 +3277,10 @@
"seccomp_profile": "default",
"user": 1000
},
"upstream": {
"kind": "manual",
"url": "no public listing for lightninglabs/lightning-stack — verify by hand"
},
"version": "0.12.0",
"volumes": [
{
@@ -3625,7 +3731,7 @@
"key": "/var/lib/archipelago/netbird/tls.key"
}
],
"image": "docker.io/library/nginx:1.31.3-alpine",
"image": "docker.io/library/nginx:1.31.4-alpine",
"network": "netbird-net",
"pull_policy": "if-not-present"
},
@@ -4315,7 +4421,7 @@
"key": "/var/lib/archipelago/pine/tls.key"
}
],
"image": "docker.io/library/nginx:1.31.3-alpine",
"image": "docker.io/library/nginx:1.31.4-alpine",
"network": "archy-net",
"network_aliases": [
"pine"
@@ -4701,6 +4807,10 @@
"no_new_privileges": true,
"readonly_root": false
},
"upstream": {
"kind": "dockerhub",
"repo": "rhasspy/wyoming-whisper"
},
"version": "3.4.2",
"volumes": [
{
@@ -4998,7 +5108,7 @@
"manifest": {
"app": {
"container": {
"image": "dockurr/strfry:1.1.1",
"image": "dockurr/strfry:1.1.2",
"image_signature": "cosign://...",
"pull_policy": "verify-signature"
},
@@ -5055,7 +5165,7 @@
"kind": "github",
"repo": "hoytech/strfry"
},
"version": "1.1.1",
"version": "1.1.2",
"volumes": [
{
"options": [
@@ -5076,7 +5186,7 @@
]
}
},
"version": "1.1.1"
"version": "1.1.2"
},
"tailscale": {
"image": "source.archipelago-foundation.org/lfg2025/tailscale:stable",
@@ -5256,7 +5366,7 @@
}
},
"schema": 1,
"signature": "97628de24e3ffa17f639c663e19881cf6dea8c79aab272fe9c5442a4e951b3f0d257fee21aa9ce6158e3824b56a337acb71805f6fd34245e48345b86b46ec007",
"signature": "da5b6b183ac46c062945c27abdc06affb558e805e1ccf67ac0ee17e5e3dd85cc05a0656dd83bdacb1e1d237445145d00995f55e77209e1cbb2b6d8ce47084e0a",
"signed_by": "did:key:z6Mkfu5LT8d4DjETtrkATvHh9Dvcbnr7zBCUwfau8Sw7DLWT",
"updated": "2026-08-19"
"updated": "2026-08-30"
}
+45 -1
View File
@@ -39,6 +39,7 @@ import os
import re
import sys
import urllib.error
import urllib.parse
import urllib.request
from dataclasses import dataclass, field
from pathlib import Path
@@ -191,7 +192,50 @@ def _highest(tags: list[str], current: str = "") -> str:
return max(ranked)[1]
FETCHERS = {"github": latest_github, "dockerhub": latest_dockerhub}
def latest_gitlab(project: str, current: str = "") -> str:
"""Newest release tag for a GitLab `group/project`.
Some projects publish releases only on GitLab with no GitHub mirror
(bark lives at ark-bitcoin/bark and nowhere else). GitLab release tags
sometimes carry the project name as a prefix (`bark-0.6.2`); strip it so
version ordering can see the number.
"""
esc = urllib.parse.quote(project, safe="")
releases = http_json(
f"https://gitlab.com/api/v4/projects/{esc}/releases?per_page=100"
)
tags = [str(r["tag_name"]) for r in releases]
prefix = project.rsplit("/", 1)[-1].lower() + "-"
tags = [t[len(prefix):] if t.lower().startswith(prefix) else t for t in tags]
return _highest(tags, current)
def latest_ghcr(repo: str, current: str = "") -> str:
"""Newest version-like tag on GitHub's container registry.
Some images exist only on ghcr.io (immich-app/postgres publishes there
and nowhere else), so neither the GitHub-release nor the Docker Hub
fetcher can see them. Anonymous pull token first, then the tag list —
the same handshake any `docker pull ghcr.io/...` performs.
"""
token = http_json(
f"https://ghcr.io/token?scope=repository:{repo}:pull&service=ghcr.io"
)["token"]
req = urllib.request.Request(
f"https://ghcr.io/v2/{repo}/tags/list",
headers={"User-Agent": USER_AGENT, "Authorization": f"Bearer {token}"},
)
with urllib.request.urlopen(req, timeout=TIMEOUT) as res: # noqa: S310
tags = [str(t) for t in json.loads(res.read().decode()).get("tags", [])]
return _highest(tags, current)
FETCHERS = {
"github": latest_github,
"dockerhub": latest_dockerhub,
"gitlab": latest_gitlab,
"ghcr": latest_ghcr,
}
# ── Manifest reading ───────────────────────────────────────────────────────
+1 -1
View File
@@ -37,7 +37,7 @@ MEMPOOL_WEB_IMAGE="$ARCHY_REGISTRY/mempool-frontend:v3.3.1"
MARIADB_IMAGE="$ARCHY_REGISTRY/mariadb:11.4.10"
# BTCPay
BTCPAY_IMAGE="docker.io/btcpayserver/btcpayserver:2.4.2"
BTCPAY_IMAGE="docker.io/btcpayserver/btcpayserver:2.4.3"
NBXPLORER_IMAGE="$ARCHY_REGISTRY/nbxplorer:2.6.0"
POSTGRES_IMAGE="$ARCHY_REGISTRY/postgres:15.17"
BTCPAY_POSTGRES_IMAGE="$ARCHY_REGISTRY/postgres:15.17"
+42
View File
@@ -9,6 +9,8 @@
# C. FM-guards — the concrete failure modes that have bitten the
# fleet: port-drift (FM8), secret-completeness (FM2),
# orphaned container states (FM9), OTA wedge (FM12)
# D. Host capture (#144) — kdump + rasdaemon baseline: crash dumps configured
# and reserved, hang policy live, ECC recording running
#
# Everything here is READ-ONLY: no install/stop/start/uninstall, no service bounce.
# Safe to run against a live production node. It is the per-boot building block the
@@ -226,6 +228,43 @@ section_c() {
fi
}
# ══ Section D — host capture (#144): kdump + rasdaemon ═══════════════════════
section_d() {
echo
echo "== D. Host capture — crash + hardware-error evidence (#144) =="
if [[ "$ARCHY_LOCAL" != "1" ]]; then
record WARN "kdump + rasdaemon baseline" "remote node — host checks skipped"
return
fi
# D1. kdump enabled in config (image bakes it in; OTA host fixups converge)
if grep -qE '^USE_KDUMP=.?1' /etc/default/kdump-tools 2>/dev/null; then
record PASS "kdump-tools configured" "USE_KDUMP=1, dumps to /var/crash"
else
record FAIL "kdump-tools configured" "/etc/default/kdump-tools missing USE_KDUMP=1 — host fixup didn't land"
fi
# D2. crashkernel reservation — memory is reserved at BOOT, so a node that
# took the OTA fixup but hasn't rebooted yet is WARN, not FAIL.
if grep -q 'crashkernel=' /proc/cmdline 2>/dev/null; then
record PASS "crashkernel reserved" "$(grep -oE 'crashkernel=[^ ]+' /proc/cmdline | head -1)"
elif grep -q 'crashkernel=' /etc/default/grub 2>/dev/null; then
record WARN "crashkernel reserved" "written to GRUB — applies on next reboot"
else
record FAIL "crashkernel reserved" "absent from /proc/cmdline AND /etc/default/grub"
fi
# D3. hang/panic capture policy — runtime-settable, expected immediately
if [[ "$(cat /proc/sys/kernel/hung_task_panic 2>/dev/null)" == "1" ]]; then
record PASS "hang-capture policy live" "kernel.hung_task_panic=1"
else
record FAIL "hang-capture policy live" "kernel.hung_task_panic!=1 — sysctl drop-in not applied"
fi
# D4. rasdaemon recording hardware errors (ECC/AER events → sqlite)
if systemctl is-active --quiet rasdaemon 2>/dev/null; then
record PASS "rasdaemon active" "hardware-error events recorded to /var/lib/rasdaemon"
else
record FAIL "rasdaemon active" "service not running — package missing or host fixup failed"
fi
}
# ── run ────────────────────────────────────────────────────────────────────────
echo "=============================================================="
echo " OS-wide audit — ${BASE_URL} ($(date '+%Y-%m-%d %H:%M:%S'))"
@@ -237,6 +276,9 @@ if (( FAIL == 0 )) || [[ -n "$SESSION" ]]; then
section_b
section_c
fi
# Host-capture baseline is independent of RPC health: a wedged backend must
# not mask that the node also stopped capturing evidence.
section_d
echo
echo "=============================================================="