Files
archy/docs/UNIFIED-TASK-TRACKER.md
T
archipelagoandClaude Opus 5 0d513a0ef7 docs(10-05): record the Core-wallet fleet census — 4 nodes clear, 6 unchecked (D-07b)
Task 3 of plan 10-05, run by the operator over Tailscale on 2026-08-02 using the
read-only procedure in KEY-03-SIGNING-POSTURE.md. No escalation: nothing found.

Examined and CLEAR (4): archi-dev-box, shorty-s/.228, archy-x250-beta,
archy-x250-pa. On every one there is no wallet named `archipelago` — the deleted
handler's default wallet_name — `listwallets` returns only the unnamed default,
and that default reports blank=true, keypoolsize=0, txcount=0, balance=0. The
only named wallets are Fedimint gatewayd-*. The result holds across two
container vintages (bitcoin-knots and bitcoin-core), so it is not four copies of
one image behaving identically.

Not examined (6), recorded with reasons rather than omitted: framework-pt,
archipelago-1, archipelago and archy-dev-pa (SSH permission denied — password
rotated/not held), archipelago-5 (timed out during banner exchange), and
archy-x250-dev (offline). Password auth was deliberately not attempted: several
fleet nodes lock PAM quickly on a wrong password, and locking out an in-use
production node is a worse outcome than an incomplete census.

The conclusion is stated at the strength the evidence supports — no *examined*
node holds a wallet the deleted handler created, and no examined node holds any
wallet with keys or funds. It is deliberately NOT generalised to "the fleet is
clear" while six nodes are unknown. F-13 is closed by deletion regardless: the
code that could create such a wallet is gone from every future build.

No key material appeared in any output and `listdescriptors true` was never run.

Also corrects the now-stale R-04/F-13 entry in UNIFIED-TASK-TRACKER.md, which
still described `handle_bitcoin_init_wallet_from_seed` and a watch-only
migration as pending work — that code no longer exists. Marks it done-by-
deletion and adds the six unchecked nodes as a standing item, flagged as a
natural fold-in for KEY-04's on-node work but tracked independently so it does
not vanish if KEY-04 is re-scoped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 11:32:50 -04:00

31 KiB
Raw Blame History

Unified Task Tracker — OTA 1.8.0 + Master Plan

Single working list for everything left before 1.8.0 ships and the next master-plan exit criteria (multinode + workstreams B/C/D) are met. Supersedes the open-task sections of docs/archive/SESSION-1.8.0-OTA-PROGRESS.md and docs/PRODUCTION-MASTER-PLAN.md as the day-to-day tracker — those docs remain the historical record / detailed narrative and are still linked from here where useful. Ordered fastest/simplest first so we work top-down instead of hunting across docs.

Verified against actual code state on 2026-07-01 (not just doc text — several items the source docs still listed as "open" turned out to already be shipped; those are marked below with the commit that did it, so we stop re-litigating them).


Tier 0 — Quick / mechanical, no blockers

  • Ship the lightning payment false-failure fix in the next release (fixed on main 2026-07-27, needs OTA). Slow multi-hop payments (>15s) surfaced as "Payment failed" while LND settled them in the background — the shared LND REST client's 15s timeout aborted the synchronous /v1/channels/transactions wait. Now: payinvoice decodes the invoice first for its payment hash, waits up to 120s on a dedicated client, returns status: "pending" (never a failure) on timeout, and the new lnd.paymentstatus RPC + frontend payLightningInvoice() helper poll to a real terminal state (all 5 UI call sites migrated). Verify on Framework PT with a real multi-hop payment.

  • Show the app version on the companion mobile-app banner in the app store and on its install/pairing modal (user request 2026-07-27) — so it's obvious at a glance whether the node is serving the latest APK build.

  • Optimise the companion QR scan — quicker + better (user request 2026-07-27; deferred to a later session on purpose). The pairing/scan QR flow works (user-verified on-device 2026-07-27) but should get faster and smoother: quicker camera start + decode (scan resolution/framerate, continuous autofocus), more forgiving in low light / at an angle, and snappier feedback once the code locks. Touch the native-scan path from PR #104 and the in-app scan modal together so both benefit.

  • Update tests/lifecycle/TESTING.md's stale Release Gates checklist (lines 289296) — several boxes are unchecked but actually true now:

    • #1 bitcoin-stops: covered by tests/lifecycle/bats/bitcoin-knots.bats stop/restart tier, included in the 5/5 green gate run.
    • #2 ARCHY_ITERATIONS=5 on .228: GREEN 2026-06-23 per CLAUDE.md — check the box.
    • #5 cargo 0 warnings: confirmed 0 warnings on cargo build --release (2026-07-01).
    • #7 layman changelog: CHANGELOG.md is backfilled with layman-readable entries through v1.8.00-alpha — check the box.
    • Leave #3 (multinode), #4 (backend-survives-restart / Phase-3 default-on), #6 (LoC decision), #8 (tag pushed) unchecked — genuinely still open, see Tier 2/3.
  • Finish the archival/full-node manifest generalization — investigated 2026-07-01: the hardcoded fallback names in dependencies.rs:48-52 (electrs, mempool-electrs, mempool-web) are legacy alias ids for electrumx/mempool, resolved via id-mapping in a dozen other places (install.rs, runtime.rs, config.rs, etc.), not separate un-migrated apps with their own manifests. electrumx and mempool themselves already declare bitcoin:archival. The fallback is correct as-is — not tech debt, closing this item rather than risk breaking alias resolution.

  • Confirm/close the Portainer image-pin item — confirmed 2026-07-01: 146.59.87.168:3000/lfg2025/portainer:2.19.4 is present in podman images on all 3 LAN nodes (.116/.198/.228), i.e. actually resolvable/pulled from the mirror. Not a live bug.

  • grafana Quadlet "stuck activating" — checked live on .116 (2026-07-01): grafana.service is active (running), container Up 2 hours (healthy). The 2026-06-21 report is stale for grafana. strfry still unconfirmed — not installed on any of .116/.198/.228 to check directly; low priority until someone actually needs it installed.

  • Add cargo audit / cargo deny to CI, failing on duplicate rand majors (entropy audit R-05, finding F-07 — docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md). cargo-audit is not installed anywhere, so no RustSec check has ever run against this tree. Separately, cargo tree shows both rand 0.8.5 (direct, all first-party key generation) and rand 0.9.2 (transitive via totp-rs and tungstenite 0.26.2) resolved into one binary. rand 0.9.0 removed ThreadRng fork protection and the orchestrator forks constantly, so a future bump must be visible rather than silent — add a bans rule so the duplicate majors show up in CI, not in an incident.

  • Harden the release signing ceremony's mnemonic input (entropy audit R-08, finding F-06). ceremony gen prints the release master mnemonic to stdout (core/archipelago/src/ceremony.rs:71-77) and load_release_root_key prefers the RELEASE_MASTER_MNEMONIC environment variable over stdin (:157-160) — both leak into shell history, /proc/<pid>/environ, tmux scrollback and terminal recordings. This is the seed that derives the fleet release-root signing key, so a leak means forged signed manifests fleet-wide. Make stdin/TTY the only supported input for sign/pubkey; write gen's output to a 0600 file rather than the terminal. Small change, but schedule it deliberately — it is the signing ceremony.

  • Small entropy-audit hygiene batch (entropy audit R-09 R-12, R-14 — docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md). Five independent one-liners, each closing a Low/Informational finding:

    • Persist the CSPRNG-readiness verdict (seed.rs:85-91) as a durable structured event, so any node can answer post-hoc "was the entropy pool ready when this seed was born?" — the question Coldcard owners cannot answer today.
    • Add a test asserting the getrandom crate uses the blocking syscall, making seed.rs:52-57's invariant mechanical instead of a comment.
    • Clear _seed_words from sessionStorage on route-leave from onboarding, not only on successful verify (OnboardingSeedVerify.vue:251), plus a wall-clock expiry mirroring the server's 10-minute MNEMONIC_TTL.
    • Replace % charset.len() in totp.rs:305 with SliceRandom::choose(&mut OsRng). (No bias today — 32 divides 256 — but any future charset edit introduces one silently. The audit refutes the research's claim that this is currently biased.)
    • Comment pickRandomIndices (OnboardingSeedVerify.vue:157) to record that its Math.random() picks a UX challenge, not key material, so the next auditor does not re-derive that it is benign.
  • Swap container generated_secrets to explicit OsRng (entropy audit R-13, finding F-10) — two-line change in container/secrets.rs:90-102 SUPERSEDED 2026-08-02 by R-16 / KEY-05. The audit scoped this at 2 call sites; the real surface is 41 across 15 files — see the audit's new §F-10a. secrets.rs is 2 of them, and a two-line fix there while 39 other sites inherit the same dependency default is not a fix.

  • Crate-wide CSPRNG enforcement — a defaulted RNG cannot be inherited anywhere (entropy audit R-16 / F-10a, Medium) — tracked as KEY-05 in Phase 10, so plan and execute it there rather than as a standalone item. session.rs (16 sites), pine_ha.rs (6), wallet/bdhke.rs (2 prod — Cashu proof secret + blinding factor, genuine key material), storage_crypto.rs (1 — AEAD nonce), mesh/x3dh.rs (2 — prekey identifiers, not key material — corrected 2026-08-02), +10 more files. Nothing is broken today (rand::random()/thread_rng() are ChaCha12 from getrandom(2)), but it is the T1 shape that produced the COLDCARD defect, now with key material in the blast radius. Five layers: sealed allowlist trait at key-gen seams; clippy.toml disallowed-methods ban (compile-time, CI-enforced — no clippy.toml exists yet); cargo-deny on duplicate rand majors (absorbs R-05); degenerate-entropy runtime check; persist the CSPRNG-readiness verdict (absorbs R-09). Also retires the impl rand::CryptoRng for CountingRng false promise at seed.rs:656. Gated: do not start until the concurrent Phase 1 agent is done and synced.

Tier 1 — Medium effort, unblocked

  • Fix the fail-open first-boot secret regeneration in the ISO (entropy audit R-02 + R-03, finding F-03 — docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md). The installed rootfs is a cached container export shared by every node (image-recipe/_archived/build-auto-installer-iso.sh:717-726, extracted at :2303), and it bakes SSH host keys (via the openssh-server install at :345) and a TLS keypair (:463-469). archipelago-first-boot-secrets.service correctly regenerates both per device — but both branches are fail-open (:1647, :1659) and touch "$MARKER" at :1663 runs unconditionally, so a single transient failure permanently leaves that node on the image-wide shared SSH host key and TLS private key, with the failure visible only in a log file. Fix: (a) set the marker only when both regenerations succeeded, so it retries next boot; (b) surface the failure in the UI/doctor, not just the log; (c) strip the baked keys from the rootfs tar so a failure degrades to "no key" rather than "shared key". Needs an ISO rebuild and two fresh flashes to verify.

  • Reconcile Argon2::default() with ADR-005 (entropy audit R-06, finding F-05). ADR-005 states 64 MB / 3 iterations (docs/adr/005-chacha20-backup-encryption.md:31); Argon2::default() in argon2 0.5.3 is Argon2id at 19 MiB / t=2 / p=1. Used at core/archipelago/src/seed.rs:249 and :285, backup/identity.rs:38/:93, backup/full.rs:618/:650. Either raise the parameters behind a versioned envelope with a migration (an existing master_seed.enc was encrypted under the old parameters and will not decrypt under new ones) or amend the ADR to state the real numbers. Do not change them silently.

  • Run the on-node entropy verification checklist (entropy audit R-15, §6 of docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md). Everything in that section is explicitly UNVERIFIED — it needs real hardware this environment cannot reach. Highest value first: C-3 (are SSH host-key and TLS fingerprints actually different across two nodes flashed from the same ISO?) and C-5 (the cross-node same-ISO seed collision test — the empirical check that would have caught the Coldcard defect). Also C-1 (crng init done vs seed-generation timestamp), C-2 (machine-id uniqueness), C-4 (what the rootfs tar actually contains, run on the build host), C-6 (is /rpc/v1 reachable unauthenticated from the LAN). Use a disposable node — C-5 overwrites node identity.

  • immich → Quadlet migration — investigated 2026-07-01, turned out already done: immich uses the same install_stack_via_orchestrator primitive as netbird/btcpay (immich_stack_app_ids() in stacks.rs:690), and is confirmed running as real Quadlet units live on .228 (immich_server.container, immich_postgres.container, immich_redis.container, all active). Not a legacy in-cgroup app — the only remaining piece is the fleet-wide Phase-3 default-flip, already tracked in Tier 2.

  • Netbird reinstall adoption path — investigated 2026-07-01, not a bug, by design. adopt_stack_if_exists() (stacks.rs:140-198) is only used as a fallback when the orchestrator has no manifest for the app — there's nothing to render certs/config from in that case, so skipping rendering is correct. When the orchestrator does have the manifest (the normal path), the reconcile loop already re-renders certs even for adopted-running containers, fixed in 4519dbf0 (prod_orchestrator.rs:1707-1708).

  • TanStack Query (or equivalent) investigation — spike complete 2026-07-01, recommendation: don't adopt / close as not needed. Only 3 stores actually fetch data, WebSocket push already handles hot data (server-info/package-data), no cache-invalidation or stale-data bugs found, migration would touch 62 RPC call sites for no concrete payoff. If boilerplate ever bothers us, extract a usePolling() composable instead — much cheaper than a query-cache migration.

Tier 2 — High effort, mostly unblocked (the actual next exit criteria)

  • 🔴 Gate the unauthenticated seed RPCs (entropy audit R-01, finding F-01, Criticaldocs/security/ENTROPY-SEED-AUDIT-2026-07-31.md). seed.generate, seed.verify, seed.restore and seed.save-encrypted are in UNAUTHENTICATED_METHODS (core/archipelago/src/api/rpc/middleware.rs:24-28), which skips session, RBAC and CSRF (api/rpc/mod.rs:263, :295, :326). Neither handler checks whether onboarding is already complete (api/rpc/seed_rpc.rs:93-159, :226-305), and NodeIdentity::from_seed overwrites node_key, nostr_secret and the FIPS mesh key unconditionally (identity.rs:79-114). There is no rate limit (rate_limit.rs:60-97 has no seed.* entry). The endpoint is proxied to the LAN over plaintext HTTP (image-recipe/configs/nginx-archipelago.conf:11, :165, :192) and mesh peers can reach it too (server.rs:2080 asserts /rpc/v1 passes the peer path filter). Net: one unauthenticated POST can take over or destroy a live node's identity, and seed.restore lets the attacker choose the mnemonic. The guard already exists and is simply never called — NodeIdentity::key_exists (identity.rs:117). Fix: bail when a node key exists and no onboarding mnemonic is pending; prefer also gating on auth_manager.is_onboarding_complete(); add rate limits at auth.changePassword strictness; narrow the peer path filter. Changes an authentication boundary on a live fleet — needs its own /gsd-plan-phase with a federation re-verify, not an opportunistic patch.

  • PSBT-first signing: Phase 1 — move the Bitcoin private key out of CoreDONE 2026-08-02 by deletion, not conversion (entropy audit R-04, finding F-13; Phase 10 plan 10-05, decision D-07b). The handler that imported the BIP-84 account private key into Core's wallet.dat had no caller anywhere, LND is the wallet the UI drives, and the endpoint was authenticated and password-gated — so it was deleted outright rather than rewritten watch-only. bitcoin.rs's wallet-init handler and its dispatcher.rs arm are gone; no daemon code path writes the BIP-84 private key into Bitcoin Core. No migration was performed or is needed — a 4-node fleet census found no wallet the handler created. D-09's key-origin requirement moved to the PSBT itself: lnd.create-psbt now reports key_origin (psbt_key_origin_report, api/rpc/lnd/wallet.rs). Read docs/security/KEY-03-SIGNING-POSTURE.md for the current state — it also records the verdict that no fleet node is provisioned watch-only, so what ships today is PSBT transport, not air-gapped custody.

  • Finish the Core-wallet fleet census — 6 nodes unchecked (Phase 10 plan 10-05, Task 3; standing item). The 2026-08-02 census examined 4 nodes (archi-dev-box, shorty-s/.228, archy-x250-beta, archy-x250-pa) and found no wallet created by the deleted handler and no wallet holding keys or funds. Six were not examined: framework-pt, archipelago-1, archipelago, archy-dev-pa and archipelago-5 (SSH auth/connectivity) and archy-x250-dev (offline). Re-run the read-only procedure in docs/security/KEY-03-SIGNING-POSTURE.md § Fleet census when credentials or connectivity allow — a natural fold-in for KEY-04's on-node work. Never run listdescriptors true (it returns private keys). If any node reports a wallet named archipelago, or any descriptor wallet with private_keys_enabled: true that is not blank/empty, stop and escalate — do not migrate or modify it (D-07b).

  • PSBT-first signing: Phases 2-7 rollout (docs/security/PSBT-SIGNING-ARCHITECTURE.md §8) — the spec is written to be consumed directly by /gsd-plan-phase, with per-phase goals, dependencies, candidate requirements and hardware gating. Sequence: PSBT construct/export → external-signer import + finalize → air-gap transport (BC-UR v2 primary, BBQr for Coldcard, file fallback always) → wsh(sortedmulti) multisig on BIP-48 → LND remote signing → hot-wallet spend limits and cold/warm/hot tiering. Two hard rules the spec fixes in place: a channel-funding PSBT must never be self-broadcast (funds can be lost), and no UI copy may imply a routing node's Lightning channel keys are cold — they are necessarily hot. Phases 3-6 need real hardware.

  • Confine the seed-bearing RPCs to loopback/TLS (entropy audit R-07, finding F-04 / [ARCHY-4]). The 24-word master mnemonic is returned to the browser over JSON-RPC (core/archipelago/src/api/rpc/seed_rpc.rs:147, :156-158), held in process memory under a 10-minute TTL (:27) and deliberately not cleared at verify time (:205-211, with a documented and defensible rationale about client retries) — over a transport that is plaintext HTTP on LAN by design (api/rpc/mod.rs:227-241). Anyone with LAN traffic visibility during onboarding reads the phrase that unlocks the wallet and the node identity. Fix: force TLS or loopback for seed methods, shrink the TTL, and clear on an acknowledged verify with a short grace window. Touches the onboarding transport — needs a phase.

  • [~] Multinode test pass (docs/multinode-testing-plan.md) — worked the preconditions on .198 2026-07-01:

    • cleared 2 stale failed-unit records (archy-mempool-db.service, meshtastic.service — both not-found/dead since 6 and 5 days ago, harmless bookkeeping, systemctl --user reset-failed).
    • nginx /app/lnd/ proxy target confirmed correct (→ 18083, matches the running archy-lnd-ui port) — the plan's "stale proxy target" concern doesn't apply here.
    • .198 disk (448GB) is below the 1TB archival threshold + was only 21% through IBD — user chose to swap in a different node rather than wait/add storage. .116 ruled out (no bitcoin container installed at all, just the UI companion). .120 ruled out (reserved for another developer). .5 (archy-x250-beta, Tailscale 100.72.136.5) chosen: also sub-1TB (472GB, so still pruned — that ceiling is shared by every non-.228 node), but fully synced (ibd:false, blocks==headers 956,240). Bootstrapped bats 1.11.1 + jq 1.7.1 onto it 2026-07-01 and launched the 5× destructive gate (ARCHY_ITERATIONS=5 ARCHY_ALLOW_DESTRUCTIVE=1) — running now, log at /tmp/gate.log on .5, background poller watching for the RESULTS banner.
    • Once .5's gate reports: bring the rest of the fleet to precondition, then the cross-node federation/mesh/transport suites. This is the literal "next exit criterion" called out in CLAUDE.md.
  • Phase-3 Quadlet default-flip — code is validated + opt-in via ARCHIPELAGO_USE_QUADLET_BACKENDS=true on .228/.198 already (confirmed live 2026-07-01). Ready to flip (config.rs:256 + its test) the moment the .5 gate reports clean — deliberately NOT staged uncommitted in the tree (a prior attempt left an uncommitted flip sitting around and that caused confusion; it's a 2-line change, faster to just do it fresh once confirmed).

  • ~Per-app test coverage for the 30 apps with zero automated coveragereframed 2026-07-01, mostly a non-issue. all-apps-matrix.bats + all-apps-lifecycle.bats already give EVERY installed app generic baseline coverage (no stuck state, no error state, stop/start/restart survives, UI reachable). The real gap is narrower: 34 apps lack app-specific assertions (health endpoints, API queryability, data integrity) beyond that baseline — aiui, bitcoin-core, botfights, core-lightning, did-wallet, fedimint-clientd, fedimint-gateway, fips-ui, gitea, grafana, home-assistant, indeedhub (+5 sub-containers), jellyfin, lightning-stack, lnd-ui, morphos-server, netbird (+2 sub-containers), nextcloud, nostr-rs-relay, photoprism, portainer, router, searxng, strfry, uptime-kuma, vaultwarden. Not urgent — baseline coverage is real safety net; treat as a backlog "nice to harden further," not a gate item.

  • Convert remaining multi-container legacy stacks to the manifest-owned modelinvestigated 2026-07-01, DONE, nothing left. All 5 real multi-container stacks (btcpay, mempool, immich, netbird, indeedhub) are on the install_stack_via_orchestrator pattern (stacks.rs). saleor was removed from the codebase; portainer/home-assistant/grafana are single-container manifest-driven apps, never stacks; fedimint/fedimint-gateway/fedimint-clientd are 3 separate single-container apps with manifest dependency edges, not a coordinated stack. Workstream A's stack-migration tail is fully closed.

  • Container thrashing/flapping + reconciler churn (added 2026-07-04 — was implicit across other tracks, now an explicit pre-tag concern). The root cause of restart-storm flapping is pre-Quadlet architecture: restarting archipelago.service SIGKILLs every container in its cgroup, then the reconciler rebuilds the world over several minutes (the post-OTA health check deliberately skips per-app container assertions because of exactly this). Consolidated lever list, in order of impact:

    • Phase-3 Quadlet default-flip (tracked above) — removes the SIGKILL-the-world behavior entirely; the single biggest fix.
    • Workstream F lifecycle items — immich/grafana uninstall hangs + ghost containers, grafana reinstall stops, fedimint guardian sync (docs/PRODUCTION-MASTER-PLAN.md workstream F).
    • Reconciler churn observability — no metric/log today distinguishes "settling after restart" from "flapping"; add a per-app restart counter + log line when an app restarts >N times in M minutes so thrash is visible instead of anecdotal.
    • Failed-unit self-healing gap (observed live 2026-07-06 on .228): fedimint's quadlet unit exited 255 at 21:21 and sat failed for 7+ hours — the reconciler never revived it (it repairs missing/drifted containers but doesn't reset-failed+start failed .services). Same for the indeedhub trio after the gate run. The health monitor also can't help (container is gone when the unit fails). Add a reconcile step: quadlet-backed app whose .service is failed and not user-stopped → reset-failed + start, with backoff.
    • Already landed, don't re-do: boot-reconciler circuit breaker (2026-07-01), indeedhub crashloop fix (2026-07-01), async blocking-Command pass (4c75bb3d, removes executor stalls that made the API janky under reconcile load), quadlet entrypoint-split false-drift fix (2026-07-08 — container_command_drifted compared entrypoint/cmd halves separately, but quadlet folds sh -lc into Entrypoint=sh + Exec=-lc …, so every quadlet-created app with a multi-element entrypoint read as permanently drifted; electrumx on .228 recreated 114×/6h until the comparator was switched to concatenated argv).
    • Perf polish riding along: 93 MB frontend dist shrink (hardening plan §D 🟡).
  • Developer tooling CLI suite (validate/render/local-install/lifecycle-test) — APP-PACKAGING-MIGRATION-PLAN.md step 5, needed before external devs can publish.

  • Consolidated deploy 2026-07-01: merged PR #67 (reticulum daemon process-group fix, 469b0203), the UI/UX work (8256fde1 — mesh/web5/apps layout, modal, search UX), and archy-openwrt (TollGate/OpenWrt gateway integration — new core/openwrt crate, RPC surface, OpenWrtGateway.vue) into main, alongside the indeedhub self-heal fix — all merged clean, no conflicts. Found + fixed 2 real build-breaking issues during verification, not caught by whoever authored them: a vestigial unused ref in Web5ConnectedNodes.vue that broke vue-tsc, and a stale MeshMap.test.ts mock missing federatedPositions (predated this session's Mesh Map feature) that crashed on mount. Full test suite green (667 passed) after fixes. Deployed fleet-wide 2026-07-01, all 5 nodes sha256-verified: .116, .198, .228, .5 (recovered cleanly from one truncated-transfer hiccup, caught via checksum before it hit the live service), 100.82.34.38 (non-Quadlet node — all containers survived the restart intact, unlike the worst-case risk flagged beforehand). Also built an unbundled installer ISO from this same merged source (archipelago-installer-1.7.99-alpha-unbundled-x86_64.iso, 2.4GB) — the ISO pipeline was archived from the release process at v1.7.43-alpha (OTA tarballs are now primary) but the wrapper script still works.

  • ⚠️ NOT YET DEPLOYED — start here next session. After the fleet deploy above, found that PR #67 ("kill whole daemon process group on drop", branch fix/reticulum-daemon-process-group, head be50c886) is a different, separate reticulum-daemon fix from the one already deployed (469b0203 on fix/reticulum-daemon-pdeathsig) — I'd conflated the two by topic similarity and only merged/deployed the Python-level pdeathsig fix, missing PR #67's Rust-level kill-whole-process-group-on-Drop fix entirely. Merged PR #67 into main (7a7fec21, clean, cargo check green, complementary not conflicting with the already-deployed fix) and separately fixed a real bug found live: OpenWrtGateway.vue's back button had no @click handler at all (7d7ba573, vue-tsc clean). Both committed + pushed to main but genuinely NOT deployed to any node — user asked to hold off deploying to restart their computer. Also spot-checked openwrt.scan live on .116: RPC plumbing works, but no physical OpenWrt router was available to confirm true-positive detection, and detect::scan_subnet does blocking TCP/SSH calls inside an async fn with no .await — untested at scale, worth hardening. Next steps: build release binary + frontend from current main, deploy to all 5 fleet nodes (.116/.198/.228/.5/100.82.34.38) the same way as the earlier consolidated deploy, then verify the back button + (if a real OpenWrt router is available) router detection live.

  • [~] Cross-node federation/mesh/transport suitesbig find 2026-07-01: these already exist, just aren't wired into the gate or documented as existing: tests/multinode/smoke.sh (federation pairing/sync, FIPS anchor, peer content browse, tombstone-removal regression tests), tests/multinode/meshtastic.sh (8-stage on-air mesh test), harness in tests/multinode/lib/multinode.bash. Actually ran smoke.sh live against .116↔.228 2026-07-01: 14 passed, 1 failed, 1 skipped. Confirms federation pairing (both directions), FIPS anchor connectivity (both nodes), and peer-content-browse-over-mesh (the v1.7.95 fix) all genuinely work node-to-node right now.

    • ⚠️ Real robustness gap found: node_rpc() in tests/multinode/lib/multinode.bash has no --max-time on its curl calls — a slow server-side RPC hangs the whole suite with zero feedback (this is what looked like a hang before it eventually completed on its own). Cheap fix, not yet applied.
    • 🐛 Real regression found and root-caused: removing a federation node (federation.remove-node) doesn't reliably stick — B reappeared in A's peer list after removal in the live test. Root cause: remove_node() (core/archipelago/src/federation/storage.rs:187) does let _ = tombstone_did(data_dir, did).awaitsilently swallows the tombstone write's errors. If that write fails (disk I/O, permission, transient issue), the peer is removed from nodes.json but never actually tombstoned, so the next background sync/notify-join re-adds it — the tombstone check at handlers.rs:592-599 passes because the DID was never recorded as removed. Diagnosed as a pre-existing logic gap, not a fresh regression from the v1.7.95 fix. Not fixed yet — this is federation/trust code, deliberately not touching it blind; needs a careful fix (surface the tombstone-write failure instead of swallowing it, and/or retry) plus re-verification with smoke.sh before considering it closed.

Tier 3 — Blocked on a decision or resource only you can supply

  • Version naming decisiondecided 2026-07-08: 1.8.0-alpha. Remaining work is the mechanical bump + tag + push once the pre-tag items above close.
  • Workstream B signing ceremonydone 2026-07-02. anchor.rs pins RELEASE_ROOT_PUBKEY_HEX = 5d15cbee…9951 (signer did:key:z6MkkidEnEpo6qHMCNSZoNKWtvQvxq3whnaME9wGgEFhq7ur); mnemonic held offline per docs/workstream-b-signing-runbook.md.
  • Bitcoin multi-version fleet-wide OTA.228 fully working on branch, per your prior gating this rollout is explicitly held for your decision on timing (docs/bitcoin-version-bulletproof-rollout.md).
  • 3ccc stock-Meshtastic RF validation — needs a live send/receive test with physical radios in your hands; code fix is in place, just unverified live.

Backlog — deferred, no scope decided, low priority

  • Marketplace protocol (workstream C) — design-only (docs/marketplace-protocol.md), no tooling/trust UX built. Future work, not urgent.
  • DHT distribution (workstream D) — confirmed design-only, no code (docs/dht-distribution-design.md explicitly says "Status: Design (no code yet)"); an experimental iroh provider skeleton exists behind a feature flag for future PoC measurement, nothing fleet-facing.
  • Custom live voice-call protocol — deprioritized 2026-07-01 per user request; scope not yet decided. Revisit after the tiers above are worked down.

Historical narrative and detailed per-session logs remain in docs/archive/SESSION-1.8.0-OTA-PROGRESS.md and docs/PRODUCTION-MASTER-PLAN.md §6/§8b — this doc is the live "what's left, in priority order" list. Update it (don't just append to the old docs) as items close or new ones surface.