Files
archy/docs/UNIFIED-TASK-TRACKER.md
T
archipelagoandClaude Opus 5 0d513a0ef7 docs(10-05): record the Core-wallet fleet census — 4 nodes clear, 6 unchecked (D-07b)
Task 3 of plan 10-05, run by the operator over Tailscale on 2026-08-02 using the
read-only procedure in KEY-03-SIGNING-POSTURE.md. No escalation: nothing found.

Examined and CLEAR (4): archi-dev-box, shorty-s/.228, archy-x250-beta,
archy-x250-pa. On every one there is no wallet named `archipelago` — the deleted
handler's default wallet_name — `listwallets` returns only the unnamed default,
and that default reports blank=true, keypoolsize=0, txcount=0, balance=0. The
only named wallets are Fedimint gatewayd-*. The result holds across two
container vintages (bitcoin-knots and bitcoin-core), so it is not four copies of
one image behaving identically.

Not examined (6), recorded with reasons rather than omitted: framework-pt,
archipelago-1, archipelago and archy-dev-pa (SSH permission denied — password
rotated/not held), archipelago-5 (timed out during banner exchange), and
archy-x250-dev (offline). Password auth was deliberately not attempted: several
fleet nodes lock PAM quickly on a wrong password, and locking out an in-use
production node is a worse outcome than an incomplete census.

The conclusion is stated at the strength the evidence supports — no *examined*
node holds a wallet the deleted handler created, and no examined node holds any
wallet with keys or funds. It is deliberately NOT generalised to "the fleet is
clear" while six nodes are unknown. F-13 is closed by deletion regardless: the
code that could create such a wallet is gone from every future build.

No key material appeared in any output and `listdescriptors true` was never run.

Also corrects the now-stale R-04/F-13 entry in UNIFIED-TASK-TRACKER.md, which
still described `handle_bitcoin_init_wallet_from_seed` and a watch-only
migration as pending work — that code no longer exists. Marks it done-by-
deletion and adds the six unchecked nodes as a standing item, flagged as a
natural fold-in for KEY-04's on-node work but tracked independently so it does
not vanish if KEY-04 is re-scoped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 11:32:50 -04:00

436 lines
31 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Unified Task Tracker — OTA 1.8.0 + Master Plan
Single working list for everything left before 1.8.0 ships and the next master-plan
exit criteria (multinode + workstreams B/C/D) are met. Supersedes the open-task
sections of `docs/archive/SESSION-1.8.0-OTA-PROGRESS.md` and `docs/PRODUCTION-MASTER-PLAN.md`
as the day-to-day tracker — those docs remain the historical record / detailed
narrative and are still linked from here where useful. **Ordered fastest/simplest
first** so we work top-down instead of hunting across docs.
Verified against actual code state on 2026-07-01 (not just doc text — several
items the source docs still listed as "open" turned out to already be shipped;
those are marked ✅ below with the commit that did it, so we stop re-litigating them).
---
## Tier 0 — Quick / mechanical, no blockers
- [ ] **Ship the lightning payment false-failure fix in the next release** (fixed
on main 2026-07-27, needs OTA). Slow multi-hop payments (>15s) surfaced as
"Payment failed" while LND settled them in the background — the shared LND
REST client's 15s timeout aborted the synchronous `/v1/channels/transactions`
wait. Now: payinvoice decodes the invoice first for its payment hash, waits
up to 120s on a dedicated client, returns `status: "pending"` (never a
failure) on timeout, and the new `lnd.paymentstatus` RPC + frontend
`payLightningInvoice()` helper poll to a real terminal state (all 5 UI call
sites migrated). Verify on Framework PT with a real multi-hop payment.
- [ ] **Show the app version on the companion mobile-app banner in the app store
and on its install/pairing modal** (user request 2026-07-27) — so it's
obvious at a glance whether the node is serving the latest APK build.
- [ ] **Optimise the companion QR scan — quicker + better** (user request
2026-07-27; deferred to a later session on purpose). The pairing/scan QR
flow works (user-verified on-device 2026-07-27) but should get faster and
smoother: quicker camera start + decode (scan resolution/framerate,
continuous autofocus), more forgiving in low light / at an angle, and
snappier feedback once the code locks. Touch the native-scan path from
PR #104 and the in-app scan modal together so both benefit.
- [ ] **Update `tests/lifecycle/TESTING.md`'s stale Release Gates checklist** (lines
289296) — several boxes are unchecked but actually true now:
- #1 bitcoin-stops: covered by `tests/lifecycle/bats/bitcoin-knots.bats` stop/restart
tier, included in the 5/5 green gate run.
- #2 `ARCHY_ITERATIONS=5` on .228: **GREEN 2026-06-23 per CLAUDE.md** — check the box.
- #5 cargo 0 warnings: confirmed 0 warnings on `cargo build --release` (2026-07-01).
- #7 layman changelog: `CHANGELOG.md` is backfilled with layman-readable entries
through v1.8.00-alpha — check the box.
- Leave #3 (multinode), #4 (backend-survives-restart / Phase-3 default-on), #6
(LoC decision), #8 (tag pushed) unchecked — genuinely still open, see Tier 2/3.
- [x] ~~Finish the archival/full-node manifest generalization~~ — investigated 2026-07-01:
the hardcoded fallback names in `dependencies.rs:48-52` (`electrs`, `mempool-electrs`,
`mempool-web`) are legacy **alias** ids for `electrumx`/`mempool`, resolved via
id-mapping in a dozen other places (`install.rs`, `runtime.rs`, `config.rs`, etc.),
not separate un-migrated apps with their own manifests. `electrumx` and `mempool`
themselves already declare `bitcoin:archival`. The fallback is correct as-is —
not tech debt, closing this item rather than risk breaking alias resolution.
- [x] ~~Confirm/close the Portainer image-pin item~~ — confirmed 2026-07-01:
`146.59.87.168:3000/lfg2025/portainer:2.19.4` is present in `podman images` on
all 3 LAN nodes (.116/.198/.228), i.e. actually resolvable/pulled from the mirror.
Not a live bug.
- [x] ~~grafana Quadlet "stuck activating"~~ — checked live on .116 (2026-07-01):
`grafana.service` is `active (running)`, container `Up 2 hours (healthy)`. The
2026-06-21 report is stale for grafana. **strfry still unconfirmed** — not
installed on any of .116/.198/.228 to check directly; low priority until someone
actually needs it installed.
- [ ] **Add `cargo audit` / `cargo deny` to CI, failing on duplicate `rand` majors**
(entropy audit R-05, finding F-07 —
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). `cargo-audit` is not installed
anywhere, so no RustSec check has ever run against this tree. Separately,
`cargo tree` shows **both** `rand 0.8.5` (direct, all first-party key generation)
and `rand 0.9.2` (transitive via `totp-rs` and `tungstenite 0.26.2`) resolved into
one binary. `rand 0.9.0` removed `ThreadRng` fork protection and the orchestrator
forks constantly, so a future bump must be visible rather than silent — add a
`bans` rule so the duplicate majors show up in CI, not in an incident.
- [ ] **Harden the release signing ceremony's mnemonic input** (entropy audit R-08,
finding F-06). `ceremony gen` prints the release master mnemonic to **stdout**
(`core/archipelago/src/ceremony.rs:71-77`) and `load_release_root_key` prefers the
`RELEASE_MASTER_MNEMONIC` **environment variable** over stdin (`:157-160`) — both
leak into shell history, `/proc/<pid>/environ`, tmux scrollback and terminal
recordings. This is the seed that derives the fleet release-root signing key, so a
leak means forged signed manifests fleet-wide. Make stdin/TTY the only supported
input for `sign`/`pubkey`; write `gen`'s output to a `0600` file rather than the
terminal. Small change, but schedule it deliberately — it is the signing ceremony.
- [ ] **Small entropy-audit hygiene batch** (entropy audit R-09 R-12, R-14 —
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). Five independent one-liners,
each closing a Low/Informational finding:
- Persist the CSPRNG-readiness verdict (`seed.rs:85-91`) as a durable structured
event, so any node can answer post-hoc "was the entropy pool ready when this seed
was born?" — the question Coldcard owners cannot answer today.
- Add a test asserting the `getrandom` crate uses the **blocking** syscall, making
`seed.rs:52-57`'s invariant mechanical instead of a comment.
- Clear `_seed_words` from `sessionStorage` on route-leave from onboarding, not only
on successful verify (`OnboardingSeedVerify.vue:251`), plus a wall-clock expiry
mirroring the server's 10-minute `MNEMONIC_TTL`.
- Replace `% charset.len()` in `totp.rs:305` with `SliceRandom::choose(&mut OsRng)`.
(No bias today — 32 divides 256 — but any future charset edit introduces one
silently. The audit refutes the research's claim that this is currently biased.)
- Comment `pickRandomIndices` (`OnboardingSeedVerify.vue:157`) to record that its
`Math.random()` picks a UX challenge, not key material, so the next auditor does
not re-derive that it is benign.
- [ ] ~~**Swap container `generated_secrets` to explicit `OsRng`** (entropy audit R-13,
finding F-10) — two-line change in `container/secrets.rs:90-102`~~
**SUPERSEDED 2026-08-02 by R-16 / KEY-05.** The audit scoped this at 2 call sites; the
real surface is **41 across 15 files** — see the audit's new §F-10a. `secrets.rs` is 2
of them, and a two-line fix there while 39 other sites inherit the same dependency
default is not a fix.
- [ ] **Crate-wide CSPRNG enforcement — a defaulted RNG cannot be inherited anywhere**
(entropy audit **R-16 / F-10a**, Medium) — tracked as **KEY-05 in Phase 10**, so plan
and execute it there rather than as a standalone item. `session.rs` (16 sites),
`pine_ha.rs` (6), `wallet/bdhke.rs` (2 prod — **Cashu proof secret + blinding factor,
genuine key material**), `storage_crypto.rs` (1 — **AEAD nonce**), `mesh/x3dh.rs` (2 —
prekey *identifiers*, **not** key material — corrected 2026-08-02), +10 more files.
Nothing is broken today (`rand::random()`/`thread_rng()` are ChaCha12 from
`getrandom(2)`), but it is the T1 shape that produced the COLDCARD defect, now with key
material in the blast radius. Five layers: sealed allowlist trait at key-gen seams;
`clippy.toml` `disallowed-methods` ban (compile-time, CI-enforced — no `clippy.toml`
exists yet); `cargo-deny` on duplicate `rand` majors (absorbs R-05); degenerate-entropy
runtime check; persist the CSPRNG-readiness verdict (absorbs R-09). Also retires the
`impl rand::CryptoRng for CountingRng` false promise at `seed.rs:656`.
**Gated: do not start until the concurrent Phase 1 agent is done and synced.**
## Tier 1 — Medium effort, unblocked
- [ ] **Fix the fail-open first-boot secret regeneration in the ISO** (entropy audit
R-02 + R-03, finding F-03 — `docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`).
The installed rootfs is a **cached container export shared by every node**
(`image-recipe/_archived/build-auto-installer-iso.sh:717-726`, extracted at
`:2303`), and it bakes SSH host keys (via the `openssh-server` install at `:345`)
and a TLS keypair (`:463-469`). `archipelago-first-boot-secrets.service` correctly
regenerates both per device — but both branches are **fail-open** (`:1647`,
`:1659`) and `touch "$MARKER"` at `:1663` runs **unconditionally**, so a single
transient failure permanently leaves that node on the image-wide shared SSH host
key and TLS private key, with the failure visible only in a log file. Fix:
(a) set the marker only when both regenerations succeeded, so it retries next
boot; (b) surface the failure in the UI/doctor, not just the log; (c) strip the
baked keys from the rootfs tar so a failure degrades to "no key" rather than
"shared key". Needs an ISO rebuild and two fresh flashes to verify.
- [ ] **Reconcile `Argon2::default()` with ADR-005** (entropy audit R-06, finding F-05).
ADR-005 states 64 MB / 3 iterations
(`docs/adr/005-chacha20-backup-encryption.md:31`); `Argon2::default()` in
argon2 0.5.3 is Argon2id at **19 MiB / t=2 / p=1**. Used at
`core/archipelago/src/seed.rs:249` and `:285`, `backup/identity.rs:38`/`:93`,
`backup/full.rs:618`/`:650`. Either raise the parameters behind a versioned
envelope **with a migration** (an existing `master_seed.enc` was encrypted under
the old parameters and will not decrypt under new ones) or amend the ADR to state
the real numbers. Do not change them silently.
- [ ] **Run the on-node entropy verification checklist** (entropy audit R-15, §6 of
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). Everything in that section is
explicitly **UNVERIFIED** — it needs real hardware this environment cannot reach.
Highest value first: **C-3** (are SSH host-key and TLS fingerprints actually
different across two nodes flashed from the same ISO?) and **C-5** (the cross-node
same-ISO seed collision test — the empirical check that would have caught the
Coldcard defect). Also C-1 (`crng init done` vs seed-generation timestamp), C-2
(`machine-id` uniqueness), C-4 (what the rootfs tar actually contains, run on the
build host), C-6 (is `/rpc/v1` reachable unauthenticated from the LAN). Use a
disposable node — C-5 overwrites node identity.
- [x] ~~immich → Quadlet migration~~ — investigated 2026-07-01, turned out already done:
immich uses the same `install_stack_via_orchestrator` primitive as netbird/btcpay
(`immich_stack_app_ids()` in `stacks.rs:690`), and is confirmed running as real
Quadlet units live on .228 (`immich_server.container`, `immich_postgres.container`,
`immich_redis.container`, all active). Not a legacy in-cgroup app — the only
remaining piece is the fleet-wide Phase-3 default-flip, already tracked in Tier 2.
- [x] ~~Netbird reinstall adoption path~~ — investigated 2026-07-01, **not a bug, by
design.** `adopt_stack_if_exists()` (`stacks.rs:140-198`) is only used as a
fallback when the orchestrator has no manifest for the app — there's nothing to
render certs/config from in that case, so skipping rendering is correct. When
the orchestrator *does* have the manifest (the normal path), the reconcile loop
already re-renders certs even for adopted-running containers, fixed in
`4519dbf0` (`prod_orchestrator.rs:1707-1708`).
- [x] ~~TanStack Query (or equivalent) investigation~~ — spike complete 2026-07-01,
**recommendation: don't adopt / close as not needed.** Only 3 stores actually fetch
data, WebSocket push already handles hot data (server-info/package-data), no
cache-invalidation or stale-data bugs found, migration would touch 62 RPC call
sites for no concrete payoff. If boilerplate ever bothers us, extract a
`usePolling()` composable instead — much cheaper than a query-cache migration.
## Tier 2 — High effort, mostly unblocked (the actual next exit criteria)
- [ ] **🔴 Gate the unauthenticated seed RPCs** (entropy audit R-01, finding **F-01,
Critical** — `docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). `seed.generate`,
`seed.verify`, `seed.restore` and `seed.save-encrypted` are in
`UNAUTHENTICATED_METHODS` (`core/archipelago/src/api/rpc/middleware.rs:24-28`),
which skips session, RBAC **and** CSRF (`api/rpc/mod.rs:263`, `:295`, `:326`).
Neither handler checks whether onboarding is already complete
(`api/rpc/seed_rpc.rs:93-159`, `:226-305`), and `NodeIdentity::from_seed`
overwrites `node_key`, `nostr_secret` and the FIPS mesh key **unconditionally**
(`identity.rs:79-114`). There is no rate limit (`rate_limit.rs:60-97` has no
`seed.*` entry). The endpoint is proxied to the LAN over plaintext HTTP
(`image-recipe/configs/nginx-archipelago.conf:11`, `:165`, `:192`) and mesh peers
can reach it too (`server.rs:2080` asserts `/rpc/v1` passes the peer path filter).
Net: **one unauthenticated POST can take over or destroy a live node's identity**,
and `seed.restore` lets the attacker choose the mnemonic. The guard already exists
and is simply never called — `NodeIdentity::key_exists` (`identity.rs:117`).
Fix: bail when a node key exists and no onboarding mnemonic is pending; prefer
also gating on `auth_manager.is_onboarding_complete()`; add rate limits at
`auth.changePassword` strictness; narrow the peer path filter. Changes an
authentication boundary on a live fleet — **needs its own `/gsd-plan-phase` with a
federation re-verify**, not an opportunistic patch.
- [x] **PSBT-first signing: Phase 1 — move the Bitcoin private key out of Core** — **DONE
2026-08-02 by deletion, not conversion** (entropy audit R-04, finding **F-13**;
Phase 10 plan 10-05, decision **D-07b**). The handler that imported the BIP-84
account **private** key into Core's `wallet.dat` had no caller anywhere, LND is the
wallet the UI drives, and the endpoint was authenticated *and* password-gated — so
it was deleted outright rather than rewritten watch-only. `bitcoin.rs`'s wallet-init
handler and its `dispatcher.rs` arm are gone; **no daemon code path writes the
BIP-84 private key into Bitcoin Core.** No migration was performed or is needed —
a 4-node fleet census found no wallet the handler created. D-09's key-origin
requirement moved to the PSBT itself: `lnd.create-psbt` now reports
`key_origin` (`psbt_key_origin_report`, `api/rpc/lnd/wallet.rs`).
**Read `docs/security/KEY-03-SIGNING-POSTURE.md` for the current state** — it also
records the verdict that **no fleet node is provisioned watch-only**, so what ships
today is PSBT *transport*, not air-gapped custody.
- [ ] **Finish the Core-wallet fleet census — 6 nodes unchecked** (Phase 10 plan 10-05,
Task 3; standing item). The 2026-08-02 census examined 4 nodes (archi-dev-box,
shorty-s/.228, archy-x250-beta, archy-x250-pa) and found **no** wallet created by
the deleted handler and no wallet holding keys or funds. Six were not examined:
framework-pt, archipelago-1, archipelago, archy-dev-pa and archipelago-5
(SSH auth/connectivity) and archy-x250-dev (offline). Re-run the **read-only**
procedure in `docs/security/KEY-03-SIGNING-POSTURE.md` § *Fleet census* when
credentials or connectivity allow — a natural fold-in for KEY-04's on-node work.
**Never run `listdescriptors true`** (it returns private keys). If any node reports
a wallet named `archipelago`, or any descriptor wallet with
`private_keys_enabled: true` that is not blank/empty, **stop and escalate — do not
migrate or modify it** (D-07b).
- [ ] **PSBT-first signing: Phases 2-7 rollout**
(`docs/security/PSBT-SIGNING-ARCHITECTURE.md` §8) — the spec is written to be
consumed directly by `/gsd-plan-phase`, with per-phase goals, dependencies,
candidate requirements and hardware gating. Sequence: PSBT construct/export →
external-signer import + finalize → air-gap transport (BC-UR v2 primary, BBQr for
Coldcard, file fallback always) → `wsh(sortedmulti)` multisig on BIP-48 → LND
remote signing → hot-wallet spend limits and cold/warm/hot tiering. Two hard rules
the spec fixes in place: a channel-funding PSBT must **never** be self-broadcast
(funds can be lost), and no UI copy may imply a routing node's Lightning channel
keys are cold — they are necessarily hot. Phases 3-6 need real hardware.
- [ ] **Confine the seed-bearing RPCs to loopback/TLS** (entropy audit R-07, finding
F-04 / [ARCHY-4]). The 24-word master mnemonic is returned to the browser over
JSON-RPC (`core/archipelago/src/api/rpc/seed_rpc.rs:147`, `:156-158`), held in
process memory under a 10-minute TTL (`:27`) and deliberately **not** cleared at
verify time (`:205-211`, with a documented and defensible rationale about client
retries) — over a transport that is plaintext HTTP on LAN by design
(`api/rpc/mod.rs:227-241`). Anyone with LAN traffic visibility during onboarding
reads the phrase that unlocks the wallet and the node identity. Fix: force TLS or
loopback for seed methods, shrink the TTL, and clear on an acknowledged verify
with a short grace window. Touches the onboarding transport — needs a phase.
- [~] **Multinode test pass** (`docs/multinode-testing-plan.md`) — worked the
preconditions on .198 2026-07-01:
- ✅ cleared 2 stale failed-unit records (`archy-mempool-db.service`,
`meshtastic.service` — both `not-found`/dead since 6 and 5 days ago, harmless
bookkeeping, `systemctl --user reset-failed`).
- ✅ nginx `/app/lnd/` proxy target confirmed correct (→ `18083`, matches the
running `archy-lnd-ui` port) — the plan's "stale proxy target" concern doesn't
apply here.
- ⛔ .198 disk (448GB) is below the 1TB archival threshold + was only 21%
through IBD — user chose to **swap in a different node** rather than wait/add
storage. **.116 ruled out** (no bitcoin container installed at all, just the
UI companion). **.120 ruled out** (reserved for another developer). **.5**
(archy-x250-beta, Tailscale `100.72.136.5`) chosen: also sub-1TB (472GB, so
still pruned — that ceiling is shared by every non-.228 node), but **fully
synced** (`ibd:false`, blocks==headers 956,240). Bootstrapped bats 1.11.1 +
jq 1.7.1 onto it 2026-07-01 and **launched the 5× destructive gate
(`ARCHY_ITERATIONS=5 ARCHY_ALLOW_DESTRUCTIVE=1`) — running now**, log at
`/tmp/gate.log` on .5, background poller watching for the `RESULTS` banner.
- Once .5's gate reports: bring the rest of the fleet to precondition, then the
cross-node federation/mesh/transport suites. This is the literal
"next exit criterion" called out in `CLAUDE.md`.
- [ ] **Phase-3 Quadlet default-flip** — code is validated + opt-in via
`ARCHIPELAGO_USE_QUADLET_BACKENDS=true` on .228/.198 already (confirmed live
2026-07-01). Ready to flip (`config.rs:256` + its test) the moment the .5 gate
reports clean — deliberately NOT staged uncommitted in the tree (a prior attempt
left an uncommitted flip sitting around and that caused confusion; it's a 2-line
change, faster to just do it fresh once confirmed).
- [x] ~~Per-app test coverage for the ~30 apps with zero automated coverage~~ —
**reframed 2026-07-01, mostly a non-issue.** `all-apps-matrix.bats` +
`all-apps-lifecycle.bats` already give EVERY installed app generic baseline
coverage (no stuck state, no error state, stop/start/restart survives, UI
reachable). The real gap is narrower: **34 apps lack app-specific assertions**
(health endpoints, API queryability, data integrity) beyond that baseline —
aiui, bitcoin-core, botfights, core-lightning, did-wallet, fedimint-clientd,
fedimint-gateway, fips-ui, gitea, grafana, home-assistant, indeedhub (+5
sub-containers), jellyfin, lightning-stack, lnd-ui, morphos-server, netbird
(+2 sub-containers), nextcloud, nostr-rs-relay, photoprism, portainer, router,
searxng, strfry, uptime-kuma, vaultwarden. Not urgent — baseline coverage is
real safety net; treat as a backlog "nice to harden further," not a gate item.
- [x] ~~Convert remaining multi-container legacy stacks to the manifest-owned model~~
**investigated 2026-07-01, DONE, nothing left.** All 5 real multi-container
stacks (btcpay, mempool, immich, netbird, indeedhub) are on the
`install_stack_via_orchestrator` pattern (`stacks.rs`). saleor was removed from
the codebase; portainer/home-assistant/grafana are single-container
manifest-driven apps, never stacks; fedimint/fedimint-gateway/fedimint-clientd
are 3 separate single-container apps with manifest dependency edges, not a
coordinated stack. Workstream A's stack-migration tail is fully closed.
- [ ] **Container thrashing/flapping + reconciler churn** (added 2026-07-04 — was
implicit across other tracks, now an explicit pre-tag concern). The root cause
of restart-storm flapping is pre-Quadlet architecture: restarting
`archipelago.service` SIGKILLs every container in its cgroup, then the
reconciler rebuilds the world over several minutes (the post-OTA health check
deliberately skips per-app container assertions because of exactly this).
Consolidated lever list, in order of impact:
- **Phase-3 Quadlet default-flip** (tracked above) — removes the SIGKILL-the-world
behavior entirely; the single biggest fix.
- **Workstream F lifecycle items** — immich/grafana uninstall hangs + ghost
containers, grafana reinstall stops, fedimint guardian sync
(`docs/PRODUCTION-MASTER-PLAN.md` workstream F).
- **Reconciler churn observability** — no metric/log today distinguishes "settling
after restart" from "flapping"; add a per-app restart counter + log line when an
app restarts >N times in M minutes so thrash is visible instead of anecdotal.
- **Failed-unit self-healing gap (observed live 2026-07-06 on .228)**: fedimint's
quadlet unit exited 255 at 21:21 and sat `failed` for 7+ hours — the reconciler
never revived it (it repairs missing/drifted containers but doesn't
`reset-failed`+start failed .services). Same for the indeedhub trio after the
gate run. The health monitor also can't help (container is gone when the unit
fails). Add a reconcile step: quadlet-backed app whose .service is `failed` and
not user-stopped → reset-failed + start, with backoff.
- Already landed, don't re-do: boot-reconciler circuit breaker (2026-07-01),
indeedhub crashloop fix (2026-07-01), async blocking-Command pass (`4c75bb3d`,
removes executor stalls that made the API janky under reconcile load),
quadlet entrypoint-split false-drift fix (2026-07-08 — `container_command_drifted`
compared entrypoint/cmd halves separately, but quadlet folds `sh -lc` into
`Entrypoint=sh` + `Exec=-lc …`, so every quadlet-created app with a
multi-element entrypoint read as permanently drifted; electrumx on .228
recreated 114×/6h until the comparator was switched to concatenated argv).
- Perf polish riding along: 93 MB frontend dist shrink (hardening plan §D 🟡).
- [ ] **Developer tooling CLI suite** (validate/render/local-install/lifecycle-test) —
APP-PACKAGING-MIGRATION-PLAN.md step 5, needed before external devs can publish.
- [x] ~~**Consolidated deploy 2026-07-01**: merged PR #67 (reticulum daemon
process-group fix, `469b0203`), the UI/UX work (`8256fde1` — mesh/web5/apps
layout, modal, search UX), and `archy-openwrt` (TollGate/OpenWrt gateway
integration — new `core/openwrt` crate, RPC surface, `OpenWrtGateway.vue`)
into `main`, alongside the indeedhub self-heal fix~~ — all merged clean, no
conflicts. **Found + fixed 2 real build-breaking issues during
verification, not caught by whoever authored them**: a vestigial unused
`ref` in `Web5ConnectedNodes.vue` that broke `vue-tsc`, and a stale
`MeshMap.test.ts` mock missing `federatedPositions` (predated this
session's Mesh Map feature) that crashed on mount. Full test suite green
(667 passed) after fixes. **Deployed fleet-wide 2026-07-01, all 5 nodes
sha256-verified**: .116, .198, .228, .5 (recovered cleanly from one
truncated-transfer hiccup, caught via checksum before it hit the live
service), 100.82.34.38 (non-Quadlet node — all containers survived the
restart intact, unlike the worst-case risk flagged beforehand). Also
built an unbundled installer ISO from this same merged source
(`archipelago-installer-1.7.99-alpha-unbundled-x86_64.iso`, 2.4GB) —
the ISO pipeline was archived from the release process at v1.7.43-alpha
(OTA tarballs are now primary) but the wrapper script still works.
- [ ] **⚠️ NOT YET DEPLOYED — start here next session.** After the fleet deploy
above, found that PR #67 ("kill whole daemon process group on drop",
branch `fix/reticulum-daemon-process-group`, head `be50c886`) is a
**different, separate** reticulum-daemon fix from the one already
deployed (`469b0203` on `fix/reticulum-daemon-pdeathsig`) — I'd
conflated the two by topic similarity and only merged/deployed the
Python-level `pdeathsig` fix, missing PR #67's Rust-level
kill-whole-process-group-on-`Drop` fix entirely. Merged PR #67 into
`main` (`7a7fec21`, clean, `cargo check` green, complementary not
conflicting with the already-deployed fix) and separately fixed a real
bug found live: `OpenWrtGateway.vue`'s back button had no `@click`
handler at all (`7d7ba573`, `vue-tsc` clean). **Both committed + pushed
to `main` but genuinely NOT deployed to any node** — user asked to hold
off deploying to restart their computer. Also spot-checked
`openwrt.scan` live on .116: RPC plumbing works, but no physical
OpenWrt router was available to confirm true-positive detection, and
`detect::scan_subnet` does blocking TCP/SSH calls inside an `async fn`
with no `.await` — untested at scale, worth hardening. **Next steps**:
build release binary + frontend from current `main`, deploy to all 5
fleet nodes (.116/.198/.228/.5/100.82.34.38) the same way as the
earlier consolidated deploy, then verify the back button + (if a real
OpenWrt router is available) router detection live.
- [~] **Cross-node federation/mesh/transport suites** — **big find 2026-07-01: these
already exist**, just aren't wired into the gate or documented as existing:
`tests/multinode/smoke.sh` (federation pairing/sync, FIPS anchor, peer content
browse, tombstone-removal regression tests), `tests/multinode/meshtastic.sh`
(8-stage on-air mesh test), harness in `tests/multinode/lib/multinode.bash`.
**Actually ran `smoke.sh` live against .116↔.228 2026-07-01: 14 passed, 1
failed, 1 skipped.** Confirms federation pairing (both directions), FIPS
anchor connectivity (both nodes), and peer-content-browse-over-mesh (the
v1.7.95 fix) all genuinely work node-to-node right now.
- ⚠️ **Real robustness gap found**: `node_rpc()` in `tests/multinode/lib/multinode.bash`
has no `--max-time` on its curl calls — a slow server-side RPC hangs the whole
suite with zero feedback (this is what looked like a hang before it eventually
completed on its own). Cheap fix, not yet applied.
- 🐛 **Real regression found and root-caused**: removing a federation node
(`federation.remove-node`) doesn't reliably stick — B reappeared in A's peer
list after removal in the live test. Root cause: `remove_node()`
(`core/archipelago/src/federation/storage.rs:187`) does
`let _ = tombstone_did(data_dir, did).await` — **silently swallows the
tombstone write's errors.** If that write fails (disk I/O, permission,
transient issue), the peer is removed from `nodes.json` but never actually
tombstoned, so the next background sync/notify-join re-adds it — the
tombstone check at `handlers.rs:592-599` passes because the DID was never
recorded as removed. Diagnosed as a **pre-existing logic gap**, not a fresh
regression from the v1.7.95 fix. **Not fixed yet** — this is federation/trust
code, deliberately not touching it blind; needs a careful fix (surface the
tombstone-write failure instead of swallowing it, and/or retry) plus
re-verification with `smoke.sh` before considering it closed.
## Tier 3 — Blocked on a decision or resource only you can supply
- [x] ~~Version naming decision~~**decided 2026-07-08: `1.8.0-alpha`.** Remaining
work is the mechanical bump + tag + push once the pre-tag items above close.
- [x] ~~Workstream B signing ceremony~~**done 2026-07-02.** `anchor.rs` pins
`RELEASE_ROOT_PUBKEY_HEX = 5d15cbee…9951` (signer
`did:key:z6MkkidEnEpo6qHMCNSZoNKWtvQvxq3whnaME9wGgEFhq7ur`); mnemonic held
offline per `docs/workstream-b-signing-runbook.md`.
- [ ] **Bitcoin multi-version fleet-wide OTA**`.228` fully working on branch,
per your prior gating this rollout is explicitly held for your decision on
timing (`docs/bitcoin-version-bulletproof-rollout.md`).
- [ ] **3ccc stock-Meshtastic RF validation** — needs a live send/receive test with
physical radios in your hands; code fix is in place, just unverified live.
## Backlog — deferred, no scope decided, low priority
- [ ] **Marketplace protocol (workstream C)** — design-only (`docs/marketplace-protocol.md`),
no tooling/trust UX built. Future work, not urgent.
- [ ] **DHT distribution (workstream D)** — confirmed design-only, no code
(`docs/dht-distribution-design.md` explicitly says "Status: Design (no code yet)");
an experimental iroh provider skeleton exists behind a feature flag for future
PoC measurement, nothing fleet-facing.
- [ ] **Custom live voice-call protocol** — deprioritized 2026-07-01 per user request;
scope not yet decided. Revisit after the tiers above are worked down.
---
*Historical narrative and detailed per-session logs remain in
`docs/archive/SESSION-1.8.0-OTA-PROGRESS.md` and `docs/PRODUCTION-MASTER-PLAN.md` §6/§8b —
this doc is the live "what's left, in priority order" list. Update it (don't just
append to the old docs) as items close or new ones surface.*