Task 3 of plan 10-05, run by the operator over Tailscale on 2026-08-02 using the read-only procedure in KEY-03-SIGNING-POSTURE.md. No escalation: nothing found. Examined and CLEAR (4): archi-dev-box, shorty-s/.228, archy-x250-beta, archy-x250-pa. On every one there is no wallet named `archipelago` — the deleted handler's default wallet_name — `listwallets` returns only the unnamed default, and that default reports blank=true, keypoolsize=0, txcount=0, balance=0. The only named wallets are Fedimint gatewayd-*. The result holds across two container vintages (bitcoin-knots and bitcoin-core), so it is not four copies of one image behaving identically. Not examined (6), recorded with reasons rather than omitted: framework-pt, archipelago-1, archipelago and archy-dev-pa (SSH permission denied — password rotated/not held), archipelago-5 (timed out during banner exchange), and archy-x250-dev (offline). Password auth was deliberately not attempted: several fleet nodes lock PAM quickly on a wrong password, and locking out an in-use production node is a worse outcome than an incomplete census. The conclusion is stated at the strength the evidence supports — no *examined* node holds a wallet the deleted handler created, and no examined node holds any wallet with keys or funds. It is deliberately NOT generalised to "the fleet is clear" while six nodes are unknown. F-13 is closed by deletion regardless: the code that could create such a wallet is gone from every future build. No key material appeared in any output and `listdescriptors true` was never run. Also corrects the now-stale R-04/F-13 entry in UNIFIED-TASK-TRACKER.md, which still described `handle_bitcoin_init_wallet_from_seed` and a watch-only migration as pending work — that code no longer exists. Marks it done-by- deletion and adds the six unchecked nodes as a standing item, flagged as a natural fold-in for KEY-04's on-node work but tracked independently so it does not vanish if KEY-04 is re-scoped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
436 lines
31 KiB
Markdown
436 lines
31 KiB
Markdown
# Unified Task Tracker — OTA 1.8.0 + Master Plan
|
||
|
||
Single working list for everything left before 1.8.0 ships and the next master-plan
|
||
exit criteria (multinode + workstreams B/C/D) are met. Supersedes the open-task
|
||
sections of `docs/archive/SESSION-1.8.0-OTA-PROGRESS.md` and `docs/PRODUCTION-MASTER-PLAN.md`
|
||
as the day-to-day tracker — those docs remain the historical record / detailed
|
||
narrative and are still linked from here where useful. **Ordered fastest/simplest
|
||
first** so we work top-down instead of hunting across docs.
|
||
|
||
Verified against actual code state on 2026-07-01 (not just doc text — several
|
||
items the source docs still listed as "open" turned out to already be shipped;
|
||
those are marked ✅ below with the commit that did it, so we stop re-litigating them).
|
||
|
||
---
|
||
|
||
## Tier 0 — Quick / mechanical, no blockers
|
||
|
||
- [ ] **Ship the lightning payment false-failure fix in the next release** (fixed
|
||
on main 2026-07-27, needs OTA). Slow multi-hop payments (>15s) surfaced as
|
||
"Payment failed" while LND settled them in the background — the shared LND
|
||
REST client's 15s timeout aborted the synchronous `/v1/channels/transactions`
|
||
wait. Now: payinvoice decodes the invoice first for its payment hash, waits
|
||
up to 120s on a dedicated client, returns `status: "pending"` (never a
|
||
failure) on timeout, and the new `lnd.paymentstatus` RPC + frontend
|
||
`payLightningInvoice()` helper poll to a real terminal state (all 5 UI call
|
||
sites migrated). Verify on Framework PT with a real multi-hop payment.
|
||
- [ ] **Show the app version on the companion mobile-app banner in the app store
|
||
and on its install/pairing modal** (user request 2026-07-27) — so it's
|
||
obvious at a glance whether the node is serving the latest APK build.
|
||
- [ ] **Optimise the companion QR scan — quicker + better** (user request
|
||
2026-07-27; deferred to a later session on purpose). The pairing/scan QR
|
||
flow works (user-verified on-device 2026-07-27) but should get faster and
|
||
smoother: quicker camera start + decode (scan resolution/framerate,
|
||
continuous autofocus), more forgiving in low light / at an angle, and
|
||
snappier feedback once the code locks. Touch the native-scan path from
|
||
PR #104 and the in-app scan modal together so both benefit.
|
||
|
||
- [ ] **Update `tests/lifecycle/TESTING.md`'s stale Release Gates checklist** (lines
|
||
289–296) — several boxes are unchecked but actually true now:
|
||
- #1 bitcoin-stops: covered by `tests/lifecycle/bats/bitcoin-knots.bats` stop/restart
|
||
tier, included in the 5/5 green gate run.
|
||
- #2 `ARCHY_ITERATIONS=5` on .228: **GREEN 2026-06-23 per CLAUDE.md** — check the box.
|
||
- #5 cargo 0 warnings: confirmed 0 warnings on `cargo build --release` (2026-07-01).
|
||
- #7 layman changelog: `CHANGELOG.md` is backfilled with layman-readable entries
|
||
through v1.8.00-alpha — check the box.
|
||
- Leave #3 (multinode), #4 (backend-survives-restart / Phase-3 default-on), #6
|
||
(LoC decision), #8 (tag pushed) unchecked — genuinely still open, see Tier 2/3.
|
||
- [x] ~~Finish the archival/full-node manifest generalization~~ — investigated 2026-07-01:
|
||
the hardcoded fallback names in `dependencies.rs:48-52` (`electrs`, `mempool-electrs`,
|
||
`mempool-web`) are legacy **alias** ids for `electrumx`/`mempool`, resolved via
|
||
id-mapping in a dozen other places (`install.rs`, `runtime.rs`, `config.rs`, etc.),
|
||
not separate un-migrated apps with their own manifests. `electrumx` and `mempool`
|
||
themselves already declare `bitcoin:archival`. The fallback is correct as-is —
|
||
not tech debt, closing this item rather than risk breaking alias resolution.
|
||
- [x] ~~Confirm/close the Portainer image-pin item~~ — confirmed 2026-07-01:
|
||
`146.59.87.168:3000/lfg2025/portainer:2.19.4` is present in `podman images` on
|
||
all 3 LAN nodes (.116/.198/.228), i.e. actually resolvable/pulled from the mirror.
|
||
Not a live bug.
|
||
- [x] ~~grafana Quadlet "stuck activating"~~ — checked live on .116 (2026-07-01):
|
||
`grafana.service` is `active (running)`, container `Up 2 hours (healthy)`. The
|
||
2026-06-21 report is stale for grafana. **strfry still unconfirmed** — not
|
||
installed on any of .116/.198/.228 to check directly; low priority until someone
|
||
actually needs it installed.
|
||
|
||
- [ ] **Add `cargo audit` / `cargo deny` to CI, failing on duplicate `rand` majors**
|
||
(entropy audit R-05, finding F-07 —
|
||
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). `cargo-audit` is not installed
|
||
anywhere, so no RustSec check has ever run against this tree. Separately,
|
||
`cargo tree` shows **both** `rand 0.8.5` (direct, all first-party key generation)
|
||
and `rand 0.9.2` (transitive via `totp-rs` and `tungstenite 0.26.2`) resolved into
|
||
one binary. `rand 0.9.0` removed `ThreadRng` fork protection and the orchestrator
|
||
forks constantly, so a future bump must be visible rather than silent — add a
|
||
`bans` rule so the duplicate majors show up in CI, not in an incident.
|
||
|
||
- [ ] **Harden the release signing ceremony's mnemonic input** (entropy audit R-08,
|
||
finding F-06). `ceremony gen` prints the release master mnemonic to **stdout**
|
||
(`core/archipelago/src/ceremony.rs:71-77`) and `load_release_root_key` prefers the
|
||
`RELEASE_MASTER_MNEMONIC` **environment variable** over stdin (`:157-160`) — both
|
||
leak into shell history, `/proc/<pid>/environ`, tmux scrollback and terminal
|
||
recordings. This is the seed that derives the fleet release-root signing key, so a
|
||
leak means forged signed manifests fleet-wide. Make stdin/TTY the only supported
|
||
input for `sign`/`pubkey`; write `gen`'s output to a `0600` file rather than the
|
||
terminal. Small change, but schedule it deliberately — it is the signing ceremony.
|
||
|
||
- [ ] **Small entropy-audit hygiene batch** (entropy audit R-09 – R-12, R-14 —
|
||
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). Five independent one-liners,
|
||
each closing a Low/Informational finding:
|
||
- Persist the CSPRNG-readiness verdict (`seed.rs:85-91`) as a durable structured
|
||
event, so any node can answer post-hoc "was the entropy pool ready when this seed
|
||
was born?" — the question Coldcard owners cannot answer today.
|
||
- Add a test asserting the `getrandom` crate uses the **blocking** syscall, making
|
||
`seed.rs:52-57`'s invariant mechanical instead of a comment.
|
||
- Clear `_seed_words` from `sessionStorage` on route-leave from onboarding, not only
|
||
on successful verify (`OnboardingSeedVerify.vue:251`), plus a wall-clock expiry
|
||
mirroring the server's 10-minute `MNEMONIC_TTL`.
|
||
- Replace `% charset.len()` in `totp.rs:305` with `SliceRandom::choose(&mut OsRng)`.
|
||
(No bias today — 32 divides 256 — but any future charset edit introduces one
|
||
silently. The audit refutes the research's claim that this is currently biased.)
|
||
- Comment `pickRandomIndices` (`OnboardingSeedVerify.vue:157`) to record that its
|
||
`Math.random()` picks a UX challenge, not key material, so the next auditor does
|
||
not re-derive that it is benign.
|
||
|
||
- [ ] ~~**Swap container `generated_secrets` to explicit `OsRng`** (entropy audit R-13,
|
||
finding F-10) — two-line change in `container/secrets.rs:90-102`~~
|
||
**SUPERSEDED 2026-08-02 by R-16 / KEY-05.** The audit scoped this at 2 call sites; the
|
||
real surface is **41 across 15 files** — see the audit's new §F-10a. `secrets.rs` is 2
|
||
of them, and a two-line fix there while 39 other sites inherit the same dependency
|
||
default is not a fix.
|
||
|
||
- [ ] **Crate-wide CSPRNG enforcement — a defaulted RNG cannot be inherited anywhere**
|
||
(entropy audit **R-16 / F-10a**, Medium) — tracked as **KEY-05 in Phase 10**, so plan
|
||
and execute it there rather than as a standalone item. `session.rs` (16 sites),
|
||
`pine_ha.rs` (6), `wallet/bdhke.rs` (2 prod — **Cashu proof secret + blinding factor,
|
||
genuine key material**), `storage_crypto.rs` (1 — **AEAD nonce**), `mesh/x3dh.rs` (2 —
|
||
prekey *identifiers*, **not** key material — corrected 2026-08-02), +10 more files.
|
||
Nothing is broken today (`rand::random()`/`thread_rng()` are ChaCha12 from
|
||
`getrandom(2)`), but it is the T1 shape that produced the COLDCARD defect, now with key
|
||
material in the blast radius. Five layers: sealed allowlist trait at key-gen seams;
|
||
`clippy.toml` `disallowed-methods` ban (compile-time, CI-enforced — no `clippy.toml`
|
||
exists yet); `cargo-deny` on duplicate `rand` majors (absorbs R-05); degenerate-entropy
|
||
runtime check; persist the CSPRNG-readiness verdict (absorbs R-09). Also retires the
|
||
`impl rand::CryptoRng for CountingRng` false promise at `seed.rs:656`.
|
||
**Gated: do not start until the concurrent Phase 1 agent is done and synced.**
|
||
|
||
## Tier 1 — Medium effort, unblocked
|
||
|
||
- [ ] **Fix the fail-open first-boot secret regeneration in the ISO** (entropy audit
|
||
R-02 + R-03, finding F-03 — `docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`).
|
||
The installed rootfs is a **cached container export shared by every node**
|
||
(`image-recipe/_archived/build-auto-installer-iso.sh:717-726`, extracted at
|
||
`:2303`), and it bakes SSH host keys (via the `openssh-server` install at `:345`)
|
||
and a TLS keypair (`:463-469`). `archipelago-first-boot-secrets.service` correctly
|
||
regenerates both per device — but both branches are **fail-open** (`:1647`,
|
||
`:1659`) and `touch "$MARKER"` at `:1663` runs **unconditionally**, so a single
|
||
transient failure permanently leaves that node on the image-wide shared SSH host
|
||
key and TLS private key, with the failure visible only in a log file. Fix:
|
||
(a) set the marker only when both regenerations succeeded, so it retries next
|
||
boot; (b) surface the failure in the UI/doctor, not just the log; (c) strip the
|
||
baked keys from the rootfs tar so a failure degrades to "no key" rather than
|
||
"shared key". Needs an ISO rebuild and two fresh flashes to verify.
|
||
|
||
- [ ] **Reconcile `Argon2::default()` with ADR-005** (entropy audit R-06, finding F-05).
|
||
ADR-005 states 64 MB / 3 iterations
|
||
(`docs/adr/005-chacha20-backup-encryption.md:31`); `Argon2::default()` in
|
||
argon2 0.5.3 is Argon2id at **19 MiB / t=2 / p=1**. Used at
|
||
`core/archipelago/src/seed.rs:249` and `:285`, `backup/identity.rs:38`/`:93`,
|
||
`backup/full.rs:618`/`:650`. Either raise the parameters behind a versioned
|
||
envelope **with a migration** (an existing `master_seed.enc` was encrypted under
|
||
the old parameters and will not decrypt under new ones) or amend the ADR to state
|
||
the real numbers. Do not change them silently.
|
||
|
||
- [ ] **Run the on-node entropy verification checklist** (entropy audit R-15, §6 of
|
||
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). Everything in that section is
|
||
explicitly **UNVERIFIED** — it needs real hardware this environment cannot reach.
|
||
Highest value first: **C-3** (are SSH host-key and TLS fingerprints actually
|
||
different across two nodes flashed from the same ISO?) and **C-5** (the cross-node
|
||
same-ISO seed collision test — the empirical check that would have caught the
|
||
Coldcard defect). Also C-1 (`crng init done` vs seed-generation timestamp), C-2
|
||
(`machine-id` uniqueness), C-4 (what the rootfs tar actually contains, run on the
|
||
build host), C-6 (is `/rpc/v1` reachable unauthenticated from the LAN). Use a
|
||
disposable node — C-5 overwrites node identity.
|
||
|
||
- [x] ~~immich → Quadlet migration~~ — investigated 2026-07-01, turned out already done:
|
||
immich uses the same `install_stack_via_orchestrator` primitive as netbird/btcpay
|
||
(`immich_stack_app_ids()` in `stacks.rs:690`), and is confirmed running as real
|
||
Quadlet units live on .228 (`immich_server.container`, `immich_postgres.container`,
|
||
`immich_redis.container`, all active). Not a legacy in-cgroup app — the only
|
||
remaining piece is the fleet-wide Phase-3 default-flip, already tracked in Tier 2.
|
||
- [x] ~~Netbird reinstall adoption path~~ — investigated 2026-07-01, **not a bug, by
|
||
design.** `adopt_stack_if_exists()` (`stacks.rs:140-198`) is only used as a
|
||
fallback when the orchestrator has no manifest for the app — there's nothing to
|
||
render certs/config from in that case, so skipping rendering is correct. When
|
||
the orchestrator *does* have the manifest (the normal path), the reconcile loop
|
||
already re-renders certs even for adopted-running containers, fixed in
|
||
`4519dbf0` (`prod_orchestrator.rs:1707-1708`).
|
||
- [x] ~~TanStack Query (or equivalent) investigation~~ — spike complete 2026-07-01,
|
||
**recommendation: don't adopt / close as not needed.** Only 3 stores actually fetch
|
||
data, WebSocket push already handles hot data (server-info/package-data), no
|
||
cache-invalidation or stale-data bugs found, migration would touch 62 RPC call
|
||
sites for no concrete payoff. If boilerplate ever bothers us, extract a
|
||
`usePolling()` composable instead — much cheaper than a query-cache migration.
|
||
|
||
## Tier 2 — High effort, mostly unblocked (the actual next exit criteria)
|
||
|
||
- [ ] **🔴 Gate the unauthenticated seed RPCs** (entropy audit R-01, finding **F-01,
|
||
Critical** — `docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). `seed.generate`,
|
||
`seed.verify`, `seed.restore` and `seed.save-encrypted` are in
|
||
`UNAUTHENTICATED_METHODS` (`core/archipelago/src/api/rpc/middleware.rs:24-28`),
|
||
which skips session, RBAC **and** CSRF (`api/rpc/mod.rs:263`, `:295`, `:326`).
|
||
Neither handler checks whether onboarding is already complete
|
||
(`api/rpc/seed_rpc.rs:93-159`, `:226-305`), and `NodeIdentity::from_seed`
|
||
overwrites `node_key`, `nostr_secret` and the FIPS mesh key **unconditionally**
|
||
(`identity.rs:79-114`). There is no rate limit (`rate_limit.rs:60-97` has no
|
||
`seed.*` entry). The endpoint is proxied to the LAN over plaintext HTTP
|
||
(`image-recipe/configs/nginx-archipelago.conf:11`, `:165`, `:192`) and mesh peers
|
||
can reach it too (`server.rs:2080` asserts `/rpc/v1` passes the peer path filter).
|
||
Net: **one unauthenticated POST can take over or destroy a live node's identity**,
|
||
and `seed.restore` lets the attacker choose the mnemonic. The guard already exists
|
||
and is simply never called — `NodeIdentity::key_exists` (`identity.rs:117`).
|
||
Fix: bail when a node key exists and no onboarding mnemonic is pending; prefer
|
||
also gating on `auth_manager.is_onboarding_complete()`; add rate limits at
|
||
`auth.changePassword` strictness; narrow the peer path filter. Changes an
|
||
authentication boundary on a live fleet — **needs its own `/gsd-plan-phase` with a
|
||
federation re-verify**, not an opportunistic patch.
|
||
|
||
- [x] **PSBT-first signing: Phase 1 — move the Bitcoin private key out of Core** — **DONE
|
||
2026-08-02 by deletion, not conversion** (entropy audit R-04, finding **F-13**;
|
||
Phase 10 plan 10-05, decision **D-07b**). The handler that imported the BIP-84
|
||
account **private** key into Core's `wallet.dat` had no caller anywhere, LND is the
|
||
wallet the UI drives, and the endpoint was authenticated *and* password-gated — so
|
||
it was deleted outright rather than rewritten watch-only. `bitcoin.rs`'s wallet-init
|
||
handler and its `dispatcher.rs` arm are gone; **no daemon code path writes the
|
||
BIP-84 private key into Bitcoin Core.** No migration was performed or is needed —
|
||
a 4-node fleet census found no wallet the handler created. D-09's key-origin
|
||
requirement moved to the PSBT itself: `lnd.create-psbt` now reports
|
||
`key_origin` (`psbt_key_origin_report`, `api/rpc/lnd/wallet.rs`).
|
||
**Read `docs/security/KEY-03-SIGNING-POSTURE.md` for the current state** — it also
|
||
records the verdict that **no fleet node is provisioned watch-only**, so what ships
|
||
today is PSBT *transport*, not air-gapped custody.
|
||
|
||
- [ ] **Finish the Core-wallet fleet census — 6 nodes unchecked** (Phase 10 plan 10-05,
|
||
Task 3; standing item). The 2026-08-02 census examined 4 nodes (archi-dev-box,
|
||
shorty-s/.228, archy-x250-beta, archy-x250-pa) and found **no** wallet created by
|
||
the deleted handler and no wallet holding keys or funds. Six were not examined:
|
||
framework-pt, archipelago-1, archipelago, archy-dev-pa and archipelago-5
|
||
(SSH auth/connectivity) and archy-x250-dev (offline). Re-run the **read-only**
|
||
procedure in `docs/security/KEY-03-SIGNING-POSTURE.md` § *Fleet census* when
|
||
credentials or connectivity allow — a natural fold-in for KEY-04's on-node work.
|
||
**Never run `listdescriptors true`** (it returns private keys). If any node reports
|
||
a wallet named `archipelago`, or any descriptor wallet with
|
||
`private_keys_enabled: true` that is not blank/empty, **stop and escalate — do not
|
||
migrate or modify it** (D-07b).
|
||
|
||
- [ ] **PSBT-first signing: Phases 2-7 rollout**
|
||
(`docs/security/PSBT-SIGNING-ARCHITECTURE.md` §8) — the spec is written to be
|
||
consumed directly by `/gsd-plan-phase`, with per-phase goals, dependencies,
|
||
candidate requirements and hardware gating. Sequence: PSBT construct/export →
|
||
external-signer import + finalize → air-gap transport (BC-UR v2 primary, BBQr for
|
||
Coldcard, file fallback always) → `wsh(sortedmulti)` multisig on BIP-48 → LND
|
||
remote signing → hot-wallet spend limits and cold/warm/hot tiering. Two hard rules
|
||
the spec fixes in place: a channel-funding PSBT must **never** be self-broadcast
|
||
(funds can be lost), and no UI copy may imply a routing node's Lightning channel
|
||
keys are cold — they are necessarily hot. Phases 3-6 need real hardware.
|
||
|
||
- [ ] **Confine the seed-bearing RPCs to loopback/TLS** (entropy audit R-07, finding
|
||
F-04 / [ARCHY-4]). The 24-word master mnemonic is returned to the browser over
|
||
JSON-RPC (`core/archipelago/src/api/rpc/seed_rpc.rs:147`, `:156-158`), held in
|
||
process memory under a 10-minute TTL (`:27`) and deliberately **not** cleared at
|
||
verify time (`:205-211`, with a documented and defensible rationale about client
|
||
retries) — over a transport that is plaintext HTTP on LAN by design
|
||
(`api/rpc/mod.rs:227-241`). Anyone with LAN traffic visibility during onboarding
|
||
reads the phrase that unlocks the wallet and the node identity. Fix: force TLS or
|
||
loopback for seed methods, shrink the TTL, and clear on an acknowledged verify
|
||
with a short grace window. Touches the onboarding transport — needs a phase.
|
||
|
||
- [~] **Multinode test pass** (`docs/multinode-testing-plan.md`) — worked the
|
||
preconditions on .198 2026-07-01:
|
||
- ✅ cleared 2 stale failed-unit records (`archy-mempool-db.service`,
|
||
`meshtastic.service` — both `not-found`/dead since 6 and 5 days ago, harmless
|
||
bookkeeping, `systemctl --user reset-failed`).
|
||
- ✅ nginx `/app/lnd/` proxy target confirmed correct (→ `18083`, matches the
|
||
running `archy-lnd-ui` port) — the plan's "stale proxy target" concern doesn't
|
||
apply here.
|
||
- ⛔ .198 disk (448GB) is below the 1TB archival threshold + was only 21%
|
||
through IBD — user chose to **swap in a different node** rather than wait/add
|
||
storage. **.116 ruled out** (no bitcoin container installed at all, just the
|
||
UI companion). **.120 ruled out** (reserved for another developer). **.5**
|
||
(archy-x250-beta, Tailscale `100.72.136.5`) chosen: also sub-1TB (472GB, so
|
||
still pruned — that ceiling is shared by every non-.228 node), but **fully
|
||
synced** (`ibd:false`, blocks==headers 956,240). Bootstrapped bats 1.11.1 +
|
||
jq 1.7.1 onto it 2026-07-01 and **launched the 5× destructive gate
|
||
(`ARCHY_ITERATIONS=5 ARCHY_ALLOW_DESTRUCTIVE=1`) — running now**, log at
|
||
`/tmp/gate.log` on .5, background poller watching for the `RESULTS` banner.
|
||
- Once .5's gate reports: bring the rest of the fleet to precondition, then the
|
||
cross-node federation/mesh/transport suites. This is the literal
|
||
"next exit criterion" called out in `CLAUDE.md`.
|
||
- [ ] **Phase-3 Quadlet default-flip** — code is validated + opt-in via
|
||
`ARCHIPELAGO_USE_QUADLET_BACKENDS=true` on .228/.198 already (confirmed live
|
||
2026-07-01). Ready to flip (`config.rs:256` + its test) the moment the .5 gate
|
||
reports clean — deliberately NOT staged uncommitted in the tree (a prior attempt
|
||
left an uncommitted flip sitting around and that caused confusion; it's a 2-line
|
||
change, faster to just do it fresh once confirmed).
|
||
- [x] ~~Per-app test coverage for the ~30 apps with zero automated coverage~~ —
|
||
**reframed 2026-07-01, mostly a non-issue.** `all-apps-matrix.bats` +
|
||
`all-apps-lifecycle.bats` already give EVERY installed app generic baseline
|
||
coverage (no stuck state, no error state, stop/start/restart survives, UI
|
||
reachable). The real gap is narrower: **34 apps lack app-specific assertions**
|
||
(health endpoints, API queryability, data integrity) beyond that baseline —
|
||
aiui, bitcoin-core, botfights, core-lightning, did-wallet, fedimint-clientd,
|
||
fedimint-gateway, fips-ui, gitea, grafana, home-assistant, indeedhub (+5
|
||
sub-containers), jellyfin, lightning-stack, lnd-ui, morphos-server, netbird
|
||
(+2 sub-containers), nextcloud, nostr-rs-relay, photoprism, portainer, router,
|
||
searxng, strfry, uptime-kuma, vaultwarden. Not urgent — baseline coverage is
|
||
real safety net; treat as a backlog "nice to harden further," not a gate item.
|
||
- [x] ~~Convert remaining multi-container legacy stacks to the manifest-owned model~~ —
|
||
**investigated 2026-07-01, DONE, nothing left.** All 5 real multi-container
|
||
stacks (btcpay, mempool, immich, netbird, indeedhub) are on the
|
||
`install_stack_via_orchestrator` pattern (`stacks.rs`). saleor was removed from
|
||
the codebase; portainer/home-assistant/grafana are single-container
|
||
manifest-driven apps, never stacks; fedimint/fedimint-gateway/fedimint-clientd
|
||
are 3 separate single-container apps with manifest dependency edges, not a
|
||
coordinated stack. Workstream A's stack-migration tail is fully closed.
|
||
- [ ] **Container thrashing/flapping + reconciler churn** (added 2026-07-04 — was
|
||
implicit across other tracks, now an explicit pre-tag concern). The root cause
|
||
of restart-storm flapping is pre-Quadlet architecture: restarting
|
||
`archipelago.service` SIGKILLs every container in its cgroup, then the
|
||
reconciler rebuilds the world over several minutes (the post-OTA health check
|
||
deliberately skips per-app container assertions because of exactly this).
|
||
Consolidated lever list, in order of impact:
|
||
- **Phase-3 Quadlet default-flip** (tracked above) — removes the SIGKILL-the-world
|
||
behavior entirely; the single biggest fix.
|
||
- **Workstream F lifecycle items** — immich/grafana uninstall hangs + ghost
|
||
containers, grafana reinstall stops, fedimint guardian sync
|
||
(`docs/PRODUCTION-MASTER-PLAN.md` workstream F).
|
||
- **Reconciler churn observability** — no metric/log today distinguishes "settling
|
||
after restart" from "flapping"; add a per-app restart counter + log line when an
|
||
app restarts >N times in M minutes so thrash is visible instead of anecdotal.
|
||
- **Failed-unit self-healing gap (observed live 2026-07-06 on .228)**: fedimint's
|
||
quadlet unit exited 255 at 21:21 and sat `failed` for 7+ hours — the reconciler
|
||
never revived it (it repairs missing/drifted containers but doesn't
|
||
`reset-failed`+start failed .services). Same for the indeedhub trio after the
|
||
gate run. The health monitor also can't help (container is gone when the unit
|
||
fails). Add a reconcile step: quadlet-backed app whose .service is `failed` and
|
||
not user-stopped → reset-failed + start, with backoff.
|
||
- Already landed, don't re-do: boot-reconciler circuit breaker (2026-07-01),
|
||
indeedhub crashloop fix (2026-07-01), async blocking-Command pass (`4c75bb3d`,
|
||
removes executor stalls that made the API janky under reconcile load),
|
||
quadlet entrypoint-split false-drift fix (2026-07-08 — `container_command_drifted`
|
||
compared entrypoint/cmd halves separately, but quadlet folds `sh -lc` into
|
||
`Entrypoint=sh` + `Exec=-lc …`, so every quadlet-created app with a
|
||
multi-element entrypoint read as permanently drifted; electrumx on .228
|
||
recreated 114×/6h until the comparator was switched to concatenated argv).
|
||
- Perf polish riding along: 93 MB frontend dist shrink (hardening plan §D 🟡).
|
||
- [ ] **Developer tooling CLI suite** (validate/render/local-install/lifecycle-test) —
|
||
APP-PACKAGING-MIGRATION-PLAN.md step 5, needed before external devs can publish.
|
||
- [x] ~~**Consolidated deploy 2026-07-01**: merged PR #67 (reticulum daemon
|
||
process-group fix, `469b0203`), the UI/UX work (`8256fde1` — mesh/web5/apps
|
||
layout, modal, search UX), and `archy-openwrt` (TollGate/OpenWrt gateway
|
||
integration — new `core/openwrt` crate, RPC surface, `OpenWrtGateway.vue`)
|
||
into `main`, alongside the indeedhub self-heal fix~~ — all merged clean, no
|
||
conflicts. **Found + fixed 2 real build-breaking issues during
|
||
verification, not caught by whoever authored them**: a vestigial unused
|
||
`ref` in `Web5ConnectedNodes.vue` that broke `vue-tsc`, and a stale
|
||
`MeshMap.test.ts` mock missing `federatedPositions` (predated this
|
||
session's Mesh Map feature) that crashed on mount. Full test suite green
|
||
(667 passed) after fixes. **Deployed fleet-wide 2026-07-01, all 5 nodes
|
||
sha256-verified**: .116, .198, .228, .5 (recovered cleanly from one
|
||
truncated-transfer hiccup, caught via checksum before it hit the live
|
||
service), 100.82.34.38 (non-Quadlet node — all containers survived the
|
||
restart intact, unlike the worst-case risk flagged beforehand). Also
|
||
built an unbundled installer ISO from this same merged source
|
||
(`archipelago-installer-1.7.99-alpha-unbundled-x86_64.iso`, 2.4GB) —
|
||
the ISO pipeline was archived from the release process at v1.7.43-alpha
|
||
(OTA tarballs are now primary) but the wrapper script still works.
|
||
- [ ] **⚠️ NOT YET DEPLOYED — start here next session.** After the fleet deploy
|
||
above, found that PR #67 ("kill whole daemon process group on drop",
|
||
branch `fix/reticulum-daemon-process-group`, head `be50c886`) is a
|
||
**different, separate** reticulum-daemon fix from the one already
|
||
deployed (`469b0203` on `fix/reticulum-daemon-pdeathsig`) — I'd
|
||
conflated the two by topic similarity and only merged/deployed the
|
||
Python-level `pdeathsig` fix, missing PR #67's Rust-level
|
||
kill-whole-process-group-on-`Drop` fix entirely. Merged PR #67 into
|
||
`main` (`7a7fec21`, clean, `cargo check` green, complementary not
|
||
conflicting with the already-deployed fix) and separately fixed a real
|
||
bug found live: `OpenWrtGateway.vue`'s back button had no `@click`
|
||
handler at all (`7d7ba573`, `vue-tsc` clean). **Both committed + pushed
|
||
to `main` but genuinely NOT deployed to any node** — user asked to hold
|
||
off deploying to restart their computer. Also spot-checked
|
||
`openwrt.scan` live on .116: RPC plumbing works, but no physical
|
||
OpenWrt router was available to confirm true-positive detection, and
|
||
`detect::scan_subnet` does blocking TCP/SSH calls inside an `async fn`
|
||
with no `.await` — untested at scale, worth hardening. **Next steps**:
|
||
build release binary + frontend from current `main`, deploy to all 5
|
||
fleet nodes (.116/.198/.228/.5/100.82.34.38) the same way as the
|
||
earlier consolidated deploy, then verify the back button + (if a real
|
||
OpenWrt router is available) router detection live.
|
||
- [~] **Cross-node federation/mesh/transport suites** — **big find 2026-07-01: these
|
||
already exist**, just aren't wired into the gate or documented as existing:
|
||
`tests/multinode/smoke.sh` (federation pairing/sync, FIPS anchor, peer content
|
||
browse, tombstone-removal regression tests), `tests/multinode/meshtastic.sh`
|
||
(8-stage on-air mesh test), harness in `tests/multinode/lib/multinode.bash`.
|
||
**Actually ran `smoke.sh` live against .116↔.228 2026-07-01: 14 passed, 1
|
||
failed, 1 skipped.** Confirms federation pairing (both directions), FIPS
|
||
anchor connectivity (both nodes), and peer-content-browse-over-mesh (the
|
||
v1.7.95 fix) all genuinely work node-to-node right now.
|
||
- ⚠️ **Real robustness gap found**: `node_rpc()` in `tests/multinode/lib/multinode.bash`
|
||
has no `--max-time` on its curl calls — a slow server-side RPC hangs the whole
|
||
suite with zero feedback (this is what looked like a hang before it eventually
|
||
completed on its own). Cheap fix, not yet applied.
|
||
- 🐛 **Real regression found and root-caused**: removing a federation node
|
||
(`federation.remove-node`) doesn't reliably stick — B reappeared in A's peer
|
||
list after removal in the live test. Root cause: `remove_node()`
|
||
(`core/archipelago/src/federation/storage.rs:187`) does
|
||
`let _ = tombstone_did(data_dir, did).await` — **silently swallows the
|
||
tombstone write's errors.** If that write fails (disk I/O, permission,
|
||
transient issue), the peer is removed from `nodes.json` but never actually
|
||
tombstoned, so the next background sync/notify-join re-adds it — the
|
||
tombstone check at `handlers.rs:592-599` passes because the DID was never
|
||
recorded as removed. Diagnosed as a **pre-existing logic gap**, not a fresh
|
||
regression from the v1.7.95 fix. **Not fixed yet** — this is federation/trust
|
||
code, deliberately not touching it blind; needs a careful fix (surface the
|
||
tombstone-write failure instead of swallowing it, and/or retry) plus
|
||
re-verification with `smoke.sh` before considering it closed.
|
||
|
||
## Tier 3 — Blocked on a decision or resource only you can supply
|
||
|
||
- [x] ~~Version naming decision~~ — **decided 2026-07-08: `1.8.0-alpha`.** Remaining
|
||
work is the mechanical bump + tag + push once the pre-tag items above close.
|
||
- [x] ~~Workstream B signing ceremony~~ — **done 2026-07-02.** `anchor.rs` pins
|
||
`RELEASE_ROOT_PUBKEY_HEX = 5d15cbee…9951` (signer
|
||
`did:key:z6MkkidEnEpo6qHMCNSZoNKWtvQvxq3whnaME9wGgEFhq7ur`); mnemonic held
|
||
offline per `docs/workstream-b-signing-runbook.md`.
|
||
- [ ] **Bitcoin multi-version fleet-wide OTA** — `.228` fully working on branch,
|
||
per your prior gating this rollout is explicitly held for your decision on
|
||
timing (`docs/bitcoin-version-bulletproof-rollout.md`).
|
||
- [ ] **3ccc stock-Meshtastic RF validation** — needs a live send/receive test with
|
||
physical radios in your hands; code fix is in place, just unverified live.
|
||
|
||
## Backlog — deferred, no scope decided, low priority
|
||
|
||
- [ ] **Marketplace protocol (workstream C)** — design-only (`docs/marketplace-protocol.md`),
|
||
no tooling/trust UX built. Future work, not urgent.
|
||
- [ ] **DHT distribution (workstream D)** — confirmed design-only, no code
|
||
(`docs/dht-distribution-design.md` explicitly says "Status: Design (no code yet)");
|
||
an experimental iroh provider skeleton exists behind a feature flag for future
|
||
PoC measurement, nothing fleet-facing.
|
||
- [ ] **Custom live voice-call protocol** — deprioritized 2026-07-01 per user request;
|
||
scope not yet decided. Revisit after the tiers above are worked down.
|
||
|
||
---
|
||
|
||
*Historical narrative and detailed per-session logs remain in
|
||
`docs/archive/SESSION-1.8.0-OTA-PROGRESS.md` and `docs/PRODUCTION-MASTER-PLAN.md` §6/§8b —
|
||
this doc is the live "what's left, in priority order" list. Update it (don't just
|
||
append to the old docs) as items close or new ones surface.*
|