Files
archy/docs/UNIFIED-TASK-TRACKER.md
T

436 lines
31 KiB
Markdown
Raw Normal View History

# Unified Task Tracker — OTA 1.8.0 + Master Plan
Single working list for everything left before 1.8.0 ships and the next master-plan
exit criteria (multinode + workstreams B/C/D) are met. Supersedes the open-task
sections of `docs/archive/SESSION-1.8.0-OTA-PROGRESS.md` and `docs/PRODUCTION-MASTER-PLAN.md`
as the day-to-day tracker — those docs remain the historical record / detailed
narrative and are still linked from here where useful. **Ordered fastest/simplest
first** so we work top-down instead of hunting across docs.
Verified against actual code state on 2026-07-01 (not just doc text — several
items the source docs still listed as "open" turned out to already be shipped;
those are marked ✅ below with the commit that did it, so we stop re-litigating them).
---
## Tier 0 — Quick / mechanical, no blockers
- [ ] **Ship the lightning payment false-failure fix in the next release** (fixed
on main 2026-07-27, needs OTA). Slow multi-hop payments (>15s) surfaced as
"Payment failed" while LND settled them in the background — the shared LND
REST client's 15s timeout aborted the synchronous `/v1/channels/transactions`
wait. Now: payinvoice decodes the invoice first for its payment hash, waits
up to 120s on a dedicated client, returns `status: "pending"` (never a
failure) on timeout, and the new `lnd.paymentstatus` RPC + frontend
`payLightningInvoice()` helper poll to a real terminal state (all 5 UI call
sites migrated). Verify on Framework PT with a real multi-hop payment.
- [ ] **Show the app version on the companion mobile-app banner in the app store
and on its install/pairing modal** (user request 2026-07-27) — so it's
obvious at a glance whether the node is serving the latest APK build.
- [ ] **Optimise the companion QR scan — quicker + better** (user request
2026-07-27; deferred to a later session on purpose). The pairing/scan QR
flow works (user-verified on-device 2026-07-27) but should get faster and
smoother: quicker camera start + decode (scan resolution/framerate,
continuous autofocus), more forgiving in low light / at an angle, and
snappier feedback once the code locks. Touch the native-scan path from
PR #104 and the in-app scan modal together so both benefit.
- [ ] **Update `tests/lifecycle/TESTING.md`'s stale Release Gates checklist** (lines
289296) — several boxes are unchecked but actually true now:
- #1 bitcoin-stops: covered by `tests/lifecycle/bats/bitcoin-knots.bats` stop/restart
tier, included in the 5/5 green gate run.
- #2 `ARCHY_ITERATIONS=5` on .228: **GREEN 2026-06-23 per CLAUDE.md** — check the box.
- #5 cargo 0 warnings: confirmed 0 warnings on `cargo build --release` (2026-07-01).
- #7 layman changelog: `CHANGELOG.md` is backfilled with layman-readable entries
through v1.8.00-alpha — check the box.
- Leave #3 (multinode), #4 (backend-survives-restart / Phase-3 default-on), #6
(LoC decision), #8 (tag pushed) unchecked — genuinely still open, see Tier 2/3.
- [x] ~~Finish the archival/full-node manifest generalization~~ — investigated 2026-07-01:
the hardcoded fallback names in `dependencies.rs:48-52` (`electrs`, `mempool-electrs`,
`mempool-web`) are legacy **alias** ids for `electrumx`/`mempool`, resolved via
id-mapping in a dozen other places (`install.rs`, `runtime.rs`, `config.rs`, etc.),
not separate un-migrated apps with their own manifests. `electrumx` and `mempool`
themselves already declare `bitcoin:archival`. The fallback is correct as-is —
not tech debt, closing this item rather than risk breaking alias resolution.
- [x] ~~Confirm/close the Portainer image-pin item~~ — confirmed 2026-07-01:
`146.59.87.168:3000/lfg2025/portainer:2.19.4` is present in `podman images` on
all 3 LAN nodes (.116/.198/.228), i.e. actually resolvable/pulled from the mirror.
Not a live bug.
- [x] ~~grafana Quadlet "stuck activating"~~ — checked live on .116 (2026-07-01):
`grafana.service` is `active (running)`, container `Up 2 hours (healthy)`. The
2026-06-21 report is stale for grafana. **strfry still unconfirmed** — not
installed on any of .116/.198/.228 to check directly; low priority until someone
actually needs it installed.
- [ ] **Add `cargo audit` / `cargo deny` to CI, failing on duplicate `rand` majors**
(entropy audit R-05, finding F-07 —
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). `cargo-audit` is not installed
anywhere, so no RustSec check has ever run against this tree. Separately,
`cargo tree` shows **both** `rand 0.8.5` (direct, all first-party key generation)
and `rand 0.9.2` (transitive via `totp-rs` and `tungstenite 0.26.2`) resolved into
one binary. `rand 0.9.0` removed `ThreadRng` fork protection and the orchestrator
forks constantly, so a future bump must be visible rather than silent — add a
`bans` rule so the duplicate majors show up in CI, not in an incident.
- [ ] **Harden the release signing ceremony's mnemonic input** (entropy audit R-08,
finding F-06). `ceremony gen` prints the release master mnemonic to **stdout**
(`core/archipelago/src/ceremony.rs:71-77`) and `load_release_root_key` prefers the
`RELEASE_MASTER_MNEMONIC` **environment variable** over stdin (`:157-160`) — both
leak into shell history, `/proc/<pid>/environ`, tmux scrollback and terminal
recordings. This is the seed that derives the fleet release-root signing key, so a
leak means forged signed manifests fleet-wide. Make stdin/TTY the only supported
input for `sign`/`pubkey`; write `gen`'s output to a `0600` file rather than the
terminal. Small change, but schedule it deliberately — it is the signing ceremony.
- [ ] **Small entropy-audit hygiene batch** (entropy audit R-09 R-12, R-14 —
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). Five independent one-liners,
each closing a Low/Informational finding:
- Persist the CSPRNG-readiness verdict (`seed.rs:85-91`) as a durable structured
event, so any node can answer post-hoc "was the entropy pool ready when this seed
was born?" — the question Coldcard owners cannot answer today.
- Add a test asserting the `getrandom` crate uses the **blocking** syscall, making
`seed.rs:52-57`'s invariant mechanical instead of a comment.
- Clear `_seed_words` from `sessionStorage` on route-leave from onboarding, not only
on successful verify (`OnboardingSeedVerify.vue:251`), plus a wall-clock expiry
mirroring the server's 10-minute `MNEMONIC_TTL`.
- Replace `% charset.len()` in `totp.rs:305` with `SliceRandom::choose(&mut OsRng)`.
(No bias today — 32 divides 256 — but any future charset edit introduces one
silently. The audit refutes the research's claim that this is currently biased.)
- Comment `pickRandomIndices` (`OnboardingSeedVerify.vue:157`) to record that its
`Math.random()` picks a UX challenge, not key material, so the next auditor does
not re-derive that it is benign.
- [ ] ~~**Swap container `generated_secrets` to explicit `OsRng`** (entropy audit R-13,
finding F-10) — two-line change in `container/secrets.rs:90-102`~~
**SUPERSEDED 2026-08-02 by R-16 / KEY-05.** The audit scoped this at 2 call sites; the
real surface is **41 across 15 files** — see the audit's new §F-10a. `secrets.rs` is 2
of them, and a two-line fix there while 39 other sites inherit the same dependency
default is not a fix.
- [ ] **Crate-wide CSPRNG enforcement — a defaulted RNG cannot be inherited anywhere**
(entropy audit **R-16 / F-10a**, Medium) — tracked as **KEY-05 in Phase 10**, so plan
and execute it there rather than as a standalone item. `session.rs` (16 sites),
`pine_ha.rs` (6), `wallet/bdhke.rs` (2 prod — **Cashu proof secret + blinding factor,
genuine key material**), `storage_crypto.rs` (1 — **AEAD nonce**), `mesh/x3dh.rs` (2 —
prekey *identifiers*, **not** key material — corrected 2026-08-02), +10 more files.
Nothing is broken today (`rand::random()`/`thread_rng()` are ChaCha12 from
`getrandom(2)`), but it is the T1 shape that produced the COLDCARD defect, now with key
material in the blast radius. Five layers: sealed allowlist trait at key-gen seams;
`clippy.toml` `disallowed-methods` ban (compile-time, CI-enforced — no `clippy.toml`
exists yet); `cargo-deny` on duplicate `rand` majors (absorbs R-05); degenerate-entropy
runtime check; persist the CSPRNG-readiness verdict (absorbs R-09). Also retires the
`impl rand::CryptoRng for CountingRng` false promise at `seed.rs:656`.
**Gated: do not start until the concurrent Phase 1 agent is done and synced.**
## Tier 1 — Medium effort, unblocked
- [ ] **Fix the fail-open first-boot secret regeneration in the ISO** (entropy audit
R-02 + R-03, finding F-03 — `docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`).
The installed rootfs is a **cached container export shared by every node**
(`image-recipe/_archived/build-auto-installer-iso.sh:717-726`, extracted at
`:2303`), and it bakes SSH host keys (via the `openssh-server` install at `:345`)
and a TLS keypair (`:463-469`). `archipelago-first-boot-secrets.service` correctly
regenerates both per device — but both branches are **fail-open** (`:1647`,
`:1659`) and `touch "$MARKER"` at `:1663` runs **unconditionally**, so a single
transient failure permanently leaves that node on the image-wide shared SSH host
key and TLS private key, with the failure visible only in a log file. Fix:
(a) set the marker only when both regenerations succeeded, so it retries next
boot; (b) surface the failure in the UI/doctor, not just the log; (c) strip the
baked keys from the rootfs tar so a failure degrades to "no key" rather than
"shared key". Needs an ISO rebuild and two fresh flashes to verify.
- [ ] **Reconcile `Argon2::default()` with ADR-005** (entropy audit R-06, finding F-05).
ADR-005 states 64 MB / 3 iterations
(`docs/adr/005-chacha20-backup-encryption.md:31`); `Argon2::default()` in
argon2 0.5.3 is Argon2id at **19 MiB / t=2 / p=1**. Used at
`core/archipelago/src/seed.rs:249` and `:285`, `backup/identity.rs:38`/`:93`,
`backup/full.rs:618`/`:650`. Either raise the parameters behind a versioned
envelope **with a migration** (an existing `master_seed.enc` was encrypted under
the old parameters and will not decrypt under new ones) or amend the ADR to state
the real numbers. Do not change them silently.
- [ ] **Run the on-node entropy verification checklist** (entropy audit R-15, §6 of
`docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). Everything in that section is
explicitly **UNVERIFIED** — it needs real hardware this environment cannot reach.
Highest value first: **C-3** (are SSH host-key and TLS fingerprints actually
different across two nodes flashed from the same ISO?) and **C-5** (the cross-node
same-ISO seed collision test — the empirical check that would have caught the
Coldcard defect). Also C-1 (`crng init done` vs seed-generation timestamp), C-2
(`machine-id` uniqueness), C-4 (what the rootfs tar actually contains, run on the
build host), C-6 (is `/rpc/v1` reachable unauthenticated from the LAN). Use a
disposable node — C-5 overwrites node identity.
- [x] ~~immich → Quadlet migration~~ — investigated 2026-07-01, turned out already done:
immich uses the same `install_stack_via_orchestrator` primitive as netbird/btcpay
(`immich_stack_app_ids()` in `stacks.rs:690`), and is confirmed running as real
Quadlet units live on .228 (`immich_server.container`, `immich_postgres.container`,
`immich_redis.container`, all active). Not a legacy in-cgroup app — the only
remaining piece is the fleet-wide Phase-3 default-flip, already tracked in Tier 2.
- [x] ~~Netbird reinstall adoption path~~ — investigated 2026-07-01, **not a bug, by
design.** `adopt_stack_if_exists()` (`stacks.rs:140-198`) is only used as a
fallback when the orchestrator has no manifest for the app — there's nothing to
render certs/config from in that case, so skipping rendering is correct. When
the orchestrator *does* have the manifest (the normal path), the reconcile loop
already re-renders certs even for adopted-running containers, fixed in
`4519dbf0` (`prod_orchestrator.rs:1707-1708`).
- [x] ~~TanStack Query (or equivalent) investigation~~ — spike complete 2026-07-01,
**recommendation: don't adopt / close as not needed.** Only 3 stores actually fetch
data, WebSocket push already handles hot data (server-info/package-data), no
cache-invalidation or stale-data bugs found, migration would touch 62 RPC call
sites for no concrete payoff. If boilerplate ever bothers us, extract a
`usePolling()` composable instead — much cheaper than a query-cache migration.
## Tier 2 — High effort, mostly unblocked (the actual next exit criteria)
- [ ] **🔴 Gate the unauthenticated seed RPCs** (entropy audit R-01, finding **F-01,
Critical** — `docs/security/ENTROPY-SEED-AUDIT-2026-07-31.md`). `seed.generate`,
`seed.verify`, `seed.restore` and `seed.save-encrypted` are in
`UNAUTHENTICATED_METHODS` (`core/archipelago/src/api/rpc/middleware.rs:24-28`),
which skips session, RBAC **and** CSRF (`api/rpc/mod.rs:263`, `:295`, `:326`).
Neither handler checks whether onboarding is already complete
(`api/rpc/seed_rpc.rs:93-159`, `:226-305`), and `NodeIdentity::from_seed`
overwrites `node_key`, `nostr_secret` and the FIPS mesh key **unconditionally**
(`identity.rs:79-114`). There is no rate limit (`rate_limit.rs:60-97` has no
`seed.*` entry). The endpoint is proxied to the LAN over plaintext HTTP
(`image-recipe/configs/nginx-archipelago.conf:11`, `:165`, `:192`) and mesh peers
can reach it too (`server.rs:2080` asserts `/rpc/v1` passes the peer path filter).
Net: **one unauthenticated POST can take over or destroy a live node's identity**,
and `seed.restore` lets the attacker choose the mnemonic. The guard already exists
and is simply never called — `NodeIdentity::key_exists` (`identity.rs:117`).
Fix: bail when a node key exists and no onboarding mnemonic is pending; prefer
also gating on `auth_manager.is_onboarding_complete()`; add rate limits at
`auth.changePassword` strictness; narrow the peer path filter. Changes an
authentication boundary on a live fleet — **needs its own `/gsd-plan-phase` with a
federation re-verify**, not an opportunistic patch.
- [x] **PSBT-first signing: Phase 1 — move the Bitcoin private key out of Core** — **DONE
2026-08-02 by deletion, not conversion** (entropy audit R-04, finding **F-13**;
Phase 10 plan 10-05, decision **D-07b**). The handler that imported the BIP-84
account **private** key into Core's `wallet.dat` had no caller anywhere, LND is the
wallet the UI drives, and the endpoint was authenticated *and* password-gated — so
it was deleted outright rather than rewritten watch-only. `bitcoin.rs`'s wallet-init
handler and its `dispatcher.rs` arm are gone; **no daemon code path writes the
BIP-84 private key into Bitcoin Core.** No migration was performed or is needed —
a 4-node fleet census found no wallet the handler created. D-09's key-origin
requirement moved to the PSBT itself: `lnd.create-psbt` now reports
`key_origin` (`psbt_key_origin_report`, `api/rpc/lnd/wallet.rs`).
**Read `docs/security/KEY-03-SIGNING-POSTURE.md` for the current state** — it also
records the verdict that **no fleet node is provisioned watch-only**, so what ships
today is PSBT *transport*, not air-gapped custody.
- [ ] **Finish the Core-wallet fleet census — 6 nodes unchecked** (Phase 10 plan 10-05,
Task 3; standing item). The 2026-08-02 census examined 4 nodes (archi-dev-box,
shorty-s/.228, archy-x250-beta, archy-x250-pa) and found **no** wallet created by
the deleted handler and no wallet holding keys or funds. Six were not examined:
framework-pt, archipelago-1, archipelago, archy-dev-pa and archipelago-5
(SSH auth/connectivity) and archy-x250-dev (offline). Re-run the **read-only**
procedure in `docs/security/KEY-03-SIGNING-POSTURE.md` § *Fleet census* when
credentials or connectivity allow — a natural fold-in for KEY-04's on-node work.
**Never run `listdescriptors true`** (it returns private keys). If any node reports
a wallet named `archipelago`, or any descriptor wallet with
`private_keys_enabled: true` that is not blank/empty, **stop and escalate — do not
migrate or modify it** (D-07b).
- [ ] **PSBT-first signing: Phases 2-7 rollout**
(`docs/security/PSBT-SIGNING-ARCHITECTURE.md` §8) — the spec is written to be
consumed directly by `/gsd-plan-phase`, with per-phase goals, dependencies,
candidate requirements and hardware gating. Sequence: PSBT construct/export →
external-signer import + finalize → air-gap transport (BC-UR v2 primary, BBQr for
Coldcard, file fallback always) → `wsh(sortedmulti)` multisig on BIP-48 → LND
remote signing → hot-wallet spend limits and cold/warm/hot tiering. Two hard rules
the spec fixes in place: a channel-funding PSBT must **never** be self-broadcast
(funds can be lost), and no UI copy may imply a routing node's Lightning channel
keys are cold — they are necessarily hot. Phases 3-6 need real hardware.
- [ ] **Confine the seed-bearing RPCs to loopback/TLS** (entropy audit R-07, finding
F-04 / [ARCHY-4]). The 24-word master mnemonic is returned to the browser over
JSON-RPC (`core/archipelago/src/api/rpc/seed_rpc.rs:147`, `:156-158`), held in
process memory under a 10-minute TTL (`:27`) and deliberately **not** cleared at
verify time (`:205-211`, with a documented and defensible rationale about client
retries) — over a transport that is plaintext HTTP on LAN by design
(`api/rpc/mod.rs:227-241`). Anyone with LAN traffic visibility during onboarding
reads the phrase that unlocks the wallet and the node identity. Fix: force TLS or
loopback for seed methods, shrink the TTL, and clear on an acknowledged verify
with a short grace window. Touches the onboarding transport — needs a phase.
- [~] **Multinode test pass** (`docs/multinode-testing-plan.md`) — worked the
preconditions on .198 2026-07-01:
- ✅ cleared 2 stale failed-unit records (`archy-mempool-db.service`,
`meshtastic.service` — both `not-found`/dead since 6 and 5 days ago, harmless
bookkeeping, `systemctl --user reset-failed`).
- ✅ nginx `/app/lnd/` proxy target confirmed correct (→ `18083`, matches the
running `archy-lnd-ui` port) — the plan's "stale proxy target" concern doesn't
apply here.
- ⛔ .198 disk (448GB) is below the 1TB archival threshold + was only 21%
through IBD — user chose to **swap in a different node** rather than wait/add
storage. **.116 ruled out** (no bitcoin container installed at all, just the
UI companion). **.120 ruled out** (reserved for another developer). **.5**
(archy-x250-beta, Tailscale `100.72.136.5`) chosen: also sub-1TB (472GB, so
still pruned — that ceiling is shared by every non-.228 node), but **fully
synced** (`ibd:false`, blocks==headers 956,240). Bootstrapped bats 1.11.1 +
jq 1.7.1 onto it 2026-07-01 and **launched the 5× destructive gate
(`ARCHY_ITERATIONS=5 ARCHY_ALLOW_DESTRUCTIVE=1`) — running now**, log at
`/tmp/gate.log` on .5, background poller watching for the `RESULTS` banner.
- Once .5's gate reports: bring the rest of the fleet to precondition, then the
cross-node federation/mesh/transport suites. This is the literal
"next exit criterion" called out in `CLAUDE.md`.
- [ ] **Phase-3 Quadlet default-flip** — code is validated + opt-in via
`ARCHIPELAGO_USE_QUADLET_BACKENDS=true` on .228/.198 already (confirmed live
2026-07-01). Ready to flip (`config.rs:256` + its test) the moment the .5 gate
reports clean — deliberately NOT staged uncommitted in the tree (a prior attempt
left an uncommitted flip sitting around and that caused confusion; it's a 2-line
change, faster to just do it fresh once confirmed).
- [x] ~~Per-app test coverage for the ~30 apps with zero automated coverage~~ —
**reframed 2026-07-01, mostly a non-issue.** `all-apps-matrix.bats` +
`all-apps-lifecycle.bats` already give EVERY installed app generic baseline
coverage (no stuck state, no error state, stop/start/restart survives, UI
reachable). The real gap is narrower: **34 apps lack app-specific assertions**
(health endpoints, API queryability, data integrity) beyond that baseline —
aiui, bitcoin-core, botfights, core-lightning, did-wallet, fedimint-clientd,
fedimint-gateway, fips-ui, gitea, grafana, home-assistant, indeedhub (+5
sub-containers), jellyfin, lightning-stack, lnd-ui, morphos-server, netbird
(+2 sub-containers), nextcloud, nostr-rs-relay, photoprism, portainer, router,
searxng, strfry, uptime-kuma, vaultwarden. Not urgent — baseline coverage is
real safety net; treat as a backlog "nice to harden further," not a gate item.
- [x] ~~Convert remaining multi-container legacy stacks to the manifest-owned model~~
**investigated 2026-07-01, DONE, nothing left.** All 5 real multi-container
stacks (btcpay, mempool, immich, netbird, indeedhub) are on the
`install_stack_via_orchestrator` pattern (`stacks.rs`). saleor was removed from
the codebase; portainer/home-assistant/grafana are single-container
manifest-driven apps, never stacks; fedimint/fedimint-gateway/fedimint-clientd
are 3 separate single-container apps with manifest dependency edges, not a
coordinated stack. Workstream A's stack-migration tail is fully closed.
- [ ] **Container thrashing/flapping + reconciler churn** (added 2026-07-04 — was
implicit across other tracks, now an explicit pre-tag concern). The root cause
of restart-storm flapping is pre-Quadlet architecture: restarting
`archipelago.service` SIGKILLs every container in its cgroup, then the
reconciler rebuilds the world over several minutes (the post-OTA health check
deliberately skips per-app container assertions because of exactly this).
Consolidated lever list, in order of impact:
- **Phase-3 Quadlet default-flip** (tracked above) — removes the SIGKILL-the-world
behavior entirely; the single biggest fix.
- **Workstream F lifecycle items** — immich/grafana uninstall hangs + ghost
containers, grafana reinstall stops, fedimint guardian sync
(`docs/PRODUCTION-MASTER-PLAN.md` workstream F).
- **Reconciler churn observability** — no metric/log today distinguishes "settling
after restart" from "flapping"; add a per-app restart counter + log line when an
app restarts >N times in M minutes so thrash is visible instead of anecdotal.
- **Failed-unit self-healing gap (observed live 2026-07-06 on .228)**: fedimint's
quadlet unit exited 255 at 21:21 and sat `failed` for 7+ hours — the reconciler
never revived it (it repairs missing/drifted containers but doesn't
`reset-failed`+start failed .services). Same for the indeedhub trio after the
gate run. The health monitor also can't help (container is gone when the unit
fails). Add a reconcile step: quadlet-backed app whose .service is `failed` and
not user-stopped → reset-failed + start, with backoff.
- Already landed, don't re-do: boot-reconciler circuit breaker (2026-07-01),
indeedhub crashloop fix (2026-07-01), async blocking-Command pass (`4c75bb3d`,
removes executor stalls that made the API janky under reconcile load),
quadlet entrypoint-split false-drift fix (2026-07-08 — `container_command_drifted`
compared entrypoint/cmd halves separately, but quadlet folds `sh -lc` into
`Entrypoint=sh` + `Exec=-lc …`, so every quadlet-created app with a
multi-element entrypoint read as permanently drifted; electrumx on .228
recreated 114×/6h until the comparator was switched to concatenated argv).
- Perf polish riding along: 93 MB frontend dist shrink (hardening plan §D 🟡).
- [ ] **Developer tooling CLI suite** (validate/render/local-install/lifecycle-test) —
APP-PACKAGING-MIGRATION-PLAN.md step 5, needed before external devs can publish.
- [x] ~~**Consolidated deploy 2026-07-01**: merged PR #67 (reticulum daemon
process-group fix, `469b0203`), the UI/UX work (`8256fde1` — mesh/web5/apps
layout, modal, search UX), and `archy-openwrt` (TollGate/OpenWrt gateway
integration — new `core/openwrt` crate, RPC surface, `OpenWrtGateway.vue`)
into `main`, alongside the indeedhub self-heal fix~~ — all merged clean, no
conflicts. **Found + fixed 2 real build-breaking issues during
verification, not caught by whoever authored them**: a vestigial unused
`ref` in `Web5ConnectedNodes.vue` that broke `vue-tsc`, and a stale
`MeshMap.test.ts` mock missing `federatedPositions` (predated this
session's Mesh Map feature) that crashed on mount. Full test suite green
(667 passed) after fixes. **Deployed fleet-wide 2026-07-01, all 5 nodes
sha256-verified**: .116, .198, .228, .5 (recovered cleanly from one
truncated-transfer hiccup, caught via checksum before it hit the live
service), 100.82.34.38 (non-Quadlet node — all containers survived the
restart intact, unlike the worst-case risk flagged beforehand). Also
built an unbundled installer ISO from this same merged source
(`archipelago-installer-1.7.99-alpha-unbundled-x86_64.iso`, 2.4GB) —
the ISO pipeline was archived from the release process at v1.7.43-alpha
(OTA tarballs are now primary) but the wrapper script still works.
- [ ] **⚠️ NOT YET DEPLOYED — start here next session.** After the fleet deploy
above, found that PR #67 ("kill whole daemon process group on drop",
branch `fix/reticulum-daemon-process-group`, head `be50c886`) is a
**different, separate** reticulum-daemon fix from the one already
deployed (`469b0203` on `fix/reticulum-daemon-pdeathsig`) — I'd
conflated the two by topic similarity and only merged/deployed the
Python-level `pdeathsig` fix, missing PR #67's Rust-level
kill-whole-process-group-on-`Drop` fix entirely. Merged PR #67 into
`main` (`7a7fec21`, clean, `cargo check` green, complementary not
conflicting with the already-deployed fix) and separately fixed a real
bug found live: `OpenWrtGateway.vue`'s back button had no `@click`
handler at all (`7d7ba573`, `vue-tsc` clean). **Both committed + pushed
to `main` but genuinely NOT deployed to any node** — user asked to hold
off deploying to restart their computer. Also spot-checked
`openwrt.scan` live on .116: RPC plumbing works, but no physical
OpenWrt router was available to confirm true-positive detection, and
`detect::scan_subnet` does blocking TCP/SSH calls inside an `async fn`
with no `.await` — untested at scale, worth hardening. **Next steps**:
build release binary + frontend from current `main`, deploy to all 5
fleet nodes (.116/.198/.228/.5/100.82.34.38) the same way as the
earlier consolidated deploy, then verify the back button + (if a real
OpenWrt router is available) router detection live.
- [~] **Cross-node federation/mesh/transport suites** — **big find 2026-07-01: these
already exist**, just aren't wired into the gate or documented as existing:
`tests/multinode/smoke.sh` (federation pairing/sync, FIPS anchor, peer content
browse, tombstone-removal regression tests), `tests/multinode/meshtastic.sh`
(8-stage on-air mesh test), harness in `tests/multinode/lib/multinode.bash`.
**Actually ran `smoke.sh` live against .116↔.228 2026-07-01: 14 passed, 1
failed, 1 skipped.** Confirms federation pairing (both directions), FIPS
anchor connectivity (both nodes), and peer-content-browse-over-mesh (the
v1.7.95 fix) all genuinely work node-to-node right now.
- ⚠️ **Real robustness gap found**: `node_rpc()` in `tests/multinode/lib/multinode.bash`
has no `--max-time` on its curl calls — a slow server-side RPC hangs the whole
suite with zero feedback (this is what looked like a hang before it eventually
completed on its own). Cheap fix, not yet applied.
- 🐛 **Real regression found and root-caused**: removing a federation node
(`federation.remove-node`) doesn't reliably stick — B reappeared in A's peer
list after removal in the live test. Root cause: `remove_node()`
(`core/archipelago/src/federation/storage.rs:187`) does
`let _ = tombstone_did(data_dir, did).await` — **silently swallows the
tombstone write's errors.** If that write fails (disk I/O, permission,
transient issue), the peer is removed from `nodes.json` but never actually
tombstoned, so the next background sync/notify-join re-adds it — the
tombstone check at `handlers.rs:592-599` passes because the DID was never
recorded as removed. Diagnosed as a **pre-existing logic gap**, not a fresh
regression from the v1.7.95 fix. **Not fixed yet** — this is federation/trust
code, deliberately not touching it blind; needs a careful fix (surface the
tombstone-write failure instead of swallowing it, and/or retry) plus
re-verification with `smoke.sh` before considering it closed.
## Tier 3 — Blocked on a decision or resource only you can supply
- [x] ~~Version naming decision~~**decided 2026-07-08: `1.8.0-alpha`.** Remaining
work is the mechanical bump + tag + push once the pre-tag items above close.
- [x] ~~Workstream B signing ceremony~~**done 2026-07-02.** `anchor.rs` pins
`RELEASE_ROOT_PUBKEY_HEX = 5d15cbee…9951` (signer
`did:key:z6MkkidEnEpo6qHMCNSZoNKWtvQvxq3whnaME9wGgEFhq7ur`); mnemonic held
offline per `docs/workstream-b-signing-runbook.md`.
- [ ] **Bitcoin multi-version fleet-wide OTA**`.228` fully working on branch,
per your prior gating this rollout is explicitly held for your decision on
timing (`docs/bitcoin-version-bulletproof-rollout.md`).
- [ ] **3ccc stock-Meshtastic RF validation** — needs a live send/receive test with
physical radios in your hands; code fix is in place, just unverified live.
## Backlog — deferred, no scope decided, low priority
- [ ] **Marketplace protocol (workstream C)** — design-only (`docs/marketplace-protocol.md`),
no tooling/trust UX built. Future work, not urgent.
- [ ] **DHT distribution (workstream D)** — confirmed design-only, no code
(`docs/dht-distribution-design.md` explicitly says "Status: Design (no code yet)");
an experimental iroh provider skeleton exists behind a feature flag for future
PoC measurement, nothing fleet-facing.
- [ ] **Custom live voice-call protocol** — deprioritized 2026-07-01 per user request;
scope not yet decided. Revisit after the tiers above are worked down.
---
*Historical narrative and detailed per-session logs remain in
`docs/archive/SESSION-1.8.0-OTA-PROGRESS.md` and `docs/PRODUCTION-MASTER-PLAN.md` §6/§8b —
this doc is the live "what's left, in priority order" list. Update it (don't just
append to the old docs) as items close or new ones surface.*