From 9f65f1e7ae031f781bb0885c3e7becae28697acf Mon Sep 17 00:00:00 2001 From: archipelago Date: Mon, 27 Jul 2026 15:14:23 -0400 Subject: [PATCH] =?UTF-8?q?docs:=20FIPS=20near-100%=20uptime=20+=20optimis?= =?UTF-8?q?tic=20UI=20state=20plan=20=E2=80=94=20live-proven=20root=20caus?= =?UTF-8?q?es=20(nft=205679=20drop,=20.228=20daemon=20skew,=20dead=20LAN?= =?UTF-8?q?=20peering,=20no=20fast-fail)=20+=20phased=20execution?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 --- docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md | 300 ++++++++++++++++++++++++++ 1 file changed, 300 insertions(+) create mode 100644 docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md diff --git a/docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md b/docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md new file mode 100644 index 00000000..45a53f25 --- /dev/null +++ b/docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md @@ -0,0 +1,300 @@ +# FIPS near-100% uptime + optimistic UI state — implementation plan + +**Date:** 2026-07-27. **Status:** researched + root-caused live on the fleet; ready to +implement for the next release. Two workstreams: (A) make node↔node FIPS transport +succeed whenever a FIPS path physically exists, (B) stop the UI reloading everything +on every navigation (optimistic/cached cards, stale-while-revalidate) while keeping +data fresh. + +**Honesty note on "100%":** if a node's network blackholes every anchor (the .116 +WiFi case, `docs/HANDOFF-2026-07-20-fips-peer-files.md:117-133`), Tor fallback is +*correct*. The achievable target is: **FIPS wins whenever a FIPS path exists, and +fallback frequency is measured in-product so regressions are visible.** Today several +paths are 0% FIPS *by construction* regardless of network health — that's the bug. + +--- + +## Part A — why Cloud/FIPS "commonly falls back to Tor": ranked root causes + +All verified live on 2026-07-27 (.116 local, .198, .228, Framework PT, x250s) plus a +full code audit of `core/archipelago/src/{fips,transport,federation,server.rs}`. + +### RC0 — 🔥 The hardening firewall drops the peer-API port on every hardened node (PROVEN) + +The fips0 default-deny baseline (`/etc/fips/fips.nft`) is opened by archipelago's +drop-in `80-web-ui.nft` (`fips/config.rs:236-255`) for **80 + 8443 + app ports only**. +The peer-API listener — which carries *all* federation sync, cloud browse/download, +mesh envelopes, DWN, invoices — is **`PEER_PORT = 5679`** (`fips/dial.rs:35`). +**5679 is not in the allowlist.** The drop-in's own comment claims "web UI + peer +API" but the peer API port was never added. + +Live proof (2026-07-27): +- .116 nft chain: 5,965 dropped packets; .198: **28,670 dropped packets** — that's + peers' FIPS dials dying at the firewall. +- .198 → .116 `GET :5679/health`: **timeout (6s)** before; **HTTP 200 in 0.35s** + after `nft insert rule inet fips inbound iifname fips0 tcp dport 5679 accept`. + Same result in reverse direction (200 in 0.64s). +- Explains the exact fleet split in `federation/nodes.json`: hardened-baseline nodes + (Framework PT, .198, .228, x250-dev, x250-mad2) = `last_transport: tor`; + non-hardened nodes (Austin Sapien, X250-Beta, X250-PA) answer :5679 (404 from the + path allowlist = listener reachable) = `last_transport: fips`. +- Every dial to a hardened peer pays the 8s FIPS connect timeout + (`dial.rs:114`) ×2 (retry, `dial.rs:128-140`) → then Tor. That's the "Cloud takes + forever / shows Tor" experience. + +**Fix (one line + reload):** add `tcp dport 5679 accept` to the drop-in in +`fips/config.rs` (use a constant shared with `dial.rs::PEER_PORT`, not a literal). +The drop-in reinstalls on every daemon config install, so it heals fleet-wide on OTA. +⚠️ Transient manual rules were inserted on .116 and .198 during diagnosis (2026-07-27) +— they vanish on the next `nft -f /etc/fips/fips.nft` reload or reboot; the code fix +makes them permanent. + +### RC1 — .228 (Shorty's) runs fips 0.3.0-dev; the 0.4.1 fleet can't reach it + +.228's daemon: `0.3.0-dev (rev 34e00b9f6e)`, both anchor links "connected", but its +ULA is 100% unreachable from 0.4.1 nodes (ping loss 100%). FIPS wire format is not +stable across revs (`docs/HANDOFF-2026-07-23-companion-apk-deploy.md:78`). Everything +to/from .228 rides Tor no matter what else we fix. + +**Fix:** fleet fips-version audit + upgrade to v0.4.1 everywhere (in-product updater +exists: `fips/update.rs`; .deb path per `reference_vps2_fips_anchor`). Add a version +check to `fips.status` and surface a "peer daemon outdated" warning. + +### RC2 — Direct LAN/endpoint peering is dead code + wrong port + stale seed anchors + +Without direct links, all peer traffic hairpins through the vps2 anchor spanning +tree (observed: .116→.198 cold RTT 1.5–3.5s on the same LAN; also the wedged-anchor +latency-rot incident, `HANDOFF-2026-07-23:141-160`). + +- **G1 — `lan_fips_anchors()` has never run.** It needs `PeerRecord.fips_npub`, but + `PeerRegistry::set_fips_npub` (`transport/mod.rs:302`) has **zero callers** — mDNS + TXT records only carry `did`/`pubkey`/`version` (`transport/lan.rs:50-54`). So the + "co-located peers form a direct link" feature (`anchors.rs:294-305`, + `server.rs:761-766`) is a fleet-wide no-op. +- **G2 — wrong UDP port.** `anchors.rs:293` dials `8668`, but the generated + fips.yaml binds UDP **2121** (`fips/config.rs:187`, `fips/mod.rs:130`). Even if G1 + ran, it would dial a dead port. `.116`'s live `seed-anchors.json` still carries + `.198@192.168.1.198:8668` — **stale IP (LAN renumbered to 192.168.63.x) AND dead + port**; both manual entries are useless today. +- No Tailscale/alternate endpoint fallback when LAN is unreachable (the .116↔.198 + fix of 2026-07-20 was hand-applied per-node config, never productized). + +**Fix:** (a) `FIPS_UDP_PORT` → `crate::fips::PUBLISHED_UDP_PORT` + drift-guard test; +(b) hydrate `fips_npub` into the registry from federation storage (did-keyed join) so +`lan_fips_anchors` goes live with no wire change; (c) advertise the npub in the mDNS +TXT + `set_fips_npub` on resolve as the proper fix; (d) teach the LAN-anchor tick to +also try a peer's Tailscale/last-known-good endpoint when LAN fails (reviewed change +— this area got handoffs wrong twice, per memory). + +### RC3 — No fast-fail on the hottest call sites; retry silently doubles every budget + +- `content.browse-peer` — **the Cloud page** — has NO `fips_timeout` + (`api/rpc/content.rs:363-366`): a cold FIPS path burns up to ~16.6s (8s connect + + 600ms + 8s retry) before Tor even starts, against a UI deadline of 30s + (`Cloud.vue:720`) — and the frontend then retries ×3. Users see errors, not + fallback. 12 call sites total lack `fips_timeout` (browse/download/preview-peer, + `/blob`, DWN, node_message, rotation notifies). +- `dial.rs:128-140` runs 2 full-budget attempts, so `fips_timeout(6s)` really means + ~12.6s everywhere. + +**Fix:** wrap `send_with_retry` in a single `tokio::time::timeout(fips_attempt_timeout())` +(call sites `dial.rs:455`, `dial.rs:488`; halve per-attempt client timeout), then add +`.fips_timeout(...)`: `content.rs:366` (6s), `content.rs:281` (8s), `content.rs:1139` +(6s), `typed_messages.rs:822` (8s), `dwn_sync.rs:188/213/272` (6s), +`node_message.rs:376` (8s), `node_message.rs:412` (4s), `tor/mod.rs:501` (6s), +`federation/handlers.rs:869` (6s). **Skip the three 900s streaming downloads** +(`content.rs:552/870/1061`, `proxy.rs:236`) — `dial.rs:311-319` documents why; the +retry-budget wrap covers their connect phase. + +### RC4 — Two features are 100% Tor by construction (allowlist 404) + +The peer listener path allowlist (`server.rs:1219-1239`) omits `/blob/` (mesh +file sharing, `typed_messages.rs:813-822`) and `/dwn/health` (step 1 of DWN sync, +`dwn_sync.rs:186`) → deterministic 404 over FIPS (`dial.rs:44-46` treats 404 as +fall-back) → deterministic Tor, after paying the full FIPS cost. Both endpoints are +already cryptographically gated, so they meet the allowlist's stated criterion. + +**Fix:** add `|| path.starts_with("/blob/") || path.starts_with("/dwn/")`; extend the +existing test block at `server.rs:1935-1945` (assert `/blob/abc` + `/dwn/health` +allowed, `/blobber` + `/dwnx` denied). + +### RC5 — Inbound listener can't heal; anchor flap = 5-minute Tor window; probe overhead + +- `peer_late_bind_loop` returns after first successful bind (`server.rs:1203`) and + `accept_loop` `continue`s on errors forever (`server.rs:1249-1258`): a fips0 + teardown/re-key leaves the node inbound-dead until process restart → **every peer** + falls back to Tor against it. +- Nothing reacts to anchor-link drops: anchors re-apply only on the 300s tick + (`server.rs:731`); worst-case 5min Tor-only after a flap (the historic "link dead + timeout 30s" flapping made this chronic). +- `is_service_active()` spawns up to 2 `systemctl` per FIPS attempt *and* per peer + per 25s warm tick (`dial.rs:284-294`); `warm_path` skips peers without + `fips_npub` in federation storage (`fips/mod.rs:88-95`); `anchors::apply` is + serial with unbounded subprocess waits (`anchors.rs:234-283`). + +**Fix:** rebindable listener; a ~25s connectivity watcher (reuse +`service::peer_connectivity_summary`, `fips/service.rs:178-207`) that re-applies +anchors immediately on a connected→disconnected edge with bounded backoff; 10s TTL +cache for `is_service_active` (mirror `transport/fips.rs:24-107`); warm the union of +federation+registry peers; make `apply()` concurrent with per-connect timeouts. + +### RC6 — Zero observability: fallbacks are invisible, so "uptime" is unfalsifiable + +Fallbacks log at `debug!` only (`dial.rs:458,491`); no counters; `last_transport` is +written by only 7 of ~20 call sites and **never read** to influence anything +(`storage.rs:120-147`). The parallel `TransportRouter` system can't even see FIPS +(`FipsTransport` is never constructed — `server.rs:422-442` registers Tor/Mesh/LAN +only). + +**Fix:** per-reason fallback counters (F1 no-npub / F2 service-inactive / F3 +DNS-fail / F4 connect-fail / F5 404 / F6 5xx) surfaced in `fips.status` + `info!` +logs with a `reason` field; call `record_peer_transport` from all peer-dial sites; +UI: per-peer transport badge on Cloud (the response already carries `transport` — +`content.rs:392-400` — Cloud.vue currently throws it away at `:716-721`). + +--- + +## Part A — execution phases + +### Phase A0 — fleet triage (no release needed; do first, validates everything) +1. Fleet audit: `fipsctl --version` + `nft list table inet fips` + `ss -tlnp | grep 5679` + on every node (roster: `reference_test_deploy_roster`). +2. Transient `nft insert rule inet fips inbound iifname fips0 tcp dport 5679 accept` + on hardened nodes (already done on .116 + .198, 2026-07-27) — instant fleet-wide + FIPS recovery while the code fix rides the OTA. +3. Upgrade .228 (and any other 0.3.x) fips daemon to v0.4.1. +4. Regenerate/clean stale `seed-anchors.json` on .116 (dead 192.168.1.x + :8668 entries). +5. Baseline measurement: for each node pair, `content.browse-peer` time + transport. + +### Phase A1 — P0 code (one commit, mechanical, offline-testable) +1. **nft drop-in: open 5679** — `fips/config.rs` (share the constant with + `dial.rs::PEER_PORT`). ← RC0 +2. **Allowlist `/blob/`, `/dwn/`** — `server.rs:1219-1239` + tests. ← RC4 +3. **`FIPS_UDP_PORT` = `PUBLISHED_UDP_PORT` (2121)** — `anchors.rs:293` + drift-guard + test against `render_config_yaml()`. ← RC2-G2 +4. **Un-deaden `lan_fips_anchors`** — hydrate `fips_npub` from federation storage in + `server.rs:761-766`; then mDNS TXT `fips` key + `set_fips_npub` + (`transport/lan.rs:50-54`, `lan.rs:96-108`, `LanTransport::new` 4th arg via + `crate::identity::fips_npub(&data_dir.join("identity"))`). ← RC2-G1 +5. **Retry-budget wrap + `fips_timeout` on 12 call sites** (list in RC3). ← RC3 + Verify: `cd core && cargo test -p archipelago` — watch `test_rendered_yaml_exact_snapshot` + (`config.rs:419`) + `test_render_is_deterministic` (`config.rs:476`); item 3 must + not change rendered output. + +### Phase A2 — telemetry BEFORE tuning (second commit) +6. Fallback counters by reason + `fips.status` exposure + `info!` reason logs; + `record_peer_transport` from all sites. ← RC6 (gives the baseline that makes A3 + measurable and "100%" falsifiable) + +### Phase A3 — resilience (third commit, measured against A2 baseline) +7. `is_service_active` 10s TTL cache; warm-path union + `warm_path_unchecked`. +8. Link-state watcher → immediate anchor re-apply on drop (replaces waiting for the + 300s tick); concurrent `apply()` with subprocess timeouts. +9. Rebindable peer listener (`server.rs:1203`, `1249-1258`). +10. (Reviewed, separate PR) endpoint-fallback for direct peering: LAN → Tailscale → + last-known-good, npub-keyed. Mesh-routing area — needs careful review per memory. + +### Phase A4 — verification gate (on nodes, before tag) +- On .116/.198/framework-pt/.228: `content.browse-peer` to every peer must return + `transport: "fips"` with sub-second latency (LAN pairs) / <3s (WAN), 20/20 calls. +- Kill the fips daemon on one node → calls fall back to Tor gracefully within the + fast-fail budget (<8s), UI shows partial results, no errors. +- Restart daemon → FIPS recovers within one watcher tick (~25s), verified in + `fips.status` counters. +- Flap the anchor link (drop vps2 route) → direct LAN pairs keep FIPS via their + direct link (G1 fix proof). +- Add these as `tests/multinode/` cases per `docs/multinode-testing-plan.md`; also + fix the known `node_rpc()` missing `--max-time` (tracker item). + +--- + +## Part B — optimistic loading + state management (frontend) + +Full audit: Pinia exists but pages fetch-on-mount with `loading=true` spinners; +`Dashboard.vue:89` keys the router-view by `route.path`, so **every navigation +unmounts and refetches everything**; no KeepAlive/onActivated anywhere; no dedup, +no abort, no SWR layer. Four hand-rolled cache implementations already exist and +prove the pattern (`useFleetData.ts:198-231` sessionStorage hydrate; +`homeStatus.ts` sticky-ready loadState; `Home.vue:591-621` wallet localStorage +snapshot; `curatedApps.ts:21-77` TTL cache). `SkeletonCard.vue` exists, imported by +zero files. + +### B1 — one shared primitive: `useCachedResource` composable + `resources` Pinia store +Semantics (generalize `homeStatus.ts` + `useFleetData.ts`): +- Keyed resource: `{ data, loadState: idle|loading|ready|error|refreshing, fetchedAt, error }`. +- **Hydrate synchronously** from memory (Pinia, survives navigation) → sessionStorage + snapshot (survives reload) → then revalidate in background. +- Sticky-ready: once `ready`, never regress to `loading` + (`loadState = loadState==='ready' ? 'ready' : 'loading'` — the `homeStatus.ts:80` idiom); + keep-last-known-value on error with a stale badge (age from `fetchedAt`). +- TTL per resource; `revalidateOnFocus` + on WS push (debounced, the + `Home.vue:539-542` pattern); explicit `invalidate(key)` for mutations. +- Optimistic mutation helper: apply → RPC → rollback on error (generalize + `TransportPrefsCard.vue:112-127`). + +### B2 — rpc-client upgrades (`src/api/rpc-client.ts`) +- `AbortSignal` in `RPCOptions` (today the AbortController at `:87` is timeout-only) + → abort-on-unmount for fan-outs. +- In-flight dedup keyed `method+JSON(params)` — collapses duplicate concurrent calls. +- Per-call `maxRetries` override; set `maxRetries: 1` for `content.browse-peer` / + `preview-peer` (retry×3 on a 30s timeout is why one slow peer = 90s spinner). + +### B3 — Cloud page conversion (worst offender, the marquee win) +- Move `sectionCounts`, `peerNodes`, `myFiles`, `peerFiles`, `paidItems` out of + `Cloud.vue` component state (`:403,:476,:582,:689,:427`) into the cached store — + instant render on revisit, background refresh. +- **Incremental per-peer fan-in**: render each peer's card as its + `content.browse-peer` resolves (today `Promise.allSettled` at `:708-747` blocks on + the slowest peer). Per-peer states: cached/fresh/loading/unreachable. +- **Surface `transport` per peer** (already in the response, discarded at `:716-721`): + FIPS/Tor badge + latency — this is also the fleet-wide FIPS-uptime dashboard the + user asked for, for free. +- Skeleton cards (revive `SkeletonCard.vue`, copy `FileGrid.vue:3-19` shimmer) instead + of spinners for counts/folders/peer grids. +- Stop `CloudFolder.vue:307-319` calling `cloudStore.reset()` on every folder entry — + cache per-path listings, navigate renders cache + revalidates. +- `PeerFiles.vue`: persist catalog + preview cache in the store; cap the + `preview-peer` fan-out (`:832-841`, currently unbounded) with a small concurrency + queue + abort-on-unmount. + +### B4 — roll out to remaining offenders (in audit order) +PeerFiles → Web5 wallet/ecash/LND slices → Monitoring → Lightning channels +(`LightningChannelsPanel.vue:650`) → Federation (already has `{showLoader:false}` — +just adopt the store) → Server → Credentials/OpenWrtGateway/ContainerApps. +`Apps.vue`/`Marketplace.vue`/`Fleet.vue` are already good; don't touch. + +### B5 — freshness via the existing push channel +`/ws/db` firehose + `sync.ts` JSON-patch already exist. Wire `useCachedResource` +revalidation to relevant WS pushes (debounced 800ms), keep the 30s staleness +reconciliation as backstop. No new backend needed for v1; a per-topic subscribe can +come later. + +### Part B verification (on nodes) +- Navigate Cloud → Apps → Cloud: peer files render instantly from cache (0 spinner), + refresh indicator while revalidating, updated data lands without layout jump. +- One unreachable peer: its card shows stale/unreachable state; other peers render + immediately (no 30s all-or-nothing). +- Kill backend mid-view: stale data stays visible with age badge; recovery + revalidates automatically. +- Hard reload: sessionStorage hydrate paints before first RPC completes. + +--- + +## Sequencing for the next release + +1. **A0 now** (fleet triage + transient nft rules + .228 daemon upgrade + baseline). +2. **A1 + A2** land together (P0 fixes + telemetry) → deploy to .116/.198 → + Phase A4 checks on the pair → framework-pt → full fleet. +3. **B1 + B2 + B3** (composable + rpc-client + Cloud) in parallel with A-testing — + frontend-only, verifiable against .116 dev (`reference_neode_ui_dev_testing`). +4. **A3** after telemetry baseline exists; **B4/B5** ride the same or next OTA. +5. Gate: Phase A4 checklist green + Part B verification on-device + existing + single-node gate stays green → tag/OTA per ship ritual. + +## Success criteria +- `content.browse-peer` transport = fips for ≥99% of calls between healthy 0.4.1 + nodes over 24h (measured by the new counters), Tor reserved for genuinely + FIPS-unreachable peers (.116-WiFi-class networks). +- Cloud revisit paints in <100ms from cache; fresh data within one revalidate. +- Fallback counters visible in `fips.status` so regressions are caught on the + dashboard, not by users.