docs: FIPS near-100% uptime + optimistic UI state plan — live-proven root causes (nft 5679 drop, .228 daemon skew, dead LAN peering, no fast-fail) + phased execution
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
c5eeb4392f
commit
9f65f1e7ae
300
docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md
Normal file
300
docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md
Normal file
@ -0,0 +1,300 @@
|
||||
# FIPS near-100% uptime + optimistic UI state — implementation plan
|
||||
|
||||
**Date:** 2026-07-27. **Status:** researched + root-caused live on the fleet; ready to
|
||||
implement for the next release. Two workstreams: (A) make node↔node FIPS transport
|
||||
succeed whenever a FIPS path physically exists, (B) stop the UI reloading everything
|
||||
on every navigation (optimistic/cached cards, stale-while-revalidate) while keeping
|
||||
data fresh.
|
||||
|
||||
**Honesty note on "100%":** if a node's network blackholes every anchor (the .116
|
||||
WiFi case, `docs/HANDOFF-2026-07-20-fips-peer-files.md:117-133`), Tor fallback is
|
||||
*correct*. The achievable target is: **FIPS wins whenever a FIPS path exists, and
|
||||
fallback frequency is measured in-product so regressions are visible.** Today several
|
||||
paths are 0% FIPS *by construction* regardless of network health — that's the bug.
|
||||
|
||||
---
|
||||
|
||||
## Part A — why Cloud/FIPS "commonly falls back to Tor": ranked root causes
|
||||
|
||||
All verified live on 2026-07-27 (.116 local, .198, .228, Framework PT, x250s) plus a
|
||||
full code audit of `core/archipelago/src/{fips,transport,federation,server.rs}`.
|
||||
|
||||
### RC0 — 🔥 The hardening firewall drops the peer-API port on every hardened node (PROVEN)
|
||||
|
||||
The fips0 default-deny baseline (`/etc/fips/fips.nft`) is opened by archipelago's
|
||||
drop-in `80-web-ui.nft` (`fips/config.rs:236-255`) for **80 + 8443 + app ports only**.
|
||||
The peer-API listener — which carries *all* federation sync, cloud browse/download,
|
||||
mesh envelopes, DWN, invoices — is **`PEER_PORT = 5679`** (`fips/dial.rs:35`).
|
||||
**5679 is not in the allowlist.** The drop-in's own comment claims "web UI + peer
|
||||
API" but the peer API port was never added.
|
||||
|
||||
Live proof (2026-07-27):
|
||||
- .116 nft chain: 5,965 dropped packets; .198: **28,670 dropped packets** — that's
|
||||
peers' FIPS dials dying at the firewall.
|
||||
- .198 → .116 `GET :5679/health`: **timeout (6s)** before; **HTTP 200 in 0.35s**
|
||||
after `nft insert rule inet fips inbound iifname fips0 tcp dport 5679 accept`.
|
||||
Same result in reverse direction (200 in 0.64s).
|
||||
- Explains the exact fleet split in `federation/nodes.json`: hardened-baseline nodes
|
||||
(Framework PT, .198, .228, x250-dev, x250-mad2) = `last_transport: tor`;
|
||||
non-hardened nodes (Austin Sapien, X250-Beta, X250-PA) answer :5679 (404 from the
|
||||
path allowlist = listener reachable) = `last_transport: fips`.
|
||||
- Every dial to a hardened peer pays the 8s FIPS connect timeout
|
||||
(`dial.rs:114`) ×2 (retry, `dial.rs:128-140`) → then Tor. That's the "Cloud takes
|
||||
forever / shows Tor" experience.
|
||||
|
||||
**Fix (one line + reload):** add `tcp dport 5679 accept` to the drop-in in
|
||||
`fips/config.rs` (use a constant shared with `dial.rs::PEER_PORT`, not a literal).
|
||||
The drop-in reinstalls on every daemon config install, so it heals fleet-wide on OTA.
|
||||
⚠️ Transient manual rules were inserted on .116 and .198 during diagnosis (2026-07-27)
|
||||
— they vanish on the next `nft -f /etc/fips/fips.nft` reload or reboot; the code fix
|
||||
makes them permanent.
|
||||
|
||||
### RC1 — .228 (Shorty's) runs fips 0.3.0-dev; the 0.4.1 fleet can't reach it
|
||||
|
||||
.228's daemon: `0.3.0-dev (rev 34e00b9f6e)`, both anchor links "connected", but its
|
||||
ULA is 100% unreachable from 0.4.1 nodes (ping loss 100%). FIPS wire format is not
|
||||
stable across revs (`docs/HANDOFF-2026-07-23-companion-apk-deploy.md:78`). Everything
|
||||
to/from .228 rides Tor no matter what else we fix.
|
||||
|
||||
**Fix:** fleet fips-version audit + upgrade to v0.4.1 everywhere (in-product updater
|
||||
exists: `fips/update.rs`; .deb path per `reference_vps2_fips_anchor`). Add a version
|
||||
check to `fips.status` and surface a "peer daemon outdated" warning.
|
||||
|
||||
### RC2 — Direct LAN/endpoint peering is dead code + wrong port + stale seed anchors
|
||||
|
||||
Without direct links, all peer traffic hairpins through the vps2 anchor spanning
|
||||
tree (observed: .116→.198 cold RTT 1.5–3.5s on the same LAN; also the wedged-anchor
|
||||
latency-rot incident, `HANDOFF-2026-07-23:141-160`).
|
||||
|
||||
- **G1 — `lan_fips_anchors()` has never run.** It needs `PeerRecord.fips_npub`, but
|
||||
`PeerRegistry::set_fips_npub` (`transport/mod.rs:302`) has **zero callers** — mDNS
|
||||
TXT records only carry `did`/`pubkey`/`version` (`transport/lan.rs:50-54`). So the
|
||||
"co-located peers form a direct link" feature (`anchors.rs:294-305`,
|
||||
`server.rs:761-766`) is a fleet-wide no-op.
|
||||
- **G2 — wrong UDP port.** `anchors.rs:293` dials `8668`, but the generated
|
||||
fips.yaml binds UDP **2121** (`fips/config.rs:187`, `fips/mod.rs:130`). Even if G1
|
||||
ran, it would dial a dead port. `.116`'s live `seed-anchors.json` still carries
|
||||
`.198@192.168.1.198:8668` — **stale IP (LAN renumbered to 192.168.63.x) AND dead
|
||||
port**; both manual entries are useless today.
|
||||
- No Tailscale/alternate endpoint fallback when LAN is unreachable (the .116↔.198
|
||||
fix of 2026-07-20 was hand-applied per-node config, never productized).
|
||||
|
||||
**Fix:** (a) `FIPS_UDP_PORT` → `crate::fips::PUBLISHED_UDP_PORT` + drift-guard test;
|
||||
(b) hydrate `fips_npub` into the registry from federation storage (did-keyed join) so
|
||||
`lan_fips_anchors` goes live with no wire change; (c) advertise the npub in the mDNS
|
||||
TXT + `set_fips_npub` on resolve as the proper fix; (d) teach the LAN-anchor tick to
|
||||
also try a peer's Tailscale/last-known-good endpoint when LAN fails (reviewed change
|
||||
— this area got handoffs wrong twice, per memory).
|
||||
|
||||
### RC3 — No fast-fail on the hottest call sites; retry silently doubles every budget
|
||||
|
||||
- `content.browse-peer` — **the Cloud page** — has NO `fips_timeout`
|
||||
(`api/rpc/content.rs:363-366`): a cold FIPS path burns up to ~16.6s (8s connect +
|
||||
600ms + 8s retry) before Tor even starts, against a UI deadline of 30s
|
||||
(`Cloud.vue:720`) — and the frontend then retries ×3. Users see errors, not
|
||||
fallback. 12 call sites total lack `fips_timeout` (browse/download/preview-peer,
|
||||
`/blob`, DWN, node_message, rotation notifies).
|
||||
- `dial.rs:128-140` runs 2 full-budget attempts, so `fips_timeout(6s)` really means
|
||||
~12.6s everywhere.
|
||||
|
||||
**Fix:** wrap `send_with_retry` in a single `tokio::time::timeout(fips_attempt_timeout())`
|
||||
(call sites `dial.rs:455`, `dial.rs:488`; halve per-attempt client timeout), then add
|
||||
`.fips_timeout(...)`: `content.rs:366` (6s), `content.rs:281` (8s), `content.rs:1139`
|
||||
(6s), `typed_messages.rs:822` (8s), `dwn_sync.rs:188/213/272` (6s),
|
||||
`node_message.rs:376` (8s), `node_message.rs:412` (4s), `tor/mod.rs:501` (6s),
|
||||
`federation/handlers.rs:869` (6s). **Skip the three 900s streaming downloads**
|
||||
(`content.rs:552/870/1061`, `proxy.rs:236`) — `dial.rs:311-319` documents why; the
|
||||
retry-budget wrap covers their connect phase.
|
||||
|
||||
### RC4 — Two features are 100% Tor by construction (allowlist 404)
|
||||
|
||||
The peer listener path allowlist (`server.rs:1219-1239`) omits `/blob/<cid>` (mesh
|
||||
file sharing, `typed_messages.rs:813-822`) and `/dwn/health` (step 1 of DWN sync,
|
||||
`dwn_sync.rs:186`) → deterministic 404 over FIPS (`dial.rs:44-46` treats 404 as
|
||||
fall-back) → deterministic Tor, after paying the full FIPS cost. Both endpoints are
|
||||
already cryptographically gated, so they meet the allowlist's stated criterion.
|
||||
|
||||
**Fix:** add `|| path.starts_with("/blob/") || path.starts_with("/dwn/")`; extend the
|
||||
existing test block at `server.rs:1935-1945` (assert `/blob/abc` + `/dwn/health`
|
||||
allowed, `/blobber` + `/dwnx` denied).
|
||||
|
||||
### RC5 — Inbound listener can't heal; anchor flap = 5-minute Tor window; probe overhead
|
||||
|
||||
- `peer_late_bind_loop` returns after first successful bind (`server.rs:1203`) and
|
||||
`accept_loop` `continue`s on errors forever (`server.rs:1249-1258`): a fips0
|
||||
teardown/re-key leaves the node inbound-dead until process restart → **every peer**
|
||||
falls back to Tor against it.
|
||||
- Nothing reacts to anchor-link drops: anchors re-apply only on the 300s tick
|
||||
(`server.rs:731`); worst-case 5min Tor-only after a flap (the historic "link dead
|
||||
timeout 30s" flapping made this chronic).
|
||||
- `is_service_active()` spawns up to 2 `systemctl` per FIPS attempt *and* per peer
|
||||
per 25s warm tick (`dial.rs:284-294`); `warm_path` skips peers without
|
||||
`fips_npub` in federation storage (`fips/mod.rs:88-95`); `anchors::apply` is
|
||||
serial with unbounded subprocess waits (`anchors.rs:234-283`).
|
||||
|
||||
**Fix:** rebindable listener; a ~25s connectivity watcher (reuse
|
||||
`service::peer_connectivity_summary`, `fips/service.rs:178-207`) that re-applies
|
||||
anchors immediately on a connected→disconnected edge with bounded backoff; 10s TTL
|
||||
cache for `is_service_active` (mirror `transport/fips.rs:24-107`); warm the union of
|
||||
federation+registry peers; make `apply()` concurrent with per-connect timeouts.
|
||||
|
||||
### RC6 — Zero observability: fallbacks are invisible, so "uptime" is unfalsifiable
|
||||
|
||||
Fallbacks log at `debug!` only (`dial.rs:458,491`); no counters; `last_transport` is
|
||||
written by only 7 of ~20 call sites and **never read** to influence anything
|
||||
(`storage.rs:120-147`). The parallel `TransportRouter` system can't even see FIPS
|
||||
(`FipsTransport` is never constructed — `server.rs:422-442` registers Tor/Mesh/LAN
|
||||
only).
|
||||
|
||||
**Fix:** per-reason fallback counters (F1 no-npub / F2 service-inactive / F3
|
||||
DNS-fail / F4 connect-fail / F5 404 / F6 5xx) surfaced in `fips.status` + `info!`
|
||||
logs with a `reason` field; call `record_peer_transport` from all peer-dial sites;
|
||||
UI: per-peer transport badge on Cloud (the response already carries `transport` —
|
||||
`content.rs:392-400` — Cloud.vue currently throws it away at `:716-721`).
|
||||
|
||||
---
|
||||
|
||||
## Part A — execution phases
|
||||
|
||||
### Phase A0 — fleet triage (no release needed; do first, validates everything)
|
||||
1. Fleet audit: `fipsctl --version` + `nft list table inet fips` + `ss -tlnp | grep 5679`
|
||||
on every node (roster: `reference_test_deploy_roster`).
|
||||
2. Transient `nft insert rule inet fips inbound iifname fips0 tcp dport 5679 accept`
|
||||
on hardened nodes (already done on .116 + .198, 2026-07-27) — instant fleet-wide
|
||||
FIPS recovery while the code fix rides the OTA.
|
||||
3. Upgrade .228 (and any other 0.3.x) fips daemon to v0.4.1.
|
||||
4. Regenerate/clean stale `seed-anchors.json` on .116 (dead 192.168.1.x + :8668 entries).
|
||||
5. Baseline measurement: for each node pair, `content.browse-peer` time + transport.
|
||||
|
||||
### Phase A1 — P0 code (one commit, mechanical, offline-testable)
|
||||
1. **nft drop-in: open 5679** — `fips/config.rs` (share the constant with
|
||||
`dial.rs::PEER_PORT`). ← RC0
|
||||
2. **Allowlist `/blob/`, `/dwn/`** — `server.rs:1219-1239` + tests. ← RC4
|
||||
3. **`FIPS_UDP_PORT` = `PUBLISHED_UDP_PORT` (2121)** — `anchors.rs:293` + drift-guard
|
||||
test against `render_config_yaml()`. ← RC2-G2
|
||||
4. **Un-deaden `lan_fips_anchors`** — hydrate `fips_npub` from federation storage in
|
||||
`server.rs:761-766`; then mDNS TXT `fips` key + `set_fips_npub`
|
||||
(`transport/lan.rs:50-54`, `lan.rs:96-108`, `LanTransport::new` 4th arg via
|
||||
`crate::identity::fips_npub(&data_dir.join("identity"))`). ← RC2-G1
|
||||
5. **Retry-budget wrap + `fips_timeout` on 12 call sites** (list in RC3). ← RC3
|
||||
Verify: `cd core && cargo test -p archipelago` — watch `test_rendered_yaml_exact_snapshot`
|
||||
(`config.rs:419`) + `test_render_is_deterministic` (`config.rs:476`); item 3 must
|
||||
not change rendered output.
|
||||
|
||||
### Phase A2 — telemetry BEFORE tuning (second commit)
|
||||
6. Fallback counters by reason + `fips.status` exposure + `info!` reason logs;
|
||||
`record_peer_transport` from all sites. ← RC6 (gives the baseline that makes A3
|
||||
measurable and "100%" falsifiable)
|
||||
|
||||
### Phase A3 — resilience (third commit, measured against A2 baseline)
|
||||
7. `is_service_active` 10s TTL cache; warm-path union + `warm_path_unchecked`.
|
||||
8. Link-state watcher → immediate anchor re-apply on drop (replaces waiting for the
|
||||
300s tick); concurrent `apply()` with subprocess timeouts.
|
||||
9. Rebindable peer listener (`server.rs:1203`, `1249-1258`).
|
||||
10. (Reviewed, separate PR) endpoint-fallback for direct peering: LAN → Tailscale →
|
||||
last-known-good, npub-keyed. Mesh-routing area — needs careful review per memory.
|
||||
|
||||
### Phase A4 — verification gate (on nodes, before tag)
|
||||
- On .116/.198/framework-pt/.228: `content.browse-peer` to every peer must return
|
||||
`transport: "fips"` with sub-second latency (LAN pairs) / <3s (WAN), 20/20 calls.
|
||||
- Kill the fips daemon on one node → calls fall back to Tor gracefully within the
|
||||
fast-fail budget (<8s), UI shows partial results, no errors.
|
||||
- Restart daemon → FIPS recovers within one watcher tick (~25s), verified in
|
||||
`fips.status` counters.
|
||||
- Flap the anchor link (drop vps2 route) → direct LAN pairs keep FIPS via their
|
||||
direct link (G1 fix proof).
|
||||
- Add these as `tests/multinode/` cases per `docs/multinode-testing-plan.md`; also
|
||||
fix the known `node_rpc()` missing `--max-time` (tracker item).
|
||||
|
||||
---
|
||||
|
||||
## Part B — optimistic loading + state management (frontend)
|
||||
|
||||
Full audit: Pinia exists but pages fetch-on-mount with `loading=true` spinners;
|
||||
`Dashboard.vue:89` keys the router-view by `route.path`, so **every navigation
|
||||
unmounts and refetches everything**; no KeepAlive/onActivated anywhere; no dedup,
|
||||
no abort, no SWR layer. Four hand-rolled cache implementations already exist and
|
||||
prove the pattern (`useFleetData.ts:198-231` sessionStorage hydrate;
|
||||
`homeStatus.ts` sticky-ready loadState; `Home.vue:591-621` wallet localStorage
|
||||
snapshot; `curatedApps.ts:21-77` TTL cache). `SkeletonCard.vue` exists, imported by
|
||||
zero files.
|
||||
|
||||
### B1 — one shared primitive: `useCachedResource` composable + `resources` Pinia store
|
||||
Semantics (generalize `homeStatus.ts` + `useFleetData.ts`):
|
||||
- Keyed resource: `{ data, loadState: idle|loading|ready|error|refreshing, fetchedAt, error }`.
|
||||
- **Hydrate synchronously** from memory (Pinia, survives navigation) → sessionStorage
|
||||
snapshot (survives reload) → then revalidate in background.
|
||||
- Sticky-ready: once `ready`, never regress to `loading`
|
||||
(`loadState = loadState==='ready' ? 'ready' : 'loading'` — the `homeStatus.ts:80` idiom);
|
||||
keep-last-known-value on error with a stale badge (age from `fetchedAt`).
|
||||
- TTL per resource; `revalidateOnFocus` + on WS push (debounced, the
|
||||
`Home.vue:539-542` pattern); explicit `invalidate(key)` for mutations.
|
||||
- Optimistic mutation helper: apply → RPC → rollback on error (generalize
|
||||
`TransportPrefsCard.vue:112-127`).
|
||||
|
||||
### B2 — rpc-client upgrades (`src/api/rpc-client.ts`)
|
||||
- `AbortSignal` in `RPCOptions` (today the AbortController at `:87` is timeout-only)
|
||||
→ abort-on-unmount for fan-outs.
|
||||
- In-flight dedup keyed `method+JSON(params)` — collapses duplicate concurrent calls.
|
||||
- Per-call `maxRetries` override; set `maxRetries: 1` for `content.browse-peer` /
|
||||
`preview-peer` (retry×3 on a 30s timeout is why one slow peer = 90s spinner).
|
||||
|
||||
### B3 — Cloud page conversion (worst offender, the marquee win)
|
||||
- Move `sectionCounts`, `peerNodes`, `myFiles`, `peerFiles`, `paidItems` out of
|
||||
`Cloud.vue` component state (`:403,:476,:582,:689,:427`) into the cached store —
|
||||
instant render on revisit, background refresh.
|
||||
- **Incremental per-peer fan-in**: render each peer's card as its
|
||||
`content.browse-peer` resolves (today `Promise.allSettled` at `:708-747` blocks on
|
||||
the slowest peer). Per-peer states: cached/fresh/loading/unreachable.
|
||||
- **Surface `transport` per peer** (already in the response, discarded at `:716-721`):
|
||||
FIPS/Tor badge + latency — this is also the fleet-wide FIPS-uptime dashboard the
|
||||
user asked for, for free.
|
||||
- Skeleton cards (revive `SkeletonCard.vue`, copy `FileGrid.vue:3-19` shimmer) instead
|
||||
of spinners for counts/folders/peer grids.
|
||||
- Stop `CloudFolder.vue:307-319` calling `cloudStore.reset()` on every folder entry —
|
||||
cache per-path listings, navigate renders cache + revalidates.
|
||||
- `PeerFiles.vue`: persist catalog + preview cache in the store; cap the
|
||||
`preview-peer` fan-out (`:832-841`, currently unbounded) with a small concurrency
|
||||
queue + abort-on-unmount.
|
||||
|
||||
### B4 — roll out to remaining offenders (in audit order)
|
||||
PeerFiles → Web5 wallet/ecash/LND slices → Monitoring → Lightning channels
|
||||
(`LightningChannelsPanel.vue:650`) → Federation (already has `{showLoader:false}` —
|
||||
just adopt the store) → Server → Credentials/OpenWrtGateway/ContainerApps.
|
||||
`Apps.vue`/`Marketplace.vue`/`Fleet.vue` are already good; don't touch.
|
||||
|
||||
### B5 — freshness via the existing push channel
|
||||
`/ws/db` firehose + `sync.ts` JSON-patch already exist. Wire `useCachedResource`
|
||||
revalidation to relevant WS pushes (debounced 800ms), keep the 30s staleness
|
||||
reconciliation as backstop. No new backend needed for v1; a per-topic subscribe can
|
||||
come later.
|
||||
|
||||
### Part B verification (on nodes)
|
||||
- Navigate Cloud → Apps → Cloud: peer files render instantly from cache (0 spinner),
|
||||
refresh indicator while revalidating, updated data lands without layout jump.
|
||||
- One unreachable peer: its card shows stale/unreachable state; other peers render
|
||||
immediately (no 30s all-or-nothing).
|
||||
- Kill backend mid-view: stale data stays visible with age badge; recovery
|
||||
revalidates automatically.
|
||||
- Hard reload: sessionStorage hydrate paints before first RPC completes.
|
||||
|
||||
---
|
||||
|
||||
## Sequencing for the next release
|
||||
|
||||
1. **A0 now** (fleet triage + transient nft rules + .228 daemon upgrade + baseline).
|
||||
2. **A1 + A2** land together (P0 fixes + telemetry) → deploy to .116/.198 →
|
||||
Phase A4 checks on the pair → framework-pt → full fleet.
|
||||
3. **B1 + B2 + B3** (composable + rpc-client + Cloud) in parallel with A-testing —
|
||||
frontend-only, verifiable against .116 dev (`reference_neode_ui_dev_testing`).
|
||||
4. **A3** after telemetry baseline exists; **B4/B5** ride the same or next OTA.
|
||||
5. Gate: Phase A4 checklist green + Part B verification on-device + existing
|
||||
single-node gate stays green → tag/OTA per ship ritual.
|
||||
|
||||
## Success criteria
|
||||
- `content.browse-peer` transport = fips for ≥99% of calls between healthy 0.4.1
|
||||
nodes over 24h (measured by the new counters), Tor reserved for genuinely
|
||||
FIPS-unreachable peers (.116-WiFi-class networks).
|
||||
- Cloud revisit paints in <100ms from cache; fresh data within one revalidate.
|
||||
- Fallback counters visible in `fips.status` so regressions are caught on the
|
||||
dashboard, not by users.
|
||||
Loading…
x
Reference in New Issue
Block a user