Files
archy/docs/incident-2026-09-01.md
T
archipelago 0d0e2e243a
Demo images / Build & push demo images (push) Successful in 3m49s
feat(lnd): channel-peer watchdog — a dropped peer link heals itself
LND normally reconnects channel peers after a restart, but not reliably:
after long or repeated downtime (an app update, a node reboot,
reconciler churn) the peer link can stay down for hours while BOTH
endpoints keep the channel flagged disabled in the routing graph. The
node looks perfectly healthy, the wallet shows balance, and every
payment in either direction fails "no route to the recipient" —
observed live on framework-pt (2026-09-01): its only channel sat
disabled on both policy sides for ~17 hours after the LND 0.21.2
update, while shorty had 583k spendable and the user was told, by a
mis-mapped modal, that they had 'no payment channel'.

The channel graph is desired state — every open channel should have a
live peer connection. A daemon-side watchdog now enforces it:

- every 2 minutes, list channels + peers over LND REST
- for each channel whose remote peer is not connected, look the peer's
  advertised addresses up in the public graph and dial one
- per-peer retries throttled to 10 minutes so an unreachable peer is
  not hammered; 'already connected' counts as done; a peer with no
  advertised address is logged once per pass (cannot be dialed)
- no-ops quietly on nodes without LND (missing macaroon) and while a
  wallet is locked (503 body has no channels)

Unit tests pin the selection against the live REST shapes
(remote_pubkey in /v1/channels vs pub_key in /v1/peers).

v1.8.10 CHANGELOG + What's New entries staged so the next release run
is clean first time.
2026-09-01 17:51:15 -04:00

109 lines
9.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Incident + follow-up tracker — 2026-09-01 (post-HTTPS-work, post-LND-0.21.2 breakage)
Live incident spanning framework-pt and shorty-s after the HTTPS/launcher
work and the LND 0.18.4→0.21.2 pin bump. Root causes found on real nodes;
status updated as work lands. Each fix ships with a regression test so the
same class cannot silently return.
## A. Root causes (all verified live)
| # | Symptom | Root cause |
|---|---------|-----------|
| A1 | LND sends fail "Payment failed: Not Found" | LND 0.21 **removed** the deprecated `/v1/channels/transactions` REST route; backend still called it. Receive was fine; the "Failed to fetch" on framework-pt was A3 masking it. |
| A2 | Shorty NPM restart-loops (counter 3176) | Manifest conversion (fc68c5b6) dropped (a) the `/etc/letsencrypt` mount NPM's s6 boot demands, and (b) `NET_BIND_SERVICE` — its internal nginx binds 80/443/81 and the orchestrator runs `--cap-drop=ALL`. |
| A3 | framework-pt: every `/rpc/v1` fetch CORS-blocked, "Failed to fetch", dashboard "not responding", mempool/indeehub frames broken | nginx sent `Strict-Transport-Security: max-age=31536000; includeSubDomains` on **HTTPS**; browsers cached it, then silently upgraded the still-open **http** dashboard's fetches/frames to https → scheme change = cross-origin → CORS block. HTTP is a supported mode on purpose (self-signed cert, /ca.crt flow). |
| A4 | Mempool/IndeeHub/bitcoin-UI frames stay `http://` on HTTPS pages (mixed content, "does not connect") | `portAuth()` looked the launch port up under the launch alias (`mempool-web`, `lnd`, `bitcoin-knots`…); the signed catalog declares those ports under the manifest id that owns them (`archy-mempool-web`, `lnd-ui`, `bitcoin-ui`) → miss → launcher fell back to http. Cache also only warmed in Store/Discover views. |
| A5 | IndeeHub nostr sign-in dead over HTTPS | NIP-07 bridge compared `event.origin` for strict equality with the stored (http) app URL and replied to the **stored** URL as postMessage targetOrigin — both break when the frame was scheme-upgraded. |
| A6 | Portainer "disappeared" after restart/update, then demands a setup token "see server logs" | Update to 2.45.0 recreated the container; on a fresh DB Portainer ≥2.21 mints a one-time setup token printed ONLY in container logs — hostile appliance UX. The "disappearance" was the recreate + this unknown-token first screen. |
## B. Fixes (code)
| Fix | Files | Status |
|-----|-------|--------|
| B1 LND pay via `Router.SendPaymentV2` (`/v2/router/send`), pending-status + actionable failure reasons preserved | `core/archipelago/src/api/rpc/lnd/payments.rs` (+ unit tests) | ✅ code |
| B2 Portainer setup token surfaced in the existing credentials interstitial (`package.credentials` → AppSidebar card with copy) | `core/archipelago/src/api/rpc/package/install.rs` (+ unit tests) | ✅ code |
| B3 HSTS: none on :80, `max-age=0` on :443 (actively clears cached policy) | `image-recipe/configs/nginx-archipelago.conf` | ✅ code |
| B4 NPM manifest: `/etc/letsencrypt` mount + `NET_BIND_SERVICE` | `apps/nginx-proxy-manager/manifest.yml` | ✅ code |
| B5 `portAuth` alias resolution + unanimous port-wide fallback | `neode-ui/src/views/discover/curatedApps.ts` | ✅ code |
| B6 Catalog cache warmed at dashboard bootstrap | `neode-ui/src/App.vue` | ✅ |
| B7 NIP-07 bridge: host/port equality + reply to `event.origin` | `neode-ui/src/stores/appLauncher.ts` ✅ · `neode-ui/src/views/appSession/useNostrBridge.ts` ✅ | ✅ |
| B8 Stale LND 0.18.4 refs in test expectations | `tests/lifecycle/remote-lifecycle.sh` | ✅ |
## C. Regression tests ("never again")
| Test | Guards | Status |
|------|-------|--------|
| C1 Rust: router v2 response shape, nested errors, failure reasons | B1 | ✅ |
| C2 Rust: setup-token log extraction (live-captured 2.45.0 line shape) | B2 | ✅ |
| C3 bats: `lnd-api-compat` — POST `/v2/router/send` on the running LND must answer (never 404) | B1 vs image skew at gate time | ✅ (route probe verified live on shorty: HTTP 500 ≠ 404) |
| C4 bats: nginx must NOT send HSTS on :80; :443 must send `max-age=0` | B3 | ✅ |
| C5 neode-ui unit: portAuth alias + unanimous-scan (incl. bitcoin-knots→8334 https) | B5/B4-mixed-content | ✅ (6 tests) |
| C6 neode-ui unit: bridge origin equality ignores scheme | B7 | ✅ (2 tests) |
Backend suites: 34 targeted Rust tests green (payments v2 shape, setup-token
extraction, lnd wallet/info regressions); middleware/dispatcher suite green;
full neode-ui suite green (62 tests in the touched areas); production bundle
built and verified to embed the alias fix. `cargo fmt` applied.
## D. Deploy & live verification
| Step | Status |
|------|--------|
| D1 shorty NPM crash-loop stopped cleanly (user-stopped marker; public hosts keep serving via host nginx mirror) | ✅ 12:52Z |
| D2 shorty live nginx HSTS patch + reload | ✅ verified: :80 and :443 both answer `max-age=0` |
| D3 Regenerate catalog (releases/app-catalog.json + store copies) | ✅ semantic diff = exactly the two NPM fixes |
| D4 **User runs `scripts/sign-catalog.sh`** (signer built at /tmp/archy-sign-bin) | ✅ catalog signed + committed + pushed |
| D5 Commit + push (origin + gitea-vps2 OTA mirror) | ✅ 9 commits pushed |
| D6 Release v1.8.9-alpha: `scripts/create-release.sh 1.8.9-alpha` (mnemonic) → `scripts/publish-release-assets.sh 1.8.9-alpha gitea-vps2` | ✅ PUBLISHED (tag v1.8.9-alpha, releases/manifest.json live, backend+frontend assets verified by the script) |
| D7 OTA on shorty-s + framework-pt (Update button; shorty is on 1.8.8-alpha, daily check — hit Update now) | ⬜ user action |
| D8 shorty: clear the NPM user-stopped marker + Start (or it starts via the fixed catalog) | ✅ NPM LIVE-HEALED via the signed catalog: unit regenerated with both fixes, container up, admin UI HTTP 200 on :8081 (verified 15:42Z) |
| D9 framework-pt: Start Mempool — its containers are confirmed stopped (port 4080 refuses; gate answers on 7778/8334/50002/18083 so those apps will embed over https immediately) | ⬜ |
| D10 Post-deploy live checks: LND send+receive; mempool/IndeeHub/bitcoin-UI frames over https; NPM healthy + admin :8081 ✅; portainer token card on fresh DB; zero CORS errors | ⬜ after nodes update |
## E. Follow-ups discovered during the incident (ride the NEXT release, v1.8.10+)
- **LND channel-peer watchdog** (this release's headline platform fix): every
2 minutes the daemon reconnects peers of open channels that LND has not
re-established on its own (per-peer retry throttled to 10 minutes), using
the peer's advertised addresses from the public graph. Kills the whole
class this incident exposed — a channel unroutable ~17h after an LND update
while both nodes looked healthy. Unit tests pin the selection logic over the
live REST shapes.
- **Funding-modal honesty fix** (1464b1b2): the
Lightning "no channel" modal now states the node's real state — pending
channel confirming / balance on the far side / payment couldn't route /
genuinely no channels. Note the stale-direction defect it fixes: the
payment-failure mapper never set the direction, so a SEND failure showed
the RECEIVE-branch copy ("Receiving needs inbound liquidity…") — the exact
modal users saw while their node had a healthy 583k-outbound channel.
Both fixes have their v1.8.10 CHANGELOG + What's New entries staged so the
next `create-release.sh 1.8.10-alpha` runs clean first time.
- Nodes poll for OTA updates on `daily_check` — after publishing, tell the
user to hit Update rather than wait for the next check.
- `origin` remote had a stale pushurl with a dead token (pushes failed);
fixed to the canonical repo URL, stale `~/.git-credentials` entry with an
encoded port removed.
## F. Post-v1.8.9 verification on shorty-s (2026-09-01 evening)
- v1.8.9 applied; payment pipeline confirmed live: a 400,000 sat payment
SUCCEEDED through the v2 router route; the 404s are gone.
- App gate serves TLS on 4080/8334/18083/50002 (401 gate pages over https) —
https app frames now answer. Mempool over https requires a hard refresh
(PWA precaches the old bundle).
- **"No route to the recipient" on sends is real**: the invoices being tested
are from framework-pt, whose only channel (peer "Sandwich Farm",
0224c955…) is flagged `disabled` on BOTH policy sides in the routing graph
after today's node churn — the peer connection never re-established
(LND's reconnect backoff can stretch to hours). A disabled edge is
unroutable in both directions, so payments to/from framework-pt fail
regardless of shorty's 583k outbound. Fix: `lncli connect` the peer, wait
for the channel_update to re-enable the edge (~minutes), then re-test.
- The 577k attempt earlier failed for a different, correct reason: it exceeded
the channel's spendable balance (583,542 − 9,850 reserve ≈ 573k max).
framework-pt immediate workaround until its OTA lands: open the dashboard by
IP (`http://192.168.x.x`) instead of `framework-pt.local`, and/or clear the
cached policy once via `chrome://net-internals/#hsts` → Delete domain security
policies → `framework-pt.local`.