Files
archy/docs/incident-2026-09-01.md
T

109 lines
9.0 KiB
Markdown
Raw Normal View History

# Incident + follow-up tracker — 2026-09-01 (post-HTTPS-work, post-LND-0.21.2 breakage)
Live incident spanning framework-pt and shorty-s after the HTTPS/launcher
work and the LND 0.18.4→0.21.2 pin bump. Root causes found on real nodes;
status updated as work lands. Each fix ships with a regression test so the
same class cannot silently return.
## A. Root causes (all verified live)
| # | Symptom | Root cause |
|---|---------|-----------|
| A1 | LND sends fail "Payment failed: Not Found" | LND 0.21 **removed** the deprecated `/v1/channels/transactions` REST route; backend still called it. Receive was fine; the "Failed to fetch" on framework-pt was A3 masking it. |
| A2 | Shorty NPM restart-loops (counter 3176) | Manifest conversion (fc68c5b6) dropped (a) the `/etc/letsencrypt` mount NPM's s6 boot demands, and (b) `NET_BIND_SERVICE` — its internal nginx binds 80/443/81 and the orchestrator runs `--cap-drop=ALL`. |
| A3 | framework-pt: every `/rpc/v1` fetch CORS-blocked, "Failed to fetch", dashboard "not responding", mempool/indeehub frames broken | nginx sent `Strict-Transport-Security: max-age=31536000; includeSubDomains` on **HTTPS**; browsers cached it, then silently upgraded the still-open **http** dashboard's fetches/frames to https → scheme change = cross-origin → CORS block. HTTP is a supported mode on purpose (self-signed cert, /ca.crt flow). |
| A4 | Mempool/IndeeHub/bitcoin-UI frames stay `http://` on HTTPS pages (mixed content, "does not connect") | `portAuth()` looked the launch port up under the launch alias (`mempool-web`, `lnd`, `bitcoin-knots`…); the signed catalog declares those ports under the manifest id that owns them (`archy-mempool-web`, `lnd-ui`, `bitcoin-ui`) → miss → launcher fell back to http. Cache also only warmed in Store/Discover views. |
| A5 | IndeeHub nostr sign-in dead over HTTPS | NIP-07 bridge compared `event.origin` for strict equality with the stored (http) app URL and replied to the **stored** URL as postMessage targetOrigin — both break when the frame was scheme-upgraded. |
| A6 | Portainer "disappeared" after restart/update, then demands a setup token "see server logs" | Update to 2.45.0 recreated the container; on a fresh DB Portainer ≥2.21 mints a one-time setup token printed ONLY in container logs — hostile appliance UX. The "disappearance" was the recreate + this unknown-token first screen. |
## B. Fixes (code)
| Fix | Files | Status |
|-----|-------|--------|
| B1 LND pay via `Router.SendPaymentV2` (`/v2/router/send`), pending-status + actionable failure reasons preserved | `core/archipelago/src/api/rpc/lnd/payments.rs` (+ unit tests) | ✅ code |
| B2 Portainer setup token surfaced in the existing credentials interstitial (`package.credentials` → AppSidebar card with copy) | `core/archipelago/src/api/rpc/package/install.rs` (+ unit tests) | ✅ code |
| B3 HSTS: none on :80, `max-age=0` on :443 (actively clears cached policy) | `image-recipe/configs/nginx-archipelago.conf` | ✅ code |
| B4 NPM manifest: `/etc/letsencrypt` mount + `NET_BIND_SERVICE` | `apps/nginx-proxy-manager/manifest.yml` | ✅ code |
| B5 `portAuth` alias resolution + unanimous port-wide fallback | `neode-ui/src/views/discover/curatedApps.ts` | ✅ code |
| B6 Catalog cache warmed at dashboard bootstrap | `neode-ui/src/App.vue` | ✅ |
| B7 NIP-07 bridge: host/port equality + reply to `event.origin` | `neode-ui/src/stores/appLauncher.ts` ✅ · `neode-ui/src/views/appSession/useNostrBridge.ts` ✅ | ✅ |
| B8 Stale LND 0.18.4 refs in test expectations | `tests/lifecycle/remote-lifecycle.sh` | ✅ |
## C. Regression tests ("never again")
| Test | Guards | Status |
|------|-------|--------|
| C1 Rust: router v2 response shape, nested errors, failure reasons | B1 | ✅ |
| C2 Rust: setup-token log extraction (live-captured 2.45.0 line shape) | B2 | ✅ |
| C3 bats: `lnd-api-compat` — POST `/v2/router/send` on the running LND must answer (never 404) | B1 vs image skew at gate time | ✅ (route probe verified live on shorty: HTTP 500 ≠ 404) |
| C4 bats: nginx must NOT send HSTS on :80; :443 must send `max-age=0` | B3 | ✅ |
| C5 neode-ui unit: portAuth alias + unanimous-scan (incl. bitcoin-knots→8334 https) | B5/B4-mixed-content | ✅ (6 tests) |
| C6 neode-ui unit: bridge origin equality ignores scheme | B7 | ✅ (2 tests) |
Backend suites: 34 targeted Rust tests green (payments v2 shape, setup-token
extraction, lnd wallet/info regressions); middleware/dispatcher suite green;
full neode-ui suite green (62 tests in the touched areas); production bundle
built and verified to embed the alias fix. `cargo fmt` applied.
## D. Deploy & live verification
| Step | Status |
|------|--------|
| D1 shorty NPM crash-loop stopped cleanly (user-stopped marker; public hosts keep serving via host nginx mirror) | ✅ 12:52Z |
| D2 shorty live nginx HSTS patch + reload | ✅ verified: :80 and :443 both answer `max-age=0` |
| D3 Regenerate catalog (releases/app-catalog.json + store copies) | ✅ semantic diff = exactly the two NPM fixes |
| D4 **User runs `scripts/sign-catalog.sh`** (signer built at /tmp/archy-sign-bin) | ✅ catalog signed + committed + pushed |
| D5 Commit + push (origin + gitea-vps2 OTA mirror) | ✅ 9 commits pushed |
| D6 Release v1.8.9-alpha: `scripts/create-release.sh 1.8.9-alpha` (mnemonic) → `scripts/publish-release-assets.sh 1.8.9-alpha gitea-vps2` | ✅ PUBLISHED (tag v1.8.9-alpha, releases/manifest.json live, backend+frontend assets verified by the script) |
| D7 OTA on shorty-s + framework-pt (Update button; shorty is on 1.8.8-alpha, daily check — hit Update now) | ⬜ user action |
| D8 shorty: clear the NPM user-stopped marker + Start (or it starts via the fixed catalog) | ✅ NPM LIVE-HEALED via the signed catalog: unit regenerated with both fixes, container up, admin UI HTTP 200 on :8081 (verified 15:42Z) |
| D9 framework-pt: Start Mempool — its containers are confirmed stopped (port 4080 refuses; gate answers on 7778/8334/50002/18083 so those apps will embed over https immediately) | ⬜ |
| D10 Post-deploy live checks: LND send+receive; mempool/IndeeHub/bitcoin-UI frames over https; NPM healthy + admin :8081 ✅; portainer token card on fresh DB; zero CORS errors | ⬜ after nodes update |
## E. Follow-ups discovered during the incident (ride the NEXT release, v1.8.10+)
- **LND channel-peer watchdog** (this release's headline platform fix): every
2 minutes the daemon reconnects peers of open channels that LND has not
re-established on its own (per-peer retry throttled to 10 minutes), using
the peer's advertised addresses from the public graph. Kills the whole
class this incident exposed — a channel unroutable ~17h after an LND update
while both nodes looked healthy. Unit tests pin the selection logic over the
live REST shapes.
- **Funding-modal honesty fix** (1464b1b2): the
Lightning "no channel" modal now states the node's real state — pending
channel confirming / balance on the far side / payment couldn't route /
genuinely no channels. Note the stale-direction defect it fixes: the
payment-failure mapper never set the direction, so a SEND failure showed
the RECEIVE-branch copy ("Receiving needs inbound liquidity…") — the exact
modal users saw while their node had a healthy 583k-outbound channel.
Both fixes have their v1.8.10 CHANGELOG + What's New entries staged so the
next `create-release.sh 1.8.10-alpha` runs clean first time.
- Nodes poll for OTA updates on `daily_check` — after publishing, tell the
user to hit Update rather than wait for the next check.
- `origin` remote had a stale pushurl with a dead token (pushes failed);
fixed to the canonical repo URL, stale `~/.git-credentials` entry with an
encoded port removed.
## F. Post-v1.8.9 verification on shorty-s (2026-09-01 evening)
- v1.8.9 applied; payment pipeline confirmed live: a 400,000 sat payment
SUCCEEDED through the v2 router route; the 404s are gone.
- App gate serves TLS on 4080/8334/18083/50002 (401 gate pages over https) —
https app frames now answer. Mempool over https requires a hard refresh
(PWA precaches the old bundle).
- **"No route to the recipient" on sends is real**: the invoices being tested
are from framework-pt, whose only channel (peer "Sandwich Farm",
0224c955…) is flagged `disabled` on BOTH policy sides in the routing graph
after today's node churn — the peer connection never re-established
(LND's reconnect backoff can stretch to hours). A disabled edge is
unroutable in both directions, so payments to/from framework-pt fail
regardless of shorty's 583k outbound. Fix: `lncli connect` the peer, wait
for the channel_update to re-enable the edge (~minutes), then re-test.
- The 577k attempt earlier failed for a different, correct reason: it exceeded
the channel's spendable balance (583,542 − 9,850 reserve ≈ 573k max).
framework-pt immediate workaround until its OTA lands: open the dashboard by
IP (`http://192.168.x.x`) instead of `framework-pt.local`, and/or clear the
cached policy once via `chrome://net-internals/#hsts` → Delete domain security
policies → `framework-pt.local`.