feat(lnd): channel-peer watchdog — a dropped peer link heals itself
Demo images / Build & push demo images (push) Successful in 3m49s

LND normally reconnects channel peers after a restart, but not reliably:
after long or repeated downtime (an app update, a node reboot,
reconciler churn) the peer link can stay down for hours while BOTH
endpoints keep the channel flagged disabled in the routing graph. The
node looks perfectly healthy, the wallet shows balance, and every
payment in either direction fails "no route to the recipient" —
observed live on framework-pt (2026-09-01): its only channel sat
disabled on both policy sides for ~17 hours after the LND 0.21.2
update, while shorty had 583k spendable and the user was told, by a
mis-mapped modal, that they had 'no payment channel'.

The channel graph is desired state — every open channel should have a
live peer connection. A daemon-side watchdog now enforces it:

- every 2 minutes, list channels + peers over LND REST
- for each channel whose remote peer is not connected, look the peer's
  advertised addresses up in the public graph and dial one
- per-peer retries throttled to 10 minutes so an unreachable peer is
  not hammered; 'already connected' counts as done; a peer with no
  advertised address is logged once per pass (cannot be dialed)
- no-ops quietly on nodes without LND (missing macaroon) and while a
  wallet is locked (503 body has no channels)

Unit tests pin the selection against the live REST shapes
(remote_pubkey in /v1/channels vs pub_key in /v1/peers).

v1.8.10 CHANGELOG + What's New entries staged so the next release run
is clean first time.
This commit is contained in:
archipelago
2026-09-01 17:51:15 -04:00
parent 9c49b502e3
commit 0d0e2e243a
5 changed files with 276 additions and 4 deletions
+11 -4
View File
@@ -62,15 +62,22 @@ built and verified to embed the alias fix. `cargo fmt` applied.
## E. Follow-ups discovered during the incident (ride the NEXT release, v1.8.10+)
- **Funding-modal honesty fix landed after the v1.8.9 tag** (1464b1b2): the
- **LND channel-peer watchdog** (this release's headline platform fix): every
2 minutes the daemon reconnects peers of open channels that LND has not
re-established on its own (per-peer retry throttled to 10 minutes), using
the peer's advertised addresses from the public graph. Kills the whole
class this incident exposed — a channel unroutable ~17h after an LND update
while both nodes looked healthy. Unit tests pin the selection logic over the
live REST shapes.
- **Funding-modal honesty fix** (1464b1b2): the
Lightning "no channel" modal now states the node's real state — pending
channel confirming / balance on the far side / payment couldn't route /
genuinely no channels. Ships in the next release; needs its own
create-release run (one more mnemonic paste). The CHANGELOG entry for
v1.8.10 should carry it. Note the stale-direction defect it fixes: the
genuinely no channels. Note the stale-direction defect it fixes: the
payment-failure mapper never set the direction, so a SEND failure showed
the RECEIVE-branch copy ("Receiving needs inbound liquidity…") — the exact
modal users saw while their node had a healthy 583k-outbound channel.
Both fixes have their v1.8.10 CHANGELOG + What's New entries staged so the
next `create-release.sh 1.8.10-alpha` runs clean first time.
- Nodes poll for OTA updates on `daily_check` — after publishing, tell the
user to hit Update rather than wait for the next check.
- `origin` remote had a stale pushurl with a dead token (pushes failed);