Three real incidents from live-testing the flash feature on archy-x250-dev,
each traced through logs on the node and fixed at the root:
- A failed flash left the mesh listener auto-resuming into a reconnect loop
whose backoff reset to its 5s minimum on any momentary connection, even
mid-boot-loop — every retry's open() toggles DTR/RTS, which resets many
ESP32 boards, so the retries were themselves sustaining the loop. Backoff
now only resets after a session runs stably for 20s+, and the flash job
no longer auto-resumes the listener after a failure (only on success,
after a settle delay).
- The HTTP client's blanket 30s request timeout covered entire downloads
(Meshtastic's zip is ~170MB), killing large transfers mid-stream; fixed
with a per-chunk stall timeout instead of a fixed total-transfer cap. But
the download's initial request had no timeout at all, so a slow-to-start
server hung it forever — wedging the single-flash-job guard permanently
("already in progress" on every subsequent attempt). Fixed with a bounded
wait for the response to start, plus an absolute ceiling around the whole
job as a last-resort safety net.
- archy-rnodeconf's PyInstaller freeze broke its own internal esptool
invocation (sys.executable pointed at the frozen binary itself instead of
a real interpreter) — fixed via a new runtime hook. esptool's own
auto-reset-into-bootloader handshake is separately known to be flaky on
some CP2102/CH340 boards; it now retries once at a conservative baud
rate instead of failing outright, and failure logs show the full error
chain instead of just the outer context message.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
9.9 KiB
Combined test session — 2026-07-22 batch (one sitting)
Staged on framework-pt (100.65.115.109) AND archi thinkpad (this
machine's node) so everything can be tested in one pass. Items marked ✅ were
already verified by the agent on a node; ❑ items need a human.
What's in this batch: mesh message/DM persistence across restarts · first-message + DM announce fix · 15s announce poll · radio hot-swap modal (probe / keep-as-is / apply-settings) · whisper beam-1 (release-gated, see §D) · real-time wallet push (0-conf tx shows in seconds) · calm Lightning "still starting" notice · external tx-explorer fallback with consent modal + wallet-settings On-chain tab · apps open ABOVE modals with the launch animation · mempool installs no longer blocked by a resyncing ElectrumX · [pending: other agents' two push sets — section F fills in when their code lands].
H. Wallet & explorer (new — test on the thinkpad node, it's pruned)
- ❑ Real-time tx display: send a small on-chain amount to this node's wallet → the balance and the yellow "unconfirmed" transaction appear within a few seconds of broadcast, no refresh, no wallet action.
- ❑ External explorer consent: with no local Mempool app running, tap a
transaction → amber consent modal explains it opens on another node's
mempool (default tx1138.com, placeholder mempool.guide, editable) →
Open Explorer opens
<explorer>/tx/<hash>in a new tab. Tick "don't ask again" and confirm the next tap opens directly. - ❑ Wallet Settings → On-chain tab: explorer URL editable, warning shown, "don't warn" toggle; tabs now read Channels / Cashu / Fedi / Ark / On-chain and fit on one row (check mobile too).
- ❑ Modal → app animation: on a node WITH Mempool running, open Transactions and tap a tx → the Mempool app animates in ABOVE the modal (previously loaded invisibly underneath); closing it returns to the modal.
- ❑ Lightning "still starting": right after a node restart, try opening a channel → either it just works (silent retry) or a calm amber ⏳ notice appears — never the red "Failed to connect to peer" error.
A. Staged state (agent-verified before you start)
- ✅ Dev binary (persistence + announce seeder + hot-swap) on
/usr/local/bin/archipelago, service healthy, no crash-loop. - ✅ Frontend bundle with the new device modal at
/opt/archipelago/web-ui. - ✅ Seeder re-ran:
automations.yamlupgraded v1→v2 (first-message announce),configuration.yamlrest block atscan_interval: 15, HA restarted clean. - ✅
mesh-messages.jsonpersisting + restored across a service restart. - ✅
mesh.probe-devicereturns real firmware details for the plugged stick.
B. Mesh history survives restarts (the "messages go missing" fix)
- ❑ Open Mesh chat — your existing DM/channel history from today is visible.
- ❑ Send one channel message and one DM (either direction).
- ❑ Reboot the whole node (not just the service). After it's back: history still there, including the two new messages, correct timestamps/senders.
- ❑ Send a NEW message to another node right after the reboot and confirm the other side receives it (this exercises the send-seq fix — before it, the first post-reboot sends were silently dropped by peers as replays).
C. Speaker announcements
- ❑ Have another node send a public channel message → speaker announces sender + text within ~15s (was ~30s).
- ❑ Have another node send you a DM → speaker announces it the same way.
- ❑ Restart Home Assistant (or the node) → the last old message is NOT re-announced (no announce storm).
- ❑ (First-message case — the original bug — only reproducible on a node with an empty history: optional, covered by agent verification of the guard.)
D. Voice (regression + speed)
- ❑ "Hey Jarvis, what's the block height" and one fuzzy phrasing — same correct answers as before (no behavior change is the pass condition).
- ⓘ The ~45% faster speech-to-text (whisper beam-1) ships via the signed catalog in the release — it is NOT on the node during this test session. Benchmarked on this exact hardware: identical transcripts, 0.94s → 0.51s.
E. Radio hot-swap modal (your Reticulum stick is already plugged in)
- ❑ Open the web UI anywhere — within ~30s a "Mesh Radio Detected" modal
appears showing the stick on
/dev/ttyACM0, with a card of what's on it (firmware badge: Reticulum RNode / MeshCore / Meshtastic + current name/region/channels where the firmware exposes them). - ❑ Press Keep As Is → mesh connects using the radio exactly as flashed (check Mesh → Device tab: connected, firmware type correct; nothing on the radio changed).
- ❑ Unplug the stick, plug the old MeshCore one → the modal appears AGAIN (every plug re-triggers, same or different /dev path).
- ❑ This time press Set Up with Archipelago Settings → second screen
shows channel
archipelago, your region, and the node's RF params (the validated Portugal preset on this fleet) BEFORE anything is written; confirm → radio provisions and joins the mesh. - ❑ Swap sticks once more with no UI interaction except "Keep As Is" — chat still works end-to-end afterwards (hot-swap without ceremony).
F. Companion pairing + mobile onboarding (other agent — push set #1, MERGED)
- ❑ Companion app: pair with the node via the new named QR (device tokens) — pairing completes instantly, device appears in the paired-devices list.
- ❑ Remote access now rides the embedded FIPS mesh (WireGuard replaced): with the phone OFF the node's WiFi, the companion still reaches the node.
- ❑ The reworked mobile onboarding/intro overlay screens flow correctly on first launch of the new APK (in-tarball APK is the 27MB build).
- ❑ (Push set #2 from the other agents is still pending — the release waits for it; this staged build does NOT include it yet.)
G. Quick regressions
- ❑ Pine launcher page (:10380) still shows the live node card; "Connect Pine to WiFi" button loads without JS errors.
- ❑ Mobile Home: wallet card directly under My Apps (G5 from the voice epic).
- ❑ Mesh RF settings panel (Mesh → Device) still loads and saves.
H. LoRa radio firmware flashing (Heltec V3/V4, new — extends Section E)
Full v1 scope is 3 firmware families × 2 boards (6 cells); mark each cell tested on real hardware vs. code-reviewed only as this is run.
- ❑ From the hot-swap modal's step 1 (device already probed), press Flash Firmware… → new step shows firmware-family + board pickers and the erase-confirmation checkbox; "Erase & Flash Now" stays disabled until family, board, AND the checkbox are all set.
- ❑ Confirm what's currently on the test stick via the existing probe BEFORE flashing it — don't flash the only known-good device without a fallback board on hand.
- ❑ Prefer a spare Heltec V3/V4 for the first destructive erase+flash run; only exercise a primary/in-use stick once the flow is proven safe.
- ❑ MeshCore → Heltec V3: erase + write completes, progress bar and log tail update live, ends at "Flash complete".
- ❑ Meshtastic → Heltec V3: same, using the extracted
*.factory.binfrom the esp32s3 release zip. - ❑ Reticulum/RNode → Heltec V3:
archy-rnodeconf --autoinstallpath completes (no raw esptool erase/write step for this family — seemesh/flash.rsdoc comment). - ❑ Repeat 4-6 against a Heltec V4. Confirmed 2026-07-23 on real hardware: V4 uses the ESP32-S3's native-USB JTAG/serial peripheral (vid:pid 303a:1001, generic to every native-USB ESP32-S3 board, not V4-specific) — so unlike V3's CP2102 bridge chip, V4 is permanently NOT auto-matchable by vid:pid. Board auto-detect should fail closed for it every time (manual board selection required, "couldn't confirm automatically" warning shown) — this is expected steady-state behavior, not a gap to close later.
- ❑ After a successful flash, the modal automatically re-probes and shows the NEW firmware's badge/details — same as unplugging and replugging (Section E item 3), but without physically touching the cable.
- ❑ Deliberately test a failure path once (disconnect the board mid-write,
or point at a bad cached asset) — confirm the error surfaces in the
progress log AND that
docs/troubleshooting.md's "LoRa radio firmware flash failed" recovery steps (BOOT+RST bootloader entry, manual esptool/ rnodeconf command) actually get the board back to a flashable state. - ❑ Cancel button only appears (and only works) while still in the "Downloading firmware…" stage — once erasing/writing starts, no cancel affordance is offered.
- ❑ Boot-loop regression (2026-07-23 incident): after a failed flash
(e.g. kill network access mid-download to force a failure), confirm the
mesh listener does NOT auto-resume —
journalctl -u archipelagoshould show a singleLeaving mesh listener stopped after failed flashline and then go quiet for that device, not a repeatingmesh::serial: Opened serial port... Starting Meshcore handshakecycle every few seconds. Reconnect manually via the hot-swap modal afterward and confirm it connects normally (the board itself should be untouched — the download fails before esptool/rnodeconf ever runs). - ❑ Separately, force a device to flap connected/disconnected a few times
in under 20s each (e.g. a marginal USB connection) and confirm
reconnect_delayin the logs actually escalates (5s → 10s → 20s → ...) rather than resetting to 5s on every attempt — seeSTABLE_SESSION_THRESHOLDinmesh/listener/mod.rs.
After this passes: fold the batch + other agent's work into the next release
(OTA binary + frontend tarball + catalog regen/sign/publish for pine-whisper
3.4.2), then re-run tests/lifecycle/run-gate.sh on .228 (back online as
Tailscale shorty-s).