container-doctor.sh's fix_tor_permissions() exact-matches mode '700', but Tor
sets its hidden-service dirs to 2700 (setgid). Every 5-minute doctor run
'fixes' 2700->700 and restarts tor@default; Tor resets it and the cycle
repeats, so Tor never builds a usable consensus/HSDir cache. Result on a live
node: 'No more HSDir available to query', onion peers unresolvable, and with
the FIPS direct path also timing out, mesh sends failed entirely.
Planned as 01-20 in wave 1 so it ships in the same release. FIPS direct
connect_fail is a separate concern, handed to FED-03's transport review.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Deployed the poller fix to archi-dev-box and re-ran the frozen harness
(02-PERF-FINAL.json, 5 runs/surface). Web5 fixed (275ms, below both
its 566ms pre-phase-2 baseline and the 300ms pass bar). Server's
regression closed (574ms, below 738ms baseline) though not yet under
300ms. Fleet substantially improved (790ms, down from a 2631ms
regression). AppDetails restored to at/near its own baseline.
Investigated the two open items the coordinator raised:
- OpenWrtGateway: this run's 5/5 samples failed with a Chromium
"Target crashed" error cascading from an unrelated surface earlier
in the same harness run — recorded as not-measurable, not written
in as data. Separately confirmed via source (OpenWrtGateway.vue's
h1 renders unconditionally, and a "No router configured" RPC error
deterministically shows a real Connect-to-Router form) that the
prior baseline/after/remeasure numbers were measuring a genuine,
substantive disconnected-state UI render, not an empty/error page —
so the "six confirmed regressions" count is not retracted, but the
numbers are flagged as reflecting one specific code branch.
- Discover (1389ms, worst remaining, least improved): profiled
directly and found a SECOND, distinct, phase-2-class cause —
showStagger/card-stagger entrance-animation classes are baked into
the DOM at first mount and never programmatically removed (the flag
is a correctly-scoped once-per-session const, but nothing ever
re-renders to strip the class), so every KeepAlive detach/reattach
cycle restarts the CSS animation on reactivation, replaying the full
entrance cascade on every revisit. Confirmed via a diagnostic
showing DOM card count doubling transiently on every revisit and an
extended animationstart/animationend event log. Not fixed this pass
— the safe fix's blast radius spans 5+ files outside this plan's
scope (Apps.vue, Marketplace.vue, Home.vue, several Web5 sub-cards)
and needs its own real-device verification budget, matching the
precedent 02-02's original KeepAlive rollout needed for this exact
class of change. Named and evidenced, recommended as a dedicated
follow-up rather than expanded into this plan under time pressure.
REQUIREMENTS.md's PERF-02/PERF-03 rows updated to the final state.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Diagnosed on archy-x250-mad2: handle_lnd_createinvoice omits LND's private
flag, so invoices carry route_hints: [] and are unroutable for any node whose
channels are unannounced. Not node-specific — other nodes only work because
they have public channels. Planned as 01-19 in wave 1 so it ships early.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- New `## Re-measurement (gap closure)` section in 02-FINDINGS.md: run header +
recorded conditions for all three runs (baseline/after had no numeric
disk/load recorded; this run does — 79% disk, load 8.70-11.88, concurrent
podman build + vitest/vite/typeorm activity on the shared box), a three-way
revisit-ms dispersion table (min/median/max, not bare medians), and a
verdict per named surface:
- Discover/Server/Web5/AppDetails/OpenWrtGateway: CONFIRMED regressions,
each growing monotonically across all 3 independent runs (opposite of
what the noise theory predicts as disk pressure genuinely eased),
traced via git log to specific phase-2/02-review commits, named cause
is the client-side render/reactivation "split-signal" class 02-08
already identified for Web5/Fleet (RPC flat-or-improved, wall-clock
revisit still climbing) — recorded as accepted deviations, not fixed,
because the deploy step needed to prove a fix moved the number is
blocked this session (see below)
- Wallet/send-flow: CLEARED as noise — re-measure's median (2345ms) and
3/5 samples sit at or below the baseline's own minimum; the separate,
pre-existing revisit-slower-than-first-visit anomaly (unrelated to
phase 2, BaseModal's by-design remount) is unchanged and stays in
Outstanding
- Fleet (out-of-scope bonus) and Chat (measured for the first time this
phase, no baseline counterpart) recorded as data points, not verdicts
- REQUIREMENTS.md: PERF-02/PERF-03 traceability rows updated to point at
this section (scope deviation from files_modified, per plan checker note
— the coverage-table update this plan's own Task 2 text calls for)
Deploy note: Task 1 deployed archi-dev-box to 3e3159fa (frontend-only,
clean tree, dirty=false) before measuring, since 4 of 6 named surfaces are
touched by the 02-review commits the previously-deployed 8fe6217b predates.
Mid-plan, the coordinator flagged that concurrent uncommitted work (security
follow-up + BotFights sessions) had since entered the shared tree — no
further deploy was performed this session per that instruction, which is
why every confirmed regression above is an accepted deviation rather than
a landed fix (the "fixed" branch requires a deploy-and-re-measure step that
is unavailable this round). `neode-ui/e2e/perf/02-REVIEW.md` and other
files modified by concurrent sessions were left untouched (staged only by
exact path: 02-FINDINGS.md, REQUIREMENTS.md).
Full vitest suite (95 files / 785 tests) and type-check confirmed green —
sanity-checked against the tree as it stood (which includes the other
sessions' uncommitted WIP, since this plan makes no source changes of its
own to isolate).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>