security+feat: v1.3.0 — pentest remediation, container reliability, UI overhaul
Security (33 pentest findings addressed): - CRITICAL: backend binds 127.0.0.1, path traversal in tor.rs/dwn fixed - HIGH: federation requires signatures, XSS login redirect, RBAC viewer restricted - HIGH: tar slip prevention, S3 SSRF validation, backup ID validation - MEDIUM: remember-me random secret, TOTP session rotation, password re-auth - LOW: CSP unsafe-inline removed, CORS dev-only, onion/webhook validation Container reliability: - Memory limits on all 37 containers (OOM prevention) - Exited vs stopped state distinction with health-aware status badges - Crash recovery coordination (no more restart cascade) - User-stopped tracking survives reboots - Tiered boot recovery (databases → core → services → apps) UI: - Wallet TransactionsModal, health-aware app status badges - Restart button on containers, exited/crashed red state - Mesh view overhaul, glass button updates, BaseModal/ToggleSwitch - Apps sticky header removed, dev faucet, mutable mock wallet Infrastructure: - LND REST port 8080 exposed over Tor (LND Connect fix) - Nginx cookie_session fix, deploy script Tor config updated - Dev environment: podman auto-start, boot mode simulation Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
d1b48388fb
commit
1a74a930f7
@@ -350,8 +350,26 @@ async fn restart_container(name: &str) -> bool {
|
||||
/// Spawn the health monitor background task.
|
||||
pub fn spawn_health_monitor(state: Arc<StateManager>, data_dir: PathBuf) {
|
||||
tokio::spawn(async move {
|
||||
// Wait 2 minutes for containers to start up
|
||||
tokio::time::sleep(std::time::Duration::from_secs(120)).await;
|
||||
// Wait for boot recovery to complete before starting health checks.
|
||||
// This prevents the health monitor from fighting with crash_recovery
|
||||
// which is starting containers in tier order.
|
||||
info!("Health monitor: waiting for boot recovery to complete...");
|
||||
let wait_start = std::time::Instant::now();
|
||||
loop {
|
||||
if crate::crash_recovery::is_recovery_complete() {
|
||||
break;
|
||||
}
|
||||
// Safety timeout: start anyway after 5 minutes even if recovery hangs
|
||||
if wait_start.elapsed().as_secs() > 300 {
|
||||
warn!("Health monitor: boot recovery did not complete within 5 minutes, starting anyway");
|
||||
break;
|
||||
}
|
||||
tokio::time::sleep(std::time::Duration::from_secs(5)).await;
|
||||
}
|
||||
// Additional cooldown after recovery to let containers stabilize
|
||||
info!("Health monitor: recovery done, waiting 60s for containers to stabilize...");
|
||||
tokio::time::sleep(std::time::Duration::from_secs(60)).await;
|
||||
info!("Health monitor: starting health checks");
|
||||
|
||||
let mut tracker = RestartTracker::new();
|
||||
let mut mem_tracker = MemoryTracker::new();
|
||||
@@ -378,6 +396,9 @@ pub fn spawn_health_monitor(state: Arc<StateManager>, data_dir: PathBuf) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// Load user-stopped list to skip intentionally stopped containers
|
||||
let user_stopped = crate::crash_recovery::load_user_stopped(&data_dir).await;
|
||||
|
||||
// Sort containers by startup tier so databases restart before dependent services
|
||||
let mut unhealthy: Vec<&ContainerHealth> = Vec::new();
|
||||
let mut state_changed = false;
|
||||
@@ -392,6 +413,11 @@ pub fn spawn_health_monitor(state: Arc<StateManager>, data_dir: PathBuf) {
|
||||
continue;
|
||||
}
|
||||
if container.state == "exited" || container.state == "stopped" {
|
||||
// Skip user-stopped containers
|
||||
if user_stopped.contains(&container.name) {
|
||||
debug!("Skipping user-stopped container: {}", container.name);
|
||||
continue;
|
||||
}
|
||||
unhealthy.push(container);
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user