mirror of
https://github.com/openclaw/openclaw.git
synced 2026-08-13 06:03:39 -06:00
8624b9acb8
* feat(gateway): recover channels and health promptly after host sleep A dependency-free thaw detector rides the existing 30s maintenance tick: when the process resumes after being frozen >=45s beyond cadence (laptop sleep, VM pause, SIGSTOP), the gateway restarts running channel accounts (dead sockets otherwise take up to ~35 minutes to notice), refreshes health/presence, and resets the event-loop histogram so the freeze does not read as degradation. Admission is rechecked before every recovery side effect; a suspension beginning mid-recovery re-pends the thaw, and timed-out channel stops complete their two-call restart in one pass. The macOS app cooperates: NSWorkspace sleep/wake observers in GatewayConnectivityCoordinator best-effort prepare a local gateway suspension before sleep and resume it on wake, never blocking sleep. The lease is bound to the route that prepared it and always cleared on wake; route or mode changes across sleep drop it to self-expiry. Live proof: SIGSTOP 85s on an isolated dev gateway -> 'host thaw detected: process was frozen ~57683ms', channels restarted, health ok, eventLoop degraded=false after thaw. * fix(macos): resume a sleep lease whose prepare response arrives after wake A prepare completing after didWake previously discarded the lease id, fencing the gateway until the two-minute expiry after micro-sleeps; the late response now resumes immediately. Document the conservative route-token drift tradeoff. * fix(macos): retry wake resume after refreshing the dead post-sleep transport After real sleep the WebSocket is usually dead exactly when resume runs; refresh the endpoint first, then attempt resume up to three times with bounded delays, clearing the lease only on success or exhaustion. A new sleep cycle aborts in-flight retries. * fix(gateway): bound plugin stopAccount so channel stops cannot wedge recovery stopChannel awaited plugin stopAccount unbounded; a never-settling stop hung the thaw restart (and health-monitor sweeps) and held the single-flight recovery guard forever. Race it against the existing 5s stop timeout; the timed-out path flows into the established recoveryStopTimedOut two-call restart contract. Regression wedges pre-fix. * refactor(gateway): move thaw channel restart off ChannelManager and fence mid-pass restartRunningChannelAccounts is a standalone helper over the public manager surface with a shouldContinue probe checked before every stop and start, so a suspension committing while an account stop is awaited leaves later accounts untouched. Regression covers the mid-pass close. * fix(gateway): sanitize late writes from an abandoned stopAccount An abandoned (timed-out) stopAccount can settle after its replacement started; route its late setStatus writes through the existing stale-task sanitizer so they cannot repaint or tear down the replacement. Regression fails pre-fix.