Skip to content

fix(feishu): detect stalled event loops for supervised recovery - #826

Open
ytahml wants to merge 1 commit into
lsdefine:mainfrom
ytahml:fix-feishu-watchdog
Open

ytahml wants to merge 1 commit into
lsdefine:mainfrom
ytahml:fix-feishu-watchdog

Conversation

@ytahml

@ytahml ytahml commented Oct 4, 2026 •

Copy link
Copy Markdown

Fixes #825.

Problem

The Feishu frontend already retries with exponential backoff when cli.start() returns or raises. However, synchronous work that stalls the SDK event loop can leave the process alive while it stops processing incoming events, preventing both the outer retry loop and a process supervisor's exit-based restart policy from helping. This change adds an application-owned event-loop watchdog so a supervised deployment can recover from that failure mode.

One concrete trigger is endpoint discovery during startup or reconnection: the SDK versions inspected perform a synchronous HTTP request on the WebSocket event-loop thread without an explicit request timeout. See larksuite/oapi-sdk-python#169 for the dependency-level report. The frontend guard does not require an SDK upgrade or override its endpoint-discovery implementation.

Reproduction and expected behavior

  1. Start the Feishu WebSocket client under a process supervisor configured to restart it after failure.
  2. Block its event-loop thread during initial connection or reconnection, for example with a stalled endpoint-discovery request.
  3. Previously, the process could remain alive while neither event processing nor the outer frontend retry loop progressed.
  4. With this change, a loop that stops advancing for 180 seconds is detected by an independent thread, which logs the failure and exits with status 1. The supervisor can then start a fresh process.

The committed regression tests simulate both startup and runtime stalls in disposable subprocesses with shortened deadlines. Normal idle time and asynchronous retry waits continue advancing the loop and must not trigger recovery.

Changes

  • Keep production changes in frontends/fsapp.py, wrapping the existing cli.start() call.
  • Schedule a lightweight event-loop heartbeat every 5 seconds and check its monotonic timestamp from one daemon thread.
  • Stop and join that thread when cli.start() returns or raises, including KeyboardInterrupt, so retries do not accumulate watchers.
  • Preserve the existing 5-to-120-second reconnect backoff and dependency requirements.
  • Add six focused tests in frontends/tests/test_fsapp_watchdog.py.

Validation

  • python -m pytest -q frontends/tests/test_fsapp_watchdog.py: 6 passed on Python 3.11.
  • Mutation check: disabling the watcher makes both stalled-loop cases fail.
  • Additional local HTTP/WebSocket integration checks using real lark-oapi clients passed for the versions listed below: receive an event, stall endpoint discovery after disconnection, exit via the watchdog, restart in a fresh process, and receive another event. These checks used Python 3.11, dummy credentials, a shortened 0.5-second watchdog deadline, and a simulated process supervisor. They validate watchdog recovery, not every frontend feature or every SDK/dependency combination.
  • Four isolated smoke checks passed on Linux with Python 3.10: startup stall, runtime stall, normal idle/cleanup, and exception cleanup.
  • A deployment smoke check confirmed incoming message handling, an actual model response, and successful outgoing delivery.
  • The full repository test suite was not run. No network or event-loop fault was injected into the live service.

SDK compatibility and dependency declaration

This watchdog does not require lark-oapi>=1.6.8. Its integration point is lark_oapi.ws.client.loop: Client.start() must run the same module-level event loop on which the heartbeat is scheduled.

SDK version Evidence for this watchdog
1.2.10, 1.3.0, 1.4.24, 1.5.5, 1.6.7, 1.6.8, 1.7.3 Local real-SDK recovery checks passed with websockets==15.0.1.
1.6.0 The same recovery check passed with websockets==13.1, respecting this SDK release's websockets>=11,<14 constraint. This yanked release was checked for compatibility, not recommended for installation.
1.2.1 Official wheel source contains the required module-level loop and Client.start() uses it; no runtime check was performed for this version.
1.0.x and 1.2.0 Official wheels do not contain lark_oapi.ws.client; these versions cannot support the existing Feishu WebSocket frontend.

The new top-level loop import makes the last case fail earlier, at frontend import time, including --check and --check-agent. Previously, the existing lark.ws.Client startup path was already unsupported on those releases. This earlier failure is a compatibility limitation of this patch, not evidence that all versions allowed by the current dependency declaration work.

Suggested project follow-up: tighten lark-oapi>=1.0 in the all-frontends extra. The earliest stable SDK release found to provide the required WebSocket module is 1.2.1, so lark-oapi>=1.2.1 is a candidate minimum, subject to validating the complete frontend against it. It should not be raised to 1.6.8 solely for this watchdog. Before choosing the final lower bound, test configuration checks, client initialization, and message handling on the candidate oldest supported version, then keep that version and a current SDK in the compatibility checks. A clear unsupported-SDK error would also improve startup diagnostics.

This PR leaves dependency declarations unchanged; the lower-bound adjustment above is a recommendation, not an implemented or fully validated frontend support guarantee. The runtime checks are representative versions, not an exhaustive sweep of every release between them.

Scope and tradeoffs

  • This detects event-loop stalls, not every unhealthy connection. A loop that remains responsive while a coroutine waits indefinitely, or while a connection no longer receives application events, is outside this guard's scope.
  • Recovery requires an external supervisor, such as systemd with an appropriate restart policy. Standalone execution will exit rather than restart itself.
  • os._exit(1) terminates the whole frontend process without normal Python cleanup and can interrupt in-flight tasks. Raising SystemExit inside the watchdog thread would only stop that thread and would not recover the service.
  • The helper uses the SDK's module-level event loop. It avoids copying SDK request logic, but compatibility with that loop integration point remains relevant when upgrading the SDK.
  • The defaults are a 180-second timeout and a 5-second check interval; tests shorten them to keep failure checks bounded.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feishu frontend cannot recover when the SDK event loop stalls while the process stays alive

1 participant