Problem
The Feishu frontend's existing retry loop only runs after cli.start() returns or raises. If synchronous work stalls the SDK event-loop thread, the process can remain alive while it stops processing incoming events. An exit-based process supervisor cannot recover it in that state.
This report concerns event-loop stalls, not every form of silent or half-open connection failure. One concrete trigger is synchronous endpoint discovery during startup or reconnection without an explicit HTTP timeout; see larksuite/oapi-sdk-python#169.
Reproduction
- Run the Feishu frontend under a supervisor that restarts it after failure.
- Stall the WebSocket event-loop thread during initial connection or reconnection.
- Observe that the process remains alive, incoming events stop being handled, and the frontend's outer retry loop does not run.
Local HTTP/WebSocket checks using real lark-oapi 1.6.8 and 1.7.3 clients reproduce an endpoint-discovery stall after an initial event and a simulated disconnection. A fresh process can connect and receive events once the endpoint accepts new requests.
Expected behavior and proposed scope
A supervised frontend should expose a persistently stalled event loop as a process failure, allowing the supervisor to restart it. A small application-owned watchdog around cli.start() can monitor loop progress from an independent thread without requiring an SDK upgrade or duplicating endpoint-discovery code.
The proposed defaults are a 5-second heartbeat/check interval and a 180-second stall threshold. Existing reconnect backoff would stay unchanged. Normal idle time and asynchronous retry waits must not be treated as failures, and watchers must be cleaned up when start() returns or raises.
This recovery would interrupt in-flight tasks and require an external restart policy. It would not detect cases where the event loop remains responsive but the connection or business logic is unhealthy. Feedback on this limited frontend-level recovery approach is welcome.
Problem
The Feishu frontend's existing retry loop only runs after
cli.start()returns or raises. If synchronous work stalls the SDK event-loop thread, the process can remain alive while it stops processing incoming events. An exit-based process supervisor cannot recover it in that state.This report concerns event-loop stalls, not every form of silent or half-open connection failure. One concrete trigger is synchronous endpoint discovery during startup or reconnection without an explicit HTTP timeout; see larksuite/oapi-sdk-python#169.
Reproduction
Local HTTP/WebSocket checks using real
lark-oapi1.6.8 and 1.7.3 clients reproduce an endpoint-discovery stall after an initial event and a simulated disconnection. A fresh process can connect and receive events once the endpoint accepts new requests.Expected behavior and proposed scope
A supervised frontend should expose a persistently stalled event loop as a process failure, allowing the supervisor to restart it. A small application-owned watchdog around
cli.start()can monitor loop progress from an independent thread without requiring an SDK upgrade or duplicating endpoint-discovery code.The proposed defaults are a 5-second heartbeat/check interval and a 180-second stall threshold. Existing reconnect backoff would stay unchanged. Normal idle time and asynchronous retry waits must not be treated as failures, and watchers must be cleaned up when
start()returns or raises.This recovery would interrupt in-flight tasks and require an external restart policy. It would not detect cases where the event loop remains responsive but the connection or business logic is unhealthy. Feedback on this limited frontend-level recovery approach is welcome.