Skip to content

wait for the client's close frame before terminating the twisted websocket connection - #50

Merged
bentsku merged 3 commits into
mainfrom
twisted-websocket-close-handshake
Sep 30, 2026
Merged

bentsku merged 3 commits into
mainfrom
twisted-websocket-close-handshake

Conversation

@bentsku

@bentsku bentsku commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

#42 made the twisted channel terminate the TCP connection as part of the closing handshake, citing RFC 6455 section 7. For a client-initiated close that's right, and this PR doesn't change it: the server echoes the close frame and terminates the TCP connection immediately.

For a server-initiated close (wsClose), the channel also terminated the TCP connection right after sending its close frame. The RFC ties that to having both sent and received a close frame:

  • Section 5.5.1: "After both sending and receiving a Close message, an endpoint considers the WebSocket connection closed and MUST close the underlying TCP connection. The server MUST close the underlying TCP connection immediately."
  • Section 7.1.2: "Once an endpoint has both sent and received a Close control frame, that endpoint SHOULD Close the WebSocket Connection." Only at that point does section 7.1.1's "SHOULD initiate a TCP Close immediately" apply. After only sending, the handshake is just started: the CLOSING state in section 7.1.3.

Section 1.4 explains why the server has to wait, and it describes the failure this causes:

By sending a Close frame and waiting for a Close frame in response, certain cases are avoided where data may be unnecessarily lost. For instance, on some platforms, if a socket is closed with data in the receive queue, a RST packet is sent, which will then cause recv() to fail for the party that received the RST, even if there was data waiting to be read.

This shows up when the server closes while the client is still sending (for example pings), and the client hasn't read all the server's messages yet. Twisted stops reading once the connection is terminated, the client's frames stay unread, and the kernel resets the connection. The client then gets ECONNRESET and loses the messages it hadn't read. #49 moves close onto the reactor right after the queued sends, which makes this happen every time instead of sometimes.

Changes

  • wsClose sends the close frame, but keeps the TCP connection open and keeps reading. The connection is terminated in either of these cases:
    • the client echoes the close frame: dataReceived already handles that, and terminates the connection immediately, as section 5.5.1 requires;
    • WebSocketChannel.closeTimeout (5 s) runs out because the client never echoes. Section 7.1.1 allows closing "via any means available when necessary". The pending timeout is cancelled when the connection closes earlier.
  • The close timeout only starts once the transport's send buffer is no longer full. Otherwise it would run while the messages sent before the close frame are still buffered on the server, and a slow client could run out of time reading them. The channel registers as a producer on the request, so Twisted calls its pauseProducing / resumeProducing when the buffer fills and drains. The websockets library does the same, by waiting for its write buffer to drain before starting its close timer.
  • WebSocketChannel.closeAbortTimeout (30 s) bounds the whole close. If the TCP connection is still open that long after the close frame was sent, it's aborted and any buffered data is dropped. Without it, a client that stops reading would keep the send buffer full forever: the close timeout never starts, and since the request isn't finished, Twisted's HTTP timeouts don't apply either. It also covers a graceful close that stalls for the same reason.
  • The listener sees the websocket as closed as soon as the close frame is sent (poison pill), same as before.
  • Data frames received after the server's close frame are dropped instead of queued behind the poison pill. Section 5.5.1: "there is no guarantee that the endpoint that has already sent a Close frame will continue to process data". Pings are still parsed, but no Pong is sent after the close frame, since section 1.4 says a peer sends no further data after sending a Close frame.
  • WebSocketChannel takes the reactor, so the timeout uses the same reactor as the rest of the resource. Since schedule twisted websocket writes onto the reactor thread #49, all channel operations run on the reactor thread.

Testing

  • test_server_close_while_client_is_sending (twisted only; hypercorn terminates the TCP connection right away as well): the server streams 50k messages and closes while the client keeps pinging. The client must receive every message and then the close frame. It fails every run without this change (Linux: 5 out of 5 failed), and passes with it (0 out of 10 failed on Linux and on macOS).
  • test_close_handshake_server_initiated (twisted and asgi) now echoes the close frame the way a well-behaved client does, and expects the server to terminate TCP after that.
  • test_close_handshake_server_initiated_client_does_not_echo (twisted): a client that never echoes still gets its TCP connection terminated, after closeTimeout (lowered to 0.5 s in the test).
  • The full suite passes on macOS and Linux (188 tests).
  • test_server_close_client_stops_reading (twisted): the server sends 8 MiB and closes, with small socket buffers, while the client doesn't read. The server has to abort the connection after closeAbortTimeout (lowered to 0.5 s in the test). Without the abort deadline it stays connected.
  • The pinger in test_server_close_while_client_is_sending sleeps 1 ms between pings, instead of sending them in a busy loop.
  • The benchmark from schedule twisted websocket writes onto the reactor thread #49 shows no measurable change: 61.5k / 50.5k msg/s for ws / wss sends, and connect/send/close cycles of 0.67 / 1.27 ms. The early-close scenario went from 20 out of 20 connections broken to 0 out of 20.

CI failure on the first run

The first CI run failed test_server_close_while_client_is_sending once, on Python 3.13, on a slow runner (81 s for the suite, versus about 21 s locally). The connection closed while the client was still reading the stream. I couldn't reproduce it: there were no failures in about 40 runs at the original commit, including the full suite and runs in CPU-limited Linux containers. The later CI run passed on all versions.

A likely cause is the close timeout starting before the send buffer drained, which the timeout change above addresses. The busy-loop pinger may also have been what slowed that runner down. Neither is confirmed, since a test for a slow client passed with and without the change, so I left that test out. The timeout change is hardening, not a proven fix.

The same log also has 200 forceAbortClient tracebacks (AttributeError: 'NoneType' object has no attribute 'shutdown'), one per connection of #49's TLS test. When a websocket closes, Request.finish() makes the HTTPChannel schedule its 15 s force-abort. The HTTPChannel never gets connectionLost, because the websocket channel replaced it as the protocol, so the timers fire later on connections that are already closed. That dates from #42's close handling. It's harmless noise and not part of this PR.

🤖 Generated with Claude Code

@bentsku
bentsku added this pull request to stack #51 September 30, 2026 21:42
Base automatically changed from twisted-websocket-thread-safety to main September 30, 2026 21:45
…ocket connection

When the server closed the websocket, it terminated the TCP connection right after sending its close
frame. RFC 6455 section 5.5.1 only has the server terminate it once it has both sent and received a
close frame. Terminating it right away stops reading the client's frames, and closing a socket with
unread data makes the kernel reset the connection, which discards the frames the client has not read
yet (RFC 6455 section 1.4).

The channel now waits for the client's close frame, and terminates the TCP connection after
`closeTimeout` if the client never sends it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bentsku
bentsku force-pushed the twisted-websocket-close-handshake branch from bed56a1 to 4d74399 Compare September 30, 2026 21:45
bentsku and others added 2 commits October 1, 2026 00:17
The close timeout started when the close frame was queued, while the messages sent before it could
still be buffered on the server. A slow client could run out of time while reading them. The channel
now registers as a producer on the request, and only starts the timeout once the transport's send
buffer is no longer full, like the `websockets` library, which waits for the write buffer to drain.

Also throttle the pinger in test_server_close_while_client_is_sending, which sent pings in a busy loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the client stops reading after the server sent its close frame, the send buffer never drains,
so the close timeout never starts and the request never finishes, which also keeps Twisted's HTTP
timeouts from applying. The connection and its buffered data were never released.

The channel now aborts the TCP connection `closeAbortTimeout` (30 s) after sending its close frame,
if it is still open by then.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bentsku
bentsku merged commit 876669b into main Sep 30, 2026
5 checks passed
@bentsku
bentsku deleted the twisted-websocket-close-handshake branch September 30, 2026 22:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant