Conversation
A server or proxy can close a keep-alive connection while it sits idle
in a persistent client. The client expires idle connections after
keep_alive_timeout, 5 seconds by default, but a server's idle timeout
can be shorter, and the server starts counting when it sends the
response, before the client has read it. The client only checked
whether its own end of the socket was closed, so it wrote the next
request into the closed connection. That request failed with
HTTP::ResponseHeaderError ("couldn't read response headers"), or with
OpenSSL::SSL::SSLError ("unexpected eof while reading") when the server
skipped close_notify. A server that sends a response before closing,
such as a 408 Request Timeout, was worse: the next request read that
408 as its own response.
Before reusing a connection, read off the previous response, then check
the socket without blocking. Once the previous response has been read,
an idle connection has nothing left to read, so any readable data means
it can't carry another request: an EOF, a reset, a TLS close_notify or
an unsolicited response. Close it then and open a new one. urllib3
makes the same check (wait_for_read(sock, timeout=0.0)), Net::HTTP
reconnects on wait_readable(0) && eof?, and Go's transport drops idle
connections that see an EOF or an unsolicited response.
www.debian.org, lwn.net, www.postgresql.org and www.sqlite.org closed
idle connections 4.8 to 4.98 seconds after the client read a response,
and www.haproxy.org as early as 2 seconds. With the default
keep_alive_timeout, a second request 4.99 seconds later (3 seconds for
www.haproxy.org) failed 16 of 16 times and now succeeds 16 of 16 times
on a new connection. 14 servers that keep idle connections open were
still reused 28 of 28 times after 4.5 seconds idle.
The check happens before any byte of the next request is written, so it
applies to every method, POST included, and can't send a request twice.
It runs after the client is marked dirty, so an exception that
interrupts reading off the previous response still makes the next
request reconnect. It can't help when the server closes the connection
after the request was written; only idempotent requests are safe to
retry then.
Reading off a body larger than 1 MiB closes the connection instead, but
the client then wrote the next request to the closed connection and
raised HTTP::SocketWriteError ("closed stream"). The same check now
reconnects.
test_connection_reuse_enabled_socket_issue_transparently_reopens
asserted that the first request after the server closed the socket
raised. It now expects the client to reopen the connection, as its name
says, and a POST variant covers the same case.
This was referenced Oct 3, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A persistent client writes its next request into a connection the server already closed, then fails with
couldn't read response headers(#420, #459). Before reusing a connection, this branch reads off the previous response and polls the socket without blocking. If anything is readable on an idle connection, the client opens a new one. Usually that's an EOF, a reset, a TLS close_notify or an unsolicited 408, and none of those can carry another request.#459 was closed pointing at
.retriable(#775). That helps, but it's opt-in, it still writes the request into the dead socket and pays the failed round trip before retrying, and it retries non-idempotent requests too. The check here runs before any byte goes out, costs about 1 µs, and is safe for every method, so the default client gets it..retriablestill covers the case below where the close crosses the request.It also happens with the default config.
keep_alive_timeoutis 5 s, and plenty of servers close idle connections at 5 s too. They start counting when they send the response and the client starts after it reads it, so the client loses that race even with matching timeouts. Second request 4.99 s after the first (3 s for haproxy.org), default options:ResponseHeaderError3/3ResponseHeaderError3/3SSLErrorunexpected eof 3/3SSLErrorunexpected eof 3/3ResponseHeaderError4/4Servers that keep idle connections open still get reused: 14 hosts, 28 of 28 after 4.5 s idle with default options (www.samba.org was down that run), and 15 hosts, 30 of 30 after 6 s with
keep_alive_timeout: 300.Local repro, plain and TLS give the same result, MRI 3.4.8 and JRuby 10.1.2.0 too:
ResponseHeaderErrorSocketWriteError: closed streamResponseHeaderErrorResponseHeaderErrorPOSTs, 30 runs per case, counted by the server: after an idle FIN, an idle RST, or a 1 MiB body left unread, main fails all 90 and the server receives none of them. This branch delivers each one exactly once on a new connection, 90/90, on MRI 3.4.8 and JRuby 10.1.2.0. The check runs before any byte is written, so it's safe for every method.
The unread-body case doesn't need an idle server.
performmarks the client clean once headers are read, so the next request flushes the old body insideConnection#send_request, the flush closes the connection because the body is over 1 MiB, and the write fails. Flushing before the check fixes it.Other clients make the same check:
wait_for_read(sock, timeout=0.0)is truebegin_transporton@socket.io.to_io.wait_readable(0) && @socket.eof?stale?is one non-blocking poll: 1.0 µs on MRI 3.4.8, about 1% of a 95 µs loopback request, and 1.5–2.5 µs on JRuby, under 0.5%.Limits:
Connection#stale?andConnection#flush_pending_responseare public becauseClientcalls them. Happy to mark them@api privateif you'd rather not commit to them.test_connection_reuse_enabled_socket_issue_transparently_reopensasserted that the first request after the server closed the socket raised. It now expects the reopen its name describes.test_connection_reuse_enabled_raises_when_server_closes_after_receiving_requestpins down the first limit..mutant.ymlignoresHTTP::Connection*andHTTP::Client*, so CI doesn't mutate this code. Run locally with them included, the new lines have no survivors; the rest are equivalent or in lines this doesn't change.5-x-stable reuses connections the same way (
verify_connection!, client.rb:131), and it's the newest line apps on Ruby < 3.2 can use. The backport is #862.I ran into this through HTTPS proxies, which closed idle CONNECT tunnels after 10 to 120 seconds.
Local repro script
Real-host script
HOSTS=www.debian.org,lwn.net KAT=5 GAP=4.99 ROUNDS=3 bundle exec ruby -Ilib idle_reuse_real.rbKernel trace: Linux, bpftrace kprobes and uprobes on libruby and libssl; an HTTPS proxy closing idle tunnels after 10 s