Skip to content

Rewind IO request bodies after a failed write - #860

Draft
ilyazub wants to merge 1 commit into
httprb:mainfrom
serpapi:rewind-io-body-after-failed-write
Draft

ilyazub wants to merge 1 commit into
httprb:mainfrom
serpapi:rewind-io-body-after-failed-write

Conversation

@ilyazub

@ilyazub ilyazub commented Oct 3, 2026

Copy link
Copy Markdown

Body#each rewinds an IO body after IO.copy_stream has sent all of it. When a write fails partway, copy_stream raises and the rewind is skipped, so the IO stays where the failed write stopped. A retried request, e.g. from .retriable, then sends only the remainder under the original Content-Length, and the server waits for the missing bytes until the read times out.

Origin that resets connection 1 after reading about 64 KiB of the body, and counts body bytes on connection 2. Client: HTTP.timeout(connect: 2, write: 5, read: 2).retriable(tries: 2, delay: 0).post(url, body: 8 MiB):

bytes received on the retry client got
main, StringIO body (2 runs) 7,634,944 and 7,684,096 of 8,388,608 HTTP::OutOfRetriesError, read timed out after 2 s
main, String body 8,388,608 201
this branch, StringIO body (3 runs) 8,388,608 201

On main the retry followed a SocketWriteError in one run and a SocketReadError in the other. Writer#send_request rescues EPIPE, so a broken pipe only surfaces when the response is read, but the rewind was skipped on both paths.

The fix moves the copy into copy_io, which rewinds in an ensure. Body#size already reports the IO's full size as Content-Length, so rewinding to the start matches what the header declares. Pipes still can't be rewound (rewind rescues ESPIPE as before), so a retried pipe body still sends only what's left. That's unchanged.

The new test lets the first 16 KiB chunk through, fails the second with EPIPE, and checks that the next each yields the whole body. It fails on main.

  • MRI 3.4.8: 2333 runs, 0 failures, 100% line and branch coverage. rubocop and yardstick pass, steep shows the same 4 warnings as main. mutant run --since main: 76/76 mutations killed
  • JRuby 10.1.2.0: 2332 runs, 0 failures (normalizer_test's allocation test only runs on MRI). With main's body.rb the new test fails there too
Repro script
# Does http.rb resend an IO request body from where a failed write stopped?
# Usage: cd <http.rb checkout> && bundle exec ruby -Ilib io_resend.rb (BODY=string for the String control)
require "socket"
require "stringio"
require "http"

SIZE = 8 * 1024 * 1024

server = TCPServer.new("127.0.0.1", 0)
port = server.addr[1]
seen = Queue.new

origin = Thread.new do
  2.times do |i|
    sock = server.accept
    head = +""
    head << sock.readpartial(4096) until head.include?("\r\n\r\n")
    headers, rest = head.split("\r\n\r\n", 2)
    length = headers[/^content-length: *(\d+)/i, 1].to_i
    got = rest.bytesize
    if i.zero?
      got += sock.readpartial(65_536).bytesize
      sock.setsockopt(Socket::SOL_SOCKET, Socket::SO_LINGER, [1, 0].pack("ii"))
      sock.close
      seen << "connection 1: content-length #{length}, read #{got} bytes, then reset"
      next
    end

    while got < length && IO.select([sock], nil, nil, 3)
      chunk = sock.read_nonblock(65_536, exception: false)
      break if chunk.nil?

      got += chunk.bytesize unless chunk == :wait_readable
    end
    sock.write("HTTP/1.1 201 Created\r\nContent-Length: 0\r\nConnection: close\r\n\r\n") if got == length
    seen << "connection 2: content-length #{length}, got #{got} bytes (#{got == length ? 'complete, answered 201' : "short by #{length - got}, no answer"})"
    sock.close
  rescue IOError, SystemCallError => e
    seen << "connection #{i + 1}: #{e.class}"
  end
end

body = ENV["BODY"] == "string" ? "x" * SIZE : StringIO.new("x" * SIZE)
retries = []
started = Process.clock_gettime(Process::CLOCK_MONOTONIC)
outcome =
  begin
    HTTP.timeout(connect: 2, write: 5, read: 2)
        .retriable(tries: 2, delay: 0, on_retry: ->(_req, err, _res) { retries << err&.class })
        .post("http://127.0.0.1:#{port}/upload", body: body)
        .status.to_s
  rescue HTTP::Error, SystemCallError => e
    "#{e.class}: #{e.message}"
  end
elapsed = Process.clock_gettime(Process::CLOCK_MONOTONIC) - started
origin.join(10)

puts "http #{HTTP::VERSION}, #{RUBY_DESCRIPTION.split(' (').first}, body=#{body.class}"
puts seen.pop until seen.empty?
puts "retried after: #{retries.inspect}"
puts format("client: %s after %.1fs", outcome, elapsed)

Body#each rewinds an IO source after IO.copy_stream has sent it, so the
request can be sent again. When a write fails partway, for example
because the server reset the connection mid-upload, copy_stream raises
and the rewind is skipped. The IO stays where the failed write stopped,
so a retried request, such as one sent by `.retriable`, sends only the
remainder while its Content-Length still declares the full size. The
server then waits for the missing bytes until the read times out.

Writer#send_request rescues EPIPE, so the same happens when the write
fails with a broken pipe and the request fails later, reading the
response.

Rewind in an ensure, so the IO is back at its start however the copy
ends.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant