Skip to content

fix(qwp): fix sender close() and drain() timing out after the server acknowledged all data - #106

Closed
bluestreak01 wants to merge 1 commit into
mainfrom
fix/qwp-lost-final-ack
Closed

bluestreak01 wants to merge 1 commit into
mainfrom
fix/qwp-lost-final-ack

Conversation

@bluestreak01

Copy link
Copy Markdown
Member

The QWP WebSocket sender could lose the server's acknowledgement (ACK) of a frame it had just sent. When that frame was the last one before the sender went quiet, close() and drain() waited for their full timeout and then reported unacknowledged data, although the server had committed every row. QuestDB CI hit this as a 300 s close() drain timeout in QwpSenderE2ETest.testConcurrentSenders_sameTable_doubleToDecimal (build 276712). The server's DEBUG log from that run shows it committed the final frame and sent its ACK 64 µs after receiving it.

Cause

SegmentRing.appendOrFsn() made a frame's bytes visible to the I/O thread (MmapSegment.publishedCursor) before it stored the frame's FSN in publishedFsn. The I/O thread sends a frame as soon as it lies below publishedOffset(), so the server can commit and ACK it while the producer is still between the two stores, for example when the OS preempts the producer there. acknowledge() clamps every ACK at publishedFsn, so it capped that ACK one frame short and the I/O thread discarded it. Server ACKs are cumulative and only follow new frames: mid-stream the next frame's ACK covers the loss, but after the final frame nothing re-delivers it.

The window is a few instructions wide, so the hang needs the producer descheduled at exactly that point. Across 305 QuestDB PR builds and 51 macwin builds since 2026-09-07 it appeared once. The defect dates from the original store-and-forward implementation (#17).

Change

  • MmapSegment.tryAppend() splits into package-private tryWrite(), which writes the frame and bumps the frame count without publishing it, and publishWritten(), which advances publishedCursor. tryAppend() still does both, so its other callers behave as before.
  • appendOrFsn() calls tryWrite(), stores publishedFsn, then calls publishWritten(). Every frame the I/O thread can send is already covered by publishedFsn, so the clamp no longer drops a legitimate ACK. The clamp itself stays, as the defense-in-depth against bogus server sequence numbers that testAcknowledgeClampsAtPublishedFsn pins.
  • appendOrFsn() runs the high-water manager wakeup after publication rather than between the two stores.

Tradeoffs

  • publishedFsn now leads the visible bytes by at most one frame for an instant, where it previously lagged them; its javadoc says so. The readers that run concurrently with the producer are the ACK clamp and the error-range reporting. Cursor positioning uses frameCount and publishedOffset(), whose relative order is unchanged. The producer itself, recovery, engine close and the orphan drainer read it without a concurrent producer and see no difference.
  • A bogus server ACK arriving in that instant could cover a frame that is fully written but not yet sent. The clamp never guarded against acknowledging written-but-unsent frames, and the I/O thread's own clamp to frames it actually sent still applies, so this adds no exposure.
  • The change reorders the same three volatile stores and adds no work to the append path. CursorEngineAppendLatencyBenchmark (interleaved runs, pinned cores) showed no difference beyond run-to-run noise: 8.25 M vs 8.04 M appends/s with 64-byte payloads (8 runs each, overlapping ranges) and 353.3 k vs 352.9 k appends/s with 4 KiB payloads (10 runs each, about 1% standard deviation), new vs old.

Test plan

  • New SegmentRingTest.testAckOfFrameVisibleDuringAppendIsNotClampedAway: the high-water wakeup, which observes the ring once the appended frame is visible, plays the I/O thread and ACKs the newest visible frame. It fails on the previous code every run and passes with the change.
  • New SegmentRingTest.testAckOfFrameVisibleToConsumerLandsUnderConcurrentAppends: an observer thread ACKs each frame as soon as it becomes visible while the producer appends 200,000 frames across rotations. It failed 12 of 12 runs on the previous code and passed 50 of 50 with the change, 30 of them restricted to two CPUs.
  • Full client test suite: 3,477 tests pass; the 7 skips are @Ignore or OS-gated.
  • QuestDB cutlass/qwp tests against this client: 1,555 pass, including QwpSenderE2ETest 137 of 137.
  • With a 20 ms producer pause injected at the old race point, 16 of 16 QuestDB senders hung on the previous code. With the same pause injected between the two stores of the new code, or right after them, 0 of 48 hung and every row arrived.
  • Compiles at Java 8 language level (class file version 52). Not built on a real JDK 8 locally.

SegmentRing.appendOrFsn() made a frame's bytes visible to the I/O
thread (MmapSegment.publishedCursor) before it stored the frame's FSN
in publishedFsn. The I/O thread sends a frame as soon as it lies below
publishedOffset(), so the server could commit and ACK it while the
producer still sat between the two stores, for example when the OS
preempted the producer there. acknowledge() clamps every ACK at
publishedFsn, so it capped that ACK one frame short and the I/O
thread discarded it. QWP server ACKs are cumulative and only follow
new frames, so nothing re-delivered it once the producer went quiet:
after the final frame, close() and drain() waited out their full
timeout although the server had committed every row. QuestDB CI hit
this as a 300 s close() drain timeout in QwpSenderE2ETest.

MmapSegment.tryAppend() now splits into tryWrite(), which writes the
frame and bumps the frame count without publishing it, and
publishWritten(), which advances publishedCursor. appendOrFsn() calls
tryWrite(), stores publishedFsn, then calls publishWritten(), so
publishedFsn covers every frame the I/O thread can send and the clamp
never drops a legitimate ACK. publishedFsn may now run one frame
ahead of the visible bytes for an instant, never behind them; no
concurrent reader depends on the old direction. The clamp stays as
the defense-in-depth against bogus server sequence numbers.

appendOrFsn() also runs the high-water manager wakeup after
publication instead of between the two stores, so the wakeup no
longer delays the frame's FSN. tryAppend() keeps its behaviour for
every other caller, and the change adds no work to the append path:
it reorders the same volatile stores.

Two SegmentRingTest tests reproduce the lost ACK, one
deterministically through the high-water wakeup and one with a
concurrent observer that ACKs each frame as soon as it becomes
visible. Both fail on the previous code and pass with this change.
@bluestreak01 bluestreak01 added bug Something isn't working QWP labels Oct 7, 2026
@bluestreak01

Copy link
Copy Markdown
Member Author

Superseded by #107: same commit (50d1221), on branch fix/qwp-exec-done-negative-rows-affected so it pairs by name with questdb/questdb#7770. GitHub closes a PR whose head branch is renamed, so this moved to a new PR instead.

@bluestreak01
bluestreak01 deleted the fix/qwp-lost-final-ack branch October 7, 2026 18:59
@mtopolnik

Copy link
Copy Markdown
Contributor

[PR Coverage check]

😍 pass : 18 / 18 (100.00%)

file detail

path covered line new line coverage
🔵 io/questdb/client/cutlass/qwp/client/sf/cursor/MmapSegment.java 6 6 100.00%
🔵 io/questdb/client/cutlass/qwp/client/sf/cursor/SegmentRing.java 12 12 100.00%

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working QWP

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants