Repository navigation
fix(qwp): fix QWP query clients failing or reporting an invalid row count on TRUNCATE, SET and similar statements, and QWP senders timing out on close() - #7770
Conversation
QWP egress replied to every statement the compiler executes at parse time with rows_affected = -1 in EXEC_DONE: TRUNCATE, RENAME TABLE, SET, BEGIN / COMMIT / ROLLBACK, VACUUM, CHECKPOINT, REFRESH MATERIALIZED VIEW, WAL SUSPEND / RESUME, and in Enterprise the access control statements. The catch-all branch of executeNonSelect() read CompiledQuery.getAffectedRowsCount(), which CompiledQueryImpl only ever set to -1. rows_affected is an unsigned LEB128 varint, so -1 went out as ten bytes. Clients that decode it as u64 (Rust, C/C++, Node.js) read 18446744073709551615, Java and .NET read -1, and the Go client rejected the frame as a varint overflow: Exec() failed with a transport error after the server had applied the statement, and the client latched the connection as broken. The protocol spec and the client docs state 0 for statements without a row count. executeNonSelect() now leaves rows_affected at 0 for these statements, as its CREATE / DROP / ALTER branches already did. The commit removes CompiledQuery.getAffectedRowsCount(), whose only caller was this branch, and makes QwpEgressFrameWriter.writeExecDone() assert that rows_affected is not negative. QwpEgressDdlExecTest now asserts 0 rows affected for every DDL it runs and covers the parse-time statement families explicitly.
QuestDB runs with -ea in production (questdb.sh, docker-entrypoint.sh), so the assertion was not test-only. When it fired, handleQueryRequest()'s catch-all turned the AssertionError into a QUERY_ERROR after the statement had been applied. One path can still produce a negative count: OperationFutureImpl reads an asynchronously completed UPDATE's 64-bit row count with Unsafe.getInt(), so 2^31 rows or more come back truncated. The assertion would have turned that wrong count into an error for a committed write. The writeExecDone() Javadoc still states that rows_affected must not be negative, and QwpEgressDdlExecTest checks the values clients see.
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The server rejects CHECKPOINT CREATE on Windows, which lacks the sync() system call it relies on, so testParseTimeExecutedStatements failed there with a QUERY_ERROR. Guard only CREATE: CHECKPOINT RELEASE needs no prior CREATE and still runs on every platform.
Points the submodule at java-questdb-client 50d12216 (questdb/java-questdb-client#107). The QWP sender's segment ring made a frame's bytes visible to its I/O thread before it published the frame's FSN, so an ACK the server sent inside that window was clamped one frame short and dropped. After the final frame nothing re-delivered it, and close() and drain() waited out their full timeout although the server had committed every row. This branch's linux-other CI leg hit it as a 300 s close() drain timeout in QwpSenderE2ETest.testConcurrentSenders_sameTable_doubleToDecimal. The client now publishes the FSN before the frame becomes visible. Once the client PR merges, this pointer moves to its squashed commit on main.
|
I found no blocking issues in any of the three PRs, and nothing at Moderate or Minor either. ENT, OSS and client are all approve. Reviewed PR #1265 at level 3 in tandem mode, with #7770 and the nested questdb/java-questdb-client#107. Revisions: ENT Submodule scope
PR titles and descriptions: titles follow Conventional Commits with user-facing descriptions, and the labels match. Nothing to fix. FindingsNone at Critical, Moderate or Minor. Test coverageThe test gate passes with no coverage gaps. Runs at the reviewed revisions:
Summary
Behaviour changes to be aware of (deliberate, not defects):
The scratch worktree has been removed, and all three checkouts are clean at the reviewed SHAs. |
[PR Coverage check]😍 pass : 0 / 0 (0%) |
|
Tandem: questdb/java-questdb-client#107, questdb/questdb-enterprise#1265
This PR bumps the
java-questdb-clientsubmodule to the client PR's branch, for the second fix below. The Enterprise companion bumps thequestdbsubmodule to this branch and adds a test for the access-control statements. Merge order: the client PR first, then re-pointjava-questdb-clienthere at its squash commit and merge this PR, then re-point the Enterprise submodule at this PR's squash commit.Problem
QWP egress answers a non-SELECT statement with an
EXEC_DONEframe that carries the statement'sop_typeandrows_affected. For every statement the compiler executes at parse time, the server put-1intorows_affected. That coversTRUNCATE,RENAME TABLE,SET/RESET,BEGIN/COMMIT/ROLLBACK,DEALLOCATE,VACUUM,CHECKPOINT CREATE/RELEASE,REFRESH MATERIALIZED VIEW,ALTER TABLE ... SUSPEND WAL/RESUME WAL, and in Enterprise the access-control statements (CREATE USER,GRANT,REVOKE, ...).CREATE,DROP,ALTER,INSERTandUPDATEwere not affected.rows_affectedis an unsigned LEB128 varint, so-1went out as ten bytes (FFx9,01). The clients handled it as follows:rows_affectedasu6418446744073709551615bigint18446744073709551615n-1int64, rejecting values above 2^63 - 1Execfails with "varint overflow" after the server has applied the statement, and the client latches the connection as brokenThe QWP egress protocol spec and the documentation of every client state
0for statements without a row count.Cause
The catch-all branch of
QwpEgressUpgradeProcessor.executeNonSelect()readCompiledQuery.getAffectedRowsCount().CompiledQueryImplonly ever assigned-1to the backing field, and nothing else called the getter.Change
executeNonSelect()leavesrows_affectedat0for these statements, as itsCREATE/DROP/ALTERbranches already did.CompiledQuery.getAffectedRowsCount()and its field. They only ever produced-1, and egress was their only caller. Enterprise does not reference them.QwpEgressFrameWriter.writeExecDone()Javadoc states thatrows_affectedmust not be negative.Compatibility and limitations
0instead of-1for these statements. Code that treated-1as "not applicable" observes the change;op_typeremains the documented way to tell DDL from DML.-1. Clients connected to them keep seeing the old values, and the Go client keeps failing these statements against them unless it learns to tolerate the old encoding.rows_affected. QuestDB runs with-eain production, and one path can still produce a negative count:OperationFutureImplreads an asynchronously completedUPDATE's 64-bit row count withUnsafe.getInt(), so 2^31 rows or more come back truncated, on every protocol. An assertion would turn that wrong count into aQUERY_ERRORafter the write was applied. This PR does not change that path.Test plan
QwpEgressDdlExecTest:executeDdl()now asserts0rows affected for every DDL the class runs, andtestParseTimeExecutedStatementsruns 13 parse-time statements and checks the op type and0rows affected for each. On Windows it skipsCHECKPOINT CREATE, which the server rejects there ("Checkpoint is not supported on Windows");CHECKPOINT RELEASEneeds no prior checkpoint and still runs on every platform. Without the fix,testRenameTableandtestParseTimeExecutedStatementsfail withexpected:<0> but was:<-1>.executeDdl()now go throughexecuteExec()and assert their row counts.QwpEgress*Testclasses (365 tests) andCompiledQueryTypeCodeTestpass locally.testParseTimeExecutedStatementspassed five runs with different fragmentation seeds.CREATE/ALTER/DROPof users, groups and service accounts,GRANT,REVOKEandADD/REMOVE USERfails with-1without this fix and passes with it.Second fix: QWP senders timing out in
close()anddrain()Problem
This branch's
linux-otherCI leg failed once with a 300 sclose()drain timeout inQwpSenderE2ETest.testConcurrentSenders_sameTable_doubleToDecimal(build 276712). The egress change above does not cause it: the Java QWP sender could discard the server's ACK of a frame it had just sent. When that frame was the last one,close()anddrain()waited for their full timeout and then reported unacknowledged data, although the server had committed every row. The server's DEBUG log from that run shows it sent the ACK 64 µs after receiving the frame.Cause
The sender's
SegmentRing.appendOrFsn()made a frame's bytes visible to its I/O thread before it published the frame's FSN. The I/O thread sends a frame as soon as its bytes are visible, andSegmentRing.acknowledge()clamps every ACK at the published FSN, so an ACK that arrived between the two stores was capped one frame short and dropped. Server ACKs are cumulative and only follow new frames, so after the final frame nothing re-delivered it. The window spans a few instructions and needs the producer thread descheduled inside it: the hang appeared once in 305 PR builds and 51 macwin builds since 2026-09-07.Change
java-questdb-clientsubmodule to that branch. After the client PR merges, the pointer has to move to its squash commit onmainbefore this PR merges.Test plan
SegmentRingTesttests that fail on the previous client and pass with the fix, and the full client suite passes (3,477 tests).cutlass/qwppass, includingQwpSenderE2ETest(137).QwpSenderE2ETest-style run hung on the previous client. With the same pause at either point of the new order, 0 of 48 hung and every row arrived.