fix(parquet): cut data page byte-budget mini-batches on exact value counts - #10554
fix(parquet): cut data page byte-budget mini-batches on exact value counts#10554adriangb wants to merge 3 commits into
Conversation
967f1b7 to
5af13c0
Compare
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
Benchmark for this request failed. Run configurationrun benchmark arrow_writer
baseline:
ref: "bd5237f38ca3229694f42dd62139435d41eb93f5"Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
Benchmark for this request failed. Run configurationrun benchmark arrow_writer
baseline:
ref: "bd5237f38ca3229694f42dd62139435d41eb93f5"Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
Benchmark for this request failed. Run configurationrun benchmark arrow_writer
baseline:
ref: "bd5237f38ca3229694f42dd62139435d41eb93f5"Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
run benchmark arrow_writer writer_overhead baseline: |
2 similar comments
|
run benchmark arrow_writer writer_overhead baseline: |
|
run benchmark arrow_writer writer_overhead baseline: |
|
run benchmark writer_overhead baseline: |
2 similar comments
|
run benchmark writer_overhead baseline: |
|
run benchmark writer_overhead baseline: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark writer_overhead
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark writer_overhead
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
9a4c021 to
de7a544
Compare
The large-value `arrow_writer` benchmarks all build their arrays with `StringArray::from_iter_values`, so `RecordBatch::try_from_iter` marks the field non-nullable and the column's definition levels are absent. The writer's byte-budget sub-batching resolves such a chunk's value count in O(1) and never inspects levels, so no benchmark today exercises the path taken by a nullable column whose values exceed `data_page_size_limit` — the path that decides how many values share a page, and on `DELTA_BYTE_ARRAY` whether prefix deduplication survives at all. Add nullable shared-prefix and distinct variants, and a repeated (list-of-string) variant for the third level shape: records cannot span data pages, so mini-batches there must step whole records. These resolve the effect they are meant to. Against apache#10554 (which changes exactly this path), measured base/branch/base so the two base passes give a noise floor: | benchmark | delta | noise floor | | --- | --- | --- | | `large_string_shared_prefix_nullable/delta_byte_array` | -31% | 1.6% | | `large_string_shared_prefix_nullable/plain` | +25% | 8.1% | | `large_string_distinct_nullable/plain` | +24% | 4.5% | | `large_string_shared_prefix_list/*` | flat | 2.2% | The list variants staying flat is itself the expected result: a record holding several over-limit values cannot be split across pages whatever the sub-batching does.
…page limit A `BYTE_ARRAY` value larger than `data_page_size_limit` exceeds the limit on its own, so the post-write `should_add_data_page` check cuts a page after every single value. Parquet requires at least one value per data page, so that value is unsplittable and the limit is simply unsatisfiable there. For `DELTA_BYTE_ARRAY` this is destructive rather than merely wasteful: each page boundary discards the encoder's previous value, so every value gets `prefix_length = 0` and the encoding degenerates to exactly `PLAIN`. Columns of large values that share long prefixes stop being deduplicated (apache#10489). Exempt a page's mandatory first value from the byte limit when that value alone already exceeds it, so the limit applies to what follows — the bytes we can still place on another page. The exemption is gated on `ColumnValueEncoder::compresses_against_previous_value`, so only `DELTA_BYTE_ARRAY` opts in; `PLAIN` and `DELTA_LENGTH_BYTE_ARRAY` cost the same wherever a value lands and keep their existing one-value page bound. Pages therefore stay bounded by the value size rather than growing with `write_batch_size`, preserving the fix from apache#9972. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first-value exemption triggers on a page-opening mini-batch holding exactly one value. A null in a chunk changes the level:value ratio, the byte-budget chunker rounds up to two-level mini-batches, and pages that open with a two-value mini-batch miss the exemption: they are cut after two values with the first stored in full. The mini-batch pairing the null with a value has one value, so the page it opens does get the exemption and accumulates the remaining suffixes. Pin that layout ([2, 2, 2, 2, 9] for 16 identical values with a null at index 8) so the limitation is a documented decision rather than an accident, and note it on set_page_size_floor. If the trigger is later keyed on values written to the page (0 -> 1) instead of mini-batch shape, the test fails with fewer, larger pages and should be updated to pin the improved layout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ounts `byte_budget_sub_batch_size` asks the encoder how many *values* fit in a page byte budget, then converts that to a *level* count using the chunk-wide level:value ratio, rounded up. For a chunk with no nulls that is exact. With one null in 17 levels it gives `ceil(17/16) == 2`, and `write_granular_chunk` slices the chunk into uniform two-level windows — most of which carry two values, twice what the budget allowed. The mechanism dates to apache#9972; apache#10505 raises its cost but does not introduce it. Where the encoding compresses a value against its predecessor, that round-up costs whole values of output. 128 values of 2 MiB at one null in 16 write 16.78 MB instead of 2.10 MB; at 8 MiB values it is 64 MiB instead of 8 MiB. Have the chunker return the value count it already computed, and let `write_granular_chunk` end a window by walking definition levels until it has covered that many values. Apply it only where it changes the bytes written, because it roughly doubles the mini-batch count on a nullable column. Two conditions: - The budget must be the data page budget, a constant `data_page_size_limit`, so a one-value budget means the value itself overflows a page. The dictionary page budget is the limit minus what the dictionary already holds, so it shrinks toward zero and reaches a one-value budget on ordinary values; cutting exactly there measured +13.0% on `string/default` and +8.3% on `string/parquet_2`. - The encoding must compress against the previous value. `PLAIN` and `DELTA_LENGTH_BYTE_ARRAY` store a value identically wherever it lands, so value-exact windows leave their output byte for byte the same while doubling the page count — +27.6% on a nullable column for no reduction in output. What the other paths give up is bounded. A ratio-scaled window spans `ceil(values * levels / values_in_chunk)` levels, covering at most two values where one already fills the budget, whatever the null density, and exactly one wherever the ratio is a whole number. Against a 1 MiB limit that is a 4 MiB page for 2 MiB values rather than 2 MiB — and 2 MiB is the floor, since a page must hold a value. It does not scale with `write_batch_size`, which is the failure apache#9972 fixed. `test_column_writer_delta_byte_array_nullable_shared_prefix_partial_dedup` is re-pinned from `[2, 2, 2, 2, 9]` to `[17]` and renamed, the layout apache#10505 left a marker for. `test_column_writer_caps_page_size_with_sparse_nulls` pins two values per page under `PLAIN`, so it fails both if the encoding gate is dropped and if the bound is lost. Closes apache#10538
e19bdf0 to
9ac4026
Compare
Important
Stacked on #10505 — do not merge first. The two commits below it belong to that PR.
Review only this PR's own commit:
1d6ca1af17...9ac402625b.Once #10505 merges I'll rebase and this PR's diff becomes clean on its own.
The problem
byte_budget_sub_batch_sizeasks the encoder how many values fit in a page byte budget, then converts that to a level count using the chunk-wide level:value ratio, rounded up:For a chunk with no nulls this is exact. With one null in 17 levels it gives
ceil(17/16) == 2, andwrite_granular_chunkslices the chunk into uniform two-level windows — most of which carry two values, i.e. twice what the budget allowed. The mechanism predates #10505; it dates to #9972.Where the encoding compresses a value against its predecessor, that round-up costs whole values of output. 128 values of 2 MiB at one null in 16:
At 8 MiB values it is 64 MiB against 8 MiB, and the acceptance case in #10538 — 16 identical 64 KiB values with one null — goes from five pages storing ~5 values in full to one page storing ~1.
The fix
Have the chunker return the value count it already computed, and let
write_granular_chunkend a window by walking definition levels until it has covered that many values. No ratio, no rounding.Then apply it only where it changes the bytes written, because it is not free — value-exact windows roughly double the mini-batch count on a nullable column. Two conditions, both required:
The budget must be the data page budget. That one is a constant
data_page_size_limit, so a one-value budget means the value itself overflows a page. The dictionary page budget is the limit minus what the dictionary already holds, so it shrinks toward zero as the dictionary fills and reaches a one-value budget on perfectly ordinary values; cutting exactly there measured +13.0% onstring/defaultand +8.3% onstring/parquet_2.The encoding must compress against the previous value.
PLAINandDELTA_LENGTH_BYTE_ARRAYstore a value identically wherever it lands, so value-exact windows leave their output byte for byte the same while doubling the page count — measured +27.6% on a nullable column for no reduction in output at all. Thecompresses_against_previous_valueflag #10505 added already marks exactly the right set.What the other paths give up
A ratio-scaled window spans
ceil(values × levels / values_in_chunk)levels. Where one value already fills the budget that covers at most two values, whatever the null density, and exactly one wherever the ratio is a whole number. Measured against a 1 MiB limit:PLAINPLAINPLAINPLAINThe floor is what a page must hold: one value. So the concession is a factor of two above an unavoidable minimum, it does not vary with null density, and — the property #9972 exists for — it does not scale with
write_batch_size. Before #9972 a page took a whole mini-batch: 1024 × 2 MiB, roughly 2000× the limit.Measurements
Base is #10505's head (measured at
fd806be5d3; both branches have since been rebased ontomain, with the trees verified identical across the rebase), so these isolate this PR. Local, run base → branch → base on an idle machine, with the two base passes as a per-benchmark noise floor. Benchmarks are the ones added in #10561...._nullable/plain..._nullable_trailing/delta_byte_array..._nullable/delta_byte_array..._nullable_dense/delta_byte_arraylarge_string_distinct_nullable/delta_byte_arraymedium_string_shared_prefix_nullable/delta_byte_arrayTwo costs remain, both on
DELTA_BYTE_ARRAYwhere the byte win does not materialise:Verified byte-identical output — same length, same hash — between base and this PR for dictionary-encoded nullable columns (four shapes) and for repeated columns (both encodings), confirming those paths are untouched rather than merely unchanged in aggregate.
Scope
Repeated columns are unchanged: records cannot span pages, so a record holding several over-limit values still exceeds the budget. That is inherent to the format.
Tests
test_column_writer_delta_byte_array_nullable_shared_prefix_dedup— re-pinned from[2, 2, 2, 2, 9]to[17]and renamed, the layout fix(parquet): keep DELTA_BYTE_ARRAY dedup for values larger than the page size limit #10505 left a marker for.test_column_writer_caps_page_size_with_sparse_nulls— pins two values per page underPLAIN, so it fails both if the encoding gate is dropped (pages would hold one) and if the bound is lost (they would hold many).Full
parquetsuite green (1312 lib + integration),fmtandclippy -D warningsclean.Note on #10505
Its
bool_to_int_with_iftripscargo clippy -- -D warnings, which CI runs — worth fixing on that branch too. Corrected here as part of editing that test.🤖 Generated with Claude Code