fix(parquet): cut data page byte-budget mini-batches on exact value counts - #10554
fix(parquet): cut data page byte-budget mini-batches on exact value counts#10554adriangb wants to merge 4 commits into
Conversation
967f1b7 to
5af13c0
Compare
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
Benchmark for this request failed. Run configurationrun benchmark arrow_writer
baseline:
ref: "bd5237f38ca3229694f42dd62139435d41eb93f5"Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
Benchmark for this request failed. Run configurationrun benchmark arrow_writer
baseline:
ref: "bd5237f38ca3229694f42dd62139435d41eb93f5"Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
Benchmark for this request failed. Run configurationrun benchmark arrow_writer
baseline:
ref: "bd5237f38ca3229694f42dd62139435d41eb93f5"Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
run benchmark arrow_writer writer_overhead baseline: |
2 similar comments
|
run benchmark arrow_writer writer_overhead baseline: |
|
run benchmark arrow_writer writer_overhead baseline: |
|
run benchmark writer_overhead baseline: |
2 similar comments
|
run benchmark writer_overhead baseline: |
|
run benchmark writer_overhead baseline: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark writer_overhead
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench writer_overhead File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to fd806be diff Run configurationrun benchmark writer_overhead
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (5af13c0) to main diff Run configurationrun benchmark writer_overhead
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
run benchmark arrow_writer baseline: |
2 similar comments
|
run benchmark arrow_writer baseline: |
|
run benchmark arrow_writer baseline: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_writer File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to fd806be diff Run configurationrun benchmark arrow_writer
baseline:
ref: "fd806be5d3c0ac522fd9e3e0b3dd70b08e527dd8"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing claude/parquet-exact-value-windows-10538 (9a4c021) to main diff Run configurationrun benchmark arrow_writer
baseline:
ref: "main"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
…page limit A `BYTE_ARRAY` value larger than `data_page_size_limit` exceeds the limit on its own, so the post-write `should_add_data_page` check cuts a page after every single value. Parquet requires at least one value per data page, so that value is unsplittable and the limit is simply unsatisfiable there. For `DELTA_BYTE_ARRAY` this is destructive rather than merely wasteful: each page boundary discards the encoder's previous value, so every value gets `prefix_length = 0` and the encoding degenerates to exactly `PLAIN`. Columns of large values that share long prefixes stop being deduplicated (apache#10489). Exempt a page's mandatory first value from the byte limit when that value alone already exceeds it, so the limit applies to what follows — the bytes we can still place on another page. The exemption is gated on `ColumnValueEncoder::compresses_against_previous_value`, so only `DELTA_BYTE_ARRAY` opts in; `PLAIN` and `DELTA_LENGTH_BYTE_ARRAY` cost the same wherever a value lands and keep their existing one-value page bound. Pages therefore stay bounded by the value size rather than growing with `write_batch_size`, preserving the fix from apache#9972. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first-value exemption triggers on a page-opening mini-batch holding exactly one value. A null in a chunk changes the level:value ratio, the byte-budget chunker rounds up to two-level mini-batches, and pages that open with a two-value mini-batch miss the exemption: they are cut after two values with the first stored in full. The mini-batch pairing the null with a value has one value, so the page it opens does get the exemption and accumulates the remaining suffixes. Pin that layout ([2, 2, 2, 2, 9] for 16 identical values with a null at index 8) so the limitation is a documented decision rather than an accident, and note it on set_page_size_floor. If the trigger is later keyed on values written to the page (0 -> 1) instead of mini-batch shape, the test fails with fewer, larger pages and should be updated to pin the improved layout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The byte-budget chunker computes how many values fit in a page budget, then scaled that count to a level count by the chunk-wide level:value ratio, rounded up. For a chunk with no nulls this is exact. With one null in 17 levels, `ceil(17/16) == 2`, and `write_granular_chunk` sliced the chunk into uniform two-level windows that mostly carry two values — so a window could hold twice the values its budget allowed. Return the value count from the chunker and have `write_granular_chunk` end each window by walking definition levels until it has covered that many values. The walk runs only in the granular path, whose values are by definition large enough to overflow a page budget, and that path already makes full def-level and rep-level passes. Two effects, both on columns whose values exceed `data_page_size_limit`: - The apache#9972 per-page bound now holds on nullable columns, one value per page rather than two. Pinned by `test_column_writer_caps_page_size_with_sparse_nulls`, which uses PLAIN so it is independent of the first-value exemption. - The exemption that preserves DELTA_BYTE_ARRAY prefix dedup fires when a page opens with a single-value mini-batch, which is now guaranteed regardless of where nulls fall. 16 identical 64 KiB values with one null, against a 16 KiB limit, go from five pages storing ~5 values in full to one 17-level page storing ~1 — re-pinning the test that apache#10505 left as a marker for exactly this. Repeated columns are unchanged: a record still cannot span pages, so a record holding several over-limit values still exceeds the budget. Closes apache#10538
Cutting every granular mini-batch on an exact value count regressed nullable, dictionary-encoded byte-array writes: +13.0% on `string/default`, +8.3% on `string/parquet_2`, measured over 3 runs whose per-job median drift was within 0.5%. Every `*_non_null` string variant stayed flat, matching the def-level walk being reachable only on nullable columns. The two budgets the chunker works against have different shapes: - the data page budget is a constant `data_page_size_limit`, so a one-value budget means the value itself overflows a page. That is apache#10538's regime, where the level walk is noise next to writing the value. - the dictionary page budget is the limit minus what the dictionary already holds, so it shrinks toward zero as the dictionary fills. A one-value budget is routine there for ordinary values, and cutting one value per window multiplies mini-batches: at 25% nulls the old two-level windows covered 1024 levels in 512 windows, value-exact windows need 768. So scale to levels for the dictionary budget, as before, and cut on exact value counts only against the data page budget. `SubBatch` carries which, and documents why. This keeps the apache#10538 fix whole — `test_column_writer_caps_page_size_ with_sparse_nulls` and `test_column_writer_delta_byte_array_nullable_ shared_prefix_dedup` both disable dictionary encoding, as DELTA_BYTE_ ARRAY dedup requires — and leaves the dictionary page bound exactly where apache#9972 put it.
9a4c021 to
de7a544
Compare
The large-value `arrow_writer` benchmarks all build their arrays with `StringArray::from_iter_values`, so `RecordBatch::try_from_iter` marks the field non-nullable and the column's definition levels are absent. The writer's byte-budget sub-batching resolves such a chunk's value count in O(1) and never inspects levels, so no benchmark today exercises the path taken by a nullable column whose values exceed `data_page_size_limit` — the path that decides how many values share a page, and on `DELTA_BYTE_ARRAY` whether prefix deduplication survives at all. Add nullable shared-prefix and distinct variants, and a repeated (list-of-string) variant for the third level shape: records cannot span data pages, so mini-batches there must step whole records. These resolve the effect they are meant to. Against apache#10554 (which changes exactly this path), measured base/branch/base so the two base passes give a noise floor: | benchmark | delta | noise floor | | --- | --- | --- | | `large_string_shared_prefix_nullable/delta_byte_array` | -31% | 1.6% | | `large_string_shared_prefix_nullable/plain` | +25% | 8.1% | | `large_string_distinct_nullable/plain` | +24% | 4.5% | | `large_string_shared_prefix_list/*` | flat | 2.2% | The list variants staying flat is itself the expected result: a record holding several over-limit values cannot be split across pages whatever the sub-batching does.
Important
Stacked on #10505 — do not merge first. The first two commits below belong to that PR.
Review only this PR's own two commits:
fd806be5d3...9a4c021a1b.Once #10505 merges I'll rebase and this PR's diff becomes clean on its own.
The problem
byte_budget_sub_batch_sizeasks the encoder how many values fit in a page byte budget, then converts that to a level count using the chunk-wide level:value ratio, rounded up:For a chunk with no nulls this is exact. With one null in 17 levels it gives
ceil(17/16) == 2, andwrite_granular_chunkthen slices the chunk into uniform two-level windows — most of which carry two values, i.e. twice what the budget allowed.The mechanism predates #10505; it dates to #9972. #10505 raises its cost but does not introduce it.
The fix
The budget is inherently a value count — it comes from summing value byte sizes. So return it as one, and have
write_granular_chunkend each window by walking definition levels until it has covered that many values. No ratio, no rounding.Non-nullable and fixed-width columns are untouched:
Absent/Uniformdef levels take an O(1) arm.…applied to the data page budget only
The chunker works against two budgets, and they have different shapes:
data_page_size_limitValue-exact windows are nearly free where the budget is constant: they only get narrow when the values are already large enough to dominate the cost of writing them. Where the budget shrinks to near-zero they are not free at all — at 25% nulls the old two-level windows covered 1024 levels in 512 mini-batches, and value-exact windows need 768.
That is measurable (see below), so the dictionary budget keeps the ratio-scaled windowing it has today and only the data page budget cuts on exact value counts.
SubBatchcarries which, and documents why. This costs #10538 nothing:DELTA_BYTE_ARRAYdedup requires dictionary encoding to be off, so the case the issue is about is always on the data page budget.Results
Both effects apply only to values exceeding
data_page_size_limit.The #9972 page bound now holds on nullable columns. One value per page instead of two. This is independent of #10505 — pinned with
PLAIN, where the first-value exemption never fires:DELTA_BYTE_ARRAY dedup becomes complete on nullable columns. #10505's exemption fires when a page opens with a single-value mini-batch — now guaranteed regardless of where nulls fall:
DELTA_BYTE_ARRAYmainPLAIN)[2, 2, 2, 2, 9][17]That is the acceptance criterion in #10538, and it re-pins
test_column_writer_delta_byte_array_nullable_shared_prefix_partial_dedup— the test #10505 added as a marker for exactly this change, per its comment.Benchmarks
arrow_writeron the runner, against this PR's base (fd806be5d3), so the numbers isolate this PR from #10505. A run counts only if its control benchmarks stayed within 3%: the fixed-width types (decimal/*,primitive/*,bool/*), which short-circuit onstatic_always_fits, and the largePLAINbyte-array benchmarks, which opt out of #10505's exemption viacompresses_against_previous_valueand are non-nullable. Neither group can reach a line either PR changes.The regression, measured. Cutting every budget on exact value counts — the first commit alone — regressed nullable dictionary-encoded writes. Three runs, all controls clean:
string/defaultstring/parquet_2decimal/default(control)large_string_shared_prefix/plain(control)Every run is positive, and the nullable/non-null split matches the code:
string_non_null/*stayed within ±2.7% of zero, because non-null columns take the O(1)Absentdef-level arm.The fix, proven by construction. Scoping to the data page budget makes dictionary-encoded columns take a path they never leave. Writing dictionary-encoded nullable columns across four shapes — value sizes 512 B to 16 KiB, null densities from 1-in-2 to 1-in-7, dictionary page limits from 8 KiB to 128 KiB — produces byte-for-byte identical page layouts on this PR and on its base: same page sizes, same per-page value counts, same dictionary page sizes. The cost on that path is not small, it is zero, and that is a deterministic result rather than a statistical one.
The follow-up benchmark runs agree — the runs surviving the control check put
string/parquet_2at +1.0% andstring/defaultat −4.4%, i.e. straddling zero rather than consistently positive — but the runner was noisy enough that evening that most were voided on controls, so treat the byte-identity result as the load-bearing evidence and these as corroboration.large_string_*(dictionary disabled — the path #10538 is about) andwriter_overheadare flat throughout.Scope
Repeated columns are unchanged. Records cannot span pages, so a single record holding several over-limit values still exceeds the budget; that is inherent to the format, not to this heuristic.
The dictionary page keeps the bound #9972 gave it — no better, no worse.
Tests
test_column_writer_caps_page_size_with_sparse_nulls— new.PLAIN, sparse nulls, asserts one value's payload per page. Guards this fix on its own, so it still holds if fix(parquet): keep DELTA_BYTE_ARRAY dedup for values larger than the page size limit #10505 is ever reworked.test_column_writer_delta_byte_array_nullable_shared_prefix_dedup— re-pinned from[2, 2, 2, 2, 9]to[17]and renamed (no longer partial), with the old layout recorded in the comment.Both were verified to fail against the previous ratio-scaled windowing, the DELTA one reproducing
[2, 2, 2, 2, 9]exactly. Both disable dictionary encoding, so they exercise the data page budget path. Fullparquetsuite green (1312 lib + integration),fmtandclippy -D warningsclean.Note on #10505
Its
bool_to_int_with_ifin the test above tripscargo clippy -- -D warnings, which CI runs — worth fixing on that branch too. It's corrected here as part of editing that test.🤖 Generated with Claude Code