Skip to content

introduce append_array() - #10765

Open
Rich-T-kid wants to merge 4 commits into
apache:mainfrom
Rich-T-kid:rich-T-kid/bench-dict-append-array-impl
Open

introduce append_array()#10765
Rich-T-kid wants to merge 4 commits into
apache:mainfrom
Rich-T-kid:rich-T-kid/bench-dict-append-array-impl

Conversation

@Rich-T-kid

@Rich-T-kid Rich-T-kid commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

the utf8/binary -> dict<_,utf8/binary> currently calls append_value() in a loop even when it has the entire array ahead of time. this PR avoids that by working with th entire array at once as well as branching on weather or not the array contains nulls.

What changes are included in this PR?

new append_array() method on the byteDictionaryBuilder.

Are these changes tested?

yes & benchmarked

Are there any user-facing changes?

yes, new append_array() method on the byteDictionaryBuilder.

@github-actions github-actions Bot added arrow Changes to the arrow crate arrow-cast arrow-array labels Aug 19, 2026
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

env:
BENCH_FILTER: cast binary to dict

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/bench-dict-append-array-impl branch from 14e35c9 to 33cded2 Compare August 19, 2026 20:58
@adriangbot

This comment was marked as duplicate.

@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed before finishing (Kubernetes reason: BackoffLimitExceeded).

Benchmarks requested: cast_kernels

Kubernetes message
Job has reached the specified backoff limit

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

env:
BENCH_FILTER: cast binary to dict

@adriangbot

Copy link
Copy Markdown

Hi @Rich-T-kid, your benchmark configuration could not be parsed (#10765 (comment)).

Error: invalid configuration: unknown field BENCH_FILTER, expected one of env, baseline, changed at line 3 column 1

Usage:

run benchmark <name>           # run specific benchmark(s)
run benchmarks                 # run default suite
run benchmarks <name1> <name2> # run specific benchmarks

Any benchmark name is accepted: bench.sh suite names (e.g. tpch, clickbench_partitioned, wide_schema) and Criterion bench targets (e.g. sql_planner) are resolved automatically. A name that matches neither fails on the runner.

Per-side configuration (run benchmark tpch followed by):

env:
# shared env is inherited by BOTH the build and the run, so build
# flags go here. Builds default to no debuginfo for speed; opt back
# in for hung-job gdb dumps and cap jobs to stay within memory:
CARGO_PROFILE_RELEASE_DEBUG: "1"
CARGO_BUILD_JOBS: "1"
baseline:
ref: v45.0.0
env:
# per-side env only reaches the benchmark run, not the build
DATAFUSION_RUNTIME_MEMORY_LIMIT: 1G
changed:
ref: v46.0.0
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: 2G

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels
env:
BENCH_FILTER: cast binary to dict

@adriangbot

Copy link
Copy Markdown

Hi @Rich-T-kid, your benchmark configuration could not be parsed (#10765 (comment)).

Error: invalid configuration: unknown field BENCH_FILTER, expected one of env, baseline, changed at line 2 column 1

Usage:

run benchmark <name>           # run specific benchmark(s)
run benchmarks                 # run default suite
run benchmarks <name1> <name2> # run specific benchmarks

Any benchmark name is accepted: bench.sh suite names (e.g. tpch, clickbench_partitioned, wide_schema) and Criterion bench targets (e.g. sql_planner) are resolved automatically. A name that matches neither fails on the runner.

Per-side configuration (run benchmark tpch followed by):

env:
# shared env is inherited by BOTH the build and the run, so build
# flags go here. Builds default to no debuginfo for speed; opt back
# in for hung-job gdb dumps and cap jobs to stay within memory:
CARGO_PROFILE_RELEASE_DEBUG: "1"
CARGO_BUILD_JOBS: "1"
baseline:
ref: v45.0.0
env:
# per-side env only reaches the benchmark run, not the build
DATAFUSION_RUNTIME_MEMORY_LIMIT: 1G
changed:
ref: v46.0.0
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: 2G

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

env:
BENCH_FILTER: "cast binary to dict"

@adriangbot

This comment was marked as duplicate.

@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed before finishing (Kubernetes reason: BackoffLimitExceeded).

Benchmarks requested: cast_kernels

Kubernetes message
Job has reached the specified backoff limit

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

show benchmark queue

@adriangbot

Copy link
Copy Markdown

Hi @Rich-T-kid, you asked to view the benchmark queue (#10765 (comment)).

Comment Repo PR User Benchmarks Status
#5347364903 apache/arrow-rs #10690 etseidl ["arrow_reader"] running

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

@adriangbot

This comment was marked as duplicate.

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/bench-dict-append-array-impl (33cded2) to 7ec3f5a (merge-base) diff

Run configuration
run benchmark cast_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                              main                                   rich-T-kid_bench-dict-append-array-impl
-----                                                              ----                                   ---------------------------------------
"cast decimal128 to float64"                                       1.00     27.1±0.02µs        ? ?/sec    1.00     27.1±0.05µs        ? ?/sec
"cast decimal128 to int64"                                         1.00     48.0±0.47µs        ? ?/sec    1.00     47.9±0.49µs        ? ?/sec
"cast decimal128 to int8"                                          1.00     60.7±0.34µs        ? ?/sec    1.00     60.5±0.61µs        ? ?/sec
"cast decimal256 to float64"                                       1.00     68.6±0.10µs        ? ?/sec    1.00     68.5±0.08µs        ? ?/sec
"cast decimal256 to int64"                                         1.00    152.1±0.87µs        ? ?/sec    1.00    151.4±0.75µs        ? ?/sec
"cast float64 to decimal128(32, 3)"                                1.00     35.2±0.28µs        ? ?/sec    1.03     36.4±0.41µs        ? ?/sec
"cast invalid float64 to to decimal128(32, 3)"                     1.09     23.5±1.22µs        ? ?/sec    1.00     21.6±1.00µs        ? ?/sec
"cast string to decimal128(38, 3)"                                 1.07    117.8±0.46µs        ? ?/sec    1.00    110.5±0.57µs        ? ?/sec
"cast string to decimal256(76, 3)"                                 1.02    150.5±0.42µs        ? ?/sec    1.00    148.2±0.32µs        ? ?/sec
cast binary dict to string view (sparse)                           1.23     55.5±0.33µs        ? ?/sec    1.00     45.0±0.11µs        ? ?/sec
cast binary view to string                                         1.00     68.6±0.75µs        ? ?/sec    1.00     68.5±0.69µs        ? ?/sec
cast binary view to string view                                    1.00     63.2±1.47µs        ? ?/sec    1.00     63.4±0.70µs        ? ?/sec
cast binary view to wide string                                    1.00     69.5±0.66µs        ? ?/sec    1.00     69.8±0.66µs        ? ?/sec
cast date32 to date64 512                                          1.00    332.8±0.93ns        ? ?/sec    1.02    340.5±0.83ns        ? ?/sec
cast date64 to date32 512                                          1.00   1428.5±3.47ns        ? ?/sec    1.02   1455.5±4.81ns        ? ?/sec
cast decimal128 to decimal128 512                                  1.00      6.9±0.01µs        ? ?/sec    1.00      6.9±0.01µs        ? ?/sec
cast decimal128 to decimal128 512 lower precision                  1.00     19.9±0.02µs        ? ?/sec    1.01     20.2±0.03µs        ? ?/sec
cast decimal128 to decimal128 512 with lower scale (infallible)    1.00     45.9±0.11µs        ? ?/sec    1.00     45.9±0.05µs        ? ?/sec
cast decimal128 to decimal128 512 with same scale                  1.01     75.5±0.22ns        ? ?/sec    1.00     74.9±0.38ns        ? ?/sec
cast decimal128 to decimal256 512                                  1.00     26.2±0.02µs        ? ?/sec    1.00     26.2±0.03µs        ? ?/sec
cast decimal256 to decimal128 512                                  1.01    321.1±0.30µs        ? ?/sec    1.00    318.6±0.75µs        ? ?/sec
cast decimal256 to decimal256 512                                  1.00     82.0±0.08µs        ? ?/sec    1.01     83.0±0.21µs        ? ?/sec
cast decimal256 to decimal256 512 with same scale                  1.00     76.0±1.35ns        ? ?/sec    1.02     77.2±1.60ns        ? ?/sec
cast dict to string view                                           1.00     15.0±0.01µs        ? ?/sec    1.00     14.9±0.01µs        ? ?/sec
cast dict to string view (sparse)                                  1.04      5.5±0.08µs        ? ?/sec    1.00      5.3±0.07µs        ? ?/sec
cast f32 to string 512                                             1.00     12.1±0.06µs        ? ?/sec    1.00     12.1±0.05µs        ? ?/sec
cast f64 to string 512                                             1.00     15.4±0.05µs        ? ?/sec    1.01     15.5±0.02µs        ? ?/sec
cast float32 to int32 512                                          1.00   1376.8±3.25ns        ? ?/sec    1.01   1392.0±4.24ns        ? ?/sec
cast float64 to float32 512                                        1.00    666.1±3.89ns        ? ?/sec    1.05    699.3±5.12ns        ? ?/sec
cast float64 to uint64 512                                         1.00  1455.8±18.52ns        ? ?/sec    1.00  1461.2±41.61ns        ? ?/sec
cast i64 to string 512                                             1.00      8.8±0.03µs        ? ?/sec    1.00      8.8±0.03µs        ? ?/sec
cast int32 to float32 512                                          1.00    709.2±6.19ns        ? ?/sec    1.04    737.3±5.43ns        ? ?/sec
cast int32 to float64 512                                          1.00    729.9±3.36ns        ? ?/sec    1.05    764.6±3.90ns        ? ?/sec
cast int32 to int32 512                                            1.00    171.1±1.04ns        ? ?/sec    1.00    171.4±1.55ns        ? ?/sec
cast int32 to int64 512                                            1.02    705.7±4.44ns        ? ?/sec    1.00    690.5±5.35ns        ? ?/sec
cast int32 to uint32 512                                           1.00   1389.7±3.64ns        ? ?/sec    1.01   1400.6±2.19ns        ? ?/sec
cast int64 to int32 512                                            1.00   1487.9±3.66ns        ? ?/sec    1.02   1524.1±3.04ns        ? ?/sec
cast nested dict to dict                                           1.00      4.7±0.02µs        ? ?/sec    1.00      4.7±0.02µs        ? ?/sec
cast no runs of int32s to ree<int32>                               1.00     58.9±1.94µs        ? ?/sec    1.01     59.5±0.39µs        ? ?/sec
cast runs of 10 string to ree<int32>                               1.00      8.7±0.06µs        ? ?/sec    1.01      8.7±0.02µs        ? ?/sec
cast runs of 1000 int32s to ree<int32>                             1.01      3.4±0.01µs        ? ?/sec    1.00      3.4±0.01µs        ? ?/sec
cast string single run to ree<int32>                               1.00     27.5±0.02µs        ? ?/sec    1.00     27.4±0.22µs        ? ?/sec
cast string to binary view 512                                     1.00      2.3±0.02µs        ? ?/sec    1.02      2.3±0.02µs        ? ?/sec
cast string view to binary view                                    1.00     83.0±0.84ns        ? ?/sec    1.00     83.2±0.87ns        ? ?/sec
cast string view to dict                                           1.00    154.5±0.40µs        ? ?/sec    1.01    155.2±0.38µs        ? ?/sec
cast string view to string                                         1.00     42.9±0.89µs        ? ?/sec    1.02     43.9±0.77µs        ? ?/sec
cast string view to wide string                                    1.00     42.7±0.83µs        ? ?/sec    1.01     43.0±0.81µs        ? ?/sec
cast time32s to time32ms 512                                       1.00   1446.2±3.16ns        ? ?/sec    1.01   1465.7±3.72ns        ? ?/sec
cast time32s to time64us 512                                       1.00    323.9±0.51ns        ? ?/sec    1.05    340.6±0.69ns        ? ?/sec
cast time64ns to time32s 512                                       1.00    414.9±0.57ns        ? ?/sec    1.02    424.6±0.95ns        ? ?/sec
cast timestamp_ms to i64 512                                       1.01    247.2±1.42ns        ? ?/sec    1.00    244.3±0.78ns        ? ?/sec
cast timestamp_ms to timestamp_ns 512                              1.00   1859.1±2.42ns        ? ?/sec    1.00   1867.8±2.89ns        ? ?/sec
cast timestamp_ns to timestamp_s 512                               1.00    170.3±1.41ns        ? ?/sec    1.01    171.5±2.70ns        ? ?/sec
cast utf8 to date32 512                                            1.00      6.9±0.05µs        ? ?/sec    1.01      7.0±0.04µs        ? ?/sec
cast utf8 to date64 512                                            1.01     31.9±0.16µs        ? ?/sec    1.00     31.7±0.15µs        ? ?/sec
cast utf8 to f32                                                   1.00      5.6±0.03µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
cast utf8 to i32                                                   1.00      5.3±0.04µs        ? ?/sec    1.00      5.3±0.04µs        ? ?/sec
cast wide string to binary view 512                                1.00      4.0±0.08µs        ? ?/sec    1.01      4.0±0.07µs        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 575.1s
Peak memory 15.4 MiB
Avg memory 10.3 MiB
CPU user 572.3s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 575.1s
Peak memory 17.0 MiB
Avg memory 15.8 MiB
CPU user 570.7s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

env:
BENCH_FILTER: cast binary to dict

@adriangbot

This comment was marked as outdated.

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

@adriangbot

This comment was marked as duplicate.

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

show benchmark queue

@adriangbot

Copy link
Copy Markdown

Hi @Rich-T-kid, you asked to view the benchmark queue (#10765 (comment)).

Comment Repo PR User Benchmarks Status
#5349672232 apache/arrow-rs #10765 Rich-T-kid ["cast_kernels"] running

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/bench-dict-append-array-impl (d6a16fb) to 6bb5e2b (merge-base) diff

Run configuration
run benchmark cast_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                              main                                   rich-T-kid_bench-dict-append-array-impl
-----                                                              ----                                   ---------------------------------------
"cast decimal128 to float64"                                       1.00     27.1±0.03µs        ? ?/sec    1.00     27.1±0.01µs        ? ?/sec
"cast decimal128 to int64"                                         1.00     47.8±0.41µs        ? ?/sec    1.00     47.9±0.26µs        ? ?/sec
"cast decimal128 to int8"                                          1.01     60.9±0.55µs        ? ?/sec    1.00     60.3±0.76µs        ? ?/sec
"cast decimal256 to float64"                                       1.00     68.5±0.09µs        ? ?/sec    1.00     68.7±0.09µs        ? ?/sec
"cast decimal256 to int64"                                         1.00    151.9±0.69µs        ? ?/sec    1.00    151.4±0.71µs        ? ?/sec
"cast float64 to decimal128(32, 3)"                                1.00     35.4±0.33µs        ? ?/sec    1.03     36.6±0.32µs        ? ?/sec
"cast invalid float64 to to decimal128(32, 3)"                     1.00     23.9±1.14µs        ? ?/sec    1.02     24.4±1.87µs        ? ?/sec
"cast string to decimal128(38, 3)"                                 1.06    117.0±0.43µs        ? ?/sec    1.00    110.8±0.57µs        ? ?/sec
"cast string to decimal256(76, 3)"                                 1.02    150.6±0.40µs        ? ?/sec    1.00    147.6±0.41µs        ? ?/sec
cast binary dict to string view (sparse)                           1.00     45.0±0.15µs        ? ?/sec    1.00     44.9±0.17µs        ? ?/sec
cast binary to dict high cardinality                               1.15    212.6±0.33µs        ? ?/sec    1.00    185.2±0.34µs        ? ?/sec
cast binary to dict high cardinality no nulls                      1.19    194.9±0.65µs        ? ?/sec    1.00    163.6±0.34µs        ? ?/sec
cast binary to dict low cardinality                                1.18    181.2±0.25µs        ? ?/sec    1.00    153.4±0.39µs        ? ?/sec
cast binary to dict low cardinality no nulls                       1.24    163.4±0.36µs        ? ?/sec    1.00    132.1±0.25µs        ? ?/sec
cast binary to dict medium cardinality                             1.18    181.7±0.40µs        ? ?/sec    1.00    154.5±0.31µs        ? ?/sec
cast binary to dict medium cardinality no nulls                    1.23    164.6±0.45µs        ? ?/sec    1.00    133.8±0.30µs        ? ?/sec
cast binary view to string                                         1.00     68.7±0.95µs        ? ?/sec    1.00     68.4±0.86µs        ? ?/sec
cast binary view to string view                                    1.02     63.7±2.05µs        ? ?/sec    1.00     62.4±1.60µs        ? ?/sec
cast binary view to wide string                                    1.00     69.2±0.90µs        ? ?/sec    1.02     70.4±0.62µs        ? ?/sec
cast date32 to date64 512                                          1.00    330.7±5.99ns        ? ?/sec    1.04    343.0±0.39ns        ? ?/sec
cast date64 to date32 512                                          1.00   1459.5±5.03ns        ? ?/sec    1.00   1466.2±3.59ns        ? ?/sec
cast decimal128 to decimal128 512                                  1.00      6.9±0.00µs        ? ?/sec    1.00      6.9±0.00µs        ? ?/sec
cast decimal128 to decimal128 512 lower precision                  1.00     20.0±0.02µs        ? ?/sec    1.00     20.0±0.03µs        ? ?/sec
cast decimal128 to decimal128 512 with lower scale (infallible)    1.00     45.8±0.05µs        ? ?/sec    1.00     46.0±0.06µs        ? ?/sec
cast decimal128 to decimal128 512 with same scale                  1.01     75.6±0.52ns        ? ?/sec    1.00     74.9±0.28ns        ? ?/sec
cast decimal128 to decimal256 512                                  1.00     26.2±0.01µs        ? ?/sec    1.00     26.2±0.02µs        ? ?/sec
cast decimal256 to decimal128 512                                  1.00    320.3±0.31µs        ? ?/sec    1.00    318.9±0.50µs        ? ?/sec
cast decimal256 to decimal256 512                                  1.00     82.1±0.07µs        ? ?/sec    1.01     82.9±0.07µs        ? ?/sec
cast decimal256 to decimal256 512 with same scale                  1.07     81.2±3.95ns        ? ?/sec    1.00     76.0±1.33ns        ? ?/sec
cast dict to string view                                           1.11     16.5±0.01µs        ? ?/sec    1.00     14.8±0.01µs        ? ?/sec
cast dict to string view (sparse)                                  1.00      5.3±0.07µs        ? ?/sec    1.00      5.2±0.06µs        ? ?/sec
cast f32 to string 512                                             1.00     12.0±0.05µs        ? ?/sec    1.00     12.0±0.07µs        ? ?/sec
cast f64 to string 512                                             1.00     15.3±0.03µs        ? ?/sec    1.01     15.4±0.06µs        ? ?/sec
cast float32 to int32 512                                          1.00   1358.0±7.03ns        ? ?/sec    1.00   1364.2±4.26ns        ? ?/sec
cast float64 to float32 512                                        1.02    703.8±4.24ns        ? ?/sec    1.00    687.6±5.41ns        ? ?/sec
cast float64 to uint64 512                                         1.04  1473.5±11.25ns        ? ?/sec    1.00   1422.7±4.72ns        ? ?/sec
cast i64 to string 512                                             1.00      8.7±0.04µs        ? ?/sec    1.01      8.8±0.06µs        ? ?/sec
cast int32 to float32 512                                          1.00    701.4±4.33ns        ? ?/sec    1.00    701.1±5.09ns        ? ?/sec
cast int32 to float64 512                                          1.00    721.9±3.82ns        ? ?/sec    1.01    730.7±4.23ns        ? ?/sec
cast int32 to int32 512                                            1.00    171.5±0.77ns        ? ?/sec    1.00    171.1±1.02ns        ? ?/sec
cast int32 to int64 512                                            1.00    680.8±4.93ns        ? ?/sec    1.02    695.5±2.97ns        ? ?/sec
cast int32 to uint32 512                                           1.01   1405.9±1.66ns        ? ?/sec    1.00   1392.7±1.46ns        ? ?/sec
cast int64 to int32 512                                            1.00   1481.8±2.59ns        ? ?/sec    1.00   1484.2±1.34ns        ? ?/sec
cast nested dict to dict                                           1.00      4.5±0.01µs        ? ?/sec    1.01      4.6±0.01µs        ? ?/sec
cast no runs of int32s to ree<int32>                               1.03     60.7±1.40µs        ? ?/sec    1.00     58.8±1.37µs        ? ?/sec
cast runs of 10 string to ree<int32>                               1.00      8.7±0.06µs        ? ?/sec    1.01      8.8±0.02µs        ? ?/sec
cast runs of 1000 int32s to ree<int32>                             1.01      3.4±0.01µs        ? ?/sec    1.00      3.4±0.01µs        ? ?/sec
cast string single run to ree<int32>                               1.00     27.5±0.03µs        ? ?/sec    1.00     27.4±0.02µs        ? ?/sec
cast string to binary view 512                                     1.00      2.3±0.04µs        ? ?/sec    1.01      2.3±0.02µs        ? ?/sec
cast string view to binary view                                    1.00     82.6±1.09ns        ? ?/sec    1.01     83.5±1.03ns        ? ?/sec
cast string view to dict                                           1.00    154.3±0.47µs        ? ?/sec    1.01    155.1±0.34µs        ? ?/sec
cast string view to string                                         1.00     43.4±0.81µs        ? ?/sec    1.07     46.6±8.01µs        ? ?/sec
cast string view to wide string                                    1.00     43.5±0.84µs        ? ?/sec    1.08     46.7±8.80µs        ? ?/sec
cast time32s to time32ms 512                                       1.00  1454.1±12.98ns        ? ?/sec    1.01   1468.0±3.19ns        ? ?/sec
cast time32s to time64us 512                                       1.00    327.7±1.41ns        ? ?/sec    1.05    342.6±0.35ns        ? ?/sec
cast time64ns to time32s 512                                       1.00    418.4±1.40ns        ? ?/sec    1.02    427.3±0.24ns        ? ?/sec
cast timestamp_ms to i64 512                                       1.00    245.7±1.11ns        ? ?/sec    1.00    246.2±1.58ns        ? ?/sec
cast timestamp_ms to timestamp_ns 512                              1.00   1856.6±3.79ns        ? ?/sec    1.00   1865.4±3.88ns        ? ?/sec
cast timestamp_ns to timestamp_s 512                               1.02    173.6±2.91ns        ? ?/sec    1.00    170.2±1.34ns        ? ?/sec
cast utf8 to date32 512                                            1.00      7.0±0.05µs        ? ?/sec    1.00      7.0±0.05µs        ? ?/sec
cast utf8 to date64 512                                            1.00     32.0±0.11µs        ? ?/sec    1.04     33.2±0.17µs        ? ?/sec
cast utf8 to f32                                                   1.00      5.6±0.04µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
cast utf8 to i32                                                   1.00      5.3±0.05µs        ? ?/sec    1.01      5.3±0.07µs        ? ?/sec
cast wide string to binary view 512                                1.00      4.0±0.07µs        ? ?/sec    1.00      4.0±0.07µs        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 640.1s
Peak memory 16.4 MiB
Avg memory 11.3 MiB
CPU user 636.4s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 640.1s
Peak memory 15.7 MiB
Avg memory 11.4 MiB
CPU user 634.6s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author
cast binary to dict high cardinality                               1.15    212.6±0.33µs        ? ?/sec    1.00    185.2±0.34µs        ? ?/sec
cast binary to dict high cardinality no nulls                      1.19    194.9±0.65µs        ? ?/sec    1.00    163.6±0.34µs        ? ?/sec
cast binary to dict low cardinality                                1.18    181.2±0.25µs        ? ?/sec    1.00    153.4±0.39µs        ? ?/sec
cast binary to dict low cardinality no nulls                       1.24    163.4±0.36µs        ? ?/sec    1.00    132.1±0.25µs        ? ?/sec
cast binary to dict medium cardinality                             1.18    181.7±0.40µs        ? ?/sec    1.00    154.5±0.31µs        ? ?/sec
cast binary to dict medium cardinality no nulls                    1.23    164.6±0.45µs        ? ?/sec    1.00    133.8±0.30µs        ? ?/sec

🚀

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/bench-dict-append-array-impl branch from d6a16fb to faca997 Compare August 20, 2026 01:05
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

1 similar comment
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels

@adriangbot

This comment was marked as duplicate.

@adriangbot

This comment was marked as duplicate.

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/bench-dict-append-array-impl (dca89f9) to 365c760 (merge-base) diff

Run configuration
run benchmark cast_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                              main                                   rich-T-kid_bench-dict-append-array-impl
-----                                                              ----                                   ---------------------------------------
"cast decimal128 to float64"                                       1.00     27.2±0.02µs        ? ?/sec    1.00     27.1±0.02µs        ? ?/sec
"cast decimal128 to int64"                                         1.00     47.9±0.49µs        ? ?/sec    1.00     48.0±0.58µs        ? ?/sec
"cast decimal128 to int8"                                          1.01     60.8±0.39µs        ? ?/sec    1.00     60.3±0.54µs        ? ?/sec
"cast decimal256 to float64"                                       1.00     68.6±0.11µs        ? ?/sec    1.00     68.6±0.09µs        ? ?/sec
"cast decimal256 to int64"                                         1.00    152.1±0.79µs        ? ?/sec    1.00    152.8±0.70µs        ? ?/sec
"cast float64 to decimal128(32, 3)"                                1.00     35.4±0.30µs        ? ?/sec    1.05     37.0±0.38µs        ? ?/sec
"cast invalid float64 to to decimal128(32, 3)"                     1.00     23.7±0.83µs        ? ?/sec    1.00     23.7±1.14µs        ? ?/sec
"cast string to decimal128(38, 3)"                                 1.06    116.6±0.44µs        ? ?/sec    1.00    110.4±0.62µs        ? ?/sec
"cast string to decimal256(76, 3)"                                 1.02    150.7±0.39µs        ? ?/sec    1.00    147.8±0.38µs        ? ?/sec
cast binary dict to string view (sparse)                           1.07     47.8±3.71µs        ? ?/sec    1.00     44.8±0.08µs        ? ?/sec
cast binary to dict high cardinality                               1.15    211.9±0.34µs        ? ?/sec    1.00    184.8±0.31µs        ? ?/sec
cast binary to dict high cardinality no nulls                      1.20    195.2±0.31µs        ? ?/sec    1.00    163.3±0.30µs        ? ?/sec
cast binary to dict low cardinality                                1.18    180.7±0.35µs        ? ?/sec    1.00    152.9±0.31µs        ? ?/sec
cast binary to dict low cardinality no nulls                       1.24    163.4±0.34µs        ? ?/sec    1.00    132.2±0.26µs        ? ?/sec
cast binary to dict medium cardinality                             1.18    182.0±0.31µs        ? ?/sec    1.00    154.5±0.27µs        ? ?/sec
cast binary to dict medium cardinality no nulls                    1.24    164.6±0.33µs        ? ?/sec    1.00    133.3±0.26µs        ? ?/sec
cast binary view to string                                         1.00     68.8±0.86µs        ? ?/sec    1.00     68.7±1.06µs        ? ?/sec
cast binary view to string view                                    1.00     63.4±1.97µs        ? ?/sec    1.00     63.4±1.47µs        ? ?/sec
cast binary view to wide string                                    1.00     69.4±0.72µs        ? ?/sec    1.02     70.7±0.91µs        ? ?/sec
cast date32 to date64 512                                          1.00    323.5±0.52ns        ? ?/sec    1.06    343.6±0.95ns        ? ?/sec
cast date64 to date32 512                                          1.00   1435.6±3.68ns        ? ?/sec    1.04   1486.0±8.20ns        ? ?/sec
cast decimal128 to decimal128 512                                  1.00      6.9±0.00µs        ? ?/sec    1.00      6.9±0.01µs        ? ?/sec
cast decimal128 to decimal128 512 lower precision                  1.02     20.4±0.09µs        ? ?/sec    1.00     20.1±0.02µs        ? ?/sec
cast decimal128 to decimal128 512 with lower scale (infallible)    1.00     45.9±0.06µs        ? ?/sec    1.00     46.0±0.07µs        ? ?/sec
cast decimal128 to decimal128 512 with same scale                  1.01     75.9±0.68ns        ? ?/sec    1.00     75.4±0.34ns        ? ?/sec
cast decimal128 to decimal256 512                                  1.00     26.2±0.02µs        ? ?/sec    1.00     26.2±0.02µs        ? ?/sec
cast decimal256 to decimal128 512                                  1.01    320.2±0.17µs        ? ?/sec    1.00    318.2±0.17µs        ? ?/sec
cast decimal256 to decimal256 512                                  1.00     82.0±0.09µs        ? ?/sec    1.01     82.9±0.06µs        ? ?/sec
cast decimal256 to decimal256 512 with same scale                  1.00     76.4±1.27ns        ? ?/sec    1.01     77.2±1.43ns        ? ?/sec
cast dict to string view                                           1.00     14.9±0.01µs        ? ?/sec    1.00     15.0±0.01µs        ? ?/sec
cast dict to string view (sparse)                                  1.02      5.3±0.06µs        ? ?/sec    1.00      5.3±0.06µs        ? ?/sec
cast f32 to string 512                                             1.00     12.0±0.05µs        ? ?/sec    1.01     12.1±0.05µs        ? ?/sec
cast f64 to string 512                                             1.00     15.4±0.04µs        ? ?/sec    1.00     15.5±0.05µs        ? ?/sec
cast float32 to int32 512                                          1.00   1354.7±4.09ns        ? ?/sec    1.02   1381.5±2.66ns        ? ?/sec
cast float64 to float32 512                                        1.02    687.2±4.95ns        ? ?/sec    1.00    676.4±4.99ns        ? ?/sec
cast float64 to uint64 512                                         1.01   1442.6±6.22ns        ? ?/sec    1.00   1428.9±2.60ns        ? ?/sec
cast i64 to string 512                                             1.00      8.8±0.03µs        ? ?/sec    1.01      8.9±0.03µs        ? ?/sec
cast int32 to float32 512                                          1.00    695.1±4.76ns        ? ?/sec    1.00    694.5±3.94ns        ? ?/sec
cast int32 to float64 512                                          1.00    710.0±3.08ns        ? ?/sec    1.02    726.5±3.50ns        ? ?/sec
cast int32 to int32 512                                            1.00    171.4±1.28ns        ? ?/sec    1.05    180.5±4.79ns        ? ?/sec
cast int32 to int64 512                                            1.03    712.0±9.91ns        ? ?/sec    1.00    690.6±5.44ns        ? ?/sec
cast int32 to uint32 512                                           1.01   1408.3±3.11ns        ? ?/sec    1.00   1393.2±1.15ns        ? ?/sec
cast int64 to int32 512                                            1.01   1488.7±1.10ns        ? ?/sec    1.00   1478.7±1.19ns        ? ?/sec
cast nested dict to dict                                           1.00      4.5±0.01µs        ? ?/sec    1.01      4.6±0.01µs        ? ?/sec
cast no runs of int32s to ree<int32>                               1.00     56.9±1.30µs        ? ?/sec    1.00     56.7±1.29µs        ? ?/sec
cast runs of 10 string to ree<int32>                               1.00      8.7±0.07µs        ? ?/sec    1.01      8.8±0.08µs        ? ?/sec
cast runs of 1000 int32s to ree<int32>                             1.00      3.4±0.01µs        ? ?/sec    1.02      3.5±0.01µs        ? ?/sec
cast string single run to ree<int32>                               1.00     27.4±0.02µs        ? ?/sec    1.00     27.4±0.02µs        ? ?/sec
cast string to binary view 512                                     1.00      2.2±0.04µs        ? ?/sec    1.01      2.3±0.01µs        ? ?/sec
cast string view to binary view                                    1.00     82.7±0.80ns        ? ?/sec    1.04     85.9±2.57ns        ? ?/sec
cast string view to dict                                           1.00    154.3±0.46µs        ? ?/sec    1.01    155.2±0.38µs        ? ?/sec
cast string view to string                                         1.00     43.1±0.92µs        ? ?/sec    1.00     43.1±0.92µs        ? ?/sec
cast string view to wide string                                    1.00     43.1±0.84µs        ? ?/sec    1.00     43.1±0.81µs        ? ?/sec
cast time32s to time32ms 512                                       1.00   1437.9±2.96ns        ? ?/sec    1.03   1483.5±3.30ns        ? ?/sec
cast time32s to time64us 512                                       1.00    324.4±0.53ns        ? ?/sec    1.06    343.4±0.74ns        ? ?/sec
cast time64ns to time32s 512                                       1.00    412.7±0.38ns        ? ?/sec    1.02    421.7±3.83ns        ? ?/sec
cast timestamp_ms to i64 512                                       1.01    249.2±3.28ns        ? ?/sec    1.00    247.7±2.16ns        ? ?/sec
cast timestamp_ms to timestamp_ns 512                              1.00   1851.4±5.64ns        ? ?/sec    1.01   1864.3±4.43ns        ? ?/sec
cast timestamp_ns to timestamp_s 512                               1.00    170.1±2.11ns        ? ?/sec    1.08    182.9±6.99ns        ? ?/sec
cast utf8 to date32 512                                            1.00      7.0±0.06µs        ? ?/sec    1.00      6.9±0.06µs        ? ?/sec
cast utf8 to date64 512                                            1.00     32.0±0.12µs        ? ?/sec    1.00     31.8±0.15µs        ? ?/sec
cast utf8 to f32                                                   1.00      5.6±0.04µs        ? ?/sec    1.01      5.6±0.03µs        ? ?/sec
cast utf8 to i32                                                   1.00      5.3±0.05µs        ? ?/sec    1.00      5.3±0.05µs        ? ?/sec
cast wide string to binary view 512                                1.00      4.0±0.07µs        ? ?/sec    1.02      4.1±0.07µs        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 640.1s
Peak memory 15.9 MiB
Avg memory 11.3 MiB
CPU user 635.4s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 640.1s
Peak memory 17.1 MiB
Avg memory 11.5 MiB
CPU user 635.7s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/bench-dict-append-array-impl (dca89f9) to 365c760 (merge-base) diff

Run configuration
run benchmark cast_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                              main                                   rich-T-kid_bench-dict-append-array-impl
-----                                                              ----                                   ---------------------------------------
"cast decimal128 to float64"                                       1.00     27.2±0.01µs        ? ?/sec    1.00     27.1±0.10µs        ? ?/sec
"cast decimal128 to int64"                                         1.00     48.1±0.77µs        ? ?/sec    1.00     48.3±0.83µs        ? ?/sec
"cast decimal128 to int8"                                          1.00     61.1±0.50µs        ? ?/sec    1.00     61.0±0.59µs        ? ?/sec
"cast decimal256 to float64"                                       1.00     68.8±0.19µs        ? ?/sec    1.00     68.8±0.14µs        ? ?/sec
"cast decimal256 to int64"                                         1.01    151.8±0.88µs        ? ?/sec    1.00    151.0±0.64µs        ? ?/sec
"cast float64 to decimal128(32, 3)"                                1.00     35.3±0.30µs        ? ?/sec    1.05     37.0±0.25µs        ? ?/sec
"cast invalid float64 to to decimal128(32, 3)"                     1.01     23.9±1.08µs        ? ?/sec    1.00     23.6±0.94µs        ? ?/sec
"cast string to decimal128(38, 3)"                                 1.06    116.9±0.49µs        ? ?/sec    1.00    110.2±0.54µs        ? ?/sec
"cast string to decimal256(76, 3)"                                 1.02    151.4±0.44µs        ? ?/sec    1.00    147.8±0.42µs        ? ?/sec
cast binary dict to string view (sparse)                           1.00     45.2±0.38µs        ? ?/sec    1.23     55.5±0.42µs        ? ?/sec
cast binary to dict high cardinality                               1.15    211.9±0.34µs        ? ?/sec    1.00    185.0±0.38µs        ? ?/sec
cast binary to dict high cardinality no nulls                      1.19    194.6±0.33µs        ? ?/sec    1.00    163.2±0.38µs        ? ?/sec
cast binary to dict low cardinality                                1.18    181.2±0.32µs        ? ?/sec    1.00    153.1±0.27µs        ? ?/sec
cast binary to dict low cardinality no nulls                       1.24    163.2±0.34µs        ? ?/sec    1.00    132.1±0.31µs        ? ?/sec
cast binary to dict medium cardinality                             1.18    182.2±0.30µs        ? ?/sec    1.00    154.5±0.29µs        ? ?/sec
cast binary to dict medium cardinality no nulls                    1.23    164.6±0.30µs        ? ?/sec    1.00    133.6±0.32µs        ? ?/sec
cast binary view to string                                         1.01     68.5±0.77µs        ? ?/sec    1.00     68.2±0.69µs        ? ?/sec
cast binary view to string view                                    1.01     63.4±1.64µs        ? ?/sec    1.00     62.8±0.93µs        ? ?/sec
cast binary view to wide string                                    1.00     69.1±0.72µs        ? ?/sec    1.01     70.1±0.69µs        ? ?/sec
cast date32 to date64 512                                          1.00    332.5±1.03ns        ? ?/sec    1.03    343.2±1.38ns        ? ?/sec
cast date64 to date32 512                                          1.02   1468.7±4.59ns        ? ?/sec    1.00   1440.8±4.48ns        ? ?/sec
cast decimal128 to decimal128 512                                  1.00      6.9±0.00µs        ? ?/sec    1.00      6.9±0.01µs        ? ?/sec
cast decimal128 to decimal128 512 lower precision                  1.00     20.0±0.07µs        ? ?/sec    1.01     20.2±0.02µs        ? ?/sec
cast decimal128 to decimal128 512 with lower scale (infallible)    1.00     45.9±0.13µs        ? ?/sec    1.00     45.8±0.07µs        ? ?/sec
cast decimal128 to decimal128 512 with same scale                  1.02     76.6±1.12ns        ? ?/sec    1.00     75.3±0.25ns        ? ?/sec
cast decimal128 to decimal256 512                                  1.00     26.2±0.02µs        ? ?/sec    1.00     26.2±0.01µs        ? ?/sec
cast decimal256 to decimal128 512                                  1.01    321.7±0.42µs        ? ?/sec    1.00    318.2±0.17µs        ? ?/sec
cast decimal256 to decimal256 512                                  1.00     81.9±0.07µs        ? ?/sec    1.01     83.0±0.06µs        ? ?/sec
cast decimal256 to decimal256 512 with same scale                  1.01     76.9±1.23ns        ? ?/sec    1.00     76.1±1.36ns        ? ?/sec
cast dict to string view                                           1.00     14.9±0.02µs        ? ?/sec    1.00     14.9±0.12µs        ? ?/sec
cast dict to string view (sparse)                                  1.02      5.4±0.32µs        ? ?/sec    1.00      5.3±0.08µs        ? ?/sec
cast f32 to string 512                                             1.00     12.0±0.06µs        ? ?/sec    1.00     11.9±0.06µs        ? ?/sec
cast f64 to string 512                                             1.00     15.4±0.04µs        ? ?/sec    1.01     15.5±0.05µs        ? ?/sec
cast float32 to int32 512                                          1.00   1356.7±4.36ns        ? ?/sec    1.03   1403.3±6.09ns        ? ?/sec
cast float64 to float32 512                                        1.03    699.3±3.03ns        ? ?/sec    1.00    677.4±4.51ns        ? ?/sec
cast float64 to uint64 512                                         1.03   1466.8±5.80ns        ? ?/sec    1.00   1422.8±3.87ns        ? ?/sec
cast i64 to string 512                                             1.00      8.8±0.03µs        ? ?/sec    1.00      8.8±0.04µs        ? ?/sec
cast int32 to float32 512                                          1.00    686.0±3.10ns        ? ?/sec    1.01    696.1±3.28ns        ? ?/sec
cast int32 to float64 512                                          1.00    699.1±2.52ns        ? ?/sec    1.05    730.7±2.61ns        ? ?/sec
cast int32 to int32 512                                            1.00    170.7±0.88ns        ? ?/sec    1.01    172.1±1.21ns        ? ?/sec
cast int32 to int64 512                                            1.00    675.0±3.13ns        ? ?/sec    1.05    706.3±4.20ns        ? ?/sec
cast int32 to uint32 512                                           1.00   1405.8±2.11ns        ? ?/sec    1.00   1403.0±5.85ns        ? ?/sec
cast int64 to int32 512                                            1.01   1483.2±1.78ns        ? ?/sec    1.00   1468.0±1.95ns        ? ?/sec
cast nested dict to dict                                           1.00      4.5±0.01µs        ? ?/sec    1.02      4.6±0.01µs        ? ?/sec
cast no runs of int32s to ree<int32>                               1.00     57.4±1.07µs        ? ?/sec    1.03     58.9±1.04µs        ? ?/sec
cast runs of 10 string to ree<int32>                               1.00      8.8±0.05µs        ? ?/sec    1.00      8.8±0.08µs        ? ?/sec
cast runs of 1000 int32s to ree<int32>                             1.01      3.5±0.01µs        ? ?/sec    1.00      3.4±0.01µs        ? ?/sec
cast string single run to ree<int32>                               1.01     27.6±0.67µs        ? ?/sec    1.00     27.4±0.03µs        ? ?/sec
cast string to binary view 512                                     1.00      2.3±0.04µs        ? ?/sec    1.03      2.3±0.01µs        ? ?/sec
cast string view to binary view                                    1.01     84.9±1.77ns        ? ?/sec    1.00     83.6±0.99ns        ? ?/sec
cast string view to dict                                           1.00    154.3±0.40µs        ? ?/sec    1.00    154.8±0.31µs        ? ?/sec
cast string view to string                                         1.00     43.2±0.86µs        ? ?/sec    1.00     43.1±0.89µs        ? ?/sec
cast string view to wide string                                    1.00     43.1±0.84µs        ? ?/sec    1.00     43.0±0.82µs        ? ?/sec
cast time32s to time32ms 512                                       1.00   1427.7±2.75ns        ? ?/sec    1.01   1444.0±2.93ns        ? ?/sec
cast time32s to time64us 512                                       1.00    332.9±0.45ns        ? ?/sec    1.03    343.1±1.53ns        ? ?/sec
cast time64ns to time32s 512                                       1.00    415.8±0.43ns        ? ?/sec    1.02    426.1±1.36ns        ? ?/sec
cast timestamp_ms to i64 512                                       1.00    245.4±0.86ns        ? ?/sec    1.00    245.3±0.41ns        ? ?/sec
cast timestamp_ms to timestamp_ns 512                              1.00   1859.0±5.84ns        ? ?/sec    1.00   1866.8±2.66ns        ? ?/sec
cast timestamp_ns to timestamp_s 512                               1.00    170.3±0.84ns        ? ?/sec    1.01    172.2±1.67ns        ? ?/sec
cast utf8 to date32 512                                            1.00      6.9±0.04µs        ? ?/sec    1.00      6.9±0.05µs        ? ?/sec
cast utf8 to date64 512                                            1.01     32.1±0.36µs        ? ?/sec    1.00     31.8±0.15µs        ? ?/sec
cast utf8 to f32                                                   1.05      5.9±0.33µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
cast utf8 to i32                                                   1.01      5.3±0.04µs        ? ?/sec    1.00      5.3±0.04µs        ? ?/sec
cast wide string to binary view 512                                1.00      4.0±0.08µs        ? ?/sec    1.01      4.1±0.08µs        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 640.1s
Peak memory 16.4 MiB
Avg memory 11.3 MiB
CPU user 637.4s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 640.1s
Peak memory 16.5 MiB
Avg memory 11.5 MiB
CPU user 635.7s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/bench-dict-append-array-impl branch from dca89f9 to e327c25 Compare August 20, 2026 02:35
@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/bench-dict-append-array-impl branch from e327c25 to 54a655d Compare August 20, 2026 02:35
@Rich-T-kid
Rich-T-kid marked this pull request as ready for review August 20, 2026 02:39
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

reproducible results

cast binary to dict high cardinality                               1.15    211.9±0.34µs        ? ?/sec    1.00    185.0±0.38µs        ? ?/sec
cast binary to dict high cardinality no nulls                      1.19    194.6±0.33µs        ? ?/sec    1.00    163.2±0.38µs        ? ?/sec
cast binary to dict low cardinality                                1.18    181.2±0.32µs        ? ?/sec    1.00    153.1±0.27µs        ? ?/sec
cast binary to dict low cardinality no nulls                       1.24    163.2±0.34µs        ? ?/sec    1.00    132.1±0.31µs        ? ?/sec
cast binary to dict medium cardinality                             1.18    182.2±0.30µs        ? ?/sec    1.00    154.5±0.29µs        ? ?/sec
cast binary to dict medium cardinality no nulls                    1.23    164.6±0.30µs        ? ?/sec    1.00    133.6±0.32µs        ? ?/sec

this is ready for review now @Jefffrey

@Jefffrey

Copy link
Copy Markdown
Contributor

run benchmark cast_kernels
env:
BENCH_FILTER: cast binary to dict

@Jefffrey

Copy link
Copy Markdown
Contributor

fyi need that indentation on env; doesnt show in the rendered output but comment is like this

run benchmark cast_kernels
env:
  BENCH_FILTER: cast binary to dict

@adriangbot

This comment was marked as duplicate.

@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed before finishing (Kubernetes reason: BackoffLimitExceeded).

Benchmarks requested: cast_kernels

Kubernetes message
Job has reached the specified backoff limit

File an issue against this benchmark runner

let start = offsets[row_idx].as_usize();
let end = offsets[row_idx + 1].as_usize();
// SAFETY: offsets are valid by GenericByteArray invariants
let bytes = unsafe { raw_data.get_unchecked(start..end) };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would be nice if we could find a way to deduplicate this with get_or_insert_key without regressing performance 🤔

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think placing this in a anon func would be the best way to solve this.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

7be8b21 does just that

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

fyi need that indentation on env; doesnt show in the rendered output but comment is like this

run benchmark cast_kernels
env:
  BENCH_FILTER: cast binary to dict

ahh okay makes sense.

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels
env:
BENCH_FILTER: cast binary to dict

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5350920745-1787-nncrr 6.12.85+ #1 SMP Wed Jun 17 20:31:55 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-T-kid/bench-dict-append-array-impl (54a655d) to 7ec3f5a (merge-base) diff

Run configuration
run benchmark cast_kernels
env:
  BENCH_FILTER: "cast binary to dict"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench cast_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed before finishing (Kubernetes reason: BackoffLimitExceeded).

Benchmarks requested: cast_kernels

Kubernetes message
Job has reached the specified backoff limit

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

thank you @Jefffrey 🚀

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels
env:
BENCH_FILTER: cast binary to dict

@adriangbot

Copy link
Copy Markdown

Hi @Rich-T-kid, your benchmark configuration could not be parsed (#10765 (comment)).

Error: invalid configuration: unknown field BENCH_FILTER, expected one of env, baseline, changed at line 2 column 1

Usage:

run benchmark <name>           # run specific benchmark(s)
run benchmarks                 # run default suite
run benchmarks <name1> <name2> # run specific benchmarks

Any benchmark name is accepted: bench.sh suite names (e.g. tpch, clickbench_partitioned, wide_schema) and Criterion bench targets (e.g. sql_planner) are resolved automatically. A name that matches neither fails on the runner.

Per-side configuration (run benchmark tpch followed by):

env:
# shared env is inherited by BOTH the build and the run, so build
# flags go here. Builds default to no debuginfo for speed; opt back
# in for hung-job gdb dumps and cap jobs to stay within memory:
CARGO_PROFILE_RELEASE_DEBUG: "1"
CARGO_BUILD_JOBS: "1"
baseline:
ref: v45.0.0
env:
# per-side env only reaches the benchmark run, not the build
DATAFUSION_RUNTIME_MEMORY_LIMIT: 1G
changed:
ref: v46.0.0
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: 2G

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark cast_kernels
env:
BENCH_FILTER: cast binary to dict

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5374822842-1853-mfjq4 6.12.85+ #1 SMP Sat Jun 27 09:31:30 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-T-kid/bench-dict-append-array-impl (891c7ff) to 5ec9eaf (merge-base) diff

Run configuration
run benchmark cast_kernels
env:
  BENCH_FILTER: "cast binary to dict"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench cast_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/bench-dict-append-array-impl (891c7ff) to 5ec9eaf (merge-base) diff

Run configuration
run benchmark cast_kernels
env:
  BENCH_FILTER: "cast binary to dict"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                              main                                   rich-T-kid_bench-dict-append-array-impl
-----                                              ----                                   ---------------------------------------
cast binary to dict high cardinality               1.04    212.3±0.30µs        ? ?/sec    1.00    204.3±0.56µs        ? ?/sec
cast binary to dict high cardinality no nulls      1.00    194.9±0.27µs        ? ?/sec    1.00    194.8±0.41µs        ? ?/sec
cast binary to dict low cardinality                1.04    180.3±0.28µs        ? ?/sec    1.00    172.7±0.35µs        ? ?/sec
cast binary to dict low cardinality no nulls       1.00    163.2±0.31µs        ? ?/sec    1.00    162.6±0.33µs        ? ?/sec
cast binary to dict medium cardinality             1.03    181.6±0.35µs        ? ?/sec    1.00    176.2±0.27µs        ? ?/sec
cast binary to dict medium cardinality no nulls    1.00    164.7±0.26µs        ? ?/sec    1.00    164.1±0.38µs        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 65.0s
Peak memory 11.3 MiB
Avg memory 10.8 MiB
CPU user 62.9s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 65.0s
Peak memory 12.6 MiB
Avg memory 10.9 MiB
CPU user 62.1s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Introduce Append_Array() method on dictionaryByteBuilder

3 participants