Skip to content

Support off-heap group-by state for DISTINCTCOUNTULL (config-gated, default off) - #19402

Draft
xiangfu0 wants to merge 3 commits into
apache:masterfrom
xiangfu0:xiangfu0/offheap-agg-ull
Draft

xiangfu0 wants to merge 3 commits into
apache:masterfrom
xiangfu0:xiangfu0/offheap-agg-ull

Conversation

@xiangfu0

@xiangfu0 xiangfu0 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

PR flow

Off-heap group-by state for DISTINCTCOUNTULL: from config to holder creation, update, and release.

flowchart TD
  N0["Set groupByOffHeap in QueryContext #40;F1#44; F6#41;"]:::stModified
  N1["Create DefaultGroupByExecutor #40;F4#44; F5#41;"]:::stModified
  N2["Create group key generator #40;wrapped if off#45;heap#41; #40;F9#41;"]:::stModified
  N3["Create off#45;heap result holder for DISTINCTCOUNTULL #40;F8#44; F9#41;"]:::stModified
  N4["Register off#45;heap holder on resource tracker #40;F9#41;"]:::stModified
  N5["Update off#45;heap state during aggregation #40;F8#41;"]:::stModified
  N6["Close group key generator #40;releasing off#45;heap holder#41; #40;F3#44; F4#44; F5#41;"]:::stModified
  N0 -->|"passes config"| N1
  N1 -->|"creates generator"| N2
  N1 -->|"creates holder"| N3
  N3 -->|"registers holder"| N4
  N2 -->|"provides group keys"| N5
  N5 -->|"triggers close"| N6
  classDef stAdded fill:#dafbe1,stroke:#1a7f37,color:#1f2328,stroke-width:2px
  classDef stModified fill:#fff8c5,stroke:#9a6700,color:#1f2328,stroke-width:2px
  classDef stRemoved fill:#ffebe9,stroke:#cf222e,color:#1f2328,stroke-width:2px
  classDef stUnchanged fill:#f6f8fa,stroke:#656d76,color:#1f2328,stroke-width:1px
Loading

AI-generated · Green: added · Yellow: modified · Red: removed · Gray: existing

Partial evidence: 20 file patches omitted; 0 truncated.

Diff evidence
  • F1: pinot-common/src/main/java/org/apache/pinot/common/utils/config/QueryOptionsUtils.java — before · after
  • F3: pinot-core/src/main/java/org/apache/pinot/core/operator/combine/GroupByCombineOperator.java — before · after
  • F4: pinot-core/src/main/java/org/apache/pinot/core/operator/query/FilteredGroupByOperator.java — before · after
  • F5: pinot-core/src/main/java/org/apache/pinot/core/operator/query/GroupByOperator.java — before · after
  • F6: pinot-core/src/main/java/org/apache/pinot/core/plan/maker/InstancePlanMakerImplV2.java — before · after
  • F8: pinot-core/src/main/java/org/apache/pinot/core/query/aggregation/function/DistinctCountULLAggregationFunction.java — before · after
  • F9: pinot-core/src/main/java/org/apache/pinot/core/query/aggregation/groupby/DefaultGroupByExecutor.java — before · after
  • Regenerate PR flow

Stacked on #19380 (off-heap group-by key tables and result holders for SSE). Review only the last commit (ffec9cdc12); the first two commits are #19380. Will rebase once #19380 merges.

What

First per-function off-heap aggregation state on top of #19380: DISTINCTCOUNTULL / DISTINCTCOUNTRAWULL group-by keeps each group's UltraLogLog register array (2^p bytes, ~4.1KB at the default p=12) in pooled direct memory instead of one heap object per group. A 200K-group query carries ~840MB of sketch heap per segment execution today; with groupByOffHeap enabled that moves off the heap entirely.

How

  • New seam AggregationFunction#createOffHeapGroupByResultHolder(initialCapacity, maxCapacity) (default null = unchanged). DefaultGroupByExecutor consults it first inside the existing off-heap gate and registers the returned holder on the ResourceTrackingGroupKeyGenerator, so the existing generator close sites release the memory. Later function conversions (t-digest, KLL, theta, distinct sets) reuse this hook.
  • OffHeapUltraLogLogGroupByResultHolder: append-only groupKey -> slotId indirection, so direct memory grows with the number of groups actually seen (like on-heap lazy allocation), never with the group-count upper bound. Slots live in 256KB pooled chunks (never moved/resized). The hash4j 0.30.0 register-update math (add/pack/unpack) is vendored verbatim (UltraLogLog is final and heap-only) and pinned byte-identical to the library by a differential test across p=3/8/12/18/19.
  • Two modes per holder: raw input values hash straight into off-heap registers; dictionary-encoded input (dict-id bitmap) and pre-serialized-ULL BYTES input (incl. star-tree pre-aggregated columns) go through a lazy on-heap ObjectGroupByResultHolder delegate — mode exclusivity is enforced with a hard check. Untouched groups read back as null; extraction materializes a fresh heap copy per group.
  • Plan-time p validation ([3, 26]): previously only UltraLogLog.create checked the user-supplied literal; the off-heap holder sizes slots as 1 << p, so an unchecked p could allocate up to 1GB per group (p=27..30) or corrupt neighbor slots via int-shift wrap (p>30). Minor behavior change: a query with an out-of-range p now fails at planning even when its filter matches zero rows.

Benchmark (BenchmarkOffHeapGroupByUllSSE, new)

2 segments x 2M rows, dict INT group column, raw LONG input, p=12, -prof gc -wi 4 -w 5 -i 10 -r 5, Xmx10g, M-series Mac:

benchmark groups on-heap ms/op off-heap ms/op latency alloc MB/op gc count/time per op
segmentGroupBy 10K 159.0 ± 3.7 140.0 ± 7.8 -12% 298 → 337 15/36ms → 19/28ms
segmentGroupBy 200K 348.0 ± 2.3 237.8 ± 18.3 -32% 1084 → 340 27/264ms → 11/29ms
query (full) 10K 180.1 ± 3.0 160.8 ± 0.8 -11% 694 → 855 31/89ms → 42/63ms
query (full) 200K 416.2 ± 3.9 360.7 ± 9.8 -13% 2312 → 2379 63/815ms → 53/100ms

Off-heap wins latency in every cell. At 200K groups the segment-phase allocation drops 69% and GC time drops ~9x. The small-tier alloc increase (+13-23%) is extraction: off-heap materializes a 4KB heap copy per group at hand-off where on-heap returns the live object; the combine phase merges heap ULLs in both arms (off-heap combine is a later milestone).

Testing

  • OffHeapUltraLogLogGroupByResultHolderTest: state-byte differential vs hash4j (incl. edge hashes, growth, chunk boundaries, wrapper-fallback arm via setViewSizeLimitBytes(0)), untouched-null / touch-empty semantics, delegate mode, INVALID_ID, close releases direct memory + idempotent, out-of-range p rejected.
  • OffHeapGroupByQueriesTest#testDistinctCountULL + #testDistinctCountULLSerializedBytesAndStarTree: on-heap vs off-heap differential over raw/dict/MV inputs, explicit p, RAWULL serialized output (byte-exact), filtered aggregation, order-by trim path, null handling, a serialized-ULL BYTES column, and a star-tree segment (asserted to actually serve the query) — every off-heap query asserts direct memory returns to baseline.
  • Full group-by battery (228 tests) green; spotless/checkstyle/license clean.

🤖 Generated with Claude Code

@codecov-commenter

codecov-commenter commented Aug 31, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 89.91477% with 142 lines in your changes missing coverage. Please review.
✅ Project coverage is 58.66%. Comparing base (910f9d5) to head (0635c91).

Files with missing lines Patch % Lines
.../function/DistinctCountULLAggregationFunction.java 43.39% 28 Missing and 2 partials ⚠️
...gation/groupby/offheap/OffHeapBytesGroupIdMap.java 91.01% 12 Missing and 11 partials ⚠️
...offheap/OffHeapUltraLogLogGroupByResultHolder.java 87.70% 10 Missing and 5 partials ⚠️
...pby/NoDictionarySingleColumnGroupKeyGenerator.java 92.14% 7 Missing and 4 partials ⚠️
...ry/aggregation/groupby/DefaultGroupByExecutor.java 74.35% 7 Missing and 3 partials ⚠️
...tion/groupby/offheap/OffHeapGroupByBufferPool.java 83.33% 6 Missing and 3 partials ⚠️
...regation/groupby/offheap/OffHeapIntGroupIdMap.java 94.96% 2 Missing and 5 partials ⚠️
...egation/groupby/offheap/OffHeapLongGroupIdMap.java 95.17% 2 Missing and 5 partials ⚠️
...pby/offheap/ResourceTrackingGroupKeyGenerator.java 79.31% 5 Missing and 1 partial ⚠️
...t/core/operator/query/FilteredGroupByOperator.java 42.85% 4 Missing ⚠️
... and 10 more

❗ There is a different number of reports uploaded between BASE (910f9d5) and HEAD (0635c91). Click for more details.

HEAD has 14 uploads less than BASE
Flag BASE (910f9d5) HEAD (0635c91)
unittests 2 1
java-25 6 3
temurin 6 3
lane-a 2 1
integration 4 2
integration1 2 0
lane-b 2 1
unittests2 1 0
Additional details and impacted files
@@             Coverage Diff              @@
##             master   #19402      +/-   ##
============================================
- Coverage     68.63%   58.66%   -9.98%     
+ Complexity     1486        1    -1485     
============================================
  Files          3526     2726     -800     
  Lines        230143   169817   -60326     
  Branches      36562    27770    -8792     
============================================
- Hits         157963    99618   -58345     
- Misses        59837    61951    +2114     
+ Partials      12343     8248    -4095     
Flag Coverage Δ
integration 0.00% <ø> (-100.00%) ⬇️
integration1 ?
integration2 0.00% <ø> (ø)
java-25 58.66% <89.91%> (-9.98%) ⬇️
lane-a 0.00% <ø> (-100.00%) ⬇️
lane-b 0.00% <ø> (ø)
temurin 58.66% <89.91%> (-9.98%) ⬇️
unittests 58.66% <89.91%> (-9.98%) ⬇️
unittests1 58.66% <89.91%> (+0.32%) ⬆️
unittests2 ?

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@xiangfu0
xiangfu0 force-pushed the xiangfu0/offheap-agg-ull branch 6 times, most recently from 8fb0398 to 407bde5 Compare September 5, 2026 09:15
@xiangfu0 xiangfu0 added enhancement Improvement to existing functionality query Related to query processing aggregation Related to aggregation functions and operations memory Related to memory usage or optimization performance Related to performance optimization configuration Config changes (addition/deletion/change in behavior) labels Sep 24, 2026
…nerator group counts

For primitive stored types with null handling enabled, the null group lives outside the primitive
key map but still takes the next dense group id, so getNumKeys() and getCurrentGroupKeyUpperBound()
under-counted by one once a null was seen. Since DefaultGroupByExecutor sizes result holders with
ensureCapacity(getCurrentGroupKeyUpperBound()), a segment whose group count exceeds the initial
holder capacity then wrote one slot past the holder array (ArrayIndexOutOfBoundsException) on the
default on-heap path. Object stored types are unaffected (their null key lives inside the map).

Adds NoDictionaryNullGroupCountRegressionTest reproducing the AIOOBE through
DefaultGroupByExecutor.process() with a shrunk maxInitialResultHolderCapacity for all four
primitive stored types, and pinning the counts and the null-group emission.
…g-gated, default off)

Adds an opt-in off-heap storage mode for the SSE per-segment group-by state, targeting
high-cardinality group-bys whose on-heap key maps and result holders drive GC pressure:

- New pinot-core package o.a.p.core.query.aggregation.groupby.offheap:
  - OffHeapIntGroupIdMap / OffHeapLongGroupIdMap: open-addressing key->dense-id tables over direct
    memory (8/16-byte slots, load factor 0.5, linear probing, out-of-band -1/0 key), drop-in
    replacements for IntGroupIdMap / Long2IntOpenHashMap semantics.
  - OffHeapBytesGroupIdMap: DuckDB-style two-part table for var-width keys — an 8-byte-entry
    directory (16-bit salt | 48-bit payload offset) over append-only 256KB payload chunks storing
    [hash][groupId][keyLength][key bytes]; the stored hash makes directory resize free of key reads.
  - OffHeapDouble/Long/IntGroupByResultHolder: fixed-width result holders over direct memory with
    semantics identical to the on-heap holders.
  - ResourceTrackingGroupKeyGenerator: wraps the generator and owns every off-heap resource, so the
    existing generator close() call sites release all direct memory (including the shared-generator
    filtered-aggregation case).
  - OffHeapGroupByBufferPool: bounded per-thread buffer reuse across queries (mirrors the on-heap
    thread-local map caching, with an explicit cap and visible accounting), default off.
  - All hot paths use absolute-indexed direct ByteBuffer views (wrapper fallback beyond 2GB).
- Wiring: server config pinot.server.query.executor.groupby.offheap (default false), query option
  groupByOffHeap, pool cap config groupby.offheap.pool.max.bytes.per.thread (default 0). Off-heap
  RawKeyHolder variants in DictionaryBasedGroupKeyGenerator (the ARRAY_BASED tier stays on-heap);
  off-heap modes in both NoDictionary generators; holder mirroring in DefaultGroupByExecutor.
  Grouping sets stay on-heap. Group ids remain dense ints; no AggregationFunction changes; no wire
  or storage format changes.
- Close-path hardening (also fixes pre-existing on-heap leak windows): exception guards in
  GroupByOperator/FilteredGroupByOperator/DefaultGroupByExecutor and widened finally coverage in
  the group-by combine operator. The streaming combine needs no extra plumbing: since apache#19066 each
  per-segment result is detached and its generator closed on the producing worker thread, which
  releases the off-heap state promptly as well.
- Tests: differential suites comparing off-heap vs on-heap row-for-row (OffHeapGroupByQueriesTest
  end-to-end battery with per-query direct-memory leak assertions, OffHeapGroupKeyGeneratorParityTest
  at generator level incl. null-group id bookkeeping), per-structure unit tests incl. forced
  wrapper-fallback runs, and buffer pool tests.
- Benchmarks (pinot-perf): BenchmarkOffHeapGroupBySSE / -LargeSSE / -HugeSSE and
  OffHeapGroupByMemoryFootprint. Measured: retained heap for the per-segment state drops to ~0
  (e.g. 8.7GB -> 4MB at 100M string groups, with 3.7x faster build); at ~1M+ groups off-heap is
  faster end-to-end (up to -50%) because it removes the GC pressure that dominates on-heap; at
  ~80K groups (cache-resident) there is a 10-22% latency premium, which the per-query opt-in
  avoids.
…efault off)

Adds the first per-function off-heap aggregation state on top of the off-heap
group-by SSE feature: DISTINCTCOUNTULL/DISTINCTCOUNTRAWULL keep each group's
UltraLogLog register array (2^p bytes, ~4.1KB at the default p=12) in pooled
direct-memory chunks instead of one heap object per group.

- New optional AggregationFunction#createOffHeapGroupByResultHolder seam
  (default null = unchanged); DefaultGroupByExecutor consults it first inside
  the existing off-heap gate and registers the holder on the resource tracker.
- OffHeapUltraLogLogGroupByResultHolder: append-only slotId indirection so
  direct memory grows with actual groups (not the group-count upper bound),
  256KB pooled chunks, hash4j 0.30.0 register math vendored verbatim (pinned
  byte-identical by test), lazy on-heap delegate for the dictionary and
  pre-serialized-BYTES modes (mode exclusivity enforced hard), untouched
  groups read back as null, snapshot materialization at extraction.
- Validates the user-supplied p literal in [3, 26] at plan time (previously
  only UltraLogLog.create checked it; the off-heap holder sizes slots as
  1 << p, so an unchecked p could over-allocate or corrupt neighbor slots).
  A bad p now fails planning even when a filter matches zero rows.
- Tests: differential holder test vs hash4j across p=3/8/12/18/19 including
  the buffer-wrapper fallback arm; e2e battery in OffHeapGroupByQueriesTest
  (raw/dict/MV/explicit-p/RAWULL/filtered/order-by-trim/null-handling, plus a
  star-tree segment proving the pre-aggregated BYTES path in both modes) with
  per-query direct-memory leak asserts.
- Benchmark BenchmarkOffHeapGroupByUllSSE (10K/200K groups x flag): off-heap
  segment phase -12%/-32% latency, full query -11%/-13%, with the per-group
  sketch heap (~42MB/840MB per segment execution) moved off the heap.
@xiangfu0
xiangfu0 force-pushed the xiangfu0/offheap-agg-ull branch from 407bde5 to 0635c91 Compare October 9, 2026 20:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

aggregation Related to aggregation functions and operations configuration Config changes (addition/deletion/change in behavior) enhancement Improvement to existing functionality memory Related to memory usage or optimization performance Related to performance optimization query Related to query processing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants