Repository navigation
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #19402 +/- ##
============================================
- Coverage 68.63% 58.66% -9.98%
+ Complexity 1486 1 -1485
============================================
Files 3526 2726 -800
Lines 230143 169817 -60326
Branches 36562 27770 -8792
============================================
- Hits 157963 99618 -58345
- Misses 59837 61951 +2114
+ Partials 12343 8248 -4095
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
xiangfu0
force-pushed
the
xiangfu0/offheap-agg-ull
branch
6 times, most recently
from
September 5, 2026 09:15
8fb0398 to
407bde5
Compare
…nerator group counts For primitive stored types with null handling enabled, the null group lives outside the primitive key map but still takes the next dense group id, so getNumKeys() and getCurrentGroupKeyUpperBound() under-counted by one once a null was seen. Since DefaultGroupByExecutor sizes result holders with ensureCapacity(getCurrentGroupKeyUpperBound()), a segment whose group count exceeds the initial holder capacity then wrote one slot past the holder array (ArrayIndexOutOfBoundsException) on the default on-heap path. Object stored types are unaffected (their null key lives inside the map). Adds NoDictionaryNullGroupCountRegressionTest reproducing the AIOOBE through DefaultGroupByExecutor.process() with a shrunk maxInitialResultHolderCapacity for all four primitive stored types, and pinning the counts and the null-group emission.
…g-gated, default off)
Adds an opt-in off-heap storage mode for the SSE per-segment group-by state, targeting
high-cardinality group-bys whose on-heap key maps and result holders drive GC pressure:
- New pinot-core package o.a.p.core.query.aggregation.groupby.offheap:
- OffHeapIntGroupIdMap / OffHeapLongGroupIdMap: open-addressing key->dense-id tables over direct
memory (8/16-byte slots, load factor 0.5, linear probing, out-of-band -1/0 key), drop-in
replacements for IntGroupIdMap / Long2IntOpenHashMap semantics.
- OffHeapBytesGroupIdMap: DuckDB-style two-part table for var-width keys — an 8-byte-entry
directory (16-bit salt | 48-bit payload offset) over append-only 256KB payload chunks storing
[hash][groupId][keyLength][key bytes]; the stored hash makes directory resize free of key reads.
- OffHeapDouble/Long/IntGroupByResultHolder: fixed-width result holders over direct memory with
semantics identical to the on-heap holders.
- ResourceTrackingGroupKeyGenerator: wraps the generator and owns every off-heap resource, so the
existing generator close() call sites release all direct memory (including the shared-generator
filtered-aggregation case).
- OffHeapGroupByBufferPool: bounded per-thread buffer reuse across queries (mirrors the on-heap
thread-local map caching, with an explicit cap and visible accounting), default off.
- All hot paths use absolute-indexed direct ByteBuffer views (wrapper fallback beyond 2GB).
- Wiring: server config pinot.server.query.executor.groupby.offheap (default false), query option
groupByOffHeap, pool cap config groupby.offheap.pool.max.bytes.per.thread (default 0). Off-heap
RawKeyHolder variants in DictionaryBasedGroupKeyGenerator (the ARRAY_BASED tier stays on-heap);
off-heap modes in both NoDictionary generators; holder mirroring in DefaultGroupByExecutor.
Grouping sets stay on-heap. Group ids remain dense ints; no AggregationFunction changes; no wire
or storage format changes.
- Close-path hardening (also fixes pre-existing on-heap leak windows): exception guards in
GroupByOperator/FilteredGroupByOperator/DefaultGroupByExecutor and widened finally coverage in
the group-by combine operator. The streaming combine needs no extra plumbing: since apache#19066 each
per-segment result is detached and its generator closed on the producing worker thread, which
releases the off-heap state promptly as well.
- Tests: differential suites comparing off-heap vs on-heap row-for-row (OffHeapGroupByQueriesTest
end-to-end battery with per-query direct-memory leak assertions, OffHeapGroupKeyGeneratorParityTest
at generator level incl. null-group id bookkeeping), per-structure unit tests incl. forced
wrapper-fallback runs, and buffer pool tests.
- Benchmarks (pinot-perf): BenchmarkOffHeapGroupBySSE / -LargeSSE / -HugeSSE and
OffHeapGroupByMemoryFootprint. Measured: retained heap for the per-segment state drops to ~0
(e.g. 8.7GB -> 4MB at 100M string groups, with 3.7x faster build); at ~1M+ groups off-heap is
faster end-to-end (up to -50%) because it removes the GC pressure that dominates on-heap; at
~80K groups (cache-resident) there is a 10-22% latency premium, which the per-query opt-in
avoids.
…efault off) Adds the first per-function off-heap aggregation state on top of the off-heap group-by SSE feature: DISTINCTCOUNTULL/DISTINCTCOUNTRAWULL keep each group's UltraLogLog register array (2^p bytes, ~4.1KB at the default p=12) in pooled direct-memory chunks instead of one heap object per group. - New optional AggregationFunction#createOffHeapGroupByResultHolder seam (default null = unchanged); DefaultGroupByExecutor consults it first inside the existing off-heap gate and registers the holder on the resource tracker. - OffHeapUltraLogLogGroupByResultHolder: append-only slotId indirection so direct memory grows with actual groups (not the group-count upper bound), 256KB pooled chunks, hash4j 0.30.0 register math vendored verbatim (pinned byte-identical by test), lazy on-heap delegate for the dictionary and pre-serialized-BYTES modes (mode exclusivity enforced hard), untouched groups read back as null, snapshot materialization at extraction. - Validates the user-supplied p literal in [3, 26] at plan time (previously only UltraLogLog.create checked it; the off-heap holder sizes slots as 1 << p, so an unchecked p could over-allocate or corrupt neighbor slots). A bad p now fails planning even when a filter matches zero rows. - Tests: differential holder test vs hash4j across p=3/8/12/18/19 including the buffer-wrapper fallback arm; e2e battery in OffHeapGroupByQueriesTest (raw/dict/MV/explicit-p/RAWULL/filtered/order-by-trim/null-handling, plus a star-tree segment proving the pre-aggregated BYTES path in both modes) with per-query direct-memory leak asserts. - Benchmark BenchmarkOffHeapGroupByUllSSE (10K/200K groups x flag): off-heap segment phase -12%/-32% latency, full query -11%/-13%, with the per-group sketch heap (~42MB/840MB per segment execution) moved off the heap.
xiangfu0
force-pushed
the
xiangfu0/offheap-agg-ull
branch
from
October 9, 2026 20:25
407bde5 to
0635c91
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR flow
Off-heap group-by state for DISTINCTCOUNTULL: from config to holder creation, update, and release.
AI-generated · Green: added · Yellow: modified · Red: removed · Gray: existing
Partial evidence: 20 file patches omitted; 0 truncated.
Diff evidence
What
First per-function off-heap aggregation state on top of #19380:
DISTINCTCOUNTULL/DISTINCTCOUNTRAWULLgroup-by keeps each group's UltraLogLog register array (2^pbytes, ~4.1KB at the default p=12) in pooled direct memory instead of one heap object per group. A 200K-group query carries ~840MB of sketch heap per segment execution today; withgroupByOffHeapenabled that moves off the heap entirely.How
AggregationFunction#createOffHeapGroupByResultHolder(initialCapacity, maxCapacity)(defaultnull= unchanged).DefaultGroupByExecutorconsults it first inside the existing off-heap gate and registers the returned holder on theResourceTrackingGroupKeyGenerator, so the existing generator close sites release the memory. Later function conversions (t-digest, KLL, theta, distinct sets) reuse this hook.OffHeapUltraLogLogGroupByResultHolder: append-onlygroupKey -> slotIdindirection, so direct memory grows with the number of groups actually seen (like on-heap lazy allocation), never with the group-count upper bound. Slots live in 256KB pooled chunks (never moved/resized). The hash4j 0.30.0 register-update math (add/pack/unpack) is vendored verbatim (UltraLogLogis final and heap-only) and pinned byte-identical to the library by a differential test across p=3/8/12/18/19.ObjectGroupByResultHolderdelegate — mode exclusivity is enforced with a hard check. Untouched groups read back asnull; extraction materializes a fresh heap copy per group.pvalidation ([3, 26]): previously onlyUltraLogLog.createchecked the user-supplied literal; the off-heap holder sizes slots as1 << p, so an unchecked p could allocate up to 1GB per group (p=27..30) or corrupt neighbor slots via int-shift wrap (p>30). Minor behavior change: a query with an out-of-range p now fails at planning even when its filter matches zero rows.Benchmark (
BenchmarkOffHeapGroupByUllSSE, new)2 segments x 2M rows, dict INT group column, raw LONG input, p=12,
-prof gc -wi 4 -w 5 -i 10 -r 5, Xmx10g, M-series Mac:Off-heap wins latency in every cell. At 200K groups the segment-phase allocation drops 69% and GC time drops ~9x. The small-tier alloc increase (+13-23%) is extraction: off-heap materializes a 4KB heap copy per group at hand-off where on-heap returns the live object; the combine phase merges heap ULLs in both arms (off-heap combine is a later milestone).
Testing
OffHeapUltraLogLogGroupByResultHolderTest: state-byte differential vs hash4j (incl. edge hashes, growth, chunk boundaries, wrapper-fallback arm viasetViewSizeLimitBytes(0)), untouched-null / touch-empty semantics, delegate mode, INVALID_ID, close releases direct memory + idempotent, out-of-range p rejected.OffHeapGroupByQueriesTest#testDistinctCountULL+#testDistinctCountULLSerializedBytesAndStarTree: on-heap vs off-heap differential over raw/dict/MV inputs, explicit p, RAWULL serialized output (byte-exact), filtered aggregation, order-by trim path, null handling, a serialized-ULL BYTES column, and a star-tree segment (asserted to actually serve the query) — every off-heap query asserts direct memory returns to baseline.🤖 Generated with Claude Code