Skip to content

Benchmark harness for the streaming selection ORDER BY combine - #1

Draft
rohityadav1993 wants to merge 5 commits into
oss/pr1-streaming-selection-combinefrom
rohity/pr1-arm4-memory-crossover-bench
Draft

rohityadav1993 wants to merge 5 commits into
oss/pr1-streaming-selection-combinefrom
rohity/pr1-arm4-memory-crossover-bench

Conversation

@rohityadav1993

@rohityadav1993 rohityadav1993 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

This PR is the benchmark harness and the measurement report for #19120. apache#19120 adds the operator. This PR adds no operator code. It stacks on oss/pr1-streaming-selection-combine and changes pinot-perf only.

TL;DR

ON = the streaming combine from apache#19120. OFF = the existing MinMax combine (10 threads). 100 segments × 50,000 rows, ASC.

Question Result
Memory (minimum heap) At LIMIT ≥ 100,000, ON needs 2x to 14x less heap. At LIMIT 1,000,000, OFF needs about 930 MB; ON needs less than 64 MB (DISJOINT) to 237 MB (FULL). ON stays flat as LIMIT grows.
Wall time At LIMIT ≥ 100,000, ON is faster: 0.07 to 0.75 of OFF (single-shot timing). One 4 GB average-time run shows ON up to 1.15x slower at LIMIT 100,000. At DISJOINT and PARTIAL, ON is also faster at most smaller LIMIT values.
Adverse case FULL overlap and LIMIT < 100,000: ON is 1.8x to 7x slower (up to 10x with a one-column ORDER BY). ON reads all tied segments on one thread; OFF uses 10. At FULL / LIMIT 10,000, ON needs about 1.27x more heap.
CPU ON uses less CPU in 17 of 18 cases.
Fed back to apache#19120 The benchmark found a tie issue. The fix (in 6df6b87de5) took one case from 77.8x slower to 0.82x.
Limits One host, combine only (no broker or MSE), immutable segments, ASC only. Heap figures are a provisioning bound, not a live-set measurement.

Details, methods and raw sources follow.

Which PR1 code was measured. The rerun with the fix used operator code identical to 6df6b87de5. git diff 4f79f49f78 6df6b87de5 is empty on the combine, the leaf, SelectionPlanNode, CombinePlanNode, InstancePlanMakerImplV2 and QueryContext. Four suites ran on code before the fix. The FULL K=1 rerun ran on 642af405b8 (deferral only, no frontier fix). The rerun with the fix ran on 4f79f49f78.

The earlier figures still apply. Tie deferral is off under TS_VAL. Frontier pruning can only keep closed a segment that would have opened before. At FULL under TS_VAL it cannot fire, and the FULL rerun matched the earlier figures within noise. At DISJOINT and PARTIAL the earlier figures can only be pessimistic for ON. OFF is not changed by the fix.

The four later commits cannot change a result:

Commit Why it cannot change a result
1923549146 Explain output and server-config defaults. The harness sets block size 10,000 directly. Other changes are comments and tests.
8271cf9b0e Changes pruning only with null handling on. The harness leaves it off.
3b92607347 Test only.
ba2b193b32 Applies only to MutableSegment. All fixture segments are immutable. A rerun is not necessary.

What this adds

All paths are under pinot-perf. The Java files are in src/main/java/org/apache/pinot/perf/sortedmerge/. The scripts are in src/main/scripts/.

File Role
SortedMergeFixture Builds 100 OFFLINE segments of 50,000 rows. Axes: overlap and key cardinality. Caches fixtures. Measures overlap depth (G3).
SortedMergeDriver Builds the QueryContext and the plan nodes. Runs CombinePlanNode and drains the blocks. Asserts the operator class (G1). Holds the digest helpers and the OrderBy shapes.
BenchmarkSortedMergeMemoryCrossover The JMH benchmark (mergeTopK). Reports wall time, CPU through @AuxCounters, and allocation through -prof gc.
SortedMergeCpuMeter Opt-in CPU meter (-Darm4.cpuTime=true).
SortedMergeMemoryProbe CLI. gate is the correctness gate (G2). run does one merge per JVM for the heap search.
sorted-merge-bisect-xmx.sh Runs the heap search: it bisects -Xmx over Probe run.
sorted-merge-bisect-xmx-selftest.sh Self-test of the heap-search script. It stubs java. 14 cases, all pass.
3 test classes SortedMergeGateTest (15), SortedMergeDriverTest (5), SortedMergeFixtureTest (3). Total 23.

The rest of this description is the measurement report.

How to run

Run all commands from the repository root. Use the shaded benchmarks.jar. The main() in the benchmark class ignores its arguments.

./mvnw -pl pinot-perf -am install -DskipTests -Ddevelocity.cache.local.enabled=false

# Wall time, LIMIT 10 to 10,000. Add -p _keyCardinality=1,100,10000 for the K sweep.
java -jar pinot-perf/target/benchmarks.jar BenchmarkSortedMergeMemoryCrossover \
  -p _limit=10,100,1000,10000 -bm avgt -f 3 -wi 3 -i 5 -r 2s -w 2s \
  -jvmArgs "-Xmx16g -Darm4.cpuTime=true"
# Wall time, LIMIT 100,000 and 1,000,000
java -jar pinot-perf/target/benchmarks.jar BenchmarkSortedMergeMemoryCrossover \
  -p _limit=100000,1000000 -bm ss -f 10 -wi 5 -i 1 -jvmArgs "-Xmx16g -Darm4.cpuTime=true"
# Allocation (separate run)
java -jar pinot-perf/target/benchmarks.jar BenchmarkSortedMergeMemoryCrossover \
  -bm avgt -f 1 -wi 1 -i 3 -r 5s -w 5s -prof gc -jvmArgs "-Xmx16g"

./mvnw -pl pinot-perf dependency:build-classpath -Dmdep.outputFile=/tmp/arm4-cp.txt \
  -Ddevelocity.cache.local.enabled=false
# Correctness gate (separate step)
java -cp pinot-perf/target/classes:$(cat /tmp/arm4-cp.txt) \
  org.apache.pinot.perf.sortedmerge.SortedMergeMemoryProbe gate 100 50000 10000 10 TS_VAL 1
# Heap search, one K for each run
KEY_CARDINALITY=1 bash pinot-perf/src/main/scripts/sorted-merge-bisect-xmx.sh /tmp/arm4-cp.txt > bisect.csv
# Tests
bash pinot-perf/src/main/scripts/sorted-merge-bisect-xmx-selftest.sh
./mvnw -pl pinot-perf -am test -Dtest='SortedMerge*Test' -Dsurefire.failIfNoSpecifiedTests=false -Ddevelocity.cache.local.enabled=false

The flag -Ddevelocity.cache.local.enabled=false stops the build cache from reporting success when no tests ran. Run the heap search from the repository root, because it uses relative paths.

1. Objective

apache#19120 adds StreamingSelectionOrderByCombineOperator. Query option sortedSelectionMergeMode=ON selects it. This report calls it the streaming combine (ON). The existing MinMaxValueBasedSelectionOrderByCombineOperator is the MinMax combine (OFF). The target is ORDER BY plus LIMIT on a leading column that is sorted in each segment.

The benchmark answers two questions. Memory: does ON need less heap than OFF as LIMIT grows? Crossover: at which LIMIT and overlap does ON stop costing more wall time and CPU? Memory and latency need separate instruments, because allocation and elapsed time do not show live heap.

2. Definitions

Term Meaning
Overlap depth The number of segments whose tsCol range contains the global midpoint. DISJOINT=1, PARTIAL≈10, FULL=100. Depth is not measured at the LIMIT boundary, so it is exact only at FULL.
OFF, ON OFF: MinMax combine, work on 10 pool threads. ON: streaming combine, merge on the calling thread.
TS_VAL ORDER BY tsCol, valCol. valCol is unique, so the order is total.
TS_ONLY ORDER BY tsCol. Ties can cross the LIMIT boundary.
K Key cardinality. The number of consecutive rows in a segment with the same tsCol.
The fix The tie-deferral fix in apache#19120 (squashed into 6df6b87de5). See section 8.
Wall ratio, CPU difference ON ÷ OFF wall time (below 1: ON faster). ON minus OFF CPU in ms. A/A control: OFF cases rerun later, to measure noise.
Segments opened How many of the 100 segments the combine opens (segmentsProcessed log line).
Minimum surviving heap The smallest -Xmx at which a run completes. (low, high]: failed at low, survived at high. An upper bound that includes GC headroom.

3. Setup

Pipeline: SortedMergeFixture feeds SortedMergeDriver. The driver feeds JMH, SortedMergeCpuMeter and Probe run. Probe run feeds the heap search. Probe gate compares both arms in a separate step.

Item Value
Segments One OFFLINE table, no broker. 100 for each fixture, 50,000 rows each, 5,000,000 rows. About 1 MB for each segment.
Fixtures 9: three overlaps × K = 1, 100, 10,000. LIMIT 10 to 1,000,000 in JMH, 1,000 to 1,000,000 in the heap search.
Columns tsCol LONG (sorted, dictionary). valCol INT (unique). payloadCol STRING (raw, 54 to 59 characters).
Execution maxExecutionThreads=10. sortedSelectionMergeBlockSize=10,000. The block is min(limit + offset, 10,000) rows under TS_ONLY.
Host One host (rohity-pinot-oss1), 96 cores, 377 GB RAM, Temurin 25.0.3, G1. Hot page cache.
Heap 16 GB in JMH. 4 GB in one rerun. The heap search uses 64 MB to 4,096 MB.
Volume 6 suites. 144 heap measurements in 1,652 JVM launches.

4. Results

Memory: OFF ÷ ON minimum surviving heap

TS_VAL, K=1. Each figure is a lower bound: the heap where OFF failed divided by the heap where ON survived. "≥" means ON survived at the 64 MB floor, so the true ratio is not resolved.

Overlap LIMIT 1,000 LIMIT 10,000 LIMIT 100,000 LIMIT 1,000,000
DISJOINT both ≤64 MB both ≤64 MB ≥ 2.0x ≥ 14.3x
PARTIAL both ≤64 MB both ≤64 MB ≥ 2.5x ≥ 11.6x
FULL both ≤64 MB ON needs ≥ 1.27x more ≥ 3.3x ≥ 3.9x
  • OFF grows with LIMIT. ON does not. At LIMIT 1,000,000, OFF needs about 914 to 946 MB at every overlap. Section 7 has the brackets.
  • K=10,000. The ratios agree within about 15%.
  • FULL/LIMIT 10,000 is the one adverse case. Under TS_VAL, all 100 segments with the same minimum must open.
  • Quote these figures as a provisioning bound. They are not live-set measurements. Churn adds GC headroom. At 4 GB and LIMIT ≥ 100,000, OFF allocates about 4.5x more than ON at FULL/100k and DISJOINT/1M. At FULL/1M it allocates about 2.7x more (1.62 GB against 606 MB).

Wall clock: ON ÷ OFF

Below 1 means ON is faster. Bold means ON is slower beyond noise.

TS_VAL, K=1:

Overlap 10 100 1,000 10,000 100,000 1,000,000
DISJOINT 0.71 0.77 0.74 0.86 0.50 0.32
PARTIAL 0.77 0.73 0.73 1.45 0.75 0.33
FULL 1.77 2.74 4.58 7.02 0.52 0.11

TS_ONLY, FULL, with the fix. A blank cell has no measurement. The 0.64, 0.82 and 0.90 entries have error bars that include 1. Read them as parity.

K 10 100 1,000 10,000 100,000 1,000,000
1 0.64 2.87 4.58 10.1 0.52 0.074
100 1.76 8.18
10,000 0.63 0.82 0.90
  • Crossover. At DISJOINT, ON is faster at every LIMIT in SingleShotTime (0.50 at 100k, 0.32 at 1M). The 4 GB avgt run (raw/jmh-two-column-grid/r4-heap4g.json) shows ON 1.10x slower at DISJOINT/100k and 1.15x slower at PARTIAL/100k. At PARTIAL, ON is slower only at 10,000 in the 16 GB runs. At FULL, ON is faster only from 100,000 (up to 9x under TS_VAL, 13x under TS_ONLY).
  • Why ON is slower. It reads every tied segment on one thread. OFF uses 10.
  • Few tied segments needed. ON is at parity or faster, except K=100/LIMIT 1,000 (1.76). K=10,000/LIMIT 10,000 went from 77.8x to 0.82x.
  • K=1 at LIMIT ≥ 100 stays slow with the fix. The output needs key 0 from every segment, so all 100 open.

CPU and GC

  • ON uses less CPU in 17 of 18 TS_VAL cases (12 AverageTime, 6 SingleShotTime). The exception is FULL/LIMIT 10,000, at +25.6 ms. At LIMIT ≥ 100,000 the CPU ratio is 0.09 to 0.20.
  • At 4 GB, GC takes up to 10.73% of wall time for OFF and 0.57% for ON (section 7).
  • Noise. A/A wall: 1.89% median, 10.86% worst. Effects under about 11% are not distinguishable.

5. Harness

Instruments

  • Wall time: JMH 1.37, @Threads(1). Allocation: -prof gc (B/op), in a separate run.
  • CPU: SortedMergeCpuMeter sums ThreadMXBean CPU over the calling thread and its pool threads. GC and JIT threads are excluded, so OFF is understated.
  • CPU caveat. Both arms build the 100 leaf operators on the pool inside the timed region. This constant dilutes the CPU ratio at small LIMIT. Use the CPU difference there.
  • Both arms use CombinePlanNode with a ResultsBlockStreamer. The mode is explicitly ON or OFF, never AUTO, and it also selects the leaf. The streamer is a no-op. The executor is a fixed pool of 10 daemon threads.

JMH settings

Item Code default Used on the command line
Mode AverageTime, ms LIMIT 10 to 10,000: -bm avgt. LIMIT 100,000 and 1,000,000: -bm ss (SingleShotTime).
Forks 1 3 for avgt. 10 for ss. 1 for the allocation run. The rerun with the fix used 1 fork and 3 iterations, except K=10,000/LIMIT 100,000 (ss, 10 forks).
Warmup 1 × 5 s avgt: 3 × 2 s (-wi 3 -w 2s). ss: 5 (-wi 5).
Measurement 3 × 5 s avgt: 5 × 2 s (-i 5 -r 2s). ss: 1 (-i 1).

The allocation run uses -f 1 -wi 1 -i 3 -r 5s -w 5s -prof gc. The 4 GB run uses -Xms4g -Xmx4g -XX:+UseG1GC -prof gc. The default grid has 72 configs (the class Javadoc says 36). OFF is bimodal, because its segment count races across 10 threads.

Guards

Guard What it does
G1 Asserts the operator class (Streaming for ON, MinMax for OFF). Runs in JMH setup, every run and the gate.
G2 The correctness gate below. A separate Probe gate step. Not in JMH.
G3 Measures overlap depth at the global midpoint, from segment metadata.
Fixture checks Marker, row count and lock. The harness never reuses a corrupt or partial fixture.

The gate covers 3 overlaps × 6 LIMIT values, ON and OFF. Row counts must be equal. It also requires:

  • TS_VAL: equal full-row digest. At LIMIT ≤ 10,000, also equal row multiset and tsCol sorted.
  • TS_ONLY: equal digest of tsCol values only. At LIMIT ≤ 10,000, also equal tsCol multiset and tsCol sorted.

Under TS_ONLY the arms can return different tied rows. The gate proves nothing about valCol or payload there. Above LIMIT 10,000 it proves nothing about order. SortedMergeGateTest has negative controls.

Fixture cache and locking

  • Cache and trust: an in-JVM cache, and /tmp/pinot-arm4-sorted-merge/<key> on disk. The harness reuses a fixture only with a .arm4-complete marker, after it reloads and recounts every segment.
  • Lock: a per-key <key>.lock file lock covers validate-or-build, because each heap-search attempt is a new JVM. A build takes 15 to 25 s. The run drivers also used flock /tmp/arm4-bench.lock, which is not in the committed code.

Heap search

flowchart TD
  A["Probe the floor (64 MB)"] --> B{"Survives?"}
  B -->|yes| R1["AT_OR_BELOW_LOW"]
  B -->|no| C["Probe the ceiling (4,096 MB)"]
  C --> D{"Survives?"}
  D -->|no| R2["ABOVE_HIGH"]
  D -->|yes| E["Pick the midpoint of (low, high]"]
  E --> F["Run N consecutive probes<br/>(N=3 for OFF, N=2 for ON)"]
  F --> G{"All N survive?"}
  G -->|yes| H["high = midpoint"]
  G -->|no| I["low = midpoint"]
  H --> J{"high - low <= max(16 MB, 5% of high)?"}
  I --> J
  J -->|no| E
  J -->|yes| K["Output bracket (low, high]"]
Loading
  • Pass criterion. A heap survives only if all N consecutive runs survive. OFF uses N=3, because its retention depends on thread scheduling. This raises the OFF minimum, so it is conservative against the PR.
  • JVM flags. -Xms = -Xmx, G1, -XX:+ExitOnOutOfMemoryError. Fixtures are prebuilt at -Xmx8g.

Exit 3 (or 137/143 before the timeout) is an OOM. Exit 124, or 137 after the 900 s timeout, is ERROR_TIMEOUT. No [arm4] built or reused line is ERROR_FIXTURE_NOT_READY. Any other exit is ERROR_PROBE_FAILED.

An ERROR_* outcome never steers the search, so a harness bug is never recorded as "needs more heap". The harness abandons that case. Every attempt goes to an audit CSV. Each case has 2 replicates. The script hardcodes TS_VAL.

Suites

Suite Instrument Code Supplies
Two-column grid JMH Before the fix TS_VAL wall, CPU, 4 GB heap, A/A
Query shape JMH Before the fix TS_ONLY against TS_VAL
Key cardinality JMH Before the fix K sweep, OFF baselines, before-fix ON
Heap search 144 measurements Before the fix Every heap figure
FULL K=1 rerun JMH With deferral only (642af405b8) Cases the fix cannot change
Rerun with the fix JMH, ON only 4f79f49f78 ON figures for cases the fix changes

6. Fixtures

All fixtures come from SortedMergeFixture. The tsCol of each segment starts at base. In each segment tsCol ascends, so the leaf scans forward.

Overlap base of segment i Depth
DISJOINT i × 50,000 1
PARTIAL i × 5,000 ≈10
FULL 0 100

tsCol = (base + j) / K. Each key covers K rows. At FULL, every segment starts with the same K rows of key 0. This is the tie in section 8.

Measured depth is 1 / 10 / 100 at every K, except PARTIAL at K=10,000 (11).

7. Details

Heap search results (TS_VAL, MB)

Each cell is one heap-search result. ≤64 means survival at the floor. At LIMIT ≤ 10,000 every cell is ≤64 except FULL/10,000. If the two replicates disagree, the bracket spans both.

Overlap LIMIT OFF K=1 ON K=1 OFF K=10,000 ON K=10,000
DISJOINT 100,000 (127,142] ≤64 (111,142] ≤64
DISJOINT 1,000,000 (914,946] ≤64 (883,914] ≤64
PARTIAL 100,000 (158,190] ≤64 (142,174] ≤64
PARTIAL 1,000,000 (914,946] (64,79] (788,820] ≤64
FULL 10,000 (158,174] (221,237] (142,158] (190,205]
FULL 100,000 (788,820] (221,237] (662,725] (190,205]
FULL 1,000,000 (914,946] (221,237] (788,820] (190,205]
  • Flat in LIMIT and K. At FULL, ON needs at most about 205 to 237 MB at every LIMIT from 10,000 to 1,000,000. At FULL all 100 segments open, each with a live 10,000-row block. OFF holds a small capped queue for each.
  • Run quality. The 1,652 launches were each classified OK or OOM. None timed out. Survival was monotone in heap size in all 144 measurements. Replicates disagreed in 5 of 72 cases. All 5 are OFF at LIMIT 100,000, and differ by one adjacent bracket.

CPU and GC detail

  • CPU difference, TS_VAL, K=1, LIMIT ≤ 10,000: -0.09 to -7.8 ms at DISJOINT and PARTIAL. At FULL: -2.0, -1.9, -4.1, then +25.6 ms at 10,000.
  • 4 GB allocation at LIMIT ≥ 100,000: OFF 247 MB to 1.62 GB for each operation. ON 34 to 606 MB. At LIMIT 1,000: OFF 1.8 MB, ON 0.98 MB.
  • FULL/10,000 ON took 135.7, 123.6 and 96.3 ms in three runs. Do not quote it alone.

Shape, K and precision

  • TS_VAL does not get worse for ON as K rises. FULL/LIMIT 10,000: 127.9, 120.7 and 105.5 ms at K=1, 100, 10,000. TS_ONLY at DISJOINT and PARTIAL is within noise of TS_VAL.
  • Noise. Worst scoreError/score at LIMIT ≤ 10,000: 11.8%. CV at LIMIT ≥ 100,000: 3.3% to 49.2%. A/A CPU: 2.65% median, 8.74% worst. The A/A control covers 12 OFF cases at LIMIT ≤ 10,000, 90 minutes later.
  • Error. Ratio error is about sqrt(errON² + errOFF²). The rerun with the fix had up to 59% wall error. Allocation is near deterministic.

8. Finding fed back to apache#19120

How it surfaced

The heap search found the one case where ON needs more heap than OFF: FULL, LIMIT 10,000, ON (221,237] against OFF (158,174]. At FULL every segment has minimum tsCol 0. The tie defeats the lazy opening of segments. All 100 open, and each pins a 10,000-row block.

The key-cardinality suite showed the cost under TS_ONLY: at FULL, K=10,000, LIMIT 10,000, OFF opened 10 segments and ON opened 100 (78x slower).

The issue and the fix

ON visits segments in order of their minimum. It keeps a segment closed only while that minimum sorts strictly after the merge frontier. A segment whose minimum ties the frontier always opened. OFF already skips such a segment under TS_ONLY.

A second flaw hid behind the first. The frontier is the smaller of two heads (the leader's and the heap top's), but the check used the larger. This kept K=100 at 100 segments.

The fix has two parts. It defers tied segments under a single-column ORDER BY only. With two or more columns, a tie on the first column can hide an earlier row on the second. It also takes the frontier as the smaller of the two heads, for all shapes.

A deferred segment is postponed, never skipped. It opens when no other segment can supply a row. Its remaining rows tie, so they are interchangeable under a single-column sort key.

Effect

TS_ONLY, FULL, ON. "Before" is from the key-cardinality suite (K=1/LIMIT 10: query-shape suite). The wall ratio with the fix is in section 4. Allocation is in MB per operation.

K LIMIT Segments before / with fix Wall ratio before ON/OFF allocation before / with fix
1 10 100 / 10 3.654 2.61 / 0.98
100 1,000 100 / 10 15.650 4.06 / 0.49
100 10,000 100 / 100 8.375 1.11 / 1.11
10,000 1,000 100 / 1 25.901 16.15 / 0.52
10,000 10,000 100 / 1 77.815 11.20 / 0.28
10,000 100,000 100 / 10 3.851 not measured
  • Largest change. K=10,000/LIMIT 10,000: segments 100 to 1, wall 121.6 to 1.29 ms (94x), ON allocation 285.06 to 7.14 MB (OFF 25.46).
  • K=100/LIMIT 10,000 is unchanged. The output is the 100 key-0 rows of each segment, so all 100 must open. OFF also opens all 100.
  • OFF baselines predate a rebase onto newer master. They matched within noise in the FULL K=1 rerun.

Under TS_VAL, ON still opens all tied segments. The adverse case in section 4 remains.

9. Limits and open items

Limits

Limit Detail
Host and scope One host, JDK 25, G1. Combine only: no broker, no MSE. The streamer is a no-op.
Segments Immutable OFFLINE segments only. No DESC.
Heap figures A provisioning bound, not retention. 80 of 144 heap cases are at the 64 MB floor, so "at least 14x" is a lower bound.
Heap search coverage None with the fix. None under TS_ONLY. The expected 1-segment gain at K=10,000 is unmeasured.
Unswept Block size (10,000 only). The AUTO ratio (0.8). maxExecutionThreads. Segment count and rows. Payload width. Concurrent queries.
Cache Hot page cache. The two-phase payload fetch gain applies to a hot cache only.

Open items

  • Heap search with the fix. Add an ORDER_BY knob. Run a before-and-after pair at FULL under TS_ONLY. Rerun with LOW_MB=16.
  • Frontier pruning at DISJOINT and PARTIAL under TS_VAL. It was not rerun. It can only help ON.
  • depth(L). The fixture measures depth at the midpoint, exact only at FULL. It does not compute depth at the LIMIT boundary.
  • Open question. Leaf block tightening shows in OFF allocation at FULL but not in ON. The cause is unexplained.

@rohityadav1993
rohityadav1993 force-pushed the oss/pr1-streaming-selection-combine branch from f3f2db3 to d605e80 Compare September 26, 2026 18:41
@rohityadav1993
rohityadav1993 force-pushed the rohity/pr1-arm4-memory-crossover-bench branch from 97151cd to 968d36f Compare September 27, 2026 03:28
@rohityadav1993
rohityadav1993 force-pushed the oss/pr1-streaming-selection-combine branch from d605e80 to ff800d6 Compare September 28, 2026 17:47
@rohityadav1993
rohityadav1993 force-pushed the oss/pr1-streaming-selection-combine branch from ff800d6 to 8271cf9 Compare September 29, 2026 07:07
@rohityadav1993
rohityadav1993 force-pushed the rohity/pr1-arm4-memory-crossover-bench branch from 968d36f to 0143a87 Compare September 29, 2026 08:48
Adds the JMH harness used to measure the streaming selection ORDER BY
combine against the existing MinMaxValueBased combine. Confined to
pinot-perf; no production code is touched.

  SortedMergeFixture          builds and caches the segment fixture, with
                              DISJOINT/PARTIAL/FULL overlap and a
                              keyCardinality knob giving runs of K rows
                              per distinct tsCol value
  SortedMergeDriver           builds the query context and runs one
                              configuration on either arm
  SortedMergeMemoryProbe      the correctness gate: runs both arms on
                              every configuration and requires agreement
  SortedMergeCpuMeter         per-thread CPU accounting
  BenchmarkSortedMergeMemoryCrossover
                              the JMH entry point
  sorted-merge-bisect-xmx.sh  bisecting -Xmx search for the minimum heap
                              a configuration survives on

The query shape is selectable: ORDER BY tsCol, valCol keeps a total
order, and ORDER BY tsCol alone drops the tiebreaker so the leaf's scan
block tightens to the limit. Under the latter the order is not total, so
the gate compares a commutative digest over tsCol values rather than
whole rows: which tied rows come back is free, but how many rows carry
each value is not.

SortedMergeGateTest adds a test source root to pinot-perf and covers the
gate's comparison primitives with negative controls, since running the
gate by hand only ever exercises the pass path.

This is the harness as run for the published results, committed before
any follow-up changes so the numbers stay reproducible against it.
The -Xmx bisection forks a fresh JVM per attempt, so two attempts can race
to build the same fixture key. Nothing serialized them: the in-JVM
SEGMENT_CACHE only orders callers within one process.

Guard each key with a FileLock held across validation and build, and route
destroyAll() and purgeAll() through the same per-key lock instead of
deleting the base directory wholesale. Without that the teardown path
bypassed the lock entirely and could delete a directory another process was
still building.

FileLock is scoped to the JVM rather than the thread, so a teardown
colliding with a build inside one JVM throws OverlappingFileLockException.
That extends IllegalStateException, not IOException, so it escaped the
existing handler -- out of a shutdown hook, abandoning the rest of the
cleanup loop. Both sides now handle it: the teardown skips the key being
built and continues, and a build that collides with a teardown fails with
an explanation rather than a bare exception.

SortedMergeFixtureTest covers the teardown half, which is the half that
regressed. The build half needs two JVMs and is not reachable from a unit
test. Every test points the fixture at a temporary directory and asserts
the override took effect before anything destructive runs, because
purgeAll() deletes every key under the base directory and would otherwise
destroy the fixtures that published results were measured against.

The base directory override exists only for that guard and is unset on
every non-test path, so the on-disk cache built by previous runs is
unchanged.
gate() decided whether two arms agreed and formatted the result line in one
body, so the decision could only be tested by running a full merge on both
arms. The decision is the part worth pinning: under TS_ONLY the ORDER BY is
not a total order, ties can straddle the LIMIT boundary, and two correct
operators may legitimately return different rows -- so the comparison is
deliberately weaker there than under TS_VAL, and nothing tested that the
relaxation is exactly as wide as intended.

cellVerdict() now returns the failure reason or null, and gate() keeps only
the formatting. The printf output is unchanged.

SortedMergeGateTest gains nine cases over the verdict itself, including the
pair that fixes the shape of the relaxation: the same two arms that TS_ONLY
must accept because they picked different tied rows, TS_VAL must reject.

Also corrects a comment on the TS_ONLY branch claiming the full-row digest
was never computed. gate() computes it for both arms and derives
armsPickedDifferentTiedRows from it. The real reason that digest cannot
decide the verdict is that two correct arms may differ on which tied rows
fall inside the LIMIT.
TS_VAL and TS_ONLY are not two spellings of one query. TS_ONLY makes
sortedColumnsPrefixSize equal the ORDER BY length, which tightens the
leaf's maxDocsPerCall to min(limit + offset, 10000); TS_VAL leaves it
pinned at the block size. That difference is the axis two suites were run
to measure, and it rested on an untested assumption about what
buildQueryContext emits.

SortedMergeDriverTest asserts the ORDER BY of each shape, that the select
list is identical and three wide under both, that the no-arg overload still
means TS_VAL, and that the direction is ascending. No production behavior
changes.
The -Xmx search measures the smallest heap a configuration completes on. A
JVM fails at live set plus GC headroom plus fragmentation, and headroom
scales with allocation rate, which is the one thing the two arms differ in
by about 5x. So the figure is a provisioning requirement, not a retained
set, and a ratio read off the grid is not a ratio of retention. The script
now states that at the top, and the output is named accordingly.

Six defects found in review of the v3 run:

- The floor was never probed. A cell that already fit in LOW_MB converged
  to 95 MB and printed it, a heap size that was never tried. It now probes
  LOW_MB first and reports AT_OR_BELOW_LOW.
- Neither the collector nor -Xms was pinned, so the figures were specific
  to the JDK defaults that produced them and could not be read alongside
  the 4 GB JMH runs. Now -Xms = -Xmx with G1 explicit.
- Survival near the boundary is a sigmoid, and requiring N consecutive
  survivals converges on an unstated quantile that depends on the draws it
  got. Each cell is now measured REPLICATES times, every replicate is its
  own row, and the result is the half-open bracket (low_mb, high_mb] where
  both bounds were actually run.
- REPEATS_ON was 1, justified by identical segments-processed counts. That
  shows the workload is deterministic; what varies at the OOM boundary is
  GC and allocation timing, which is not. Now 2.
- No timeout, so a JVM thrashing without tripping the GC overhead limit
  hung the queue. Now wrapped, with a distinct outcome. The kill backstop
  and the kernel OOM killer share exit 137, so they are separated by
  elapsed time: recording a wedged JVM as an OOM would have manufactured
  evidence for the claim under test.
- No per-attempt record. Roughly 900 launches collapsed into 48 rows, so a
  non-monotone response was unrecoverable afterwards. Each attempt is now
  logged with its outcome, duration, and segments processed.

The tolerance is relative, since a flat 32 MB was half the floor and 0.8%
of the ceiling. TOLERANCE_MB is refused rather than ignored, because a knob
someone deliberately set must not be silently dropped.

sorted-merge-bisect-xmx-selftest.sh stubs java and injects the failure
modes a real run cannot produce on demand. The grid axes became
overridable so it can do that in seconds rather than hours; the banner
prints the grid actually used, so a run says what it measured.

The CSV schema changed and v3 output no longer parses. The peak live set
the PR's claim is really about still needs a separate instrument, which
this is deliberately not.
@rohityadav1993
rohityadav1993 force-pushed the rohity/pr1-arm4-memory-crossover-bench branch from 0143a87 to e38986b Compare October 5, 2026 18:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant