Skip to content

Configure OPEN_STRUCT indexes the way every other column is configured - #19736

Merged
raghavyadav01 merged 4 commits into
apache:masterfrom
raghavyadav01:openstruct-per-key-index-config
Oct 6, 2026
Merged

raghavyadav01 merged 4 commits into
apache:masterfrom
raghavyadav01:openstruct-per-key-index-config

Conversation

@raghavyadav01

@raghavyadav01 raghavyadav01 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

PR flow

Enables dictionary config via indexes for OPEN_STRUCT keys and allows index settings on the sparse blob column, making its configuration uniform with other columns.

flowchart TD
  N0["OpenStructIndexConfig #40;F6#41;"]:::stAdded
  N1["IndexLoadingConfig#46;withOpenStructChildConfigs #40;F2#41;"]:::stModified
  N2["FieldIndexConfigsUtil#46;fromFieldConfig #40;F5#41;"]:::stModified
  N3["ImmutableOpenStructDataSource#46;getSparseJsonIndex #40;F3#41;"]:::stModified
  N4["MapFilterOperator#46;trySparseJsonIndex#47;tryJsonIndex #40;F1#41;"]:::stModified
  N0 -->|"getSparseFieldConfig used for sparse column"| N1
  N1 -->|"per#8209;key and blob configs built via fromFieldConfig"| N2
  N2 -->|"FieldIndexConfigs stored then used to build sparse JSONIndex"| N3
  N3 -->|"MapFilterOperator retrieves and checks the index"| N4
  classDef stAdded fill:#dafbe1,stroke:#1a7f37,color:#1f2328,stroke-width:2px
  classDef stModified fill:#fff8c5,stroke:#9a6700,color:#1f2328,stroke-width:2px
  classDef stRemoved fill:#ffebe9,stroke:#cf222e,color:#1f2328,stroke-width:2px
  classDef stUnchanged fill:#f6f8fa,stroke:#656d76,color:#1f2328,stroke-width:1px
Loading

AI-generated · Green: added · Yellow: modified · Red: removed · Gray: existing

Diff evidence
  • F1: pinot-core/src/main/java/org/apache/pinot/core/operator/filter/MapFilterOperator.java — before · after
  • F2: pinot-segment-local/src/main/java/org/apache/pinot/segment/local/segment/index/loader/IndexLoadingConfig.java — before · after
  • F3: pinot-segment-local/src/main/java/org/apache/pinot/segment/local/segment/index/openstruct/ImmutableOpenStructDataSource.java — before · after
  • F5: pinot-segment-spi/src/main/java/org/apache/pinot/segment/spi/index/FieldIndexConfigsUtil.java — before · after
  • F6: pinot-spi/src/main/java/org/apache/pinot/spi/config/table/OpenStructIndexConfig.java — before · after
  • Regenerate PR flow

An OPEN_STRUCT column is split at segment build time into materialized
per-key child columns plus one shared blob column holding every key that was
not materialized. Both halves were configurable in their own idiosyncratic
way rather than the way the rest of Pinot is configured. Two fixes.

1. A key can ask for its dictionary under indexes

A dictionary is enabled on an ordinary column by putting it in indexes. An
OPEN_STRUCT key could not be: FieldIndexConfigsUtil#fromFieldConfig, which
builds a key's index configs, skipped DICTIONARY_ID outright and took the
answer from encodingType alone — and the validator then rejected an
indexes.dictionary entry that did not already agree with it. The same JSON
meant different things depending on which kind of column you wrote it on, and
on a key it failed silently: you got a raw key and no diagnostic.

An enabled inverted index hid half of this, because it forces a dictionary on
its own. A key whose only index is a range one had nothing to rescue it — the
dictionary entry was ignored, the key stayed raw, and the range index was then
asked for over a dictionary that did not exist.

indexes.dictionary can only turn a dictionary on. Turning one off stays
with encodingType, so a disabled entry that contradicts the column's own
encoding is still a contradiction and is still refused rather than quietly
winning. The forward index now follows the dictionary decision rather than the
declared encoding, for the same reason it already followed encodingType: a
key whose dictionary was enabled this way needs a dict-encoded forward index,
or a reload rebuilds one over a dictionary that is not there.

2. The sparse blob column can carry index settings

The per-key children each accept a FieldConfig. The blob column could not —
index loading skipped it before it reached the index-config builder, so the
only way to evaluate a predicate against a key living inside the blob was a
full scan over every document.

OpenStructIndexConfig gains an optional sparseFieldConfig. It is a plain
FieldConfig, shaped exactly like the one you would write for any other
column:

"openStruct": {
  "sparseFieldConfig": {
    "name": "props$__sparse__",
    "indexes": { "json": {} }
  }
}

The blob is forced to RAW whatever the config asks for. Each value is a
serialized document for one row, so a dictionary over it would be a dictionary
of whole documents — it dedupes nothing and costs a round trip per lookup.

The existing sparseJsonIndex: true flag keeps working and is now defined as
exactly indexes.json = {}, so the two spellings converge on one code path
instead of being maintained separately.

Testing

OpenStructPerKeyIndexReloadTest is now 9 tests. The one that fails when the
first production change alone is reverted is
testARangeOnlyKeyCanAskForADictionaryUnderIndexes. Two new cases build a
JSON index on the blob, via sparseJsonIndex and via sparseFieldConfig
naming the blob column directly.

369 tests green across the OPEN_STRUCT, config and loader suites; existing
per-key coverage (dictionary, range, inverted, bloom) unchanged.

Compatibility

Additive. A table config with no sparseFieldConfig behaves exactly as
before, sparseJsonIndex produces the index it always did, and the previous
10-arg OpenStructIndexConfig creator is deprecated and delegating, so
existing configs deserialize unchanged.

@raghavyadav01
raghavyadav01 requested a review from xiangfu0 October 2, 2026 04:45
@codecov-commenter

codecov-commenter commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 84.09091% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 68.25%. Comparing base (8cbbc1a) to head (8f89a8d).

Files with missing lines Patch % Lines
...local/segment/index/loader/IndexLoadingConfig.java 85.00% 0 Missing and 3 partials ⚠️
...pinot/segment/spi/index/FieldIndexConfigsUtil.java 57.14% 0 Missing and 3 partials ⚠️
.../pinot/core/operator/filter/MapFilterOperator.java 66.66% 1 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master   #19736      +/-   ##
============================================
+ Coverage     68.22%   68.25%   +0.02%     
  Complexity     1450     1450              
============================================
  Files          3520     3520              
  Lines        228899   228935      +36     
  Branches      36305    36318      +13     
============================================
+ Hits         156171   156261      +90     
+ Misses        60489    60395      -94     
- Partials      12239    12279      +40     
Flag Coverage Δ
integration 100.00% <ø> (ø)
integration1 100.00% <ø> (ø)
integration2 0.00% <ø> (ø)
java-25 68.25% <84.09%> (+0.02%) ⬆️
lane-a 100.00% <ø> (ø)
lane-b 0.00% <ø> (ø)
temurin 68.25% <84.09%> (+0.02%) ⬆️
unittests 68.25% <84.09%> (+0.02%) ⬆️
unittests1 58.10% <22.72%> (-0.04%) ⬇️
unittests2 39.97% <70.45%> (+0.05%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

…umn does

A dictionary is enabled on an ordinary column by putting it in `indexes`.
An OPEN_STRUCT key could not be configured that way: fromFieldConfig, which
builds a key's index configs, skipped the dictionary type outright and took
the answer from encodingType alone, and the validator then rejected an
indexes.dictionary entry that did not already agree with it.

So the two spellings disagreed depending on which kind of column you were
configuring, and the key spelling failed silently -- you got a raw key and
no diagnostic.

An enabled inverted index hid half of it, since it forces a dictionary on
its own. A key whose only index is a range one had nothing to rescue it:
the dictionary entry was ignored, the key stayed raw, and the range index
was asked for over a dictionary that did not exist. That case is
testARangeOnlyKeyCanAskForADictionaryUnderIndexes, and it is the one that
fails when the production change alone is reverted.

indexes.dictionary can only turn a dictionary *on*. Turning one off stays
with encodingType, so a `disabled` entry contradicting the column's own
encoding is still a contradiction and still refused rather than quietly
winning.

The forward index follows the dictionary decision rather than the declared
encoding, for the reason it already followed encodingType: a key whose
dictionary was enabled this way needs a dict-encoded forward index, or a
reload rebuilds one over a dictionary that is not there.

369 tests green across the OPEN_STRUCT, config and loader suites.
An OPEN_STRUCT column splits into materialized per-key child columns and
one shared blob column holding every key that was not materialized. The
per-key children already accept a FieldConfig each, so a key can ask for
a dictionary, a range index, a bloom filter and so on. The blob column
could not: index loading skipped it outright, so the only way to search
the keys living inside it was a full scan.

This adds a sparseFieldConfig to OpenStructIndexConfig. It is shaped
like the FieldConfig of any other column, and it configures the blob the
same way a per-key entry configures a child:

  "openStruct": {
    "sparseFieldConfig": {
      "name": "props$__sparse__",
      "indexes": { "json": {} }
    }
  }

The blob is always RAW regardless of what the config asks for -- it is a
serialized document per row, so a dictionary over it would be a
dictionary of whole documents and buys nothing. The existing
sparseJsonIndex flag still works and now means exactly indexes.json={},
so the two spellings converge on one code path instead of two.

Adds four tests covering a JSON index built on the blob through either
spelling, and keeps the existing per-key cases green.
The sparse fast path in MapFilterOperator asks the OPEN_STRUCT data source for
the blob's JSON index, and that lookup only ever returned the standard one.
An ordinary column does not stop there: tryJsonIndex falls back to the
composite JSON index, which reads as a JsonIndexReader and answers the same
predicates.

So a blob given a composite JSON index built it and was then never asked. The
key fell back to a full scan, silently costing exactly what the index was
configured to avoid -- no error, no plan difference a reader would notice,
just the slow path.

Make the blob lookup do what the column lookup already does. The index ships
as a plugin and is absent from most deployments, so it is resolved by id and
the fallback is inert where nothing registered it.
The standard JSON index indexes every path, so asking it about any key is
safe. An index that indexes a configured subset of paths is not: asked about
a path it never indexed, it answers with an empty bitmap, and an empty bitmap
is indistinguishable from "no document matches this". The query returns zero
rows, the plan says the index served the filter, and nothing anywhere says the
answer is wrong.

`JsonIndexReader.isPathIndexed` already exists to answer exactly this, and
defaults to true so a fully-indexed reader is unaffected. Nothing was calling
it. Both JSON fast paths now do -- the one over a column and the one over an
OPEN_STRUCT sparse blob -- and a key the index does not cover falls back to
the scan, which is slower and right.

The two existing tests that expect the index to serve now stub `isPathIndexed`:
Mockito answers false for an unstubbed boolean, interface default or not.
@raghavyadav01
raghavyadav01 force-pushed the openstruct-per-key-index-config branch from 50da4bf to 8f89a8d Compare October 2, 2026 04:49

@tarun11Mavani tarun11Mavani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@raghavyadav01
raghavyadav01 merged commit 21596bb into apache:master Oct 6, 2026
14 of 15 checks passed
xiangfu0 added a commit to pinot-contrib/pinot-docs that referenced this pull request Oct 6, 2026
Docs follow-up for apache/pinot#19736. Documents enabling a dictionary
for a RAW materialized key through indexes.dictionary and configuring
the shared sparse blob through sparseFieldConfig, including its raw
encoding and compatibility with sparseJsonIndex.\n\nValidation: full
docs validator passed (existing non-fatal anchor warnings only); git
diff --check passed.

Co-authored-by: Xiang Fu <xiangfu@Xiang-mac-mtv-2.local>
@xiangfu0

xiangfu0 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Docs follow-up: pinot-contrib/pinot-docs#1086 (merged). The OPEN_STRUCT reference and ingestion guide now cover per-key indexes.dictionary and sparseFieldConfig for the shared sparse column.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants