Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions build-with-pinot/ingestion/complex-type/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,8 @@ Pinot decides which keys are dense in this order:

Dense keys reuse Pinot's standard column infrastructure, so each materialized key gets a forward index and can also use vetted per-key settings for dictionary, inverted, range, and bloom-filter behavior through `valueFieldConfigs`. If you do not configure a dense key explicitly, Pinot defaults to dictionary encoding plus an inverted index for that key.

For a materialized key configured with `"encodingType": "RAW"`, add `"dictionary": {}` under its `indexes` to enable a dictionary without also enabling an inverted index. This supports a range index on that key while keeping the per-key config consistent with ordinary columns.

After changing a materialized key's index settings, reload the affected segments to apply them, as you would for an ordinary column. Reloading does not change which keys are dense: that choice is made when each segment is built.

### Define the schema
Expand Down Expand Up @@ -187,6 +189,8 @@ For a materialized key, Pinot reads the generated child column and can use its d

Keys stored in the shared sparse column are also available through the item operator. Pinot exposes each sparse key as a virtual typed data source, so projections, filters, grouping, and aggregations use the same SQL syntax as dense keys. Sparse keys use scan-based execution by default. Set `sparseJsonIndex` to `true` to build a JSON index over the sparse column; Pinot can use it for compatible string-key equality and `IN` predicates, while other predicates continue to scan the virtual data source.

You can instead configure the sparse blob with `sparseFieldConfig` inside `open_struct`. For example, `"sparseFieldConfig": {"name": "attributes$__sparse__", "indexes": {"json": {}}}` builds the same JSON index as `"sparseJsonIndex": true`. The blob remains raw-encoded even if its field config requests a dictionary. See the [table reference](../../../reference/configuration-reference/table.md#open_struct-index) for the complete config surface.

Keys with array values can be stored and queried as multi-value keys. A declared child field controls its type and single- or multi-value shape; for an undeclared key, Pinot infers the shape and element type from ingested values. You can also select the whole `OPEN_STRUCT` column (including through `SELECT *`); Pinot reconstructs it as a JSON document from dense and sparse keys.

New segments record the inferred type of undeclared sparse keys, so a key keeps its type whether it is dense or sparse in that segment. Older segments without this type metadata continue to expose undeclared sparse keys as `STRING`; rebuild those segments if you need their inferred types preserved.
Expand Down
5 changes: 3 additions & 2 deletions reference/configuration-reference/table.md
Original file line number Diff line number Diff line change
Expand Up @@ -250,14 +250,15 @@ The `open_struct` config object supports the following properties:
| `denseKeyMinFillRate` | Minimum fraction of documents that must contain a key before Pinot materializes it automatically. Default: `0.5`. |
| `maxDenseKeys` | Maximum number of dense keys to materialize. `-1` means unlimited, `0` disables dense keys entirely, and positive values keep only the highest-fill-rate qualifying keys as dense. |
| `maxNestedKeyDepth` | Maximum number of path segments exposed as keys for nested objects. Default: `1` (no flattening); `2` exposes `a.b`, and `3` exposes `a.b.c`. Nested objects remain stored as whole values at the depth limit. |
| `sparseJsonIndex` | Whether to build a JSON index over the shared sparse column. Default: `false`. When enabled, compatible string-key equality and `IN` filters can use the JSON index; other sparse-key operations use the virtual scan path. |
| `sparseJsonIndex` | Whether to build a JSON index over the shared sparse column. Default: `false`. This existing option remains supported; it is equivalent to adding `"json": {}` under `sparseFieldConfig.indexes`. Compatible string-key equality and `IN` filters can use the JSON index; other sparse-key operations use the virtual scan path. |
| `sparseFieldConfig` | Optional `FieldConfig` for the shared `<column>$__sparse__` column. Set `name` to that generated column name and configure indexes under `indexes`, for example `"indexes": {"json": {}}`. Pinot stores this blob with raw encoding even if the config requests a dictionary. |
| `perKeyMetricsEnabled` | Whether to emit `OPEN_STRUCT_LAST_SEGMENT_KEY_DOC_COUNT` for every key discovered in a sealed segment. Default: `false`, which limits the gauge to keys listed in `denseKeys`. Enabling this for an unbounded key space can create high metric cardinality. |
| `defaultValueFieldConfig` | Optional fallback `FieldConfig` applied to dense keys that do not appear in `valueFieldConfigs`. |
| `valueFieldConfigs` | Optional list of per-key `FieldConfig` entries. Each entry's `name` must match an `OPEN_STRUCT` key name. |

Per-key `FieldConfig.indexes` inside `defaultValueFieldConfig` or `valueFieldConfigs` may use only the vetted subset: `dictionary`, `forward`, `inverted`, `range`, and `bloom`. Forward indexes are always written for materialized child columns. When neither `defaultValueFieldConfig` nor `valueFieldConfigs` configures a dense key, Pinot uses dictionary encoding plus an inverted index for that key by default.

Set per-key `encodingType` to choose dictionary or raw encoding; enabling an inverted index also makes the key dictionary-encoded. Materialized keys always get a forward index, with LZ4 compression for raw keys. Per-key `indexes.forward` and `indexes.dictionary` may only restate that resulting encoding and keep the forward index enabled; they cannot tune its format or compression. Pinot rejects `codecSpec`, `compressionCodec`, `indexTypes`, `timestampConfig`, non-empty `properties` or `tierOverwrites`, and forward or dictionary options that conflict with the materialized key. These rules apply to both `valueFieldConfigs` and `defaultValueFieldConfig`, including multi-value and nested keys. Remove such settings before updating an existing table config; already-loaded tables continue to ingest.
Set per-key `encodingType` to choose dictionary or raw encoding; enabling an inverted index also makes the key dictionary-encoded. For a key declared `RAW`, `"indexes": {"dictionary": {}}` enables a dictionary without an inverted index, which is useful with a range index. Materialized keys always get a forward index, with LZ4 compression for raw keys. Per-key `indexes.forward` cannot tune its format or compression, and `indexes.dictionary` cannot disable a dictionary selected by the key's encoding. Pinot rejects `codecSpec`, `compressionCodec`, `indexTypes`, `timestampConfig`, non-empty `properties` or `tierOverwrites`, and forward or dictionary options that conflict with the materialized key. These rules apply to both `valueFieldConfigs` and `defaultValueFieldConfig`, including multi-value and nested keys. Remove such settings before updating an existing table config; already-loaded tables continue to ingest.

Example:

Expand Down
Loading