Skip to content

Add Avro and ORC parsers (data-lake siblings of Parquet) - #46

Merged
copyleftdev merged 1 commit into
mainfrom
feat/avro-orc
May 31, 2026
Merged

copyleftdev merged 1 commit into
mainfrom
feat/avro-orc

Conversation

@copyleftdev

Copy link
Copy Markdown
Owner

Implements the Avro / ORC plugins — issue #23 (the last of the 23 format issues).

Both lower to the same engine-independent Columns as the Parquet/Arrow
parsers, so no library type escapes the contract.

AvroParser (apache-avro)

Each record in the object-container file is a row; record fields become typed
columns via avro_to_value: bool/int/long/float/double/string/enum mapped;
bytes/fixed → hex Str; date/time logical types → their integer value;
unions unwrap to their held value; nested records/arrays/maps/decimals → Null
(honest absence for v1's flat-scalar lowering). Magic Obj\x01; extension .avro.

OrcParser (orc-rust → Arrow)

The file is read into Arrow record batches; each cell is rendered and run through
infer_scalar, so numbers/bools become typed columns and nulls are preserved.
Magic ORC; extension .orc. arrow is pinned to the major orc-rust uses so
the RecordBatch types unify.

Feature gating

Both behind the default-on datalake feature (binary formats); the text-only
--no-default-features build stays lean. Both builds verified.

Testing

Roundtrip tests write a tiny Avro file (apache-avro Writer) and a tiny ORC file
(orc-rust Arrow writer) in-memory — no committed binaries; avro_to_value is
unit-tested across every handled variant (incl. union unwrap, date/time, bytes);
non-format input is a clean AxError::Parse.

Gates

  • fmt / clippy -D warnings / full workspace tests green (default and
    --no-default-features).
  • Mutation gate: 0 surviving mutants on avro.rs. Deleting avro_to_value's
    explicit Null arm is a documented equivalent (it and the catch-all both
    return Null).

Closes #23

🤖 Generated with Claude Code

Two parsers in ax-normalize, both lowering to the same engine-independent
Columns as the Parquet/Arrow parsers so no library type escapes the contract.

AvroParser (apache-avro): each record in the object-container file is a row;
record fields become typed columns via avro_to_value (bool/int/long/float/double
/string/enum mapped; bytes/fixed -> hex Str; date/time logical types -> their
integer value; unions unwrap; nested records/arrays/maps/decimals -> Null, honest
absence for v1's flat-scalar lowering). Magic Obj\x01; extension .avro.

OrcParser (orc-rust -> Arrow): the file is read into Arrow record batches; each
cell is rendered and run through infer_scalar so numbers/bools become typed
columns, nulls preserved. Magic ORC; extension .orc. arrow pinned to the major
orc-rust uses so the RecordBatch types unify.

Both behind the default-on datalake feature (binary formats), so the text-only
build stays lean. Roundtrip tests write a tiny Avro file (apache-avro Writer)
and a tiny ORC file (orc-rust Arrow writer) in-memory — no committed binaries;
avro_to_value is unit-tested across all handled variants; non-format input is a
clean Parse error.

Mutation gate: 0 surviving mutants on the new file. Deleting avro_to_value's
explicit Null arm is a documented equivalent (it and the catch-all both return
Null).

Closes #23

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented May 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@copyleftdev, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 20 minutes and 15 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 028195fb-bf62-4314-bf10-9e6b935aa87d

📥 Commits

Reviewing files that changed from the base of the PR and between d2fd02f and 5940822.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (5)
  • .cargo/mutants.toml
  • crates/ax-normalize/Cargo.toml
  • crates/ax-normalize/src/parser.rs
  • crates/ax-normalize/src/parsers/avro.rs
  • crates/ax-normalize/src/parsers/mod.rs
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/avro-orc

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@copyleftdev
copyleftdev merged commit 0381b97 into main May 31, 2026
2 checks passed
@copyleftdev
copyleftdev deleted the feat/avro-orc branch May 31, 2026 18:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add Avro / ORC parsers (data-lake siblings of Parquet)

1 participant