Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 34 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,38 @@ adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

## [0.2.0] - 2026-05-31

Format explosion — anomalyx now normalizes ~30 formats spanning logs, security
telemetry, network captures, observability streams, spreadsheets, and data-lake
files, all behind the same record-model boundary and detector taxonomy.

### Added

- **Logs & observability** parsers: `logfmt`, web access logs (Combined/Common),
`syslog` (RFC 3164/5424), `systemd journal` (`journalctl -o json`),
`Prometheus`/OpenMetrics, and `OpenTelemetry` (OTLP/JSON traces).
- **Security telemetry** parsers: `CEF`/`LEEF`, Linux `auditd`, `EVTX` (Windows
Event Log), Suricata/Zeek `EVE` JSON, `osquery` results, and AWS `CloudTrail`.
- **Network** parsers: `PCAP`/`PCAPNG` (beaconing/C2 via `cadence`), `NetFlow`/
IPFIX (nfdump CSV), AWS `VPC Flow Logs`, and DNS query logs (DGA/exfil via
`point` on query-name entropy/length).
- **Structured-data** parsers: `YAML`, `TOML`/`INI`, and `XML`
(Nessus/OpenVAS/SOAP).
- **Columnar, data-lake & database** parsers: `Avro`, `ORC`, Excel/`ODS`
(`xlsx`/`xls`/`xlsb`), and `SQLite` — joining the existing Parquet/Arrow.
- Several parsers **compute detection features** (DNS name entropy/length, flow
`duration`, span durations, normalized epoch timestamps) and rename source
fields to a canonical schema.
- Binary/heavyweight parsers sit behind **default-on feature flags**
(`evtx`, `pcap`, `xlsx`, `sqlite`, `datalake`, `polars`), so
`--no-default-features` is a lean text-only normalizer.

### Notes

- 32 parser plugins total; each ships its own property/exact tests and passes
the workspace-wide 0-surviving-mutant gate.

## [0.1.0] - 2026-05-30

Initial release — a contract-first anomaly-detection CLI over arbitrary corpora.
Expand Down Expand Up @@ -47,5 +79,6 @@ Initial release — a contract-first anomaly-detection CLI over arbitrary corpor
gates on every push.
- Dual-licensed under MIT OR Apache-2.0.

[Unreleased]: https://github.com/copyleftdev/anomalyx/compare/v0.1.0...HEAD
[Unreleased]: https://github.com/copyleftdev/anomalyx/compare/v0.2.0...HEAD
[0.2.0]: https://github.com/copyleftdev/anomalyx/compare/v0.1.0...v0.2.0
[0.1.0]: https://github.com/copyleftdev/anomalyx/releases/tag/v0.1.0
48 changes: 41 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,14 @@

Contract-first anomaly detection over arbitrary corpora — a CLI built on the
thesis of [*AI Tools Need Contracts, Not Prompts*][article]: **the executable is
the contract.** Point a normalizer at any supported format, run a battery of
typed anomaly detectors, and get back a dense, versioned, machine-readable
envelope an agent can trust — not pretty text it has to scrape.
the contract.**

anomalyx meets your data where it already lives. Point it at **~30 formats** —
logs, security telemetry, packet captures, flow records, observability streams,
spreadsheets, and data-lake files — and it normalizes each into one typed record
model, runs a battery of deterministic anomaly detectors, and returns a dense,
versioned, machine-readable envelope an agent can trust — not pretty text it has
to scrape.

[article]: https://dev.to/copyleftdev/ai-tools-need-contracts-not-prompts-5ca3

Expand Down Expand Up @@ -56,17 +61,46 @@ $ ... | anomalyx explain cell:amount:8
cleanly with exit `2`.
- **Handle-based evidence** — `scan` stays compact; `explain` drills in.

## Formats

**32 built-in parsers** across five domains — each an independent plugin, each
lowered to the same typed `RecordSet`:

- **Tabular & structured** — CSV, TSV, NDJSON, JSON, YAML, TOML/INI, XML
- **Columnar, data-lake & databases** — Parquet, Arrow IPC, Avro, ORC,
Excel/ODS, SQLite
- **Logs & observability** — logfmt, web access logs, syslog (RFC 3164/5424),
systemd journal, Prometheus/OpenMetrics, OpenTelemetry (OTLP)
- **Security telemetry** — Zeek, CEF/LEEF, auditd, EVTX (Windows Event Log),
Suricata/Zeek EVE, osquery, AWS CloudTrail
- **Network** — PCAP/PCAPNG, NetFlow/IPFIX (nfdump CSV), AWS VPC Flow Logs,
DNS query logs

Several parsers compute the features the detectors want — DNS query-name entropy
& length, flow `duration`, span durations, normalized timestamps — and rename
cryptic source fields to a canonical schema. So the same taxonomy lights up
across domains: **beaconing/C2** via `cadence` on PCAP inter-arrival times,
**DGA/exfil** via `point` on DNS name entropy, **config drift** via
`struct.schema` on YAML/TOML, **exfil** via `mv.mahalanobis` on NetFlow
(bytes, packets, duration), **alert-type drift** via `dist.chi2` on EVE/CEF.

Resolution is by extension first, then deterministic content sniff (binary magic
before text signatures); an unrecognized stream is an explicit error, never a
guess. Binary/heavyweight parsers sit behind default-on feature flags, so
`--no-default-features` yields a lean, text-only normalizer. Full table:
[docs › Input & normalization](https://copyleftdev.github.io/anomalyx/formats.html).

## Architecture

```
crates/
ax-core contract types: RecordSet, anomaly taxonomy, tq1 envelope,
handles, deterministic reductions (no heavy deps — the contract
stays engine-independent and the mutation gate stays fast)
ax-normalize any input format → RecordSet (CSV/TSV/NDJSON/JSON via a lean
deterministic reader; Parquet/Arrow IPC via the Polars backbone,
behind the default-on `polars` feature — all lowered to the same
RecordSet so detectors never see a Polars type)
ax-normalize any input format → RecordSet (32 parser plugins — text via a
lean deterministic reader, binary/library-backed formats behind
default-on feature flags — all lowered to the same RecordSet so
detectors never see a library type. See "Formats" below)
ax-detect Detector trait + registry; detection math assembled from
statrs, not reinvented
anomalyx the four-verb CLI surface (the installable crate / binary)
Expand Down
108 changes: 94 additions & 14 deletions docs/src/formats.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,24 +3,104 @@
> *"Given any corpus of information regardless of its format, we'll normalize
> it."*

Every supported format is lowered to one engine-independent record model — a
`RecordSet` of named, typed columns — and detectors only ever see that. The
anomalyx meets your data where it already lives. Every supported format —
whether a packet capture, a SIEM event stream, a Kubernetes manifest, or a
data-lake file — is lowered to one engine-independent record model, a
`RecordSet` of named, typed columns, and the detectors only ever see that. The
contract stays stable while the backend underneath it changes.

## Supported formats

**32 built-in parsers** across five domains. Each is an independent plugin
(`crates/ax-normalize/src/parsers/`); adding one doesn't touch the others.

### Tabular & structured data

| Format | Extensions | Notes |
|---|---|---|
| CSV / TSV | `.csv`, `.tsv`, `.tab` | lean deterministic reader |
| NDJSON / JSON | `.ndjson`, `.jsonl`, `.json` | array, object, or one-record-per-line |
| YAML | `.yaml`, `.yml` | Kubernetes / CI manifests; multi-document |
| TOML / INI | `.toml`, `.ini`, `.cfg`, `.conf` | config drift via `struct.schema` |
| XML | `.xml`, `.nessus` | Nessus/OpenVAS, SOAP; repeated element → rows |

### Columnar, data-lake & databases

| Format | Extensions | Backend |
|---|---|---|
| CSV | `.csv` | lean deterministic reader |
| TSV | `.tsv`, `.tab` | lean deterministic reader |
| NDJSON | `.ndjson`, `.jsonl` | lean deterministic reader |
| JSON | `.json` | lean deterministic reader |
| Parquet | `.parquet`, `.pq` | Polars / Arrow |
| Arrow IPC | `.arrow`, `.ipc`, `.feather` | Polars / Arrow |
| Avro | `.avro` | `apache-avro` |
| ORC | `.orc` | `orc-rust` → Arrow |
| Excel / ODS | `.xlsx`, `.xls`, `.xlsb`, `.ods` | `calamine` (first sheet) |
| SQLite | `.db`, `.sqlite`, `.sqlite3`, `.db3` | `rusqlite` (first table, in-memory deserialize) |

### Logs & observability

| Format | Detected by | Anomaly angle |
|---|---|---|
| logfmt | `key=value` shape | structured app logs |
| Web access logs (Combined/Common) | `[time] "request" status` | status-mix `dist`, latency `point`, bursts `coll` |
| syslog (RFC 3164 / 5424) | `<PRI>` header | event-rate `dist`, off-hours `contextual` |
| systemd journal | `journalctl -o json` | event-rate `cadence`/`coll`, rare-unit `dist` |
| Prometheus / OpenMetrics | exposition lines | per-series `point` spikes, `dist` drift |
| OpenTelemetry (OTLP/JSON) | `resourceSpans` | span-duration `point`, error-rate `dist`, emit `cadence` |

### Security telemetry

| Format | Detected by | Anomaly angle |
|---|---|---|
| Zeek (`conn.log` family) | `#separator` header | connection analytics |
| CEF / LEEF | `CEF:` / `LEEF:` prefix | signature/category mix shift via `dist.chi2` |
| auditd | `msg=audit(` | exec/syscall mix `dist`, bursty activity `coll` |
| EVTX (Windows Event Log) | `ElfFile` magic | rare event-ID `point`, logon `dist`, off-hours `contextual` |
| Suricata/Zeek EVE | `event_type` + `timestamp` | alert-type drift via `dist.chi2`; new classes surface |
| osquery results | `hostIdentifier` + `columns`/`snapshot` | fleet-posture drift via `structural`/`dist` |
| AWS CloudTrail | `Records[].eventName` | off-hours `contextual`/`cadence`, rare-API `dist` |

### Network

| Format | Detected by | Anomaly angle |
|---|---|---|
| PCAP / PCAPNG | libpcap / SHB magic | **beaconing/C2 via `cadence`** on inter-arrival times |
| NetFlow / IPFIX (nfdump CSV) | nfdump header | exfil via `mv.mahalanobis` on (bytes, packets, duration) |
| AWS VPC Flow Logs | `srcaddr dstaddr dstport` header | same flow anomalies, zero new infra |
| DNS query logs (dnsmasq) | `query[TYPE] … from` | DGA/exfil via `point` on name **entropy/length** + `cadence` |

Several parsers compute the features the detectors want rather than just
extracting fields — DNS query-name Shannon entropy and length, flow `duration`
(`end - start`), span `durationNanos`, normalized epoch timestamps — and rename
cryptic source fields to a canonical schema (e.g. nfdump `ibyt`→`bytes`,
`td`→`duration`).

## Resolution

Format is resolved by **file extension first**, then by **content sniff** —
binary magic numbers (`PAR1`, `ORC`, `SQLite format 3\0`, …) are checked at high
confidence, then distinctive text signatures, then a CSV last-resort fallback.
Resolution is deterministic: the highest-confidence match wins, ties break by
registration order. An unrecognized stream is an explicit error, never a silent
guess.

Several formats deliberately claim **no extension** (Zeek, syslog content,
journald, EVE, osquery, auditd, DNS, NetFlow, VPC) because their files are
generically `*.log`/`*.json`; pipe them on stdin and the content signature
routes them.

## Feature flags & the lean build

The binary and heavyweight parsers sit behind **default-on feature flags**, so a
default build reads everything but a `--no-default-features` build is a lean,
text-only normalizer with no binary dependencies:

Format is resolved by extension first, then by content sniff (binary magic
numbers `PAR1` / `ARROW1` are checked before UTF-8 text sniffing). An
unrecognized stream is an explicit error, never a silent guess.
| Feature | Parsers |
|---|---|
| `polars` | Parquet, Arrow IPC |
| `evtx` | EVTX |
| `pcap` | PCAP / PCAPNG |
| `xlsx` | Excel / ODS |
| `sqlite` | SQLite |
| `datalake` | Avro, ORC |

## The record model

Expand All @@ -36,8 +116,8 @@ amount,tier → column "amount": Int [10, 11, 9, …]
11,b
```

Binary columnar formats live entirely behind this boundary: the Polars
`DataFrame` is converted to a `RecordSet` (integers fold to `i64`, floats to
`f64` with non-finite → `Null`, unsupported logical types preserved as their
string form), so no Polars type ever reaches a detector. Text formats never
touch Polars at all.
Binary and library-backed formats live entirely behind this boundary: a Polars
`DataFrame`, an Arrow `RecordBatch`, a calamine sheet, or a SQLite row is
converted to a `RecordSet` (integers fold to `i64`, floats to `f64` with
non-finite → `Null`, unsupported logical types preserved as their string form),
so no library type ever reaches a detector. Text formats touch none of it.
9 changes: 6 additions & 3 deletions docs/src/introduction.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,9 +8,12 @@

anomalyx is a deterministic Rust CLI built on the thesis of
[*AI Tools Need Contracts, Not Prompts*][article]: **the executable is the
contract.** Point it at any supported format, run a battery of typed anomaly
detectors, and get back a dense, versioned, machine-readable envelope an agent
(or a human) can trust — not pretty text that has to be scraped.
contract.** Point it at **~30 formats** — logs, security telemetry, packet
captures, flow records, observability streams, spreadsheets, and data-lake files
([the full set](./formats.md)) — and it normalizes each into one typed record
model, runs a battery of typed anomaly detectors, and returns a dense, versioned,
machine-readable envelope an agent (or a human) can trust — not pretty text that
has to be scraped.

[article]: https://dev.to/copyleftdev/ai-tools-need-contracts-not-prompts-5ca3

Expand Down