Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
95 commits
Select commit Hold shift + click to select a range
7c8dc91
feat(datagen): recording toolkit and hand-recorded OpenInference corpora
anticorrelator Aug 20, 2026
034e7df
feat(datagen): add OTLP corpus replayer and phoenix datagen CLI
anticorrelator Aug 20, 2026
4be0535
fix(datagen): group corpus spans across requests
anticorrelator Aug 20, 2026
524ac4b
fix(datagen): preserve replay fidelity and package corpora
anticorrelator Aug 20, 2026
9522215
fix(datagen): address acceptance findings
anticorrelator Aug 21, 2026
b12b763
feat: add optional datagen deployment recipes
anticorrelator Aug 21, 2026
c5328f5
refactor(datagen): rename corpora to datagen assets and scenarios
anticorrelator Aug 21, 2026
dd3a20a
feat(datagen): fragment banks, generation lanes, session composer, di…
anticorrelator Aug 21, 2026
2476532
fix(datagen): make generation-tooling tests importable without PYTHON…
anticorrelator Aug 21, 2026
b15145f
fix(datagen): resolve recorder environments and verify offline recording
anticorrelator Aug 21, 2026
088e35a
chore(datagen): re-record starter assets under current instrumenter pins
anticorrelator Aug 21, 2026
cff4301
feat(datagen): move assets to GCS
anticorrelator Aug 21, 2026
6b9111c
Merge branch 'dustin/datagen-assets-on-gcs' into dustin/data-generati…
anticorrelator Aug 21, 2026
ddf6259
feat(datagen): application profiles, profile-scoped matrix, structure…
anticorrelator Aug 21, 2026
6fa6dfe
feat(datagen): add customer support profiles
anticorrelator Aug 21, 2026
a83bca3
feat(datagen): add coding agent application profiles
anticorrelator Aug 21, 2026
ca9ce63
feat(datagen): add data analyst application profiles
anticorrelator Aug 21, 2026
90c25f3
feat(datagen): add deep research profiles
anticorrelator Aug 21, 2026
80e82a5
feat(datagen): deterministic seed mechanics and materialized environm…
anticorrelator Aug 22, 2026
40d7538
test(datagen): add seed mechanics to generation fixture
anticorrelator Aug 22, 2026
64f316b
feat(datagen): add deep research seed mechanics
anticorrelator Aug 22, 2026
4058bfb
feat(datagen): add data analyst seed mechanics
anticorrelator Aug 22, 2026
c1ed58f
feat(datagen): add coding agent seed mechanics
anticorrelator Aug 22, 2026
d81c30f
feat(datagen): add customer support seed mechanics
anticorrelator Aug 22, 2026
6380261
feat(datagen): judged outcomes with engagement-based routing
anticorrelator Aug 22, 2026
58cc5d6
fix(datagen): align profile and composition boundaries
anticorrelator Aug 22, 2026
470e6cd
fix(datagen): make asset publication owner-run
anticorrelator Aug 22, 2026
519a2c0
fix(datagen): reject invalid conversation structure
anticorrelator Aug 22, 2026
e1d464d
feat(datagen): replay rate schedule, backfill, and error injection
anticorrelator Aug 22, 2026
bf33c5d
feat(datagen): supplemental fault runs and bank merge
anticorrelator Aug 22, 2026
c745ce2
refactor(datagen): trim runtime verification to its floor
anticorrelator Aug 25, 2026
282fa85
refactor(datagen): remove the cost plane and the batch lane
anticorrelator Aug 25, 2026
f5791e2
refactor(datagen): one shared serialization module for the sidecar sc…
anticorrelator Aug 25, 2026
8e2daed
refactor(datagen): share the transcript hygiene names across the guards
anticorrelator Aug 25, 2026
a8dbe0c
refactor(datagen): rename bank to scenario and enforce judged outcome…
anticorrelator Aug 25, 2026
c6a9db3
refactor(datagen): scenario vocabulary and one owner per publish check
anticorrelator Aug 25, 2026
13502a8
style(datagen): format test_codex_exec.py
anticorrelator Aug 25, 2026
ff9f5af
refactor(datagen): default the destination project to phoenix-datagen
anticorrelator Aug 25, 2026
a147b1c
refactor(datagen): zero-config replay with bundled or sole published …
anticorrelator Aug 25, 2026
cacf220
refactor(datagen): drop the seven session-shape tuning flags
anticorrelator Aug 25, 2026
076c845
refactor(datagen): remove backfill, rate schedules, and the anomaly m…
anticorrelator Aug 25, 2026
2f9cec6
Relax datagen replay validation and cache checks
anticorrelator Aug 25, 2026
2b80ccc
Trim datagen generation checks and tests
anticorrelator Aug 25, 2026
211aa17
Flatten datagen's published banks into a single corpus
anticorrelator Aug 25, 2026
8b0880b
Trim the datagen replayer to its live paths
anticorrelator Aug 26, 2026
a8f3656
feat(datagen): simplify corpus archive pipeline
anticorrelator Aug 26, 2026
1a2e58b
refactor(datagen): simplify trace replay
anticorrelator Aug 26, 2026
585b13a
feat(datagen): replace generation runs with recorder fixtures
anticorrelator Aug 26, 2026
e8df019
refactor(datagen): record archetypes from fixed fixtures
anticorrelator Aug 26, 2026
b539c4f
refactor(datagen): align deployment with corpus replay
anticorrelator Aug 26, 2026
77d583c
fix(datagen): satisfy repository type checks
anticorrelator Aug 26, 2026
374c4d7
feat(datagen): add recorder condition materialization
anticorrelator Aug 26, 2026
7af90dc
feat(datagen): add conditioned live recording lane
anticorrelator Aug 27, 2026
6e6f676
fix(datagen): skip llama-index recorder test when instrumenter is absent
anticorrelator Aug 27, 2026
b9021f4
fix(datagen): skip guardrail recorder test when framework is absent
anticorrelator Aug 27, 2026
e548e13
test(datagen): trim suite to one happy path per surface
anticorrelator Aug 27, 2026
bed50ae
feat(datagen): add iterative coding tool traces
anticorrelator Aug 27, 2026
4453d11
feat(datagen): enrich authored corpus inputs
anticorrelator Aug 27, 2026
58fdefb
feat(datagen): report corpus depth statistics
anticorrelator Aug 27, 2026
023cb18
feat(datagen): simulate live chat follow-up users
anticorrelator Aug 27, 2026
095c6c3
fix(datagen): suppress simulated user spans
anticorrelator Aug 27, 2026
90bee3c
fix(datagen): resolve luna recorder model
anticorrelator Aug 27, 2026
ce47250
fix(datagen): configure luna tool calls
anticorrelator Aug 27, 2026
0dc7ac4
feat(datagen): vary simulated user dispositions
anticorrelator Aug 27, 2026
53d8791
feat(datagen): add manual agent phase spans
anticorrelator Aug 27, 2026
4982309
Weight replay session sampling by fragment count
anticorrelator Aug 27, 2026
7c35481
Add a fat-tail slow-span outlier to replay jitter
anticorrelator Aug 27, 2026
3566548
docs: replace internal vocabulary with plain terms
anticorrelator Aug 27, 2026
caaa985
Fix CI: formatting, redundant cast, and datagen script type checking
anticorrelator Aug 27, 2026
27b1348
Fix datagen container start commands for the distroless image
anticorrelator Aug 27, 2026
f3b7e26
Prefix replayed session ids with their domain
anticorrelator Aug 27, 2026
7f53499
Give each archetype its own session-length profile
anticorrelator Aug 27, 2026
bdec7ab
Steer conversation length organically and diversify coding seeds
anticorrelator Aug 27, 2026
5d519fe
Let chat sessions chain a few whole conversations
anticorrelator Aug 27, 2026
738000a
Apply ruff formatting to datagen recorder and test
anticorrelator Aug 27, 2026
1cad4d0
Vary chat conversation openings per live run
anticorrelator Aug 27, 2026
17e561d
Merge remote-tracking branch 'origin/main' into dustin/data-generatio…
anticorrelator Aug 27, 2026
5e2e4d1
Restore scripts/ in the unit-test checkout and pin the chat test opening
anticorrelator Aug 27, 2026
cc7225d
Scope the scripted coding-agent test to fixtures with scripted episodes
anticorrelator Aug 27, 2026
f224791
Format the coding-agent test selection
anticorrelator Aug 27, 2026
57dfd31
Move datagen tooling tests out of the unit suite
anticorrelator Aug 27, 2026
c1d3edf
Narrow this PR to the datagen replayer runtime
anticorrelator Aug 28, 2026
a16a356
Add the datagen corpus generation tooling
anticorrelator Aug 28, 2026
64a7025
Describe the tooling test scope in the datagen README
anticorrelator Aug 28, 2026
6bba085
fix(ci): cap pydantic-ai-slim below 2.34 in unit test requirements
anticorrelator Aug 28, 2026
1ea2018
Keep jittered token totals consistent when one component is missing
anticorrelator Aug 28, 2026
4055fd6
Merge remote-tracking branch 'origin/main' into dustin/data-generatio…
anticorrelator Aug 28, 2026
e35a25d
Drop the pydantic-ai-slim unit-test cap after the vendored re-sync
anticorrelator Aug 28, 2026
d5a070b
Inline PHOENIX_CLIENT_HEADERS parsing in the datagen command
anticorrelator Aug 28, 2026
7b4f050
Merge branch 'dustin/data-generation-sidecar' into dustin/datagen-gen…
anticorrelator Aug 28, 2026
e751dd3
Mark datagen as internal tooling and move it under experimental
anticorrelator Aug 28, 2026
bfd6cd5
Merge branch 'dustin/data-generation-sidecar' into dustin/datagen-gen…
anticorrelator Aug 28, 2026
cd16e54
Follow the datagen package move to phoenix.experimental
anticorrelator Aug 28, 2026
ac83864
Match the datagen tooling job's uv pin to the repo-wide bump
anticorrelator Aug 28, 2026
d002600
Merge remote-tracking branch 'origin/main' into dustin/datagen-genera…
anticorrelator Aug 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions .github/workflows/python-CI.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ jobs:
uv_lock: ${{ steps.filter.outputs.uv_lock }}
json_canonicalization_schema: ${{ steps.filter.outputs.json_canonicalization_schema }}
filter_dsl: ${{ steps.filter.outputs.filter_dsl }}
datagen_tooling: ${{ steps.filter.outputs.datagen_tooling }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
Expand Down Expand Up @@ -83,6 +84,9 @@ jobs:
- "js/app/src/pages/project/sessionFilterDSL.ts"
- "src/phoenix/trace/dsl/**"
- "scripts/ci/check_filter_dsl_snippets.py"
datagen_tooling:
- "scripts/datagen/**"
- "src/phoenix/experimental/datagen/**"
- name: Print Filters
env:
IPYNB: ${{ steps.filter.outputs.ipynb }}
Expand Down Expand Up @@ -563,6 +567,35 @@ jobs:
timeout-minutes: 60
run: uvx tox run -e unit_tests -- -ra --reruns 5 --db postgresql -n 16 --dist loadscope --postgresql-exec /usr/lib/postgresql/14/bin/pg_ctl

datagen-tooling-tests:
name: Datagen Tooling Tests
runs-on: ubuntu-latest
needs: changes
if: ${{ needs.changes.outputs.datagen_tooling == 'true' && github.event_name == 'pull_request' }}
steps:
- name: Checkout repository
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
sparse-checkout: |
requirements/
scripts/datagen/
src/
packages/
evals/
- name: Set up Python
uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0
with:
python-version: "3.10"
- name: Set up `uv`
uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7.6.0
with:
version: "0.12.5"
- name: Sync dependencies
run: uv sync --frozen
- name: Run datagen tooling tests
run: uv run pytest scripts/datagen/tests -ra

integration-tests:
name: Integration Tests
runs-on: ${{ matrix.os }}
Expand Down Expand Up @@ -749,6 +782,7 @@ jobs:
- check-lockfile
- type-check
- unit-tests
- datagen-tooling-tests
- integration-tests
- test-migrations
- test-json-canonicalization-schema
Expand Down
205 changes: 205 additions & 0 deletions scripts/datagen/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,205 @@
# Trace corpus recorders

These scripts record application traffic through real OpenInference instrumenters.
The resulting corpus contains raw OTLP protobuf JSON requests and the fragment rows used by the
Phoenix datagen composer. A **fragment** is one replayable unit — a conversation turn or an agent
episode — pointing at the recorded traces it produced. Recording frameworks remain outside Phoenix
runtime dependencies.

## Fixed inputs

`recorder_fixtures.json` contains the application inputs for every retained recorder:

- a stable fragment ID;
- an archetype (which kind of application produced the trace — plain chat, RAG, tool agent,
graph agent, guardrails, or structured extraction) and a domain (its subject area, such as
customer support or coding);
- direct prompts, turns, documents, or expected structured values.

The fixtures contain no sampling weights or generated text. Each recorder receives a
`RecorderFixture`, appends OTLP requests to `traces.jsonl`, and returns the trace IDs it emitted.
`record_fixture` then appends the matching fragment row (`fragment_id`, `archetype`, `domain`,
`trace_ids`) to `fragments.jsonl`.

The fixture set includes multiple examples for plain chat, RAG, tool agents, graph agents,
guardrails, and structured extraction. Success, blocked, redacted, conflicting-source, and
incomplete-input examples are represented directly in the app inputs.

Tool-agent fixtures may carry `prompt_variants`: alternative phrasings of the opening prompt.
Live recording picks one phrasing per run, so repeated runs of the same task do not open with
identical text. Coding fixtures without a deterministic scripted episode are skipped by
scripted auto-selection and record live only.

## Offline providers and tools

`ScriptedOpenAIProvider` serves a fixed sequence of text, tool-call, or HTTP responses through an
in-process `httpx` transport. It supports buffered and streaming chat completions without a network
connection or API key.

`local_tools` exposes deterministic document search, record lookup, status lookup, arithmetic,
and ticket creation over `tool_fixtures.json`. The file contains separate customer-support,
analytics, and coding data sets.

## Generate varied recordings

`organic_conditions.json` defines authored input variations: degraded versions of the base
fixture inputs that let response-quality issues arise naturally rather than by script. Each
condition names a base fixture, a unique output fragment ID, an intensity, and one payload per
intensity level. Document edits, replacements at existing fixture-input paths, and
matched local-tool result overlays are applied before the application runs. Input replacements
cannot add or remove structure. Keep every condition fragment ID distinct from the IDs in
`recorder_fixtures.json` and from other conditions. Intensity selects the level:

| Intensity | Level |
| --------- | ----- |
| below 0.2 | `subtle` |
| below 0.5 | `moderate` |
| 0.5 and above | `strong` |

All recorder commands accept `--condition` and `--append`. With neither flag, a recorder uses its
fixed fixtures and resets the output directory. `--condition` runs the one fixture with that
condition's edits applied; `--append` preserves existing rows so multiple conditions and
recorders can share a recording directory.

Plain chat, RAG, tool agent, and structured extraction also accept `--provider scripted|live`.
Scripted is the default. Live recording requires an explicit `--model`, reads `OPENAI_API_KEY`, and
uses `OPENAI_BASE_URL` when it is set. For example:

```console
export OPENAI_API_KEY="..."
uv run --script scripts/datagen/tool_agent.py \
--output-dir dist/datagen/recording \
--condition support-stale-delivery-status \
--provider live \
--model gpt-5.4 \
--append
```

Graph and guardrail recorders are deterministic applications and therefore expose condition and
append controls without provider or model options.

Live plain-chat conversations run until the simulated user closes them. A target turn count
(drawn per fixture, or set with `--target-turns`) controls when the simulator is told to wrap
up once its current concern is addressed; the conversation ends at that natural closing
message, with a hard cap at twice the target. Live model aliases: `luna` and `terra` resolve
to their provider model IDs with tool-calling options applied.

A live run records every instrumented invocation that emits trace IDs. Responses are not compared
with fixture-authored answers, and incomplete responses or traced application errors are retained.
A run fails only when it emits no trace IDs. Review or evaluate quality after recording; keep
ambiguous outcomes in the recorded set.

### Recording playbook (all recorders, one directory)

Choose one live model for the batch and run these commands in order. The first command starts a new
recording and captures every scripted tool fixture, including the longer coding sessions. Each later
command appends either all base fixtures for one recorder or its authored condition. The resulting
recording contains more than two dozen fragments across all six archetypes.

```console
export DATAGEN_MODEL="gpt-5-mini"

uv run --script scripts/datagen/tool_agent.py \
--output-dir dist/datagen/recording

uv run --script scripts/datagen/openai_chat_sessions.py \
--output-dir dist/datagen/recording \
--provider live --model "$DATAGEN_MODEL" --append
uv run --script scripts/datagen/openai_chat_sessions.py \
--output-dir dist/datagen/recording \
--condition support-late-express-ambiguity \
--provider live --model "$DATAGEN_MODEL" --append

uv run --script scripts/datagen/llama_index_rag.py \
--output-dir dist/datagen/recording \
--provider live --model "$DATAGEN_MODEL" --append
uv run --script scripts/datagen/llama_index_rag.py \
--output-dir dist/datagen/recording \
--condition research-fleet-delivery-pressure \
--provider live --model "$DATAGEN_MODEL" --append

uv run --script scripts/datagen/tool_agent.py \
--output-dir dist/datagen/recording \
--provider live --model "$DATAGEN_MODEL" --append
uv run --script scripts/datagen/tool_agent.py \
--output-dir dist/datagen/recording \
--condition support-stale-delivery-status \
--provider live --model "$DATAGEN_MODEL" --append

uv run --script scripts/datagen/structured_extraction.py \
--output-dir dist/datagen/recording \
--provider live --model "$DATAGEN_MODEL" --append
uv run --script scripts/datagen/structured_extraction.py \
--output-dir dist/datagen/recording \
--condition analytics-corrected-refund-export \
--provider live --model "$DATAGEN_MODEL" --append

uv run --script scripts/datagen/graph_multi_agent.py \
--output-dir dist/datagen/recording --append
uv run --script scripts/datagen/guardrailed_app.py \
--output-dir dist/datagen/recording --append
```

Keep every command result that reports trace IDs, including responses that are incomplete,
ambiguous, or accompanied by a traced application error. If a command reports no trace IDs, fix
that recorder before continuing so later `--append` calls do not hide the missing fragment.

Anyone — or any coding agent — running this playbook can choose the conditions, models, run
count, and command order. To use Codex with ChatGPT subscription access, authenticate once and ask
the non-interactive command to inspect the playbook and invoke recorder commands:

```console
codex login
codex exec 'Read scripts/datagen/README.md and scripts/datagen/organic_conditions.json. Choose varied conditions and models, run the applicable recorders into dist/datagen/recording with --append, and retain every run that emits trace IDs.'
```

`codex exec` chooses and runs commands in this workflow; it is not a provider implemented by the
recorders. Direct `--provider live` recorder calls use API credentials from the environment.

## Recorder environments

Every framework recorder has a PEP 723 dependency block and must be run with `uv run --script`.
This keeps recorder dependencies out of the Phoenix package and pins the instrumenter stack used
to create stored traces. Shared modules imported by those entry points use only the Python standard
library unless their dependency is declared in every importing script.

Each JSONL line in `traces.jsonl` is one protobuf-JSON `ExportTraceServiceRequest`. A single trace
may span multiple rows.

Tests for the packaging pipeline (conditions, packer, fetcher roundtrip) live in `tests/` next to
this file and run with `uv run pytest scripts/datagen/tests`. They are separate from the Phoenix
unit test suite: CI runs them in the Datagen Tooling Tests job when files under `scripts/datagen/`
or `src/phoenix/experimental/datagen/` change. Recorder behavior has no unit tests; verify recorders by
generating a corpus.

## Package a corpus

After all selected fixtures and conditions have been recorded into one directory, package the
recording. The printed statistics include `opening_diversity_by_domain` — distinct opening
inputs per domain — so low seed variety is visible before publication:

```console
uv run python -m scripts.datagen.corpus <recording-dir> \
--archive dist/datagen/corpus.tar.gz
```

The archive contains only `fragments.jsonl` and `traces.jsonl`.

Validate or stage the archive for manual publication:

```console
uv run python -m scripts.datagen.publish validate \
--archive dist/datagen/corpus.tar.gz

uv run python -m scripts.datagen.publish prepare-archive \
--archive dist/datagen/corpus.tar.gz \
--output-dir dist/datagen-publication
```

Preparation prints the exact commands for uploading the digest-addressed archive first and the
public pointer second.

## Freshness

Re-record and review the corpus whenever a pinned instrumenter version changes. This keeps stored
span shapes aligned with the framework and instrumenter versions declared by each recorder.
Loading
Loading