Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
245 changes: 245 additions & 0 deletions skills/data-loading-bottleneck/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,245 @@
---
Comment thread
rostan-t marked this conversation as resolved.
name: data-loading-bottleneck
description: "Diagnose input-bound PyTorch training. Use for low or bursty GPU utilization, slow batches, num_workers tuning, preprocessing regressions, or input stalls. Not for model/kernel optimization."
Comment thread
mdabek-nvidia marked this conversation as resolved.
license: Apache-2.0
compatibility: Requires Linux, Python, Git, an NVIDIA GPU and driver, and CUDA-enabled PyTorch.
permissions:
- file_read
- file_write
- shell
- network
- env
metadata:
author: "DALI Team <dali-team@nvidia.com>"
tags:
- pytorch
- training
- performance
- data-loading
- profiling
- dali
languages:
- python
team: dali
domain: deep-learning
version: "1.0.0"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is the intent to update version number manually? Can it re-use DALI versioning ?

@rostan-t rostan-t Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since skills are distributed independently from DALI (consumed by NVIDIA/skills), I wouldn't be sure what version to use here. Should it be 2.3.0 or the current version in VERSION 2.4.0dev? Do we open a PR to update the skill versions each time we update DALI?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are right, that since it is distributed independently it would be difficult to share the versioning. I think that we need to keep in mind, that the skill itself may need to be updated after changes in DALI. I would feel safer knowing what is minimal version of DALI that would be supported by the skill.

---

# Data Loading Bottleneck

## Purpose

Determine whether PyTorch training is input-bound and, if so, localize one cause without
changing the production path, lifecycle, data flow, or topology.

## Prerequisites

Requires Linux, Git, CUDA, and PyTorch. DALI replay and Nsight/NVTX profiling are optional.
Preflight attempts their allowed setup and routes around anything unavailable.

## Instructions

### 1. Preflight and route

Create an artifact directory for commands and raw output, then run preflight with the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The skill assumes that the platform that we run the preflight is configured so that it can run training.
Would it be possible to extend it to create virtual environment and execute pip install -f requiremnets.txt if such file exist?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's generally safe to assume that somebody interested in evaluating if training performance is bottlenecked by data loading has a working environment. We want to use this environment with the exact pinned library versions they are using. Those might not be the same as a potential requirements.txt file or it might not tell the full picture because it's not uncommon for projects to have other files like requirements.dev.txt.

Additionally, supporting requirements.txt only would be incomplete, pyproject.toml is also often used (with e.g. uv or poetry), some projects have a conda configuration instead, others rely on Docker containers, etc.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's generally safe to assume that somebody interested in evaluating if training performance is bottlenecked by data loading has a working environment.

This is bold assumption. There are cases where production environment is isolated from development environment with running agent. I think that adding this assumption (or asking a user to point to an active environment) to the skill would make it more user friendly.

production Python:

```bash
<production-python> <skill-dir>/scripts/collect_preflight.py \
--source-dir <checkout> --artifact-dir <artifact-dir> \
--data-path <local-dataset> \
[--expected-visible-gpus N]
```

For remote or custom input, use `--data-source <description>` instead of `--data-path`.
Preflight records only the description, while the production run validates access. Stop on a
hard blocker. If a sandbox hides the accelerator, rerun preflight and GPU work in production.
CPU execution is not a substitute. Record any environment change made in response to a
warning.

Start with the replay-support result from preflight. If replay is unavailable, use
`torch.version.cuda` to choose one package for a single isolated installation attempt:

- CUDA 12.x: `nvidia-dali-cuda120`
- CUDA 13.x: `nvidia-dali-cuda130`
- Other or unknown: record the unsupported runtime and skip replay.

```bash
Comment thread
mdabek-nvidia marked this conversation as resolved.
<production-python> -m pip install --target <artifact-dir>/dali-deps <dali-package>
```

After a successful installation, append the target with `site.addsitedir()` and retry
`from nvidia.dali.plugin.pytorch.loader_evaluator import LoaderEvaluator`. Keep
target-installed dependencies behind production packages. If the import succeeds, use the
same setup for Real and Replay. If it fails, preserve the failure, leave production unchanged,
and profile after Real. Do not try another installation strategy. Use the production Python
and worktree `PYTHONPATH` for every run.

After preflight, record the original checkout status and diff, then create a disposable Git
worktree. Reproduce the canonical code and configuration, including staged, unstaged, and
relevant untracked changes. Put its package root or `src` directory first on `PYTHONPATH`,
point out-of-tree builds to it, and record representative module `__file__` paths. Stop if it
cannot reproduce the workload. Leave the original checkout untouched. Remove the worktree
after diagnosis and keep the artifacts.

### 2. Instrument one bounded production run

Instrument the production training path in the disposable worktree. Use a standalone harness
only when that path cannot be bounded or instrumented, as described in Troubleshooting.
Locate the last blocking loader retrieval at the intended replay boundary and the complete
optimizer update that consumes its batch. Record device transfer relative to the boundary,
batch/sample accounting, distributed topology, and implicit defaults. When
`LoaderEvaluator` is available, read `references/pytorch-dali.md` and build its paired Real
and Replay loaders. Without it, use the bounded production loader directly.

Set `prefetch_depth` to the batches that can be ready at the replay boundary, including
production wrapper buffers. Per rank, it must be at least `num_workers * prefetch_factor`.
Use 2 when `prefetch_factor` is unset and one when `num_workers == 0`. Choose
`measured_batches > prefetch_depth`. Set warmup and drain to at least that depth, then set
`total_batches = warmup_batches + measured_batches + drain_batches`. Bound the source to
that total, warm the first part, time the measured window, and drain outside it.

Keep the same workload, integration boundary, and timed window through localization. Add
only semantic ranges. Preserve work-affecting batch
structure, routing, topology, synchronization, and update behavior. Time complete steps and
every blocking loader retrieval. Synchronize the device and count samples at both window
boundaries. With multiple ranks, add boundary barriers and record each rank. Include normal
wait variability. In Real and Profile, do not manipulate the page cache or replace production
input with synthetic, repeated, modified, or deliberately pre-cached data.

If instrumentation fails, follow Troubleshooting. If no valid Real window remains after
those attempts, report `INCONCLUSIVE` and continue at §6.

Runs with material changes to the data source, sampling rules, batching, preprocessing, or
work-affecting input distribution are substitutes and cannot support canonical
`NOT DETECTED`.

### 3. Run Real

In a fresh process, run the bounded window through `LoaderEvaluator(mode="log")` when
available. Otherwise, use the bounded production loader. Per rank, record measured
samples, complete-step time, exposed loader wait, and the timestamp immediately before each
window barrier. Classify the source cache state as known warm before Real, warmed only by
this run's normal access, or unknown.

```text
aggregate throughput = sum(samples across ranks) / max(rank window duration)
```

Do not infer balanced ranks from final arrival skew alone. Collectives can repeatedly
reconverge imbalanced ranks. Continue at §4 when `LoaderEvaluator` is available. Otherwise,
continue at §5.

### 4. Run Replay and classify

Start another fresh process with the same `LoaderEvaluator` wrapper in `replay` mode. Change
only the mode. Any other change invalidates the comparison. Apply the post-boundary work
equivalence and lifecycle checks in `references/pytorch-dali.md` before classification.

If either run emits fewer than `total_batches`, report `INCONCLUSIVE` and continue at §6.
For any other unavailable or invalid Replay, continue at §5.

For a valid comparison:

```text
speedup = replay aggregate throughput / real aggregate throughput
```

| speedup | Verdict |
|---|---|
| `>1.50x` | `DETECTED` |
| `>1.10x` and `<=1.50x` | `POTENTIAL` |
| `<=1.10x` | `NOT DETECTED` |

These fixed heuristics follow DALI's [Data Loading Bottleneck Detection tutorial](https://docs.nvidia.com/deeplearning/dali/user-guide/docs/examples/frameworks/pytorch/loader_evaluator/pytorch_data_loader_evaluator.html).
They are not estimates of run-to-run noise.

Do not repeat Real or Replay. Continue at §6 after `NOT DETECTED`, and at §5 after
`DETECTED` or `POTENTIAL`.

### 5. Profile and localize

When §3 or §4 routes here, read `references/profiling.md`. If Nsight Systems, NVTX, or profile
summarization is still unavailable after its allowed setup, skip capture. Keep a valid Replay
verdict and mark localization unavailable. Without valid Replay, report `INCONCLUSIVE`.
Continue at §6.

Otherwise, reuse a valid Real run as the unprofiled baseline. Follow the reference's
single-capture workflow and shared limit of one recapture for any reason.

Apply the reference's structural and Real-baseline checks to each conclusion. A structurally
invalid capture is unusable. A work mismatch restricts only the conclusions it could affect,
while a valid Replay verdict remains authoritative. Without valid Replay, report `DETECTED`
only when validated profile evidence shows an input-path stage materially delaying full-step
progress in the measured window. Otherwise, report `INCONCLUSIVE`.

Use that delay as detection evidence. Attribute a cause only to a measured concrete stage or
supported capacity limit. Leave an unsplit limiting stage as `unresolved composite`. When
localization supports an optimization, give one ranked, evidence-backed recommendation,
mark it untested, and do not implement or benchmark it. Otherwise, state the missing evidence.

### 6. Report

Complete `assets/report-template.md` and return it in the final response, *not* as a path to a
Markdown file. Remove unused sections and placeholders. Always keep the decision table,
Workload, Detection, and Confidence and scope. Keep Workload to one paragraph when the
canonical command ran as-is.

- For `DETECTED` or `POTENTIAL`, give Cause and Next action. Include Localization when
profiling ran and Recommendation when supported. Without localization, set Cause to
`unresolved composite`.
- For `NOT DETECTED`, set Cause and Next action to `not applicable`, then omit Localization
and Recommendation.
- For `INCONCLUSIVE`, set Cause to `unresolved composite`, name the missing evidence under
Next action, and add Localization only when profiling ran.

Add Substitutions only when the measured run differed from canonical. Add Distributed
behavior and Trace navigation only for `world_size > 1`. Use Real timing for every rank,
profiled GPU or NCCL values only for captured ranks, and map profiled PIDs to ranks. Mark
missing measurements `invalid` or `unavailable`. Report excluded input costs in Detection
with their value and frequency, never as Cause. A recommendation must cite its measurement,
mechanism, feasibility constraint, and `untested` status.

Under Confidence and scope, keep only facts that could change the interpretation. Include
invalid and superseded artifacts with their rejection reasons. Use absolute artifact paths,
including the full `.nsys-rep` path. Cite the evidence for the verdict and recommendation,
and limit the result to the measured workload and recorded source cache state.

## Available Scripts

| Script | Purpose | Arguments |
|---|---|---|
| `scripts/collect_preflight.py` | Record environment, source, data, and optional-tool readiness in `preflight.json` | `--source-dir`, `--artifact-dir`, one of `--data-path` or `--data-source`; optional `--expected-visible-gpus` |
| `scripts/summarize_nsys.py` | Summarize the measured NVTX window, CUDA activity, and loader-wait/GPU-idle overlap | input `.nsys-rep` and required `--output` JSON path |
Comment thread
mdabek-nvidia marked this conversation as resolved.

Both scripts write to the requested path, print status to stdout, and send diagnostics to
stderr. Preflight returns `0` when ready, `1` when blocked, and `2` on an operational or
usage error. The summarizer returns `0` on success, `1` for an invalid report or output, and
`2` for invalid arguments. Where a `run_script` helper exists, use it with the same
repo-relative script and arguments. Otherwise, use the documented production-Python commands.

## Examples

- “Find out whether data loading is causing bursty GPU utilization in this PyTorch training
job, and identify the limiting stage.”
- “Optimize this model's attention kernels” is outside this skill's scope.

## Limitations

- Covers steady-state CUDA-enabled PyTorch training. Inference, startup, epoch transitions,
checkpointing, and offline data preparation are outside the workflow.
- Determines whether the input path limits full-step training throughput. It does not
optimize model, kernel, optimizer, or communication performance.
- Produces workload-specific conclusions. It does not predict behavior under another
topology, data path, or cache state.

## Troubleshooting

- **Instrumented run failure:** rerun the canonical command without instrumentation and save
both commands and outputs to separate workload failure from instrumentation failure.
- **Resource failure:** make at most one nearest runnable attempt that changes one resource
setting. Label it a substitute and restrict the verdict accordingly.
- **Production path cannot be bounded or instrumented:** save the failed production attempt
before using a standalone harness. Treat the harness as a substitute. State what differs or
is missing and limit every conclusion that depends on it. Without the complete training
step, it cannot establish that production is input-bound.
4 changes: 4 additions & 0 deletions skills/data-loading-bottleneck/agents/openai.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
interface:
display_name: "Data Loading Bottleneck"
short_description: "Measure, profile, and diagnose training input stalls"
default_prompt: "Use $data-loading-bottleneck to check whether this training workload is input-bound."
93 changes: 93 additions & 0 deletions skills/data-loading-bottleneck/assets/report-template.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# Data-loading result

| Decision | Result |
|---|---|
| **Verdict** | **[DETECTED / POTENTIAL / NOT DETECTED / INCONCLUSIVE]** |
| **Cause** | [Supported concrete stage or capacity limit, or `unresolved composite`] |
| **Next action** | [One evidence-backed recommendation, or the evidence needed to resolve the result] |

## Workload

[Command, model, settings, environment, dataset source, format, cardinality, and cache state]

**Substitutions**

| Change | Why it was required | Expected bias and verdict scope |
|---|---|---|
| [canonical -> measured] | [reason] | [effect and where the conclusion applies] |

## Detection

| Run | Full-step window | Throughput | Exposed loader wait |
|---|---:|---:|---:|
| Real | [value, invalid, or unavailable] | [value, invalid, or unavailable] | [value and percent, invalid, or unavailable] |
| Replay | [value, invalid, or unavailable] | [value, invalid, or unavailable] | [value and percent, invalid, or unavailable] |

**Result:** [Measured speedup, wait-only prediction from Real, their difference, and threshold
Comment thread
mdabek-nvidia marked this conversation as resolved.
interpretation. If Replay is invalid or unavailable, report the primary effect]

**Replay boundary:** [What replay bypassed and what remained, or why it was unavailable]

**Distributed behavior**

| Rank | Samples | Full-step window | Exposed loader wait | Pre-barrier timestamp | Useful GPU / NCCL |
|---:|---:|---:|---:|---:|---:|
| [rank] | [value] | [value] | [value and percent] | [value] | [value / value or unavailable] |

**Rank result:** [Slowest-window aggregation, wait and arrival skew, and whether collectives
mask starvation]

## Localization

**Primary Nsight Systems report**

```text
/absolute/path/to/trace.nsys-rep
```

**Profile summary**

```text
/absolute/path/to/profile-summary.json
```

**Trace navigation**

| Domain | PID(s) | Range names |
|---|---|---|
| [domain] | [PID list] | [names] |

**GPU timeline**

- [Active and idle result]
- [Transfer or other critical-path result]

**End-to-end attribution**

| Stage (main / worker / GPU) | Status | P50 | Critical-path evidence | Artifact |
|---|---|---:|---|---|
| [stage] | [measured / absent / composite / unavailable] | [value, N, and unit] | [dependency, observed overlap and denominator, or why unavailable] | [summary or timeline] |

**Cause result:** [Supported actionable cause, unresolved composite and limit, or missing
evidence.]

## Recommendation

[Recommendation]

## Confidence and scope

- **Equivalence:** [Real <=> Replay: `pass` (`exact` or `work-equivalent`), `fail`, or
`unavailable`. Real <=> Profile: give one of those statuses for each conclusion. Cite the
relevant identities or batch signatures, steps, lifecycle, operating-regime evidence, and
any restrictions]
- **Missing coverage:** [Main-process ranges or worker PIDs without ranges]
- **Profiler perturbation:** [Profiled versus unprofiled difference, or unavailable]
- **Limits:** [For INCONCLUSIVE, name the evidence needed]

**Artifacts**

| Absolute path | Status | Purpose / reason |
|---|---|---|
| [/path/to/primary-artifact] | primary | [detection or diagnosis] |
| [/path/to/artifact] | [supporting / restricted / topology only / superseded / invalid] | [purpose, restriction, or rejection reason] |
16 changes: 16 additions & 0 deletions skills/data-loading-bottleneck/evals/config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
schema_version: 1

harbor:
task_source: evals_json
custom_dockerfile_mode: preserve
base_image_mode: disabled
n_attempts: 1
n_concurrent: 1
max_agents: 1
stop_on_pass: false
timeout_multiplier: 6
auto_scale_timeout: true
sandbox:
template: harbor-eval-claude-code-gpu
pre_agent_setup:
- cp -a /opt/eval/workloads /workspace/workloads
Loading
Loading