The control plane your coding agent uses to run physical-AI workloads on Nebius.
Quickstart · Guides · Workbench docs · Benchmarks · Operator tools · CLI reference · Python & API · Cookbooks · Contributing
For shared Kubernetes clusters, use team namespaces to configure
namespace selection and private SkyPilot contexts with npa workbench namespace.
Workbench is the agent-facing control plane for physical AI on Nebius. You
describe the outcome; Codex, Claude Code, or another coding agent with terminal
access to this checkout uses npa to configure the project, check access, plan
resources, run containerized tools and workflows, and inspect the result with
you. You can drive the same operations directly through the CLI, and supported
tools also expose Python interfaces.
%%{init: {"theme": "base", "fontFamily": "Arial, sans-serif", "themeVariables": {"fontFamily": "Arial, sans-serif", "fontSize": "16px", "primaryColor": "#E0FF4F", "primaryTextColor": "#052B42", "primaryBorderColor": "#052B42", "lineColor": "#526575", "edgeLabelBackground": "#FFFFFF", "clusterBkg": "#EEF3FF", "clusterBorder": "#B4C9DC", "clusterTextColor": "#052B42", "tertiaryTextColor": "#052B42"}, "flowchart": {"htmlLabels": true, "curve": "linear", "nodeSpacing": 32, "rankSpacing": 48}}}%%
flowchart TB
accTitle: How Workbench runs a task on Nebius
accDescr: You use a coding agent or the CLI and supported Python interfaces to control npa. Workflow YAML drives the workflow engine and SkyPilot, which launch containerized GPU and CPU jobs on Nebius Kubernetes. Tools read and write S3 artifacts; npa reads and writes durable run state. Token Factory is a separate hosted inference API. Results return through npa for inspection.
access["You + your coding agent<br/>or direct CLI / supported Python interfaces"]
control["npa control plane<br/>Configure · preflight · plan · submit · inspect"]
spec["npa.workflow YAML<br/>States · toolRefs · resource profiles"]
workflow["Workflow engine + SkyPilot<br/>Plan waves · submit jobs · track status"]
access <-->|"requests / results"| control
spec --> control
control <-->|"submit / monitor / resume"| workflow
subgraph nebius["Nebius AI Cloud"]
tools["Containerized tools · Kubernetes<br/>GPU + CPU<br/>Simulate · train · generate<br/>Curate · reconstruct · evaluate"]
storage["Object Storage · S3<br/>Inputs · checkpoints · media<br/>Reports · durable run state"]
hosted["Token Factory<br/>Hosted inference API<br/>No customer GPU required"]
tools <-->|"read / write"| storage
end
workflow <-->|"jobs / status / logs"| tools
control <-->|"direct API"| hosted
control <-.->|"artifacts / durable state"| storage
classDef access fill:#FFFFFF,stroke:#B4C9DC,color:#052B42,stroke-width:1.5px,rx:12,ry:12;
classDef control fill:#E0FF4F,stroke:#052B42,color:#052B42,stroke-width:2px,rx:12,ry:12;
classDef orchestration fill:#EEF3FF,stroke:#052B42,color:#052B42,stroke-width:1.5px,rx:12,ry:12;
classDef service fill:#FFFFFF,stroke:#052B42,color:#052B42,stroke-width:1.5px,rx:12,ry:12;
classDef data fill:#052B42,stroke:#052B42,color:#FFFFFF,stroke-width:1.5px,rx:12,ry:12;
class access,spec access;
class control control;
class workflow orchestration;
class tools,hosted service;
class storage data;
The architecture overview maps selected NVIDIA and open-ecosystem solutions to
NPA's public GHCR container catalog
at ghcr.io/nebius/nebius-physical-ai and the Nebius services that run and support
them. Some images fetch models or vendor runtimes at execution time; the
packaging contract records those boundaries.
The task flow above shows the standard Kubernetes workflow path. The workflow engine plans execution waves, submits them through SkyPilot, and uses S3 artifacts and durable run state to evaluate decisions and resume runs. The operator-side SkyPilot environment submits the jobs; the workload containers run on Nebius GPU or CPU nodes.
Token Factory is a separate hosted inference API. Direct calls need no cluster; workflow stages can also call it from their containers. Individual tools additionally support the runtime modes documented for that tool. Inspect the resulting artifacts with Rerun, Foxglove, media viewers, or reports. You remain in the loop for choices and human-bound approvals such as accepting model terms. The diagram reference records the scope and source files behind the visuals.
Start with a workload guide. The guide tells you which data, model access, GPU, and output to expect.
Using a coding agent? Open this checkout in your agent and paste one of the maintained prompts:
The prompts keep credentials private, stop at human approval gates, validate the plan before provisioning, and stay with the run through artifact inspection. The manual path below exposes the same control-plane steps.
Use Python 3.10+ on macOS, Linux, or WSL2 Ubuntu. Clone the repository, then run the remaining commands from its root:
git clone https://github.com/nebius/nebius-physical-ai.git
cd nebius-physical-ai
python3 -m venv .venv
source .venv/bin/activate
pip install -e npa
npa --versionYou should see the installed npa version. Remote GPU workloads use their
container dependencies; you do not need CUDA or a local GPU to use the CLI.
Installation covers platform setup and optional dependencies.
Install the tested Nebius CLI and configure your project:
curl -fsSL https://storage.eu-north1.nebius.cloud/cli/install.sh \
| NEBIUS_CLI_VERSION=0.12.254 bash
export PATH="${HOME}/.nebius/bin:${PATH}"
npa configure
npa workbench health preflight --checks nebius --jsonThe preflight should report working Nebius authentication. It does not reserve
GPUs or verify model access. Interactive configuration provisions object storage
by default and needs project admin permission for that setup. Use
npa configure --no-provision to save settings without creating storage.
See configuration for SSO, existing projects, and tokens.
Check a project's saved storage with its configured alias in place of <alias>:
npa workbench health preflight --project '<alias>' --checks s3,nebius --jsonMissing or invalid project storage fails without borrowing shell or host-file
credentials or changing configuration. S3 readiness uses a bounded listing call;
it does not enumerate the dataset or verify write access. Without --project,
the S3 check keeps the existing host credential selection.
Choose one guide below and follow it through input preparation, GPU setup, submission, and output inspection. Workbench setup provides the shared cluster and SkyPilot steps; complete them once per project. Keep the returned run ID: you use it to inspect logs, find artifacts, and resume.
For video augmentation, the PAIDF + Cosmos 3 runbook includes a public source-video example and full submission commands.
| Result | Guide | Main requirement |
|---|---|---|
| Generated images or video | Cosmos 3 generation | Compatible GPU and model access |
| Augmented video with evaluation and curation reports | PAIDF + Cosmos 3 | Source video, GPU, S3, hosted inference |
| A labeled video dataset | Physical AI Data Factory | Source video, RT-core GPU, S3 |
| A reconstructed 3D scene and novel views | Neural reconstruction | Sensor capture and RT-core GPU |
| A robot-policy training checkpoint | Reachy 2 + LeRobot | Matching recorded dataset and GPU |
| A Franka training and evaluation exercise | Franka + Genesis | GPU; recorded run did not solve the task |
| Quadruped reinforcement learning | Isaac Lab | RT-core GPU |
| A G1 locomotion evaluation or training path | G1 + SONIC | Runtime-specific GPU and checkpoints |
| A browser workbench and artifact viewer | Deploy the agent | Terraform, S3, SSH key, Token Factory key |
The guide index records validation scope. Cookbooks cover longer training and data pipelines. GPU names alone do not establish image compatibility; check the image/GPU matrix before provisioning.
A workflow is a YAML state graph: each toolRef selects a Workbench operation;
outputs become inputs to later stages. SkyPilot schedules the work on Nebius.
These commands validate and plan a checked-in example locally:
npa workbench workflow validate-spec workflows/testing/cosmos3-generate.yaml
npa workbench workflow plan-spec workflows/testing/cosmos3-generate.yaml --run-id demoExpect a valid specification and a plan containing the generate stage.
The example's storage values are placeholders. Planning creates no workload and
does not prove its inputs, credentials, image, or GPU are ready for execution.
For a real run, follow the selected guide's prepare-run, image preflight,
submit --runtime, and monitoring instructions with your own project and input.
See the workflow catalog,
authoring guide, and
run lifecycle. The canonical
14-stage Sim2Real pipeline uses this
same runtime at workflows/main/sim2real.yaml
and requires its own prepared images, task data, and resource profiles.
Its data contracts preserve
simulator episode resets across sparse samples and exclude reset intervals from
training credit. The older sim2real/runbook.yaml is a legacy path.
| Task | Reference |
|---|---|
| Discover tools by task | Workbench docs |
| Use FA4 in your own RTX PRO 6000 container | RTX PRO 6000 FA4 adoption guide |
| Compare standalone FA2 and tuned FA4 | RTX PRO 6000 measurements and actual renders |
| Find a command or option | CLI index, then npa workbench <tool> --help |
| Call a tool from Python or HTTP | CLI / SDK walkthrough — supported interfaces vary by tool |
| Develop a native Ray application | Ray guide |
| Inspect or share run outputs | Rerun · Foxglove/MCAP |
| Diagnose a failed run | Troubleshooting |
Cancel active jobs before deleting their infrastructure. Preview project cleanup:
npa destroy --project "<alias>" --allThis is read-only until you pass --yes. Review the listed resources and save
needed artifacts before deleting storage. Teardown covers
service, controller, cluster, and storage cleanup. Local npa cleanup alone
does not stop cloud resources.
Use the public image catalog to pull
available GHCR images. Repository-owned runtime defaults already select GHCR;
npa configure needs no registry setup. Some images fetch separately licensed
runtimes or model weights at startup, so check their access requirements too.
For modified or private images, follow the
build and packaging guide and select
the image explicitly. NPA_REGISTRY alone does not change runtime defaults.
The image/GPU matrix records hardware tests and distinguishes them from expected compatibility. Each guide states what its run established: a working container, a generated artifact, a training checkpoint, or measured task performance. A smoke test is not evidence of policy convergence. See also the LeRobot benchmarks and planned partner capabilities.
| Directory | Contents |
|---|---|
npa/ |
Python package, CLI, SDK, and tests |
workflows/ |
Declarative workflow catalog and runbooks |
docs/ |
Setup, tool guides, cookbooks, and references |
deploy/cluster/ |
Managed Kubernetes Terraform wrapper |
skills/ |
Operating and contribution instructions for agents |
workbench/mlflow/ |
Local MLflow and Postgres stack |
research/ |
Historical standalone deployment research |
Use the documentation index to find setup, operations, and
contributor references. Report a broken example with its command, npa version,
and redacted error in GitHub Issues.
Keep credentials and private infrastructure identifiers out of issue text.
For dedicated CI capacity, operators can set the repository Actions variables
NPA_CI_PRIORITY_RUNNER and NPA_CI_TEST_RUNNER to approved Ubuntu runner labels.
Both default to ubuntu-latest; neither reserves capacity by itself. See
validation concurrency for routing and setup.
The temporary CPU runner guide covers disposable
Nebius workers. make ci-runners-down restores routing and safely drains the
configured pool; make ci-runners-status reports its state.
See CONTRIBUTING.md for the development environment, required
checks, and PR process. The package README
has the shortest test commands. Update the relevant documentation and
root skill when changing behavior.
Run make precheck for fast local CI checks. After committing, fetch main and run
make merge-precheck to check the combined dependency inputs. The
merge-readiness guide
also explains the automatic PR comments for merge-queue rejections.
Security disclosures: SECURITY.md.
Apache License 2.0. Third-party software, models, and datasets retain their own licenses and access terms.
MJLab integration provides native training and resume,
measured evaluation, ONNX export, a scoped authenticated service, and CLI/SDK
clients. Use the train/evaluate workflow
on a Nebius GPU with an explicitly built MJLab image. Native GPU acceptance
results are recorded in the guide; public image promotion remains gated.
eval --video publishes the rendered MP4 and a self-contained HTML report with
measured episode results and checkpoint provenance.
The former deterministic scoring placeholder has been removed.