Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
49 commits
Select commit Hold shift + click to select a range
a03823d
feat(inferencex_atom): add multinode params and scaling efficiency me…
hnimra-amd Jul 6, 2026
658efd3
feat(inferencex_atom): run distributed ATOM serve across cluster ranks
hnimra-amd Jul 6, 2026
bd5a225
feat(inferencex_atom): add MI300X 2-node cluster and W1 perf_multi va…
hnimra-amd Jul 6, 2026
cb33d56
test(inferencex_atom): cover multinode job routing and config loading
hnimra-amd Jul 6, 2026
ac05ac4
docs(inferencex_atom): document multinode cluster and variant params
hnimra-amd Jul 6, 2026
b69d698
feat(orchestrator): install openssh-server when image lacks sshd
hnimra-amd Jul 6, 2026
8699fd8
fix(inference): restore test_launch_container in inference_suite_life…
hnimra-amd Jul 6, 2026
b6c58ad
fix(inferencex_atom): use ATOM-native multinode serve instead of vLLM…
hnimra-amd Jul 6, 2026
4f55c31
fix(inferencex_atom): isolate ATOM multinode argv from vLLM distribut…
hnimra-amd Jul 6, 2026
1edba1e
config(inferencex_atom): expand multinode sweep to 5 shapes x conc 16…
hnimra-amd Jul 6, 2026
3ea052a
feat(inferencex_atom): enforce multinode scaling thresholds in gates …
hnimra-amd Jul 7, 2026
624d794
fix(inference): repair syntax error in test_launch_container
hnimra-amd Jul 7, 2026
126863a
fix(inference): tolerate CollectReport in html_metric_table_row
hnimra-amd Jul 7, 2026
ae2b647
refactor(inferencex_atom): rename suite from inferencex_atom_single
hnimra-amd Jul 7, 2026
81fbe06
docs(inferencex_atom): run make install before activating venv
hnimra-amd Jul 7, 2026
36cea47
style(inferencex_atom): fix ruff format and lint on multinode changes
hnimra-amd Jul 13, 2026
5f68413
feat(inferencex_atom): add MI355X multinode cluster and perf_multi va…
hnimra-amd Jul 13, 2026
979a41e
config(inferencex_atom): add MI300X baseline perf sweep variant
hnimra-amd Jul 15, 2026
b44c9b8
config(inferencex_atom): add MI355X baseline perf sweep variant
hnimra-amd Jul 15, 2026
8657ad3
test(inferencex_atom): cover baseline perf sweep config loading
hnimra-amd Jul 15, 2026
3a9f83a
docs(inferencex_atom): document baseline perf sweep runbook
hnimra-amd Jul 15, 2026
5dbc306
config(inferencex_atom): point models_dir at /home/models
hnimra-amd Jul 15, 2026
a69ef0b
config(cluster): note /home/models expectation for IX-atom templates
hnimra-amd Jul 15, 2026
ac38280
docs(inferencex_atom): document /home/models cache layout
hnimra-amd Jul 15, 2026
164c6ad
config(cluster): consolidate IX-atom cluster templates
hnimra-amd Jul 15, 2026
3b041bd
feat(inferencex_atom): add multinode baseline sweep and drop inferenc…
hnimra-amd Jul 15, 2026
2b37fd0
refactor(inferencex_atom): rename config stems from inferencex-atom-s…
hnimra-amd Jul 16, 2026
b71e9d8
chore(inferencex_atom): refresh multinode configs, docs, and loader t…
hnimra-amd Jul 21, 2026
d7726bf
fix(inferencex_atom): calibrate perf_multi C=16 thresholds from lab r…
hnimra-amd Jul 22, 2026
0f412be
fix(container): fail fast when openssh-server is missing from image
hnimra-amd Jul 22, 2026
9876b02
feat(inferencex_atom): add vllm_atom and sglang drivers for true PP m…
hnimra-amd Jul 22, 2026
0cb8a31
test(inferencex_atom): cover vllm_atom PP2 and sglang distributed launch
hnimra-amd Jul 22, 2026
d660459
config(inferencex_atom): switch multinode variants to vllm_atom PP=2
hnimra-amd Jul 22, 2026
44ba9fa
config(inferencex_atom): add SGLang PP=2 multinode W1 variant
hnimra-amd Jul 22, 2026
0c45f6d
docs(inferencex_atom): document vllm_atom and sglang PP multinode paths
hnimra-amd Jul 22, 2026
3f3dc1b
docs(inferencex_atom): align plan and runbooks with PP=2 driver model
hnimra-amd Jul 22, 2026
af127ef
docs(cluster): improve inferencex_atom cluster template for multinode…
hnimra-amd Jul 22, 2026
155c1a0
Rename inferencex_atom configs to flat single/distributed layout.
hnimra-amd Jul 22, 2026
021ade4
Fix multinode vLLM-ATOM readiness and align configs with flat layout.
hnimra-amd Jul 22, 2026
881193b
Fix multinode NCCL/Gloo networking for inferencex_atom.
hnimra-amd Jul 22, 2026
e46c76a
fix(container): repair setup_sshd docstring and drop sshd preflight c…
hnimra-amd Jul 23, 2026
200ff0b
Auto-discover socket netdev for multinode inferencex_atom.
hnimra-amd Jul 23, 2026
0e73bd3
Coerce legacy mlx5 ib_netdev configs to auto discovery.
hnimra-amd Jul 23, 2026
0227f00
Resolve multinode fabric lazily when topology test is skipped in smok…
hnimra-amd Jul 23, 2026
6ea2829
Probe multinode fabric on host OS, not inside container.
hnimra-amd Jul 23, 2026
d15edfa
Fix host topology probes: valid bash ip syntax and ibv banner parsing.
hnimra-amd Jul 23, 2026
0975c96
Recalibrate MI300X multinode thresholds from 20260723 lab run.
hnimra-amd Jul 23, 2026
c0d2cbc
Recalibrate CONC=16 TTFT gates from 20260724 multinode run.
hnimra-amd Jul 24, 2026
ac88bed
Raise W1 multinode max_model_length to 8192 for 2k sweep cells.
hnimra-amd Jul 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions cvs/core/orchestrators/container.py
Original file line number Diff line number Diff line change
Expand Up @@ -618,6 +618,10 @@ def exec(self, cmd, hosts=None, timeout=None, detailed=False):

return self.runtime.exec(self.container_id, cmd, hosts, timeout, detailed)

def exec_on_host(self, cmd, hosts=None, timeout=None, detailed=False):
"""Execute command on the cluster host OS (SSH), not inside the container."""
return super().exec(cmd, hosts=hosts, timeout=timeout, detailed=detailed)

def exec_on_head(self, cmd, timeout=None):
"""
Execute command directly on head node (baremetal).
Expand Down
17 changes: 16 additions & 1 deletion cvs/core/orchestrators/unittests/test_container.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@
# patched once in setUp (not per method); _make() returns a fresh orch + runtime mock.

import unittest
from unittest.mock import MagicMock, patch
from unittest.mock import MagicMock, call, patch

from cvs.core.orchestrators.factory import OrchestratorConfig, _resolve_container_lifetime
from cvs.core.orchestrators.container import ContainerOrchestrator
Expand Down Expand Up @@ -258,6 +258,21 @@ def test_setup_sshd_multinode_attempts_setup(self):
}
self.assertTrue(orch.setup_sshd())
self.assertTrue(runtime.exec.called)
cmds = [call.args[0] for call in runtime.exec.call_args_list]
self.assertFalse(any("apt-get" in cmd for cmd in cmds))

@patch("time.sleep", lambda *_a, **_k: None)
def test_setup_sshd_multinode_fails_when_sshd_missing(self):
orch, runtime = self._make(lifetime="per_run")
orch.container_id = "cvs_iter_test"

def _exec_side_effect(cmd, **kwargs):
if "command -v /usr/sbin/sshd" in cmd:
return {"10.0.0.1": {"exit_code": 127, "output": "missing sshd"}}
return {"10.0.0.1": {"exit_code": 0}}

runtime.exec.side_effect = _exec_side_effect
self.assertFalse(orch.setup_sshd())

# ------------------------------------------------------------------
# teardown_containers lifetime branching
Expand Down
35 changes: 35 additions & 0 deletions cvs/input/cluster_file/inferencex_atom_cluster.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
{
"_comment": "InferenceX ATOM cluster template (container backend). Copy to ~/input/cluster_file/inferencex_atom_cluster.json and edit placeholders. head_node_dict.mgmt_ip MUST be the rank-0 GPU node VPC IP (same as variant params.master_addr for nnodes>1) — not the pytest jumphost. Trim node_dict to one host for single-node variants. Variant config overrides container.image/name/volumes; cluster container block is the fallback default.",
"orchestrator": "container",
"username": "{user-id}",
"priv_key_file": "/home/{user-id}/.ssh/cluster_id_ed25519",
"head_node_dict": {
"mgmt_ip": "{head-node-ip}"
},
"env_vars": {},
"_env_vars_comment": "Optional host env exported on each GPU node before container setup (e.g. ROCm install on host). Usually empty when the workload image carries ROCm/vLLM.",
"node_dict": {
"{head-node-ip}": {
"bmc_ip": "NA",
"vpc_ip": "{head-node-ip}"
},
"{worker-node-ip}": {
"bmc_ip": "NA",
"vpc_ip": "{worker-node-ip}"
}
},
"container": {
"lifetime": "per_run",
"image": "rocm/atom-dev:latest",
"name": "inferencex_atom",
"runtime": {
"name": "docker",
"args": {
"network": "host",
"ipc": "host",
"privileged": true,
"shm_size": "128G"
}
}
}
}
30 changes: 0 additions & 30 deletions cvs/input/cluster_file/mi300x_atom_single.json

This file was deleted.

30 changes: 0 additions & 30 deletions cvs/input/cluster_file/mi355x_atom_single.json

This file was deleted.

Loading