Skip to content
Open
Show file tree
Hide file tree
Changes from 121 commits
Commits
Show all changes
142 commits
Select commit Hold shift + click to select a range
02a05ae
docs(dtni): add vllm_single PoC plan and DTNI suite developer guide
atnair-amd Jun 9, 2026
438aafc
docs(dtni): expand dev guide with Background, before/after, config/th…
atnair-amd Jun 9, 2026
9926ab4
docs(dtni): rewrite dev guide with lifecycle flowcharts, test skeleto…
atnair-amd Jun 9, 2026
9466ce7
fix(cli): exclude conftest.py and _-prefixed files from cvs list [AIM…
atnair-amd Jun 15, 2026
5878aa7
vLLM single node refactor (#223)
atnair-amd Jun 16, 2026
8b3c756
feat(dtni): move vLLM bench client to stock vllm bench serve (Spec 0)…
atnair-amd Jun 17, 2026
cfd59f9
Merge pull request #225 from ROCm/hnimrama/inferencemax-uplift
hnimra-amd Jun 18, 2026
0207833
Revert "Merge pull request #225 from ROCm/hnimrama/inferencemax-uplif…
hnimra-amd Jun 18, 2026
831b19f
Restructure shared/inference libs + sweep selector + artifact-based v…
atnair-amd Jun 24, 2026
15b1ec3
Updating sglanf for multinode
amd-droy Jun 15, 2026
1f6a0f3
Updating smoke tests-1
amd-droy Jun 15, 2026
89894bd
Updating smoke tests-2
amd-droy Jun 15, 2026
de777a7
Updating smoke tests-3
amd-droy Jun 15, 2026
5120a41
Updating smoke tests-4
amd-droy Jun 15, 2026
e4c6cd3
Updating smoke tests-5
amd-droy Jun 15, 2026
5e430cd
Updating smoke tests-6
amd-droy Jun 15, 2026
e1840dd
Updating smoke tests-7
amd-droy Jun 15, 2026
e991b4a
Updating smoke tests-8
amd-droy Jun 16, 2026
2fd8228
Updating smoke tests-9
amd-droy Jun 16, 2026
ce7ce0b
Updating smoke tests-10
amd-droy Jun 16, 2026
341c537
Updating smoke tests-11
amd-droy Jun 16, 2026
917333d
Updating smoke tests-12
amd-droy Jun 16, 2026
dd0fe25
Updating smoke tests-13
amd-droy Jun 16, 2026
f168ae5
Updating smoke tests-14
amd-droy Jun 16, 2026
0b4af44
Updating smoke tests-15
amd-droy Jun 16, 2026
a6e7f26
Updating smoke tests-16
amd-droy Jun 16, 2026
0fa1356
Updating hellaswag-2
amd-droy Jun 17, 2026
d8c9e18
Updating hellaswag-3
amd-droy Jun 17, 2026
ae347f1
Updating hellaswag-4
amd-droy Jun 17, 2026
d1e7fed
Updating hellaswag-5
amd-droy Jun 17, 2026
6134bcf
Updating hellaswag-6
amd-droy Jun 17, 2026
bcaef40
Updating hellaswag-7
amd-droy Jun 17, 2026
f006160
Updating hellaswag-8
amd-droy Jun 17, 2026
97d035a
framework update -1
amd-droy Jun 17, 2026
4d34a69
framework update -2
amd-droy Jun 17, 2026
07a4954
Framework update -3
amd-droy Jun 17, 2026
e78cd77
Framework update -5
amd-droy Jun 18, 2026
5d343c1
Framework update -6
amd-droy Jun 18, 2026
c0773d1
Framework update -7'
amd-droy Jun 18, 2026
c212ae3
Framework update -8
amd-droy Jun 18, 2026
8d25045
Framework update -9
amd-droy Jun 18, 2026
48ffde1
Framework update -10
amd-droy Jun 18, 2026
d546aed
Framework update -11
amd-droy Jun 18, 2026
67d86a8
Framework update -13
amd-droy Jun 19, 2026
a6fff21
Framework update -14
amd-droy Jun 19, 2026
5ede28b
Framework update -15
amd-droy Jun 19, 2026
ff2e9a1
Framework update -17
amd-droy Jun 19, 2026
9453fec
Framework update -18
amd-droy Jun 20, 2026
c0cce0a
Framework update -18
amd-droy Jun 20, 2026
11a8be0
Framework update -19
amd-droy Jun 20, 2026
e98cf90
Framework update -20
amd-droy Jun 21, 2026
f707dde
Framework update -20
amd-droy Jun 21, 2026
f48d7d7
Framework update -21
amd-droy Jun 21, 2026
00333ab
Framework update -22
amd-droy Jun 21, 2026
4fbebf3
Framework update -23
amd-droy Jun 21, 2026
7ed71da
Framework update -24
amd-droy Jun 21, 2026
e2469a1
Framework update -25
amd-droy Jun 21, 2026
be762a9
Framework update -26
amd-droy Jun 21, 2026
be86f73
Framework update -27
amd-droy Jun 21, 2026
ab98d11
Framework update -27
amd-droy Jun 22, 2026
44396a2
Framework update -28
amd-droy Jun 22, 2026
d21ed9f
Framework update -29
amd-droy Jun 22, 2026
c65b784
Framework update -30
amd-droy Jun 22, 2026
369f315
Framework update -31
amd-droy Jun 22, 2026
3e1e031
xdit
amd-droy Jun 22, 2026
3d7bc88
Framework update -1
amd-droy Jun 22, 2026
13347a1
Framework update -2
amd-droy Jun 23, 2026
9c68d39
Framework update -3
amd-droy Jun 23, 2026
093502c
Framework update -4
amd-droy Jun 23, 2026
6db0dc9
Framework update -5
amd-droy Jun 23, 2026
474efc8
Framework update -6
amd-droy Jun 23, 2026
ff23d34
Framework update -7
amd-droy Jun 23, 2026
932ce55
Framework update -8
amd-droy Jun 23, 2026
7d44703
Framework update -17
amd-droy Jun 24, 2026
5a862ae
Framework update -18
amd-droy Jun 24, 2026
dce67d9
Framework update -19
amd-droy Jun 24, 2026
48b36b6
Framework update -20
amd-droy Jun 24, 2026
cd4fc82
Framework update -21
amd-droy Jun 25, 2026
1925377
Framework update -22
amd-droy Jun 25, 2026
1c0a1c1
Framework update -22
amd-droy Jun 25, 2026
8790cb2
Framework update -23
amd-droy Jun 25, 2026
78a303b
Framework update -25
amd-droy Jun 25, 2026
fce5e24
Framework update -26
amd-droy Jun 25, 2026
17077a1
Framework update -27
amd-droy Jun 25, 2026
c9cba51
Hnimrama/ix atom (#238) AIMVT-244/Add ATOM framework and inference …
hnimra-amd Jun 29, 2026
ea77382
docs(plans): add inference suite architecture changes presentation doc
atnair-amd Jul 1, 2026
1d3f7ea
Revert "docs(plans): add inference suite architecture changes present…
atnair-amd Jul 1, 2026
43ae95b
Hnimrama/Added ix atom reporting run deck module (#244)
hnimra-amd Jul 8, 2026
4332727
refactor(inference): move shared suite helpers under inference/utils …
hnimra-amd Jul 14, 2026
1fc1adb
[CVS] GPU metrics polling integration for inference validation suites…
atnair-amd Jul 15, 2026
8e3fc46
Integrating nodesmoke tier 1 tests using Primus cli (#250)
urtiwari Jul 16, 2026
d50fa0a
feat(vllm): unify vllm_single + vllm_distributed into one topology-pa…
atnair-amd Jul 17, 2026
e7bce93
feat(vllm): wire unified vLLM suite into inference report engine (#261)
atnair-amd Jul 20, 2026
a03823d
feat(inferencex_atom): add multinode params and scaling efficiency me…
hnimra-amd Jul 6, 2026
658efd3
feat(inferencex_atom): run distributed ATOM serve across cluster ranks
hnimra-amd Jul 6, 2026
bd5a225
feat(inferencex_atom): add MI300X 2-node cluster and W1 perf_multi va…
hnimra-amd Jul 6, 2026
cb33d56
test(inferencex_atom): cover multinode job routing and config loading
hnimra-amd Jul 6, 2026
ac05ac4
docs(inferencex_atom): document multinode cluster and variant params
hnimra-amd Jul 6, 2026
b69d698
feat(orchestrator): install openssh-server when image lacks sshd
hnimra-amd Jul 6, 2026
8699fd8
fix(inference): restore test_launch_container in inference_suite_life…
hnimra-amd Jul 6, 2026
b6c58ad
fix(inferencex_atom): use ATOM-native multinode serve instead of vLLM…
hnimra-amd Jul 6, 2026
4f55c31
fix(inferencex_atom): isolate ATOM multinode argv from vLLM distribut…
hnimra-amd Jul 6, 2026
1edba1e
config(inferencex_atom): expand multinode sweep to 5 shapes x conc 16…
hnimra-amd Jul 6, 2026
3ea052a
feat(inferencex_atom): enforce multinode scaling thresholds in gates …
hnimra-amd Jul 7, 2026
624d794
fix(inference): repair syntax error in test_launch_container
hnimra-amd Jul 7, 2026
126863a
fix(inference): tolerate CollectReport in html_metric_table_row
hnimra-amd Jul 7, 2026
ae2b647
refactor(inferencex_atom): rename suite from inferencex_atom_single
hnimra-amd Jul 7, 2026
81fbe06
docs(inferencex_atom): run make install before activating venv
hnimra-amd Jul 7, 2026
36cea47
style(inferencex_atom): fix ruff format and lint on multinode changes
hnimra-amd Jul 13, 2026
5f68413
feat(inferencex_atom): add MI355X multinode cluster and perf_multi va…
hnimra-amd Jul 13, 2026
979a41e
config(inferencex_atom): add MI300X baseline perf sweep variant
hnimra-amd Jul 15, 2026
b44c9b8
config(inferencex_atom): add MI355X baseline perf sweep variant
hnimra-amd Jul 15, 2026
8657ad3
test(inferencex_atom): cover baseline perf sweep config loading
hnimra-amd Jul 15, 2026
3a9f83a
docs(inferencex_atom): document baseline perf sweep runbook
hnimra-amd Jul 15, 2026
5dbc306
config(inferencex_atom): point models_dir at /home/models
hnimra-amd Jul 15, 2026
a69ef0b
config(cluster): note /home/models expectation for IX-atom templates
hnimra-amd Jul 15, 2026
ac38280
docs(inferencex_atom): document /home/models cache layout
hnimra-amd Jul 15, 2026
164c6ad
config(cluster): consolidate IX-atom cluster templates
hnimra-amd Jul 15, 2026
3b041bd
feat(inferencex_atom): add multinode baseline sweep and drop inferenc…
hnimra-amd Jul 15, 2026
2b37fd0
refactor(inferencex_atom): rename config stems from inferencex-atom-s…
hnimra-amd Jul 16, 2026
b71e9d8
chore(inferencex_atom): refresh multinode configs, docs, and loader t…
hnimra-amd Jul 21, 2026
d7726bf
fix(inferencex_atom): calibrate perf_multi C=16 thresholds from lab r…
hnimra-amd Jul 22, 2026
0f412be
fix(container): fail fast when openssh-server is missing from image
hnimra-amd Jul 22, 2026
9876b02
feat(inferencex_atom): add vllm_atom and sglang drivers for true PP m…
hnimra-amd Jul 22, 2026
0cb8a31
test(inferencex_atom): cover vllm_atom PP2 and sglang distributed launch
hnimra-amd Jul 22, 2026
d660459
config(inferencex_atom): switch multinode variants to vllm_atom PP=2
hnimra-amd Jul 22, 2026
44ba9fa
config(inferencex_atom): add SGLang PP=2 multinode W1 variant
hnimra-amd Jul 22, 2026
0c45f6d
docs(inferencex_atom): document vllm_atom and sglang PP multinode paths
hnimra-amd Jul 22, 2026
3f3dc1b
docs(inferencex_atom): align plan and runbooks with PP=2 driver model
hnimra-amd Jul 22, 2026
af127ef
docs(cluster): improve inferencex_atom cluster template for multinode…
hnimra-amd Jul 22, 2026
155c1a0
Rename inferencex_atom configs to flat single/distributed layout.
hnimra-amd Jul 22, 2026
021ade4
Fix multinode vLLM-ATOM readiness and align configs with flat layout.
hnimra-amd Jul 22, 2026
881193b
Fix multinode NCCL/Gloo networking for inferencex_atom.
hnimra-amd Jul 22, 2026
e46c76a
fix(container): repair setup_sshd docstring and drop sshd preflight c…
hnimra-amd Jul 23, 2026
200ff0b
Auto-discover socket netdev for multinode inferencex_atom.
hnimra-amd Jul 23, 2026
0e73bd3
Coerce legacy mlx5 ib_netdev configs to auto discovery.
hnimra-amd Jul 23, 2026
0227f00
Resolve multinode fabric lazily when topology test is skipped in smok…
hnimra-amd Jul 23, 2026
6ea2829
Probe multinode fabric on host OS, not inside container.
hnimra-amd Jul 23, 2026
d15edfa
Fix host topology probes: valid bash ip syntax and ibv banner parsing.
hnimra-amd Jul 23, 2026
0975c96
Recalibrate MI300X multinode thresholds from 20260723 lab run.
hnimra-amd Jul 23, 2026
c0d2cbc
Recalibrate CONC=16 TTFT gates from 20260724 multinode run.
hnimra-amd Jul 24, 2026
ac88bed
Raise W1 multinode max_model_length to 8192 for 2k sweep cells.
hnimra-amd Jul 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -17,3 +17,6 @@ docs/sphinx/_toc.yml

# Build distributions
dist/

# Local sample output
sample_reports/
9 changes: 8 additions & 1 deletion cvs/cli_plugins/list_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,14 @@ def discover_tests():
test_map[pkg_name] = {}
for root, dirs, files in os.walk(tests_dir):
for file in files:
if file.endswith(".py") and file != "__init__.py":
# Skip pytest infra (conftest.py) and private helpers
# (e.g. _shared.py): they are not selectable suites.
if (
file.endswith(".py")
and file != "__init__.py"
and file != "conftest.py"
and not file.startswith("_")
):
rel_path = os.path.relpath(os.path.join(root, file), tests_dir)
module_parts = os.path.splitext(rel_path)[0].split(os.sep)
# Module path: <tests_path>.<test_name>
Expand Down
8 changes: 6 additions & 2 deletions cvs/cli_plugins/run_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,12 @@ def get_parser(self, subparsers):
)
parser.add_argument(
"--log-file",
default="/tmp/cvs/test.log",
help="Pytest: Path to file for logging output (default: /tmp/cvs/test.log)",
default=None,
metavar="PATH",
help=(
"Pytest: write logging output to this file (optional). "
"Parent directories are created automatically when set."
),
)
parser.add_argument(
"--log-level",
Expand Down
29 changes: 29 additions & 0 deletions cvs/cli_plugins/unittests/test_run_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,35 @@ def test_run_test_multiple_functions(self, mock_exit, mock_pytest_main):
mock_pytest_main.assert_called_once_with(expected_args)
mock_exit.assert_called_once_with(0)

@patch("cvs.cli_plugins.run_plugin.pytest.main")
@patch("cvs.cli_plugins.run_plugin.sys.exit")
def test_run_test_omits_log_file_when_not_set(self, mock_exit, mock_pytest_main):
"""No --log-file is passed to pytest when the user does not request file logging."""
args = MagicMock()
args.test = "agfhc_cvs"
args.function = []
args.cluster_file = "/path/to/cluster.json"
args.config_file = "/path/to/config.json"
args.html = None
args.self_contained_html = False
args.log_file = None
args.log_level = None
args.capture = None
args.extra_pytest_args = []

mock_pytest_main.return_value = 0

with patch.object(self.plugin, "get_test_file", return_value="/mock/path/test.py"):
self.plugin.run(args)

expected_args = [
"/mock/path/test.py",
"--cluster_file=/path/to/cluster.json",
"--config_file=/path/to/config.json",
]
mock_pytest_main.assert_called_once_with(expected_args)
mock_exit.assert_called_once_with(0)


if __name__ == "__main__":
unittest.main()
117 changes: 112 additions & 5 deletions cvs/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,16 +6,19 @@
"""

import importlib.metadata
import logging
import sys
from pathlib import Path

import pytest

from cvs.lib.report_plugins import HtmlReportManager

log = logging.getLogger(__name__)

@pytest.hookimpl(tryfirst=True)
def pytest_configure(config):

def _sync_suite_name_from_args(config):
"""Derive suite stem from the first ``*.py`` target in ``config.args``."""
suite_name = "test"
for arg in config.args:
bare = arg.split("::")[0]
Expand All @@ -24,7 +27,96 @@ def pytest_configure(config):
break
config._suite_name = suite_name
config._test_html_dir = f"{suite_name}_html"


def _ensure_html_report_manager(config):
"""Create ``HtmlReportManager`` once; safe if ``pytest_configure`` did not run."""
_sync_suite_name_from_args(config)
mgr = getattr(config, "_html_report_manager", None)
if mgr is not None:
return mgr

config._html_report_manager = HtmlReportManager(config)
return config._html_report_manager


def _auto_register_inference_suite_report(config):
from cvs.lib.report.auto_register import try_auto_register_inference_suite_report

_sync_suite_name_from_args(config)
return try_auto_register_inference_suite_report(config)


@pytest.hookimpl(tryfirst=True)
def pytest_configure(config):
_ensure_html_report_manager(config)
_auto_register_inference_suite_report(config)


@pytest.fixture(scope="session", autouse=True)
def _cvs_inference_suite_report_session(request):
"""Initialize the session report store when a suite preset is registered."""
from cvs.lib.report.registry import clear_session_results, get_suite_report_config
from cvs.lib.report.types import InferenceReportConfig

if not isinstance(get_suite_report_config(request.config), InferenceReportConfig):
yield
return

clear_session_results()
yield


@pytest.fixture(scope="module", autouse=True)
def _cvs_inference_suite_report_bind_module(request, _cvs_inference_suite_report_session):
"""Bind module-scoped suite fixtures into the session store at module teardown."""
from cvs.lib.report.registry import bind_session_results, get_suite_report_config
from cvs.lib.report.types import InferenceReportConfig

if not isinstance(get_suite_report_config(request.config), InferenceReportConfig):
yield
return

inf_res_dict = None
variant_config = None
lifecycle = None
try:
inf_res_dict = request.getfixturevalue("inf_res_dict")
except pytest.FixtureLookupError:
log.warning(
"Inference suite report preset registered but inf_res_dict fixture is missing; "
"session-end report will be skipped"
)
yield
return
try:
variant_config = request.getfixturevalue("variant_config")
except pytest.FixtureLookupError:
log.warning(
"Inference suite report preset registered but variant_config fixture is missing; "
"session-end report will be skipped"
)
yield
return
try:
lifecycle = request.getfixturevalue("lifecycle")
except pytest.FixtureLookupError:
log.warning(
"Inference suite report preset registered but lifecycle fixture is missing; "
"session-end report will be skipped"
)
yield
return

def _bind_at_module_end():
bind_session_results(
inf_res_dict=inf_res_dict,
variant_config=variant_config,
lifecycle=lifecycle,
)

request.addfinalizer(_bind_at_module_end)
yield


# Add all additional cmd line arguments for the script
Expand Down Expand Up @@ -93,15 +185,28 @@ def pytest_metadata(metadata):

# Prepare a clean per-run log directory before tests start.
def pytest_sessionstart(session):
session.config._html_report_manager.setup_log_dir()
_auto_register_inference_suite_report(session.config)
_ensure_html_report_manager(session.config).setup_log_dir()


# Capture each test report and attach a per-test external log link.
@pytest.hookimpl(hookwrapper=True)
def pytest_runtest_makereport(item, call): # noqa: ARG001
outcome = yield
report = outcome.get_result()
report.extras = item.config._html_report_manager.write_test_log(report, item.originalname)
report.extras = _ensure_html_report_manager(item.config).write_test_log(report, item.originalname)

from cvs.lib.report.registry import get_suite_report_config
from cvs.lib.report.types import InferenceReportConfig

if isinstance(get_suite_report_config(item.config), InferenceReportConfig):
from cvs.lib.report.inference_wiring import (
attach_inference_suite_lifecycle_table,
attach_inference_suite_report_row_extra,
)

attach_inference_suite_lifecycle_table(item, report)
attach_inference_suite_report_row_extra(item, report)


# Replace inline pytest-html log content with a short externalized-log message.
Expand All @@ -118,4 +223,6 @@ def pytest_html_results_summary(prefix, summary, postfix):
@pytest.hookimpl(hookwrapper=True)
def pytest_sessionfinish(session, exitstatus): # noqa: ARG001
yield # wait for pytest-html and all other plugins to finish writing the report
session.config._html_report_manager.create_zip_bundle(session)
mgr = _ensure_html_report_manager(session.config)
mgr.generate_suite_reports(session)
mgr.create_zip_bundle(session)
7 changes: 6 additions & 1 deletion cvs/core/orchestrators/container.py
Original file line number Diff line number Diff line change
Expand Up @@ -477,11 +477,16 @@ def setup_sshd(self):
"chown -R root:root /root/.ssh",
"bash -c 'chmod 700 /root/.ssh && chmod 600 /root/.ssh/*'",
"mkdir -p /run/sshd", # Create privilege separation directory for sshd
"bash -c 'which sshd > /dev/null 2>&1 || (apt-get update -qq && apt-get install -y -q openssh-server iproute2)'",
"/usr/sbin/sshd -p2224",
]

timeouts = {
"bash -c 'which sshd > /dev/null 2>&1 || (apt-get update -qq && apt-get install -y -q openssh-server iproute2)'": 120,
}
Comment thread
amd-droy marked this conversation as resolved.
Outdated

for cmd in ssh_setup_commands:
result = self.exec(cmd, timeout=10, detailed=True)
result = self.exec(cmd, timeout=timeouts.get(cmd, 10), detailed=True)
# Check if command succeeded on all hosts
for hostname, output in result.items():
if output['exit_code'] != 0:
Expand Down
17 changes: 16 additions & 1 deletion cvs/core/runtimes/docker.py
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,14 @@ def setup_containers(
else:
self.log.info(f"Image {container_config['image']} already exists, skipping tar load")

if not self.check_image_exists(image):
self.log.info(f"Image {image} not present on all hosts; pulling before start")
pull_result = self.pull_image(image, timeout=600)
failed_pull = [host for host, res in pull_result.items() if res.get('exit_code') != 0]
if failed_pull:
self.log.error(f"Failed to pull image on hosts: {failed_pull}")
return False

cmd = f"sudo docker run -d --name {container_name} {all_args_str} {image} sleep infinity"

self.log.info(f"Starting long-running containers on {len(self.orchestrator.hosts)} nodes: {container_name}")
Expand All @@ -130,7 +138,7 @@ def setup_containers(
remove_cmd = f"sudo docker rm -f {container_name} || true"
self.orchestrator.all.exec(remove_cmd, timeout=30, print_console=False)

result = self.orchestrator.all.exec(cmd, timeout=60, detailed=True)
result = self.orchestrator.all.exec(cmd, timeout=120, detailed=True)

# Check if all hosts started successfully
success = all(output['exit_code'] == 0 for output in result.values())
Expand Down Expand Up @@ -278,6 +286,13 @@ def _build_runtime_args(runtime_args_config):

return args

def pull_image(self, image_name, timeout=None):
"""Pull container image on all hosts."""
timeout = timeout or 600
cmd = f"sudo docker pull {shlex.quote(image_name)}"
self.log.info(f"Pulling image on all hosts: {image_name}")
return self.orchestrator.all.exec(cmd, timeout=timeout, detailed=True)

def load_image(self, tar_path, timeout=None):
"""Load container image from tar file on all hosts."""
cmd = f"sudo docker load < {tar_path}"
Expand Down
23 changes: 23 additions & 0 deletions cvs/core/runtimes/unittests/test_docker.py
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,29 @@ def test_cmd_never_contains_gpus_all(self):
f"[{label}] '--gpus all' must never appear in docker cmd:\n{captured[0]}",
)

def test_pulls_image_when_missing_before_run(self):
calls = []

def _fake_exec(cmd, timeout=None, detailed=False, print_console=True):
calls.append(cmd)
if "docker images" in cmd and "grep" in cmd:
return {"host1": {"output": "", "exit_code": 1}}
return {"host1": {"output": "", "exit_code": 0}}

orchestrator = MagicMock()
orchestrator.hosts = ["host1"]
orchestrator.all.exec.side_effect = _fake_exec
rt = DockerRuntime(MagicMock(), orchestrator)

result = rt.setup_containers(
container_config=_container_config(),
container_name="cvs_iter_test",
volumes=["/home/u:/workspace"],
)
self.assertTrue(result)
self.assertTrue(any("docker pull" in c for c in calls))
self.assertTrue(any(c.startswith("sudo docker run") for c in calls))


if __name__ == "__main__":
unittest.main()
34 changes: 34 additions & 0 deletions cvs/input/cluster_file/inferencex_atom_cluster.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
{
"_comment": "InferenceX ATOM cluster template. Edit node_dict to match variant params.nnodes (1 host for single-node sweeps; 2+ for multinode). Set head_node_dict.mgmt_ip to the rank-0 VPC IP (params.master_addr for nnodes>1). Model weights under /home/models on each GPU node. Variant config overrides container image/name.",
"orchestrator": "container",
"username": "{user-id}",
"priv_key_file": "/home/{user-id}/.ssh/id_rsa",
"head_node_dict": {
"mgmt_ip": "{head-node-ip}"
},
"env_vars": {},
"node_dict": {
"{head-node-ip}": {
"bmc_ip": "NA",
"vpc_ip": "{head-node-ip}"
},
"{worker-node-ip}": {
"bmc_ip": "NA",
"vpc_ip": "{worker-node-ip}"
}
},
"container": {
"lifetime": "per_run",
"image": "rocm/atom-dev:latest",
"name": "inferencex_atom",
"runtime": {
"name": "docker",
"args": {
"network": "host",
"ipc": "host",
"privileged": true,
"shm_size": "128G"
}
}
}
}
Loading