Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -195,9 +195,31 @@ See [`docs/nop-guide.md`](docs/nop-guide.md) for full end-to-end walkthroughs pe
| `--account-pubkey` | Account public key (secure workflow) | `full`, `network` |
| `--mlnode-image` | Custom MLNode Docker image (overrides auto-detection) | `full`, `mlnode` |
| `--attention-backend` | vLLM attention backend: `FLASHINFER` or `FLASH_ATTN` | `full`, `mlnode` |
| `--model` | Governance-approved model ID (rejects mismatched GPU arch before any files are written) | `full`, `mlnode` |
| `-y, --yes` | Non-interactive mode | All |
| `-o, --output` | Output directory (default: `./gonka-node`) | All |

## Supported Models

NOP auto-recommends a model based on detected GPUs and the chain's governance
list. In interactive setup it shows the recommendation and offers a picker
filtered to **only the models your GPU architecture can serve** — pick any
of those, or accept the recommendation by hitting Enter. Override
non-interactively with `--model <id>`; NOP rejects mismatches before
writing any config.

| Model ID | Min VRAM/GPU | Min Total VRAM | GPU arch | Image variant |
|---|---|---|---|---|
| `Qwen/Qwen3-235B-A22B-Instruct-2507-FP8` | 40 GB | 320 GB | any (A100+) | standard |
| `Qwen/Qwen3-32B-FP8` | 20 GB | 40 GB | any | standard |
| `Qwen/QwQ-32B` | 20 GB | 80 GB | any | standard |
| `moonshotai/Kimi-K2.6` | 80 GB | 320 GB | **Blackwell only** (sm_100 / sm_103 / sm_120) | blackwell |

Picking a model your hardware cannot serve (e.g. Kimi on A100) is a hard
error at setup time. To still earn weight for a model you cannot run, see
the **per-model governance opt-in** flow in [`docs/nop-guide.md`](docs/nop-guide.md)
— delegate to a Blackwell host instead of refusing (~5% loss vs ~10%).

## Manual vs Automated

| Task | Manual | With gonka-nop |
Expand Down
39 changes: 39 additions & 0 deletions docs/nop-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ matches your setup and follow the section start to finish.

- [Topologies at a glance](#topologies-at-a-glance)
- [Prerequisites](#prerequisites)
- [Picking a model](#picking-a-model)
- [Topology 1 — Full (single server)](#topology-1--full-single-server)
- [Topology 2 — Multi-server (1 validator + N ML slaves)](#topology-2--multi-server-1-validator--n-ml-slaves)
- [Topology 3 — Network-only validator (no local GPU)](#topology-3--network-only-validator-no-local-gpu)
Expand Down Expand Up @@ -54,6 +55,44 @@ HuggingFace cache (`HF_HOME`) needs ~220 GB for Qwen3-235B FP8.

---

## Picking a model

NOP auto-recommends a model from the chain's governance-approved list
based on the detected hardware.

- **Interactive setup**: after GPU detection, you see the recommendation
and a picker. The picker is **filtered by GPU architecture** —
Kimi-K2.6 is hidden on A100/H100/H200, visible on Blackwell. Hit
Enter to accept the recommendation, or arrow-keys to pick another
supported model. Switching models re-runs the recommender so TP/PP
and memory settings match the new choice.
- **Non-interactive (`--yes`)**: override with `--model <id>`. NOP
rejects mismatched architectures before writing any config — no
half-deployed state.

| Model ID | Min VRAM/GPU | Min Total VRAM | GPU arch | Notes |
|---|---|---|---|---|
| `Qwen/Qwen3-235B-A22B-Instruct-2507-FP8` | 40 GB | 320 GB | any (A100+, H100, H200, Blackwell) | Default for most fleets; ~220 GB HF cache |
| `Qwen/Qwen3-32B-FP8` | 20 GB | 40 GB | any | Smaller model, lower VRAM bar |
| `Qwen/QwQ-32B` | 20 GB | 80 GB | any | Reasoning variant |
| `moonshotai/Kimi-K2.6` | 80 GB | 320 GB | **Blackwell only** (sm_100, sm_103, sm_120) | Refuses to deploy on Ampere/Hopper — NOP will hard-error |

**Hardware ↔ model decision matrix**

| Your hardware | Recommended model |
|---|---|
| 8×A100 80GB | Qwen3-235B-A22B-Instruct-2507-FP8 |
| 8×H100 / H200 80GB | Qwen3-235B-A22B-Instruct-2507-FP8 |
| 4–8× B200 / B300 | Kimi-K2.6 (best weight); Qwen3-235B also works |
| 2× B200 / 4× consumer 24GB | Qwen3-32B-FP8 or QwQ-32B |

For models you **cannot** run (e.g. Kimi on A100), see
[Per-model governance opt-in](#per-model-governance-opt-in) — delegating
to a Blackwell host that runs Kimi costs ~5% weight vs ~10% for refuse,
~15% for doing nothing.

---

## Topology 1 — Full (single server)

Single host runs everything. Simplest path.
Expand Down
34 changes: 30 additions & 4 deletions internal/phases/02_gpu_detection.go
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@
return !state.IsPhaseComplete(p.Name())
}

func (p *GPUDetection) Run(ctx context.Context, state *config.State) error {

Check failure on line 44 in internal/phases/02_gpu_detection.go

View workflow job for this annotation

GitHub Actions / Lint

cyclomatic complexity 29 of func `(*GPUDetection).Run` is high (> 15) (gocyclo)
gpus, err := p.detectGPUs(ctx)
if err != nil {
return err
Expand Down Expand Up @@ -88,12 +88,38 @@
ui.Error("%v", err)
return err
}
// Interactive model picker: when the operator did NOT pass --model and we
// have ≥2 arch-compatible recipes, let them override the auto-recommended
// choice. Skips in --yes (non-interactive) and when --model was supplied.
// If a non-default is chosen, re-run recommendConfig so TP/PP/memory and
// image-variant selection match the new model.
if state.SelectedModel == "" && !ui.IsNonInteractive() {
options := BuildModelMenuOptions(gpus[0].Architecture)
if len(options) > 1 {
chosen, perr := ui.SelectWithDefault(
"Select model to deploy",
options,
rec.Model,
)
if perr != nil {
return fmt.Errorf("model selection prompt: %w", perr)
}
if chosen != rec.Model {
state.SelectedModel = chosen
rec, err = recommendConfig(len(gpus), gpus[0].MemoryMB, gpus[0].Architecture, topology.HasNVLink, state.SelectedModel)
if err != nil {
ui.Error("%v", err)
return err
}
}
}
}
state.TPSize = rec.TP
state.PPSize = rec.PP
// Preserve pre-set SelectedModel (from --model flag or saved state) so
// the operator's choice survives phase 02 (Review Notes F5). When the
// flag was passed, recommendConfig already validated arch/VRAM and rec.Model
// equals state.SelectedModel; this guard makes the no-op explicit.
// Preserve pre-set SelectedModel (from --model flag, picker choice, or
// saved state) so the operator's choice survives phase 02 (Review Notes
// F5). When the flag was passed, recommendConfig already validated
// arch/VRAM and rec.Model equals state.SelectedModel.
if state.SelectedModel == "" {
state.SelectedModel = rec.Model
}
Expand Down Expand Up @@ -452,7 +478,7 @@
defer cancel()

// Get anonymous token for GHCR
tokenURL := "https://ghcr.io/token?scope=repository:product-science/mlnode:pull"

Check failure on line 481 in internal/phases/02_gpu_detection.go

View workflow job for this annotation

GitHub Actions / Lint

G101: Potential hardcoded credentials (gosec)
req, err := http.NewRequestWithContext(ctx, "GET", tokenURL, nil)
if err != nil {
return ""
Expand Down
55 changes: 55 additions & 0 deletions internal/ui/prompt.go
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,61 @@ func Select(message string, options []string) (string, error) {
return result, err
}

// SelectWithDefault prompts the user to select from a list of options with
// a default highlighted at the top. Behaves like Select() with respect to
// override resolution (--yes mode looks up an exact / substring-matching
// override value in the option set) and non-interactive fallback (uses
// defaultVal if present in options, else first option).
//
// When defaultVal is in options, survey highlights it as the preselected
// entry — the operator can hit Enter to accept.
func SelectWithDefault(message string, options []string, defaultVal string) (string, error) {
if val, ok := findOverride(message); ok {
lowerVal := strings.ToLower(val)
for _, opt := range options {
if strings.Contains(strings.ToLower(opt), lowerVal) {
return opt, nil
}
}
for _, opt := range options {
if strings.EqualFold(opt, val) {
return opt, nil
}
}
return "", fmt.Errorf("override value %q does not match any option: %v", val, options)
}
if IsNonInteractive() {
for _, opt := range options {
if opt == defaultVal {
return defaultVal, nil
}
}
if len(options) > 0 {
return options[0], nil
}
return "", fmt.Errorf("no options available for prompt: %s", message)
}

defaultExists := false
for _, opt := range options {
if opt == defaultVal {
defaultExists = true
break
}
}

var result string
prompt := &survey.Select{
Message: message,
Options: options,
}
if defaultExists {
prompt.Default = defaultVal
}
err := survey.AskOne(prompt, &result)
return result, err
}

// Confirm prompts user for yes/no confirmation
func Confirm(message string, defaultVal bool) (bool, error) {
if IsNonInteractive() {
Expand Down
Loading