Skip to content

[Feature] Establish Simpler as a vLLM program execution backend #2429

Description

@Crane-Liu

Summary

Establish Simpler's program execution path as a safe, measurable backend for vLLM serving. The work starts from the Pipeline A run-to-run enqueue capability and ends with a fixed Qwen workload running through a vLLM eager adapter, followed by capture/replay support where the target SDK and platform contracts are proven.

The device remains serial at the operator level. The host may prepare and enqueue later runs while an earlier run is still executing; every run must retain independent completion, result, error, and resource ownership until its last consumer is finished.

Why this issue exists

Pipeline A and the P3 contracts provide the runtime foundation for this work:

This issue is the end-to-end integration tracker. Individual implementation PRs should remain focused and link back here.

Scope and boundaries

The work is split into a program-path milestone and a later vLLM/capture milestone.

Program-path milestone

  1. Validate the current P4 capability: real Worker.submit early enqueue, whole-operator device ordering, sustained refill, and depth-one default behavior.
  2. Revalidate P1–P3 contracts under the new admission path:
    • run-owned parameter, runtime/image, arena, result, timing, and diagnostic resources;
    • per-run completion/drain, error attribution, event references, and generation retirement;
    • parent-run dispatch FIFO, prepared-successor admission, build/recorder ownership, and host accessor ordering;
    • safe default, unsupported-shape fallback, diagnostic exclusivity, direct-control, close, and stop-admission paths.
  3. Extend the proven path only where the target workload requires it, with an explicit capability matrix for A3/A5 HBG, TMR, HOST/DEVICE tensors, group/SUB/communication/provider shapes, and endpoint scope.
  4. Replace fixed-slot assumptions with bounded resource accounting and backpressure only after run identity, last-consumer retirement, workspace growth, and failure handling are proven.
  5. Integrate continuous DFX collection by collector and runtime, preserving run identity, buffer ownership, accounting, bounded flush/close, and diagnostic isolation.

vLLM milestone

  1. Freeze the Qwen workload and ABI: weights, prompt/token hash, prefill/decode shape, KV/page layout, sampling, output contract, and device/software versions.
  2. Add a fixed-shape Simpler-backed vLLM eager path through the public model-runner boundary, then cover the agreed dynamic batch and request-lifecycle range.
  3. Add capture/replay only after the preparation-result, parameter-binding, graph-reference lifetime, and target SDK contracts are established. Capture/replay is a separate capability boundary from the first program-path milestone.
  4. Run end-to-end correctness and performance comparisons with identical workload and device conditions.

Current acceptance gate

The first P4 acceptance phase passed on the bounded A3 HBG scope: local L3, same endpoint, HOST tensors, single non-group NEXT_LEVEL task, launch_depth=2, with depth one as the default control.

The P1–P3 regression phase is currently partial:

  • CPU contract tests passed, including admission/lifecycle coverage after the current runtime was built.
  • A3 worker_async_fifo passed for HBG and TMR.
  • worker_async_endpoint failed on both device 0 and device 1 with the same FFTSPLUS AICore fault (507018, 0x4000000000000000), followed by scheduler -100 and force reset.

The original endpoint failure was investigated with controlled depth-one/depth-two runs and five-round pressure tests on two additional devices. The fault did not reproduce with the same main, callable, and runtime. It is retained as a non-stable device/runtime anomaly; no code repair is opened from this incident. The bounded program-path milestone is now complete for the documented scope, while vendor-level root-cause analysis remains optional follow-up if maintainers request it.

Milestones

  • P4 bounded capability and evidence audit complete
  • P1–P3 contracts pass under the new admission path (bounded program-path scope)
  • Endpoint fault dispositioned as a non-stable device/runtime anomaly; bounded supported scope recorded
  • Qwen fixed-workload Worker.submit path is correct for the bounded Step 2 scope (depth one real qualification; depth-two safe fallback)
  • Step 3 clone-request L3 contract and A3 early-enqueue evidence accepted (Test clone requests on the early-enqueue path #2471)
  • PR-S3-1 clone-request contract implementation and complete CI validation passed; merge pending
  • Qwen clone-request path is connected on the merged Step 2 baseline
  • Same-request DEVICE feedback joined enqueue is accepted for the target shape
  • Required A3/A5 HBG and TMR capability matrix is accepted
  • Resource lifetime, workspace budget, backpressure, and continuous DFX contracts are accepted
  • vLLM fixed-shape eager path is correct
  • Agreed dynamic request range is correct
  • Capture/replay capability is accepted for the target platform/version
  • End-to-end correctness and performance report is complete

Step 2 closeout and Step 3 handoff

The bounded Qwen Step 2 scope is complete in two review PRs:

  • #2447: real Qwen depth-one Worker.submit, 127-step autoregressive token/KV qualification, adapter and lifecycle contract.
  • #2456: workload manifest, dependency matrix, depth-two request evidence, and explicit safe serial fallback when host sampled-token feedback prevents joined enqueue.

Both PRs are merged into main. The Step 2 depth-two conclusion is a safe fallback; it does not claim joined native early enqueue.

Step 3 starts with the L3 clone-request path on the current program runtime. The first proof uses multiple logically independent requests with identical workload content:

  • request content may be identical;
  • request/run/generation identities must differ;
  • live KV pages, block tables, input staging, sampled-token/output buffers, events, errors, results, and resource leases must be request-private;
  • request B must be prepared and enqueued while request A is still executing;
  • operator execution remains serial, parent dispatch FIFO remains ordered, and A remains readable after B starts;
  • physical capacity may be reused only after the previous request's last consumer retires.

PR #2471 is the first Step 3 contract/evidence PR. It adds the clone-request identity/resource checks to the existing A3 early-enqueue path and records A3 hardware evidence. Its current head is 0095497523c33fb86766d5e90e01de649ea578a3, based on main at 833327f60ea8f4772bf1673ea8edb349314a350d. The complete GitHub CI matrix passed, including A2A3/A5 simulation and onboard checks, UT, pre-commit, build, and DeepSeek/network validation. The PR remains open pending merge. It does not add Qwen adapter code or change runtime admission.

After the clone-request contract is accepted, the project will connect the real Qwen adapter on the post-merge Step 2 baseline. Same-request adjacent decode feedback (sampled token/KV/sequence metadata) is a separate DEVICE producer/consumer capability and remains depth-one fallback until proven. A5 HBG, TMR, multi-worker extensions, and distinct-request HBG numerical correctness remain separate capability work. The existing program-path boundaries remain in force: operator execution stays serial, host access has no implicit cross-run synchronization, and unsupported shapes must fall back or be rejected explicitly.

Tracking rules

Every implementation PR must:

  • link this issue in the PR body, using Part of #<issue-number> unless the PR is intended to close the entire issue;
  • state the exact phase and capability it changes;
  • record the base/head, supported scope, tests, hardware task IDs and evidence paths;
  • preserve the distinction between accepted, prepared, enqueued, completed, and retired states;
  • document unsupported or fallback paths when the change affects admission.

Progress updates on this issue should summarize the phase, current commit, evidence, open blocker, and next action. Individual PR merge does not close this issue. Close this issue only after the milestone checklist and final vLLM acceptance are complete.

Related design boundaries

The first milestone follows the existing program-path Pipeline A design. Kernel API work, tiling caches, ACLGraph, and capture-specific resource contracts are separate follow-on capabilities and are not implicit prerequisites for the bounded program-path acceptance. No stage in this issue changes device execution from serial operator ordering to device concurrency, adds implicit host tensor synchronization, or promises a performance gain before controlled measurements are complete.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions