Skip to content

feat(vllm_dissag): Kimi-K3-MXFP4 2P/2D disagg framework integration - #237

Open
MIR-AMD wants to merge 12 commits into
ROCm:developfrom
MIR-AMD:mir/kimik3-framework-only
Open

MIR-AMD wants to merge 12 commits into
ROCm:developfrom
MIR-AMD:mir/kimik3-framework-only

Conversation

@MIR-AMD

@MIR-AMD MIR-AMD commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds Kimi-K3-MXFP4 (MI300X gfx942, MXFP4 MoE, 2.8T params / 896 experts) to the vllm_dissag framework as a first-class 2P/2D disaggregated model — implemented as a generic framework capability, not a model-specific fork.

The core change: the old hardcoded MODEL_NAME=="Kimi-K3-MXFP4" branches in shared code are replaced by a generic EP_TP_SIZE knob (TP-within-EP). Unset / 1 → plain wideEP (-tp 1; today's behavior for DeepSeek/Llama/Qwen); N>1 → TP inside each EP pool (K3 = 2 → TP2×DP8 → EP16). K3 sets it only in its models.yaml recipe env — shared code stays model-agnostic.

Changes (11 files, +832 −27)

  • models.yaml: K3 recipe env anchor + model entry (EP_TP_SIZE, RDMA fabric, per-role cudagraph/all2all, KV budget, PR#241-parity env)
  • connectors/moriio.sh: _moriio_is_kimik3 branches → generic (( EP_TP_SIZE>1 )) gates; per-role all2all (prefill mori_high_throughput / decode mori_low_latency) + per-role cudagraph
  • connectors/rixl.sh: unsafe eval → shared _model_config_to_array helper
  • vllm_disagg.sh: early EP_TP_SIZE resolve + divisibility guard + PREFILL/DECODE_POD_HOSTS export
  • run_xPyD_models.slurm: generic NUM_NODES=xP+yD (no K3 topology lock) + EP_TP_SIZE>1-gated JIT cache role-split + allowlists
  • tests/run_interactive.sh, tests/drive_cell.sh: EP_TP_SIZE resolve on the interactive path
  • tests/argv_assert.sh: 40 → 56 offline assertions (K3 TP2×DP8 + non-K3 wideEP dormancy + parametric EP_TP_SIZE fuzz)
  • README.MD / ARCHITECTURE.md: docs for EP_TP_SIZE, generic topology, in-tree Dockerfile
  • docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile: new in-tree K3 disagg image (K3-specialized sibling of the generalized vllm_disagg_inference.ubuntu.amd.Dockerfile)

Regression safety

No model-name decision branch remains in shared code — K3 specifics are confined to models.yaml. Non-K3 models (DeepSeek-V3/R1, Llama, Qwen) resolve EP_TP_SIZE → 1: plain wideEP, JIT split OFF, no topology lock.

  • Offline argv_assert.sh: 56 passed / 0 failed
  • 64-cell develop-vs-branch dry-run diff: 64/64 byte-identical for non-K3 cells (dense TP moriio/rixl + DeepSeek wideEP mori/deepep, across ranks, 2P/2D)

Live validation (K3 2P/2D, MI300X + CX7)

  • Single-stream NIAH correctness: 12/12 at 50K–280K context
  • Concurrency NIAH: 57/57 across con = {1, 8, 16, 32}
  • Recipe pinned to the PR#241-validated config: gpu-memory-utilization 0.68, kv-cache-memory-bytes 8e9, per-role mori_high_throughput/mori_low_latency, decode cudagraph_mode FULL_AND_PIECEWISE / prefill NONE

Performance

Measured on 4× MI300X (2P/2D), MI300X + CX7 8-rail RoCE, 64 output tokens, temp 0.

End-to-end request latency (single stream, by context length):

Context Latency
50K ~25 s (first-request warmup ~38 s)
100K ~53 s
200K ~121 s
280K ~187 s

Throughput under concurrency (50K context, distinct-needle recall):

Concurrency Wall time Recall
1 24.9 s 1/1
8 44.7 s 8/8
16 85.0 s 16/16
32 162.9 s 32/32

32 concurrent 50K-context requests complete in ~163 s (≈4.9× faster than serial execution) at full recall (57/57).

Image

Built in-tree from the repo root:

docker build -f docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile \
  --build-arg VLLM_REF=kimi-k3-wideep-disagg-fullsource-v3 \
  -t kimik3-wideep-disagg:latest .

Related

Test plan

  • bash -n on all .sh files
  • YAML lint on models.yaml
  • argv_assert.sh offline tests (56/56)
  • 64-cell dry-run parity vs develop (non-K3 argv unchanged)
  • 4-node live smoke: NIAH single-stream 12/12 (to 280K) + concurrency 57/57

Note on RDMA rails: the recipe ships an 8-rail NCCL_IB_HCA/MORI_RDMA_DEVICES list; it is fabric-specific and documented as adjustable in models.yaml.

MIR-AMD and others added 9 commits August 27, 2026 20:34
Adds Kimi-K3-MXFP4 (MI300X gfx942 MXFP4 MoE) to the vllm_dissag framework:
- models.yaml: K3 recipe env anchor + model entry (TP2×DP8, MoRI-EP)
- run_xPyD_models.slurm: allowlists + 2P/2D validation + JIT cache split
- moriio.sh: K3-gated topology (pod_hosts, api-server-count, headless KV)
- vllm_disagg.sh: pod hosts export + _model_config_to_array helper
- argv_assert.sh: K3 offline tests

Config-only integration; image build documented in README.

Co-Authored-By: Claude <noreply@anthropic.com>
models.yaml recipe env: MI300X+CX7 RoCE fabric (RDMA rails, GID 3, socket ifname); PR241 env parity (NCCL/TORCH_NCCL watchdog, AITER MHA, MoRI relaxed-ordering/SL, QP_PER_PE=8, QP_PER_TRANSFER=2, NCCL_IB_TIMEOUT); KV write-visibility fence switched from image-incompatible readback to sleep fence (K3_WRITE_READBACK=0, K3_WRITE_FENCE=1, K3_WRITE_DEVSYNC=1, K3_WRITE_FENCE_MS=20); serve/cudagraph tuning (max-num-seqs=32, FULL_AND_PIECEWISE capture [1,2,4], prefill mori_high_throughput). run_interactive.sh + run_xPyD_models.slurm: fix GPU_MEMORY_UTILIZATION shadow so recipe value propagates (conditional -e injection).
The 0.85 recipe value fails the vLLM startup free-memory check on 192GB
MI300X GPUs (needs 163.19 GiB free; ~31 GiB framework baseline leaves
only ~160 GiB), aborting boot. 0.80 is the proven value (prefill boots:
KV cache allocated, all-8 RDMA rails init, MoRIIO connector up). The
0.85 value was never actually exercised because the launcher previously
shadowed it with a hardcoded 0.8 default.
Match the PR#241 (435641, 12/12 NIAH) golden config, cross-validated via
/proc<pid>/cmdline+environ on live run 435862 (12/12 NIAH 50K-280K x depths):
- rails: drop mlx5_9 -> 7-rail (mlx5_9 RoCE rail caused MoRI-EP barrier crash)
- kv-cache-memory-bytes: 8e9
- gpu-memory-utilization: 0.68
- prefill all2all-backend: mori_high_throughput (decode stays mori_low_latency)
- decode cudagraph_mode: FULL_AND_PIECEWISE (prefill stays NONE)
- int4 MoE requant (--quantization-config) + prefetch + max-num-seqs 8
Add mlx5_9 to NCCL_IB_HCA and MORI_RDMA_DEVICES (7-rail -> 8-rail). NCCL_IB_GID_INDEX=3 already correct for mlx5_9; MoRI auto-selects the same GID.

Validated on a fresh 4-node 2P2D quad (warm JIT cache):
- Boot: NCCL init + MoRI shmem/EpDispatchCombine all OK on mlx5_9; router Add Prefill/Add Decode; served P->D completions.
- Accuracy: single-request NIAH 12/12 (50K/100K/200K/280K x depths 0.1/0.5/0.9); concurrency NIAH 57/57 (con=1/8/16/32 @50k) — matches PR#241.

The prior 7-rail workaround was a false positive: prefill (CUDAGRAPH_MODE=NONE) comes up fast and busy-spins in the DP coordinator (~200%% CPU, GPU 0%%) while decode (FULL_AND_PIECEWISE) captures cudagraphs (~5-7 min, longer on cold cache) — mistaken for a hang. mlx5_9 is healthy cluster-wide (RoCEv2 IPv4 GID idx 3, ACTIVE on all probed nodes; ib_write_bw ~385 Gb/s).
Replace hardcoded MODEL_NAME==Kimi-K3-MXFP4 checks in shared connector/
topology code with a generic EP_TP_SIZE recipe knob (TP within each EP
pool; unset/1 = plain wideEP, the historical -tp 1 layout).

P0:
- models.yaml: add EP_TP_SIZE=2 to the Kimi-K3 recipe env block
- vllm_disagg.sh: early-resolve EP_TP_SIZE from the recipe (env/-e wins) so
  topology math can size per-node DP ranks; generic WIDE_EP+EP_TP_SIZE>1 gate
- connectors/moriio.sh: drop _moriio_is_kimik3(); gate kv-transfer-config,
  tensor-parallel sizing, api-server-count, and router dp on (( EP_TP_SIZE>1 ))
P2: fix broken out-of-tree doc links (README/ARCHITECTURE -> PR#241 refs)
P4: adjacency-based argv_assert checks; forward PREFILL_CUDAGRAPH_MODE

Validated: argv_assert 37/37 (byte-identical to pre-refactor K3 EP16; DeepSeek
stays -tp 1). Live 4-node MI300X 8-rail disagg (TP2xDP8->EP16): NIAH single
12/12 (2K-100K x depths 0.1/0.5/0.9), concurrency 56/57 (the one miss a
greedy-decode formatting artifact, correct needle).
…y guard

Follow-up to 6e4ab58: remove the remaining MODEL_NAME==Kimi-K3-MXFP4 control-flow
coupling from the shared slurm/interactive launchers and harden the generic knob.

P1 (run_xPyD_models.slurm):
- resolve EP_TP_SIZE from the recipe at submit time (env/-e wins), mirroring
  vllm_disagg.sh; gate the wideEP+moriio combo check on (( EP_TP_SIZE>1 ))
- drop the K3-specific xP=2/yD=2 and NUM_NODES==4 locks; topology is generic
  (xP+yD) and correctness is enforced by the divisibility guard below

P3 (run_xPyD_models.slurm + tests/run_interactive.sh):
- role-split the persistent JIT cache when (( EP_TP_SIZE>1 )) instead of by model
  name; knob renamed JIT_CACHE_SPLIT_K3 -> JIT_CACHE_SPLIT_ROLE

Guard (vllm_disagg.sh):
- fail fast if EP_TP_SIZE does not divide GPUS_PER_NODE or the per-pool DP sizes

Tests (tests/argv_assert.sh):
- assert -e EP_TP_SIZE=1 override beats the recipe (K3 -> plain wideEP -tp 1)
- assert the divisibility guard rejects an indivisible EP_TP_SIZE (=3 on 8 GPUs)

Verified offline: argv_assert 40/40; bash -n on all touched files; resolve/split
simulation shows K3->EP_TP_SIZE=2 (JIT role-split ON, unchanged), DeepSeek/Llama->1
(split OFF, unchanged), -e EP_TP_SIZE=1 forces OFF. No residual name branching in
shared code (only the standard VALID_MODELS allowlist entries remain).
…logy/JIT

- README/ARCHITECTURE: reference in-tree docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile
- ARCHITECTURE: correct stale claims (topology no longer slurm-locked; JIT split gated on EP_TP_SIZE>1, not MODEL_NAME)
- argv_assert.sh: add non-K3 wideEP dormancy + parametric EP_TP_SIZE fuzz rows (40->56 tests)
K3-specialized sibling of docker/vllm_disagg_inference.ubuntu.amd.Dockerfile.
Byte-identical to the PR#241 standalone recipe (MoRI v1.2.2 + AITER 0.1.19 + K3-tuned FlyDSL).
@MIR-AMD
MIR-AMD marked this pull request as ready for review September 22, 2026 17:16
Copilot AI lite review requested due to automatic review settings September 22, 2026 17:16

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved credential exposure, runtime topology/env propagation, test coverage, and Docker reproducibility issues remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 5 Medium severity · 1 Low severity

Open (7)
What changed in this PR

Adds Kimi-K3-MXFP4 support to vLLM disaggregated serving with MoRI-EP/MoRIIO, TP2×DP8 topology, launcher integration, tests, documentation, and a specialized Docker image.

Changes:

  • Adds K3 recipe, topology handling, allowlists, and cache splitting.
  • Adds MoRIIO KV routing and interactive launch support.
  • Adds offline assertions, architecture documentation, and Docker build support.
File Reviewed changes
scripts/​vllm_dissag/​vllm_disagg.sh K3 pod-host handling, DP sizing, and model configuration parsing
scripts/​vllm_dissag/​tests/​run_interactive.sh Interactive cache and container setup; EP_TP_SIZE forwarding requires correction
scripts/​vllm_dissag/​tests/​drive_cell.sh Remote environment forwarding; topology and override variables require correction
scripts/​vllm_dissag/​tests/​argv_assert.sh K3 and regression argument assertions
scripts/​vllm_dissag/​run_xPyD_models.slurm Allowlists, topology validation, cache handling, and container launch; gate coverage and env forwarding require changes
scripts/​vllm_dissag/​README.MD K3 usage and image documentation; Slurm invocation and model-path guidance require correction
scripts/​vllm_dissag/​models.yaml K3 recipe and runtime configuration; GID parity requires correction
scripts/​vllm_dissag/​connectors/​rixl.sh Safe model-argument parsing
scripts/​vllm_dissag/​connectors/​moriio.sh K3 MoRIIO gating, KV routing, and router topology
scripts/​vllm_dissag/​ARCHITECTURE.md K3 worker topology documentation
docker/​vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile K3 runtime image build; credential handling, reproducibility, dependencies, and toolchain validation require changes

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile
Comment thread docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile
Comment thread scripts/vllm_dissag/run_xPyD_models.slurm
Comment thread scripts/vllm_dissag/run_xPyD_models.slurm
Comment thread scripts/vllm_dissag/tests/drive_cell.sh
Comment thread scripts/vllm_dissag/tests/run_interactive.sh
Comment thread docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile
The K3 image bakes TRITON_CACHE_DIR/VLLM_CACHE_ROOT/COMGR_CACHE_DIR/AITER_JIT_DIR
to /opt/vllm_cache, so this connector fallback never fires for K3. Reverting to
develop-s /tmp default keeps non-K3 moriio models (DeepSeek/Llama) byte-identical
to develop. No effect on the JIT role-split (image env + /opt mount, not this fallback).
Copilot AI review requested due to automatic review settings September 22, 2026 18:52

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Unresolved moderate findings affect credential safety, reproducibility, runtime configuration, and launch validation.

Review effort: Lite
Findings: 1 High severity · 5 Medium severity · 1 Low severity

Open (7)
Previously missed (5)

In code that hasn't changed since last review

Medium severity Mutable dependency branches make the default image unreproducible

docker/​vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile:191

The default image build is not reproducible: VLLM_REF is a mutable branch, and this Dockerfile also builds from a router branch plus the default branches of DeepEP and rocm-systems. A later rebuild can silently change the validated kernels and runtime behavior even though versions.txt only records the new result after checkout. Pin the defaults to immutable commits (or verify them against a lock) before treating this image as the validated K3 artifact.

Medium severity Offline gate allowlists omit the newly supported K3 model

scripts/​vllm_dissag/​run_xPyD_models.slurm:90

This adds K3 to the runtime allowlists, but tests/gate_check.sh still mirrors the old VALID_MODELS, MORI_EP_VALID_MODELS, and WIDE_EP_ONLY_MODELS lists without K3 (tests/gate_check.sh:30-34). The documented offline gate will therefore reject the newly supported K3/moriio-wideEP combination; update the mirror and its cases.

This issue also appears on line 652 of the same file.

Medium severity drive_cell.sh strips EP_TP_SIZE from the interactive environment

scripts/​vllm_dissag/​tests/​drive_cell.sh:42

drive_cell.sh reconstructs the environment from this allowlist before invoking run_interactive.sh, but EP_TP_SIZE is omitted. Thus a caller's explicit TP-within-EP override is stripped before the interactive launcher can forward it, even after the container-side propagation is fixed. Include EP_TP_SIZE in FWD.

Medium severity Interactive docker launch omits EP_TP_SIZE

scripts/​vllm_dissag/​tests/​run_interactive.sh:126

The interactive launcher uses the resolved value to split the host JIT cache, but this docker run environment list does not pass EP_TP_SIZE into the container. Consequently run_interactive.sh ... EP_TP_SIZE=1 still launches the K3 worker with the recipe's TP2 topology, while using the host's unsplit cache. Add the variable to the container environment.

Low severity Documented MODEL_PATH is ignored by the Slurm launcher

scripts/​vllm_dissag/​README.MD:75

The documented command passes an arbitrary MODEL_PATH, but run_xPyD_models.slurm overwrites that variable and only searches its hard-coded roots plus MODEL_DIR (lines 322-346). As written, this example cannot launch weights from /path/to/... unless they happen to be in one of those roots; make explicit MODEL_PATH take precedence or document the supported MODEL_DIR form.

rixl is not on the wideEP/K3 moriio path; the eval->helper swap was an unrelated
drive-by cleanup. Reverting to develop keeps the PR diff K3-scoped. The shared
_model_config_to_array helper stays (used by moriio.sh). Behaviorally identical
(64-cell dry-run parity showed rixl/deepep argv byte-identical).
Copilot AI review requested due to automatic review settings September 22, 2026 19:17

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved configuration, catalog integration, connector compatibility, and dependency pinning issues remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Low severity

Open (1)
Resolved since last review (7)
Previously missed (3)

In code that hasn't changed since last review

Medium severity Pin mutable source refs to validated immutable commits

docker/​vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile:191

VLLM_REF defaults to the mutable kimi-k3-wideep-disagg-fullsource-v3 branch, while this image contains the validated vLLM correctness fixes; the default build can therefore silently change after validation. The same Dockerfile also pulls mutable DeepEP/default and rocm-systems develop sources. Pin these defaults to immutable validated commits (or require immutable refs) and update them deliberately.

Medium severity Apply the K3 MORI_IB_GID_INDEX value on Slurm launches

scripts/​vllm_dissag/​models.yaml:267

The K3 recipe's MORI_IB_GID_INDEX: "3" does not take effect on the documented Slurm path. run_xPyD_models.slurm forwards connectors/moriio.env as -e MORI_IB_GID_INDEX=1 before the container starts, and vllm_disagg.sh deliberately skips any YAML key already present in the environment, so moriio.sh retains 1 instead of the validated K3 value 3. Align the connector default with the K3 recipe or add an explicit model-aware forwarding/precedence path; otherwise the advertised K3 Slurm launch uses the wrong RDMA GID.

Low severity Update gate_check allowlists and K3 cases

scripts/​vllm_dissag/​run_xPyD_models.slurm:90

tests/gate_check.sh is a hand-maintained mirror of these allowlists, but its VALID_MODELS, MORI_EP_VALID_MODELS, and WIDE_EP_ONLY_MODELS arrays still omit Kimi-K3-MXFP4. The production gate is therefore not exercised for this newly enabled model, so the gate test can remain green while its K3 allow/reject behavior regresses. Update the mirror and add K3 combo cases, or make the test execute the production gate directly.

Comment thread scripts/vllm_dissag/vllm_disagg.sh
Cemberk
Cemberk previously approved these changes Sep 22, 2026

@Cemberk Cemberk left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM 👍

…k3-framework-only

# Conflicts:
#	scripts/vllm_dissag/README.MD
#	scripts/vllm_dissag/connectors/moriio.sh
#	scripts/vllm_dissag/models.yaml
#	scripts/vllm_dissag/run_xPyD_models.slurm
Copilot AI review requested due to automatic review settings September 25, 2026 15:23

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Model discovery and gate allowlists are incomplete, with additional environment-forwarding and topology-validation issues to resolve.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity · 1 Medium severity · 1 Low severity

Open (4)

Comment thread scripts/vllm_dissag/README.MD
Comment on lines 90 to +99
@@ -95,6 +96,7 @@ MORI_EP_VALID_MODELS=( \
"DeepSeek-V3" \
"DeepSeek-V3-5layer" \
"DeepSeek-R1" \
"Kimi-K3-MXFP4" \
Comment on lines +121 to +129
if [[ "${WIDE_EP:-0}" == "1" ]] && (( ${EP_TP_SIZE:-1} > 1 )); then
if (( _GPUS_PER_NODE % EP_TP_SIZE != 0 || PREFILL_DP_SIZE % EP_TP_SIZE != 0 || DECODE_DP_SIZE % EP_TP_SIZE != 0 )); then
echo "Error: EP_TP_SIZE=${EP_TP_SIZE} must divide GPUS_PER_NODE=${_GPUS_PER_NODE} and per-pool DP sizes (prefill=${PREFILL_DP_SIZE}, decode=${DECODE_DP_SIZE})." >&2
exit 1
fi
DP_PARALLEL_SIZE_LOCAL=$(( _GPUS_PER_NODE / ${EP_TP_SIZE} ))
PREFILL_DP_START_RANK=$(( NODE_RANK * DP_PARALLEL_SIZE_LOCAL ))
DECODE_DP_START_RANK=$(( (NODE_RANK - xP) * DP_PARALLEL_SIZE_LOCAL ))
fi
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants