You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Adds Kimi-K3-MXFP4 (MI300X gfx942, MXFP4 MoE, 2.8T params / 896 experts) to the vllm_dissag framework as a first-class 2P/2D disaggregated model — implemented as a generic framework capability, not a model-specific fork.
The core change: the old hardcoded MODEL_NAME=="Kimi-K3-MXFP4" branches in shared code are replaced by a generic EP_TP_SIZE knob (TP-within-EP). Unset / 1 → plain wideEP (-tp 1; today's behavior for DeepSeek/Llama/Qwen); N>1 → TP inside each EP pool (K3 = 2 → TP2×DP8 → EP16). K3 sets it only in its models.yaml recipe env — shared code stays model-agnostic.
README.MD / ARCHITECTURE.md: docs for EP_TP_SIZE, generic topology, in-tree Dockerfile
docker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile: new in-tree K3 disagg image (K3-specialized sibling of the generalized vllm_disagg_inference.ubuntu.amd.Dockerfile)
Regression safety
No model-name decision branch remains in shared code — K3 specifics are confined to models.yaml. Non-K3 models (DeepSeek-V3/R1, Llama, Qwen) resolve EP_TP_SIZE → 1: plain wideEP, JIT split OFF, no topology lock.
Offline argv_assert.sh: 56 passed / 0 failed
64-cell develop-vs-branch dry-run diff: 64/64 byte-identical for non-K3 cells (dense TP moriio/rixl + DeepSeek wideEP mori/deepep, across ranks, 2P/2D)
Live validation (K3 2P/2D, MI300X + CX7)
Single-stream NIAH correctness: 12/12 at 50K–280K context
Concurrency NIAH: 57/57 across con = {1, 8, 16, 32}
Recipe pinned to the PR#241-validated config: gpu-memory-utilization 0.68, kv-cache-memory-bytes 8e9, per-role mori_high_throughput/mori_low_latency, decode cudagraph_mode FULL_AND_PIECEWISE / prefill NONE
The 0.85 recipe value fails the vLLM startup free-memory check on 192GB
MI300X GPUs (needs 163.19 GiB free; ~31 GiB framework baseline leaves
only ~160 GiB), aborting boot. 0.80 is the proven value (prefill boots:
KV cache allocated, all-8 RDMA rails init, MoRIIO connector up). The
0.85 value was never actually exercised because the launcher previously
shadowed it with a hardcoded 0.8 default.
Add mlx5_9 to NCCL_IB_HCA and MORI_RDMA_DEVICES (7-rail -> 8-rail). NCCL_IB_GID_INDEX=3 already correct for mlx5_9; MoRI auto-selects the same GID.
Validated on a fresh 4-node 2P2D quad (warm JIT cache):
- Boot: NCCL init + MoRI shmem/EpDispatchCombine all OK on mlx5_9; router Add Prefill/Add Decode; served P->D completions.
- Accuracy: single-request NIAH 12/12 (50K/100K/200K/280K x depths 0.1/0.5/0.9); concurrency NIAH 57/57 (con=1/8/16/32 @50k) — matches PR#241.
The prior 7-rail workaround was a false positive: prefill (CUDAGRAPH_MODE=NONE) comes up fast and busy-spins in the DP coordinator (~200%% CPU, GPU 0%%) while decode (FULL_AND_PIECEWISE) captures cudagraphs (~5-7 min, longer on cold cache) — mistaken for a hang. mlx5_9 is healthy cluster-wide (RoCEv2 IPv4 GID idx 3, ACTIVE on all probed nodes; ib_write_bw ~385 Gb/s).
…y guard
Follow-up to 6e4ab58: remove the remaining MODEL_NAME==Kimi-K3-MXFP4 control-flow
coupling from the shared slurm/interactive launchers and harden the generic knob.
P1 (run_xPyD_models.slurm):
- resolve EP_TP_SIZE from the recipe at submit time (env/-e wins), mirroring
vllm_disagg.sh; gate the wideEP+moriio combo check on (( EP_TP_SIZE>1 ))
- drop the K3-specific xP=2/yD=2 and NUM_NODES==4 locks; topology is generic
(xP+yD) and correctness is enforced by the divisibility guard below
P3 (run_xPyD_models.slurm + tests/run_interactive.sh):
- role-split the persistent JIT cache when (( EP_TP_SIZE>1 )) instead of by model
name; knob renamed JIT_CACHE_SPLIT_K3 -> JIT_CACHE_SPLIT_ROLE
Guard (vllm_disagg.sh):
- fail fast if EP_TP_SIZE does not divide GPUS_PER_NODE or the per-pool DP sizes
Tests (tests/argv_assert.sh):
- assert -e EP_TP_SIZE=1 override beats the recipe (K3 -> plain wideEP -tp 1)
- assert the divisibility guard rejects an indivisible EP_TP_SIZE (=3 on 8 GPUs)
Verified offline: argv_assert 40/40; bash -n on all touched files; resolve/split
simulation shows K3->EP_TP_SIZE=2 (JIT role-split ON, unchanged), DeepSeek/Llama->1
(split OFF, unchanged), -e EP_TP_SIZE=1 forces OFF. No residual name branching in
shared code (only the standard VALID_MODELS allowlist entries remain).
Adds Kimi-K3-MXFP4 support to vLLM disaggregated serving with MoRI-EP/MoRIIO, TP2×DP8 topology, launcher integration, tests, documentation, and a specialized Docker image.
Changes:
Adds K3 recipe, topology handling, allowlists, and cache splitting.
Adds MoRIIO KV routing and interactive launch support.
Adds offline assertions, architecture documentation, and Docker build support.
File
Reviewed changes
scripts/vllm_dissag/vllm_disagg.sh
K3 pod-host handling, DP sizing, and model configuration parsing
scripts/vllm_dissag/tests/run_interactive.sh
Interactive cache and container setup; EP_TP_SIZE forwarding requires correction
scripts/vllm_dissag/tests/drive_cell.sh
Remote environment forwarding; topology and override variables require correction
scripts/vllm_dissag/tests/argv_assert.sh
K3 and regression argument assertions
scripts/vllm_dissag/run_xPyD_models.slurm
Allowlists, topology validation, cache handling, and container launch; gate coverage and env forwarding require changes
scripts/vllm_dissag/README.MD
K3 usage and image documentation; Slurm invocation and model-path guidance require correction
scripts/vllm_dissag/models.yaml
K3 recipe and runtime configuration; GID parity requires correction
The K3 image bakes TRITON_CACHE_DIR/VLLM_CACHE_ROOT/COMGR_CACHE_DIR/AITER_JIT_DIR
to /opt/vllm_cache, so this connector fallback never fires for K3. Reverting to
develop-s /tmp default keeps non-K3 moriio models (DeepSeek/Llama) byte-identical
to develop. No effect on the JIT role-split (image env + /opt mount, not this fallback).
The default image build is not reproducible: VLLM_REF is a mutable branch, and this Dockerfile also builds from a router branch plus the default branches of DeepEP and rocm-systems. A later rebuild can silently change the validated kernels and runtime behavior even though versions.txt only records the new result after checkout. Pin the defaults to immutable commits (or verify them against a lock) before treating this image as the validated K3 artifact.
Offline gate allowlists omit the newly supported K3 model
scripts/vllm_dissag/run_xPyD_models.slurm:90
This adds K3 to the runtime allowlists, but tests/gate_check.sh still mirrors the old VALID_MODELS, MORI_EP_VALID_MODELS, and WIDE_EP_ONLY_MODELS lists without K3 (tests/gate_check.sh:30-34). The documented offline gate will therefore reject the newly supported K3/moriio-wideEP combination; update the mirror and its cases.
This issue also appears on line 652 of the same file.
drive_cell.sh strips EP_TP_SIZE from the interactive environment
scripts/vllm_dissag/tests/drive_cell.sh:42
drive_cell.sh reconstructs the environment from this allowlist before invoking run_interactive.sh, but EP_TP_SIZE is omitted. Thus a caller's explicit TP-within-EP override is stripped before the interactive launcher can forward it, even after the container-side propagation is fixed. Include EP_TP_SIZE in FWD.
The interactive launcher uses the resolved value to split the host JIT cache, but this docker run environment list does not pass EP_TP_SIZE into the container. Consequently run_interactive.sh ... EP_TP_SIZE=1 still launches the K3 worker with the recipe's TP2 topology, while using the host's unsplit cache. Add the variable to the container environment.
Documented MODEL_PATH is ignored by the Slurm launcher
scripts/vllm_dissag/README.MD:75
The documented command passes an arbitrary MODEL_PATH, but run_xPyD_models.slurm overwrites that variable and only searches its hard-coded roots plus MODEL_DIR (lines 322-346). As written, this example cannot launch weights from /path/to/... unless they happen to be in one of those roots; make explicit MODEL_PATH take precedence or document the supported MODEL_DIR form.
rixl is not on the wideEP/K3 moriio path; the eval->helper swap was an unrelated
drive-by cleanup. Reverting to develop keeps the PR diff K3-scoped. The shared
_model_config_to_array helper stays (used by moriio.sh). Behaviorally identical
(64-cell dry-run parity showed rixl/deepep argv byte-identical).
VLLM_REF defaults to the mutable kimi-k3-wideep-disagg-fullsource-v3 branch, while this image contains the validated vLLM correctness fixes; the default build can therefore silently change after validation. The same Dockerfile also pulls mutable DeepEP/default and rocm-systemsdevelop sources. Pin these defaults to immutable validated commits (or require immutable refs) and update them deliberately.
Apply the K3 MORI_IB_GID_INDEX value on Slurm launches
scripts/vllm_dissag/models.yaml:267
The K3 recipe's MORI_IB_GID_INDEX: "3" does not take effect on the documented Slurm path. run_xPyD_models.slurm forwards connectors/moriio.env as -e MORI_IB_GID_INDEX=1 before the container starts, and vllm_disagg.sh deliberately skips any YAML key already present in the environment, so moriio.sh retains 1 instead of the validated K3 value 3. Align the connector default with the K3 recipe or add an explicit model-aware forwarding/precedence path; otherwise the advertised K3 Slurm launch uses the wrong RDMA GID.
Update gate_check allowlists and K3 cases
scripts/vllm_dissag/run_xPyD_models.slurm:90
tests/gate_check.sh is a hand-maintained mirror of these allowlists, but its VALID_MODELS, MORI_EP_VALID_MODELS, and WIDE_EP_ONLY_MODELS arrays still omit Kimi-K3-MXFP4. The production gate is therefore not exercised for this newly enabled model, so the gate test can remain green while its K3 allow/reject behavior regresses. Update the mirror and add K3 combo cases, or make the test execute the production gate directly.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds Kimi-K3-MXFP4 (MI300X gfx942, MXFP4 MoE, 2.8T params / 896 experts) to the
vllm_dissagframework as a first-class 2P/2D disaggregated model — implemented as a generic framework capability, not a model-specific fork.The core change: the old hardcoded
MODEL_NAME=="Kimi-K3-MXFP4"branches in shared code are replaced by a genericEP_TP_SIZEknob (TP-within-EP). Unset /1→ plain wideEP (-tp 1; today's behavior for DeepSeek/Llama/Qwen);N>1→ TP inside each EP pool (K3 =2→ TP2×DP8 → EP16). K3 sets it only in itsmodels.yamlrecipe env — shared code stays model-agnostic.Changes (11 files, +832 −27)
models.yaml: K3 recipe env anchor + model entry (EP_TP_SIZE, RDMA fabric, per-role cudagraph/all2all, KV budget, PR#241-parity env)connectors/moriio.sh:_moriio_is_kimik3branches → generic(( EP_TP_SIZE>1 ))gates; per-role all2all (prefillmori_high_throughput/ decodemori_low_latency) + per-role cudagraphconnectors/rixl.sh: unsafeeval→ shared_model_config_to_arrayhelpervllm_disagg.sh: earlyEP_TP_SIZEresolve + divisibility guard +PREFILL/DECODE_POD_HOSTSexportrun_xPyD_models.slurm: genericNUM_NODES=xP+yD(no K3 topology lock) +EP_TP_SIZE>1-gated JIT cache role-split + allowliststests/run_interactive.sh,tests/drive_cell.sh:EP_TP_SIZEresolve on the interactive pathtests/argv_assert.sh: 40 → 56 offline assertions (K3 TP2×DP8 + non-K3 wideEP dormancy + parametricEP_TP_SIZEfuzz)README.MD/ARCHITECTURE.md: docs forEP_TP_SIZE, generic topology, in-tree Dockerfiledocker/vllm_disagg_inference.kimik3.ubuntu.amd.Dockerfile: new in-tree K3 disagg image (K3-specialized sibling of the generalizedvllm_disagg_inference.ubuntu.amd.Dockerfile)Regression safety
No model-name decision branch remains in shared code — K3 specifics are confined to
models.yaml. Non-K3 models (DeepSeek-V3/R1, Llama, Qwen) resolveEP_TP_SIZE → 1: plain wideEP, JIT split OFF, no topology lock.argv_assert.sh: 56 passed / 0 faileddevelop-vs-branch dry-run diff: 64/64 byte-identical for non-K3 cells (dense TP moriio/rixl + DeepSeek wideEP mori/deepep, across ranks, 2P/2D)Live validation (K3 2P/2D, MI300X + CX7)
gpu-memory-utilization 0.68,kv-cache-memory-bytes 8e9, per-rolemori_high_throughput/mori_low_latency, decodecudagraph_mode FULL_AND_PIECEWISE/ prefillNONEPerformance
Measured on 4× MI300X (2P/2D), MI300X + CX7 8-rail RoCE, 64 output tokens, temp 0.
End-to-end request latency (single stream, by context length):
Throughput under concurrency (50K context, distinct-needle recall):
32 concurrent 50K-context requests complete in ~163 s (≈4.9× faster than serial execution) at full recall (57/57).
Image
Built in-tree from the repo root:
Related
models.yaml/moriio.sh/slurm/tests — additive, resolvable (K3 first, GLM rebases)Test plan
bash -non all.shfilesmodels.yamlargv_assert.shoffline tests (56/56)develop(non-K3 argv unchanged)