Skip to content

feat(modelexpress-p2p): add Intel XPU (B60) host-staged transport guide - #2461

Draft
yao531441 wants to merge 1 commit into
llm-d:mainfrom
yao531441:feat/modelexpress-p2p-xpu-support
Draft

yao531441 wants to merge 1 commit into
llm-d:mainfrom
yao531441:feat/modelexpress-p2p-xpu-support

Conversation

@yao531441

Copy link
Copy Markdown
Contributor

Summary

  • Add an Intel XPU (Arc Pro B60) variant of the ModelExpress P2P guide (guides/modelexpress-p2p/README.xpu.md), following the repo's existing README.<variant>.md naming convention (e.g. pd-disaggregation/README.tpu.md, workload-autoscaling/README.hpa-epp.md) rather than a bare xpu.md.
  • Add a build image (image/Dockerfile.xpu) pinning PyTorch 2.12.0+xpu, vLLM 0.26.0+xpu, UCX (built with verbs + Level Zero support), NIXL 1.3.0, and ModelExpress 0.5.0, with build-time checks that the ZE modules, NIXL UCX plugin linkage, and package versions all match. The image inherits its ENTRYPOINT/CMD from the base image; every Deployment in this guide sets an explicit command.
  • Add base and dranet kustomize overlays under modelserver/xpu/vllm/ for a cross-node one-seed/one-receiver Qwen3-0.6B deployment: base requests one Intel GPU DRA claim per pod on the default pod network, dranet additionally requests an RDMA-capable NIC DRA claim to carry the transfer over a faster fabric. The dranet-only ConfigMap/ResourceClaimTemplate are name-prefixed the same way as the guide's other resources to avoid collisions with other guides sharing a namespace.
  • Update guides/modelexpress-p2p/README.md and guides/README.md to link the new Intel XPU variant.

The guide validates only UCX_TLS/NIXL_UCX_TLS=tcp,ze_copy and explicitly avoids
rc_mlx5/rc/dc/ib, because direct rc_mlx5 reads from B60 device memory silently corrupt
content-dependent regions: the RDMA operation completes with UCS_OK, but bytes read from
compressible/constant/repetitive source regions (padding, sparse, and quantized values are common
in model weights) come back wrong, while random regions in the same allocation pass SHA-256
verification. This reproduces for both GET and PUT, on multiple workers, and even with same-node
forced RC — it is not a network/IOMMU/firmware/routing configuration issue.

Why this is a draft

This PR ships only the TCP-based configuration above, which is correct and safe to merge as-is —
it does not itself depend on any external fix. It is marked draft because it tracks an in-flight
upstream dependency that changes what happens next:

  • openucx/ucx#11903 (depends on
    #11902) targets the same failure mode — the
    ze_copy memory domain advertises DMA-BUF export support for ze-device memory without
    requesting an export descriptor at allocation time. Both PRs are currently open/draft and
    unmerged upstream.
  • Plan: once a UCX release containing the fix is available, re-run the byte-for-byte checksum
    check (all-zero, constant, repeating, sparse, random, and mixed-content allocations, plus an
    independent checksum for every Qwen3-0.6B tensor) on this cluster's B60/ConnectX-7 hardware with
    direct RDMA enabled. If it passes, this PR is updated (or followed up) to switch the guide's
    default to direct XPU RDMA before it comes out of draft.

Testing

  • npx markdownlint-cli2 guides/modelexpress-p2p/README.xpu.md guides/modelexpress-p2p/README.md: 0 errors.
  • kubectl kustomize guides/modelexpress-p2p/modelserver/xpu/vllm/base and .../dranet both render cleanly, with all resource names and cross-references (resourceClaimTemplateName, configMap.name) resolving correctly, no double-prefixed or dangling references.
  • pytest scripts/tests/: 36 passed, 1 skipped (unchanged baseline).
  • scripts/sync-nightly-matrix.py --check: no diff (no CI workflow changes in this PR).
  • Deployed end-to-end on a 2-node Intel Arc Pro B60 cluster: the base overlay transferred 339 Qwen3-0.6B tensors (1.20 GB) in 10.73 s (0.9 Gbps) over a 1 Gbps management network; the dranet overlay completed the same transfer in 2.20 s (4.4 Gbps) over a 400 Gbps fabric while still forcing tcp,ze_copy (no verbs data lane). Both replicas stayed Ready with zero restarts, and inference output matched between seed and receiver.

Add an Intel XPU variant of the ModelExpress P2P guide that uses
host-staged TCP + Level Zero copies instead of direct XPU RDMA, since
direct mlx5 reads from B60 device memory can silently corrupt
content-dependent regions. Includes the validated build image
(Dockerfile.xpu), base and dranet kustomize overlays, and updates to
the top-level guide README to reference the new README.xpu.md guide.
yao531441 added a commit to yao531441/llm-d that referenced this pull request Sep 16, 2026
Add modelserver/xpu/vllm/base/ mirroring the GPU OffloadingConnector
P2P-tier setup, adapted for Intel XPU:

- DRA ResourceClaimTemplate (gpu.intel.com), 1 XPU per pod, matching
  the pattern used by precise-prefix-cache-routing and
  tiered-prefix-cache's Intel XPU overlays, instead of the GPU
  overlay's nvidia.com/gpu device-plugin request.
- CI-sized functional check: 2 replicas (1 source + 1 receiver) of
  Qwen/Qwen3-0.6B instead of the GPU overlay's 16x gpt-oss-120b -
  validates the same P2P pull mechanism without requiring 24GB+ of
  device memory per pod or reproducing the GPU benchmark's scale.
  cpu_bytes_to_use and the shm tier are sized down to match.
- Same routing-sidecar (patch-sidecar.yaml) and P2P/kv-events port
  layout as the GPU overlay - both are accelerator-agnostic.

No RDMA overlay (xpu/vllm/rdma) is added. Direct RDMA over Intel XPU
device memory currently hits an upstream UCX ze_copy/DMA-BUF
data-correctness bug (silently wrong bytes, reported success) that
guides/modelexpress-p2p's Intel XPU PR (llm-d#2461) already ran into and
documented; tracked at openucx/ucx#11902 and #11903. The TCP-only
side channel here does not depend on that fix and is safe to ship now.

README changes:
- Supported Hardware Backends: document the Intel XPU variant, its
  scope (functional check, not a benchmark), and why RDMA is deferred.
- Step 3 (Deploy the Model Server): ACCELERATOR_TYPE now documents
  xpu, and TRANSPORT=base is called out as the only option for xpu.

Verified guides/p2p-kv-cache-sharing/modelserver/xpu/vllm/base renders
cleanly with 'kubectl kustomize' (ServiceAccount, Deployment with
sidecar + engine container merged correctly, port 8000->8200 renamed,
ResourceClaimTemplate name-referenced correctly). Not yet deployed to
a live cluster.

Signed-off-by: Yao, Qing <qing.yao@intel.com>
yao531441 added a commit to yao531441/llm-d that referenced this pull request Sep 16, 2026
- Add UCX_TLS=tcp to the XPU overlay's vLLM env, matching the
  pd-disaggregation XPU precedent, so this TCP-only overlay can't
  silently fall back to the buggy ze_copy/DMA-BUF transport.
- Stop citing guides/modelexpress-p2p's README as documenting the
  Intel XPU UCX bug: that content is still an unmerged draft
  (llm-d#2461), not on main. Reference the draft PR directly.
- Scope the '16 replicas, TP=1' summary to the GPU overlay; the XPU
  overlay is 2 replicas of a different model.
- Note that this guide's router values hard-code modelName:
  openai/gpt-oss-120b, and must be changed to Qwen/Qwen3-0.6B before
  installing the router when deploying the XPU overlay, or the render
  Service and every downstream verification command target a model
  the XPU pods never load.

All four issues were raised by automated review on the upstream PR;
verified against the actual repo state before fixing.

Signed-off-by: Yao, Qing <qing.yao@intel.com>
yao531441 added a commit to yao531441/llm-d that referenced this pull request Sep 16, 2026
Traced the actual vLLM OffloadingConnector/TieringOffloadingSpec P2P
data path: every transfer is staged through its CPU-mmap-backed
offload tier, and NIXL registers only that host memory (DRAM) with
UCX. XPU device memory is never handed to UCX/NIXL directly, unlike
guides/modelexpress-p2p's Intel XPU variant (llm-d#2461), which
does register device memory for its weight transfer and genuinely
needs UCX_TLS=tcp,ze_copy.

So for this overlay:
- UCX_TLS=tcp was already the right pin, but the comment overstated
  the risk (implying UCX could still reach device memory here).
- The claim that a missing RDMA overlay is 'blocked' by the upstream
  UCX ze_copy/DMA-BUF bug (openucx/ucx#11902, #11903) doesn't hold for
  this connector: an RDMA transport here would still only move
  CPU-to-CPU DRAM, not touch XPU VRAM directly. It just hasn't been
  built/validated yet, for unrelated reasons.
- 'TCP-only side channel' conflated the P2P tier's ZMQ control channel
  with the NIXL/UCX data plane; described them separately.

Reworded the patch-vllm.yaml comment, kustomization.yaml comment, and
README Supported Hardware Backends section accordingly, keeping a
pointer to the real UCX bug and llm-d#2461 for readers evaluating direct
XPU-device RDMA elsewhere.

Signed-off-by: Yao, Qing <qing.yao@intel.com>
maugustosilva pushed a commit that referenced this pull request Sep 16, 2026
* feat(p2p-kv-cache-sharing): add Intel XPU (TCP-only) variant

Add modelserver/xpu/vllm/base/ mirroring the GPU OffloadingConnector
P2P-tier setup, adapted for Intel XPU:

- DRA ResourceClaimTemplate (gpu.intel.com), 1 XPU per pod, matching
  the pattern used by precise-prefix-cache-routing and
  tiered-prefix-cache's Intel XPU overlays, instead of the GPU
  overlay's nvidia.com/gpu device-plugin request.
- CI-sized functional check: 2 replicas (1 source + 1 receiver) of
  Qwen/Qwen3-0.6B instead of the GPU overlay's 16x gpt-oss-120b -
  validates the same P2P pull mechanism without requiring 24GB+ of
  device memory per pod or reproducing the GPU benchmark's scale.
  cpu_bytes_to_use and the shm tier are sized down to match.
- Same routing-sidecar (patch-sidecar.yaml) and P2P/kv-events port
  layout as the GPU overlay - both are accelerator-agnostic.

No RDMA overlay (xpu/vllm/rdma) is added. Direct RDMA over Intel XPU
device memory currently hits an upstream UCX ze_copy/DMA-BUF
data-correctness bug (silently wrong bytes, reported success) that
guides/modelexpress-p2p's Intel XPU PR (#2461) already ran into and
documented; tracked at openucx/ucx#11902 and #11903. The TCP-only
side channel here does not depend on that fix and is safe to ship now.

README changes:
- Supported Hardware Backends: document the Intel XPU variant, its
  scope (functional check, not a benchmark), and why RDMA is deferred.
- Step 3 (Deploy the Model Server): ACCELERATOR_TYPE now documents
  xpu, and TRANSPORT=base is called out as the only option for xpu.

Verified guides/p2p-kv-cache-sharing/modelserver/xpu/vllm/base renders
cleanly with 'kubectl kustomize' (ServiceAccount, Deployment with
sidecar + engine container merged correctly, port 8000->8200 renamed,
ResourceClaimTemplate name-referenced correctly). Not yet deployed to
a live cluster.

* p2p-kv-cache-sharing: document XPU cluster verification results

Tested the xpu/vllm/base overlay on XPU-8xB60-817225 (real Intel Arc Pro
B60 hardware). DRA allocation, sidecar init, and vLLM boot with
OffloadingConnector + NIXL/UCX + P2P secondary tier all verified working.

However, the P2P pull itself is not yet verified: the pinned XPU image
(llm-d-xpu:v0.9.0, vLLM 0.26.0) has no remote_kv_source handling in its
OffloadingConnector (only max_offload_tokens), unlike the GPU overlay's
vllm-openai:v0.27.1. Documented this version-gap limitation in the guide
README.

* p2p-kv-cache-sharing: confirm XPU P2P pull works with nightly vLLM image

Manual pull test on real Intel Arc Pro B60 hardware (XPU-8xB60-817225)
confirms the P2P pull mechanism itself works correctly once the pinned
vLLM version carries vllm/v1/kv_offload/tiering/p2p/. Swapping the test
deployment's image from ghcr.io/llm-d/llm-d-xpu:v0.9.0 (vLLM 0.26.0, no
remote_kv_source support) to docker.io/vllm/vllm-openai-xpu:nightly
(vLLM 0.29.1rc1) made a 4096-token prefix pull succeed end-to-end
(external_prefix_cache_hits_total incremented by exactly 4096,
reproduced twice).

Updates the README's Intel XPU warning to reflect this: the gap is
purely the pinned image's vLLM version, not the manifests or the pull
mechanism, with a workaround (temporarily point the xpu-vllm component
at nightly) documented until ghcr.io/llm-d/llm-d-xpu is rebuilt against
a newer vLLM release.

* p2p-kv-cache-sharing: default XPU overlay to nightly xpu-vllm image

ghcr.io/llm-d/llm-d-xpu:v0.9.0 (vLLM 0.26.0) has no remote_kv_source
handling in its OffloadingConnector, so the P2P pull this guide exists
to demonstrate cannot work on it. Point the overlay at the existing
xpu-vllm/nightly component (vLLM main) instead, which was verified
end-to-end on real Intel Arc Pro B60 hardware, rather than shipping an
overlay that starts cleanly but can't do the one thing the guide is
about. Switch back to the llm-d component once ghcr.io/llm-d/llm-d-xpu
is rebuilt against a vLLM release carrying
vllm/v1/kv_offload/tiering/p2p/.

Also trims the README's Intel XPU warning now that this is the
default rather than a documented manual workaround.

Signed-off-by: Yao, Qing <qing.yao@intel.com>

* p2p-kv-cache-sharing: fix XPU review findings

- Add UCX_TLS=tcp to the XPU overlay's vLLM env, matching the
  pd-disaggregation XPU precedent, so this TCP-only overlay can't
  silently fall back to the buggy ze_copy/DMA-BUF transport.
- Stop citing guides/modelexpress-p2p's README as documenting the
  Intel XPU UCX bug: that content is still an unmerged draft
  (#2461), not on main. Reference the draft PR directly.
- Scope the '16 replicas, TP=1' summary to the GPU overlay; the XPU
  overlay is 2 replicas of a different model.
- Note that this guide's router values hard-code modelName:
  openai/gpt-oss-120b, and must be changed to Qwen/Qwen3-0.6B before
  installing the router when deploying the XPU overlay, or the render
  Service and every downstream verification command target a model
  the XPU pods never load.

All four issues were raised by automated review on the upstream PR;
verified against the actual repo state before fixing.

* p2p-kv-cache-sharing: correct UCX bug attribution for XPU overlay

Traced the actual vLLM OffloadingConnector/TieringOffloadingSpec P2P
data path: every transfer is staged through its CPU-mmap-backed
offload tier, and NIXL registers only that host memory (DRAM) with
UCX. XPU device memory is never handed to UCX/NIXL directly, unlike
guides/modelexpress-p2p's Intel XPU variant (#2461), which
does register device memory for its weight transfer and genuinely
needs UCX_TLS=tcp,ze_copy.

So for this overlay:
- UCX_TLS=tcp was already the right pin, but the comment overstated
  the risk (implying UCX could still reach device memory here).
- The claim that a missing RDMA overlay is 'blocked' by the upstream
  UCX ze_copy/DMA-BUF bug (openucx/ucx#11902, #11903) doesn't hold for
  this connector: an RDMA transport here would still only move
  CPU-to-CPU DRAM, not touch XPU VRAM directly. It just hasn't been
  built/validated yet, for unrelated reasons.
- 'TCP-only side channel' conflated the P2P tier's ZMQ control channel
  with the NIXL/UCX data plane; described them separately.

Reworded the patch-vllm.yaml comment, kustomization.yaml comment, and
README Supported Hardware Backends section accordingly, keeping a
pointer to the real UCX bug and #2461 for readers evaluating direct
XPU-device RDMA elsewhere.

* p2p-kv-cache-sharing: tighten XPU overlay wording

Trim hedging phrasing ('simply because', 'nothing about this path
blocks it', 'anyway') from the README, kustomization.yaml, and
patch-vllm.yaml so the Intel XPU transport description reads as a
final statement of fact rather than a running commentary.

---------

Signed-off-by: Yao, Qing <qing.yao@intel.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant