Conversation
Add an Intel XPU variant of the ModelExpress P2P guide that uses host-staged TCP + Level Zero copies instead of direct XPU RDMA, since direct mlx5 reads from B60 device memory can silently corrupt content-dependent regions. Includes the validated build image (Dockerfile.xpu), base and dranet kustomize overlays, and updates to the top-level guide README to reference the new README.xpu.md guide.
yao531441
added a commit
to yao531441/llm-d
that referenced
this pull request
Sep 16, 2026
Add modelserver/xpu/vllm/base/ mirroring the GPU OffloadingConnector P2P-tier setup, adapted for Intel XPU: - DRA ResourceClaimTemplate (gpu.intel.com), 1 XPU per pod, matching the pattern used by precise-prefix-cache-routing and tiered-prefix-cache's Intel XPU overlays, instead of the GPU overlay's nvidia.com/gpu device-plugin request. - CI-sized functional check: 2 replicas (1 source + 1 receiver) of Qwen/Qwen3-0.6B instead of the GPU overlay's 16x gpt-oss-120b - validates the same P2P pull mechanism without requiring 24GB+ of device memory per pod or reproducing the GPU benchmark's scale. cpu_bytes_to_use and the shm tier are sized down to match. - Same routing-sidecar (patch-sidecar.yaml) and P2P/kv-events port layout as the GPU overlay - both are accelerator-agnostic. No RDMA overlay (xpu/vllm/rdma) is added. Direct RDMA over Intel XPU device memory currently hits an upstream UCX ze_copy/DMA-BUF data-correctness bug (silently wrong bytes, reported success) that guides/modelexpress-p2p's Intel XPU PR (llm-d#2461) already ran into and documented; tracked at openucx/ucx#11902 and #11903. The TCP-only side channel here does not depend on that fix and is safe to ship now. README changes: - Supported Hardware Backends: document the Intel XPU variant, its scope (functional check, not a benchmark), and why RDMA is deferred. - Step 3 (Deploy the Model Server): ACCELERATOR_TYPE now documents xpu, and TRANSPORT=base is called out as the only option for xpu. Verified guides/p2p-kv-cache-sharing/modelserver/xpu/vllm/base renders cleanly with 'kubectl kustomize' (ServiceAccount, Deployment with sidecar + engine container merged correctly, port 8000->8200 renamed, ResourceClaimTemplate name-referenced correctly). Not yet deployed to a live cluster. Signed-off-by: Yao, Qing <qing.yao@intel.com>
yao531441
added a commit
to yao531441/llm-d
that referenced
this pull request
Sep 16, 2026
- Add UCX_TLS=tcp to the XPU overlay's vLLM env, matching the pd-disaggregation XPU precedent, so this TCP-only overlay can't silently fall back to the buggy ze_copy/DMA-BUF transport. - Stop citing guides/modelexpress-p2p's README as documenting the Intel XPU UCX bug: that content is still an unmerged draft (llm-d#2461), not on main. Reference the draft PR directly. - Scope the '16 replicas, TP=1' summary to the GPU overlay; the XPU overlay is 2 replicas of a different model. - Note that this guide's router values hard-code modelName: openai/gpt-oss-120b, and must be changed to Qwen/Qwen3-0.6B before installing the router when deploying the XPU overlay, or the render Service and every downstream verification command target a model the XPU pods never load. All four issues were raised by automated review on the upstream PR; verified against the actual repo state before fixing. Signed-off-by: Yao, Qing <qing.yao@intel.com>
yao531441
added a commit
to yao531441/llm-d
that referenced
this pull request
Sep 16, 2026
Traced the actual vLLM OffloadingConnector/TieringOffloadingSpec P2P data path: every transfer is staged through its CPU-mmap-backed offload tier, and NIXL registers only that host memory (DRAM) with UCX. XPU device memory is never handed to UCX/NIXL directly, unlike guides/modelexpress-p2p's Intel XPU variant (llm-d#2461), which does register device memory for its weight transfer and genuinely needs UCX_TLS=tcp,ze_copy. So for this overlay: - UCX_TLS=tcp was already the right pin, but the comment overstated the risk (implying UCX could still reach device memory here). - The claim that a missing RDMA overlay is 'blocked' by the upstream UCX ze_copy/DMA-BUF bug (openucx/ucx#11902, #11903) doesn't hold for this connector: an RDMA transport here would still only move CPU-to-CPU DRAM, not touch XPU VRAM directly. It just hasn't been built/validated yet, for unrelated reasons. - 'TCP-only side channel' conflated the P2P tier's ZMQ control channel with the NIXL/UCX data plane; described them separately. Reworded the patch-vllm.yaml comment, kustomization.yaml comment, and README Supported Hardware Backends section accordingly, keeping a pointer to the real UCX bug and llm-d#2461 for readers evaluating direct XPU-device RDMA elsewhere. Signed-off-by: Yao, Qing <qing.yao@intel.com>
maugustosilva
pushed a commit
that referenced
this pull request
Sep 16, 2026
* feat(p2p-kv-cache-sharing): add Intel XPU (TCP-only) variant Add modelserver/xpu/vllm/base/ mirroring the GPU OffloadingConnector P2P-tier setup, adapted for Intel XPU: - DRA ResourceClaimTemplate (gpu.intel.com), 1 XPU per pod, matching the pattern used by precise-prefix-cache-routing and tiered-prefix-cache's Intel XPU overlays, instead of the GPU overlay's nvidia.com/gpu device-plugin request. - CI-sized functional check: 2 replicas (1 source + 1 receiver) of Qwen/Qwen3-0.6B instead of the GPU overlay's 16x gpt-oss-120b - validates the same P2P pull mechanism without requiring 24GB+ of device memory per pod or reproducing the GPU benchmark's scale. cpu_bytes_to_use and the shm tier are sized down to match. - Same routing-sidecar (patch-sidecar.yaml) and P2P/kv-events port layout as the GPU overlay - both are accelerator-agnostic. No RDMA overlay (xpu/vllm/rdma) is added. Direct RDMA over Intel XPU device memory currently hits an upstream UCX ze_copy/DMA-BUF data-correctness bug (silently wrong bytes, reported success) that guides/modelexpress-p2p's Intel XPU PR (#2461) already ran into and documented; tracked at openucx/ucx#11902 and #11903. The TCP-only side channel here does not depend on that fix and is safe to ship now. README changes: - Supported Hardware Backends: document the Intel XPU variant, its scope (functional check, not a benchmark), and why RDMA is deferred. - Step 3 (Deploy the Model Server): ACCELERATOR_TYPE now documents xpu, and TRANSPORT=base is called out as the only option for xpu. Verified guides/p2p-kv-cache-sharing/modelserver/xpu/vllm/base renders cleanly with 'kubectl kustomize' (ServiceAccount, Deployment with sidecar + engine container merged correctly, port 8000->8200 renamed, ResourceClaimTemplate name-referenced correctly). Not yet deployed to a live cluster. * p2p-kv-cache-sharing: document XPU cluster verification results Tested the xpu/vllm/base overlay on XPU-8xB60-817225 (real Intel Arc Pro B60 hardware). DRA allocation, sidecar init, and vLLM boot with OffloadingConnector + NIXL/UCX + P2P secondary tier all verified working. However, the P2P pull itself is not yet verified: the pinned XPU image (llm-d-xpu:v0.9.0, vLLM 0.26.0) has no remote_kv_source handling in its OffloadingConnector (only max_offload_tokens), unlike the GPU overlay's vllm-openai:v0.27.1. Documented this version-gap limitation in the guide README. * p2p-kv-cache-sharing: confirm XPU P2P pull works with nightly vLLM image Manual pull test on real Intel Arc Pro B60 hardware (XPU-8xB60-817225) confirms the P2P pull mechanism itself works correctly once the pinned vLLM version carries vllm/v1/kv_offload/tiering/p2p/. Swapping the test deployment's image from ghcr.io/llm-d/llm-d-xpu:v0.9.0 (vLLM 0.26.0, no remote_kv_source support) to docker.io/vllm/vllm-openai-xpu:nightly (vLLM 0.29.1rc1) made a 4096-token prefix pull succeed end-to-end (external_prefix_cache_hits_total incremented by exactly 4096, reproduced twice). Updates the README's Intel XPU warning to reflect this: the gap is purely the pinned image's vLLM version, not the manifests or the pull mechanism, with a workaround (temporarily point the xpu-vllm component at nightly) documented until ghcr.io/llm-d/llm-d-xpu is rebuilt against a newer vLLM release. * p2p-kv-cache-sharing: default XPU overlay to nightly xpu-vllm image ghcr.io/llm-d/llm-d-xpu:v0.9.0 (vLLM 0.26.0) has no remote_kv_source handling in its OffloadingConnector, so the P2P pull this guide exists to demonstrate cannot work on it. Point the overlay at the existing xpu-vllm/nightly component (vLLM main) instead, which was verified end-to-end on real Intel Arc Pro B60 hardware, rather than shipping an overlay that starts cleanly but can't do the one thing the guide is about. Switch back to the llm-d component once ghcr.io/llm-d/llm-d-xpu is rebuilt against a vLLM release carrying vllm/v1/kv_offload/tiering/p2p/. Also trims the README's Intel XPU warning now that this is the default rather than a documented manual workaround. Signed-off-by: Yao, Qing <qing.yao@intel.com> * p2p-kv-cache-sharing: fix XPU review findings - Add UCX_TLS=tcp to the XPU overlay's vLLM env, matching the pd-disaggregation XPU precedent, so this TCP-only overlay can't silently fall back to the buggy ze_copy/DMA-BUF transport. - Stop citing guides/modelexpress-p2p's README as documenting the Intel XPU UCX bug: that content is still an unmerged draft (#2461), not on main. Reference the draft PR directly. - Scope the '16 replicas, TP=1' summary to the GPU overlay; the XPU overlay is 2 replicas of a different model. - Note that this guide's router values hard-code modelName: openai/gpt-oss-120b, and must be changed to Qwen/Qwen3-0.6B before installing the router when deploying the XPU overlay, or the render Service and every downstream verification command target a model the XPU pods never load. All four issues were raised by automated review on the upstream PR; verified against the actual repo state before fixing. * p2p-kv-cache-sharing: correct UCX bug attribution for XPU overlay Traced the actual vLLM OffloadingConnector/TieringOffloadingSpec P2P data path: every transfer is staged through its CPU-mmap-backed offload tier, and NIXL registers only that host memory (DRAM) with UCX. XPU device memory is never handed to UCX/NIXL directly, unlike guides/modelexpress-p2p's Intel XPU variant (#2461), which does register device memory for its weight transfer and genuinely needs UCX_TLS=tcp,ze_copy. So for this overlay: - UCX_TLS=tcp was already the right pin, but the comment overstated the risk (implying UCX could still reach device memory here). - The claim that a missing RDMA overlay is 'blocked' by the upstream UCX ze_copy/DMA-BUF bug (openucx/ucx#11902, #11903) doesn't hold for this connector: an RDMA transport here would still only move CPU-to-CPU DRAM, not touch XPU VRAM directly. It just hasn't been built/validated yet, for unrelated reasons. - 'TCP-only side channel' conflated the P2P tier's ZMQ control channel with the NIXL/UCX data plane; described them separately. Reworded the patch-vllm.yaml comment, kustomization.yaml comment, and README Supported Hardware Backends section accordingly, keeping a pointer to the real UCX bug and #2461 for readers evaluating direct XPU-device RDMA elsewhere. * p2p-kv-cache-sharing: tighten XPU overlay wording Trim hedging phrasing ('simply because', 'nothing about this path blocks it', 'anyway') from the README, kustomization.yaml, and patch-vllm.yaml so the Intel XPU transport description reads as a final statement of fact rather than a running commentary. --------- Signed-off-by: Yao, Qing <qing.yao@intel.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
guides/modelexpress-p2p/README.xpu.md), following the repo's existingREADME.<variant>.mdnaming convention (e.g.pd-disaggregation/README.tpu.md,workload-autoscaling/README.hpa-epp.md) rather than a barexpu.md.image/Dockerfile.xpu) pinning PyTorch2.12.0+xpu, vLLM0.26.0+xpu, UCX (built with verbs + Level Zero support), NIXL1.3.0, and ModelExpress0.5.0, with build-time checks that the ZE modules, NIXL UCX plugin linkage, and package versions all match. The image inherits itsENTRYPOINT/CMDfrom the base image; every Deployment in this guide sets an explicitcommand.baseanddranetkustomize overlays undermodelserver/xpu/vllm/for a cross-node one-seed/one-receiver Qwen3-0.6B deployment:baserequests one Intel GPU DRA claim per pod on the default pod network,dranetadditionally requests an RDMA-capable NIC DRA claim to carry the transfer over a faster fabric. Thedranet-onlyConfigMap/ResourceClaimTemplateare name-prefixed the same way as the guide's other resources to avoid collisions with other guides sharing a namespace.guides/modelexpress-p2p/README.mdandguides/README.mdto link the new Intel XPU variant.The guide validates only
UCX_TLS/NIXL_UCX_TLS=tcp,ze_copyand explicitly avoidsrc_mlx5/rc/dc/ib, because directrc_mlx5reads from B60 device memory silently corruptcontent-dependent regions: the RDMA operation completes with
UCS_OK, but bytes read fromcompressible/constant/repetitive source regions (padding, sparse, and quantized values are common
in model weights) come back wrong, while random regions in the same allocation pass SHA-256
verification. This reproduces for both GET and PUT, on multiple workers, and even with same-node
forced RC — it is not a network/IOMMU/firmware/routing configuration issue.
Why this is a draft
This PR ships only the TCP-based configuration above, which is correct and safe to merge as-is —
it does not itself depend on any external fix. It is marked draft because it tracks an in-flight
upstream dependency that changes what happens next:
#11902) targets the same failure mode — the
ze_copymemory domain advertises DMA-BUF export support forze-devicememory withoutrequesting an export descriptor at allocation time. Both PRs are currently open/draft and
unmerged upstream.
check (all-zero, constant, repeating, sparse, random, and mixed-content allocations, plus an
independent checksum for every Qwen3-0.6B tensor) on this cluster's B60/ConnectX-7 hardware with
direct RDMA enabled. If it passes, this PR is updated (or followed up) to switch the guide's
default to direct XPU RDMA before it comes out of draft.
Testing
npx markdownlint-cli2 guides/modelexpress-p2p/README.xpu.md guides/modelexpress-p2p/README.md: 0 errors.kubectl kustomize guides/modelexpress-p2p/modelserver/xpu/vllm/baseand.../dranetboth render cleanly, with all resource names and cross-references (resourceClaimTemplateName,configMap.name) resolving correctly, no double-prefixed or dangling references.pytest scripts/tests/: 36 passed, 1 skipped (unchanged baseline).scripts/sync-nightly-matrix.py --check: no diff (no CI workflow changes in this PR).baseoverlay transferred 339 Qwen3-0.6B tensors (1.20 GB) in 10.73 s (0.9 Gbps) over a 1 Gbps management network; thedranetoverlay completed the same transfer in 2.20 s (4.4 Gbps) over a 400 Gbps fabric while still forcingtcp,ze_copy(no verbs data lane). Both replicas stayed Ready with zero restarts, and inference output matched between seed and receiver.