Conversation
…emory Protocols which access the payload directly do not set a memtype operation, and ucp_proto_common_check_mem_access() explicitly permits UCT_EP_OP_LAST when the memory is accessible from the CPU. The performance estimation path did not follow that contract: - ucp_proto_buffer_copy_factor_id() asserted that the operation is GET_ZCOPY or PUT_ZCOPY, so any CPU-accessible memory type other than host aborted with "memtype_op=16". - ucp_proto_init_add_buffer_copy_time() modeled a plain memcpy only when both buffers are UCS_MEMORY_TYPE_HOST. Other CPU-accessible types fell through to the memtype endpoint lookup and returned UCS_ERR_UNSUPPORTED, which silently dropped the protocol. Both affect the memory types in UCS_MEMORY_TYPES_CPU_ACCESSIBLE which are not plain host memory: ze-host, ze-managed and rocm-managed. Add ucp_proto_buffer_copy_is_memcpy() to decide when the copy is performed by the CPU, and use it for the memcpy estimation. Keep the explicit UCT_EP_OP_LAST branch in ucp_proto_buffer_copy_factor_id() so that a non-CPU-accessible pair still reports the memory types rather than the misleading zero-copy assertion. With rc_x and ze_copy, initiating a 4 MB tag send from ze-host memory aborted while initializing protocol candidates at proto_init.c:292. After the fix, protocol initialization completes. For messages in the 0..2038-byte range, ze-host and ze-managed select "eager short" instead of "eager copy-in copy-out", confirming that direct-access protocols remain eligible. ze-device is not CPU accessible and is unchanged. Add test_ucp_proto_ze.cpu_accessible_direct_proto_eligible, which asserts that egr/short is selected for ze-host and ze-managed. Reverting only proto_init.c makes it reproduce the original assertion failure. Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
|
🤖 CI Triage Agent — TL;DR: The GPU gtest job failed on a single test — Full analysisSummary: Root cause: The failing test ( The commit under test changes exactly that selection logic:
Widening eligibility and making direct‑access protocols look cheaper changes the winner/thresholds of protocol selection, so with an explicit Implicated commit: File: Suggested fix:
Related: PR #11902 (this PR); prior CI triage of the same tests: #11685, #11858, #11861, #11864
|
This analysis is based on a change that isn't in this PR. The failing test uses No costing or selection change for CUDA — this cannot alter the protocol chosen for a 1 MB CUDA AM send. The failure is pre-existing: same test failed on #11685 ( Suggested fix (1) would regress this PR: pair-gating the eligibility check makes direct-access protocols ineligible for |
Add modelserver/xpu/vllm/base/ mirroring the GPU OffloadingConnector P2P-tier setup, adapted for Intel XPU: - DRA ResourceClaimTemplate (gpu.intel.com), 1 XPU per pod, matching the pattern used by precise-prefix-cache-routing and tiered-prefix-cache's Intel XPU overlays, instead of the GPU overlay's nvidia.com/gpu device-plugin request. - CI-sized functional check: 2 replicas (1 source + 1 receiver) of Qwen/Qwen3-0.6B instead of the GPU overlay's 16x gpt-oss-120b - validates the same P2P pull mechanism without requiring 24GB+ of device memory per pod or reproducing the GPU benchmark's scale. cpu_bytes_to_use and the shm tier are sized down to match. - Same routing-sidecar (patch-sidecar.yaml) and P2P/kv-events port layout as the GPU overlay - both are accelerator-agnostic. No RDMA overlay (xpu/vllm/rdma) is added. Direct RDMA over Intel XPU device memory currently hits an upstream UCX ze_copy/DMA-BUF data-correctness bug (silently wrong bytes, reported success) that guides/modelexpress-p2p's Intel XPU PR (llm-d#2461) already ran into and documented; tracked at openucx/ucx#11902 and #11903. The TCP-only side channel here does not depend on that fix and is safe to ship now. README changes: - Supported Hardware Backends: document the Intel XPU variant, its scope (functional check, not a benchmark), and why RDMA is deferred. - Step 3 (Deploy the Model Server): ACCELERATOR_TYPE now documents xpu, and TRANSPORT=base is called out as the only option for xpu. Verified guides/p2p-kv-cache-sharing/modelserver/xpu/vllm/base renders cleanly with 'kubectl kustomize' (ServiceAccount, Deployment with sidecar + engine container merged correctly, port 8000->8200 renamed, ResourceClaimTemplate name-referenced correctly). Not yet deployed to a live cluster. Signed-off-by: Yao, Qing <qing.yao@intel.com>
Traced the actual vLLM OffloadingConnector/TieringOffloadingSpec P2P data path: every transfer is staged through its CPU-mmap-backed offload tier, and NIXL registers only that host memory (DRAM) with UCX. XPU device memory is never handed to UCX/NIXL directly, unlike guides/modelexpress-p2p's Intel XPU variant (llm-d#2461), which does register device memory for its weight transfer and genuinely needs UCX_TLS=tcp,ze_copy. So for this overlay: - UCX_TLS=tcp was already the right pin, but the comment overstated the risk (implying UCX could still reach device memory here). - The claim that a missing RDMA overlay is 'blocked' by the upstream UCX ze_copy/DMA-BUF bug (openucx/ucx#11902, #11903) doesn't hold for this connector: an RDMA transport here would still only move CPU-to-CPU DRAM, not touch XPU VRAM directly. It just hasn't been built/validated yet, for unrelated reasons. - 'TCP-only side channel' conflated the P2P tier's ZMQ control channel with the NIXL/UCX data plane; described them separately. Reworded the patch-vllm.yaml comment, kustomization.yaml comment, and README Supported Hardware Backends section accordingly, keeping a pointer to the real UCX bug and llm-d#2461 for readers evaluating direct XPU-device RDMA elsewhere. Signed-off-by: Yao, Qing <qing.yao@intel.com>
* feat(p2p-kv-cache-sharing): add Intel XPU (TCP-only) variant Add modelserver/xpu/vllm/base/ mirroring the GPU OffloadingConnector P2P-tier setup, adapted for Intel XPU: - DRA ResourceClaimTemplate (gpu.intel.com), 1 XPU per pod, matching the pattern used by precise-prefix-cache-routing and tiered-prefix-cache's Intel XPU overlays, instead of the GPU overlay's nvidia.com/gpu device-plugin request. - CI-sized functional check: 2 replicas (1 source + 1 receiver) of Qwen/Qwen3-0.6B instead of the GPU overlay's 16x gpt-oss-120b - validates the same P2P pull mechanism without requiring 24GB+ of device memory per pod or reproducing the GPU benchmark's scale. cpu_bytes_to_use and the shm tier are sized down to match. - Same routing-sidecar (patch-sidecar.yaml) and P2P/kv-events port layout as the GPU overlay - both are accelerator-agnostic. No RDMA overlay (xpu/vllm/rdma) is added. Direct RDMA over Intel XPU device memory currently hits an upstream UCX ze_copy/DMA-BUF data-correctness bug (silently wrong bytes, reported success) that guides/modelexpress-p2p's Intel XPU PR (#2461) already ran into and documented; tracked at openucx/ucx#11902 and #11903. The TCP-only side channel here does not depend on that fix and is safe to ship now. README changes: - Supported Hardware Backends: document the Intel XPU variant, its scope (functional check, not a benchmark), and why RDMA is deferred. - Step 3 (Deploy the Model Server): ACCELERATOR_TYPE now documents xpu, and TRANSPORT=base is called out as the only option for xpu. Verified guides/p2p-kv-cache-sharing/modelserver/xpu/vllm/base renders cleanly with 'kubectl kustomize' (ServiceAccount, Deployment with sidecar + engine container merged correctly, port 8000->8200 renamed, ResourceClaimTemplate name-referenced correctly). Not yet deployed to a live cluster. * p2p-kv-cache-sharing: document XPU cluster verification results Tested the xpu/vllm/base overlay on XPU-8xB60-817225 (real Intel Arc Pro B60 hardware). DRA allocation, sidecar init, and vLLM boot with OffloadingConnector + NIXL/UCX + P2P secondary tier all verified working. However, the P2P pull itself is not yet verified: the pinned XPU image (llm-d-xpu:v0.9.0, vLLM 0.26.0) has no remote_kv_source handling in its OffloadingConnector (only max_offload_tokens), unlike the GPU overlay's vllm-openai:v0.27.1. Documented this version-gap limitation in the guide README. * p2p-kv-cache-sharing: confirm XPU P2P pull works with nightly vLLM image Manual pull test on real Intel Arc Pro B60 hardware (XPU-8xB60-817225) confirms the P2P pull mechanism itself works correctly once the pinned vLLM version carries vllm/v1/kv_offload/tiering/p2p/. Swapping the test deployment's image from ghcr.io/llm-d/llm-d-xpu:v0.9.0 (vLLM 0.26.0, no remote_kv_source support) to docker.io/vllm/vllm-openai-xpu:nightly (vLLM 0.29.1rc1) made a 4096-token prefix pull succeed end-to-end (external_prefix_cache_hits_total incremented by exactly 4096, reproduced twice). Updates the README's Intel XPU warning to reflect this: the gap is purely the pinned image's vLLM version, not the manifests or the pull mechanism, with a workaround (temporarily point the xpu-vllm component at nightly) documented until ghcr.io/llm-d/llm-d-xpu is rebuilt against a newer vLLM release. * p2p-kv-cache-sharing: default XPU overlay to nightly xpu-vllm image ghcr.io/llm-d/llm-d-xpu:v0.9.0 (vLLM 0.26.0) has no remote_kv_source handling in its OffloadingConnector, so the P2P pull this guide exists to demonstrate cannot work on it. Point the overlay at the existing xpu-vllm/nightly component (vLLM main) instead, which was verified end-to-end on real Intel Arc Pro B60 hardware, rather than shipping an overlay that starts cleanly but can't do the one thing the guide is about. Switch back to the llm-d component once ghcr.io/llm-d/llm-d-xpu is rebuilt against a vLLM release carrying vllm/v1/kv_offload/tiering/p2p/. Also trims the README's Intel XPU warning now that this is the default rather than a documented manual workaround. Signed-off-by: Yao, Qing <qing.yao@intel.com> * p2p-kv-cache-sharing: fix XPU review findings - Add UCX_TLS=tcp to the XPU overlay's vLLM env, matching the pd-disaggregation XPU precedent, so this TCP-only overlay can't silently fall back to the buggy ze_copy/DMA-BUF transport. - Stop citing guides/modelexpress-p2p's README as documenting the Intel XPU UCX bug: that content is still an unmerged draft (#2461), not on main. Reference the draft PR directly. - Scope the '16 replicas, TP=1' summary to the GPU overlay; the XPU overlay is 2 replicas of a different model. - Note that this guide's router values hard-code modelName: openai/gpt-oss-120b, and must be changed to Qwen/Qwen3-0.6B before installing the router when deploying the XPU overlay, or the render Service and every downstream verification command target a model the XPU pods never load. All four issues were raised by automated review on the upstream PR; verified against the actual repo state before fixing. * p2p-kv-cache-sharing: correct UCX bug attribution for XPU overlay Traced the actual vLLM OffloadingConnector/TieringOffloadingSpec P2P data path: every transfer is staged through its CPU-mmap-backed offload tier, and NIXL registers only that host memory (DRAM) with UCX. XPU device memory is never handed to UCX/NIXL directly, unlike guides/modelexpress-p2p's Intel XPU variant (#2461), which does register device memory for its weight transfer and genuinely needs UCX_TLS=tcp,ze_copy. So for this overlay: - UCX_TLS=tcp was already the right pin, but the comment overstated the risk (implying UCX could still reach device memory here). - The claim that a missing RDMA overlay is 'blocked' by the upstream UCX ze_copy/DMA-BUF bug (openucx/ucx#11902, #11903) doesn't hold for this connector: an RDMA transport here would still only move CPU-to-CPU DRAM, not touch XPU VRAM directly. It just hasn't been built/validated yet, for unrelated reasons. - 'TCP-only side channel' conflated the P2P tier's ZMQ control channel with the NIXL/UCX data plane; described them separately. Reworded the patch-vllm.yaml comment, kustomization.yaml comment, and README Supported Hardware Backends section accordingly, keeping a pointer to the real UCX bug and #2461 for readers evaluating direct XPU-device RDMA elsewhere. * p2p-kv-cache-sharing: tighten XPU overlay wording Trim hedging phrasing ('simply because', 'nothing about this path blocks it', 'anyway') from the README, kustomization.yaml, and patch-vllm.yaml so the Intel XPU transport description reads as a final statement of fact rather than a running commentary. --------- Signed-off-by: Yao, Qing <qing.yao@intel.com>
…essible-memtype-eligibility
Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
|
Residual coverage gap: |
Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
Eager bcopy protocols set a SHORT memtype operation, but pack through ucp_dt_contig_pack(), which uses memcpy for CPU-accessible memory and reaches the memtype endpoint only for non-CPU-accessible memory. Estimating those copies with the copy interface bandwidth priced eager out of small messages, so rendezvous was selected from the first byte. Keep ZCOPY operations on the copy interface estimation, since rendezvous mtype protocols perform a real memtype endpoint copy. Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
|
🤖 CI Triage Agent — TL;DR: The ASan gtest run failed on exactly one test — Full analysisSummary: Root cause: A receive hang, not a slow test. Notably this is the proto_v1 variant ( Implicated commit: unknown (not PR #11902). Most recent changes touching the implicated reassembly machinery: File: src/ucp/core/ucp_am.c:1440-1477 ( Suggested fix: Do not change the time limit and do not modify PR #11902 — re-run the ASan job to confirm the test is flaky on the target branch as well (ideally also run the test standalone: Related: PR #11902 (branch under test), #11944 (interval tree → rbtree extraction), #11452 (AM resend flag / partial message eviction), #11179 (zero-length interval tree insert protection). |
Replace ucp_proto_buffer_copy_is_memcpy() with a condition derived from the factor selected by ucp_proto_buffer_copy_factor_id(), so the operation classification lives in a single place rather than being duplicated. The derived condition preserves host/host ZCOPY as memcpy while keeping non-host ZCOPY on the memtype copy factor. Preserve the UCT_EP_OP_LAST assertion. The CPU-accessible protocol selection checks only build a select_param and never allocate a payload buffer, so drop the hardware support and copy transport gates and run them over plain rc. This covers ze-host, ze-managed, and rocm-managed without GPU hardware. Keep the ze-device negative path in a separate fixture because it requires the copy transport. Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
|
🤖 CI Triage Agent — TL;DR: The BlueField gtest job failed on a single unrelated, racy assertion in Full analysisSummary: Root cause: Test-side race, not a product bug and not related to this PR. The log shows the recovery flow worked ( Implicated commit: db208ee — "UCP/FT: probe-gated lane recovery via aux uct_ep_check (#11563)", Evgeny Leksikov (added File: test/gtest/ucp/test_ucp_fault_tolerance.cc:941 (sampling loop at 918–939; same pattern at 853–874) Suggested fix: Make probe observation edge-triggered instead of state-sampled:
Meanwhile, re-run the BlueField job for PR #11902; this failure should not block it. Related: PR #11902 (the PR under test, unrelated to the failure); PR #11563 / commit db208ee introduced the flaky assertion. |
Correct the memcpy estimation comment: a host-to-host copy returns the CPU factor for any memtype_op, so only zero-copy involving non-host memory keeps the copy interface estimation. Derive the CPU-accessible memory type list in the protocol selection test from UCS_MEMORY_TYPES_CPU_ACCESSIBLE instead of hardcoding it, so it does not go stale when a new CPU-accessible memory type is added. Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
Signed-off-by: Yaser Afshar <yaser.afshar@intel.com>
|
🤖 Starting review — findings will be posted here when done. |
|
🤖 CI Triage Agent — TL;DR: Full analysisSummary: "Tests roce on worker 2" (Azure build 137315) was killed by agent shutdown after Root cause: This is a hang, not a slow test. The last application output is at The branch under test rewrites protocol-cost estimation. In
Result: memtype-copy/pipeline rendezvous protocols are now advertised as very cheap for plain host buffers even though Implicated commit: File: Suggested fix: Gate the memcpy short-circuit on the copy actually being a CPU copy, not merely on the memory being CPU-accessible — e.g. only take it when if ((memtype_op == UCT_EP_OP_PUT_SHORT) || (memtype_op == UCT_EP_OP_GET_SHORT) ||
(memtype_op == UCT_EP_OP_LAST)) {
if ((buffer_copy_factor_id == ucp_proto_buffer_copy_cpu_factor_id(local)) &&
UCP_MEM_IS_ACCESSIBLE_FROM_CPU(local_mem_type) &&
UCP_MEM_IS_ACCESSIBLE_FROM_CPU(remote_mem_type)) {
/* memcpy estimation */
}
}Alternatively, keep the Related: #11902 (the PR under test); prior art for the bypassed guard: commit
|
What?
ucp_proto_init_add_buffer_copy_time()modeled a plain memcpy only when both buffers areUCS_MEMORY_TYPE_HOST. For the other memory types inUCS_MEMORY_TYPES_CPU_ACCESSIBLE— ze-host, ze-managed and rocm-managed — this went wrong in two different ways:Direct-access protocols, which set
memtype_op == UCT_EP_OP_LAST, fell through to the memtype endpoint lookup and returnedUCS_ERR_UNSUPPORTED, so the candidate could be rejected.ucp_proto_buffer_copy_factor_id()alsoasserted that the operation is
GET_ZCOPYorPUT_ZCOPY, which aborted:proto_init.c:292 Assertion `(memtype_op == UCT_EP_OP_GET_ZCOPY) || (memtype_op == UCT_EP_OP_PUT_ZCOPY)' failed: memtype_op=16Eager bcopy protocols, which set
UCT_EP_OP_GET_SHORTorUCT_EP_OP_PUT_SHORT, were costed with the copy interface bandwidth, so the candidate was mis-ranked.Neither matches runtime behavior:
ucp_dt_contig_pack()/unpack()use memcpy for CPU-accessible memory and reach the memtype endpoint only for non-CPU-accessible memory.Add
ucp_proto_buffer_copy_is_memcpy()to decide when the copy is performed by the CPU, and use it for the memcpy estimation. Keep the explicitUCT_EP_OP_LASTbranch inucp_proto_buffer_copy_factor_id()so a contract violation reports both memory types instead of producing the misleading zero-copy-operation assertion.Why?
With
ze_copy, initiating a 4 MB tag send from ze-host memory aborted while initializing protocol candidates. Even with the assertion silenced, the direct-access candidate was dropped, and the eager bcopy candidates were priced by the copy interface rather than by memcpy.How?
The predicate treats the copy as a memcpy for host/host, and for
UCT_EP_OP_LAST,UCT_EP_OP_GET_SHORTandUCT_EP_OP_PUT_SHORTwhen both memory types are CPU accessible.GET_ZCOPY/PUT_ZCOPYkeep the copy interfaceestimation, because the rendezvous mtype protocols perform a real memtype endpoint copy.
This changes protocol selection, not only the abort. In the tested configuration, tag send from ze-host previously selected
eager shortat size 0 and rendezvous from one byte onward. It now selects:The crossover values are hardware dependent; the shape of the change is that eager is no longer priced out of small and mid-size messages.
Tests, in separate ZE and ROCm fixtures so an unsupported environment reports a skip rather than passing without assertions:
cpu_accessible_direct_proto_eligibleverifiesegr/shortis selected for each supported CPU-accessible memory type underRNDV_THRESH=inf. Reverting only the eligibility change reproduces the original assertion failure.cpu_accessible_eager_costed_as_memcpyverifies an eager protocol is selected at one byte with rendezvous enabled. Reverting only the SHORT handling makes it fail withproto=tag/rndv.device_memory_uses_mtype_copyverifies non-CPU-accessible memory still selects rendezvous, covering the CPU-accessible gate.The ZE fixture covers ze-host and ze-managed; a separate fixture covers rocm-managed over
rc,rocm_copy. This is protocol-selection coverage only; it does not exercise data movement.Tested on Intel Arc Pro B60,
rc_mlx5over RoCE:ucx_perftest tag_lat -Vpasses for ze-host and ze-managed, which aborted before.Remaining gaps: CPU-accessible
GET_ZCOPY/PUT_ZCOPYcosting is not covered, because protocol selection reaches that path only above a hardware-dependent threshold and neitherRNDV_THRESHnorRNDV_SCHEMEreliably forces it forthese memory types; covering it would need a lower-level test that constructs and inspects buffer-copy performance directly with a configured memtype endpoint. The same SHORT reclassification applies to rocm-managed, but its protocol-selection impact was not validated on ROCm hardware.