Skip to content

UCT/CUDA_IPC: Publish the chunk layout of multi-allocation VMM ranges - #11991

Open
tomerg-nvidia wants to merge 1 commit into
openucx:masterfrom
tomerg-nvidia:vmm-multi-sender
Open

tomerg-nvidia wants to merge 1 commit into
openucx:masterfrom
tomerg-nvidia:vmm-multi-sender

Conversation

@tomerg-nvidia

Copy link
Copy Markdown
Contributor

What?

A VMM range backed by several physical allocations cannot be described by the one allocation handle a key carries today. Add the exporter side of a multi-chunk key: discover the allocations behind a range, write their descriptors to a GPU buffer shared by fabric handle, and keep that buffer on the local key until deregistration.

Nothing packs such a key yet, the pack path is wired up separately.

Why?

A step in fixing multi-allocation VMM transfers with cuda_ipc. Currently they are not working.

How?

During pack, if a multi-allocation VMM is detected (currently not wired but will be in a followup PR):

  1. Chunks are discovered by looping over the address range, in each iteration the current chunk boundery is fetched, and a fabric handle is exported. These handles are saved in an array.
  2. Then a metadata block is allocated on the device, with the exported handles array, their size, and buffer id. A header in the rkey contains the number of chunks, the metadata block size and some other versioning information.
  3. This metadata block lifetime is bound to the memory registration. The lkey contains a list of all published metadata blocks.

Note: This PR contains only the pack-side functions, and does not wire them in yet. A new handle type will be introduced for it when it is wired up.

@svc-nvidia-pr-review

Copy link
Copy Markdown

🤖 Starting review — findings will be posted here when done.

Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_md.h
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_md.h Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_md.h Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.h
@svc-nvidia-pr-review

Copy link
Copy Markdown

🤖 Starting review — findings will be posted here when done.

Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.h
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_md.h Outdated
@svc-nvidia-pr-review

Copy link
Copy Markdown

uct_cuda_ipc_mkey_pack_vmm_multi_chunk() has no caller in the library — only gtest. Is the mkey_pack/unpack side coming in a follow-up PR?

Coverage gaps: the new gtest only covers the exporter side on a fabric-capable GPU (skipped otherwise); there is no coverage for the discover/export failure paths (non-fabric multi-chunk memory) or for the importer, which does not exist yet.

@svc-ucx

svc-ucx commented Sep 28, 2026

Copy link
Copy Markdown

🤖 CI Triage Agent — UCX PR (Coverity coverity devel on coverity_rh7) · commit 74f47882

TL;DR: The Coverity devel stage failed on a single new CHECKED_RETURN defect: cuMemRelease() is called without checking its return value in the new cuda_vmm_mem_buffer::cleanup() loop (test/gtest/uct/cuda/cuda_vmm_mem_buffer.h:171). Assign and check/log the CUresult (as all other cuMemRelease call sites in src/uct/cuda/... do) to clear the gate.

Full analysis

Summary: Azure Pipelines job "Coverity coverity devel on coverity_rh7" (build 137279) ended with ##[error]Coverity found 1 issues; the build and analysis themselves succeeded (750 TUs, 100% compiled, analysis 00:12:42), the step failed only because nerrors=1 > 0.

Root cause: Coverity's CHECKED_RETURN checker flagged the newly added cleanup path in the multi-allocation VMM test helper:

/__w/1/s/test/gtest/uct/cuda/cuda_vmm_mem_buffer.h:171
  Type: Unchecked return value (CHECKED_RETURN)
  7. check_return: Calling "cuMemRelease" without checking return value
     (as is done elsewhere 9 out of 11 times).

Coverity's supporting evidence cites src/uct/cuda/cuda_copy/cuda_copy_md.c:502, cuda_ipc_cache.c:351, cuda_ipc_md.c:296, cuda_ipc_vmm_multi.c:75 and :160, all of which assign the result of cuMemRelease() and compare it to CUDA_SUCCESS — so the unchecked call in the new loop becomes a statistical outlier. Source confirms it:

164:    void cleanup()
165:    {
166:        for (size_t i = 0; i < m_num_mapped; ++i) {
167:            cuMemUnmap(m_ptr + (i * m_chunk_size), m_chunk_size);
168:        }
169:
170:        for (auto alloc_handle : m_alloc_handles) {
171:            cuMemRelease(alloc_handle);   /* <-- flagged */
172:        }

This is a real (if benign) code-quality finding introduced by this PR, not infrastructure flakiness — no timeouts, SIGTERM or log gaps appear anywhere in the run.

Implicated commit: [REDACTED:Hex High Entropy String] — "UCT/CUDA_IPC: Publish the chunk layout of multi-allocation VMM ranges", Tomer Gilad (the only commit touching this file in this PR; it introduced the multi-handle m_alloc_handles loop in cleanup())

File: test/gtest/uct/cuda/cuda_vmm_mem_buffer.h:171 (loop starting at :170; also see unchecked cuMemUnmap at :167 and cuMemAddressFree at :175, which were not flagged but are the same pattern)

Suggested fix: Capture and check the driver return value in cleanup(), matching the convention in src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c. Since cleanup() is called from the destructor and from the alloc() unwind path, don't abort — log instead:

for (auto alloc_handle : m_alloc_handles) {
    CUresult ret = cuMemRelease(alloc_handle);
    if (ret != CUDA_SUCCESS) {
        UCS_TEST_MESSAGE << "cuMemRelease() failed: " << ret;
    }
}

Alternatively wrap the call with the existing UCT_CUDADRV_FUNC_LOG_ERR()/UCT_CUDADRV_FUNC_LOG_WARN() helper (used throughout src/uct/cuda), which both checks and logs — that also satisfies Coverity. For consistency and to pre-empt a future finding, apply the same treatment to the cuMemUnmap (line 167) and cuMemAddressFree (line 175) calls. Re-run the Coverity job afterwards; the report artifact is published as coverity_devel (output/errors/index.html) if you want to confirm the count drops to 0.

Related: PR #11991 (openucx/ucx), branch vmm-multi-sender; no existing issue found for this Coverity signature.

🛡️ This comment had 1 potential secret(s) redacted (Hex High Entropy String). See request_id cbfbb8b2-eb30-4f1d-b523-0dea03d3b47d in the triage console for the audit trail.

@svc-nvidia-pr-review

Copy link
Copy Markdown

🤖 Starting review — findings will be posted here when done.

Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
}

/* Peers may unpack any published metadata until memh deregistration. */
ucs_list_add_tail(&key->vmm_multi_list, &meta->link);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unbounded metadata accumulation: every range not covered by an existing record adds another metadata buffer of at least the allocation granularity (2MB on most GPUs) that lives until dereg, so repeated packs of growing ranges on the same memh keep consuming device memory. Is there a bound, or can the record be published once for the whole registered region?

@svc-nvidia-pr-review

Copy link
Copy Markdown

Residual coverage gap (unchanged from the earlier note): the new gtest exercises only the exporter on a fabric-capable GPU (skipped otherwise); the discover/export failure paths (multi-chunk range with a non-fabric chunk) and the importer are untested, and uct_cuda_ipc_mkey_pack_vmm_multi_chunk() still has no in-library caller.

@svc-nvidia-pr-review

Copy link
Copy Markdown

🤖 Starting review — findings will be posted here when done.

}

/* Peers may unpack any published metadata until memh deregistration. */
ucs_list_add_tail(&key->vmm_multi_list, &meta->link);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Each published record pins a granularity-sized (>=2MB) device allocation until mem_dereg(), and since a newer range never supersedes an older one, a growing sequence of ranges on the same key accumulates several of them. Can we reuse one buffer sized for the largest range so far, or is the assumption at most one range per key?

Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_md.h
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c Outdated
A VMM range backed by several physical allocations cannot be described by
the one allocation handle a key carries today. Add the exporter side of a
multi-chunk key: discover the allocations behind a range, write their
descriptors to a GPU buffer shared by fabric handle, and keep that buffer on
the local key until deregistration.

Nothing packs such a key yet; the pack path is wired up separately.
@svc-nvidia-pr-review

Copy link
Copy Markdown

🤖 Starting review — findings will be posted here when done.

Comment thread src/uct/cuda/cuda_ipc/cuda_ipc_vmm_multi.c
if (!(allowed_handle_types & CU_MEM_HANDLE_TYPE_FABRIC)) {
ucs_debug("VMM chunk 0x%llx does not allow fabric handles",
chunk_base);
status = UCS_ERR_UNSUPPORTED;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes "chunk is not fabric-exportable" indistinguishable from the "single allocation" meaning documented for uct_cuda_ipc_mkey_pack_vmm_multi_chunk(), so a caller falling back to the plain pack would publish a key covering only the first chunk. Can we return a different status here (or document both cases)?

* @return UCS_OK on success, UCS_ERR_UNSUPPORTED for a single allocation, or
* another error status on failure
*/
ucs_status_t uct_cuda_ipc_mkey_pack_vmm_multi_chunk(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor: pls rename to uct_cuda_ipc_vmm_multi_mkey_pack() for consistency with the other uct_cuda_ipc_vmm_multi_* functions in this file.

@tomerg-nvidia
tomerg-nvidia marked this pull request as ready for review September 29, 2026 10:45

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants