Skip to content

Trim sample and topology results to the returned count - #212

Merged
tariq1890 merged 1 commit into
NVIDIA:mainfrom
MaxFreedomPollard:fix/returned-sample-topology-counts
Oct 6, 2026
Merged

tariq1890 merged 1 commit into
NVIDIA:mainfrom
MaxFreedomPollard:fix/returned-sample-topology-counts

Conversation

@MaxFreedomPollard

@MaxFreedomPollard MaxFreedomPollard commented Sep 20, 2026 •

Copy link
Copy Markdown

GetSamples, GetTopologyNearestGpus, and SystemGetTopologyGpuSet allocate from an initial count query, but NVML can return fewer entries on the second call. The wrappers currently expose the unused tail as zero-valued samples or device handles.

Trim each slice to the count returned on a successful second call. Error returns retain their existing behavior, including ERROR_INSUFFICIENT_SIZE with an increased required count. This follows the returned-count handling already used for compute-instance placements.

The change is limited to these three wrappers and adds no test-only NVML stubs.

Validation: go test ./pkg/nvml -count=1 passes on macOS. Hardware-dependent tests skip here because libnvidia-ml is unavailable, so the reduced-count response is not directly exercised in this local run.

Comment thread pkg/nvml/device.go Outdated
@MaxFreedomPollard
MaxFreedomPollard force-pushed the fix/returned-sample-topology-counts branch from 3907e0c to 34d53dd Compare September 30, 2026 02:56
@tariq1890

Copy link
Copy Markdown
Contributor

can you squash your commit history?

Signed-off-by: Max Freedom Pollard <272618364+MaxFreedomPollard@users.noreply.github.com>

nvml: remove test-only stubs from returned-count fix
@tariq1890
tariq1890 force-pushed the fix/returned-sample-topology-counts branch from 34d53dd to 6d66908 Compare October 6, 2026 15:17
@tariq1890
tariq1890 merged commit 1a1e86d into NVIDIA:main Oct 6, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants