Skip to content

[BUG]: cuda.compute: gpu_struct / types.from_numpy_dtype discard a structured dtype's offsets and itemsize and rebuild it with align=True #11348

Description

@NaderAlAwar

Is this a duplicate?

Type of Bug

Silent Failure

Component

Thrust

Describe the bug

When a structured numpy dtype with explicit offsets and itemsize is passed to gpu_struct or types.from_numpy_dtype, only the field names and field dtypes survive: gpu_struct flattens the dtype to {name: field_dtype} (python/cuda_cccl/cuda/compute/struct.py:42-48), from_numpy_dtype does the same for the recursive path (types.py:192-199), and StructTypeDescriptor.__init__ rebuilds the record with np.dtype(list, align=True) (types.py:82-94 calling _build_struct_dtype at types.py:145-160). A dtype {a: int32 @ 0, b: float32 @ 8, itemsize 16} therefore becomes {a @ 0, b @ 4, itemsize 8}, and a complex128 member the caller placed at offset 16 (the real cuda::std::complex<double> position) moves to offset 8. The Python operator compiled from the rebuilt descriptor reads b from the padding at offset 4 while the device data is laid out at the caller's offsets, so the reduction returns (10, 0.0) instead of (10, 5.0) with no error. A CuPy caller hits this whenever the array dtype mirrors a C struct with explicit padding, or when they try to hand-correct the complex layout from issue 1 by supplying the device offsets themselves. The isalignedstruct check in StructTypeDescriptor.__init__ only validates the rebuilt dtype, not the caller's.

Verified in both environments (second case not in the script): gpu_struct of {idx: int64 @ 0, val: complex128 @ 16, itemsize 32} becomes offsets [0, 8], itemsize 24, and a reduce_into with a Python operator on it returns (9, 1+0j) instead of (9, 1+9j).

How to Reproduce

python -m venv venv && . venv/bin/activate
pip install 'cuda-cccl[cu12]' cupy-cuda12x   # cuda-cccl 1.1.1
python issue_2_struct_layout_discarded.py     # exits 1

issue_2_struct_layout_discarded.py:

"""gpu_struct / types.from_numpy_dtype rebuild a structured dtype with align=True and
discard the caller's offsets and itemsize. The reduce then reads field 'b' from padding.
"""
import sys

import cupy as cp
import numpy as np

import cuda.compute
from cuda.compute import gpu_struct, types

# Caller's layout: b at offset 8, record padded to 16 bytes.
padded = np.dtype({"names": ["a", "b"], "formats": ["<i4", "<f4"], "offsets": [0, 8], "itemsize": 16})


def layout(dt):
    return f"offsets={[int(dt.fields[n][1]) for n in dt.names]} itemsize={dt.itemsize}"


print("given                       ", layout(padded))
print("types.from_numpy_dtype(dt)  ", layout(types.from_numpy_dtype(padded).dtype))
print("gpu_struct(dt).dtype        ", layout(gpu_struct(padded).dtype))


def add(x, y):
    return x.a + y.a, x.b + y.b


n = 10
h_in = np.zeros(n, dtype=padded)
h_in["a"] = 1
h_in["b"] = 0.5
d_out = cp.zeros(1, dtype=padded)
cuda.compute.reduce_into(d_in=cp.asarray(h_in), d_out=d_out, num_items=n, op=add,
                         h_init=np.zeros(1, dtype=padded))
actual = tuple(d_out.get()[0].tolist())
expected = (10, 5.0)
print("actual  ", actual)
print("expected", expected)
sys.exit(0 if actual == expected else 1)

Expected behavior

gpu_struct(dt).dtype and types.from_numpy_dtype(dt).dtype keep the caller's offsets [0, 8] and itemsize 16, and the reduction returns (10, 5.0). Alternatively, a dtype whose layout differs from the one cuda.compute will use is rejected with a clear error instead of being silently rebuilt.

Reproduction link

No response

Operating System

No response

nvidia-smi output

No response

NVCC version

No response

Activity

  1. added theissue type on Sep 10, 2026
  2. moved this from Todo to In Progress in CCCLon Sep 17, 2026
  3. moved this from In Progress to In Review in CCCLon Sep 17, 2026
  4. added
    cuda.computeFor all items related to the cuda.parallel Python module
    on Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    cuda.computeFor all items related to the cuda.parallel Python module

    Type

    Projects

    • Status
      In Review

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions