Is this a duplicate?
Type of Bug
Silent Failure
Component
Thrust
Describe the bug
When a structured numpy dtype with explicit offsets and itemsize is passed to gpu_struct or types.from_numpy_dtype, only the field names and field dtypes survive: gpu_struct flattens the dtype to {name: field_dtype} (python/cuda_cccl/cuda/compute/struct.py:42-48), from_numpy_dtype does the same for the recursive path (types.py:192-199), and StructTypeDescriptor.__init__ rebuilds the record with np.dtype(list, align=True) (types.py:82-94 calling _build_struct_dtype at types.py:145-160). A dtype {a: int32 @ 0, b: float32 @ 8, itemsize 16} therefore becomes {a @ 0, b @ 4, itemsize 8}, and a complex128 member the caller placed at offset 16 (the real cuda::std::complex<double> position) moves to offset 8. The Python operator compiled from the rebuilt descriptor reads b from the padding at offset 4 while the device data is laid out at the caller's offsets, so the reduction returns (10, 0.0) instead of (10, 5.0) with no error. A CuPy caller hits this whenever the array dtype mirrors a C struct with explicit padding, or when they try to hand-correct the complex layout from issue 1 by supplying the device offsets themselves. The isalignedstruct check in StructTypeDescriptor.__init__ only validates the rebuilt dtype, not the caller's.
Verified in both environments (second case not in the script): gpu_struct of {idx: int64 @ 0, val: complex128 @ 16, itemsize 32} becomes offsets [0, 8], itemsize 24, and a reduce_into with a Python operator on it returns (9, 1+0j) instead of (9, 1+9j).
How to Reproduce
python -m venv venv && . venv/bin/activate
pip install 'cuda-cccl[cu12]' cupy-cuda12x # cuda-cccl 1.1.1
python issue_2_struct_layout_discarded.py # exits 1
issue_2_struct_layout_discarded.py:
"""gpu_struct / types.from_numpy_dtype rebuild a structured dtype with align=True and
discard the caller's offsets and itemsize. The reduce then reads field 'b' from padding.
"""
import sys
import cupy as cp
import numpy as np
import cuda.compute
from cuda.compute import gpu_struct, types
# Caller's layout: b at offset 8, record padded to 16 bytes.
padded = np.dtype({"names": ["a", "b"], "formats": ["<i4", "<f4"], "offsets": [0, 8], "itemsize": 16})
def layout(dt):
return f"offsets={[int(dt.fields[n][1]) for n in dt.names]} itemsize={dt.itemsize}"
print("given ", layout(padded))
print("types.from_numpy_dtype(dt) ", layout(types.from_numpy_dtype(padded).dtype))
print("gpu_struct(dt).dtype ", layout(gpu_struct(padded).dtype))
def add(x, y):
return x.a + y.a, x.b + y.b
n = 10
h_in = np.zeros(n, dtype=padded)
h_in["a"] = 1
h_in["b"] = 0.5
d_out = cp.zeros(1, dtype=padded)
cuda.compute.reduce_into(d_in=cp.asarray(h_in), d_out=d_out, num_items=n, op=add,
h_init=np.zeros(1, dtype=padded))
actual = tuple(d_out.get()[0].tolist())
expected = (10, 5.0)
print("actual ", actual)
print("expected", expected)
sys.exit(0 if actual == expected else 1)
Expected behavior
gpu_struct(dt).dtype and types.from_numpy_dtype(dt).dtype keep the caller's offsets [0, 8] and itemsize 16, and the reduction returns (10, 5.0). Alternatively, a dtype whose layout differs from the one cuda.compute will use is rejected with a clear error instead of being silently rebuilt.
Reproduction link
No response
Operating System
No response
nvidia-smi output
No response
NVCC version
No response
Is this a duplicate?
Type of Bug
Silent Failure
Component
Thrust
Describe the bug
When a structured numpy dtype with explicit
offsetsanditemsizeis passed togpu_structortypes.from_numpy_dtype, only the field names and field dtypes survive:gpu_structflattens the dtype to{name: field_dtype}(python/cuda_cccl/cuda/compute/struct.py:42-48),from_numpy_dtypedoes the same for the recursive path (types.py:192-199), andStructTypeDescriptor.__init__rebuilds the record withnp.dtype(list, align=True)(types.py:82-94calling_build_struct_dtypeattypes.py:145-160). A dtype{a: int32 @ 0, b: float32 @ 8, itemsize 16}therefore becomes{a @ 0, b @ 4, itemsize 8}, and a complex128 member the caller placed at offset 16 (the realcuda::std::complex<double>position) moves to offset 8. The Python operator compiled from the rebuilt descriptor readsbfrom the padding at offset 4 while the device data is laid out at the caller's offsets, so the reduction returns(10, 0.0)instead of(10, 5.0)with no error. A CuPy caller hits this whenever the array dtype mirrors a C struct with explicit padding, or when they try to hand-correct the complex layout from issue 1 by supplying the device offsets themselves. Theisalignedstructcheck inStructTypeDescriptor.__init__only validates the rebuilt dtype, not the caller's.Verified in both environments (second case not in the script):
gpu_structof{idx: int64 @ 0, val: complex128 @ 16, itemsize 32}becomes offsets[0, 8], itemsize 24, and areduce_intowith a Python operator on it returns(9, 1+0j)instead of(9, 1+9j).How to Reproduce
issue_2_struct_layout_discarded.py:Expected behavior
gpu_struct(dt).dtypeandtypes.from_numpy_dtype(dt).dtypekeep the caller's offsets[0, 8]and itemsize 16, and the reduction returns(10, 5.0). Alternatively, a dtype whose layout differs from the one cuda.compute will use is rejected with a clear error instead of being silently rebuilt.Reproduction link
No response
Operating System
No response
nvidia-smi output
No response
NVCC version
No response