tb_idma_mxroundtrip corrupts the heap during model construction above DataWidth 256, so the compute roundtrip cannot be simulated at the widths the datapath is built for.
Symptom
DataWidth 512, verilator 5.020: simv: malloc.c:2415: sysmalloc: Assertion ... failed. (rc=-6)
DataWidth 256, verilator 5.046: malloc(): unaligned tcache chunk detected (rc=-6)
Deterministic, 6/6 runs, empty log; the abort lands wherever glibc next touches its metadata, so the reported width moves with the toolchain. Both builds use g++-13.2.0, so this is not the coroutine miscompile that affects g++ 11.
Where it actually happens
Valgrind puts every write in the generated root constructor, before time 0:
Invalid write of size 8
at Vtb_idma_mxroundtrip___024root::Vtb_idma_mxroundtrip___024root(...)
by Vtb_idma_mxroundtrip__Syms::Vtb_idma_mxroundtrip__Syms(...)
by main
Address 0x606a680 is 1,728 bytes inside an unallocated block of size 4,100,128
71 errors from 55 contexts, all in that constructor.
Control
Same width, same verilator, same flags, tb_idma_mxquant instead:
| binary, DW 512 |
valgrind |
mxquant_512 (passes) |
0 errors from 0 contexts |
mxroundtrip_512 (aborts) |
71 errors from 55 contexts |
So it is specific to this testbench, not verilator noise at that width in general.
Ruled out
- The DPI golden:
gm_in/gm_out are static and bounds-checked, and blk[32] is indexed 0..31 (MX blocks are 32 elements regardless of width). Under ASAN the DW 512 run completes with [MXRT] ALL PASS (128 blocks, StrbWidth=64) and no overflow report.
NumAxInFlight(StrbWidth): tb_idma_mxquant passes the identical parameter and is valgrind-clean.
- Our stimulus: the corruption is at construction time, before any of it runs.
Not root-caused. The evidence points at the generated model rather than the RTL or the testbench body, and verilator 5.046 relocates the failure to DW 256 rather than fixing it.
Impact
IDMA_MXRT_WIDTHS is capped at 32 64 256 while IDMA_MXQUANT_WIDTHS runs 32 64 256 512 1024.
- Blocks a verilator 5.046 bump, which would otherwise restore
mxneg cases 4 and 12 (DataWidth 1024, they fire ComputeMxFp16Width) but breaks mxroundtrip_256.
Reproduce
make idma_verify_sim_mxroundtrip IDMA_MXRT_WIDTHS=512
tb_idma_mxroundtripcorrupts the heap during model construction above DataWidth 256, so the compute roundtrip cannot be simulated at the widths the datapath is built for.Symptom
Deterministic, 6/6 runs, empty log; the abort lands wherever glibc next touches its metadata, so the reported width moves with the toolchain. Both builds use g++-13.2.0, so this is not the coroutine miscompile that affects g++ 11.
Where it actually happens
Valgrind puts every write in the generated root constructor, before time 0:
71 errors from 55 contexts, all in that constructor.
Control
Same width, same verilator, same flags,
tb_idma_mxquantinstead:mxquant_512(passes)mxroundtrip_512(aborts)So it is specific to this testbench, not verilator noise at that width in general.
Ruled out
gm_in/gm_outare static and bounds-checked, andblk[32]is indexed 0..31 (MX blocks are 32 elements regardless of width). Under ASAN the DW 512 run completes with[MXRT] ALL PASS (128 blocks, StrbWidth=64)and no overflow report.NumAxInFlight(StrbWidth):tb_idma_mxquantpasses the identical parameter and is valgrind-clean.Not root-caused. The evidence points at the generated model rather than the RTL or the testbench body, and verilator 5.046 relocates the failure to DW 256 rather than fixing it.
Impact
IDMA_MXRT_WIDTHSis capped at32 64 256whileIDMA_MXQUANT_WIDTHSruns32 64 256 512 1024.mxnegcases 4 and 12 (DataWidth 1024, they fireComputeMxFp16Width) but breaksmxroundtrip_256.Reproduce