Repository navigation
[EPIC] Performance portable Python tasks #992
Description
Activity
- changed the title
[-]How to write performance portable Python tasks[/-][+]Performance portable Python tasks[/+]on May 8, 2025 A bit more research:
NumPy does NOT support parallelizing any operations, except those that would be parallelized by an underlying BLAS layer:
NumPy itself is normally intentionally limited to a single thread during function calls, however it does support multiple Python threads running at the same time. Note that for performant linear algebra NumPy uses a BLAS backend such as OpenBLAS or MKL, which may use multiple threads that may be controlled by environment variables such as
OMP_NUM_THREADSdepending on what is used.
https://numpy.org/devdocs/reference/global_state.html#number-of-threads-used-for-linear-algebraAny solution would need to support NumPy ufuncs (i.e., whole-array operations like element-wise
+,-,*, etc.), so I think my needs exceed what can be done automatically via BLAS.One possibly sustainable solution would be to add support for Legate task launching to work with
PhysicalStoresdirectly (and in that case just jump to the task entry directly). Then we can have the same cuPyNumeric calls working both at the top-level, and in tasks (in the latter case only executing on the device or set of cores managed by the current processor).Some of this path already exists, in the "fast-path" experimental mode, but more refactoring may be required.
I'd like @Jacobfaib's advice here. Do you think 2 weeks would be sufficient (assuming we prioritized this) to have a proof of concept? If so, then we can at least present this is a WIP solution to the customer.
If we dropped everything else, we could maybe have a very very rough proof of concept in 2-3 weeks, sure.
My two cents are this is not worth dropping everything else for. However it would be nice to have a roadmap and timeline I can present to the customer.
Reacted by Jacob Faibussowitsch@Jacobfaib promised to put together an outline for this task, so we should have a plan to share by end of the week (?)
Python tasks are not permitted to use cuPyNumeric themselves, so they have to use something else. My initial version used CuPy for the implementation of my GPU Python task. As a first pass, it was easy for me to add a top-level switch to enable NumPy as an alternative to CuPy.
@elliottslaughter , Is it possible to share a sample code how you are doing this? We can start here and work backward to add necessary legate support.
What do you want a code sample for? The switch in the quote is literally just an
ifat top-level:if ...: import numpy as np else: import cupy as np
It's a bit more complicated than that because there's the condition and a few differences I have to smooth over, but otherwise it works (besides the fact that NumPy is largely single-core only).
Reacted by sbahirnvThis is coming up again because I have a design decision to make based on whether this is available or not.
Basically, if I can assume we'll have this, there's some existing code I can reuse for a new feature. If not, I'll have to design the code to work around it. So I'll end up with more code and/or more complexity, which will impact downstream users who later need to come and modify this (and/or reuse it for other stuff).
Code portability also allows the code to be more modular because I can reuse pieces more aggressively without duplicating stuff in the code base.
@sbahirnv here is a concrete example to target, based on what @elliottslaughter described above
import cupynumeric as cn from legate.core.task import task, InputStore, OutputStore from legate.core import align, StoreTarget, VariantCode @task( variants=(VariantCode.CPU, VariantCode.GPU), constraints=(align("out", "arg1", "arg2"),), ) def sum_task( out: OutputStore, arg1: InputStore, arg2: InputStore, ) -> None: print(f"I am a worker. I am processing elements {arg1.domain}") if out.target == StoreTarget.FBMEM or out.target == StoreTarget.ZCMEM: import cupy as np else: import numpy as np arr1_np = np.asarray(arg1) arr2_np = np.asarray(arg2) out_arr_np = np.asarray(out) out_arr_np[:] = arr1_np + arr2_np a = cn.arange(2 ** 25) b = cn.arange(2 ** 25) c = cn.empty_like(a) sum_task(c, a, b) assert c.max() == a.max() + b.max()Reacted by sbahirnv@manopapad I have a simpler example - unary negation - working. While I work toward above binary and other cuPyNumeric operations, I'm sharing the example below for early feedback.
#!/usr/bin/env python """Minimal test for unary negation in nested task execution.""" import sys import numpy as np from legate.core.task import task, InputStore, OutputStore from legate.core import VariantCode import cupynumeric as cn @task(variants=(VariantCode.CPU,)) def test_unary_negation(inp: InputStore, out: OutputStore): arr_in = cn.asarray(inp) arr_out = cn.asarray(out) cn.negative(arr_in, out=arr_out) if __name__ == "__main__": # Create test data data = np.array([1.0, -2.0, 3.0, -4.0, 5.0], dtype=np.float32) expected = -data # Convert to cuPyNumeric a = cn.array(data) b = cn.empty_like(a) # Launch nested task test_unary_negation(a, b) # Validate result result = np.array(b) if np.allclose(result, expected): print("[PASS] Negation test passed") print(f" Input: {data}") print(f" Expected: {expected}") print(f" Got: {result}") sys.exit(0) else: print("[FAIL] Negation test failed") print(f" Expected: {expected}") print(f" Got: {result}") sys.exit(1)The output looks like,
[PASS] Negation test passed Input: [ 1. -2. 3. -4. 5.] Expected: [-1. 2. -3. 4. -5.] Got: [-1. 2. -3. 4. -5.]@elliottslaughter let us know if this is useful or some more functionality is needed.
This looks reasonable to me, thanks!
Reacted by sbahirnv- changed the title
[-]Performance portable Python tasks[/-][+][EPIC][ Performance portable Python tasks[/+]on Mar 23, 2026 - changed the title
[-][EPIC][ Performance portable Python tasks[/-][+][EPIC] Performance portable Python tasks[/+]on Mar 23, 2026 - added sub-issues
on Jul 8, 2026
Some of my code uses Python tasks, because I launch a large number of operations sized for an individual GPU each, and I want to take separate streams of these and place them on different GPUs. (E.g., the first K * 1,000 operations go to the first GPU, the second K * 1,000 go to the second GPU, etc. up to hundreds of GPUs.)
Now the question is how do I make this performance portable.
Python tasks are not permitted to use cuPyNumeric themselves, so they have to use something else. My initial version used CuPy for the implementation of my GPU Python task. As a first pass, it was easy for me to add a top-level switch to enable NumPy as an alternative to CuPy.
That gets me CPU and GPU, but still leaves out the multi-CPU/OpenMP case. I cannot afford to run a rank per CPU core (nor would that likely be efficient), and vanilla NumPy has very limited support for OpenMP (essentially only what BLAS will do on its own). Keep in mind I run a variety of NumPy operations, not limited to what ultimately boils down to a BLAS call.
I would like to do this without writing my code multiple times. CuPy and NumPy are close enough I can stick to one implementation and stub out the differences. Other solutions like Numba would require a code rewrite, and would not actually be more ergonomic (i.e., the CuPy/NumPy-style implementation is actually how I want to write this code, I do not want to be writing loops, especially with all the implicit broadcasting in the code).
I can think of a couple options:
(For LANL/SLAC.)