Skip to content
This repository was archived by the owner on Oct 5, 2026. It is now read-only.
This repository was archived by the owner on Oct 5, 2026. It is now read-only.

[EPIC] Performance portable Python tasks #992

Description

@elliottslaughter

Some of my code uses Python tasks, because I launch a large number of operations sized for an individual GPU each, and I want to take separate streams of these and place them on different GPUs. (E.g., the first K * 1,000 operations go to the first GPU, the second K * 1,000 go to the second GPU, etc. up to hundreds of GPUs.)

Now the question is how do I make this performance portable.

Python tasks are not permitted to use cuPyNumeric themselves, so they have to use something else. My initial version used CuPy for the implementation of my GPU Python task. As a first pass, it was easy for me to add a top-level switch to enable NumPy as an alternative to CuPy.

That gets me CPU and GPU, but still leaves out the multi-CPU/OpenMP case. I cannot afford to run a rank per CPU core (nor would that likely be efficient), and vanilla NumPy has very limited support for OpenMP (essentially only what BLAS will do on its own). Keep in mind I run a variety of NumPy operations, not limited to what ultimately boils down to a BLAS call.

I would like to do this without writing my code multiple times. CuPy and NumPy are close enough I can stick to one implementation and stub out the differences. Other solutions like Numba would require a code rewrite, and would not actually be more ergonomic (i.e., the CuPy/NumPy-style implementation is actually how I want to write this code, I do not want to be writing loops, especially with all the implicit broadcasting in the code).

I can think of a couple options:

  1. Somehow get NumPy to run on OpenMP. As far as I'm aware NumPy does not have complete support for this and I would only get a subset of operations parallelized this way.
  2. Permit Python tasks to call cuPyNumeric operations, and somehow set things up so that, inside of a Python task, these calls directly call the leaf implementations for the appropriate architecture. See Nested parallelism #993 for details.

(For LANL/SLAC.)

Activity

  1. changed the title [-]How to write performance portable Python tasks[/-] [+]Performance portable Python tasks[/+] on May 8, 2025
  2. elliottslaughter commented on May 13, 2025

    @elliottslaughter
    Author

    A bit more research:

    NumPy does NOT support parallelizing any operations, except those that would be parallelized by an underlying BLAS layer:

    NumPy itself is normally intentionally limited to a single thread during function calls, however it does support multiple Python threads running at the same time. Note that for performant linear algebra NumPy uses a BLAS backend such as OpenBLAS or MKL, which may use multiple threads that may be controlled by environment variables such as OMP_NUM_THREADS depending on what is used.
    https://numpy.org/devdocs/reference/global_state.html#number-of-threads-used-for-linear-algebra

    Any solution would need to support NumPy ufuncs (i.e., whole-array operations like element-wise +, -, *, etc.), so I think my needs exceed what can be done automatically via BLAS.

  3. manopapad commented on May 22, 2025

    @manopapad
    Contributor

    One possibly sustainable solution would be to add support for Legate task launching to work with PhysicalStores directly (and in that case just jump to the task entry directly). Then we can have the same cuPyNumeric calls working both at the top-level, and in tasks (in the latter case only executing on the device or set of cores managed by the current processor).

    Some of this path already exists, in the "fast-path" experimental mode, but more refactoring may be required.

    I'd like @Jacobfaib's advice here. Do you think 2 weeks would be sufficient (assuming we prioritized this) to have a proof of concept? If so, then we can at least present this is a WIP solution to the customer.

  4. removed theissue type on May 23, 2025
  5. Jacobfaib commented on May 23, 2025

    @Jacobfaib
    Contributor

    If we dropped everything else, we could maybe have a very very rough proof of concept in 2-3 weeks, sure.

  6. elliottslaughter commented on May 23, 2025

    @elliottslaughter
    Author

    My two cents are this is not worth dropping everything else for. However it would be nice to have a roadmap and timeline I can present to the customer.

  7. manopapad commented on Jun 10, 2025

    @manopapad
    Contributor

    @Jacobfaib promised to put together an outline for this task, so we should have a plan to share by end of the week (?)

  8. sbahirnv commented on Jul 7, 2025

    @sbahirnv
    Contributor

    Python tasks are not permitted to use cuPyNumeric themselves, so they have to use something else. My initial version used CuPy for the implementation of my GPU Python task. As a first pass, it was easy for me to add a top-level switch to enable NumPy as an alternative to CuPy.

    @elliottslaughter , Is it possible to share a sample code how you are doing this? We can start here and work backward to add necessary legate support.

  9. elliottslaughter commented on Jul 7, 2025

    @elliottslaughter
    Author

    What do you want a code sample for? The switch in the quote is literally just an if at top-level:

    if ...:
        import numpy as np
    else:
        import cupy as np

    It's a bit more complicated than that because there's the condition and a few differences I have to smooth over, but otherwise it works (besides the fact that NumPy is largely single-core only).

  10. elliottslaughter commented on Sep 16, 2025

    @elliottslaughter
    Author

    This is coming up again because I have a design decision to make based on whether this is available or not.

    Basically, if I can assume we'll have this, there's some existing code I can reuse for a new feature. If not, I'll have to design the code to work around it. So I'll end up with more code and/or more complexity, which will impact downstream users who later need to come and modify this (and/or reuse it for other stuff).

    Code portability also allows the code to be more modular because I can reuse pieces more aggressively without duplicating stuff in the code base.

  11. manopapad commented on Oct 16, 2025

    @manopapad
    Contributor

    @sbahirnv here is a concrete example to target, based on what @elliottslaughter described above

    import cupynumeric as cn
    
    from legate.core.task import task, InputStore, OutputStore
    from legate.core import align, StoreTarget, VariantCode
    
    @task(
        variants=(VariantCode.CPU, VariantCode.GPU),
        constraints=(align("out", "arg1", "arg2"),),
    )
    def sum_task(
        out: OutputStore, arg1: InputStore, arg2: InputStore,
    ) -> None:
        print(f"I am a worker. I am processing elements {arg1.domain}")
        if out.target == StoreTarget.FBMEM or out.target == StoreTarget.ZCMEM:
            import cupy as np
        else:
            import numpy as np
        arr1_np = np.asarray(arg1)
        arr2_np = np.asarray(arg2)
        out_arr_np = np.asarray(out)
        out_arr_np[:] = arr1_np + arr2_np
    
    
    a = cn.arange(2 ** 25)
    b = cn.arange(2 ** 25)
    c = cn.empty_like(a)
    
    sum_task(c, a, b)
    
    assert c.max() == a.max() + b.max()
    
  12. sbahirnv commented on Oct 30, 2025

    @sbahirnv
    Contributor

    @manopapad I have a simpler example - unary negation - working. While I work toward above binary and other cuPyNumeric operations, I'm sharing the example below for early feedback.

    #!/usr/bin/env python
    """Minimal test for unary negation in nested task execution."""
    import sys
    import numpy as np
    from legate.core.task import task, InputStore, OutputStore
    from legate.core import VariantCode
    import cupynumeric as cn
    
    
    @task(variants=(VariantCode.CPU,))
    def test_unary_negation(inp: InputStore, out: OutputStore):
        arr_in = cn.asarray(inp)
        arr_out = cn.asarray(out)
        cn.negative(arr_in, out=arr_out)
    
    
    if __name__ == "__main__":
        # Create test data
        data = np.array([1.0, -2.0, 3.0, -4.0, 5.0], dtype=np.float32)
        expected = -data
        
        # Convert to cuPyNumeric
        a = cn.array(data)
        b = cn.empty_like(a)
        
        # Launch nested task
        test_unary_negation(a, b)
        
        # Validate result
        result = np.array(b)
        
        if np.allclose(result, expected):
            print("[PASS] Negation test passed")
            print(f"  Input:    {data}")
            print(f"  Expected: {expected}")
            print(f"  Got:      {result}")
            sys.exit(0)
        else:
            print("[FAIL] Negation test failed")
            print(f"  Expected: {expected}")
            print(f"  Got:      {result}")
            sys.exit(1)
    
    

    The output looks like,

    [PASS] Negation test passed
      Input:    [ 1. -2.  3. -4.  5.]
      Expected: [-1.  2. -3.  4. -5.]
      Got:      [-1.  2. -3.  4. -5.]
    

    @elliottslaughter let us know if this is useful or some more functionality is needed.

  13. elliottslaughter commented on Nov 3, 2025

    @elliottslaughter
    Author

    This looks reasonable to me, thanks!

  14. changed the title [-]Performance portable Python tasks[/-] [+][EPIC][ Performance portable Python tasks[/+] on Mar 23, 2026
  15. changed the title [-][EPIC][ Performance portable Python tasks[/-] [+][EPIC] Performance portable Python tasks[/+] on Mar 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions