Skip to content

Add stream callback support #311

Description

@leofang
No description provided.

Activity

  1. added
    cuda.coreEverything related to the cuda.core module
    featureNew feature or request
    P1Medium priority - Should do
    on Dec 17, 2024
  2. added this to the cuda.core beta 3 milestone on Dec 17, 2024
  3. leofang commented on Jan 16, 2025

    @leofang
    MemberAuthor

    deprioritizing this as it could have some perf considerations when it comes to CUDA graphs

  4. added
    cuda.coreEverything related to the cuda.core module
    and removed
    cuda.coreEverything related to the cuda.core module
    on Apr 2, 2026
  5. removed this from the cuda.core backlog milestone on Apr 2, 2026
  6. 0z5a commented on Sep 19, 2026

    @0z5a

    @leofang @rparolin
    I’m interested in helping with this. I noticed that #2058 now covers the concrete cuLaunchHostFunc / host_launch path and is already assigned.
    Is there still an independent scope intended for #311 beyond #2058?

  7. 0z5a commented on Sep 20, 2026

    @0z5a

    Following up on my question from yesterday, now that I have L20 time set aside for this.

    #2058 covers the cuLaunchHostFunc / host_launch path and is assigned to @Andy-Jost. Is any independent scope still intended for #311, or should #311 be treated as covered by #2058? If it is covered, I will leave it alone and not open a duplicate implementation.

    If there is independent scope, this is the slice I would take — tests and contract only, no second public API and no host_launch work:

    1. Ordinary (non-graph) stream callbacks: prior work → callback → later work ordering on the same stream, and that enqueue returning is not treated as callback completion.
    2. Lifetime/ownership: the Python callable and user data stay valid while borrowed, owning vs borrowed stream handles do not double-free or free early, and the no-callback path does not regress.
    3. Multi-device / multi-stream isolation (2× L20 available here): per-stream local order asserted, with no claim about global order or parallelism between callbacks.
    4. Exception reporting through whatever mechanism the owner settles on for this path (unraisable hook, error sink, or a documented no-propagation contract), asserted in a subprocess so a failure cannot take the test runner with it.

    @Andy-Jost: if you would rather keep the stream-callback regression tests with #2058, please say so and I will drop #311 entirely. If some of the four items above are outside your #2058 scope, I can take exactly that slice and nothing else.

  8. Andy-Jost commented on Oct 5, 2026

    @Andy-Jost
    Contributor

    Covered by #2058, which adds host_launch for the stream path and host-callback nodes for graphs. Closing as a duplicate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Medium priority - Should docuda.coreEverything related to the cuda.core modulefeatureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions