Skip to content

[BUG]: NVVM backend silently ignores 12 ProgramOptions fields that NVRTC emits #2954

Description

@Arthur031221

Is this a duplicate?

Type of Bug

Silent Failure

Component

cuda.core

Describe the bug

_prepare_nvvm_options_impl in cuda_core/cuda/core/_program.pyx raises CUDAError for 31 ProgramOptions fields that libNVVM can't use. Twelve other fields, which the NVRTC path does emit, are neither emitted nor rejected on the NVVM path. Setting one of them has no effect, and there's no error or warning:

no_cache, fdevice_time_trace, device_float128, frandom_seed, ofast_compile, pch, create_pch, use_pch, pch_dir, pch_verbose, pch_messages, instantiate_templates_in_pch

All of them except no_cache are documented as "(NVRTC only)" in the ProgramOptions docstring. device_int128 is rejected on NVVM, but device_float128 compiles without a word.

The rejection list walks the fields in almost the same order as _prepare_nvrtc_options_impl and stops at minimal, which is the field just before no_cache in the NVRTC builder. My guess is the fields after that point, and device_float128, were never added. I haven't found anything saying it was deliberate.

Two other places assume NVVM rejects some of these:

  • cuda_core/cuda/core/utils/_program_cache/_keys.py says NVVM "explicitly rejects all three at compile time" about create_pch, time and fdevice_time_trace, and gates the side-effect check on NVRTC for that reason. It also says "NVVM rejects them" about the external-content options (include_path, pre_include, pch, use_pch, pch_dir). Those two comments cover eight options. NVVM rejects three of them (time, include_path, pre_include) and silently drops the other five. With create_pch or fdevice_time_trace set, make_program_cache_key(code_type="nvvm", ...) returns a key, while code_type="c++" raises ValueError. I don't think that produces a wrong cache hit, since the compile drops the option as well.
  • test_nvvm_options_reject_each_unsupported_flag in cuda_core/tests/test_program.py says its table "mirrors _prepare_nvvm_options_impl's rejection list one-for-one", so a field missing from both can't fail it.

How to Reproduce

import dataclasses
import os
import tempfile

from cuda.bindings import nvvm
from cuda.core import CUDAError, Device, Program, ProgramOptions

# 1. Each field set on its own, then as_bytes("nvvm") compared with arch alone.
values = {
    "no_cache": True, "fdevice_time_trace": "trace.json",
    "device_float128": True, "frandom_seed": "1", "ofast_compile": "max", "pch": True,
    "create_pch": "out.pch", "use_pch": "in.pch", "pch_dir": "pch-cache", "pch_verbose": True,
    "pch_messages": True, "instantiate_templates_in_pch": True,
}
base = ProgramOptions(arch="sm_80").as_bytes("nvvm")
for f in dataclasses.fields(ProgramOptions):
    if f.name not in values:
        continue
    opts = ProgramOptions(arch="sm_80", **{f.name: values[f.name]})
    nvrtc = set(opts.as_bytes("nvrtc")) - set(ProgramOptions(arch="sm_80").as_bytes("nvrtc"))
    print(f"{f.name:30} nvvm unchanged={opts.as_bytes('nvvm') == base}  nvrtc adds {sorted(nvrtc)}")

# 2. End to end through Program.
Device().set_current()
major, minor, dmajor, dminor = nvvm.ir_version()
ir = f"""target triple = "nvptx64-unknown-cuda"
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-i128:128:128-f32:32:32-f64:64:64-v16:16:16-v32:32:32-v64:64:64-v128:128:128-n16:32:64"
define void @k(i32* %p) {{
  store i32 1, i32* %p, align 4
  ret void
}}
!nvvm.annotations = !{{!0}}
!0 = !{{void (i32*)* @k, !"kernel", i32 1}}
!nvvmir.version = !{{!1}}
!1 = !{{i32 {major}, i32 {minor}, i32 {dmajor}, i32 {dminor}}}
"""
for kw in ({"device_int128": True}, {"device_float128": True}):
    try:
        Program(ir, "nvvm", ProgramOptions(arch="sm_80", **kw)).compile("ptx")
        print(kw, "compiled")
    except CUDAError as e:
        print(kw, "CUDAError:", e)

with tempfile.TemporaryDirectory() as d:
    pch = os.path.join(d, "out.pch")
    Program(ir, "nvvm", ProgramOptions(arch="sm_80", create_pch=pch)).compile("ptx")
    print("nvvm create_pch compiled, pch written:", os.path.exists(pch))

Output with cuda.core built from main at f9ed2bd:

no_cache                       nvvm unchanged=True  nvrtc adds [b'--no-cache']
fdevice_time_trace             nvvm unchanged=True  nvrtc adds [b'--fdevice-time-trace=trace.json']
device_float128                nvvm unchanged=True  nvrtc adds [b'--device-float128']
frandom_seed                   nvvm unchanged=True  nvrtc adds [b'--frandom-seed=1']
ofast_compile                  nvvm unchanged=True  nvrtc adds [b'--Ofast-compile=max']
pch                            nvvm unchanged=True  nvrtc adds [b'--pch']
create_pch                     nvvm unchanged=True  nvrtc adds [b'--create-pch=out.pch']
use_pch                        nvvm unchanged=True  nvrtc adds [b'--use-pch=in.pch']
pch_dir                        nvvm unchanged=True  nvrtc adds [b'--pch-dir=pch-cache']
pch_verbose                    nvvm unchanged=True  nvrtc adds [b'--pch-verbose=true']
pch_messages                   nvvm unchanged=True  nvrtc adds [b'--pch-messages=true']
instantiate_templates_in_pch   nvvm unchanged=True  nvrtc adds [b'--instantiate-templates-in-pch=true']
{'device_int128': True} CUDAError: The following options are not supported by NVVM backend: device_int128
{'device_float128': True} compiled
nvvm create_pch compiled, pch written: False

Expected behavior

Each of these should either reach libNVVM or tell the user it was dropped. Which of the two is your call. The existing list raises CUDAError. On #2573, though, the review preferred a UserWarning over a new error in 1.x, following #2658, and the same reasoning looks like it applies here.

ofast_compile may be one to pass through rather than reject, though I'm not sure how useful it is there. Calling libNVVM directly on the IR above (nvidia-nvvm 13.4.92, nvvm.version() returns (2, 0)), -Ofast-compile=0 compiles. =min, =mid and =max each return ERROR_COMPILATION (9) with the log parse Can't read textual IR with a Context that discards named Values. So only 0 worked on this textual IR, but unlike -pch and -device-float128, which return ERROR_INVALID_OPTION (7), it isn't rejected as an unknown option.

link_time_optimization is also neither emitted nor rejected on NVVM. I left it off the list because test_nvvm_program_options passes it to NVVM on purpose, and target_type="ltoir" adds -gen-lto anyway.

Operating System

Ubuntu 24.04.4 LTS

nvidia-smi output

NVIDIA GeForce RTX 5090, driver 610.43.02. cuda.core built from source at f9ed2bd against cuda-bindings 13.4.3 and the cuda-toolkit 13.4.2 wheels, Python 3.12.13.

Activity

  1. 0z5a commented on Sep 28, 2026

    @0z5a

    I independently reproduced the silent-option behavior from source at f9ed2bdaede7b66e6323dc7772a93953af39dfc7 on Jetson Thor (aarch64, SM110), with cuda-bindings 13.2.0, CUDA Toolkit 13.2.78, and libNVVM reporting version 2.0 / IR version 2.0.3.2.

    Results:

    • An independently enumerated inventory accounts for all 55 public ProgramOptions dataclass fields. The twelve fields listed in this issue produce no NVVM serializer change or warning, for both sm_80 and sm_110 serialization.
    • I used this checkout's cuda_python_test_helpers.nvvm_bitcode fixtures, checking the baseline first for both textual IR and bitcode. Through public Program.compile("ptx"), the twelve fields were exercised separately with None and explicit values (including False/True for booleans), with caching disabled, cache miss, and cache hit: 182 compilations/cached returns succeeded without warnings. No PCH or trace files were created in the isolated temporary directory.
    • A wrapper around the real _program_compile_uncached function confirmed one underlying compilation for each uncached/miss case and zero for each hit. This establishes the tested cache paths; it is not evidence of a wrong-cache-hit bug.
    • Direct libNVVM calls accepted -Ofast-compile=0, min, mid, and max for both fixture formats here. Direct -device-float128 and -pch calls returned NVVM_ERROR_INVALID_OPTION. The Ofast result is specific to these minimal fixtures and this compiler version, and differs from the text/bitcode distinction reported with CUDA 13.4; it does not prove every optimization level works for arbitrary modules.
    • Existing NVVM option tests: 42 passed, 7 skipped (six need a cuda-bindings utility absent in 13.2; one needs NVRTC 13.3). NVRTC option regression selection: 58 passed, 1 skipped (selected architecture restriction).

    For 1.x compatibility, would a warning for explicitly set NVRTC-only options be preferable to adding hard errors? I would keep existing rejections intact, document options handled outside serialization, and handle ofast_compile separately with compiler/version coverage. I have not changed the public behavior while this contract remains unresolved.

    A coverage test should compare an explicit, independently maintained contract table against the full dataclass field set, so adding a new field without a backend decision fails the test. It should also check explicit false/default values and validation on cache-hit paths.

    Runnable compiler/cache audit and raw results.

  2. added
    cuda.coreEverything related to the cuda.core module
    on Sep 30, 2026
  3. leofang commented on Sep 30, 2026

    @leofang
    Member

    I think these new options should be added to the negativity check here:

    # Check for unsupported options and raise error if they are set
    unsupported = []
    if opts.relocatable_device_code is not None:
    unsupported.append("relocatable_device_code")
    if opts.extensible_whole_program is not None and opts.extensible_whole_program:
    unsupported.append("extensible_whole_program")
    if opts.lineinfo is not None and opts.lineinfo:
    unsupported.append("lineinfo")
    if opts.ptxas_options is not None:
    unsupported.append("ptxas_options")
    if opts.max_register_count is not None:
    unsupported.append("max_register_count")
    if opts.use_fast_math is not None and opts.use_fast_math:
    unsupported.append("use_fast_math")
    if opts.extra_device_vectorization is not None and opts.extra_device_vectorization:
    unsupported.append("extra_device_vectorization")
    if opts.gen_opt_lto is not None and opts.gen_opt_lto:
    unsupported.append("gen_opt_lto")
    if opts.define_macro is not None:
    unsupported.append("define_macro")
    if opts.undefine_macro is not None:
    unsupported.append("undefine_macro")
    if opts.include_path is not None:
    unsupported.append("include_path")
    if opts.use_bundled_headers:
    unsupported.append("use_bundled_headers")
    if opts.pre_include is not None:
    unsupported.append("pre_include")
    if opts.no_source_include is not None and opts.no_source_include:
    unsupported.append("no_source_include")
    if opts.std is not None:
    unsupported.append("std")
    if opts.builtin_move_forward is not None:
    unsupported.append("builtin_move_forward")
    if opts.builtin_initializer_list is not None:
    unsupported.append("builtin_initializer_list")
    if opts.disable_warnings is not None and opts.disable_warnings:
    unsupported.append("disable_warnings")
    if opts.restrict is not None and opts.restrict:
    unsupported.append("restrict")
    if opts.device_as_default_execution_space is not None and opts.device_as_default_execution_space:
    unsupported.append("device_as_default_execution_space")
    if opts.device_int128 is not None and opts.device_int128:
    unsupported.append("device_int128")
    if opts.optimization_info is not None:
    unsupported.append("optimization_info")
    if opts.no_display_error_number is not None and opts.no_display_error_number:
    unsupported.append("no_display_error_number")
    if opts.diag_error is not None:
    unsupported.append("diag_error")
    if opts.diag_suppress is not None:
    unsupported.append("diag_suppress")
    if opts.diag_warn is not None:
    unsupported.append("diag_warn")
    if opts.brief_diagnostics is not None:
    unsupported.append("brief_diagnostics")
    if opts.time is not None:
    unsupported.append("time")
    if opts.split_compile is not None:
    unsupported.append("split_compile")
    if opts.fdevice_syntax_only is not None and opts.fdevice_syntax_only:
    unsupported.append("fdevice_syntax_only")
    if opts.minimal is not None and opts.minimal:
    unsupported.append("minimal")
    if unsupported:
    raise CUDAError(f"The following options are not supported by NVVM backend: {', '.join(unsupported)}")

    PR is welcomed.

    I suppose another thing we can do is to document that certain options only work for certain code_type.

  4. added
    bugSomething isn't working
    P2Low priority - Nice to have
    and removed
    triageNeeds the team's attention
    on Sep 30, 2026
  5. idhanth commented on Sep 30, 2026

    @idhanth

    Hi, I'd like to take this one if it's still open.

    Plan, following @leofang's pointer:

    • Add the NVRTC-only fields to the unsupported check in _prepare_nvvm_options_impl (fdevice_time_trace, device_float128, frandom_seed, the seven pch* fields), using the same None/truthy pattern as the existing entries.
    • Extend test_nvvm_options_reject_each_unsupported_flag to cover them, and fix the two comments in utils/_program_cache/_keys.py that assume NVVM already rejects these.
    • Add a short note to the ProgramOptions docstring about which options apply to which code_type.

    Two questions before I start:

    1. no_cache isn't marked NVRTC-only. Should NVVM reject it too, or just ignore it?
    2. ofast_compile: @0z5a found libNVVM accepts -Ofast-compile. Should NVVM forward it rather than reject it?

    I don't have a local CUDA/Linux setup right now, so I'll rely on CI for the GPU tests.

  6. leofang commented on Sep 30, 2026

    @leofang
    Member

    I don't have a local CUDA/Linux setup right now, so I'll rely on CI for the GPU tests.

    We cannot accept such PRs unfortunately. GPU CI resources are scarce so please do make sure PRs are verified locally.

    In the case of NVRTC and NVVM, they should work on CPU-only machines today (it's not a guarantee, they just happen to work), so you should be able to develop/test/debug locally.

  7. idhanth commented on Sep 30, 2026

    @idhanth

    Thanks @leofang, I ran the full test_program.py + test_program_cache.py on a GPU with this change applied on top of e8d9075:

    • Tesla T4 (sm_75), driver 580.82.07, CUDA 12.x toolkit, cuda-bindings 12.9.9, Python 3.12
    • 380 passed, 0 failed, 12 skipped
    • All 40 NVVM rejection rows that ran passed, including the 10 new ones. The use_bundled_headers row was skipped because it needs NVRTC ≥ 13.3.

    The other skips were all version or hardware gates: numba_debug / use_bundled_headers need a newer NVRTC, device_float128 needs sm_100+, the fdevice_time_trace compile test is skipped as buggy on NVRTC < 13.0, this libNVVM doesn't recognize -numba-debug, and 6 cases skipped because check_nvvm_compiler_options isn't in this cuda.bindings build.

    I also tried CUDA 13.4 (cuda-bindings 13.4.3) on a CPU-only Linux machine: the 10 new rejection rows fail on main and pass with the fix. The full files can't be collected there without libcuda, so that run was a subset. libNVVM 12.9 and 13.4 both return NVVM_ERROR_INVALID_OPTION for each of those 10 flags too, so they were never doing anything on NVVM.

    Two things I left out of this change since I wasn't sure how you'd want them handled:

    1. ofast_compile: libNVVM accepts -Ofast-compile, but on 13.4 min/mid/max fail with textual IR (bitcode is fine). Should it be forwarded or rejected?
    2. no_cache: reject on NVVM, or just ignore it?

    Happy to open a PR as soon as I'm assigned.

  8. added this to the cuda.core next milestone on Oct 8, 2026
  9. added theissue type on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

P2Low priority - Nice to havebugSomething isn't workingcuda.coreEverything related to the cuda.core module

Type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions