Repository navigation
[BUG]: NVVM backend silently ignores 12 ProgramOptions fields that NVRTC emits #2954
Description
Activity
I independently reproduced the silent-option behavior from source at
f9ed2bdaede7b66e6323dc7772a93953af39dfc7on Jetson Thor (aarch64, SM110), with cuda-bindings 13.2.0, CUDA Toolkit 13.2.78, and libNVVM reporting version 2.0 / IR version 2.0.3.2.Results:
- An independently enumerated inventory accounts for all 55 public
ProgramOptionsdataclass fields. The twelve fields listed in this issue produce no NVVM serializer change or warning, for bothsm_80andsm_110serialization. - I used this checkout's
cuda_python_test_helpers.nvvm_bitcodefixtures, checking the baseline first for both textual IR and bitcode. Through publicProgram.compile("ptx"), the twelve fields were exercised separately withNoneand explicit values (includingFalse/Truefor booleans), with caching disabled, cache miss, and cache hit: 182 compilations/cached returns succeeded without warnings. No PCH or trace files were created in the isolated temporary directory. - A wrapper around the real
_program_compile_uncachedfunction confirmed one underlying compilation for each uncached/miss case and zero for each hit. This establishes the tested cache paths; it is not evidence of a wrong-cache-hit bug. - Direct libNVVM calls accepted
-Ofast-compile=0,min,mid, andmaxfor both fixture formats here. Direct-device-float128and-pchcalls returnedNVVM_ERROR_INVALID_OPTION. TheOfastresult is specific to these minimal fixtures and this compiler version, and differs from the text/bitcode distinction reported with CUDA 13.4; it does not prove every optimization level works for arbitrary modules. - Existing NVVM option tests: 42 passed, 7 skipped (six need a cuda-bindings utility absent in 13.2; one needs NVRTC 13.3). NVRTC option regression selection: 58 passed, 1 skipped (selected architecture restriction).
For 1.x compatibility, would a warning for explicitly set NVRTC-only options be preferable to adding hard errors? I would keep existing rejections intact, document options handled outside serialization, and handle
ofast_compileseparately with compiler/version coverage. I have not changed the public behavior while this contract remains unresolved.A coverage test should compare an explicit, independently maintained contract table against the full dataclass field set, so adding a new field without a backend decision fails the test. It should also check explicit false/default values and validation on cache-hit paths.
- An independently enumerated inventory accounts for all 55 public
- addedcuda.coreEverything related to the cuda.core moduleEverything related to the cuda.core module
on Sep 30, 2026 I think these new options should be added to the negativity check here:
cuda-python/cuda_core/cuda/core/_program.pyx
Lines 1401 to 1466 in e8d9075
# Check for unsupported options and raise error if they are set unsupported = [] if opts.relocatable_device_code is not None: unsupported.append("relocatable_device_code") if opts.extensible_whole_program is not None and opts.extensible_whole_program: unsupported.append("extensible_whole_program") if opts.lineinfo is not None and opts.lineinfo: unsupported.append("lineinfo") if opts.ptxas_options is not None: unsupported.append("ptxas_options") if opts.max_register_count is not None: unsupported.append("max_register_count") if opts.use_fast_math is not None and opts.use_fast_math: unsupported.append("use_fast_math") if opts.extra_device_vectorization is not None and opts.extra_device_vectorization: unsupported.append("extra_device_vectorization") if opts.gen_opt_lto is not None and opts.gen_opt_lto: unsupported.append("gen_opt_lto") if opts.define_macro is not None: unsupported.append("define_macro") if opts.undefine_macro is not None: unsupported.append("undefine_macro") if opts.include_path is not None: unsupported.append("include_path") if opts.use_bundled_headers: unsupported.append("use_bundled_headers") if opts.pre_include is not None: unsupported.append("pre_include") if opts.no_source_include is not None and opts.no_source_include: unsupported.append("no_source_include") if opts.std is not None: unsupported.append("std") if opts.builtin_move_forward is not None: unsupported.append("builtin_move_forward") if opts.builtin_initializer_list is not None: unsupported.append("builtin_initializer_list") if opts.disable_warnings is not None and opts.disable_warnings: unsupported.append("disable_warnings") if opts.restrict is not None and opts.restrict: unsupported.append("restrict") if opts.device_as_default_execution_space is not None and opts.device_as_default_execution_space: unsupported.append("device_as_default_execution_space") if opts.device_int128 is not None and opts.device_int128: unsupported.append("device_int128") if opts.optimization_info is not None: unsupported.append("optimization_info") if opts.no_display_error_number is not None and opts.no_display_error_number: unsupported.append("no_display_error_number") if opts.diag_error is not None: unsupported.append("diag_error") if opts.diag_suppress is not None: unsupported.append("diag_suppress") if opts.diag_warn is not None: unsupported.append("diag_warn") if opts.brief_diagnostics is not None: unsupported.append("brief_diagnostics") if opts.time is not None: unsupported.append("time") if opts.split_compile is not None: unsupported.append("split_compile") if opts.fdevice_syntax_only is not None and opts.fdevice_syntax_only: unsupported.append("fdevice_syntax_only") if opts.minimal is not None and opts.minimal: unsupported.append("minimal") if unsupported: raise CUDAError(f"The following options are not supported by NVVM backend: {', '.join(unsupported)}")
PR is welcomed.I suppose another thing we can do is to document that certain options only work for certain
code_type.- addedbugSomething isn't workingSomething isn't workingP2Low priority - Nice to haveLow priority - Nice to haveand removedtriageNeeds the team's attentionNeeds the team's attention
on Sep 30, 2026 Hi, I'd like to take this one if it's still open.
Plan, following @leofang's pointer:
- Add the NVRTC-only fields to the
unsupportedcheck in_prepare_nvvm_options_impl(fdevice_time_trace,device_float128,frandom_seed, the sevenpch*fields), using the sameNone/truthy pattern as the existing entries. - Extend
test_nvvm_options_reject_each_unsupported_flagto cover them, and fix the two comments inutils/_program_cache/_keys.pythat assume NVVM already rejects these. - Add a short note to the
ProgramOptionsdocstring about which options apply to whichcode_type.
Two questions before I start:
no_cacheisn't marked NVRTC-only. Should NVVM reject it too, or just ignore it?ofast_compile: @0z5a found libNVVM accepts-Ofast-compile. Should NVVM forward it rather than reject it?
I don't have a local CUDA/Linux setup right now, so I'll rely on CI for the GPU tests.
- Add the NVRTC-only fields to the
I don't have a local CUDA/Linux setup right now, so I'll rely on CI for the GPU tests.
We cannot accept such PRs unfortunately. GPU CI resources are scarce so please do make sure PRs are verified locally.
In the case of NVRTC and NVVM, they should work on CPU-only machines today (it's not a guarantee, they just happen to work), so you should be able to develop/test/debug locally.
Thanks @leofang, I ran the full
test_program.py+test_program_cache.pyon a GPU with this change applied on top ofe8d9075:- Tesla T4 (sm_75), driver 580.82.07, CUDA 12.x toolkit,
cuda-bindings12.9.9, Python 3.12 - 380 passed, 0 failed, 12 skipped
- All 40 NVVM rejection rows that ran passed, including the 10 new ones. The
use_bundled_headersrow was skipped because it needs NVRTC ≥ 13.3.
The other skips were all version or hardware gates:
numba_debug/use_bundled_headersneed a newer NVRTC,device_float128needs sm_100+, thefdevice_time_tracecompile test is skipped as buggy on NVRTC < 13.0, this libNVVM doesn't recognize-numba-debug, and 6 cases skipped becausecheck_nvvm_compiler_optionsisn't in thiscuda.bindingsbuild.I also tried CUDA 13.4 (
cuda-bindings13.4.3) on a CPU-only Linux machine: the 10 new rejection rows fail onmainand pass with the fix. The full files can't be collected there withoutlibcuda, so that run was a subset. libNVVM 12.9 and 13.4 both returnNVVM_ERROR_INVALID_OPTIONfor each of those 10 flags too, so they were never doing anything on NVVM.Two things I left out of this change since I wasn't sure how you'd want them handled:
ofast_compile: libNVVM accepts-Ofast-compile, but on 13.4min/mid/maxfail with textual IR (bitcode is fine). Should it be forwarded or rejected?no_cache: reject on NVVM, or just ignore it?
Happy to open a PR as soon as I'm assigned.
- Tesla T4 (sm_75), driver 580.82.07, CUDA 12.x toolkit,
Is this a duplicate?
Type of Bug
Silent Failure
Component
cuda.core
Describe the bug
_prepare_nvvm_options_implincuda_core/cuda/core/_program.pyxraisesCUDAErrorfor 31ProgramOptionsfields that libNVVM can't use. Twelve other fields, which the NVRTC path does emit, are neither emitted nor rejected on the NVVM path. Setting one of them has no effect, and there's no error or warning:no_cache,fdevice_time_trace,device_float128,frandom_seed,ofast_compile,pch,create_pch,use_pch,pch_dir,pch_verbose,pch_messages,instantiate_templates_in_pchAll of them except
no_cacheare documented as "(NVRTC only)" in theProgramOptionsdocstring.device_int128is rejected on NVVM, butdevice_float128compiles without a word.The rejection list walks the fields in almost the same order as
_prepare_nvrtc_options_impland stops atminimal, which is the field just beforeno_cachein the NVRTC builder. My guess is the fields after that point, anddevice_float128, were never added. I haven't found anything saying it was deliberate.Two other places assume NVVM rejects some of these:
cuda_core/cuda/core/utils/_program_cache/_keys.pysays NVVM "explicitly rejects all three at compile time" aboutcreate_pch,timeandfdevice_time_trace, and gates the side-effect check on NVRTC for that reason. It also says "NVVM rejects them" about the external-content options (include_path,pre_include,pch,use_pch,pch_dir). Those two comments cover eight options. NVVM rejects three of them (time,include_path,pre_include) and silently drops the other five. Withcreate_pchorfdevice_time_traceset,make_program_cache_key(code_type="nvvm", ...)returns a key, whilecode_type="c++"raisesValueError. I don't think that produces a wrong cache hit, since the compile drops the option as well.test_nvvm_options_reject_each_unsupported_flagincuda_core/tests/test_program.pysays its table "mirrors _prepare_nvvm_options_impl's rejection list one-for-one", so a field missing from both can't fail it.How to Reproduce
Output with
cuda.corebuilt frommainat f9ed2bd:Expected behavior
Each of these should either reach libNVVM or tell the user it was dropped. Which of the two is your call. The existing list raises
CUDAError. On #2573, though, the review preferred aUserWarningover a new error in 1.x, following #2658, and the same reasoning looks like it applies here.ofast_compilemay be one to pass through rather than reject, though I'm not sure how useful it is there. Calling libNVVM directly on the IR above (nvidia-nvvm13.4.92,nvvm.version()returns(2, 0)),-Ofast-compile=0compiles.=min,=midand=maxeach returnERROR_COMPILATION (9)with the logparse Can't read textual IR with a Context that discards named Values. So only0worked on this textual IR, but unlike-pchand-device-float128, which returnERROR_INVALID_OPTION (7), it isn't rejected as an unknown option.link_time_optimizationis also neither emitted nor rejected on NVVM. I left it off the list becausetest_nvvm_program_optionspasses it to NVVM on purpose, andtarget_type="ltoir"adds-gen-ltoanyway.Operating System
Ubuntu 24.04.4 LTS
nvidia-smi output
NVIDIA GeForce RTX 5090, driver 610.43.02.
cuda.corebuilt from source at f9ed2bd against cuda-bindings 13.4.3 and the cuda-toolkit 13.4.2 wheels, Python 3.12.13.