Update CUDA to version 13.3.0#10617
Conversation
|
enable gpu |
|
please test |
|
A new Pull Request was created by @fwyzard for branch IB/CMSSW_20_1_X/master. @akritkbehera, @iarspider, @raoatifshad, @smuzaffar can you please review it and eventually sign? Thanks. |
|
cms-bot internal usage |
|
-1 Summary: https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-33cbca/53787/summary.html Failed External BuildI found compilation error when building: CMSBUILD_PACKAGE=onnxruntime; export _CMSBUILD_BUILD_ENV_=1; export TMPDIR=/data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271/cmsdist-tmp; source /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/el9_amd64_gcc13/rpm-env.sh; source /data/cmsbld/jenkins/workspace/ib-run-pr-tests/pkgtools/build-system-env.sh 16; rpmbuild --noclean --buildroot /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/tmp/BUILDROOT/3d403218bee3a6bbf94b3a90e3dc8271 -bc --short-circuit --define '_smp_build_nthreads 16' --define '_topdir /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir' --define "compiling_processes 16" --define '_builddir /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271' --define "cmsdist_directory /data/cmsbld/jenkins/workspace/ib-run-pr-tests/cmsdist" --define "compilerv 1340" --define "cmscompilerv 13" --define "cmsos el9_amd64" --define "__bootstrap ~bootstrap" --define "almalinux_ver 9" --define "almalinux 9" --define "centos_ver 9" --define "centos 9" --define "rhel 9" --define "dist %{!?distprefix0:%{?distprefix}}%{expand:%{lua:for i=0,9999 do print("%{?distprefix" .. i .."}") end}}.el9%{?with_bootstrap:~bootstrap}" --define "el9 1" --define "dist_vendor AlmaLinux OS Foundation" --define "dist_name AlmaLinux" --define "dist_home_url https://almalinux.org/" --define "dist_bug_report_url https://bugs.almalinux.org/" --define "package_vectorization x86-64-v2" --define "cmsswdata_version_link 1" --define "archfirst yes" --define "cmsBuild_bootstrap 1" /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/SPECS/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271/spec >>/data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271/log 2>&1
build-build-external+onnxruntime+1.26.0-3d403218bee3a6bbf94b3a90e3dc8271 failed.
Failed to build onnxruntime. Log file in /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271/log. Final lines of the log file:
171 | ABSL_INTERNAL_X(insert_or_assign, insert_or_assign_impl, const &, &&, false,
| ^
/data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271/build/_deps/abseil_cpp-src/absl/container/internal/raw_hash_map.h:180:200: error: using template type parameter 'absl::lts_20250814::container_internal::IfRRef::AddPtr' after 'typename'
180 | ABSL_INTERNAL_X(insert_or_assign, insert_or_assign_impl, &&, const &, false,
| ^
/data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-3d403218bee3a6bbf94b3a90e3dc8271/build/_deps/abseil_cpp-src/absl/container/internal/raw_hash_map.h:180:211: error: template argument 6 is invalid
180 | ABSL_INTERNAL_X(insert_or_assign, insert_or_assign_impl, &&, const &, false,
| ^
|
8eabad3 to
4e5e2bb
Compare
|
please test |
|
Pull request #10617 was updated. |
|
please test |
|
-1 Summary: https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-33cbca/53809/summary.html Failed External BuildI found compilation error when building: cwd: /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/py3-torch-cluster-cuda/1.6.3-256da426cb7b958a8355fcf01af22be7/cmsdist-pip-src/torch_cluster-1.6.3
Building wheel for torch_cluster (pyproject.toml): finished with status 'error'
ERROR: Failed building wheel for torch_cluster
Failed to build torch_cluster
ERROR: Failed to build one or more wheels
error: Bad exit status from /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/tmp/rpm-tmp.ZWcyEt (%build)
RPM build warnings:
Macro expanded in comment on line 631: %{pkginstroot}/bin/*
Macro expanded in comment on line 636: %{pkginstroot}/${PYTHON3_LIB_SITE_PACKAGES}
|
4e5e2bb to
b10b055
Compare
|
Pull request #10617 was updated. |
|
Pull request #10617 was updated. |
|
please test |
|
-1 Summary: https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-33cbca/53818/summary.html Failed External BuildI found compilation error when building: install-cms+cms-git-tools+251202.0-c4424a3c103e232b58c2aed3dbff41ee done build-build-external+onnxruntime+1.26.0-92eddeb1b3f2986c95279595476fcca1 failed. Failed to build onnxruntime. Log file in /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-92eddeb1b3f2986c95279595476fcca1/log. Final lines of the log file: 171 | ABSL_INTERNAL_X(insert_or_assign, insert_or_assign_impl, const &, &&, false, | ^ /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-92eddeb1b3f2986c95279595476fcca1/build/_deps/abseil_cpp-src/absl/container/internal/raw_hash_map.h:180:200: error: using template type parameter 'absl::lts_20250814::container_internal::IfRRef::AddPtr' after 'typename' 180 | ABSL_INTERNAL_X(insert_or_assign, insert_or_assign_impl, &&, const &, false, | ^ /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/BUILD/el9_amd64_gcc13/external/onnxruntime/1.26.0-92eddeb1b3f2986c95279595476fcca1/build/_deps/abseil_cpp-src/absl/container/internal/raw_hash_map.h:180:211: error: template argument 6 is invalid 180 | ABSL_INTERNAL_X(insert_or_assign, insert_or_assign_impl, &&, const &, false, | ^ |
|
please test |
|
please abort |
Fix invalid C++ syntax in cub/device/device_transform.cuh in CUDA 13. Backport NVIDIA/cccl#8771 from the main branch. Drop the Volta (sm 7.0) architecture, that is not supported by CUDA 13.
Major updates: - support for DMA-BUF mmap backend with CUDA 13.3 - support for AVX2, AVX-512, MOVDIR64B, NEON CPU instructions - support for Linux kernel 6.15+ The updates, fixes and changes in v2.5.2 can be found at https://github.com/NVIDIA/gdrcopy/releases/tag/v2.5.2 . The updates, fixes and changes in v2.6 can be found at https://github.com/NVIDIA/gdrcopy/releases/tag/v2.6 .
Update CMake to the latest 3.x bugfix version, v3.31.12. Backport cmake!12162 CUDA: Add support for cuda_std_23 for nvcc 13.3+.
Backport pytorch/pytorch#180996 from the main branch
Backport microsoft/onnxruntime#28736 Patch Abseil to workaround NVIDIA bug #6302392 Do not include <ciso646> in c++20/c++23 mode
f4eda55 to
7463906
Compare
|
Pull request #10617 was updated. |
|
please test |
|
@smuzaffar @mandrenguyen @ftenchini some non-gpu tests are still running, but all NVIDIA gpu tests passed, finally this looks like it's ready to be merged. |
|
please test for CMSSW_20_1_CPP23_X lets see if this works for gcc14/c++23 |
|
+1 Summary: https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-33cbca/53940/summary.html The following merge commits were also included on top of IB + this PR after doing git cms-merge-topic:
You can see more details here: Comparison SummarySummary:
NVIDIA_H100 Comparison SummarySummary:
NVIDIA_L40S Comparison SummarySummary:
NVIDIA_T4 Comparison SummarySummary:
Max Memory Comparisons exceeding threshold NVIDIA_T4@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
|
|
+externals lets get this in IBs |
|
This pull request is fully signed and it will be integrated in one of the next IB/CMSSW_20_1_X/master IBs (tests are also fine). This pull request will now be reviewed by the release team before it's merged. @ftenchini, @sextonkennedy, @mandrenguyen (and backports should be raised in the release meeting by the corresponding L2) |
|
@fwyzard , with cuda 13.3.0 , we do get a lot of warnings (from alpaka, I think) like below currently as these are not directly generated by cmssw code so we ignore these. Is there any thing we can do to avoid these ... may be new version of alpaka? |
|
@smuzaffar a quick fix would be to add Support for CUDA 13.3.0 is still being merged in alpaka, I can look into fixing or silencing the warnings in the coming few days. |
|
Actually the |
Uh oh!
There was an error while loading. Please reload this page.