diff options
| author | Kirill Vedernikov <kvedernikov@nvidia.com> | 2025-11-21 13:13:52 +0100 |
|---|---|---|
| committer | GitHub <noreply@github.com> | 2025-11-21 17:43:52 +0530 |
| commit | 2f627c1878a3dba594c872773107c556992af3a1 (patch) | |
| tree | 8025610c667972b3fe1e4878902f0357f04d300b /lldb/test/API/functionalities/scripted_process/TestStackCoreScriptedProcess.py | |
| parent | 18d3db4bcd42e21e45b499a2999834904a925af0 (diff) | |
[NVPTX] Support for dense and sparse MMA intrinsics with block scaling. (#163561)
This change adds dense and sparse MMA intrinsics with block scaling. The
implementation is based on [PTX ISA version
9.0](https://docs.nvidia.com/cuda/parallel-thread-execution/). Tests for
new intrinsics are added for PTX 8.7 and SM 120a and are generated by
`llvm/test/CodeGen/NVPTX/wmma-ptx87-sm120a.py`. The tests have been
verified with ptxas from CUDA-13.0 release.
Dense MMA intrinsics with block scaling were supported by
@schwarzschild-radius.
Diffstat (limited to 'lldb/test/API/functionalities/scripted_process/TestStackCoreScriptedProcess.py')
0 files changed, 0 insertions, 0 deletions
