diff options
| author | Nicolai Hähnle <nicolai.haehnle@amd.com> | 2025-11-21 11:33:13 -0800 |
|---|---|---|
| committer | GitHub <noreply@github.com> | 2025-11-21 19:33:13 +0000 |
| commit | 69589dd2c0b34a664c24f7ffbb084d2eea848ab6 (patch) | |
| tree | 9ddbc5604ff047683dd8119e7c45cf9ee896db27 /lldb/test/API/functionalities/scripted_process/TestStackCoreScriptedProcess.py | |
| parent | 52f9a57b2961da168b2a5ffe9eee687fe9068c2b (diff) | |
AMDGPU: Improve getShuffleCost accuracy for 8- and 16-bit shuffles (#168818)
These shuffles can always be implemented using v_perm_b32, and so this
rewrites the analysis from the perspective of "how many v_perm_b32s does
it take to assemble each register of the result?"
The test changes in Transforms/SLPVectorizer/reduction.ll are
reasonable: VI (gfx8) has native f16 math, but not packed math.
Diffstat (limited to 'lldb/test/API/functionalities/scripted_process/TestStackCoreScriptedProcess.py')
0 files changed, 0 insertions, 0 deletions
