llvm-project.git - Unnamed repository; edit this file 'description' to name the repository.

Age	Commit message (Collapse)	Author
2025-05-23	[NFC][CodeGen] Adopt MachineFunctionProperties convenience accessors (#141101)	Rahul Joshi

2025-05-14	[GlobalISel] Add a GISelValueTracker printing pass (#139687)	David Green
	This adds a GISelValueTrackingPrinterPass that can print the known bits and sign bit of each def in a function. It is built on the new pass manager and so adds a NPM GISelValueTrackingAnalysis, renaming the older class to GISelValueTrackingAnalysisLegacy. The first 2 functions from the AArch64GISelMITest are ported over to an mir test to show it working. It also runs successfully on all files in llvm/test/CodeGen/AArch64/GlobalISel/*.mir that are not invalid. It can hopefully be used to test GlobalISel known bits analysis more directly in common cases, without jumping through the hoops that the C++ tests requires.
2025-04-07	[NFC][LLVM][AMDGPU] Cleanup pass initialization for AMDGPU (#134410)	Rahul Joshi
	- Remove calls to pass initialization from pass constructors. - https://github.com/llvm/llvm-project/issues/111767
2025-03-29	[GlobalISel][NFC] Rename GISelKnownBits to GISelValueTracking (#133466)	Tim Gymnich
	- rename `GISelKnownBits` to `GISelValueTracking` to analyze more than just `KnownBits` in the future
2025-01-06	[AMDGPU] [GlobalIsel] Combine Fmul with Select into ldexp instruction. (#120104)	Vikash Gupta
	This combine pattern perform the below transformation. fmul x, select(y, A, B) -> fldexp (x, select i32 (y, a, b)) fmul x, select(y, -A, -B) -> fldexp ((fneg x), select i32 (y, a, b)) where, A=2^a & B=2^b ; a and b are integers. It is a follow-up PR to implement the above combine for globalIsel, as the corresponding DAG combine has been done for SelectionDAG Isel (#111109)
2024-08-22	[AMDGPU][GlobalISel] Disable fixed-point iteration in all Combiners (#105517)	Jay Foad
	Disable fixed-point iteration in all AMDGPU Combiners after #102163. This saves around 2% compile time in ad hoc testing on some large graphics shaders. I did not notice any regressions in the generated code, just a bunch of harmless differences in instruction selection and register allocation.
2024-06-21	[NFC] Fix laod -> load typos. NFC	David Green

2024-06-11	[CodeGen][NewPM] Split `MachineDominatorTree` into a concrete analysis ↵	paperchalice
	result (#94571) Prepare for new pass manager version of `MachineDominatorTreeAnalysis`. We may need a machine dominator tree version of `DomTreeUpdater` to handle `SplitCriticalEdge` in some CodeGen passes.
2024-05-06	[AMDGPU] Fix typo in function name	Jay Foad

2024-05-05	[AMDGPU] Improve MIR pattern for FMinFMaxLegacy combine. NFC. (#90968)	Jay Foad

2024-05-03	[AMDGPU] Use replaceOpcodeWith instead of applyCombine_s_mul_u64. NFC.	Jay Foad

2024-05-03	[AMDGPU] Remove unneeded calls to setInstrAndDebugLoc in matchers. NFC.	Jay Foad

2024-05-03	[AMDGPU] Simplify applySelectFCmpToFMinToFMaxLegacy. NFC.	Jay Foad

2024-02-22	[AMDGPU][GlobalISel] Add fdiv / sqrt to rsq combine (#78673)	Nick Anderson
	Fixes #64743
2024-01-17	[AMDGPU] CodeGen for GFX12 8/16-bit SMEM loads (#77633)	Jay Foad

2024-01-10	[AMDGPU] Fix broken sign-extended subword buffer load combine (#77470)	Jay Foad

2024-01-08	[AMDGPU] Add CodeGen support for GFX12 s_mul_u64 (#75825)	Jay Foad

2023-09-14	[NFC][CodeGen] Change CodeGenOpt::Level/CodeGenFileType into enum classes ↵	Arthur Eubanks
	(#66295) This will make it easy for callers to see issues with and fix up calls to createTargetMachine after a future change to the params of TargetMachine. This matches other nearby enums. For downstream users, this should be a fairly straightforward replacement, e.g. s/CodeGenOpt::Aggressive/CodeGenOptLevel::Aggressive or s/CGFT_/CodeGenFileType::
2023-09-05	[GlobalISel] Refactor Combiner API	pvanhout
	Remove CodeGen leftovers from the old combiner backend and adapt the API to fit the new backend better. It's now quite a bit closer to how InstructionSelector works. - `CombinerInfo` is now a simple "options" struct. - `Combiner` is now the base class of all TableGen'd combiner implementation. - Many fields have been moved from derived classes into that class. - It has been refactored to create & own the Observer and Builder. - `tryCombineAll` TableGen'd method can now be renamed, which allows targets to implement the actual `tryCombineAll` call manually and do whatever they want to do before/after it. Note: `CombinerHelper` needs to be mutable because none of its methods are const. This can be revisited later. Depends on D158710 Reviewed By: aemerson, dsanders Differential Revision: https://reviews.llvm.org/D158713
2023-08-23	AMDGPU: Fix more unsafe rsq formation	Matt Arsenault
	Introducing rsq contract flags is wrong, and also requires some level of approximate functions. AMDGPUCodeGenPrepare already should handle the f32 cases with appropriate flags, and I don't see how new situations to handle would arise during legalization (other than cases involving the rcp intrinsic, which instcombine tries to handle). AMDGPUCodeGenPrepare does need to learn better handling of rcp/rsq for f64 though, which we never bothered to handle well. Removes another obstacle to correctly lowering sqrt. https://reviews.llvm.org/D158099
2023-07-31	[GlobalISel] convergent intrinsics	Sameer Sahasrabuddhe
	Introduced the convergent equivalent of the existing G_INTRINSIC opcodes: - G_INTRINSIC_CONVERGENT - G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS Out of the targets that currently have some support for GlobalISel, the patch assumes that the convergent intrinsics only relevant to SPIRV and AMDGPU. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D154766
2023-07-27	Restore "[GlobalISel] GIntrinsic subclass to represent intrinsics in Generic ↵	Sameer Sahasrabuddhe
	Machine IR" Some opcodes in generic MIR represent calls to intrinsics, where the intrinsic ID is the first non-def operand to the instruction. These are now represented as a subclass of GenericMachineInstr, and the method MachineInstr::getIntrinsicID() is now moved to this subclass GIntrinsic. Some target-defined instructions behave like GMIR intrinsics, and have an Intrinsic::ID operand. But they should not be recognized as generic intrinsics, and should not use GIntrinsic::getIntrinsicID(). Separated these out by introducing a new AMDGPU::getIntrinsicID(). Reviewed By: arsenm, Pierre-vh Differential Revision: https://reviews.llvm.org/D155556 This restores commit baa3386edb11a2f9bcadda8cf58d56f3707c39fa. Originally reverted in d0f7850b01cf17e50a4f4b00e3b84dded94df6b8.
2023-07-27	Revert "[GlobalISel] GIntrinsic subclass to represent intrinsics in Generic ↵	Sameer Sahasrabuddhe
	Machine IR" This reverts commit baa3386edb11a2f9bcadda8cf58d56f3707c39fa. The changes did not cover all occurrences of the deteleted function MachineInstr::getIntrinsicID().
2023-07-27	[GlobalISel] GIntrinsic subclass to represent intrinsics in Generic Machine IR	Sameer Sahasrabuddhe
	Some opcodes in generic MIR represent calls to intrinsics, where the intrinsic ID is the first non-def operand to the instruction. These are now represented as a subclass of GenericMachineInstr, and the method MachineInstr::getIntrinsicID() is now moved to this subclass GIntrinsic. Some target-defined instructions behave like GMIR intrinsics, and have an Intrinsic::ID operand. But they should not be recognized as generic intrinsics, and should not use GIntrinsic::getIntrinsicID(). Separated these out by introducing a new AMDGPU::getIntrinsicID(). Reviewed By: arsenm, Pierre-vh Differential Revision: https://reviews.llvm.org/D155556
2023-07-11	[AMDGPU] Use GlobalISel MatchTable Combiner Backend	pvanhout
	Use the new matchtable-based combiner backend for all AMDGPU combiners. This drop-in from the user's perspective; there are no test changes, the new combiner behaves exactly like the old one. Depends on D153757 NOTE: This would land iff D153757 (RFC) lands too. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153758
2023-02-22	[AMDGPU] Improve the lowering of raw_buffer_load_{i8,i16} and ↵	Konstantina Mitropoulou
	struct_buffer_load_{i8,i16} intrinsics Currently, raw_buffer_load_{i8,i16} and struct_buffer_load_{i8,i16} intrinsics are lowered as buffer_load_{u8,u16}. This patch combines buffer_load_{u8,u16} and sign extension instructions in order to generate buffer_load_{i8,i16} instructions. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D144313
2023-02-02	AMDGPU: Try to unfold fneg source when matching legacy fmin/fmax	Matt Arsenault
	This is NFC as it stands, since other combines will effectively prevent this from being reachable. This will avoid regressions in a future change which tries to make better use of select source modifiers. Didn't bother with the GlobalISel part for now, since the baseline combine doesn't seem to work on the existing test.
2022-10-26	[GlobalISel] Add Predicates to GICombineRule	Pierre van Houtryve
	Small QoL change to allow Predicates to be used in GICombineRule. Currently only one combine in the AMDGPU backend makes use of it. The implementation is pretty simple to get started but of course we can expand this later on and optimize predicate checking better if needed. Reviewed By: dsanders Differential Revision: https://reviews.llvm.org/D136681
2022-10-07	[GlobalISel] Mark mi_match as nodiscard	Jessica Paquette
	Typically when you match something, you want to check the result. Fix a couple warnings in the AMDGPUPostLegalizerCombiner which appear as a result of this. Differential Revision: https://reviews.llvm.org/D135491
2022-10-03	[GlobalISel] Allow prelegalizer combiners to have access to LegalizerInfo.	Amara Emerson
	Before, the isPreLegalize() query in CombinerHelper only checked for the presence of a LegalizerInfo object. This is problematic when we want to have a combine actually check for legality in a pre-legalizer combine pass, since if we pass a LegalizerInfo object to the constructor it causes the combines to think that we're running post legalizer, which isn't true. This change fixes it to instead check an explicit bool that passes to signal whether the pass will be run before or after legalization. Doing so exposed a bug in the extending loads combine, which tried to check for legality of candidate extending loads if LegalizerInfo was present. Since we only ran it pre-legalizer and therefore with a null LegalizerInfo, it never actually ran. Also fixes the legality checks to keep the tests passing. Differential Revision: https://reviews.llvm.org/D135044
2021-11-30	Code quality: Combine V_RSQ	Mateja Marjanovic
	Combine V_RCP and V_SQRT into V_RSQ on AMDGPU for GlobalISel. Change-Id: I93c5dcb412483156a6e8b68c4085cbce83ac9703
2021-11-17	[AMDGPU][GlobalISel] Fold G_FNEG above when users cannot fold mods	Mirko Brkusanin
	If possible fold fneg into instruction above if users cannot fold mods and we know it will decrease instruction count. Follows same logic as SDAG combiner in choosing opportunities to combine. Differential Revision: https://reviews.llvm.org/D112827
2021-04-27	AMDGPU/GlobalISel: Remove redundant G_FCANONICALIZE	Petar Avramovic
	Add basic version of isCanonicalized for global-isel. Copied from sdag. Add post legalizer combine that deletes G_FCANONICALIZE when its input is already Canonicalized. Differential Revision: https://reviews.llvm.org/D96605
2021-02-02	Fixed includes.	Thomas Symalla
	Differential Revision: https://reviews.llvm.org/D93708
2021-02-02	Reverted whitespace changes.	Thomas Symalla
	Differential Revision: https://reviews.llvm.org/D90968
2021-02-02	Formatting changes	Thomas Symalla

2021-02-02	Formatting changes.	Thomas Symalla

2021-02-02	Updating formatting changes.	Thomas Symalla

2021-02-02	Resolve formatting changes.	Thomas Symalla

2021-02-02	Move Combiner to PreLegalize step	Thomas Symalla

2021-02-02	Reverted unintended git-format change.	Thomas Symalla

2021-02-02	Fixed the lit tests and a bug in the implementation.	Thomas Symalla

2021-02-02	Refactored the pattern matching.	Thomas Symalla

2021-02-02	Added early exit.	Thomas Symalla

2021-02-02	Added comments.	Thomas Symalla

2021-02-02	clang-format	Thomas Symalla

2021-02-02	Added clamp i64 to i16 global isel pattern.	Thomas Symalla

2021-01-20	[NFC][AMDGPU] Split AMDGPUSubtarget.h to R600 and GCN subtargets	dfukalov
	... to reduce headers dependency. Reviewed By: rampitec, arsenm Differential Revision: https://reviews.llvm.org/D95036
2021-01-07	[NFC][AMDGPU] Reduce include files dependency.	dfukalov
	Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D93813
2020-11-03	AMDGPU/GlobalISel: Use same builder/observer in post-legalizer-combiner	Petar Avramovic
	Change match/apply functions into methods of new target specific combiner helper class. Use reference to MachineIRBuilder from helper instead of constructing new MachineIRBuilder each time new instruction needs to made. Allows correct tracking of newly created instructions. Differential Revision: https://reviews.llvm.org/D90623