AscVF passes
-ascvf-dispatch-vf-fusion
Inserts LocalMemBarOp synchronization barriers at required points within VecScopeOp regions
Inserts LocalMemBarOp synchronization barriers at the points within VecScopeOp regions where store-to-load dependencies must be preserved. Determines those points by marking every operation with a BarrierOp, cloning the function, running the vector fusion optimization pipeline and EraseBarrier on the clone, and mapping the barriers that survive back to LocalMemBarOp operations in the original function before running the same pipeline on it.
-ascvf-eliminate-common-mask
Eliminates duplicate mask operations within VFGroupOp
Removes redundant UpdateMaskOp and CreateMaskOp operations inside VFGroupOp regions by replacing uses of duplicate masks with the first occurrence and deleting the duplicates.
-ascvf-eliminate-data-transfer
Eliminates redundant load/store operations in VFGroupOp
Performs three optimizations: eliminates loads after stores by replacing the loaded value with the stored value, removes stores that are immediately overwritten, and merges repeated loads from the same memory location.
-ascvf-erase-barrier
Erases redundant BarrierOp operations within VecScopeOp that protect no pending stores
Erases redundant BarrierOp operations within VecScopeOp regions by sorting LoadOp, StoreOp, and BarrierOp operations by their relative position, collapsing consecutive barriers to the last one, and removing barriers with no preceding active stores. Emits a warning when a LoadOp encounters unsynchronized stores.
-ascvf-find-vf-group
Finds and groups L2 operations into VFGroupOp for vector fusion
Identifies groups of fusible L2 operations (binary, unary, reduce, duplicate) that share the same calCount and element type, and wraps them into VFGroupOp operations for subsequent vector fusion optimization.
-ascvf-fuse-vf-for
Fuses consecutive VFForOp loops within VecScopeOp
Merges adjacent VFForOp loops into a single loop by copying operations from subsequent loops into the first loop, reducing loop overhead.
-ascvf-hoist-cast
Hoists LocalTensorReinterpretCastOp operations immediately after their defining operations
Relocates each LocalTensorReinterpretCastOp to immediately after the operation that defines its input tensor, placing the cast right after its source.
-ascvf-inline-vf-group
Replaces ascvf.vf_group operation with its body
-ascvf-lower-to-reg
Lowers L2 operations to register-level operations with loops
Converts L2 operations (binary, unary, reduce, duplicate) into micro-level operations by wrapping them in VecScopeOp, creating VFForOp loops, allocating register tensors, and inserting load/store micro operations.
-ascvf-materialize-load-store
Materializes abstract load/store micro ops into data copy operations
Converts LoadOp and StoreOp operations into concrete data copy operations with physical addresses, creating LocalTensorGetPhyAddrV2Op and DataCopyLoadOp/DataCopyStoreOp operations.
-ascvf-reorder-ops-in-vec-scope
Hoists and reorders operations within VecScopeOp blocks
Reorders operations inside VecScopeOp blocks by moving constants, register tensors, variables, mask operations, and duplicate operations to the beginning in a specific order for better code generation.