AscVF passes

-ascvf-dispatch-vf-fusion

Inserts LocalMemBarOp synchronization barriers at required points within VecScopeOp regions

Inserts LocalMemBarOp synchronization barriers at the points within VecScopeOp regions where store-to-load dependencies must be preserved. Determines those points by marking every operation with a BarrierOp, cloning the function, running the vector fusion optimization pipeline and EraseBarrier on the clone, and mapping the barriers that survive back to LocalMemBarOp operations in the original function before running the same pipeline on it.

-ascvf-eliminate-common-mask

Eliminates duplicate mask operations within VFGroupOp

Removes redundant UpdateMaskOp and CreateMaskOp operations inside VFGroupOp regions by replacing uses of duplicate masks with the first occurrence and deleting the duplicates.

-ascvf-eliminate-data-transfer

Eliminates redundant load/store operations in VFGroupOp

Performs three optimizations: eliminates loads after stores by replacing the loaded value with the stored value, removes stores that are immediately overwritten, and merges repeated loads from the same memory location.

-ascvf-erase-barrier

Erases redundant BarrierOp operations within VecScopeOp that protect no pending stores

Erases redundant BarrierOp operations within VecScopeOp regions by sorting LoadOp, StoreOp, and BarrierOp operations by their relative position, collapsing consecutive barriers to the last one, and removing barriers with no preceding active stores. Emits a warning when a LoadOp encounters unsynchronized stores.

-ascvf-find-vf-group

Finds and groups L2 operations into VFGroupOp for vector fusion

Identifies groups of fusible L2 operations (binary, unary, reduce, duplicate) that share the same calCount and element type, and wraps them into VFGroupOp operations for subsequent vector fusion optimization.

-ascvf-fuse-vf-for

Fuses consecutive VFForOp loops within VecScopeOp

Merges adjacent VFForOp loops into a single loop by copying operations from subsequent loops into the first loop, reducing loop overhead.

-ascvf-hoist-cast

Hoists LocalTensorReinterpretCastOp operations immediately after their defining operations

Relocates each LocalTensorReinterpretCastOp to immediately after the operation that defines its input tensor, placing the cast right after its source.

-ascvf-inline-vf-group

Replaces ascvf.vf_group operation with its body

-ascvf-lower-to-reg

Lowers L2 operations to register-level operations with loops

Converts L2 operations (binary, unary, reduce, duplicate) into micro-level operations by wrapping them in VecScopeOp, creating VFForOp loops, allocating register tensors, and inserting load/store micro operations.

-ascvf-materialize-load-store

Materializes abstract load/store micro ops into data copy operations

Converts LoadOp and StoreOp operations into concrete data copy operations with physical addresses, creating LocalTensorGetPhyAddrV2Op and DataCopyLoadOp/DataCopyStoreOp operations.

-ascvf-reorder-ops-in-vec-scope

Hoists and reorders operations within VecScopeOp blocks

Reorders operations inside VecScopeOp blocks by moving constants, register tensors, variables, mask operations, and duplicate operations to the beginning in a specific order for better code generation.