AscendC (from AscTile extension) passes
-ascendc-allocate-tensor
Assign memory addresses to LocalTensorAutoOp placeholders for static UB allocation
Converts LocalTensorAutoOp placeholders to LocalTensorV3Op with concrete memory addresses and tile sizes.
This pass performs static memory allocation for on-chip buffers by assigning consecutive aligned addresses
to tensors based on their normalized memory position (A1/L1, A2/L0A, B2/L0B, CO1/L0C, VECCALC/UB).
Position normalization: B1→A1, VECIN/VECOUT→VECCALC. Addresses are aligned to ubBlockSize (256 bytes).
After allocation, LocalTensorAutoOp is replaced with LocalTensorV3Op containing position, address, and tile size.
-ascendc-compute-memory-consumption
Calculate total memory consumption per buffer location and attach to module
Analyzes all LocalTensorV3Op operations and computes total memory usage for each hardware buffer location:
L1, L0A, L0B, L0C, and UB. The results are stored as a dictionary attribute memoryConsumed on the module,
with keys like “L1”, “L0A”, etc. and values in bytes.
This pass is useful for verifying that kernel memory usage stays within hardware limits.
-ascendc-compute-reuse-group
Compute reuse group indices for unrolled loop iterations to guide tensor allocation reuse
Computes ascendc.reuse_group indices for operations in unrolled loops to guide downstream tensor
allocation reuse. The pass analyzes scf.execute_region operations with asctile.unroll_factor and
asctile.unroll_iter attributes, and remaps iteration indices into reuse groups.
The reuse group is computed as (unroll_iter + startIndex) % unrollFactor, where startIndex accounts
for nesting level and epilogue iterations. An epilogue iteration is detected when unroll_iter equals
unroll_factor (i.e., the remainder-handling tail of an unrolled loop). Epilogue iterations shift the
start index for subsequent loops at the same nesting level, ensuring correct reuse group alignment.
After processing, asctile.unroll_iter attributes are replaced with ascendc.reuse_group on all
operations within unrolled loops.
-ascendc-dispatch-alloc
Dispatch tensor allocation to static or TPipe-backed strategy based on target architecture
Selects between static allocation and TPipe-backed allocation for on-chip tensor buffers. The decision
is based on the asc.static_alloc module attribute, target architecture, and presence of incompatible
operations. When targeting C310 architectures without BroadcastOp, static allocation is preferred;
otherwise TPipe-backed allocation is used.
-ascendc-fill-asc-operands
Fill default cal_count, mask, repeat_times, and repeat_params operands for L0/L2 Ascend C operations
Populates default operands for vector operations based on tensor shape:
L0 operations: Fills
mask,repeat_times, andrepeat_params(stride values)L2 operations: Fills
cal_countwith number of elements
Default stride values: block_stride=1, repeat_stride=8. Mask is computed based on element type size (1, 2, 4, 8). Repeat times calculated as ceil(numElements / numPerRepeat).
-ascendc-fixup-mmad-acc-params
Fix accumulator initialization for sequential mmad operations using the same destination tensor
Handles matrix multiplication accumulation where multiple mmad operations write to the same destination tensor.
Creates a runtime boolean variable to track whether this is the first mmad to the destination:
First mmad:
cmatrixInitVal=true(initialize accumulator)Subsequent mmads:
cmatrixInitVal=false(accumulate into existing values)
Only processes mmad operations that have cmatrixInitVal parameter in their mmad_params.
Example transformation:
// Before: mmad with fixed cmatrixInitVal in loop
%acc = ascendc.local_tensor_auto co1() : <16x16xf32>
scf.for %i = 0 to %n {
%params = emitasc.init_struct !ascendc.mmad_params(
"cmatrixInitVal" = %c1 // always true (reinit each iteration)
)
ascendc.mmad %acc, %a, %b, %params
}
// After: runtime variable controls initialization
%acc = ascendc.local_tensor_auto co1() : <16x16xf32>
%var = emitasc.variable true, memref<1xi1> // init flag
scf.for %i = 0 to %n {
%init_val = memref.load %var[%c0]
%params = emitasc.init_struct !ascendc.mmad_params(
"cmatrixInitVal" = %init_val // true first iter, false after
)
ascendc.mmad %acc, %a, %b, %params
%false = arith.constant false
memref.store %false, %var[%c0] // set false for next iter
}
-ascendc-fuse-bufid-sync
Remove redundant get_buf/rls_buf synchronization between consecutive operations with matching bufId
Optimizes BufId-based synchronization by removing redundant get_buf/rls_buf operations:
For consecutive operations with the same
bufIdattribute and same pipeline type:Remove
get_bufbefore middle operations (only first op needs it)Remove
rls_bufafter middle operations (only last op needs it)
Pipeline types tracked: PIPE_V (vector), PIPE_MTE2 (GM→UB load), PIPE_MTE3 (UB→GM store)
After optimization, removes the bufId attribute from all operations. Used on C310 architecture.
-ascendc-insert-bias-bufid-sync
Insert get_buf/rls_buf synchronization around MmadOp for bias tensors copied to BT
Inserts buffer synchronization (get_buf/rls_buf) around MmadOp operations that use bias tensors
in C2 position (bias tensor memory).
This pass ensures proper synchronization between bias tensor access and matrix multiplication operations on C310 architecture. The pass:
Finds
LocalTensorV3Opin C2 position with abufIdattributeExtracts the
bufIdvalue from this tensorFor each
MmadOpwithcmatrixSourceparameter:Inserts
get_buf pipe_m, <bufId>before theMmadOpInserts
rls_buf pipe_m, <bufId>after theMmadOp
Important: This pass assumes there is only one bias tensor with bufId per function. The UnifyBiasTensor
pass consolidates all C2 tensors into a single tensor, ensuring only one bufId exists.
-ascendc-insert-bufid-sync
Insert get_buf/rls_buf synchronization around operations for BufId-based tracking on C310 arch
Inserts BufId synchronization for Ascend hardware with C310 architecture using buffer ID tracking:
Assigns unique
bufIdto each tensor allocation (LocalTensorV3Op,TBufGetTensorOp)For each operation using a tensor, inserts
get_bufbefore andrls_bufafterTracks pipeline type per operation: PIPE_MTE1 (L1 load), PIPE_MTE2 (GM load), PIPE_MTE3 (GM store), PIPE_FIX (fixpipe), PIPE_S (scalar), PIPE_M (matmul), PIPE_V (vector)
For
scf.forloops (exceptvec_scope_loopor parallel), inserts MTE3-MTE2/MTE3-MTE1 sync at loop end
Sets bufId array attribute on operations for downstream FuseBufIdSync optimization.
-ascendc-insert-cross-core-sync
Insert CrossCoreSetFlag and CrossCoreWaitFlag synchronization between if_aic and if_aiv operations
This pass inserts ascendc.cross_core_set_flag and ascendc.cross_core_wait_flag operations
to synchronize data transfers between AIV (AI Vector) and AIC (AI Cube) cores. It is a no-op
unless the function contains at least two if_aic/if_aiv core groups.
The pass identifies sync-trigger operations that transfer data between cores through an on-chip local tensor:
DataCopyOpwith direction VECCALC -> A1/B1 (UB -> L1, in AIV groups)FixpipeOpwith direction CO1 -> VECCALC (L0C -> UB, in AIC groups)
For each sync-trigger op, the pass:
Finds consumers of the trigger’s destination tensor on the opposite core (only those appearing after the trigger in IR order, via function walk). If no consumers are found, no synchronization flags are inserted for this trigger.
Allocates a flag ID (sequential, starting from 0 and wrapping at 16).
Inserts a backward
CrossCoreWaitFlagbefore the trigger if this is a second trigger for the same tensor (tensor reuse) or if inside a loop. This waits for the previous consumer’s backwardCrossCoreSetFlag(loop-carried dependency).Inserts a producer
CrossCoreSetFlagafter the trigger (forward direction), using the trigger op’s pipe, only when consumers exist. Signals “data is ready” to consumers.Inserts consumer flags only in the first consumer group after the trigger (cross-core flags are binary semaphores – one set, one wait):
A
CrossCoreWaitFlagat the start of the consumer group’s body: the consumer waits for the producer’s “data ready” signal from step 4.A
CrossCoreSetFlagbefore the consumer group’s terminator (backward direction): the consumer signals that the buffer can be safely overwritten. Both use the consumer op’s pipe (determined bygetOpPipeExt(consumerOp)). Consumer groups after the first receive no flags – they rely on intra-core ordering, since the data is already loaded by the first group’s wait.
Seeds backward synchronization to prevent the producer from overwriting a shared buffer before the consumer finishes reading. This happens for a first trigger inside a loop, or for a second trigger whose enclosing loop differs from the one recorded previously:
An initial
CrossCoreSetFlagbefore the loop (consumer’s core, PIPE_S) – seeds the backward flag to prevent first-iteration deadlock.A finalization
CrossCoreWaitFlagafter the loop (producer’s core, producer’s pipe) – consumes the leftover flag.
When the same destination tensor is reused by a different trigger op (second trigger),
the pass recognizes it via a destination-to-flag map and reuses the same flag ID:
the second trigger’s CrossCoreSetFlag re-sets the forward flag (when it has
consumers), allowing subsequent consumers to wait on it. This avoids exhausting
the 16-flag hardware limit.
-ascendc-insert-cross-core-sync-gm
Insert cross-core synchronization for CO1->GM and VECCALC->GM data transfers
This pass inserts ascendc.cross_core_set_flag and ascendc.cross_core_wait_flag
operations to synchronize data transfers from on-chip buffers to Global Memory (GM)
across AIC (producer) and AIV (consumer) cores. It handles two transfer types:
FixpipeOp from L0C (CO1) to GM, and DataCopyOp from UB (VECCALC) to GM.
GM transfers are synchronized independently of the sync-trigger mechanism handled by
InsertCrossCoreSync and of loop-carried forward/backward sync. Walking the
if_aic/if_aiv groups in pre-order, the pass:
Producer (the group containing the GM-writing op): inserts
CrossCoreSetFlagoperations at the end of the group (before its terminator), on the op’s pipe (determined bygetOpPipeExt), and records the root GlobalTensor – traced back throughglobal_tensor.subindexchains viagetRootGlobalTensor:CO1 -> GM (
FixpipeOp): inserts two flags with idsidandid + 16.VECCALC -> GM (
DataCopyOp): inserts a single flag with idid. Flags are only emitted when a subsequent consumer group exists (a later group whose operand traces to the same root); otherwise theCrossCoreSetFlagand the flag ID allocation are skipped, since noCrossCoreWaitFlagwould ever consume the flag.
Consumer (a later group whose operand traces, via
getRootGlobalTensor, to a recorded root): insertsCrossCoreWaitFlagoperations at the start of the group body (before any group operation), onPIPE_S:For CO1 -> GM producers: waits on
idonly.For VECCALC -> GM producers: waits on both
idandid + 16.
The flag pair asymmetry between transfer types reflects the different pipeline dependencies of each path. No loop-carried synchronization is inserted for the GM path.
Flag ID allocation starts from the value stored in the ascendc.cross_core_flag_id
attribute on the function (set by the preceding InsertCrossCoreSync pass, or 0
if absent) and wraps at 16 (maxTensorId); the id + 16 values occupy the
upper flag bank (16-31). After processing, the pass removes the
ascendc.cross_core_flag_id attribute, as flag ID allocation is complete.
-ascendc-insert-init-dump
Insert InitDump call and dump_addr kernel argument for kernels containing debug operations
Inserts AscendC debug-dump initialization for kernels that use debug operations. Walks the function for
ascendc.printf or ascendc.dump_tensor; if any are found, appends a dump_addr kernel argument
(__gm__ uint8_t*, emitasc.kernel_arg = DumpAddr) and inserts an ascendc.init_dump operation at
the function entry. The operation is parameterized by the kernel isMixed flag (derived from the
asc.kernel_type module attribute), the dump address, and a fixed dump size of 1 MB (1024*1024 bytes)
per core.
-ascendc-insert-subblock-guard
Insert early return guard for non-zero sub-block indices in mixed kernel functions
Inserts a guard at the beginning of kernel functions that returns early if the current sub-block index is not zero. This is useful for multi-core kernels where only sub-block 0 should execute the main kernel logic.
The pass only operates on functions within modules that have asc.kernel_type = "mixed" attribute.
For other kernel types (vector, cube) or modules without the attribute, the pass is a no-op.
The pass inserts the following C++ code at the function entry:
if (AscendC::GetSubBlockIdx() != 0) return;
This ensures that only sub-block 0 executes the kernel body, while other sub-blocks return immediately. The pass is idempotent and skips functions that already have the guard inserted.
-ascendc-lower-to-l0
Lower L2 vector operations to L0 intrinsics with default repeat parameters
Converts L2-level vector operations to L0-level hardware intrinsics:
Unary L2→L0:
abs_l2→abs_l0,exp_l2→exp_l0, etc. with mask, repeat_times=0, empty repeat_paramsBinary L2→L0:
add_l2→add_l0,mul_l2→mul_l0, etc. with similar parametersSpecial ops:
duplicate_l2→duplicate_l0,cast_l2→cast_l0,compare_l2→compare_l0Scalar ops:
adds_l2→adds_l0,muls_l2→muls_l0
L0 operations provide direct hardware control but require explicit parameters; this pass provides defaults.
L2 operations are simpler API with automatic parameter calculation (done by FillAscOperands).
-ascendc-promote-cv-block
Detect kernel type and promote conditional AIC/AIV blocks for pure cube or vector kernels
Inspects each function for if_aic (AI Cube) and if_aiv (AI Vector) conditional blocks and sets the
module-level asc.kernel_type attribute to cube, vector, or mixed accordingly. For pure cube or
vector kernels, unwraps the conditional blocks by moving their body operations into the parent block and
erasing the conditional wrapper. Mixed kernels retain their conditional structure since operations target
different compute units.
-ascendc-refine-cube-position
Refine A1 cube tensor positions to B1 or C1 based on L0 consumer destinations
Refines the memory position of LocalTensorAutoOp tensors initially placed in A1 (L1) by analyzing how
they are consumed by L0 load and copy operations (LoadDataL0V2Op, DataCopyL0Op).
The pass inspects the destination tensor position of each consumer and remaps the source tensor position:
Destination in A2 (L0A) -> source stays in A1
Destination in B2 (L0B) -> source refined to B1
Destination in C2 (bias) -> source refined to C1
Only tensors with position A1 are processed. This refinement ensures cube tensors are placed in the correct hardware buffer (L1, B1, or C1) matching the actual data flow to L0A, L0B, or bias memory.
-ascendc-remove-debug-ops
Remove debug operations (PrintfOp, DumpTensorOp, AssertOp) when debug mode is disabled
Removes ascendc.printf, asctile.dump_tensor, and asctile.assert operations from the module.
This pass should run immediately after codegen to remove debug ops when the user has not explicitly
enabled debug mode.
-ascendc-reuse-tensor-allocation
Reuse freed on-chip allocations based on unroll_iter lifetime analysis
Reuses freed on-chip buffer allocations by inserting LocalTensorReinterpretCastOp.
Covers VECCALC (UB), A1 (L1), A2 (L0A), B2 (L0B), and CO1 (L0C) positions.
Reuse is restricted to tensors of the same position, since each position maps to
a distinct hardware buffer.
Reuse decisions are based on asctile.unroll_iter attribute: if top tensor’s max
unroll_iter is less than bottom tensor’s min unroll_iter, their lifetimes don’t
overlap and reuse is valid.
Input/output tensors cannot be reused across different unroll_iters because each iteration loads/stores different data. Reusing would cause data races. Input/output tensors can only be reused with tensors that have the same unroll_iter or no unroll_iter at all.
For in/out tensors: all users share the same unroll_iter (guaranteed by upstream passes), so max == min. Two in/out tensors at different iterations can reuse each other. Two in/out tensors at the same iteration overlap and cannot reuse.
For tensors without unroll_iter: fall back to position-based ordering within their scope.
For cube positions (A1, A2, B2), allocation size is computed with cube-block alignment to match the downstream AllocateTensor pass.
When reusing: larger tensor wraps smaller via reinterpret_cast, or smaller moved before larger and cast.
-ascendc-reuse-ub-allocation
Reuse freed UB allocations via reinterpret_cast to reduce memory consumption
Analyzes tensor lifetimes and reuses freed UB allocations by inserting LocalTensorReinterpretCastOp:
For each block, collects
LocalTensorAutoOptensors and their last usersSorts tensors by last user position (earlier-freed tensors can be reused first)
Checks reusability: same block, non-overlapping lifetimes, static shape, VECCALC position
When reusing: larger tensor wraps smaller via reinterpret_cast, or smaller moved before larger and cast
Special handling for control flow: tensors used across different regions (if/else, while before/after) can be reused.
Input/output tensors in same loop are not reused unless reuse-in-out=true.
Example: two tensors with non-overlapping lifetimes reused:
// Before: separate allocations for sequential tensors
%0 = ascendc.local_tensor_auto veccalc() : <333xi32>
ascendc.data_copy_l2 %0, %gm, %c1
%1 = ascendc.local_tensor_auto veccalc() : <333xi32>
ascendc.data_copy_l2 %1, %gm, %c1
// After: single allocation with reinterpret_cast
%0 = ascendc.local_tensor_auto veccalc() : <333xi32>
%1 = ascendc.reinterpret_cast %0 : !ascendc.local_tensor<333xi32> to !ascendc.local_tensor<333xi32>
ascendc.data_copy_l2 %0, %gm, %c1
ascendc.data_copy_l2 %1, %gm, %c1
Example: tensors in different loops can be reused:
// Before: output tensor in loop 1, input tensor in loop 2
scf.for { %output = ascendc.local_tensor_auto veccalc() output }
scf.for { %input = ascendc.local_tensor_auto veccalc() input }
// After: output reused as input across different loops
%output = ascendc.local_tensor_auto veccalc() output
%input = ascendc.reinterpret_cast %output
scf.for { use %output }
scf.for { use %input }
Options
-reuse-in-out : Allow reuse between input/output tensors even in same loop
-ascendc-unify-bias-tensor
Unify C2 (bias) tensor with identical type and size, setting addr to 0
Consolidates all LocalTensorV3Op operations in C2 position (bias tensor memory) into a single tensor.
When multiple C2 tensors exist (typically created during loop unrolling or tiling), this pass:
Keeps the first encountered C2 tensor
Replaces all uses of subsequent C2 tensors with the first tensor
Erases the duplicate tensor operations
Sets the address of the unified tensor to 0
This optimization ensures a single bias tensor allocation per function, which is required for
correct bufId-based synchronization in InsertBiasBufIdSync pass.