AscendC passes
-ascendc-declare-py-struct
Insert emitasc.declare_py_struct operations for all Python struct types used in the module
Collects all emitasc.py_struct types referenced in function arguments, block arguments, and operation results,
then inserts emitasc.declare_py_struct operations at the module beginning to declare them for C++ code emission.
The pass deduplicates struct types while preserving the order of first occurrence. Nested struct types (structs containing other structs) are also declared.
-ascendc-define-cube-only
Define ASCENDC_CUBE_ONLY macro for cube-only (matrix multiplication) kernels
Inserts #define ASCENDC_CUBE_ONLY verbatim at module start and sets matmul_cube_only attribute.
This macro enables cube-only optimizations in the Ascend C runtime for kernels that only perform
matrix multiplication without vector operations.
-ascendc-detect-enable-debug
Detect debug utility usage (printf, dump_tensor) and set enable_debug attribute
Scans the module for PrintfOp or DumpTensorOp operations. If any are found, sets the enable_debug
unit attribute on the module to enable debug runtime support during kernel execution.
-ascendc-detect-kernel-type
Classify kernel as vector, cube, or mixed based on operation types present
Analyzes operations in the module to determine kernel type and sets kernel_type string attribute:
“vector”: Only vector operations present (no
MmadOporRegistMatmulObjOp)“cube”: Only matrix multiplication operations present (no
VectorOp)“mixed”: Both vector and cube operations present
This classification affects synchronization strategy and runtime configuration.
-ascendc-erase-sync
Remove TQueBind synchronization operations and replace deque_tensor with allocated tensors
Removes intra-core synchronization infrastructure for static allocation mode:
Erases
TQueBindEnqueTensorOp,SetFlagOp,WaitFlagOp,PipeBarrierOpReplaces
TQueBindDequeTensorOpwith the tensor from correspondingTQueBindAllocTensorOp
This pass is used when tensors are statically allocated and queue-based synchronization is not needed.
-ascendc-generate-boilerplate
Insert required C++ include headers based on operations used in the kernel
Adds necessary Ascend C header includes at module start:
Always includes
kernel_operator.h(core runtime)If
ListTensorDescOporListTensorDescV2Oppresent: includeskernel_operator_list_tensor_intf.hIf
RegistMatmulObjOppresent: includeslib/matmul_intf.hIf
TensorDescOppresent: includeskernel_operator_list_tensor_intf.h
-ascendc-hoist-que-bind
Move queue and buffer initialization operations to the function entry block
Hoists TQueBind-related initialization ops from loops to function root:
QueueOp,QueBindOp,TBufOpdeclarationsTPipeInitQueueOp,TPipeInitBufferOpinitializationTBufGetTensorOptensor retrieval
This optimization ensures queue/buffer setup happens once at kernel entry rather than repeatedly in loops.
-ascendc-hoist-tensor-allocation
Move LocalTensorAutoOp allocations from loops to the function entry block
Hoists LocalTensorAutoOp tensor allocations to the function root to reduce repeated allocations in loops.
When exclude-in-out=true: only hoists tensors without input or output flags (pure temporary tensors).
When exclude-in-out=false: hoists all tensors regardless of input/output status.
Input/output tensors are those connected to GM via DataCopyOp with GlobalToLocal/LocalToGlobal direction.
Options
-exclude-in-out : Keep input/output tensors inside loops, only hoist temporary tensors
-ascendc-input-output-tensor
Mark tensors as input/output based on GM data copy direction and handle input-only to output copies
Analyzes DataCopyOp to set input/output flags on LocalTensorAutoOp:
gm_ubufdirection → tensor is input (loaded from GM)ubuf_gmdirection → tensor is output (stored to GM)
Also handles special case: if an input-only tensor is used for ubuf_gm copy, creates a separate output tensor
and inserts DataCopyL2Op to copy input tensor to output tensor before the GM store.
Example transformation:
// Before: unmarked tensors
%0 = ascendc.local_tensor_auto vecin() : <64xf32>
ascendc.data_copy_l2 %0, %gm_in, %c64 // gm→ub (input)
%1 = ascendc.local_tensor_auto vecout() : <64xf32>
ascendc.data_copy_l2 %gm_out, %1, %c64 // ub→gm (output)
// After: tensors marked with input/output flags
%0 = ascendc.local_tensor_auto vecin() input : <64xf32>
%1 = ascendc.local_tensor_auto vecout() output : <64xf32>
-ascendc-insert-que-sync
Insert TQueBind enqueue/dequeue synchronization and scalar get/set_value barriers
Inserts intra-core synchronization for TQueBind-managed tensors and scalar operations:
Enqueue tensors: After each
OpWithDst, insertsTQueBindEnqueTensorOporPipeBarrierOp(pipe_v)Dequeue tensors: For enqueued tensors, inserts
TQueBindDequeTensorOpbefore first userGet/Set value sync: Wraps
get_value/set_valuewith V_S/S_V event sync (fetch_event_id, set_flag, wait_flag)Loop handling: Special sync placement for
set_valueinside single-op loopsCanonicalization: Adds
PipeBarrierOp(pipe_all)at function end and runs canonicalization
TQueBind synchronization enables double-buffering and overlapping compute with data transfer.
Example: enqueue/dequeue synchronization:
// Before: tensor used directly after allocation
%tensor = ascendc.que_bind.alloc_tensor %queue
ascendc.add_l2 %dst, %tensor, %tensor, %count
// After: tensor dequeued before use, enqueued after compute
%tensor = ascendc.que_bind.alloc_tensor %queue
%dequeued = ascendc.que_bind.deque_tensor %queue
ascendc.add_l2 %dst, %dequeued, %dequeued, %count
ascendc.que_bind.enque_tensor %queue, %dst
Example: scalar get/set_value with V_S/S_V sync:
// Before: direct scalar access
%val = ascendc.local_tensor.get_value %tensor, %offset
// After: wrapped with event sync
%event = ascendc.pipe.fetch_event_id %pipe, v_s
ascendc.set_flag v_s, %event
ascendc.wait_flag v_s, %event
%val = ascendc.local_tensor.get_value %tensor, %offset
%event2 = ascendc.pipe.fetch_event_id %pipe, s_v
ascendc.set_flag s_v, %event2
ascendc.wait_flag s_v, %event2
-ascendc-legalize-kernel-args
Attach kernel argument attributes and insert ffts_addr handling for multi-core kernels
Processes kernel entry functions (marked with ascendc.global attribute):
Marks all existing arguments as
Explicitkernel argumentsIf
setFftsAddr=true: addsffts_addrmemref argument and callsSetFftsBaseAddrOpIf matmul present and not cube-only: inserts
AscendIsAICOpcheck withFftsCrossCoreSyncOpfor AIC mode
Kernel argument attributes (emitasc.kernel_arg) control how arguments are handled by the runtime.
Options
-set-ffts-addr : Append ffts_addr kernel argument and call set_ffts_base_addr for cross-core sync
-ascendc-materialize-tensor
Convert LocalTensorAutoOp placeholders to concrete TBuf or TQueBind allocations
Materializes tensor allocation placeholders into concrete Ascend C buffer objects:
TQueBind path (for input/output tensors when
alwaysBuf=false): CreatesQueueOp+TPipeInitQueueOp+TQueBindAllocTensorOp+TQueBindFreeTensorOpat function endTBuf path (for temporary tensors or when
alwaysBuf=true): CreatesTBufOp+TPipeInitBufferOp+TBufGetTensorOp
Input tensors use VECIN queue position, output tensors use VECOUT. Temporary tensors use VECCALC TBuf.
Buffer size computed from tensor shape (static or dynamic via shape operands).
Example transformation (TQueBind path for input tensor):
// Before: placeholder
%0 = ascendc.local_tensor_auto vecin() input : <64xf32>
// After: queue-based allocation
%pipe = ascendc.pipe
%queue = ascendc.queue : <vecin, 1>
ascendc.pipe.init_queue %pipe, %queue, %c1, %c256
%0 = ascendc.que_bind.alloc_tensor %queue : !ascendc.queue<vecin, 1>, !ascendc.local_tensor<64xf32>
// ... at function end:
ascendc.que_bind.free_tensor %queue, %0
Example transformation (TBuf path for temporary tensor):
// Before: placeholder
%0 = ascendc.local_tensor_auto veccalc() : <64xf32>
// After: buffer-based allocation
%pipe = ascendc.pipe
%tbuf = ascendc.tbuf : <veccalc>
ascendc.pipe.init_buffer %pipe, %tbuf, %c256
%0 = ascendc.tbuf.get_tensor %tbuf : !ascendc.tbuf<veccalc>, !ascendc.local_tensor<64xf32>
Options
-always-buf : Use TBuf for all tensors; otherwise use TQueBind for input/output tensors
-ascendc-noop
Placeholder pass that performs no transformations
A no-operation pass useful for testing, pipeline debugging, or as a placeholder in pass schedules. The pass walks the function but performs no modifications to the IR.
-ascendc-privatize-func
Mark non-kernel functions as private and kernel functions as public for emission
Adjusts function visibility based on kernel status:
Functions without
ascendc.globalattribute → set asprivate(internal helper functions)Functions with
ascendc.globaland body → set aspublic(kernel entry points)Functions with
ascendc.globalbut no body (declarations) → remain as-is
This ensures only kernel entry points are exported while helper functions remain internal.
-ascendc-unify-pipe
Replace multiple PipeOp instances with a single unified pipe at function entry
Consolidates multiple PipeOp declarations into a single pipe object at function entry.
All uses of individual pipes are replaced with the unified pipe, and original PipeOp instances are erased.
This ensures consistent pipe management across the kernel and reduces redundant TPipe object creation.
-ascendc-verify-sync
Validate TQueBind synchronization correctness and emit warnings for mismatches
Verifies proper pairing of TQueBind operations and emits warnings for incorrect usage:
alloc_tensorwithout correspondingfree_tensorfree_tensorfor already-freed tensor or withoutalloc_tensorenque_tensorwithout correspondingdeque_tensordeque_tensorwithoutenque_tensorin queueUnexpected tensor uses between
enque_tensoranddeque_tensor
This is a verification-only pass that helps detect synchronization bugs during development.