AscendC passes

-ascendc-declare-py-struct

Insert emitasc.declare_py_struct operations for all Python struct types used in the module

Collects all emitasc.py_struct types referenced in function arguments, block arguments, and operation results, then inserts emitasc.declare_py_struct operations at the module beginning to declare them for C++ code emission.

The pass deduplicates struct types while preserving the order of first occurrence. Nested struct types (structs containing other structs) are also declared.

-ascendc-define-cube-only

Define ASCENDC_CUBE_ONLY macro for cube-only (matrix multiplication) kernels

Inserts #define ASCENDC_CUBE_ONLY verbatim at module start and sets matmul_cube_only attribute. This macro enables cube-only optimizations in the Ascend C runtime for kernels that only perform matrix multiplication without vector operations.

-ascendc-detect-enable-debug

Detect debug utility usage (printf, dump_tensor) and set enable_debug attribute

Scans the module for PrintfOp or DumpTensorOp operations. If any are found, sets the enable_debug unit attribute on the module to enable debug runtime support during kernel execution.

-ascendc-detect-kernel-type

Classify kernel as vector, cube, or mixed based on operation types present

Analyzes operations in the module to determine kernel type and sets kernel_type string attribute:

  • “vector”: Only vector operations present (no MmadOp or RegistMatmulObjOp)

  • “cube”: Only matrix multiplication operations present (no VectorOp)

  • “mixed”: Both vector and cube operations present

This classification affects synchronization strategy and runtime configuration.

-ascendc-erase-sync

Remove TQueBind synchronization operations and replace deque_tensor with allocated tensors

Removes intra-core synchronization infrastructure for static allocation mode:

  • Erases TQueBindEnqueTensorOp, SetFlagOp, WaitFlagOp, PipeBarrierOp

  • Replaces TQueBindDequeTensorOp with the tensor from corresponding TQueBindAllocTensorOp

This pass is used when tensors are statically allocated and queue-based synchronization is not needed.

-ascendc-generate-boilerplate

Insert required C++ include headers based on operations used in the kernel

Adds necessary Ascend C header includes at module start:

  • Always includes kernel_operator.h (core runtime)

  • If ListTensorDescOp or ListTensorDescV2Op present: includes kernel_operator_list_tensor_intf.h

  • If RegistMatmulObjOp present: includes lib/matmul_intf.h

  • If TensorDescOp present: includes kernel_operator_list_tensor_intf.h

-ascendc-hoist-que-bind

Move queue and buffer initialization operations to the function entry block

Hoists TQueBind-related initialization ops from loops to function root:

  • QueueOp, QueBindOp, TBufOp declarations

  • TPipeInitQueueOp, TPipeInitBufferOp initialization

  • TBufGetTensorOp tensor retrieval

This optimization ensures queue/buffer setup happens once at kernel entry rather than repeatedly in loops.

-ascendc-hoist-tensor-allocation

Move LocalTensorAutoOp allocations from loops to the function entry block

Hoists LocalTensorAutoOp tensor allocations to the function root to reduce repeated allocations in loops. When exclude-in-out=true: only hoists tensors without input or output flags (pure temporary tensors). When exclude-in-out=false: hoists all tensors regardless of input/output status.

Input/output tensors are those connected to GM via DataCopyOp with GlobalToLocal/LocalToGlobal direction.

Options

-exclude-in-out : Keep input/output tensors inside loops, only hoist temporary tensors

-ascendc-input-output-tensor

Mark tensors as input/output based on GM data copy direction and handle input-only to output copies

Analyzes DataCopyOp to set input/output flags on LocalTensorAutoOp:

  • gm_ubuf direction → tensor is input (loaded from GM)

  • ubuf_gm direction → tensor is output (stored to GM)

Also handles special case: if an input-only tensor is used for ubuf_gm copy, creates a separate output tensor and inserts DataCopyL2Op to copy input tensor to output tensor before the GM store.

Example transformation:

// Before: unmarked tensors
%0 = ascendc.local_tensor_auto vecin() : <64xf32>
ascendc.data_copy_l2 %0, %gm_in, %c64  // gm→ub (input)
%1 = ascendc.local_tensor_auto vecout() : <64xf32>
ascendc.data_copy_l2 %gm_out, %1, %c64  // ub→gm (output)

// After: tensors marked with input/output flags
%0 = ascendc.local_tensor_auto vecin() input : <64xf32>
%1 = ascendc.local_tensor_auto vecout() output : <64xf32>

-ascendc-insert-que-sync

Insert TQueBind enqueue/dequeue synchronization and scalar get/set_value barriers

Inserts intra-core synchronization for TQueBind-managed tensors and scalar operations:

  1. Enqueue tensors: After each OpWithDst, inserts TQueBindEnqueTensorOp or PipeBarrierOp(pipe_v)

  2. Dequeue tensors: For enqueued tensors, inserts TQueBindDequeTensorOp before first user

  3. Get/Set value sync: Wraps get_value/set_value with V_S/S_V event sync (fetch_event_id, set_flag, wait_flag)

  4. Loop handling: Special sync placement for set_value inside single-op loops

  5. Canonicalization: Adds PipeBarrierOp(pipe_all) at function end and runs canonicalization

TQueBind synchronization enables double-buffering and overlapping compute with data transfer.

Example: enqueue/dequeue synchronization:

// Before: tensor used directly after allocation
%tensor = ascendc.que_bind.alloc_tensor %queue
ascendc.add_l2 %dst, %tensor, %tensor, %count

// After: tensor dequeued before use, enqueued after compute
%tensor = ascendc.que_bind.alloc_tensor %queue
%dequeued = ascendc.que_bind.deque_tensor %queue
ascendc.add_l2 %dst, %dequeued, %dequeued, %count
ascendc.que_bind.enque_tensor %queue, %dst

Example: scalar get/set_value with V_S/S_V sync:

// Before: direct scalar access
%val = ascendc.local_tensor.get_value %tensor, %offset

// After: wrapped with event sync
%event = ascendc.pipe.fetch_event_id %pipe, v_s
ascendc.set_flag v_s, %event
ascendc.wait_flag v_s, %event
%val = ascendc.local_tensor.get_value %tensor, %offset
%event2 = ascendc.pipe.fetch_event_id %pipe, s_v
ascendc.set_flag s_v, %event2
ascendc.wait_flag s_v, %event2

-ascendc-legalize-kernel-args

Attach kernel argument attributes and insert ffts_addr handling for multi-core kernels

Processes kernel entry functions (marked with ascendc.global attribute):

  1. Marks all existing arguments as Explicit kernel arguments

  2. If setFftsAddr=true: adds ffts_addr memref argument and calls SetFftsBaseAddrOp

  3. If matmul present and not cube-only: inserts AscendIsAICOp check with FftsCrossCoreSyncOp for AIC mode

Kernel argument attributes (emitasc.kernel_arg) control how arguments are handled by the runtime.

Options

-set-ffts-addr : Append ffts_addr kernel argument and call set_ffts_base_addr for cross-core sync

-ascendc-materialize-tensor

Convert LocalTensorAutoOp placeholders to concrete TBuf or TQueBind allocations

Materializes tensor allocation placeholders into concrete Ascend C buffer objects:

  • TQueBind path (for input/output tensors when alwaysBuf=false): Creates QueueOp + TPipeInitQueueOp + TQueBindAllocTensorOp + TQueBindFreeTensorOp at function end

  • TBuf path (for temporary tensors or when alwaysBuf=true): Creates TBufOp + TPipeInitBufferOp + TBufGetTensorOp

Input tensors use VECIN queue position, output tensors use VECOUT. Temporary tensors use VECCALC TBuf. Buffer size computed from tensor shape (static or dynamic via shape operands).

Example transformation (TQueBind path for input tensor):

// Before: placeholder
%0 = ascendc.local_tensor_auto vecin() input : <64xf32>

// After: queue-based allocation
%pipe = ascendc.pipe
%queue = ascendc.queue : <vecin, 1>
ascendc.pipe.init_queue %pipe, %queue, %c1, %c256
%0 = ascendc.que_bind.alloc_tensor %queue : !ascendc.queue<vecin, 1>, !ascendc.local_tensor<64xf32>
// ... at function end:
ascendc.que_bind.free_tensor %queue, %0

Example transformation (TBuf path for temporary tensor):

// Before: placeholder
%0 = ascendc.local_tensor_auto veccalc() : <64xf32>

// After: buffer-based allocation
%pipe = ascendc.pipe
%tbuf = ascendc.tbuf : <veccalc>
ascendc.pipe.init_buffer %pipe, %tbuf, %c256
%0 = ascendc.tbuf.get_tensor %tbuf : !ascendc.tbuf<veccalc>, !ascendc.local_tensor<64xf32>

Options

-always-buf : Use TBuf for all tensors; otherwise use TQueBind for input/output tensors

-ascendc-noop

Placeholder pass that performs no transformations

A no-operation pass useful for testing, pipeline debugging, or as a placeholder in pass schedules. The pass walks the function but performs no modifications to the IR.

-ascendc-privatize-func

Mark non-kernel functions as private and kernel functions as public for emission

Adjusts function visibility based on kernel status:

  • Functions without ascendc.global attribute → set as private (internal helper functions)

  • Functions with ascendc.global and body → set as public (kernel entry points)

  • Functions with ascendc.global but no body (declarations) → remain as-is

This ensures only kernel entry points are exported while helper functions remain internal.

-ascendc-unify-pipe

Replace multiple PipeOp instances with a single unified pipe at function entry

Consolidates multiple PipeOp declarations into a single pipe object at function entry. All uses of individual pipes are replaced with the unified pipe, and original PipeOp instances are erased.

This ensures consistent pipe management across the kernel and reduces redundant TPipe object creation.

-ascendc-verify-sync

Validate TQueBind synchronization correctness and emit warnings for mismatches

Verifies proper pairing of TQueBind operations and emits warnings for incorrect usage:

  • alloc_tensor without corresponding free_tensor

  • free_tensor for already-freed tensor or without alloc_tensor

  • enque_tensor without corresponding deque_tensor

  • deque_tensor without enque_tensor in queue

  • Unexpected tensor uses between enque_tensor and deque_tensor

This is a verification-only pass that helps detect synchronization bugs during development.