Note
Go to the end to download the full example code.
Vector Addition
Your first AscTile kernel: an element-wise vector addition out = x + y that runs on the Ascend NPU.
In this tutorial you will learn about:
The
@asctile.jitdecorator and how a Python function becomes an NPU kernel.The SPMD programming model: many AI cores run the same kernel, each identified by
asctile.block_idx()out ofasctile.block_num().Moving data between global memory (GM) and local (on-chip) memory.
Tiling a long tensor into per-core chunks and per-tile blocks, and overlapping loads with compute (multi-buffering).
How to launch a kernel from the host and verify its result against PyTorch.
Compute Kernel
An AscTile kernel is an ordinary function marked with asctile.jit. This decorator captures the function AST
and lowers it to Ascend IR (MLIR) at runtime, then compiles it to an Ascend C binary with the Bisheng compiler. You
never write Ascend C yourself – you write Python code (similar to NumPy or Triton), and the toolchain does the rest.
Kernel arguments follow a simple type contract:
Tensor arguments are annotated
asctile.GlobalAddress. At launch time you pass a CPUtorch.Tensor(or numpy array); the runtime hands the kernel a pointer into the device buffer.Scalar arguments annotated with a plain Python type (e.g.
size: int) become runtime values – they can change from call to call without recompiling.Scalar arguments annotated
asctile.ConstExpr[int]are compile-time constants: their value is baked into the generated IR, so each distinct value produces a separately compiled (and cached) kernel. Tile sizes must beConstExprbecause the on-chip buffer they describe must have a static, compile-time-known shape.
50 from asc.experimental import asctile
51
52
53 @asctile.jit
54 def vector_add(x_ptr: asctile.GlobalAddress, # pointer to input tensor x
55 y_ptr: asctile.GlobalAddress, # pointer to input tensor y
56 out_ptr: asctile.GlobalAddress, # pointer to output tensor out
57 size: int, # number of elements (runtime value, may change between launches)
58 tile_size: asctile.ConstExpr[int], # elements per tile (compile-time constant)
59 ):
60
61 # A ``GlobalTensor`` is a descriptor for an array living in global memory (GM). It carries a pointer and a (possibly
62 # dynamic) shape -- here ``[size]`` -- but owns no storage of its own; the storage is the buffer you passed at
63 # launch time.
64 x_gm = asctile.global_tensor(x_ptr, [size])
65 y_gm = asctile.global_tensor(y_ptr, [size])
66 out_gm = asctile.global_tensor(out_ptr, [size])
67
68 # AscTile is SPMD: every launched core runs this same function. We partition the work so that each core owns a
69 # *contiguous slice* of the tensor. First split ``size`` evenly across all cores, then divide each core's share into
70 # ``tiles_per_core`` tiles with ``tile_size`` elements each:
71 #
72 # core 0 | core 1 | ... | core N-1
73 # [tile0][tile1]...[tileM] | [tile0][tile1]...[tileM] | | [tile0][tile1]...[tileM]
74 tiles_per_core = asctile.ceildiv(asctile.ceildiv(size, asctile.block_num()), tile_size)
75 core_length = tile_size * tiles_per_core
76 core_offset = asctile.block_idx() * core_length
77
78 # ``asctile.range`` is the tiled loop construct. ``unroll_factor=2`` asks the compiler to software-pipeline the loop
79 # (multi-buffering): while i-th iteration computes ``x + y`` in UB, the load for i+1-th iteration is already started
80 # from GM. This hides the memory latency behind the compute.
81 for i in asctile.range(tiles_per_core, unroll_factor=2):
82 offset = core_offset + i * tile_size
83 # ``copy_in`` moves a tile from GM into local (on-chip) memory. For vector arithmetic the destination is the
84 # Unified Buffer (UB); other local memories (L1, L0A, ...) serve the cube unit (see the matmul tutorial). It
85 # returns a new ``LocalTensor``.
86 x = asctile.copy_in(x_gm, [offset], [tile_size])
87 y = asctile.copy_in(y_gm, [offset], [tile_size])
88 # Element-wise add runs in UB. Operator overloads -- like + - * / -- call the underlying ``asctile.add``, etc.
89 # Each operation returns a *new* tensor (so the language satisfies SSA semantics internally).
90 out = x + y
91 # ``copy_out`` moves the result tile from local memory (UB) back to GM at the same offset.
92 asctile.copy_out(out, out_gm, [offset])
Launch and Verify
To run a kernel we first select a target with asctile.set_platform().
Backend.Model is the cycle-accurate Ascend simulator (no hardware needed); NPU runs on a real device.
The platform Ascend950PR_9599 is a C310-family chip used throughout AscTile tutorials.
Launch syntax: kernel[block_num](args...).
The bracketed value is the number of cores; it sets asctile.block_num() seen by the kernel.
ConstExpr parameters (here tile_size) are passed as plain Python literals.
107 if __name__ == "__main__":
108 import torch
109
110 asctile.set_platform(asctile.Backend.Model, asctile.Platform.Ascend950PR_9599)
111
112 torch.manual_seed(0)
113 # size is chosen so that 16 cores * 4 tiles * 128 elements == 8192 divide it evenly to keep it aligned on purpose.
114 size = 8192
115 block_num = 16 # number of AI cores to launch -- becomes asctile.block_num() inside
116 tile_size = 128 # elements per loop iteration (a ConstExpr at compile time)
117 dtype = torch.float32
118
119 x = torch.randn(size, dtype=dtype)
120 y = torch.randn(size, dtype=dtype)
121 out = torch.empty_like(x)
122 vector_add[block_num](x, y, out, size, tile_size)
123
124 # Verify against the reference computed by PyTorch on the CPU.
125 reference = x + y
126 torch.testing.assert_close(out, reference, atol=1e-5, rtol=1e-5)
127 max_diff = (out - reference).abs().max().item()
128 print(f"vector_add: PASSED (size={size}, blocks={block_num}, tile={tile_size}, max diff={max_diff:.2e})")