Vector Addition

Your first AscTile kernel: an element-wise vector addition out = x + y that runs on the Ascend NPU.

In this tutorial you will learn about:

  • The @asctile.jit decorator and how a Python function becomes an NPU kernel.

  • The SPMD programming model: many AI cores run the same kernel, each identified by asctile.block_idx() out of asctile.block_num().

  • Moving data between global memory (GM) and local (on-chip) memory.

  • Tiling a long tensor into per-core chunks and per-tile blocks, and overlapping loads with compute (multi-buffering).

  • How to launch a kernel from the host and verify its result against PyTorch.

Compute Kernel

An AscTile kernel is an ordinary function marked with asctile.jit. This decorator captures the function AST and lowers it to Ascend IR (MLIR) at runtime, then compiles it to an Ascend C binary with the Bisheng compiler. You never write Ascend C yourself – you write Python code (similar to NumPy or Triton), and the toolchain does the rest.

Kernel arguments follow a simple type contract:

  • Tensor arguments are annotated asctile.GlobalAddress. At launch time you pass a CPU torch.Tensor (or numpy array); the runtime hands the kernel a pointer into the device buffer.

  • Scalar arguments annotated with a plain Python type (e.g. size: int) become runtime values – they can change from call to call without recompiling.

  • Scalar arguments annotated asctile.ConstExpr[int] are compile-time constants: their value is baked into the generated IR, so each distinct value produces a separately compiled (and cached) kernel. Tile sizes must be ConstExpr because the on-chip buffer they describe must have a static, compile-time-known shape.

50 from asc.experimental import asctile
51
52
53 @asctile.jit
54 def vector_add(x_ptr: asctile.GlobalAddress,  # pointer to input tensor x
55                y_ptr: asctile.GlobalAddress,  # pointer to input tensor y
56                out_ptr: asctile.GlobalAddress,  # pointer to output tensor out
57                size: int,  # number of elements (runtime value, may change between launches)
58                tile_size: asctile.ConstExpr[int],  # elements per tile (compile-time constant)
59                ):
60
61     # A ``GlobalTensor`` is a descriptor for an array living in global memory (GM). It carries a pointer and a (possibly
62     # dynamic) shape -- here ``[size]`` -- but owns no storage of its own; the storage is the buffer you passed at
63     # launch time.
64     x_gm = asctile.global_tensor(x_ptr, [size])
65     y_gm = asctile.global_tensor(y_ptr, [size])
66     out_gm = asctile.global_tensor(out_ptr, [size])
67
68     # AscTile is SPMD: every launched core runs this same function. We partition the work so that each core owns a
69     # *contiguous slice* of the tensor. First split ``size`` evenly across all cores, then divide each core's share into
70     # ``tiles_per_core`` tiles with ``tile_size`` elements each:
71     #
72     #   core 0                   | core 1                   | ... |  core N-1
73     #   [tile0][tile1]...[tileM] | [tile0][tile1]...[tileM] |     | [tile0][tile1]...[tileM]
74     tiles_per_core = asctile.ceildiv(asctile.ceildiv(size, asctile.block_num()), tile_size)
75     core_length = tile_size * tiles_per_core
76     core_offset = asctile.block_idx() * core_length
77
78     # ``asctile.range`` is the tiled loop construct. ``unroll_factor=2`` asks the compiler to software-pipeline the loop
79     # (multi-buffering): while i-th iteration computes ``x + y`` in UB, the load for i+1-th iteration is already started
80     # from GM. This hides the memory latency behind the compute.
81     for i in asctile.range(tiles_per_core, unroll_factor=2):
82         offset = core_offset + i * tile_size
83         # ``copy_in`` moves a tile from GM into local (on-chip) memory. For vector arithmetic the destination is the
84         # Unified Buffer (UB); other local memories (L1, L0A, ...) serve the cube unit (see the matmul tutorial). It
85         # returns a new ``LocalTensor``.
86         x = asctile.copy_in(x_gm, [offset], [tile_size])
87         y = asctile.copy_in(y_gm, [offset], [tile_size])
88         # Element-wise add runs in UB. Operator overloads -- like + - * / -- call the underlying ``asctile.add``, etc.
89         # Each operation returns a *new* tensor (so the language satisfies SSA semantics internally).
90         out = x + y
91         # ``copy_out`` moves the result tile from local memory (UB) back to GM at the same offset.
92         asctile.copy_out(out, out_gm, [offset])

Launch and Verify

To run a kernel we first select a target with asctile.set_platform(). Backend.Model is the cycle-accurate Ascend simulator (no hardware needed); NPU runs on a real device. The platform Ascend950PR_9599 is a C310-family chip used throughout AscTile tutorials.

Launch syntax: kernel[block_num](args...). The bracketed value is the number of cores; it sets asctile.block_num() seen by the kernel. ConstExpr parameters (here tile_size) are passed as plain Python literals.

107 if __name__ == "__main__":
108     import torch
109
110     asctile.set_platform(asctile.Backend.Model, asctile.Platform.Ascend950PR_9599)
111
112     torch.manual_seed(0)
113     # size is chosen so that 16 cores * 4 tiles * 128 elements == 8192 divide it evenly to keep it aligned on purpose.
114     size = 8192
115     block_num = 16  # number of AI cores to launch -- becomes asctile.block_num() inside
116     tile_size = 128  # elements per loop iteration (a ConstExpr at compile time)
117     dtype = torch.float32
118
119     x = torch.randn(size, dtype=dtype)
120     y = torch.randn(size, dtype=dtype)
121     out = torch.empty_like(x)
122     vector_add[block_num](x, y, out, size, tile_size)
123
124     # Verify against the reference computed by PyTorch on the CPU.
125     reference = x + y
126     torch.testing.assert_close(out, reference, atol=1e-5, rtol=1e-5)
127     max_diff = (out - reference).abs().max().item()
128     print(f"vector_add: PASSED (size={size}, blocks={block_num}, tile={tile_size}, max diff={max_diff:.2e})")

Gallery generated by Sphinx-Gallery