asc.experimental.asctile.copy

asc.experimental.asctile.copy(src: LocalTensor, offsets: Iterable[PlainValue | int] | None = None, shape: Iterable[int] | None = None, location: asctile_TensorLocation | Literal['BT', 'FIX', 'L0A', 'L0B', 'L0C', 'L1', 'UB'] | None = None, split: asctile_SplitMode | None = None) → LocalTensor

Copy a local tensor to a new local tensor, optionally reshaping and relocating.

Rationale: Unlike frameworks with simpler memory hierarchies (e.g., CUDA’s global/shared/registers), Ascend NPUs expose multiple local memory levels (L1, L0A, L0B, L0C, UB) where local-to-local transfers are common. copy_in and copy_out have clear directional semantics when one endpoint is global memory (“copy in = to local”, “copy out = to global”), but this breaks down for local-to-local transfers: the same L0C→L1 operation is a “copy out” from L0C’s perspective yet a “copy in” from L1’s. Local copy eliminates this ambiguity by providing a direction-agnostic operation that clearly expresses intent regardless of which memory level you’re reasoning from.

Parameters:
  • src – The source tensor to copy.

  • offsets – The offsets into the source tensor for each dimension. Default is zeros.

  • shape – The shape of the resulting tensor. If None, uses the source tensor’s shape. Must contain static values (e.g., ConstExpr or compile-time constants).

  • location – The memory location for the destination tensor. Default is src.location. Supported location transfers: L1 to L0A, L1 to L0B, L1 to BT, L0C to L1, L0C to UB, UB to L1.

  • split – The split strategy when copying from L0C to UB. If None, no split is applied. Only supported when src.location is L0C and location is UB (can be resolved automatically). FullVec0 and FullVec1 copy the full source tensor to vector sub-block 0 or 1, respectively. SplitByM and SplitByN split a 2D source tensor in half along the M (first) or N (second) axis; each vector sub-block receives one half. The source tensor must have rank 2 and the given shape must be omitted or equal src.shape. For SplitByM, the M dimension must be a multiple of 2. For SplitByN, the N dimension must be a multiple of 32. The resulting tensor shape would be already reduced by half along the selected axis.

Returns:

A new tensor that is a copy of the source tensor

Return type:

LocalTensor

Raises:
  • TypeError – If src is not a LocalTensor, split is not a SplitMode, or location is not a TensorLocation-like

  • RuntimeError – If shape is invalid, data alignment check fails, offsets rank mismatch, or split constraints are violated

Examples

Copy a tensor with the same shape:

src = asctile.copy_in(x_gm, [0], [128])
result = asctile.copy(src)

Copy a sub-tensor from a larger tensor with explicit shape and offsets:

src = asctile.copy_in(x_gm, [0, 0], [64, 64])
result = asctile.copy(src, [16, 16], [32, 32])

Copy a tensor to a different memory location (e.g., L0A for matrix multiplication):

a_l1 = asctile.copy_in(a_gm, [0, 0], [64, 128], asctile.TensorLocation.L1)
a_l0a = asctile.copy(a_l1, [0, 0], [64, 32], asctile.TensorLocation.L0A)
b_l0b = asctile.copy(b_l1, [0, 0], [32, 64], asctile.TensorLocation.L0B)

Copy accumulator result from L0C to L1:

acc = asctile.zeros_acc([64, 64], dtype=asctile.float32)
asctile.matmul_acc(acc, a_l0a, b_l0b)
result_l1 = asctile.copy(acc, location=asctile.TensorLocation.L1)

Copy matmul result from L0C to UB for further processing:

result = asctile.matmul(a_l0a, b_l0b)
result_ub = asctile.copy(result, location="UB")

Copy half of each row of a matmul result to each sub-block when cv_ratio=2. Each sub-block receives shape [64, 32] and can copy its half to global memory using e.g. [0, 32 * asctile.sub_block_idx()] offsets:

result = asctile.matmul(a_l0a, b_l0b)  # shape [64, 64], located in L0C
result_ub = asctile.copy(result, location="UB", split=asctile.SplitMode.SplitByN)

Alternatively, the to method can be used to transform the tensor location:

ub_tensor = asctile.copy_in(x_gm, [0], [128], asctile.TensorLocation.UB)
l1_tensor = ub_tensor.to(asctile.TensorLocation.L1)