asc.experimental.asctile.cv_strategy
- asc.experimental.asctile.cv_strategy(split: asctile_SplitMode) CVStrategyContext
[Experimental] Declare a split strategy for vector sub-blocks.
This context manager is intended for kernels launched with
cv_ratio=2. It marks a 2D copy fromL0CtoUBand its dependent operations for splitting:SplitByMhalves the first (M) axis, whileSplitByNhalves the second (N) axis. Internally, each vector sub-block would receive and process one half of the tensor.Supported vector scenarios:
reduction operations (e.g.
reduce_sum(),reduce_max(), and other),reshape operations (
reshape(),expand_dims(),ravel(),squeeze()),broadcast operations (
broadcast_to(),broadcast_tensors()).
Supported data transfer operations:
copy_out()(UB-to-GM),copy()(L0C-to-UB, UB-to-L1). Other operations are not eligible for splitting.- Parameters:
split – The axis along which to split tensor work. Must be
SplitMode.SplitByMorSplitMode.SplitByN.- Raises:
TypeError – If
splitis not aSplitModeValueError – If
splitdoes not request splitting by an axis
Note
For every
copy()operation with L0C-to-UB context inside, the copied tensor must have rank 2. ForSplitByM, its M axis must be a multiple of 2. ForSplitByN, its N axis must be a multiple of 32. The resulting tensors of all dependent operations must satisfy these requirements as well. In addition, these operations are not allowed to change the dimension along which the tensor is split.Examples
Split a matmul result by M, reduce each sub-block’s partial result, and broadcast it to the bigger shape:
addend = asctile.copy_in(input_tensor, offsets=[0, 0], shape=[32, 64]) with asctile.cv_strategy(asctile.SplitMode.SplitByM): # actual shape | internally split shape matmul = asctile.copy(a @ b, location="UB") # [32, 64] | [16, 64] elwise = (matmul + addend) * 3 # [32, 64] | [16, 64] reduce = elwise.sum(1) # [32] | [16] rshape = reduce.expand_dims(1) # [32, 1] | [16, 1] brcast = rshape.broadcast_to(32, 128) # [32, 128] | [16, 128] asctile.copy_out(brcast, output_tensor, [0, 0]) # each sub-block copies [16, 128] half of [32, 128]
Split a matmul result by N and use it both inside and outside of the
cv_strategycontext:with asctile.cv_strategy(asctile.SplitMode.SplitByN): c_l0 = asctile.matmul(a, b) c_ub = c_l0.to("UB") * 3.0 asctile.copy_out(c_ub, output_tensor, [0, 0]) c_l1 = asctile.copy(c_ub, location=asctile.TensorLocation.L1) # similar to 'c_ub.to("L1")'