Latest Results
RISC-V: use unit-stride loads and stores in the ZVL256B ROTM kernel
The dflag dispatch selects one of three loop bodies in rotm_rvv.c, one for
each form of the modified Givens transform. All three are reached only
after the guard
if (!(incx == incy && incx > 0)) goto L70;
so throughout the dflag-dispatched part of the routine the two increments
are equal and positive. When that common increment is 1 the per-element
distance is sizeof(FLOAT) and the operands are contiguous, but the loops
still issued the strided forms vlse/vsse. Unit-stride vle/vse are
substantially cheaper on the same data, and the sibling kernel
rot_vector.c already branches on inc_x == 1 && inc_y == 1 for exactly this
reason.
Give each of the three bodies a unit-stride loop, entered when its
increment variable is 1, and keep the strided loop as the fallback for any
other increment.
The change does not alter which elements are visited or the order of the
floating-point operations, so the results are unchanged. An exhaustive
differential run against the unmodified kernel -- n over a set spanning 1
to 257, incx and incy from {1,2,3,4,5,7,8,16,-1,-2,-3}, dflag from
{1,0,-1,-2} -- matches
bit for bit in both single and double precision (31944 cases each, with no
out-of-bounds writes), including the cases where the increments differ and
the strided fallback is taken.
The drotm and srotm benchmarks fix param[0] = 1.0, so they only ever
measure the dflag > 0 body. For that body they report, on SpaceMiT
K3/X100 at 2.2 GHz:
drotm n=256 +56.29% n=512 +65.57% n=1024 +69.86%
srotm n=256 +180.06% n=512 +222.42% n=1024 +239.85%
The dflag == 0 and dflag < 0 bodies receive the same substitution over the
same access pattern.
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
Co-authored-by: Ning Tian <tianning24@iscas.ac.cn>
Signed-off-by: jiakai xu <xujiakai2025@iscas.ac.cn>6eanut:riscv64-rotm-unit-stride Latest Branches
0%
0%
0%
hugomeiland:riscv64-zvl1024b © 2026 CodSpeed Technology