You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 149ae47
Browse filesBrowse the repository at this point in the historyBrowse files
perf(kernels): edge blocks so matmul blocking covers any N; shared unroll skeleton (#103)
Two changes, one optimization and one refactor.
1. Edge blocks (bodies/matmul._blocked)
The blocked path used to fall back to the generic triple loop whenever N was
not a multiple of the block size. That was invisible while the data points
were dense ladders, but once they became arbitrary integers it cost most of
the benefit: on a declared sample (range:4:64:8) six of eight points took the
fallback and scored 1.00x against it.
Now the blocked path is three segments, with mainI = N rounded down to mr and
mainJ likewise:
main blocks [0, mainI) x [0, mainJ) full mr x nr
right band [0, mainI) x [mainJ, N) one mr x 1 block per column
bottom band [mainI, N) x [0, N) scalar triple loop
Measured on that sample:
N before after
4 520 536 0.97x (was already blocked)
13 20151 9101 2.21x
21 80460 33434 2.41x
30 228741 98059 2.33x
38 459271 190905 2.41x
47 861499 365786 2.36x
55 1373241 572307 2.40x
64 793767 793768 1.00x (was already blocked)
total 3817650 -> 2063896 1.85x
Against the -O0 reference the affected points went from about 3.4x to
5.4-9.0x, i.e. into the same band as the points that were already blocked.
A defect was found and fixed while doing this: all three segments test at the
bottom, so when a band is empty (mainJ == N, or mainI == N) the body ran once
anyway and wrote past the end. N = 4, 52 and 64 segfaulted until each band
got an explicit skip. Same class of mistake as the remainder tail loop in
5.1: the "remainder is zero" path has to be handled explicitly.
2. Shared unroll skeleton (loopgen.unrolled_loop)
add and reducesum each carried their own copy of the unroll-plus-remainder-tail
scaffolding (power-of-two check, round count, unrolled-segment end address,
tail loop, pointer advance). It is now one function parameterised by the three
things that actually differed: the per-round pointer list, the single-element
body, and the exit label.
Verified equivalent rather than assumed: normalised output is identical for
all 18 generated artifacts (6 problems x unroll 1/8/32), the token sequences
match, and clang's .text for the changed pair is byte-identical. The only
source-level difference was whitespace -- the original used one space after
the store mnemonic in the unrolled body and two in the tail, and the refactor
unifies them.
Also: the CLI now defaults -o to ./<problem>.s instead of printing to stdout,
which is what scattered 190 scratch .s files through /tmp; `-o -` still prints.
Docs updated with both findings (DEVELOPMENT.md 6.3, OPTIMIZATION.md 5 and a
new 5.1 on why "fall back the whole problem" is the most expensive kind of
conservative, and pipeline.py's table comments).
All six problems still pass 10/10 on the platform evaluator.
Co-authored-by: Claude <noreply@anthropic.com>
0 commit comments