Repository navigation
docs(kernels): add kernel framework docs and the LLVM interop design - #102
Conversation
scratchv/backend/kernels/ (new):
- ARCHITECTURE.md layered design, the four axes, and which interfaces are
implemented vs reserved vs explicitly out of scope
- DEVELOPMENT.md newcomer tutorial: build the three fp32 problems of
contest 2 from scratch, then the O0 -> O2 path
- USAGE.md platform contract, and how to add problems/dtypes/targets
- OPTIMIZATION.md cost breakdown, measured matmul data, ranked entry points
docs/reference/ARCHITECTURE_DESIGN.md (new):
- per-layer swappability with LLVM, the four things LLVM structurally cannot
do, and the one irreversible decision (shape as IR type information) with
its two consequences
- annotates, but does not overwrite, the contradictions in the existing
ARCHITECTURE.md and CLAUDE.md
Verified while writing DEVELOPMENT.md: the O0 fp32 kernels pass 10/10 data
points on the platform evaluator (add 9.0 instr/elem, reducesum 5.0,
matmul 8.17 instr/MAC, all correct against the -O0 reference).
Co-Authored-By: Claude <noreply@anthropic.com>
🤖 AI Code Review
📁
|
Implements the framework described in DEVELOPMENT.md ch.2-3: a dtype/ISA agnostic kernel generator whose bodies contain no instruction mnemonics. target.py TargetDesc + TARGETS (rv32im with int regs, rv32imf + fp regs) dtypes.py DtypePolicy + POLICIES (q16, f32), mac_instrs, zero_acc loopgen.py prologue / epilogue / unrolled_body bodies/ add, reducesum, matmul + the problem registry pipeline.py problem name -> complete .s __main__.py CLI (--problem / -o / --list) Verified on the platform evaluator, 6 problems x 10 data points, all pass: contest 2 (fp32) add 30/30 reducesum 40/40 matmul 30/30 contest 1 (q16) add 30/30 reducesum 40/40 matmul 30/30 Two defects were found by that verification and are fixed here: 1. The zero-accumulator helper took a hardcoded scratch register (t6), which in matmul holds the row stride N*4. Zeroing it silently produced wrong results on all 10 data points with no assembler error. The scratch register is now supplied by the call site. 2. DtypePolicy modelled a multiply as a single mnemonic, which is the wrong granularity: one multiply-accumulate is fmul.s+fadd.s for f32 but mul+srai 16+add for q16, because Q16.16 rescales each product before accumulating. Only matmul was affected (0/10); modelling the whole MAC as a `mac` instruction template fixes it and keeps q16 a data-only change. DEVELOPMENT.md is updated to match, including both of these as worked examples of the failure modes they illustrate. Co-Authored-By: Claude <noreply@anthropic.com>
Adds an `unroll` parameter to the add and reducesum bodies: an unrolled main loop with immediate offsets, plus a remainder tail loop so the kernel stays correct for any N (the platform does not disclose the per-point sizes). Measured on the platform evaluator (total cost over the 10 data points): unroll 1 4 8 16 32 64 128 add-fp32 181787 121113 113615 110091 108779 109023 110796 reducesum 91989 53959 50240 48478 47822 47944 48854 Both curves turn at unroll=32: past that, the extra instruction-cache lines cost more than the loop overhead saved. d_miss is constant across all unroll values, so the whole difference is instruction count vs i-cache pressure. The per-point data is why a single unroll value is only a compromise: at N=64 u=8 wins (635 vs 701) because u=32's 128-instruction body only runs twice and never amortises its fetch cost, while from N=256 up u=32 wins. The tail loop never executes for the contest data points (every N is a multiple of 64), so the platform's 10/10 says nothing about it. It was verified separately against N = 3, 7, 70, 103, 4095, 4097, 5000 using the platform's own wrapper/compile/run path. DEVELOPMENT.md gains the measured tables, the per-point analysis, and a new section 5.3 on testing code paths the contest never reaches. Co-Authored-By: Claude <noreply@anthropic.com>
Adds a (mr x nr) register-blocked matmul, selected by dtype through a plan
table. Data registers come from the bank named by DtypePolicy.bank (f32 uses
the fp bank, q16 the int bank) rather than from a hardcoded count, so the
capacity check is a real constraint instead of a comment:
(4,4) needs 16 acc + 4 A + 4 B + 1 product = 25 data registers
f32 has 32 in the fp bank -> fits
q16 has 23 in the int bank -> does not fit, stays on the generic loop
Measured on the platform evaluator (10 data points):
total cost N=64 instr N=64 instr/MAC
generic triple loop 4,129,318 2,142,559 8.17
(4,4) blocked 1,524,768 782,007 2.98
2.71x faster, and d_miss is unchanged at N=64 (773 both ways), so the whole
gain is instruction count. Section 6.4 previously predicted a cache cliff
here; measurement says there is none in this version, so the cache-blocking
step is skipped rather than implemented.
A third defect was found by verification and is fixed: the mac template wrote
the product into an operand register (fmul {t1}, {t1}, {t2}). That is
invisible in the generic loop, where t1 is reloaded every k, but fatal in a
blocked kernel where an A value must survive nr consecutive MACs. It showed
up as column 0 correct and columns 1-3 wrong on all 10 data points. The
template now takes a separate product register.
Also verified: the generic fallback still handles N not divisible by 4
(N = 1,2,3,5,6,7,10,63,65 all correct), and all six problems still pass.
Co-Authored-By: Claude <noreply@anthropic.com>
Auditing the tutorial against the code that actually shipped turned up four
classes of gap.
A. The tutorial taught things the repo does not do:
- section 2.3 showed a DtypePolicy without `prod`/`bank`, and with the
mac template that writes the product into an operand (the bug fixed in
the previous commit) - a reader following it would build the broken one
- section 3.3 showed a single `build()`; the file is now `_generic` +
`_blocked` + a dispatching `build()`
- section 5.1's `unrolled_body` hardcoded ft0/ft1 instead of taking the
registers from the policy
- appendix B was missing the top-level __init__.py and under-described
three other files
B. Implemented but undocumented:
- section 6.3 described the 2.71x register-blocking win without showing a
single line of it. Now contains the register-allocation table, the loop
skeleton, the full `_blocked()`, the dispatcher, and the BLOCKING table.
C. Described as if actionable but never built:
- sections 5.2 and 6.2 (immediate offsets, whole-range coverage) now carry
an explicit "not implemented in this repo" marker, and 6.2 spells out what
would have to be added. This also explains why UNROLL can only hold one
global value instead of one per N.
D. A note at the top now says the document is built up in stages, that
intermediate code is not final code, and that the repo wins on any
disagreement - plus that this is a "how to write kernels" tutorial, not a
"how to place well" one.
Also fixed: a duplicated paragraph in 6.4, and a pseudo-code block that was
fenced as python and did not compile.
All 17 python blocks in the document now compile.
Co-Authored-By: Claude <noreply@anthropic.com>
Chapter 4 covered reading the numbers but not the two things a newcomer actually needs when stuck. 4.4 "how to read an unfamiliar .s" - four steps: find the skeleton (the same four parts every kernel has), find the hot loop (the innermost backward branch), count instructions per element/MAC, and trace each register's lifetime. The fourth step is the one that finds bugs: the t6 clobber from 3.3 is invisible if you read the offending line alone, since the line itself is correct and the mistake is dozens of lines earlier. Also covers noting which paths the contest never executes, since those are the unverified ones. 4.5 "where to start when it is slow" - a decision tree whose first cut is always the miss share of cost. Below 20% it is instruction-bound (check against the per-problem floor first); above it, compare d_miss against the compulsory count, because being at compulsory means there is nothing to fix. Three measured cases land on the three different branches. Includes how to estimate both the instruction floor and the compulsory miss count. Chapter 4 is retitled from "measurement" to "diagnosis" accordingly, and its checkpoint grows from 2 questions to 5 with answers. Numbers in the new material are the measured ones (reducesum 2.07 instr/element, add N=4096 d_miss 773 vs compulsory 768). Co-Authored-By: Claude <noreply@anthropic.com>
The development scope was the three fp32 problems of contest 2, with everything else left as interfaces. Adding a contest-3 body went past that, and the demo contest itself is closed. Code removed: - bodies/fwht.py deleted - bodies/__init__.py back to the six contest-1/2 entries - dtypes.py: the `sub` field added for fwht is reverted (it was unused outside fwht; if a future problem needs subtraction, add it back then) Docs edited (all pre-dating this change, hence this commit rather than a revert of files): - DEVELOPMENT.md: chapter 7 loses its fwht and winograd/spmm sections; what remains is the same-family case, plus an explicit warning that no body in the repo implements a different loop shape, so there is nothing to copy from if a new problem needs one - USAGE.md: two lines no longer name contest-3 problems, keeping the platform facts they were illustrating - ARCHITECTURE.md: the planned bodies/fwht.py entry and the phasing row that named winograd/spmm are gone; the phasing row keeps the general capability No contest-3 identifier remains anywhere in the repository. Co-Authored-By: Claude <noreply@anthropic.com>
Both were written before the code existed and had drifted. USAGE.md is rewritten as a reference rather than a second tutorial. It now says the code is landed (the old header still claimed nothing was built), points to DEVELOPMENT.md for the walkthrough, and its field tables match the real TargetDesc and DtypePolicy. Fixed while in there: - the "how to add" step pointed at pipeline.py for the problem registry; the registry is bodies/__init__.py, and the tuple is (body, dtype, target) - the TargetDesc and DtypePolicy snippets showed fields that do not exist (banks=, alu=, accumulate=, reduce_rules=) and omitted ones that do (mac, prod, bank, comparison) - referenced CoveragePlan, which does not exist - the CLI section was missing --unroll - the troubleshooting table lacked the two failure modes that actually bit us: a register clobbered where it should not be, and a register reused where an operand had to survive - eval_local.py does not pass a contest, which is why contest-2 problems reported "not open for evaluation" ARCHITECTURE.md section 8 listed seven files marked as planned-but-this-phase that were never created (plan.py, cost.py, coverage.py, kir.py, isel.py, layout.py, _scaffold.py) and claimed each pass would implement CompilerPass. Replaced with what exists, where the unimplemented concepts ended up, and an explicit note that the pass_manager integration was never done. Section 9 now carries real progress: S0/S0.5/S0.6 done, S2 measured and skipped, S1/S3/S4 open. Section 4's CacheBlocking row no longer calls itself the biggest win, since the measurement says there is no miss cliff to fix here. All cross-references between the four documents were verified to resolve. Co-Authored-By: Claude <noreply@anthropic.com>
…oints The tutorial taught "sweep the unroll factor and take the lowest total cost", but never said what that total was over. It was over the platform's ten internal sizes -- which contestants cannot see, which the platform deliberately stopped publishing, and which have since been replaced entirely. Any conclusion derived that way is both unavailable to a contestant and invalidated by the next data point change. Section 5.1 now requires two things to be written down before sweeping: the sampling of the published range, and the objective. All measurements in the section were redone on a declared sample (range:64:4096:8) rather than the platform's sizes, and the tables say so. The re-measured numbers changed the answer: u total cost worst-case deviation 4 131,279 112.7% 8 123,079 105.5% 16 119,159 102.2% <- most robust 32 117,559 110.4% 64 117,479 128.7% <- lowest total reducesum agrees (u=16 most robust at 102.0%). The two objectives pick different values, and the total-cost optimum itself moved when the sample changed -- which is the point: "lowest total" is not a stable criterion when you do not know the sizes. Also re-based: section 3.6 and 4.2 examples, and section 6.3 gains a per-point table showing that register blocking only engages when N % 4 == 0. On the declared sample just 2 of 8 points take the fast path; the rest fall back to the generic loop at 1.00x. That constraint was written when the sizes were all multiples of four, and is much more costly now. The pipeline.py tables now record value + range + sample + objective together, so a future reader can tell when a value has gone stale. The measuring tool used throughout is probe.py (in the contest workspace, outside this repo): it measures an arbitrary N given on the command line, unlike measure.py / eval_local.py which read the platform's sizes. Co-Authored-By: Claude <noreply@anthropic.com>
scratchv/backend/kernels/ (new):
implemented vs reserved vs explicitly out of scope
contest 2 from scratch, then the O0 -> O2 path
docs/reference/ARCHITECTURE_DESIGN.md (new):
Verified while writing DEVELOPMENT.md: the O0 fp32 kernels pass 10/10 data points on the platform evaluator (add 9.0 instr/elem, reducesum 5.0, matmul 8.17 instr/MAC, all correct against the -O0 reference).