Skip to content

docs(kernels): add kernel framework docs and the LLVM interop design - #102

Merged
jizhenjun merged 9 commits into
mainfrom
dev/scratch
Oct 7, 2026
Merged

jizhenjun merged 9 commits into
mainfrom
dev/scratch

Conversation

@jizhenjun

Copy link
Copy Markdown
Collaborator

scratchv/backend/kernels/ (new):

  • ARCHITECTURE.md layered design, the four axes, and which interfaces are
    implemented vs reserved vs explicitly out of scope
  • DEVELOPMENT.md newcomer tutorial: build the three fp32 problems of
    contest 2 from scratch, then the O0 -> O2 path
  • USAGE.md platform contract, and how to add problems/dtypes/targets
  • OPTIMIZATION.md cost breakdown, measured matmul data, ranked entry points

docs/reference/ARCHITECTURE_DESIGN.md (new):

  • per-layer swappability with LLVM, the four things LLVM structurally cannot do, and the one irreversible decision (shape as IR type information) with its two consequences
  • annotates, but does not overwrite, the contradictions in the existing ARCHITECTURE.md and CLAUDE.md

Verified while writing DEVELOPMENT.md: the O0 fp32 kernels pass 10/10 data points on the platform evaluator (add 9.0 instr/elem, reducesum 5.0, matmul 8.17 instr/MAC, all correct against the -O0 reference).

scratchv/backend/kernels/ (new):
- ARCHITECTURE.md   layered design, the four axes, and which interfaces are
                    implemented vs reserved vs explicitly out of scope
- DEVELOPMENT.md    newcomer tutorial: build the three fp32 problems of
                    contest 2 from scratch, then the O0 -> O2 path
- USAGE.md          platform contract, and how to add problems/dtypes/targets
- OPTIMIZATION.md   cost breakdown, measured matmul data, ranked entry points

docs/reference/ARCHITECTURE_DESIGN.md (new):
- per-layer swappability with LLVM, the four things LLVM structurally cannot
  do, and the one irreversible decision (shape as IR type information) with
  its two consequences
- annotates, but does not overwrite, the contradictions in the existing
  ARCHITECTURE.md and CLAUDE.md

Verified while writing DEVELOPMENT.md: the O0 fp32 kernels pass 10/10 data
points on the platform evaluator (add 9.0 instr/elem, reducesum 5.0,
matmul 8.17 instr/MAC, all correct against the -O0 reference).

Co-Authored-By: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

🤖 AI Code Review

共审查 10 个变更文件
⚠️ 另有 3 个文件超过上限(最多 10 个)未审查

📁 docs/reference/ARCHITECTURE_DESIGN.md

🔴 矛盾未收敛:附录 C 标注了 7 处现有文档的冲突,但没指定谁修、什么时候修。 §4.2 的性能表、CLAUDE.md 的 4.2x 结论、三处 triple 不一致——这些冲突会在下次有人读旧文档时继续被当成事实引用。建议加一行:「附录 C 的每条由 X 在 P0 结束后一周内重测并更新对应文件。」

🟡 §7 给了可编译的 dataclass 定义,但 §6 说「只预留接口,不实现」。 TargetDesc 和 CostModelKey 已经带 @dataclass(frozen=True) 和字段约束,读起来像实施规范。如果 P0 不落地这些类型,建议降级为伪代码并标注「参考形态,不要求本次提交」,避免读者以为必须立刻写。

🟡 §3.1 的「A 的做法 + C 的存储方式」在澄清段落前需要一次前置判断。 现在读者要先接受一个看似矛盾的推荐,再花 10 行理解为什么它不矛盾。建议把那张「两个独立决定」的分解表提到推荐结论之前,让读者先看到「这是两个维度」再看到「A+C 是选 ①=A, ②=C」。

🟡 §5.3 的「内测比赛2 加 rv32imf 要改四处」缺少第五处:llvm_codegen.py 产出的 .ll 头部 target triple。 §5.1 已经指出 triple 写的是 riscv64 而实际跑 riscv32,但 §5.3 列改四处时没包含它——说明这条「静默错误」在 §5.1 和 §5.3 之间还没完全串起来。

🟡 §6 内核调用钩子的默认返回 None 判据写「所有现有产物逐字节不变」,但没定义怎么验。 §6.1 的「改完之后逐字节不变」是同一句话。建议补一条:验法 = git stash && make test && git stash pop && make test && diff 之类的具体步骤,否则「逐字节不变」在 CI 里无法自动检查。

💭 §1 表格「能换吗」列最后一行写「★ 见 §2」,但 §2 讲的是「LLVM 做不到的四件事」。 表格标题是「能换吗」,填「★ 见 §2」会让读者以为 §2 在回答「能不能换」,实际 §2 在回答「为什么不能换」。建议把这一列改为「可换 / 不可换(见 §2)」或直接把最后一行归到 §2 单独说明。

💭 附录 B.4 的「两边必须完全相同」和 §4.1「转换器便宜的前提是两边表示已对齐」之间有一个隐含假设:对齐是静态的。 如果 LLVM 的 !range 语义和 ScratchV 的值域语义在某次版本升级后产生分歧,转换器会静默丢失信息。建议在 §4.1 末尾加一句「转换器上线后需定期对拍语义」。


📁 scratchv/backend/kernels/ARCHITECTURE.md

🔴 Bug: Dangling CostModel reference — §2.4 引用 CostModel 的 CoverageRisk 项来估算 expected_cost,但 §8.2 明确声明"CostModel 对象没有"。读者无法判断这条引用是死代码还是未来计划。建议在 §2.4 加 [TBD — CostModel 未实现] 标记,或把 expected_cost 从 CoveragePlan 签名中移除。

🟡 Orthogonality claim contradicts §2.3 — §2 开头声称"四条轴互相正交",紧接着 §2.3 自己解释"f32 的合法分块集合与 int32 不同"(因为寄存器银行容量来自 TargetDesc),即 ②→③ 有耦合。建议把正交性声明限定为"轴之间无数据依赖"或"修改一个轴不强制修改另一个轴的 定义",避免读者误认为正交意味着可独立调参。

🟡 "三题" vs "六题" 不一致 — §0 写"内测比赛2(fp32 三题)",S0 状态写"六题 10/10"。如果每题有正反/边界变体,建议在 §0 补一句说明计数口径;否则改为同一数字。

🟡 Predicate(f) 的表示形式未定义 — §2.1 列了 Predicate(f),但没说 f 是 callable、AST、字符串还是数值区间。这直接影响 LegalityCheck 能否静态求值。建议在 §2.1 或 §8.2 标注当前实际用的是哪种。

🟡 Runtime(header_layout) 是孤立项 — 只在 §2.1 定义、S4 提及,全文没有任何 header_layout 的 schema、读取路径或测试描述。当前 S4 是"未做",建议要么补一段最小 schema,要么加 [TBD] 标记让它不会被误当作可用功能。

🟡 引用具体行号会腐烂 — §5 的 scratchv/compiler.py:527-531、regalloc_linear.py:212-218、§6 I3 的 riscv_runner.isa_violation() 都是行号/函数名引用。行号在下一次改动上游后立刻失效。建议改为"函数/类名 + 行为描述",行号放脚注或删除。

🟡 imm_max=2047 注释不精确 — §2.3 写"12 位有符号立即数上限",但 2047 只是 I-type 的正数上限;B-type(分支)是 13 位有符号(±4096)。如果这个值只用于 load/store/addi,注释应写"I-type 有符号立即数正数上限";如果是统一上限,数值本身有误。

💭 缺少 ShapeKnowledge → CoveragePlan 的具体映射示例 — 两个核心数据结构都有定义和取值表,但没有展示一个端到端的例子(如 Range(64,4096) → [Residue(32, …), GenericVariant])。这是 §1 里"看不清规模"的痛点解法,值得一个具体例子。


📁 scratchv/backend/kernels/USAGE.md

🔴 Bug: int_regs 定义不完整 — §5.1: "不含 a0/a1/a2" 但未说明 a3–a7 是否可用。a7 在 §7 末尾被提到"调用方只碰 a0/a1/a2/a7/t0",暗示 a7 不可用,但 TargetDesc 表里看不出来。建议在 int_regs 含义里明确排除 a3–a7。

🟡 §2.2 栈溢出风险未覆盖 — sp 向下生长落入 workspace,但 workspace 只有 W = max(1024, 8·N) 字节。大 N 下内核若用大量栈帧会越界到 guard_mid,表现为"无效写"而非直观的栈溢出。排查手册里没提这条,建议加一行。

🟡 §2.6 计分公式的表述有歧义 — "领先 2 倍和领先 0.1% 是同一个分数" 是对的(都是 1.0),但紧跟的 "目标是逐点追到榜首" 容易被读成"追到就够、不必领先"。实际效果是 cost ≤ 全场最优 就满分,cost > 全场最优 就线性扣分。建议补一句:本队cost ≤ 全场最优 → 满分;本队cost > 全场最优 → points_max × (全场最优/本队cost)。

🟡 §5.3 blocked_regs_needed() 缺算例 — 说"(4,4) 需要 25 个,放不下"但没给推导。25 = ?。读者看到异常抛错后无法自行验证。建议括号内加一句推导(如 mr × nr + mr + nr = 4×4 + 4 + 4 = 28,或说明实际公式)。

🟡 §2.8 日期"2026-10-07" — 如果这是计划中的未来变更,文档标题/状态栏应标注"即将生效"或"已生效"。当前写法读者无法判断这条约束是已生效还是待生效。

💭 §7 排查表"标签重名"建议加一行例子 — 比如 .Lloop 同时被展开体和尾循环引用,这是最常见的撞标签场景。

💭 §3 grep 自查的正则可能漏项 — 当前只列了 flw|fsw|lw|sw|fadd\.s|fmul\.s,如果 bodies/ 里出现 blez/beqz/jal 等控制流助记符,说明也耦合了。建议补一个"至少能抓到访存和算术"的说明,让读者知道这是底线检查而非完备检查。


📁 scratchv/backend/kernels/__init__.py

🔴 CLI entrypoint referenced but not implemented — Docstring documents python -m scratchv.backend.kernels --list / --problem / --o usage, but this __init__.py contains no argument parsing or __main__ logic. Either there's a separate __main__.py that should exist, or this will raise SystemExit(0) silently when invoked. Consider adding an if __name__ == "__main__" block or pointing to the actual entrypoint module.

🟡 Package is a dead shell — No imports, no re-exports, no __all__. Consumers have no discoverable API from this package. If there's intended submodules (e.g., add_fp32, riscv_asmgen), add explicit imports so from scratchv.backend.kernels import ... works and IDEs can resolve symbols.

💭 # -*- coding: utf-8 -*- is unnecessary in Python 3 — UTF-8 is the default source encoding. Safe to remove for cleanliness.

💭 Docstring references 4 docs files (ARCHITECTURE.md, DEVELOPMENT.md, USAGE.md, OPTIMIZATION.md) — worth verifying these exist in-tree, otherwise the references become dead links over time.


📁 scratchv/backend/kernels/__main__.py

🔴 Bug: Unvalidated --problem — Line 32: build_program(args.problem, ...) — 若用户传入不存在的题名,异常会直接穿透到调用者。建议在此处 if args.problem not in PROBLEMS 并用 ap.error() 提示可用列表。

🟡 Suggestion: --unroll 未校验 2 的幂 — Help text 声明"2 的幂"但无运行时检查。可在 CLI 层加校验:

if args.unroll is not None and (args.unroll & (args.unroll - 1)) != 0:
    ap.error('--unroll 必须是 2 的幂')

💭 Nit: open() 缺少 error handling — 若 --output 路径不可写(权限/目录不存在),用户只看到 traceback。可在 main() 中 try/except OSError 后 ap.error(str(e))。


📁 scratchv/backend/kernels/bodies/__init__.py

🔴 Key name collision: ambiguous 'add' lookup — Line 17: PROBLEMS['add'] returns the q16/rv32im version, while 'add-fp32' returns the fp32/rv32imf version. The bare name 'add' is the "default" for the competition, but there's no naming convention making this clear. Callers reading PROBLEMS['add'] would reasonably expect the standard/fp32 variant. Suggestion: rename to 'add-q16' / 'reducesum-q16' / 'matmul-q16' for consistency, or document the default-selection rule explicitly.

🟡 No referential integrity check — Nothing verifies that the body name in each PROBLEMS tuple (e.g. 'add') exists as a key in BODIES. A typo like 'reducesum2' would only surface at runtime lookup, with a confusing KeyError. Consider adding a module-level assertion or a small helper that validates PROBLEMS against BODIES on import.

💭 BODIES is redundant given PROBLEMS is the real config — If the intent is "problem → (body, type, machine)", the BODIES dict is only an indirection layer. If every PROBLEMS entry always uses its own name as the body key (which it currently does), the first tuple element could be derived from the key name instead of duplicated. If body names will ever differ from problem names, keep as-is — but add a comment explaining the indirection.


📁 scratchv/backend/kernels/bodies/add.py

🔴 Bug: unroll=0 passes validation — Line if unroll & (unroll - 1):
0 & (0 - 1) → 0 & -1 → 0 (false), so the check is skipped. Subsequent code generates invalid assembly (sh = -1, step = 0).

Suggestion: if unroll <= 0 or unroll & (unroll - 1):


🟡 Register clobbering risk in unrolled path — t1 (B pointer) and t3 (loop end address) are live across the unrolled_body(dtype, unroll) call. If dtype.tmp1 or dtype.tmp2 aliases t1 or t3, the B pointer or loop bound is silently destroyed.

Suggestion: Add an assertion or docstring on the dtype interface:

# Invariant: dtype.tmp1, dtype.tmp2 must not be in {'t1', 't3'}

💭 Nit — The unroll==1 path decrements a2 (the N counter) while the unrolled path reuses a2 for the remainder AND as a tail-loop counter. This is fine but subtle; a comment distinguishing "a2 = original N" vs "a2 = remaining count" would help future readers.


📁 scratchv/backend/kernels/bodies/matmul.py

🔴 Register collision when dtype.bank == 'int' — _blocked allocates data regs from target.int_regs but hard-codes control regs as s0–s10, t0–t4, a3. For integer dtypes these two sets overlap. Example: with mr=4, nr=4, need=25; if int_regs starts at s0, then acc[0]=s0 clobbers a_base[0] which was set to a0 two instructions earlier. FP dtypes are fine (two separate banks), but q16 will silently corrupt pointers. Fix: subtract control regs from pool before slicing, or assert set(pool) & set(control_regs) == set().

🔴 zero_acc(dtype, 's0') hard-codes s0 but MAC uses dtype.acc — generic path line ~41. If dtype.acc != 's0', the accumulator is never zeroed. Either pass ACC to zero_acc or assert dtype.acc == 's0'.

🟡 mr != 4 validation runs after 60+ lines of assembly generation — _blocked line ~148. Move to the top of the function alongside the register-count check.

🟡 No input validation on blocking tuple — build() accepts blocking=(-1, 3) or blocking=(4, 0) silently. Add assert mr > 0 and nr > 0 before dispatching.

🟡 blocked_regs_needed doesn't subtract control regs from pool size — the check len(pool) < need overstates available capacity for int-bank dtypes (same root cause as 🔴 above).

💭 Magic number 4 (bytes per f32 element) appears 12+ times. Extract as a constant derived from the dtype, e.g. ELEM_SZ = dtype.elem_bytes, so q16 (2-byte elements) can reuse the same code without copy-paste errors.

💭 Comment on unit = max(mr, nr) says "最小公倍数" — true only because both are constrained to powers of 2. If that constraint ever loosens, the mask unit-1 becomes wrong. Add a one-line assertion or guard comment.


📁 scratchv/backend/kernels/bodies/reducesum.py

🔴 Bug: N=0 不写输出 — blez a2, .Lret 在 N≤0 时直接跳到 .Lret(由 epilogue 提供),跳过 store。空 sum 语义应为 0,但 out[0] 未被写入,调用方读到未初始化内存。
建议:blez 分支也走 zero + store,或把 store 放到 .Lret 之前 / 让 epilogue 负责写默认值。

🔴 Bug: unroll=0 越权通过校验 — unroll & (unroll-1) 对 0 结果为 0,被判定为 2 的幂;接着 bit_length()-1 == -1,srli t2, a2, -1 是非法指令。
建议:加 if unroll < 2: raise ValueError(...),或在分支入口先 assert unroll >= 2。

🟡 寄存器依赖风险 — 展开路径使用 t2、t3 做循环变量,但注释只声明 t6 安全。若 dtype.acc 或 dtype.tmp1 恰好是 t2/t3,主循环会踩掉累加器或临时寄存器。
建议:加显式断言 dtype.acc/dtype.tmp1 not in {'t2','t3','t6'},或在 dtypes 层就固定好互斥的寄存器集。

🟡 有符号/无符号语义 — blez a2, .Lret 按有符号比较,若上游把 N 当 size_t 传入,负值(如溢出或错误)会被误判为"零长度",绕过所有检查。
建议:改成 beqz a2, .Lret,或在调用约定里明确 N 是有符号 int 并在此处加断言文档。

🟡 未测试的路径 — 关键边界(N=0、N=1、N<unroll、N 恰好 unroll 倍数、N 大余数)目前看不到对应测试。尤其 N<unroll 时走 beqz t2, .Ltail 快路径,最容易漏。

💭 注释小问题 — 文档写 "不需要知道 N",实际仍依赖 N 计算 t2 = N/unroll。改成"不需要单独知道 N%unroll"更准确。

💭 尾循环可复用主循环结构 — .Ltail_loop 与 unroll==1 分支主体几乎重复,若后续再改加法/累加逻辑容易漏改。可抽成 _emit_body(tail=True/False)。


📁 scratchv/backend/kernels/dtypes.py

🟡 No runtime guard for the prod ≠ t1/t2 invariant — the docstring documents this as critical, but it's only enforced by comment. If a future caller passes prod=dtype.tmp1, the mul {p}, {t1}, {t2} line silently corrupts the source operand.

def mac_instrs(dtype, acc=None, t1=None, t2=None, prod=None):
    _t1 = t1 or dtype.tmp1
    _t2 = t2 or dtype.tmp2
    _p = prod or dtype.prod
    assert _p != _t1, f"prod {_p} must differ from t1"
    assert _p != _t2, f"prod {_p} must differ from t2"
    ...

🟡 add field is defined but never referenced in this module — if other modules consume it, fine; otherwise it's a dead attribute that invites confusion about which mnemonics the policy actually controls.

💭 or default is slightly fragile — acc or dtype.acc means an intentional empty string "" silently falls through to the default. If callers could pass "", use an explicit None check instead; if not, it's fine as-is.



⚠️ 未审查的文件

  • scratchv/backend/kernels/loopgen.py
  • scratchv/backend/kernels/pipeline.py
  • scratchv/backend/kernels/target.py

jizhenjun and others added 8 commits October 7, 2026 16:04
Implements the framework described in DEVELOPMENT.md ch.2-3: a dtype/ISA
agnostic kernel generator whose bodies contain no instruction mnemonics.

  target.py    TargetDesc + TARGETS (rv32im with int regs, rv32imf + fp regs)
  dtypes.py    DtypePolicy + POLICIES (q16, f32), mac_instrs, zero_acc
  loopgen.py   prologue / epilogue / unrolled_body
  bodies/      add, reducesum, matmul + the problem registry
  pipeline.py  problem name -> complete .s
  __main__.py  CLI (--problem / -o / --list)

Verified on the platform evaluator, 6 problems x 10 data points, all pass:

  contest 2 (fp32)  add 30/30  reducesum 40/40  matmul 30/30
  contest 1 (q16)   add 30/30  reducesum 40/40  matmul 30/30

Two defects were found by that verification and are fixed here:

1. The zero-accumulator helper took a hardcoded scratch register (t6), which
   in matmul holds the row stride N*4. Zeroing it silently produced wrong
   results on all 10 data points with no assembler error. The scratch
   register is now supplied by the call site.

2. DtypePolicy modelled a multiply as a single mnemonic, which is the wrong
   granularity: one multiply-accumulate is fmul.s+fadd.s for f32 but
   mul+srai 16+add for q16, because Q16.16 rescales each product before
   accumulating. Only matmul was affected (0/10); modelling the whole MAC
   as a `mac` instruction template fixes it and keeps q16 a data-only change.

DEVELOPMENT.md is updated to match, including both of these as worked
examples of the failure modes they illustrate.

Co-Authored-By: Claude <noreply@anthropic.com>
Adds an `unroll` parameter to the add and reducesum bodies: an unrolled main
loop with immediate offsets, plus a remainder tail loop so the kernel stays
correct for any N (the platform does not disclose the per-point sizes).

Measured on the platform evaluator (total cost over the 10 data points):

  unroll        1       4       8      16      32      64     128
  add-fp32  181787  121113  113615  110091  108779  109023  110796
  reducesum  91989   53959   50240   48478   47822   47944   48854

Both curves turn at unroll=32: past that, the extra instruction-cache lines
cost more than the loop overhead saved. d_miss is constant across all unroll
values, so the whole difference is instruction count vs i-cache pressure.

The per-point data is why a single unroll value is only a compromise: at N=64
u=8 wins (635 vs 701) because u=32's 128-instruction body only runs twice and
never amortises its fetch cost, while from N=256 up u=32 wins.

The tail loop never executes for the contest data points (every N is a
multiple of 64), so the platform's 10/10 says nothing about it. It was
verified separately against N = 3, 7, 70, 103, 4095, 4097, 5000 using the
platform's own wrapper/compile/run path.

DEVELOPMENT.md gains the measured tables, the per-point analysis, and a new
section 5.3 on testing code paths the contest never reaches.

Co-Authored-By: Claude <noreply@anthropic.com>
Adds a (mr x nr) register-blocked matmul, selected by dtype through a plan
table. Data registers come from the bank named by DtypePolicy.bank (f32 uses
the fp bank, q16 the int bank) rather than from a hardcoded count, so the
capacity check is a real constraint instead of a comment:

  (4,4) needs 16 acc + 4 A + 4 B + 1 product = 25 data registers
  f32 has 32 in the fp bank  -> fits
  q16 has 23 in the int bank -> does not fit, stays on the generic loop

Measured on the platform evaluator (10 data points):

                        total cost   N=64 instr   N=64 instr/MAC
  generic triple loop    4,129,318    2,142,559        8.17
  (4,4) blocked          1,524,768      782,007        2.98

2.71x faster, and d_miss is unchanged at N=64 (773 both ways), so the whole
gain is instruction count. Section 6.4 previously predicted a cache cliff
here; measurement says there is none in this version, so the cache-blocking
step is skipped rather than implemented.

A third defect was found by verification and is fixed: the mac template wrote
the product into an operand register (fmul {t1}, {t1}, {t2}). That is
invisible in the generic loop, where t1 is reloaded every k, but fatal in a
blocked kernel where an A value must survive nr consecutive MACs. It showed
up as column 0 correct and columns 1-3 wrong on all 10 data points. The
template now takes a separate product register.

Also verified: the generic fallback still handles N not divisible by 4
(N = 1,2,3,5,6,7,10,63,65 all correct), and all six problems still pass.

Co-Authored-By: Claude <noreply@anthropic.com>
Auditing the tutorial against the code that actually shipped turned up four
classes of gap.

A. The tutorial taught things the repo does not do:
   - section 2.3 showed a DtypePolicy without `prod`/`bank`, and with the
     mac template that writes the product into an operand (the bug fixed in
     the previous commit) - a reader following it would build the broken one
   - section 3.3 showed a single `build()`; the file is now `_generic` +
     `_blocked` + a dispatching `build()`
   - section 5.1's `unrolled_body` hardcoded ft0/ft1 instead of taking the
     registers from the policy
   - appendix B was missing the top-level __init__.py and under-described
     three other files

B. Implemented but undocumented:
   - section 6.3 described the 2.71x register-blocking win without showing a
     single line of it. Now contains the register-allocation table, the loop
     skeleton, the full `_blocked()`, the dispatcher, and the BLOCKING table.

C. Described as if actionable but never built:
   - sections 5.2 and 6.2 (immediate offsets, whole-range coverage) now carry
     an explicit "not implemented in this repo" marker, and 6.2 spells out what
     would have to be added. This also explains why UNROLL can only hold one
     global value instead of one per N.

D. A note at the top now says the document is built up in stages, that
   intermediate code is not final code, and that the repo wins on any
   disagreement - plus that this is a "how to write kernels" tutorial, not a
   "how to place well" one.

Also fixed: a duplicated paragraph in 6.4, and a pseudo-code block that was
fenced as python and did not compile.

All 17 python blocks in the document now compile.

Co-Authored-By: Claude <noreply@anthropic.com>
Chapter 4 covered reading the numbers but not the two things a newcomer
actually needs when stuck.

4.4 "how to read an unfamiliar .s" - four steps: find the skeleton (the same
four parts every kernel has), find the hot loop (the innermost backward
branch), count instructions per element/MAC, and trace each register's
lifetime. The fourth step is the one that finds bugs: the t6 clobber from
3.3 is invisible if you read the offending line alone, since the line itself
is correct and the mistake is dozens of lines earlier. Also covers noting
which paths the contest never executes, since those are the unverified ones.

4.5 "where to start when it is slow" - a decision tree whose first cut is
always the miss share of cost. Below 20% it is instruction-bound (check
against the per-problem floor first); above it, compare d_miss against the
compulsory count, because being at compulsory means there is nothing to fix.
Three measured cases land on the three different branches. Includes how to
estimate both the instruction floor and the compulsory miss count.

Chapter 4 is retitled from "measurement" to "diagnosis" accordingly, and its
checkpoint grows from 2 questions to 5 with answers.

Numbers in the new material are the measured ones (reducesum 2.07
instr/element, add N=4096 d_miss 773 vs compulsory 768).

Co-Authored-By: Claude <noreply@anthropic.com>
The development scope was the three fp32 problems of contest 2, with
everything else left as interfaces. Adding a contest-3 body went past that,
and the demo contest itself is closed.

Code removed:
- bodies/fwht.py deleted
- bodies/__init__.py back to the six contest-1/2 entries
- dtypes.py: the `sub` field added for fwht is reverted (it was unused
  outside fwht; if a future problem needs subtraction, add it back then)

Docs edited (all pre-dating this change, hence this commit rather than a
revert of files):
- DEVELOPMENT.md: chapter 7 loses its fwht and winograd/spmm sections; what
  remains is the same-family case, plus an explicit warning that no body in
  the repo implements a different loop shape, so there is nothing to copy
  from if a new problem needs one
- USAGE.md: two lines no longer name contest-3 problems, keeping the
  platform facts they were illustrating
- ARCHITECTURE.md: the planned bodies/fwht.py entry and the phasing row that
  named winograd/spmm are gone; the phasing row keeps the general capability

No contest-3 identifier remains anywhere in the repository.

Co-Authored-By: Claude <noreply@anthropic.com>
Both were written before the code existed and had drifted.

USAGE.md is rewritten as a reference rather than a second tutorial. It now
says the code is landed (the old header still claimed nothing was built),
points to DEVELOPMENT.md for the walkthrough, and its field tables match the
real TargetDesc and DtypePolicy. Fixed while in there:

- the "how to add" step pointed at pipeline.py for the problem registry; the
  registry is bodies/__init__.py, and the tuple is (body, dtype, target)
- the TargetDesc and DtypePolicy snippets showed fields that do not exist
  (banks=, alu=, accumulate=, reduce_rules=) and omitted ones that do
  (mac, prod, bank, comparison)
- referenced CoveragePlan, which does not exist
- the CLI section was missing --unroll
- the troubleshooting table lacked the two failure modes that actually bit
  us: a register clobbered where it should not be, and a register reused
  where an operand had to survive
- eval_local.py does not pass a contest, which is why contest-2 problems
  reported "not open for evaluation"

ARCHITECTURE.md section 8 listed seven files marked as planned-but-this-phase
that were never created (plan.py, cost.py, coverage.py, kir.py, isel.py,
layout.py, _scaffold.py) and claimed each pass would implement CompilerPass.
Replaced with what exists, where the unimplemented concepts ended up, and an
explicit note that the pass_manager integration was never done. Section 9 now
carries real progress: S0/S0.5/S0.6 done, S2 measured and skipped, S1/S3/S4
open. Section 4's CacheBlocking row no longer calls itself the biggest win,
since the measurement says there is no miss cliff to fix here.

All cross-references between the four documents were verified to resolve.

Co-Authored-By: Claude <noreply@anthropic.com>
…oints

The tutorial taught "sweep the unroll factor and take the lowest total
cost", but never said what that total was over.  It was over the platform's
ten internal sizes -- which contestants cannot see, which the platform
deliberately stopped publishing, and which have since been replaced entirely.
Any conclusion derived that way is both unavailable to a contestant and
invalidated by the next data point change.

Section 5.1 now requires two things to be written down before sweeping:
the sampling of the published range, and the objective.  All measurements in
the section were redone on a declared sample (range:64:4096:8) rather than
the platform's sizes, and the tables say so.

The re-measured numbers changed the answer:

  u     total cost   worst-case deviation
  4       131,279       112.7%
  8       123,079       105.5%
  16      119,159       102.2%   <- most robust
  32      117,559       110.4%
  64      117,479       128.7%   <- lowest total

reducesum agrees (u=16 most robust at 102.0%).  The two objectives pick
different values, and the total-cost optimum itself moved when the sample
changed -- which is the point: "lowest total" is not a stable criterion
when you do not know the sizes.

Also re-based: section 3.6 and 4.2 examples, and section 6.3 gains a
per-point table showing that register blocking only engages when N % 4 == 0.
On the declared sample just 2 of 8 points take the fast path; the rest fall
back to the generic loop at 1.00x.  That constraint was written when the
sizes were all multiples of four, and is much more costly now.

The pipeline.py tables now record value + range + sample + objective
together, so a future reader can tell when a value has gone stale.

The measuring tool used throughout is probe.py (in the contest workspace,
outside this repo): it measures an arbitrary N given on the command line,
unlike measure.py / eval_local.py which read the platform's sizes.

Co-Authored-By: Claude <noreply@anthropic.com>

@FeelTheBeats FeelTheBeats left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

g

@jizhenjun
jizhenjun merged commit 8ea5d24 into main Oct 7, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants