Repository navigation
feat(ir): first-class ops for FWHT/WinogradConv/CSR-SpMM with C and machine-layer lowering - #99
FeelTheBeats wants to merge 5 commits into
Conversation
Add FWHT/WINOGRAD_CONV/SPMM_CSR opcodes, builder helpers and verifier signatures, plus dtype-dispatched kernels (Q16.16 on INT32, native FP32 on FLOAT32) in ir_numpy_ops. Wire the ONNX frontend with the org.scratchv domain whitelist, self-contained shape wiring, optional Winograd bias and compile-time F(2,3) kernel folding that reuses the standalone transform. Tests: IR kernel/verifier coverage and frontend parse+run vs the standalone references (tests/_op_ref.py).
Add FP32 _infer/_emit for the three operators: FWHT butterfly with 1/N inverse scaling, CSR SpMM with status-4 index checks, and the Winograd F(2,3) tile pipeline with tracked static scratch (workspace_bytes now accounts for it). Tests execute host-compiled C against the shared references, cover static rejection, and run a CompilerDriver end-to-end path.
…e 3) Add --platform-asm / CompilerConfig.platform_asm, which reuses the verified standalone FP32/rv32imf kernels for single-operator graphs (Fwht, Conv/ WinogradConv, SpmmCsr) instead of lowering IR through the scalar allocator/ ABI-frame path. Unsupported graphs are rejected explicitly. Tests cover the cnn_entry contract, size independence, multi-op rejection, the CLI flag wiring, and CompilerDriver end-to-end emission.
Fill machine-layer gaps needed by tensor-kernel lowering: immediate shift opcodes, a RET terminator with explicit semantics, and AsmEmitter formatting for float load/store memory operands.
Rewrite platform_emit to build MachineInstr kernels (FWHT / CSR SpMM / direct Conv) and render them with AsmEmitter, dropping the dependency on the standalone string generators. Add a distinct FMV_W_X opcode so float zeroing emits the standalone/interpreter mnemonic fmv.w.x. Numeric tests execute the emitted listings in the shared in-test RV32IMF interpreter.
🤖 AI Code Review
📁
|
概述
把 FWHT / WinogradConv / CSR-SpMM 三个自定义张量算子从 standalone 生成器提升为
一等 IR 算子,并打通两条 lowering 路径:可移植 C 后端、平台汇编后端(经机器层
由 AsmEmitter 渲染)。主编译管线可直接产出尺寸无关的 .s。
改动分层
1. IR / 前端(ae4bf03)
以及编译期 F(2,3) kernel folding(复用 standalone 变换)
2. 可移植 C 后端(7669610)
status-4 索引校验、Winograd F(2,3) tile 流水线(workspace_bytes 计入静态 scratch)
3. 平台汇编后端(d42976b / 0a2c043 / bb63617)
--platform-asm/CompilerConfig.platform_asm,单算子图复用已验证的standalone FP32/rv32imf kernel,不支持的多算子图显式拒绝(d42976b)
AsmEmitter 渲染,去除对 standalone 字符串生成器的依赖;新增 FMV_W_X opcode
以产出
fmv.w.x(bb63617)测试
tests/_op_ref.py共享参考实现,IR / C / 汇编路径统一对拍依赖与说明
--platform-asm(显式报错,后续可扩展)