llama32: parallelize single-row fused RMS norm across two harts - #215
llama32: parallelize single-row fused RMS norm across two harts#215ChiruGuru99 wants to merge 3 commits into
Conversation
|
The previously tested runtime revision ( |
|
Revalidated the exact current PR head rather than relying on the earlier Outer PR head: Complete candidate runtime stack:
Exact paired trusted Llama:
All ten shared-runtime regression models were rerun using the exact candidate runtime and completed with benchmark status Exact paired SmolVLM2 trusted gate passed:
Please approve and rerun PR workflows for exact head |
|
the first six commits (9f3629e3b through 334ec451b) are inherited runtime-base commits and are listed only for complete reproducibility. The changes contributed by this submission are 26b101151, 1c58f0cd9, and ec53d5620. No authorship is claimed for the inherited commits. |
Summary
llama.cpp-etto26b1011518a33e15050796630e829c315ab6704f.[2048,1,1,1]fused RMS-norm/multiply workload.Runtime commit:
ChiruGuru99/llama.cpp@26b1011