Repository navigation
[RFC][Gemma-4-31B-it][CUDA/ROCm] WS1/WS2 kernel roadmap, ablation matrix and integration plan #415
Description
Activity
- addedmultimodalFeatures, bugs, or optimizations specific to multimodal support.Features, bugs, or optimizations specific to multimodal support.
on Sep 17, 2026 - changed the title
[-][RFC][Gemma-4-31B-it][CUDA/ROCm] Operator-level train–rollout consistency plan (WS1/WS2/VIME)[/-][+][RFC][Gemma-4-31B-it][CUDA/ROCm] Operator-level train–rollout consistency plan (WS1/WS2)[/+]on Sep 17, 2026 - changed the title
[-][RFC][Gemma-4-31B-it][CUDA/ROCm] Operator-level train–rollout consistency plan (WS1/WS2)[/-][+][RFC][Gemma-4-31B-it][CUDA/ROCm] WS1/WS2 kernel roadmap, ablation matrix and integration plan[/+]on Sep 17, 2026 - addedplatform: cudaSpecific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)platform: rocmSpecific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA)Specific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA)
on Sep 17, 2026 I'm interested in this. Can I take the first one(gemma4_arch_fingerprint) as a try/start? Thank you!
Hi @haotianxue31 , Sure, thanks for picking it up! Please go ahead with gemma4_arch_fingerprint. The goal of this task is to freeze the exact Gemma-4-31B-it architecture and runtime assumptions before we start the kernel work. Please open a PR with the first version and we can review the details there. I’ll mark this item as claimed by you. Thanks!
Reacted by haotianxue31Hi @Flink-ddd, as discussed, I'd like to take on
ffn_gate_up_gemm,gelu_tanh_and_mul, andffn_down_gemm. Thanks~Reacted by VensenHi @Flink-ddd, as discussed offline, can I lead this issue as a try? If permitted, I'd like to take on tasks from gemma4_operator_trace to post_attention_residual myself for now. Please let me know if you think it sounds good to you. Thanks!
Reacted by Vensensure, assigned. Thank you.
1 remaining item
sure, assigned. Thank you.
- added a commit that references this issue
on Sep 25, 2026 Hi @Flink-ddd, I'd like to claim post_ffn_residual. This is my first operator contribution, and I have access to both CUDA and ROCm GPUs. I'll coordinate with the norm/MLP owners and include validation tests. Is it still available?
Reacted by VensenHi @FED4 , sure, assigned. Thank you.
Hi @cwgan @FED4 @Jungle430 @haotianxue31 , Please note that all PRs should target the test-gemma branch instead of main. Once the CI and validation tests pass successfully, the changes will be merged into main. Thanks for your contribution!
Reacted by Jungle, haotianxue31 and cwganHi @Flink-ddd, following my
final_logit_softcapPR #444, I'd also like to take onsoftcapped_selected_logprob.I plan to build on the existing logprob infrastructure and implement the Triton forward/backward path with fixed-order FP32 vocabulary reduction, correctness and batch-invariance tests, and CUDA benchmarks. I have access to an NVIDIA GPU and would need help with ROCm hardware validation.
Please let me know if this scope works. Thanks!
Reacted by JungleHi everyone! Please note that all PRs should target the test-gemma branch instead of main. Once the CI and validation tests pass successfully, the changes will be merged into main. Thanks for your contribution!
Reacted by JungleImportant Notice: Repository Refactor
Hi everyone! The repository refactor has been merged into main and synced to test-gemma. Please merge the latest test-gemma from RL-Align/RL-Kernel into your development branch, resolve any conflicts, and update your code and tests to follow the new directory structure. Please rerun the relevant tests after updating. Existing PRs can be updated in place and should continue targeting test-gemma, there is no need to open new PRs. Thank you.
Status: Proposed
Target: post-v0.1.0 community roadmap
Related: #386
Reference Qwen3-8B Dense track: #180, #204, #228, #230, #240, #243, #280, #315, #322, #336, #338, #343, #352, #360, #361, #377.
1. Motivation
RL-Kernel has already closed the dense Qwen3-8B train/rollout consistency loop on CUDA and ROCm, including WS1 operator contracts, WS2 distributed execution, VIME/vLLM integration, ablation, and end-to-end experiments.
google/gemma-4-31B-it is the next dense model target. The goal is the same:
This RFC covers CUDA and ROCm, and splits the work into two phases:
The split keeps the text-path consistency work unblocked while VIME-side Gemma 4 multimodal support is brought up.
2. Checkpoint fingerprint
The implementation must be pinned to the exact checkpoint/runtime revision used by each validation run. The current google/gemma-4-31B-it architecture relevant to this RFC is:
Vision-side fingerprint:
Upstream references:
Upstream runtime status
3. Architecture map
Figure 1 — Gemma-4-31B-it architecture and RL-Kernel consistency boundaries. The text backbone is closed first; the image path joins through deterministic vision soft-token construction and embedding merge.
The central difference from Qwen3-8B is that Gemma 4 does not have one uniform attention layer. Sliding and global layers have different head geometry, KV ownership, RoPE, mask semantics, cache behavior, and K/V projection semantics. They must therefore have separate runtime identities and validation rows.
One decoder layer
Figure 2 — One Gemma 4 decoder layer. WS1 qualifies each arithmetic boundary; WS2 wraps the same boundary with deterministic TP/CP/SP ownership and communication order.
4. Numerical contract
The existing RL-Kernel WS1 numerical standard remains authoritative. Gemma 4 does not get a weaker consistency definition.
For every strict reduction-bearing boundary:
Platform acceptance:
5. Kernel and work-item table
Status: 🙋 open → ⏳ in progress → 👀 in review → ✅ merged
To claim a task, put your handle in the GitHub column and open a PR. A row is not complete until both the implementation and the row-specific validation contract are included.
Legend for reuse:
6. Recommended claim order
For contributors, the recommended order is:
7. Contribution notes
if you are interested just ping below!