Skip to content

[RFC][Muse-Glimmer-30B][CUDA/ROCm] ROADMAP #501

Description

@ryankert01

Status: Proposed
Target: post-v0.1.0 community roadmap
Owner: @ryankert01


1. Motivation

meta-models/Muse-Glimmer-30B is proposed as the first gated, normalized hybrid-attention target. It is a 29.6B dense model: 52 decoder layers in a repeating local, local, local, global pattern, served by vLLM and SGLang and trainable through Megatron Bridge. It adds three things the repo does not cover today:

  1. Gated, scale-normalized attention. Every attention layer applies Q/K normalization with a fixed scale factor and a per-head sigmoid gate on the attention output. The two sides already disagree on where the projection boundary sits: the Megatron bridge fuses Q/K/V and the gate into one projection, while vLLM fuses Q/K/V only.
  2. Two attention identities with two KV heads. Local layers use RoPE and a 2,048-token sliding window; global layers use full causal attention with no positional encoding. For local layers, sequence-split invariance has to cover window eviction, not only chunk boundaries. With 2 KV heads, any TP above 2 needs replicated KV-head ownership.
  3. A normalized input/output path. Token embeddings are normalized at lookup, and logits pass through an output multiplier and a soft cap before the log-probability.
    It is also the right size for a single-node closeout. The BF16 checkpoint is 59.58 GB, so every gate in this RFC except full-parameter end-to-end training runs on one 8× H100 node.

Drift is already reported in the wild: one community card notes that greedy decoding of this model in vLLM is not bit-reproducible between runs, and vLLM builds without the native implementation fall back to a generic Transformers path without failing.

Goal:

Given the same logical tokens, weights, positions, masks and sampling context, training and rollout must execute the same declared arithmetic contract and produce exactly equal selected-token log-probabilities in strict mode.

Two phases:

  • Phase A: full text model. Fused strict operators and TP/CP invariance for all 52 layers, plus a LoRA-scale end-to-end run.
  • Phase B: multimodal path. The perception encoder, projection into the text stream, and visual-token placement under the same strict contract.
    Speculative rollout with the bundled DFlash drafter is left to a follow-up RFC (§9). Strict mode in this RFC runs with speculative decoding off.

Relationship to other RFCs

Every operator row here extends the landed Qwen3-8B dense track at Muse Glimmer's own geometry. Where another open RFC has a row in the same family, it is cited as "if landed": sliding-window attention and the logit soft cap in #415, and the NoPE GQA core in #434, which has the same 32 / 2 / 128 head geometry. No row is blocked on another RFC.


2. Checkpoint fingerprint

Pin the implementation to the exact checkpoint revision used by each validation run. Architecture values are from the official model card. Config-level values are the defaults published in the Transformers and NeMo AutoModel muse_glimmer configs; museglimmer_arch_fingerprint must confirm each one against config.json and the modeling code at the pinned revision.

Item Value
Architecture MuseGlimmerForConditionalGeneration (model_type: muse_glimmer)
dtype / size BF16, 29.6B params including the vision encoder, 59.58 GB
Hidden size 6656
Decoder layers 52
Layer schedule [Local, Local, Local, Global] repeating (every_n_layers_nope: 4)
Layer counts (derived) 39 × local, 13 × global
Global layer indices (0-based, derived) 3, 7, 11, …, 47, 51 (NoPE layers are counted backward from the last layer)
Norm RMSNorm, rms_norm_eps: 1e-5; post_norm_eps: 1e-8
Vocabulary 202,048 (200,000 BPE + 2,048 special); bos: 200000, eos: 200001
Input/output embedding untied; normalize_tok_embeddings: true
Final logits output_multiplier: 0.19611613513818404 (= 1/√26), output_soft_cap_temp: 20.0
Context 131,072+
Attention (both layer types)
Q heads / KV heads / head dim 32 / 2 / 128 (GQA, 16 Q heads per KV head)
Q / K / V / O shapes (derived) 6656→4096, 6656→256, 6656→256, 4096→6656
Q/K norm use_qk_norm: true, qk_scale_factor: 43.7840518911
Output gate use_attn_output_gate: true; per-head sigmoid gate from self_attn.gate_proj
Local layers
Mask causal, sliding_window: 2048
Positional encoding RoPE, rope_theta: 500000
Global layers
Mask full causal
Positional encoding none (NoPE)
FFN
Type / shapes SwiGLU, 6656→19,968 (gate and up), 19,968→6656
Vision (Phase B)
Perception encoder about 1.8B ViT-G/14, 50 layers, width 1536, 16 heads, patch 14; frozen
Projection vision_output_dim: 6144 into the text stream; up to 4,096 visual tokens per image
Not covered by this RFC DFlash drafter and speculative verify (§9); quantized checkpoints; video input
To freeze before kernel work gate width and where it is applied (before or after head merge); Q/K norm form and where the scale enters; norm placement and what post_norm_eps governs; embedding normalization form and dtype; soft-cap formula, dtype and order relative to the multiplier; RoPE rotated-channel count; max_position_embeddings; local-layer KV cache eviction rule

Upstream references:

Upstream runtime status as of 2026-10-08 (re-verify at claim time)

  • vLLM: native inference exists (recipe lists vLLM 0.27.0+). RL-Kernel needs strict operator binding, cache policy and provenance, not a new serving implementation.
  • SGLang: inference exists.
  • Megatron: Megatron Bridge has a builder-backed Muse Glimmer model on the MCore hybrid model, HF ↔ Megatron conversion, an 8× H100 LoRA recipe and 32× H100 full-SFT recipes. RL-Kernel needs strict operator binding into that model.
  • VIME: a Muse Glimmer provider and model script must be verified at claim time. Unless found, treat it as new integration work on top of the Megatron Bridge model.

Reference-implementation hazards (never use as the strict oracle)

  • Different fused projections on each side. The Megatron bridge maps Q/K/V/gate into one gated-QKV projection; vLLM uses one fused QKV projection with the gate separate. Different GEMM shapes mean different accumulation.
  • Silent implementation fallback. A vLLM build without the native model resolves to a generic Transformers wrapper and still serves.
  • Non-reproducible greedy decode reported on the native vLLM path.
  • Window eviction. A local layer's output for a token must not depend on when older positions were evicted from the cache, or on how prefill was chunked around the 2,048 boundary.
  • Layer identity. Applying RoPE on a global layer, or a window mask on the wrong layer index, runs without error.
  • Logit post-processing. Multiplier and soft cap applied in a different order or dtype change every log-probability.
  • Two KV heads under TP. Native runtimes replicate KV heads in runtime-specific ways.
    An FP64 re-run of the Transformers reference remains useful as an independent numerical reference for allclose checks.

3. Architecture map

Layer schedule

Image

Figure 1: Muse Glimmer layer schedule. Local = RoPE + 2,048-token window. Global = full causal, NoPE. Every layer has gated GQA attention and a SwiGLU FFN.

One decoder layer and its fused-operator boundaries

Image

Figure 2: One decoder layer. [G] boxes reuse the deterministic GEMM track. [F*] boxes are fused strict operators; each is a single work item (§5). The sigmoid output gate is part of [F3] and [F4].


4. Acceptance contract

This RFC uses the same two kinds of acceptance as #434: operator acceptance for every fused operator and parallelism acceptance for TP and CP.

4.1 Operator acceptance: batch invariance + performance

Every operator row in §5 lands in one PR that includes:

  1. Batch invariance (forward and backward where applicable). The output for a logical row is bitwise-identical (torch.equal) under:
    • different batch sizes and batch positions;
    • for stateful operators (attention KV cache, local window cache), sequence-split invariance: the same token's output is identical whether it was produced in training (full sequence), prefill (any prompt length) or decode (any step, any preceding chunked-prefill split, any eviction timing).
  2. Performance. Benchmarked on real Muse Glimmer shapes (rollout decode batch sizes, prefill lengths on both sides of the 2,048 boundary, training micro-batch token counts) against the native production path.
    The existing WS1 rules define how invariance is achieved: fixed accumulator precision and reduction order; no Split-K / Stream-K / split-KV / cross-CTA atomics unless the merge tree is fixed by contract; no TF32 or fast-math reassociation; casts only at declared epilogue points; fail closed on unsupported strict geometry.

4.2 Projection-boundary rule

The contract declares one canonical grouping for the Q, K, V and gate projections. Both the Megatron-side and the vLLM-side adapters bind to that grouping. A runtime that keeps its native fused projection is not in strict mode.

4.3 Parallelism acceptance: TP and CP invariance

  • TP invariance: selected-token logprobs (and gradients, on the training side) at TP ∈ {2, 4, 8} are bitwise-equal to TP1. TP4 and TP8 use a declared KV-head replication policy.
  • CP invariance: CP ∈ {2, 4} bitwise-equal to CP1. Local layers start with a correctness-first contract and move to deterministic halo exchange only after parity; global layers use a fixed-order merge.
  • Combined: at least one TP × CP topology.
  • Provenance records topology, backend IDs, layer type per layer, window and eviction policy, projection grouping, KV replication policy and fallback state.

4.4 Platform and hardware

  • CUDA / ROCm: exact train/rollout selected-token logprob equality under the pinned contract on each platform; cross-platform equality is tested and reported, not assumed.
  • Hardware envelope: all operator, chain, TP/CP and Phase B gates fit one 8× H100 node. The end-to-end RL run is declared as LoRA at that scale. A full-parameter end-to-end run needs a larger cluster (upstream full-SFT recipes use 32 H100s) and is tracked as a separate optional row.

5. Work-item table

Status: 🙋 open → ⏳ in progress → 👀 in review → ✅ merged

One row × one platform = one PR. Modules that can be fused are grouped into a single fused-operator row, with independent CUDA and ROCm tracks. Every operator row must satisfy §4.1, every parallelism row §4.3. A row is not done until its validation lands with the implementation.

Reuse legend:

  • Reuse: an existing Qwen3 kernel, after Muse Glimmer shape qualification.
  • Extend: existing infrastructure reused, plus Muse-specific semantics.
  • New: a new strict operator or integration boundary.

5.1 Foundation

Work item What it does Reuse / dependency CUDA owner CUDA PR CUDA status ROCm owner ROCm PR ROCm status
museglimmer_arch_fingerprint Freeze checkpoint revision, 52-layer local/global schedule, every "to freeze" item in §2, projection grouping (§4.2), cache and eviction rules; per-layer operator trace; strict loader with full key coverage; land §4 as a versioned contract doc (maintainer-owned) New @ryankert01 ⏳ (same PR, platform-agnostic) 🙋

5.2 Fused operators (acceptance: batch invariance + performance)

Work item Fuses Reuse / dependency CUDA owner CUDA PR CUDA status ROCm owner ROCm PR ROCm status
normalized_embedding Token lookup for vocab 202,048 + embedding normalization, declared dtype and cast; fwd/bwd Extend #243 🙋 🙋
fused_add_rmsnorm [F1] Residual add + RMSNorm (eps 1e-5), the post-norm governed by post_norm_eps, and the final norm; fwd/bwd Extend existing norm infra 🙋 🙋
dense_gemm_qualification [G] Shape qualification of the deterministic GEMM for the canonical Q/K/V/gate grouping, O 4096→6656, and FFN 6656→19,968 / 19,968→6656 Reuse #180 / #322 / #343 + ROCm det-GEMM 🙋 🙋
qknorm_scaled_rope [F2] Q/K normalization with the fixed scale factor, then RoPE (theta 500,000) on local layers and identity on global layers Extend #228 and QK-norm infra 🙋 🙋
local_gated_attention [F3] Causal sliding-window GQA core (window 2048, Q32/KV2, D128) + per-head sigmoid output gate + window KV cache with declared eviction; prefill / chunked-prefill / decode; fwd/bwd Extend #240 / #230; #415 sliding_attention_core if landed 🙋 🙋
global_gated_attention [F4] Full causal NoPE GQA core (Q32/KV2, D128) + per-head sigmoid output gate + KV cache; prefill / chunked-prefill / decode; fwd/bwd Extend #240 / #230; #434 nope_gqa_attention if landed 🙋 🙋
swiglu_ffn [F5] Gate/up GEMM with FP32-math SwiGLU epilogue + down GEMM, single cast; fwd/bwd Reuse #280 / #322 / #343 🙋 🙋
lm_head_softcap_logprob Untied LM head 6656→202,048 + output multiplier + soft cap (temp 20) + selected-token logprob with fixed vocab reduction Extend #243 / #204 / #336; #415 final_logit_softcap if landed 🙋 🙋

5.3 Model closeout

Work item What it does Reuse / dependency CUDA owner CUDA PR CUDA status ROCm owner ROCm PR ROCm status
full_model_chain Full 52-layer text model with real weights: train vs prefill vs decode parity; sequence-split sweeps (lengths 2047 / 2048 / 2049, chunked-prefill splits across the window, long decode past eviction); first-drift localization per layer type (maintainer-owned) Extend #315 🙋 🙋

5.4 Parallelism (acceptance: TP / CP invariance)

Work item What it does Reuse / dependency CUDA owner CUDA PR CUDA status ROCm owner ROCm PR ROCm status
tp_invariance TP2/TP4/TP8 == TP1: 32 Q-head ownership, replicated KV-head policy for 2 KV heads, gate sharding, FFN 19,968 sharding, vocab-parallel soft-capped logprob, fixed reduction order Extend Qwen3 WS2, #241, #336 🙋 🙋
cp_invariance CP2/CP4 == CP1: correctness-first local-window contract then deterministic halo exchange, fixed-order global merge; one TP × CP topology Extend #235; #415 cp_sliding_window_policy if landed 🙋 🙋

5.5 Integration and end-to-end

Work item What it does Reuse / dependency Status
vllm_runtime_adapter Bind strict operators into vLLM's native Muse Glimmer model with the canonical projection grouping; fail closed on the Transformers fallback; provenance read-back Extend #338 / #360 🙋
megatron_vime_provider Bind strict operators into the Megatron Bridge Muse Glimmer model; VIME model script and provider; HF ↔ Megatron mapping checked against the canonical grouping; weight sync to vLLM New; extend #352 🙋
museglimmer_ablation Probes on top of #230: layer-type swap, RoPE on global, window size and eviction timing, gate on/off, Q/K scale, embedding normalization, soft-cap order, projection grouping, KV replication Extend #230 🙋
final_e2e_lora End-to-end RL on one 8× H100 node with LoRA: native path vs RL-Kernel + VIME, CUDA and ROCm Extend #377 methodology 🙋
final_e2e_full_param Optional: the same experiment with full-parameter training on a larger cluster Depends on hardware access 🙋

5.6 Phase B

Work item What it does Reuse / dependency Status
vision_input_path Perception-encoder contract, projection into the text stream, visual-token placement and masks under the same strict logprob contract New 🙋

6. Recommended claim order

  1. museglimmer_arch_fingerprint (maintainer; other rows start once it lands)
  2. normalized_embedding / fused_add_rmsnorm / dense_gemm_qualification / swiglu_ffn (mostly reuse; unblocks the chain harness)
  3. qknorm_scaled_rope
  4. global_gated_attention (simpler: no window, no RoPE)
  5. local_gated_attention (critical path: window cache and eviction)
  6. lm_head_softcap_logprob
  7. full_model_chain (maintainer)
  8. tp_invariance → cp_invariance
  9. vllm_runtime_adapter / megatron_vime_provider → museglimmer_ablation → final_e2e_lora
  10. Phase B: vision_input_path
    Steps 2, 3 and 6 are independent and can be claimed in parallel. Within each step, CUDA and ROCm PRs can proceed in parallel.

7. Contribution notes

  • One row × one platform = one PR. A large fused row may land as a short stack, but the row is marked done only when the whole fused operator passes §4.1.
  • CUDA and ROCm PRs implement the same versioned contract and share one test suite. A strict kernel written once in Triton may fill both columns in a single PR, but must pass full acceptance on each platform.
  • Local and global attention are two acceptance items. Covering one does not close the other, even though they share head geometry and the gate.
  • The gate belongs to the attention rows. A PR that lands the attention core without the gate does not close the row.
  • local_gated_attention scope. Train, prefill and decode share device functions, and the window cache with its eviction rule is part of the row.
  • Oracles. Parity uses torch.equal between two strict paths. Correctness uses allclose against an independent FP64 reference. Never test a kernel against an oracle that shares its code.
  • Fail closed on a Transformers-fallback model class, a non-canonical projection grouping, an unsupported window or eviction mode, an unaligned CP shard, an undeclared KV replication policy, or speculative decoding enabled.

8. Open questions

  1. Canonical projection grouping. One fused Q/K/V/gate GEMM (matches Megatron), fused Q/K/V plus a separate gate (matches vLLM), or four separate GEMMs?
  2. Gate geometry. What is the gate's width, and is it applied per head before the O projection?
  3. Row sharing with [RFC][Gemma-4-31B-it][CUDA/ROCm] WS1/WS2 kernel roadmap, ablation matrix and integration plan #415 and [RFC][Nemotron-3-Nano-30B-A3B][CUDA/ROCm] ROADMAP #434. Can local_gated_attention share its window core with Gemma's sliding attention, and global_gated_attention share verbatim with Nemotron's NoPE core, with the gate as an epilogue?
  4. Eviction policy. Evict eagerly at the window edge, or keep a fixed over-allocation so the follow-up speculative-rollout contract (§9) never needs to restore evicted entries?
  5. Soft cap placement. Strict FP32 after the multiplier, as with Gemma's soft cap, or exactly as the reference orders it?
  6. End-to-end scale. Is a LoRA end-to-end run sufficient for closeout, or is the full-parameter run required and someone can provide the hardware?
  7. Performance targets. A fixed ratio to native per operator, or report-only until the first perf closeout?

9. Out of scope

Follow-up RFC: speculative rollout (DFlash)

Muse Glimmer ships a DFlash drafter (meta-models/Muse-Glimmer-30B-assistant; 5 layers, 16-token blocks; arXiv:2602.06036) that reads target hidden features at layers {1, 13, 25, 37, 49} of 52. The target verifies each block in one forward, which changes the rollout forward shape and needs cache rollback. This RFC runs strict mode with speculation off. The follow-up would add:

  • Verify equals decode: one verify forward over k drafted tokens is bitwise-equal to k single-token decodes, including blocks that straddle the 2,048 window boundary.
  • Rollback leaves no trace: after a rejected tail, every KV and window cache is byte-equal to the non-speculative state.
  • Hidden-feature tap: exposing target features to the drafter changes no target output byte.
    The Phase A contracts (§4.1, and the window cache and eviction rule in local_gated_attention) are written so the follow-up only needs new rows, not revised contracts.

Also out of scope

  • 4-bit K-quant, NVFP4, MXFP4 and other quantized checkpoints.
  • Video input (the model processes video as individual frames).
  • Sampling RNG contract (after logprob exactness).

If you are interested, just ping below!

Activity

  1. self-assigned this
    on Oct 9, 2026
  2. Flink-ddd commented on Oct 10, 2026

    @Flink-ddd
    Collaborator

    Hi everyone! Please note that all PRs should target the test-museglimmer branch instead of main. Once the CI and validation tests pass successfully, the changes will be merged into main. Thanks for your contribution!

  3. added
    platform: cudaSpecific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
    platform: rocmSpecific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA)
    multimodalFeatures, bugs, or optimizations specific to multimodal support.
    on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Muse-GlimmermultimodalFeatures, bugs, or optimizations specific to multimodal support.new-modelnext-phaseplatform: cudaSpecific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)platform: rocmSpecific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA)

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions