You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Status: Proposed
Target: post-v0.1.0 community roadmap
Owner: @ryankert01
1. Motivation
meta-models/Muse-Glimmer-30B is proposed as the first gated, normalized hybrid-attention target. It is a 29.6B dense model: 52 decoder layers in a repeating local, local, local, global pattern, served by vLLM and SGLang and trainable through Megatron Bridge. It adds three things the repo does not cover today:
Gated, scale-normalized attention. Every attention layer applies Q/K normalization with a fixed scale factor and a per-head sigmoid gate on the attention output. The two sides already disagree on where the projection boundary sits: the Megatron bridge fuses Q/K/V and the gate into one projection, while vLLM fuses Q/K/V only.
Two attention identities with two KV heads. Local layers use RoPE and a 2,048-token sliding window; global layers use full causal attention with no positional encoding. For local layers, sequence-split invariance has to cover window eviction, not only chunk boundaries. With 2 KV heads, any TP above 2 needs replicated KV-head ownership.
A normalized input/output path. Token embeddings are normalized at lookup, and logits pass through an output multiplier and a soft cap before the log-probability.
It is also the right size for a single-node closeout. The BF16 checkpoint is 59.58 GB, so every gate in this RFC except full-parameter end-to-end training runs on one 8× H100 node.
Drift is already reported in the wild: one community card notes that greedy decoding of this model in vLLM is not bit-reproducible between runs, and vLLM builds without the native implementation fall back to a generic Transformers path without failing.
Goal:
Given the same logical tokens, weights, positions, masks and sampling context, training and rollout must execute the same declared arithmetic contract and produce exactly equal selected-token log-probabilities in strict mode.
Two phases:
Phase A: full text model. Fused strict operators and TP/CP invariance for all 52 layers, plus a LoRA-scale end-to-end run.
Phase B: multimodal path. The perception encoder, projection into the text stream, and visual-token placement under the same strict contract.
Speculative rollout with the bundled DFlash drafter is left to a follow-up RFC (§9). Strict mode in this RFC runs with speculative decoding off.
Relationship to other RFCs
Every operator row here extends the landed Qwen3-8B dense track at Muse Glimmer's own geometry. Where another open RFC has a row in the same family, it is cited as "if landed": sliding-window attention and the logit soft cap in #415, and the NoPE GQA core in #434, which has the same 32 / 2 / 128 head geometry. No row is blocked on another RFC.
2. Checkpoint fingerprint
Pin the implementation to the exact checkpoint revision used by each validation run. Architecture values are from the official model card. Config-level values are the defaults published in the Transformers and NeMo AutoModel muse_glimmer configs; museglimmer_arch_fingerprint must confirm each one against config.json and the modeling code at the pinned revision.
vision_output_dim: 6144 into the text stream; up to 4,096 visual tokens per image
Not covered by this RFC
DFlash drafter and speculative verify (§9); quantized checkpoints; video input
To freeze before kernel work
gate width and where it is applied (before or after head merge); Q/K norm form and where the scale enters; norm placement and what post_norm_eps governs; embedding normalization form and dtype; soft-cap formula, dtype and order relative to the multiplier; RoPE rotated-channel count; max_position_embeddings; local-layer KV cache eviction rule
Megatron Bridge bridge.models.muse_glimmer (bridge, builder, config) and bridge.recipes.muse_glimmer.h100
NeMo AutoModel muse_glimmer model and recipes
Perception Encoder: arXiv:2504.13181
Upstream runtime status as of 2026-10-08 (re-verify at claim time)
vLLM: native inference exists (recipe lists vLLM 0.27.0+). RL-Kernel needs strict operator binding, cache policy and provenance, not a new serving implementation.
SGLang: inference exists.
Megatron: Megatron Bridge has a builder-backed Muse Glimmer model on the MCore hybrid model, HF ↔ Megatron conversion, an 8× H100 LoRA recipe and 32× H100 full-SFT recipes. RL-Kernel needs strict operator binding into that model.
VIME: a Muse Glimmer provider and model script must be verified at claim time. Unless found, treat it as new integration work on top of the Megatron Bridge model.
Reference-implementation hazards (never use as the strict oracle)
Different fused projections on each side. The Megatron bridge maps Q/K/V/gate into one gated-QKV projection; vLLM uses one fused QKV projection with the gate separate. Different GEMM shapes mean different accumulation.
Silent implementation fallback. A vLLM build without the native model resolves to a generic Transformers wrapper and still serves.
Non-reproducible greedy decode reported on the native vLLM path.
Window eviction. A local layer's output for a token must not depend on when older positions were evicted from the cache, or on how prefill was chunked around the 2,048 boundary.
Layer identity. Applying RoPE on a global layer, or a window mask on the wrong layer index, runs without error.
Logit post-processing. Multiplier and soft cap applied in a different order or dtype change every log-probability.
Two KV heads under TP. Native runtimes replicate KV heads in runtime-specific ways.
An FP64 re-run of the Transformers reference remains useful as an independent numerical reference for allclose checks.
3. Architecture map
Layer schedule
Figure 1: Muse Glimmer layer schedule. Local = RoPE + 2,048-token window. Global = full causal, NoPE. Every layer has gated GQA attention and a SwiGLU FFN.
One decoder layer and its fused-operator boundaries
Figure 2: One decoder layer.[G] boxes reuse the deterministic GEMM track. [F*] boxes are fused strict operators; each is a single work item (§5). The sigmoid output gate is part of [F3] and [F4].
4. Acceptance contract
This RFC uses the same two kinds of acceptance as #434: operator acceptance for every fused operator and parallelism acceptance for TP and CP.
Every operator row in §5 lands in one PR that includes:
Batch invariance (forward and backward where applicable). The output for a logical row is bitwise-identical (torch.equal) under:
different batch sizes and batch positions;
for stateful operators (attention KV cache, local window cache), sequence-split invariance: the same token's output is identical whether it was produced in training (full sequence), prefill (any prompt length) or decode (any step, any preceding chunked-prefill split, any eviction timing).
Performance. Benchmarked on real Muse Glimmer shapes (rollout decode batch sizes, prefill lengths on both sides of the 2,048 boundary, training micro-batch token counts) against the native production path.
The existing WS1 rules define how invariance is achieved: fixed accumulator precision and reduction order; no Split-K / Stream-K / split-KV / cross-CTA atomics unless the merge tree is fixed by contract; no TF32 or fast-math reassociation; casts only at declared epilogue points; fail closed on unsupported strict geometry.
4.2 Projection-boundary rule
The contract declares one canonical grouping for the Q, K, V and gate projections. Both the Megatron-side and the vLLM-side adapters bind to that grouping. A runtime that keeps its native fused projection is not in strict mode.
4.3 Parallelism acceptance: TP and CP invariance
TP invariance: selected-token logprobs (and gradients, on the training side) at TP ∈ {2, 4, 8} are bitwise-equal to TP1. TP4 and TP8 use a declared KV-head replication policy.
CP invariance: CP ∈ {2, 4} bitwise-equal to CP1. Local layers start with a correctness-first contract and move to deterministic halo exchange only after parity; global layers use a fixed-order merge.
Combined: at least one TP × CP topology.
Provenance records topology, backend IDs, layer type per layer, window and eviction policy, projection grouping, KV replication policy and fallback state.
4.4 Platform and hardware
CUDA / ROCm: exact train/rollout selected-token logprob equality under the pinned contract on each platform; cross-platform equality is tested and reported, not assumed.
Hardware envelope: all operator, chain, TP/CP and Phase B gates fit one 8× H100 node. The end-to-end RL run is declared as LoRA at that scale. A full-parameter end-to-end run needs a larger cluster (upstream full-SFT recipes use 32 H100s) and is tracked as a separate optional row.
5. Work-item table
Status: 🙋 open → ⏳ in progress → 👀 in review → ✅ merged
One row × one platform = one PR. Modules that can be fused are grouped into a single fused-operator row, with independent CUDA and ROCm tracks. Every operator row must satisfy §4.1, every parallelism row §4.3. A row is not done until its validation lands with the implementation.
Reuse legend:
Reuse: an existing Qwen3 kernel, after Muse Glimmer shape qualification.
Extend: existing infrastructure reused, plus Muse-specific semantics.
New: a new strict operator or integration boundary.
5.1 Foundation
Work item
What it does
Reuse / dependency
CUDA owner
CUDA PR
CUDA status
ROCm owner
ROCm PR
ROCm status
museglimmer_arch_fingerprint
Freeze checkpoint revision, 52-layer local/global schedule, every "to freeze" item in §2, projection grouping (§4.2), cache and eviction rules; per-layer operator trace; strict loader with full key coverage; land §4 as a versioned contract doc (maintainer-owned)
Full 52-layer text model with real weights: train vs prefill vs decode parity; sequence-split sweeps (lengths 2047 / 2048 / 2049, chunked-prefill splits across the window, long decode past eviction); first-drift localization per layer type (maintainer-owned)
CP2/CP4 == CP1: correctness-first local-window contract then deterministic halo exchange, fixed-order global merge; one TP × CP topology
Extend #235; #415cp_sliding_window_policy if landed
🙋
🙋
5.5 Integration and end-to-end
Work item
What it does
Reuse / dependency
Status
vllm_runtime_adapter
Bind strict operators into vLLM's native Muse Glimmer model with the canonical projection grouping; fail closed on the Transformers fallback; provenance read-back
Bind strict operators into the Megatron Bridge Muse Glimmer model; VIME model script and provider; HF ↔ Megatron mapping checked against the canonical grouping; weight sync to vLLM
Phase B: vision_input_path
Steps 2, 3 and 6 are independent and can be claimed in parallel. Within each step, CUDA and ROCm PRs can proceed in parallel.
7. Contribution notes
One row × one platform = one PR. A large fused row may land as a short stack, but the row is marked done only when the whole fused operator passes §4.1.
CUDA and ROCm PRs implement the same versioned contract and share one test suite. A strict kernel written once in Triton may fill both columns in a single PR, but must pass full acceptance on each platform.
Local and global attention are two acceptance items. Covering one does not close the other, even though they share head geometry and the gate.
The gate belongs to the attention rows. A PR that lands the attention core without the gate does not close the row.
local_gated_attention scope. Train, prefill and decode share device functions, and the window cache with its eviction rule is part of the row.
Oracles. Parity uses torch.equal between two strict paths. Correctness uses allclose against an independent FP64 reference. Never test a kernel against an oracle that shares its code.
Fail closed on a Transformers-fallback model class, a non-canonical projection grouping, an unsupported window or eviction mode, an unaligned CP shard, an undeclared KV replication policy, or speculative decoding enabled.
8. Open questions
Canonical projection grouping. One fused Q/K/V/gate GEMM (matches Megatron), fused Q/K/V plus a separate gate (matches vLLM), or four separate GEMMs?
Gate geometry. What is the gate's width, and is it applied per head before the O projection?
Eviction policy. Evict eagerly at the window edge, or keep a fixed over-allocation so the follow-up speculative-rollout contract (§9) never needs to restore evicted entries?
Soft cap placement. Strict FP32 after the multiplier, as with Gemma's soft cap, or exactly as the reference orders it?
End-to-end scale. Is a LoRA end-to-end run sufficient for closeout, or is the full-parameter run required and someone can provide the hardware?
Performance targets. A fixed ratio to native per operator, or report-only until the first perf closeout?
9. Out of scope
Follow-up RFC: speculative rollout (DFlash)
Muse Glimmer ships a DFlash drafter (meta-models/Muse-Glimmer-30B-assistant; 5 layers, 16-token blocks; arXiv:2602.06036) that reads target hidden features at layers {1, 13, 25, 37, 49} of 52. The target verifies each block in one forward, which changes the rollout forward shape and needs cache rollback. This RFC runs strict mode with speculation off. The follow-up would add:
Verify equals decode: one verify forward over k drafted tokens is bitwise-equal to k single-token decodes, including blocks that straddle the 2,048 window boundary.
Rollback leaves no trace: after a rejected tail, every KV and window cache is byte-equal to the non-speculative state.
Hidden-feature tap: exposing target features to the drafter changes no target output byte.
The Phase A contracts (§4.1, and the window cache and eviction rule in local_gated_attention) are written so the follow-up only needs new rows, not revised contracts.
Also out of scope
4-bit K-quant, NVFP4, MXFP4 and other quantized checkpoints.
Video input (the model processes video as individual frames).
Hi everyone! Please note that all PRs should target the test-museglimmer branch instead of main. Once the CI and validation tests pass successfully, the changes will be merged into main. Thanks for your contribution!
Status: Proposed
Target: post-v0.1.0 community roadmap
Owner: @ryankert01
1. Motivation
meta-models/Muse-Glimmer-30Bis proposed as the first gated, normalized hybrid-attention target. It is a 29.6B dense model: 52 decoder layers in a repeating local, local, local, global pattern, served by vLLM and SGLang and trainable through Megatron Bridge. It adds three things the repo does not cover today:It is also the right size for a single-node closeout. The BF16 checkpoint is 59.58 GB, so every gate in this RFC except full-parameter end-to-end training runs on one 8× H100 node.
Drift is already reported in the wild: one community card notes that greedy decoding of this model in vLLM is not bit-reproducible between runs, and vLLM builds without the native implementation fall back to a generic Transformers path without failing.
Goal:
Two phases:
Speculative rollout with the bundled DFlash drafter is left to a follow-up RFC (§9). Strict mode in this RFC runs with speculative decoding off.
Relationship to other RFCs
Every operator row here extends the landed Qwen3-8B dense track at Muse Glimmer's own geometry. Where another open RFC has a row in the same family, it is cited as "if landed": sliding-window attention and the logit soft cap in #415, and the NoPE GQA core in #434, which has the same 32 / 2 / 128 head geometry. No row is blocked on another RFC.
2. Checkpoint fingerprint
Pin the implementation to the exact checkpoint revision used by each validation run. Architecture values are from the official model card. Config-level values are the defaults published in the Transformers and NeMo AutoModel
muse_glimmerconfigs;museglimmer_arch_fingerprintmust confirm each one againstconfig.jsonand the modeling code at the pinned revision.MuseGlimmerForConditionalGeneration(model_type: muse_glimmer)[Local, Local, Local, Global]repeating (every_n_layers_nope: 4)rms_norm_eps: 1e-5;post_norm_eps: 1e-8bos: 200000,eos: 200001normalize_tok_embeddings: trueoutput_multiplier: 0.19611613513818404(= 1/√26),output_soft_cap_temp: 20.0use_qk_norm: true,qk_scale_factor: 43.7840518911use_attn_output_gate: true; per-head sigmoid gate fromself_attn.gate_projsliding_window: 2048rope_theta: 500000vision_output_dim: 6144into the text stream; up to 4,096 visual tokens per imagepost_norm_epsgoverns; embedding normalization form and dtype; soft-cap formula, dtype and order relative to the multiplier; RoPE rotated-channel count;max_position_embeddings; local-layer KV cache eviction ruleUpstream references:
muse_glimmerimplementation (vllm#51655) and recipe: https://recipes.vllm.ai/meta-models/Muse-Glimmer-30Bbridge.models.muse_glimmer(bridge, builder, config) andbridge.recipes.muse_glimmer.h100muse_glimmermodel and recipesUpstream runtime status as of 2026-10-08 (re-verify at claim time)
Reference-implementation hazards (never use as the strict oracle)
An FP64 re-run of the Transformers reference remains useful as an independent numerical reference for
allclosechecks.3. Architecture map
Layer schedule
Figure 1: Muse Glimmer layer schedule. Local = RoPE + 2,048-token window. Global = full causal, NoPE. Every layer has gated GQA attention and a SwiGLU FFN.
One decoder layer and its fused-operator boundaries
Figure 2: One decoder layer.
[G]boxes reuse the deterministic GEMM track.[F*]boxes are fused strict operators; each is a single work item (§5). The sigmoid output gate is part of[F3]and[F4].4. Acceptance contract
This RFC uses the same two kinds of acceptance as #434: operator acceptance for every fused operator and parallelism acceptance for TP and CP.
4.1 Operator acceptance: batch invariance + performance
Every operator row in §5 lands in one PR that includes:
torch.equal) under:The existing WS1 rules define how invariance is achieved: fixed accumulator precision and reduction order; no Split-K / Stream-K / split-KV / cross-CTA atomics unless the merge tree is fixed by contract; no TF32 or fast-math reassociation; casts only at declared epilogue points; fail closed on unsupported strict geometry.
4.2 Projection-boundary rule
The contract declares one canonical grouping for the Q, K, V and gate projections. Both the Megatron-side and the vLLM-side adapters bind to that grouping. A runtime that keeps its native fused projection is not in strict mode.
4.3 Parallelism acceptance: TP and CP invariance
4.4 Platform and hardware
5. Work-item table
Status: 🙋 open → ⏳ in progress → 👀 in review → ✅ merged
One row × one platform = one PR. Modules that can be fused are grouped into a single fused-operator row, with independent CUDA and ROCm tracks. Every operator row must satisfy §4.1, every parallelism row §4.3. A row is not done until its validation lands with the implementation.
Reuse legend:
5.1 Foundation
5.2 Fused operators (acceptance: batch invariance + performance)
post_norm_eps, and the final norm; fwd/bwdsliding_attention_coreif landednope_gqa_attentionif landedfinal_logit_softcapif landed5.3 Model closeout
5.4 Parallelism (acceptance: TP / CP invariance)
cp_sliding_window_policyif landed5.5 Integration and end-to-end
5.6 Phase B
6. Recommended claim order
Steps 2, 3 and 6 are independent and can be claimed in parallel. Within each step, CUDA and ROCm PRs can proceed in parallel.
7. Contribution notes
local_gated_attentionscope. Train, prefill and decode share device functions, and the window cache with its eviction rule is part of the row.torch.equalbetween two strict paths. Correctness usesallcloseagainst an independent FP64 reference. Never test a kernel against an oracle that shares its code.8. Open questions
local_gated_attentionshare its window core with Gemma's sliding attention, andglobal_gated_attentionshare verbatim with Nemotron's NoPE core, with the gate as an epilogue?9. Out of scope
Follow-up RFC: speculative rollout (DFlash)
Muse Glimmer ships a DFlash drafter (
meta-models/Muse-Glimmer-30B-assistant; 5 layers, 16-token blocks; arXiv:2602.06036) that reads target hidden features at layers {1, 13, 25, 37, 49} of 52. The target verifies each block in one forward, which changes the rollout forward shape and needs cache rollback. This RFC runs strict mode with speculation off. The follow-up would add:The Phase A contracts (§4.1, and the window cache and eviction rule in
local_gated_attention) are written so the follow-up only needs new rows, not revised contracts.Also out of scope
If you are interested, just ping below!