Skip to content

Notes/qwen3 30b a3b fp8 turbo campaign - #1246

Draft
Xiaoming-AMD wants to merge 14 commits into
mainfrom
notes/qwen3-30b-a3b-fp8-turbo-campaign
Draft

Xiaoming-AMD wants to merge 14 commits into
mainfrom
notes/qwen3-30b-a3b-fp8-turbo-campaign

Conversation

@Xiaoming-AMD

Copy link
Copy Markdown
Collaborator

Description

Qwen3-30B-A3B FP8 pretraining on 8× MI355X goes from a memory fault at iteration 3 to 40,283 tokens/s/GPU (6507.6 ms/iter), 47% above the BF16 config. This PR switches the FP8 example config to Primus-Turbo tensorwise FP8 and turns on every Turbo fusion that got there. Two of those fusions are new in this PR:

  • head_dim 128 for the fused QKV split + Q/K RMSNorm + RoPE patch. The patch was GPT-OSS-only before.
  • turbo_fp8_permute: the DeepEP permute hands the fused grouped MLP FP8 tokens, and takes its gradient back in FP8.

The PR also fixes enable_turbo_attention_float8, which failed on its first call.

Primus-Turbo requirements.

  • turbo_fp8_permute needs a Primus-Turbo whose moe_permute takes quantize_dtype (AMD-AGI/Primus-Turbo#555). With an older Turbo, the first dispatch fails.
  • The head_dim 128 fusion needs AMD-AGI/Primus-Turbo#556. Without it, the patch keeps the unfused path.
  • The other Turbo rows in the table below come from #554, #557 and #558. They need no Primus change.

End-to-end. Qwen3-30B-A3B FP8 pretrain on 8× MI355X (EP8, MBS 8, GBS 512, seq 4096, even routing), mean of iterations 11–20 of 20-iteration runs. Each row adds one change on top of the row above:

run from ms/iter TFLOP/s/GPU tokens/s/GPU mem loss@20
BF16 config (legacy grouped GEMM) 9587 622.1 27,344
FP8 config before this PR memory fault at iteration 3
FP8 tensorwise config (Turbo e2d9f1d7) this PR 8894.0 670.6 29,474 277 GB 11.33393
+ use_turbo_fused_act_with_probs this PR 7915.6 753.5 33,117 223 GB 11.33435
Turbo 1103b2df Turbo main 7925.4 752.5 33,077 222 GB 11.33427
+ flat tensorwise FP8 quant kernel Turbo #558 7870.9 757.7 33,305 222 GB 11.33415
+ per-token get_dispatch_layout Turbo #557 7774.2 767.2 33,720 222 GB 11.33431
+ HIP permute in DeepEPTokenDispatcher Turbo #554 7643.5 780.3 34,296 222 GB 11.33423
+ turbo_deepep_num_cu 80 → 160 this PR 7353.4 811.1 35,649 223 GB 11.33456
+ fused QKV split + Q/K RMSNorm + RoPE, head_dim 128 this PR + Turbo #556 6987.1 853.6 37,518 224 GB 11.33373
+ turbo_fused_grouped_gemm (mean of 2 runs) this PR 6817.7 874.9 38,451 225 GB 11.33412
+ turbo_fp8_permute (mean of 4 runs) this PR + Turbo #555 6507.6 916.7 40,283 223 GB 11.33404

From the FP8 tensorwise config to the last row: −26.8% ms/iter, +36.7% tokens/s/GPU. Loss differences are within run-to-run noise (about 4e-4 at iteration 20). This branch, rebased on main 0a68e5c, with the switches on: 6518.7 ms/iter, 40,214 tokens/s/GPU, loss@20 11.33415.

KNOWN_ISSUE

Changes

  • fix(turbo): make enable_turbo_attention_float8 runnable again
    • primus/backends/megatron/core/extensions/primus_turbo.py:
      • Pass sink= only when a sink tensor exists; flash_attn_fp8_func has no sink parameter.
      • Make q/k/v bshd-contiguous on the FP8 path. flash_attn_fp8_func permutes sbhd storage as if it were the logical layout (fixed in Turbo by fix/fp8-attn-strided-layout).
  • perf(qwen3): tune Qwen3-30B-A3B FP8 MI355X config to Turbo tensorwise
    • examples/megatron/configs/MI355X/qwen3_30B_A3B-FP8-pretrain.yaml: fp8_recipe: tensorwise, use_turbo_gemm / use_turbo_grouped_gemm, and the TE grouped-MLP path instead of the legacy grouped GEMM.
  • perf(megatron): enable the fused packed-QKV RMSNorm + RoPE patch for head_dim 128
    • primus/backends/megatron/patches/turbo/qk_rmsnorm_rope_patches.py:
      • Accept any head_dim in Turbo's QK_RMSNORM_ROPE_HEAD_DIMS, falling back to (64,) on older Turbo.
      • Accept TE RMSNorm as well as PrimusTurboRMSNorm. The fused path only reads weight / eps; zero_centered_gamma is still rejected.
    • Fused QKV split + Q/K RMSNorm + RoPE at the Qwen3 attention shape: fwd+bwd 1243.6 → 369.9 µs (3.36×).
  • perf(megatron): hand the Turbo fused grouped MLP FP8 tokens from the DeepEP permute
    • New opt-in flag turbo_fp8_permute (default false in primus_turbo.yaml).
    • Under Turbo FP8 tensorwise current scaling, PrimusTurboDeepEPTokenDispatcher passes quantize_dtype to _post_dispatch and grad_quantize_dtype to _pre_combine. The dtypes follow the FP8 format (HYBRID: e4m3 input, e5m2 gradient), and results are bit-identical to the default path. Any other recipe, or a layer with Turbo FP8 off, keeps the bf16 path.
    • Argument validation requires enable_primus_turbo, use_turbo_deepep and turbo_fused_grouped_gemm, and rejects selective recompute of moe_act.
    • Dispatcher kernel time per call at the Qwen3 EP8 shape: forward 645 → 167 µs, backward 687 → 174 µs.
  • perf(qwen3): turn on the Turbo fusions in the Qwen3-30B-A3B FP8 MI355X config
    • Top-level env: PRIMUS_FUSED_QK_RMSNORM_ROPE: "1".
    • turbo_fused_grouped_gemm, use_turbo_fused_act_with_probs and turbo_fp8_permute set to true.
    • turbo_deepep_num_cu 80 → 160. 192 gave a NaN loss on one rank at iteration 5; 256 is slower (7028 ms).

The force-"even" routing fix that this work also needed is already on main (#1243), so it is not part of this PR.

Tests

  • New in tests/unit_tests/backends/megatron/test_rocm_arg_validation.py: validate_turbo_fp8_permute (each required flag, selective recompute of moe_act, disabled no-op).
  • New tests/unit_tests/backends/megatron/test_turbo_fp8_permute_dtypes.py: _fp8_permute_dtypes returns (e4m3, e5m2) for HYBRID and (e4m3, e4m3) for E4M3 tensorwise, and (None, None) with the flag off, Turbo FP8 off or blockwise scaling.
  • test_qk_rmsnorm_rope_patches.py, test_rocm_arg_validation.py, test_validate_args_patches.py, test_router_force_even_routing.py, test_turbo_fp8_permute_dtypes.py: all pass on MI355X.
  • End-to-end runs above.

…ripts

Handoff for the 8x MI355X Qwen3-30B-A3B FP8 tensorwise campaign
(29,474 -> 40,283 tokens/s/GPU): cumulative e2e table, the Primus switches and
Turbo/Primus branches the final stack needs, PR order and ready-to-paste PR
bodies, how to rebuild the stack, rejected options and open items, plus the
round log, the microbenchmark / trace / e2e scripts and the task prompt.
Copilot AI balanced review requested due to automatic review settings October 10, 2026 15:35
print("|---|---|---|---|---|---|---|")
for name in sys.argv[1:]:
rows = []
for line in open(f"/tmp/q3_e2e_{name}.log", errors="ignore"):

gout = torch.randn(T, H, device=dev, dtype=torch.bfloat16)

def bwd(qd):
for backend in (BackendType.TURBO, BackendType.TRITON):
for fill, idx_dtype in ((7.0, torch.int32), (7.0, torch.int64), (-3.0, torch.int64)):
torch.cuda.empty_cache()
junk = torch.full((rows * H,), fill, device="cuda") # noqa: F841 (poison the allocator)
K.flydsl_qkv_rmsnorm_rope_backward(
torch.randn_like(q), torch.randn_like(k), torch.randn_like(v), qkv, qg, kg, freqs, qr, kr, split
)
except TypeError:

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The advertised Primus implementation is absent, and several validation utilities can report misleading success.

5 open findings
What changed in this PR

Adds an archival handoff for the Qwen3 FP8 optimization campaign, including benchmarks, diagnostics, run orchestration, results, and draft Turbo PR descriptions. It does not include the Primus implementation advertised in the PR description.

Changes:

  • Documents performance findings and required optimization branches.
  • Adds profiling, benchmarking, and correctness utilities.
  • Adds end-to-end experiment scripts and draft upstream PR bodies.
File Description
README.md Campaign results and handoff
prompt.md Original optimization requirements
trace_shapes.py Shape-based kernel attribution
trace_kernels.py Kernel-time aggregation
trace_attrib.py CPU-stack attribution
qv_real.py Quantization microbenchmark
quant_bitexact.py Byte-exact comparison utility
quant_variants/​run_variants.py Quantization variant runner
quant_variants/​quant_variants.hip HIP quantization variants
quant_variants/​quant_variants_hip.hip Alternate HIP launch variant
pr_bodies/​perf-quantization-tensorwise-fp8-qwen3.md Quantization PR draft
pr_bodies/​perf-moe-permute-default-hip.md HIP permute PR draft
pr_bodies/​perf-moe-fp8-permute-tensorwise.md FP8 permute PR draft
pr_bodies/​perf-flydsl-qk-rmsnorm-rope-hd128.md Fused RoPE PR draft
pr_bodies/​perf-deep-ep-dispatch-layout-per-token.md DeepEP layout PR draft
pr_bodies/​fix-moe-permute-padding-rows.md Padding fix PR draft
pr_bodies/​fix-fp8-attn-strided-layout.md Attention layout PR draft
e2e/​parse_e2e.py Run-log summarizer
e2e/​env_turbo_opt.sh Optimized Turbo environment
e2e/​env_turbo_opt_qknorm.sh Fused-QK environment
e2e/​env_turbo_base.sh Baseline Turbo environment
e2e/​e2e_run.sh Single-run launcher
e2e/​e2e_queue5.sh Fused-MLP experiment queue
e2e/​e2e_queue4.sh DeepEP CU experiment queue
e2e/​e2e_queue3.sh Profiling queue
e2e/​e2e_queue2.sh MegaMoE experiment queue
e2e/​e2e_chain.sh Attribution experiment chain
e2e/​e2e_ab.sh Sequential A/B runner
e2e/​amend_msg.py Commit-message result generator
debug_fp8_worst.py Worst-token FP8 reproducer
dbg_rope128b.py RoPE backward-write debugger
dbg_rope128.py RoPE gradient debugger
dbg_rope128_isa.py RoPE ISA compiler helper
dbg_permute_probs.py Permuted-probability debugger
dbg_aiter_fwd.py AITER attention benchmark
bench_quant_chain.py Quantization-chain benchmark
bench_qk_rmsnorm_rope.py Fused QK/RoPE benchmark
bench_permute_qwen3.py Permute backend benchmark
bench_grouped_gemm_qwen3.py Grouped-GEMM benchmark
bench_fp8_permute.py FP8 permute validation
bench_dense_layouts.py Dense GEMM layout benchmark
bench_deepep_qwen3.py DeepEP benchmark
bench_attention_qwen3.py Attention backend benchmark

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +28 to +32
The final configuration is `examples/megatron/configs/MI355X/qwen3_30B_A3B-FP8-pretrain.yaml` plus:
- `--use_turbo_fused_act_with_probs True`
- `--turbo_deepep_num_cu 160`
- `--turbo_fused_grouped_gemm True`
- `--turbo_fp8_permute True`
(cd /home/xiaompen/Primus-qwen3-qknorm && bash runner/primus-cli direct --env "$E2E/env_turbo_opt_qknorm.sh" -- \
train pretrain --config "$CFG" --train_iters 20 --use_turbo_fused_act_with_probs True \
--turbo_deepep_num_cu 160 "$@") > "/tmp/q3_e2e_${name}.log" 2>&1
echo "[e2e] $name exit=$? $(date +%T)"
Comment on lines +19 to +40
rows = []
for line in open(f"/tmp/q3_e2e_{name}.log", errors="ignore"):
m = ITER.search(line)
if m and int(m.group(1)) > 10:
t, loss, mem = TFLOPS.search(line), LOSS.search(line), MEM.search(line)
rows.append(
(
float(m.group(2)),
float(t.group(1)) if t else 0.0,
loss.group(1) if loss else "-",
float(mem.group(1)) if mem else 0.0,
)
)
if not rows:
print(f"| {name} | no iterations |")
continue
ms = statistics.mean(r[0] for r in rows)
print(
f"| {name} | {ms:.1f} | {statistics.pstdev(r[0] for r in rows):.1f} | "
f"{statistics.mean(r[1] for r in rows):.1f} | {TOKENS_PER_GPU_ITER / ms * 1e3:,.0f} | "
f"{max(r[3] for r in rows):.1f} | {float(rows[-1][2]):.5f} |"
)
Comment on lines +41 to +48
ref = torch.load(PATH)
bad = 0
for c, (ry, rs) in zip(CASES, ref):
y, s = run(c)
if not (torch.equal(y, ry) and torch.equal(s, rs)):
bad += 1
print("MISMATCH", c)
print(f"checked {len(CASES)} cases, {bad} mismatches")
def stack_at(tid, ts, depth=4):
ops = ops_by_tid.get(tid, [])
i = bisect.bisect_right(starts.get(tid, []), ts)
chain = [name for (s, e, name) in ops[max(0, i - 400) : i] if s <= ts <= e]
Copilot AI balanced review requested due to automatic review settings October 10, 2026 15:42

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The claimed implementation is absent, the selected configuration has unresolved NaN failures, and multiple runners mask failed training exits.

7 open findings
Previously missed (2)

In code that hasn't changed since last review

Medium severity Benchmark continues after failed training due to masked exit status

agent/​workspace/​qwen3_fp8_tw_gfx950_20261010/​e2e/​e2e_ab.sh:15

run always returns the successful status of this echo, so the && attribution chain continues and can report completion after a failed training run. Capture the command status before logging and abort the script on failure so later benchmark rows are not attributed to an invalid base.

Low severity Unstable optimization is enabled by default despite NaN loss

agent/​workspace/​qwen3_fp8_tw_gfx950_20261010/​README.md:118

The final configuration should not be enabled by default while identical runs still produce NaN loss (1/13 at the selected 160-CU setting). This is a training-correctness failure, not benchmark jitter; keep the unstable optimization opt-in or revert to the known-stable setting until the proposed stress test isolates the cause and a regression test covers it.

🧠 Review effort: Balanced

echo "[e2e] $name start $(date +%T)"
(cd /home/xiaompen/Primus-qwen3-pr && bash runner/primus-cli direct --env "$E2E/env_turbo_opt_cfg.sh" -- \
train pretrain --config "$CFG" --train_iters 20 "$@") > "/tmp/q3_e2e_${name}.log" 2>&1
echo "[e2e] $name exit=$? $(date +%T)"
train pretrain --config "$CFG" --train_iters 20 --use_turbo_fused_act_with_probs True \
--turbo_deepep_num_cu 160 --turbo_fused_grouped_gemm True --turbo_fp8_permute True "$@") \
> "/tmp/q3_e2e_${name}.log" 2>&1
echo "[e2e] $name exit=$? $(date +%T)"
Copilot AI balanced review requested due to automatic review settings October 10, 2026 15:47

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The advertised production changes are absent, and several validation/reporting utilities can return incorrect results or fail on supported inputs.

8 open findings
Previously missed (1)

In code that hasn't changed since last review

Medium severity Missing loss sentinel causes summary conversion failure

agent/​workspace/​qwen3_fp8_tw_gfx950_20261010/​e2e/​parse_e2e.py:40

The parser records "-" when an iteration line has no loss, but then unconditionally converts that sentinel to float, causing the summary command to fail instead of reporting the missing value. Format the loss conditionally, as already done conceptually for absent fields.

🧠 Review effort: Balanced

Comment on lines +3 to +5
Qwen3-30B-A3B FP8 pretraining on 8× MI355X goes from a memory fault at iteration 3 to **40,283 tokens/s/GPU** (6507.6 ms/iter), 47% above the BF16 config. This PR switches the FP8 example config to Primus-Turbo tensorwise FP8 and turns on every Turbo fusion that got there. Two of those fusions are new in this PR:
- **head_dim 128 for the fused QKV split + Q/K RMSNorm + RoPE patch.** The patch was GPT-OSS-only before.
- **`turbo_fp8_permute`**: the DeepEP permute hands the fused grouped MLP FP8 tokens, and takes its gradient back in FP8.
Copilot AI balanced review requested due to automatic review settings October 11, 2026 05:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

- `perf(qwen3): turn on the Turbo fusions in the Qwen3-30B-A3B FP8 MI355X config`
- Top-level `env: PRIMUS_FUSED_QK_RMSNORM_ROPE: "1"`.
- `turbo_fused_grouped_gemm`, `use_turbo_fused_act_with_probs` and `turbo_fp8_permute` set to true.
- `turbo_deepep_num_cu` 80 → 160. 192 gave a NaN loss on one rank at iteration 5; 256 is slower (7028 ms).
Copilot AI balanced review requested due to automatic review settings October 11, 2026 06:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +63 to +65
| `perf/qwen3-30b-a3b-fp8-turbo` | bef41e8a | opening | the one Primus PR, on main 0a68e5cd; body: [pr_bodies/primus-perf-qwen3-30b-a3b-fp8-turbo.md](pr_bodies/primus-perf-qwen3-30b-a3b-fp8-turbo.md) |

It holds five commits: the FP8 attention fix, the tensorwise config, the head_dim 128 qk-norm patch, `turbo_fp8_permute` (with new unit tests), and the config turning every switch on. It replaces `perf/qwen3-30b-a3b-tuning`, `perf/qwen3-qk-rmsnorm-rope-hd128` and `perf/megatron/turbo-fp8-permute`. Their force-"even" routing fix is already on main (#1243).
Record the Primus DISABLE_CHEAP_FENCE config fix and the Turbo high-CU
fence default branch for the intermittent forward-loss NaN.

Co-authored-by: Cursor <cursoragent@cursor.com>
Copilot AI balanced review requested due to automatic review settings October 11, 2026 09:06
Co-authored-by: Cursor <cursoragent@cursor.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The claimed production changes are absent, and several validation benchmarks can report misleading results.

10 open findings
Previously missed (2)

In code that hasn't changed since last review

Medium severity Clear leaf gradients before timing backward operations

agent/​workspace/​qwen3_fp8_tw_gfx950_20261010/​bench_permute_qwen3.py:82

The timed callbacks reuse leaf tensors without clearing their gradients, so after warmup every measured backward includes large AccumulateGrad additions into existing buffers. This inflates the reported permute/unpermute backward timings independently of the backend; use autograd.grad to measure only the operation's backward path.

Low severity Update GLM5 documentation to match the executed batch size

agent/​workspace/​qwen3_fp8_tw_gfx950_20261010/​bench_qk_rmsnorm_rope_ab.py:8

The documented GLM5 shape says B=4, but the benchmark actually runs B=8 in SHAPES. This makes the advertised workload and any comparison based on it ambiguous; update the docstring to match the executed shape.

🧠 Review effort: Balanced

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants