Skip to content

feat(ascend): add MXFP8 rollout support - #448

Open
PierceZhou wants to merge 3 commits into
vllm-project:ascendfrom
PierceZhou:feature/mxfp8
Open

PierceZhou wants to merge 3 commits into
vllm-project:ascendfrom
PierceZhou:feature/mxfp8

Conversation

@PierceZhou

@PierceZhou PierceZhou commented Sep 24, 2026 •

Copy link
Copy Markdown

Motivation

Add MXFP8 rollout support for the Ascend backend.

This enables vime to support rollout with MXFP8 models and improves compatibility with MXFP8 weight update workflows.

Changes

  • Add MXFP8 support in weight update path
  • Update tensor-based weight update logic to handle MXFP8 format
  • Update distributed weight update handling for MXFP8 rollout
  • Adapt vLLM engine weight transfer configuration for Ascend backend

Implementation Details

  • MXFP8-related weight synchronization is handled in the rollout weight update path.
  • The existing weight transfer flow is preserved while extending support for MXFP8 models.
  • Ascend-specific backend handling is maintained.

Test

Tested with Ascend NPU environment.

Checklist

  • Code format checked
  • Tested on Ascend backend

Integration evaluation (Ascend NPU)

Related quantization implementation: vllm-ascend#17644. These BF16/MXFP8 measurements exercise the integrated vIME rollout and weight-update workflow with online quantization in vLLM-Ascend.

Workload BF16 MXFP8 Change
Qwen3-8B, GSM8K-500 rollout throughput (tokens/GPU/s) 649.5 805.7 +24.0%
Qwen3-8B, GSM8K-500 rollout time (s) 46.6 39.7 -14.8%
Qwen3-8B, GSM8K-500 step time (s) 129.3 122.8 -5.0%
Qwen3-8B, GSM8K-500 pass@1 92.7% 91.7% -1.0 percentage point

Five short Qwen3-8B/32B training benchmarks showed 6%-11% lower total wall time with MXFP8 rollout. In the GSM8K-500 comparison, pass@4 changed from 95.0% to 93.8%, mis_kl from 0.0256 to 0.0349, and normalized ESS from 0.985 to 0.954. The math benchmark showed a larger quality trade-off: flexible extraction accuracy fell by about 6% relative for both Qwen3-8B and Qwen3-32B, despite higher inference throughput.

A 400-step validation-mode run exercised rollout -> train -> weight_update over 8 epochs. It used perform_rl_step=False, so it demonstrates execution of the workflow rather than policy improvement. These integrated measurements complement the PR's unit tests and CI; they are not a standalone measurement of the metric instrumentation added in this diff.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces comprehensive performance metrics tracking for weight updates across both distributed and tensor-based backends, updates the weight transfer backend configuration to support Ascend (NPU) specific backends (npu_ipc and hccl), and aligns the weight update lifecycle with the vLLM 0.27 API. Additionally, it adds an address guard for online MXFP8 reloads to verify that graph-captured tensor storage is preserved. The reviewer feedback highlights a critical bug in update_weight_from_tensor.py where calculating tensor sizes can exhaust a lazy generator before it is sent, as well as timing metric inaccuracies in both backends due to lazy evaluation occurring outside timed blocks. Finally, the reviewer advises against leaving commented-out code in the megatron bridge initialization without explanation.

Comment on lines +183 to +189
while True:
export_started = time.perf_counter()
try:
hf_named_tensors = next(iterator)
except StopIteration:
break
export_seconds += time.perf_counter() - export_started

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

If the iterator returned by get_hf_weight_chunks yields generators or lazy iterables, calling sum(...) on hf_named_tensors at line 190 will exhaust the generator. This leaves hf_named_tensors empty when passed to _send_hf_params at line 194, silently failing the weight update. Additionally, the time spent converting/exporting the weights is currently excluded from export_seconds because the lazy evaluation happens during sum(...) outside the timed block. Converting next(iterator) to a list inside the timed block resolves both the generator exhaustion bug and the timing inaccuracy.

Suggested change
while True:
export_started = time.perf_counter()
try:
hf_named_tensors = next(iterator)
except StopIteration:
break
export_seconds += time.perf_counter() - export_started
while True:
export_started = time.perf_counter()
try:
hf_named_tensors = list(next(iterator))
except StopIteration:
break
export_seconds += time.perf_counter() - export_started

Comment on lines 238 to 246
while True:
export_started = time.perf_counter()
try:
hf_named_tensors = next(iterator)
except StopIteration:
break
self._perf_export_seconds += time.perf_counter() - export_started
if self._is_pp_src_rank:
hf_named_tensors = list(hf_named_tensors)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If get_hf_weight_chunks yields generators, the actual weight conversion/export work is deferred until the generator is evaluated (e.g., via list(hf_named_tensors)). Currently, list(hf_named_tensors) is called on line 246, which is outside the timed block for self._perf_export_seconds. This causes the export time metric to be highly inaccurate (underreported). Moving the list(...) conversion inside the timed block (only for the PP source rank to preserve the optimization on other ranks) ensures accurate performance metrics.

        while True:
            export_started = time.perf_counter()
            try:
                hf_named_tensors = next(iterator)
            except StopIteration:
                break
            if self._is_pp_src_rank:
                hf_named_tensors = list(hf_named_tensors)
            self._perf_export_seconds += time.perf_counter() - export_started

@@ -1 +1 @@
import vime_plugins.megatron_bridge.glm4v_moe # noqa: F401 # register GLM-4.6V bridge
#import vime_plugins.megatron_bridge.glm4v_moe # noqa: F401 # register GLM-4.6V bridge

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Commenting out code instead of removing it or providing an explanatory comment/TODO is a maintainability anti-pattern. If the GLM-4.6V bridge registration is no longer needed, please delete this line entirely. If it is temporarily disabled, please add a TODO or an explanatory comment explaining why it is commented out and when it should be re-enabled.

@read-the-docs-community

read-the-docs-community Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Signed-off-by: PierceZhou <1342578551@qq.com>
Signed-off-by: PierceZhou <1342578551@qq.com>
Signed-off-by: PierceZhou <1342578551@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant