Skip to content

feat(models): support Qwen Flash Next vision and chunked training - #1751

Merged
sitabulaixizawaluduo merged 36 commits into
mainfrom
feature/qwen-flash-next-model
Oct 8, 2026
Merged

sitabulaixizawaluduo merged 36 commits into
mainfrom
feature/qwen-flash-next-model

Conversation

@dingzhiqiang

@dingzhiqiang dingzhiqiang commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Support Qwen Flash Next through the official ModelScope bridge, vision/THD/mRoPE paths, chunked loss, and preservation of frozen visual/PLE state during AWEX exchange.

Generated image/video special tokens remain text in mixed multimodal batches, including after PPO/MOPD and SFT train/eval shift the loss mask. Preserve a separate input-aligned mask for this provenance. Carry full packed modality provenance through CP and mask only the bridge visual-scatter ID copy, preserving original token embeddings, labels and gradients. Qwen weight converters, tensor layouts and frozen declarations live in AWEX; AReaL supplies live engine binding and checkpoint-evidence validation.

Keep two example configurations: one SFT smoke recipe and one multimodal Arena RL recipe that also accepts text SWE-bench Verified streams. Evaluation reuses the MM configuration in swe-eval mode: zero training steps, recovery disabled, one sample per task, and task count derived from the explicit manifest. Runtime compatibility and diagnostic helpers remain where required.

New regression tests are organized into three focused files: training/bridge, AWEX integration, and vision/runtime. They retain CP layout, TP/PP transfer reconstruction, frozen-state lifecycle, upstream mRoPE comparisons and PLE chunk forward/gradient checks. Independent GDN/QKV coordinate tests move with the converter implementation to AWEX. These checks do not establish full-model LM-head chunk/full numerical parity.

PLE CP hidden-state reconstruction uses an autograd all-gather, so causal-convolution input gradients include remote CP consumers. The override is scoped to the pinned bridge's PLE binding; ID/vision gathers and SP handling are unchanged.

Vision initialization requires the Qwen4Exp mRoPE methods, validated with Transformers 5.16.1. Standard locked Transformers 5.3.0/5.7.0 environments do not provide this model; use a compatible training runtime. AWEX 0.8.2 is declared in the project dependencies and lockfiles; no temporary AWEX installation is added to the unit-test workflow. Dependency declarations and lockfiles differ from main only by the AWEX 0.8.2 upgrade. The pre-commit workflow is unchanged from main.

Dependencies and comparison

Stacked on #1750; #1749 is already merged into main. This branch includes main at 0fa9f62 and the current #1750 parent at 7130323; retarget after #1750 merges. AWEX dependency: released awex==0.8.2 in both declarations and lockfiles, including merged inclusionAI/Awex#121. Transformers 5.16.1 and the generated-vision-token bridge fix at dingzhiqiang/mcore-bridge@557aaf93b16d083fdec4f82a8251d47d47c76ccb (based on ModelScope bc58ea9) are supplied through training PYTHONPATH, not AReaL dependencies or dependency groups. Set QWEN_TRAIN_EXTRA_PYTHONPATH and MCORE_BRIDGE_ROOT to shared, prepared runtime directories. Submission validates the bridge checkout and propagates training paths to the controller and actor workers; rollout retains its own image interpreter and inference paths. No package installation or virtualenv creation runs during submission.

Split from #1731 and reviewed against feature/qwen38-flash-next-main@856703760. Core model/bridge behavior is retained along with newer main fixes. Fork-specific startup checks and fixed deployment/benchmark artifacts are excluded; generic rank validation lives in #1749. The stack now merges main at 3c4be16bed3ab50bd549bc0900d23e952b9d2d21, including #1746 generation defaults and episode/session/turn/harness metrics. #1748 is also in main.

Validation

  • AWEX 0.8.2 release: 110 targeted tests passed; both CI installer variants verified the released package. Both lockfiles pass offline validation, with all other package records unchanged.
  • Latest review fixes: 155 targeted tests passed with actual Transformers 5.16.1 and bridge bc58ea9; 2 CUDA-only MOPD tests skipped on CPU. Includes all three Qwen files, PPO GAE, AWEX device IDs, and Megatron MOPD teacher tests.
  • Two real Gloo CP ranks exercise the installed PLE reconstruction binding and causal convolution: outputs, input gradients and summed weight gradients match CP1; the original gather fails the input-gradient comparison. This covers BSHD, not GPU THD or combined TP/SP.
  • After merging main: 159 targeted training/Arena/grouped-rollout/proxy tests passed, 2 optional-runtime skips. Includes training-only timeout defaults and preservation of explicit overrides; evaluation no longer inherits the training timeout.
  • Consolidated tests plus related existing model, Arena and proxy regressions: 557 passed, 9 optional-runtime skips in the cached development environment.
  • AWEX transport and migrated layout tests: 41 passed.
  • Full pre-commit and normal commit hooks passed; shell syntax and deleted-file references checked.
  • The official AWEX 0.8.2 wheel matches the previous pinned commit across all 100 Python files except the version string.

Remaining validation before merge

GPU THD/TP/SP gradient parity, NCCL/IPC, repeated weight equality, optimizer recovery and real nonzero-gradient RL still require combined validation. Schema 2 frozen-manifest generation tooling is not supplied; recipes require a validated manifest. CPU tests and the original zero-gradient replay are not proof of effective learning. Full docs-site build not run.

Type of Change

  • New feature

Checklist

  • Pre-commit hooks pass (pre-commit run --all-files)
  • Relevant targeted CPU tests pass
  • Self-reviewed against the pre-split snapshot
  • Created by a coding agent using the create-pr workflow
  • GPU integration and effective RL learning validated
  • Frozen-manifest generation workflow complete

Latest review fixes and validation

  • Keep rollout QSA patching and RPC startup on the rollout image interpreter. Stage live bridge exports on the sender device with contiguous layout, preserving dtype and existing bucket lifetime before NCCL synchronization.
  • After merging main, reject unsupported ModelScope-bridge MTP training early; the supplied recipes already disable it.
  • Full pre-commit and commit hooks passed. CPU regression suites: 99 passed / 3 skipped and 212 passed / 17 skipped (overlapping suites; not additive). GPU/NCCL, multi-node integration and actual RL learning were not validated in this review.

Runtime setup simplified in 36dc981: removed Qwen dependency groups, scoped overrides and the CI uv upgrade. Only AWEX changes in dependency declarations/locks relative to main. Six submission/startup regression tests passed; full pre-commit and commit hooks passed. The earlier model CPU suites remain applicable; GPU/multi-node validation has not been rerun.

Vision tensor memory fix

Share byte-identical CPU vision tensors across repeated images in multi-turn trajectories, while preserving context-dependent processor outputs and per-turn token masks. Added regression coverage for changed image content, processor identity, and trajectory assembly.

Validation: 53 vision/trajectory/streaming tests passed; full pre-commit passed. The separate NUMA, gradient-finalization cache, and checkpoint host-cache fixes are tracked in #1756.

@dingzhiqiang dingzhiqiang added the safe-to-test Ready to run unit-tests in a PR. label Sep 22, 2026
Comment thread areal/engine/megatron_utils/qwen4_exp_mrope.py Outdated
Comment thread areal/engine/megatron_utils/qwen4_exp_mrope.py Outdated
Comment thread examples/swe/qwen38_flash_next/swe_mm_rl.yaml
chaokunyang pushed a commit to inclusionAI/Awex that referenced this pull request Sep 22, 2026
## Description
Keep transfer planning as tensor views and allocate contiguous
send/receive buffers only for the active batch, including strided column
slices. Copy receives back before releasing that batch. Pack compatible
expert operations within configured bounds, check packing limits across
inference ranks before P2P starts, and provide CUDA IPC staging
allocation that restores the allocator settings used for training. The
packing byte limit is not a hard cap for a single oversized tensor.

Add a reader factory so integrations can explicitly choose the bounded
transport; the AWEX default remains unchanged. All participating readers
must choose the same transport. This extracts the shared transport
implementation from areal-project/AReaL#1731.

Own Qwen4Exp weight converters, GDN/gated-QKV layouts and
frozen-parameter declarations in AWEX. Registration is explicit and
accepts caller-supplied binders; checkpoint evidence and live engine
ownership stay in the integration. AReaL consumes these APIs in
areal-project/AReaL#1751 without a second converter implementation.

## Validation
- 41 targeted CPU tests passed, including real two-rank Gloo
agreement/disagreement, strided buffer allocation/release and
independent GDN/QKV coordinate checks across TP sizes.
- AReaL integration regression suite: 557 passed, 9 optional-runtime
skips; separate Transformers 5.16.1 vision run: 31 passed.
- Ruff check and format passed for changed files.
- GPU NCCL transfers, repeated weight equality, and cross-image CUDA IPC
remain unverified and require validation before merge.
Comment thread examples/swe/qwen38_flash_next/swe_mm_rl.yaml Outdated
Comment thread areal/models/mcore/mcore_bridge_adapter.py
Comment thread examples/swe/qwen38_flash_next/gdn_cp_compat.py
Reuse byte-identical CPU vision tensors for repeated images while
preserving context-dependent processor output and distinct token masks.
Cover changed image content, processor identity, and trajectory assembly.
fishcrap
fishcrap previously approved these changes Sep 24, 2026

@fishcrap fishcrap left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

The current dev CI images contain AWEX 0.8.1, so tests importing the new cuda_ipc module fail during collection and block the rest of the suite. Remove only those direct AWEX 0.8.2 tests until the standard images catch up.
Base automatically changed from feat/group-reward-mean-only to main September 26, 2026 01:05
@sitabulaixizawaluduo
sitabulaixizawaluduo dismissed fishcrap’s stale review September 26, 2026 01:05

The base branch was changed.

Bring the PR up to date with main while preserving its Arena gateway
renewal. Adopt the stricter Harness failure classification and tests
that landed in main.

Key changes:
- Retain Arena registration renewal and its tests
- Use main's stricter Harness failure admission
- Include main's Megatron and CUDA fixes

Refs: #1751
@dingzhiqiang dingzhiqiang added safe-to-test Ready to run unit-tests in a PR. and removed safe-to-test Ready to run unit-tests in a PR. labels Sep 26, 2026
@sitabulaixizawaluduo sitabulaixizawaluduo added safe-to-test Ready to run unit-tests in a PR. and removed safe-to-test Ready to run unit-tests in a PR. labels Sep 27, 2026
@sitabulaixizawaluduo
sitabulaixizawaluduo merged commit 6fd6652 into main Oct 8, 2026
22 of 30 checks passed
@sitabulaixizawaluduo
sitabulaixizawaluduo deleted the feature/qwen-flash-next-model branch October 8, 2026 02:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

safe-to-test Ready to run unit-tests in a PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants