Repository navigation
feat(models): support Qwen Flash Next vision and chunked training - #1751
Merged
Merged
Conversation
dingzhiqiang
marked this pull request as ready for review
September 22, 2026 11:11
dingzhiqiang
requested review from
CormickKneey,
HwVanICI,
Le8r0nJames,
PrometheusComing,
TaoZex,
fishcrap,
garrett4wade,
geshi001,
guozhihao-224,
nuzant,
rchardx and
sitabulaixizawaluduo
as code owners
September 22, 2026 11:11
dingzhiqiang
had a problem deploying
to
AReaL-unittests
September 22, 2026 12:00 — with
GitHub Actions
Failure
dingzhiqiang
had a problem deploying
to
AReaL-unittests
September 22, 2026 12:00 — with
GitHub Actions
Failure
chaokunyang
pushed a commit
to inclusionAI/Awex
that referenced
this pull request
Sep 22, 2026
## Description Keep transfer planning as tensor views and allocate contiguous send/receive buffers only for the active batch, including strided column slices. Copy receives back before releasing that batch. Pack compatible expert operations within configured bounds, check packing limits across inference ranks before P2P starts, and provide CUDA IPC staging allocation that restores the allocator settings used for training. The packing byte limit is not a hard cap for a single oversized tensor. Add a reader factory so integrations can explicitly choose the bounded transport; the AWEX default remains unchanged. All participating readers must choose the same transport. This extracts the shared transport implementation from areal-project/AReaL#1731. Own Qwen4Exp weight converters, GDN/gated-QKV layouts and frozen-parameter declarations in AWEX. Registration is explicit and accepts caller-supplied binders; checkpoint evidence and live engine ownership stay in the integration. AReaL consumes these APIs in areal-project/AReaL#1751 without a second converter implementation. ## Validation - 41 targeted CPU tests passed, including real two-rank Gloo agreement/disagreement, strided buffer allocation/release and independent GDN/QKV coordinate checks across TP sizes. - AReaL integration regression suite: 557 passed, 9 optional-runtime skips; separate Transformers 5.16.1 vision run: 31 passed. - Ruff check and format passed for changed files. - GPU NCCL transfers, repeated weight equality, and cross-image CUDA IPC remain unverified and require validation before merge.
fishcrap
reviewed
Sep 23, 2026
fishcrap
reviewed
Sep 23, 2026
6 tasks done
Reuse byte-identical CPU vision tensors for repeated images while preserving context-dependent processor output and distinct token masks. Cover changed image content, processor identity, and trajectory assembly.
dingzhiqiang
had a problem deploying
to
AReaL-unittests
September 24, 2026 06:42 — with
GitHub Actions
Failure
dingzhiqiang
had a problem deploying
to
AReaL-unittests
September 24, 2026 06:42 — with
GitHub Actions
Failure
The current dev CI images contain AWEX 0.8.1, so tests importing the new cuda_ipc module fail during collection and block the rest of the suite. Remove only those direct AWEX 0.8.2 tests until the standard images catch up.
dingzhiqiang
had a problem deploying
to
AReaL-unittests
September 24, 2026 07:20 — with
GitHub Actions
Failure
dingzhiqiang
had a problem deploying
to
AReaL-unittests
September 24, 2026 07:20 — with
GitHub Actions
Error
sitabulaixizawaluduo
dismissed
fishcrap’s stale review
September 26, 2026 01:05
The base branch was changed.
Bring the PR up to date with main while preserving its Arena gateway renewal. Adopt the stricter Harness failure classification and tests that landed in main. Key changes: - Retain Arena registration renewal and its tests - Use main's stricter Harness failure admission - Include main's Megatron and CUDA fixes Refs: #1751
sitabulaixizawaluduo
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Support Qwen Flash Next through the official ModelScope bridge, vision/THD/mRoPE paths, chunked loss, and preservation of frozen visual/PLE state during AWEX exchange.
Generated image/video special tokens remain text in mixed multimodal batches, including after PPO/MOPD and SFT train/eval shift the loss mask. Preserve a separate input-aligned mask for this provenance. Carry full packed modality provenance through CP and mask only the bridge visual-scatter ID copy, preserving original token embeddings, labels and gradients. Qwen weight converters, tensor layouts and frozen declarations live in AWEX; AReaL supplies live engine binding and checkpoint-evidence validation.
Keep two example configurations: one SFT smoke recipe and one multimodal Arena RL recipe that also accepts text SWE-bench Verified streams. Evaluation reuses the MM configuration in
swe-evalmode: zero training steps, recovery disabled, one sample per task, and task count derived from the explicit manifest. Runtime compatibility and diagnostic helpers remain where required.New regression tests are organized into three focused files: training/bridge, AWEX integration, and vision/runtime. They retain CP layout, TP/PP transfer reconstruction, frozen-state lifecycle, upstream mRoPE comparisons and PLE chunk forward/gradient checks. Independent GDN/QKV coordinate tests move with the converter implementation to AWEX. These checks do not establish full-model LM-head chunk/full numerical parity.
PLE CP hidden-state reconstruction uses an autograd all-gather, so causal-convolution input gradients include remote CP consumers. The override is scoped to the pinned bridge's PLE binding; ID/vision gathers and SP handling are unchanged.
Vision initialization requires the Qwen4Exp mRoPE methods, validated with Transformers 5.16.1. Standard locked Transformers 5.3.0/5.7.0 environments do not provide this model; use a compatible training runtime. AWEX 0.8.2 is declared in the project dependencies and lockfiles; no temporary AWEX installation is added to the unit-test workflow. Dependency declarations and lockfiles differ from main only by the AWEX 0.8.2 upgrade. The pre-commit workflow is unchanged from main.
Dependencies and comparison
Stacked on #1750; #1749 is already merged into main. This branch includes main at 0fa9f62 and the current #1750 parent at 7130323; retarget after #1750 merges. AWEX dependency: released
awex==0.8.2in both declarations and lockfiles, including merged inclusionAI/Awex#121. Transformers 5.16.1 and the generated-vision-token bridge fix atdingzhiqiang/mcore-bridge@557aaf93b16d083fdec4f82a8251d47d47c76ccb(based on ModelScope bc58ea9) are supplied through trainingPYTHONPATH, not AReaL dependencies or dependency groups. SetQWEN_TRAIN_EXTRA_PYTHONPATHandMCORE_BRIDGE_ROOTto shared, prepared runtime directories. Submission validates the bridge checkout and propagates training paths to the controller and actor workers; rollout retains its own image interpreter and inference paths. No package installation or virtualenv creation runs during submission.Split from #1731 and reviewed against
feature/qwen38-flash-next-main@856703760. Core model/bridge behavior is retained along with newer main fixes. Fork-specific startup checks and fixed deployment/benchmark artifacts are excluded; generic rank validation lives in #1749. The stack now merges main at3c4be16bed3ab50bd549bc0900d23e952b9d2d21, including #1746 generation defaults and episode/session/turn/harness metrics. #1748 is also in main.Validation
bc58ea9; 2 CUDA-only MOPD tests skipped on CPU. Includes all three Qwen files, PPO GAE, AWEX device IDs, and Megatron MOPD teacher tests.Remaining validation before merge
GPU THD/TP/SP gradient parity, NCCL/IPC, repeated weight equality, optimizer recovery and real nonzero-gradient RL still require combined validation. Schema 2 frozen-manifest generation tooling is not supplied; recipes require a validated manifest. CPU tests and the original zero-gradient replay are not proof of effective learning. Full docs-site build not run.
Type of Change
Checklist
pre-commit run --all-files)Latest review fixes and validation
Runtime setup simplified in 36dc981: removed Qwen dependency groups, scoped overrides and the CI uv upgrade. Only AWEX changes in dependency declarations/locks relative to main. Six submission/startup regression tests passed; full pre-commit and commit hooks passed. The earlier model CPU suites remain applicable; GPU/multi-node validation has not been rerun.
Vision tensor memory fix
Share byte-identical CPU vision tensors across repeated images in multi-turn trajectories, while preserving context-dependent processor outputs and per-turn token masks. Added regression coverage for changed image content, processor identity, and trajectory assembly.
Validation: 53 vision/trajectory/streaming tests passed; full pre-commit passed. The separate NUMA, gradient-finalization cache, and checkpoint host-cache fixes are tracked in #1756.