Repository navigation
feat: add bounded colocate transfer and Qwen4Exp converters - #121
Merged
chaokunyang merged 2 commits intoSep 22, 2026
Merged
Conversation
5 of 7 tasks
dingzhiqiang
marked this pull request as ready for review
September 22, 2026 11:11
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Keep transfer planning as tensor views and allocate contiguous send/receive buffers only for the active batch, including strided column slices. Copy receives back before releasing that batch. Pack compatible expert operations within configured bounds, check packing limits across inference ranks before P2P starts, and provide CUDA IPC staging allocation that restores the allocator settings used for training. The packing byte limit is not a hard cap for a single oversized tensor.
Add a reader factory so integrations can explicitly choose the bounded transport; the AWEX default remains unchanged. All participating readers must choose the same transport. This extracts the shared transport implementation from areal-project/AReaL#1731.
Own Qwen4Exp weight converters, GDN/gated-QKV layouts and frozen-parameter declarations in AWEX. Registration is explicit and accepts caller-supplied binders; checkpoint evidence and live engine ownership stay in the integration. AReaL consumes these APIs in areal-project/AReaL#1751 without a second converter implementation.
Validation