Skip to content

Kimi-K3 lora RL day-0 support - #1825

Open
yueming-yuan wants to merge 1 commit into
yueming/cp-layout-utils-to-pluginsfrom
kimi-k3
Open

Kimi-K3 lora RL day-0 support#1825
yueming-yuan wants to merge 1 commit into
yueming/cp-layout-utils-to-pluginsfrom
kimi-k3

Conversation

@yueming-yuan

@yueming-yuan yueming-yuan commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Blog

check details and experiment results in:
https://www.lmsys.org/blog/2026-07-27-kimi-k3-day0-support

Usage

sglang branch: https://github.com/sgl-project/sglang/tree/sglang-miles-k3
Docker image will be uploaded within 24h.

About

Megatron training backend for Kimi K3 plus the colocated SGLang rollout path
it needs for RL. KDA and MLA attention per layer, the attention-residual
snapshot bank, and native LoRA adapters under TP/EP/PP/CP — with shared-A /
per-expert-B factors for the 896-expert MoE, exported to the rollout engines
as HF-named chunks over CUDA IPC.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

Megatron training backend for Kimi K3 plus the colocated SGLang rollout path
it needs for RL.

Model: KDA (delta-rule) and MLA attention per layer, the attention-residual
snapshot bank, and situ activation, under TP/EP/PP/CP. Native LoRA adapters
are applied in-model, with shared-A/per-expert-B factors for the 896-expert
MoE, and exported to the rollout engines as HF-named chunks over CUDA IPC.

Also included: mbridge and megatron_bridge plugins, megatron->HF conversion,
MXFP4 pack/unpack plus an MXFP4->BF16 checkpoint tool, and the launchers.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant