Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 25 additions & 6 deletions docker/Dockerfile.npu
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,11 @@ FROM ${BASE_IMAGE}:${BASE_IMAGE_TAG}
SHELL ["/bin/bash", "-o", "pipefail", "-c"]
WORKDIR /root

ARG MEGATRON_COMMIT=3714d81d418c9f1bca4594fc35f9e8289f652862
ARG MEGATRON_COMMIT=1dcf0dafa884ad52ffb243625717a3471643e087
ARG MEGATRON_BRIDGE_COMMIT=3fd3768045422d0aa5c97e90a4e6c659aea9acb9
ARG MINDSPEED_COMMIT=fc63de5c48426dd019c3b3f39e65f5bdf56e4086
ARG MEGATRON_ADAPTOR_COMMIT=15582addff3f3d4680e350826fa70d012b475509
ARG TRANSFORMER_ENGINE_NPU_COMMIT=d743c83d060d5edc48867ecb9e93ec80d81860e4
ARG MBRIDGE_COMMIT=89eb10887887bc74853f89a4de258c0702932a1c

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please remove the megatron_comm.patch and mindspeed.patch entries from docker/npu_patch/series.conf? This keeps the declared patch stack consistent with the Dockerfile and available patch files.

ARG SOC_VERSION="ascend910_9391"
ARG PIP_INDEX_URL="https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple"
Expand All @@ -28,11 +30,12 @@ ENV SOC_VERSION=$SOC_VERSION \
HCCL_NPU_SOCKET_PORT_RANGE=61000-61050 \
PYTORCH_NPU_ALLOC_CONF=expandable_segments:True \
HYDRA_FULL_ERROR=1 \
PYTHONPATH=/root/Megatron-Bridge/src:/root/Megatron-LM:/root/vime
PYTHONPATH=/root/Megatron-Bridge/src:/root/Megatron-LM:/root/MegatronAdaptor:/root/TransformerEngineNPU:/root/vime

# PATCH MAINTENANCE: keep patch COPY/apply operations and
# docker/npu_patch/series.conf synchronized.
COPY docker/npu_patch /opt/npu_patch
COPY docker/patch/latest/megatron.patch /opt/vime_patch/megatron.patch

RUN git config --global http.sslVerify false

Expand Down Expand Up @@ -75,24 +78,40 @@ RUN git clone --branch bridge https://github.com/radixark/Megatron-Bridge.git \
RUN git clone https://gitcode.com/Ascend/MindSpeed.git /root/MindSpeed && \
git -C /root/MindSpeed checkout "${MINDSPEED_COMMIT}"

RUN git clone https://gitcode.com/Ascend/MegatronAdaptor.git /root/MegatronAdaptor && \
git -C /root/MegatronAdaptor checkout "${MEGATRON_ADAPTOR_COMMIT}" && \
git clone https://gitcode.com/Ascend/TransformerEngineNPU.git /root/TransformerEngineNPU && \
git -C /root/TransformerEngineNPU checkout "${TRANSFORMER_ENGINE_NPU_COMMIT}"

RUN git clone https://github.com/ISEEKYAN/mbridge.git /root/mbridge && \
git -C /root/mbridge checkout "${MBRIDGE_COMMIT}"

# Apply NPU training-stack patches from the build-context snapshot.
RUN git -C /root/Megatron-LM apply --whitespace=nowarn \
/opt/npu_patch/megatron_comm.patch && \
# The NPU Megatron patch is based on the common Vime Megatron patch, so the
# common patch must be applied first.
RUN git -C /root/Megatron-LM apply --check --whitespace=nowarn \
/opt/vime_patch/megatron.patch && \
git -C /root/Megatron-LM apply --whitespace=nowarn \
/opt/vime_patch/megatron.patch && \
git -C /root/Megatron-LM apply --check --whitespace=nowarn \
/opt/npu_patch/megatron.patch && \
git -C /root/Megatron-LM apply --whitespace=nowarn \
/opt/npu_patch/megatron.patch && \
git -C /root/Megatron-Bridge apply --check --whitespace=nowarn \
/opt/npu_patch/megatron-bridge.patch && \
git -C /root/Megatron-Bridge apply --whitespace=nowarn \
/opt/npu_patch/megatron-bridge.patch && \
git -C /root/MindSpeed apply --whitespace=nowarn \
/opt/npu_patch/mindspeed.patch

# Megatron-Bridge is used directly from PYTHONPATH. Installing its package
# metadata would pull CUDA-only dependencies into the Ascend environment.
RUN pip install --no-build-isolation "nvidia-modelopt[torch]>=0.37.0" && \
RUN pip install --constraint /tmp/vime-npu-constraints.txt --no-build-isolation \
"nvidia-modelopt==0.46.0" "nvdlfw-inspect==0.2.2" && \
pip install --no-deps --no-build-isolation -e /root/mbridge && \
pip install --no-deps --no-build-isolation -e /root/Megatron-LM && \
pip install --no-deps --no-build-isolation -e /root/TransformerEngineNPU && \
pip install --no-deps --no-build-isolation -e /root/MegatronAdaptor && \
pip install --no-deps --no-build-isolation -e /root/MindSpeed

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NPU CI (#193) likely fails during image build: Dockerfile.npu removes the MindSpeed clone but still runs pip install -e /root/MindSpeed.
Drop that line? This migration also requires image-build (not smk-only on the legacy MindSpeed image), and the common docker/patch/latest/megatron.patch is Dockerfile-only and not in series.conf.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Offline weight updates depend on MindSpeed, so the MindSpeed code has been re-added.


# Defaults to the ascend branch for local builds. Release workflows should
Expand Down Expand Up @@ -123,7 +142,7 @@ RUN git clone --depth 1 --branch 2026.6.0 \

# Minimal import check.
RUN source /usr/local/Ascend/ascend-toolkit/set_env.sh && \
python3 -c 'import megatron, mindspeed, torch_memory_saver, vime, vllm, vllm_ascend;'
python3 -c 'import megatron, mindspeed, megatron_adaptor, transformer_engine, torch_memory_saver, vime, vllm, vllm_ascend;'

WORKDIR /root/vime
ENTRYPOINT []
Expand Down
26 changes: 16 additions & 10 deletions docker/npu_patch/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,9 @@ This guide provides instructions for installing Vime with NPU support, including
| --------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| vime | main | [GitHub](https://github.com/vllm-project/vime/tree/main) |
| Megatron-Bridge | 3fd3768045422d0aa5c97e90a4e6c659aea9acb9 | [GitHub](https://github.com/radixark/Megatron-Bridge) |
| Megatron-LM | 3714d81d418c9f1bca4594fc35f9e8289f652862 | [GitHub](https://github.com/NVIDIA/Megatron-LM) |
| Megatron-LM | 1dcf0dafa884ad52ffb243625717a3471643e087 | [GitHub](https://github.com/NVIDIA/Megatron-LM) |
| MegatronAdaptor | main | [GitCode](https://gitcode.com/Ascend/MegatronAdaptor) |
| TransformerEngineNPU | main | [GitCode](https://gitcode.com/Ascend/TransformerEngineNPU) |
| MindSpeed | fc63de5c48426dd019c3b3f39e65f5bdf56e4086 | [GitCode](https://gitcode.com/Ascend/MindSpeed) |
| HDK | 25.3.RC1 | [Ascend](https://www.hiascend.com/hardware/firmware-drivers/commercial?product=7\&model=33) |
| CANN | 9.0.0 | [Ascend](https://www.hiascend.com/developer/download/community/result?module=cann\&cann=9.0.0\&product=7\&model=33) |
Expand All @@ -31,7 +33,6 @@ git clone --branch ascend https://github.com/vllm-project/vime.git "${WORKSPACE}
export PATCH_DIR="${WORKSPACE}/vime/docker/npu_patch"
```


#### 1. Megatron-Bridge

Used via `PYTHONPATH` (no editable install); it requires `nvidia-modelopt`.
Expand All @@ -48,21 +49,29 @@ git -C "${WORKSPACE}/Megatron-Bridge" apply --whitespace=nowarn "${PATCH_DIR}/me
pip install --no-build-isolation "nvidia-modelopt[torch]>=0.37.0"
```


#### 2. Megatron-LM

```bash
export MEGATRON_COMMIT=3714d81d418c9f1bca4594fc35f9e8289f652862
export MEGATRON_COMMIT=1dcf0dafa884ad52ffb243625717a3471643e087
git clone https://github.com/NVIDIA/Megatron-LM.git "${WORKSPACE}/Megatron-LM"
git -C "${WORKSPACE}/Megatron-LM" checkout "${MEGATRON_COMMIT}"

git -C "${WORKSPACE}/Megatron-LM" apply --whitespace=nowarn "${PATCH_DIR}/megatron_comm.patch"
git -C "${WORKSPACE}/Megatron-LM" apply --whitespace=nowarn "${WORKSPACE}/vime/docker/patch/latest/megatron.patch"
git -C "${WORKSPACE}/Megatron-LM" apply --whitespace=nowarn "${PATCH_DIR}/megatron.patch"

pip install --no-deps --no-build-isolation -e "${WORKSPACE}/Megatron-LM"
```

#### 3. MindSpeed
#### 3. MegatronAdaptor and TransformerEngineNPU

The NPU training stack now uses the two source repositories directly. The mainline Megatron patch is applied first; `docker/npu_patch/megatron.patch` contains only the NPU-specific changes rebased onto that mainline patch:

pip install --no-deps --no-build-isolation -e ${WORKSPACE}/MegatronAdaptor
pip install --no-deps --no-build-isolation -e ${WORKSPACE}/TransformerEngineNPU

@floatlibai floatlibai Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to clone these two repos or not? I think this step could use more detail.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

need install


Do not install the CUDA TransformerEngine package in the same environment.

#### 4. MegatronAdaptor and TransformerEngineNPU

```bash
export MINDSPEED_COMMIT=fc63de5c48426dd019c3b3f39e65f5bdf56e4086
Expand All @@ -74,8 +83,7 @@ git -C "${WORKSPACE}/MindSpeed" apply --whitespace=nowarn "${PATCH_DIR}/mindspee
pip install --no-deps --no-build-isolation -e "${WORKSPACE}/MindSpeed"
```


#### 4. Vime
#### 5. Vime

```bash
pip install -r "${WORKSPACE}/vime/requirements.txt"
Expand All @@ -98,7 +106,6 @@ pip install --no-deps output/torch_memory_saver-0.0.8-cp312-cp312-linux_aarch64.

#### 5. Install vLLM and vLLM Ascend


```bash
export VLLM_COMMIT=9090368b650896bf5fc990c921df7eb4c20355a5

Expand Down Expand Up @@ -126,4 +133,3 @@ pip install torch-npu==2.10.0
pip install torchvision==0.25.0
pip install numpy==1.26.4
```

Loading