Repository navigation
Apply-MegatronAdaptor-NPU-migration-to-clean-branch #385
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
4057878
dbbe100
9e2079d
0d5e894
2b7386e
3608e57
2c250e4
177b983
a18b20f
9843fce
535f9d4
5611e37
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -7,9 +7,11 @@ FROM ${BASE_IMAGE}:${BASE_IMAGE_TAG} | |
| SHELL ["/bin/bash", "-o", "pipefail", "-c"] | ||
| WORKDIR /root | ||
|
|
||
| ARG MEGATRON_COMMIT=3714d81d418c9f1bca4594fc35f9e8289f652862 | ||
| ARG MEGATRON_COMMIT=1dcf0dafa884ad52ffb243625717a3471643e087 | ||
| ARG MEGATRON_BRIDGE_COMMIT=3fd3768045422d0aa5c97e90a4e6c659aea9acb9 | ||
| ARG MINDSPEED_COMMIT=fc63de5c48426dd019c3b3f39e65f5bdf56e4086 | ||
| ARG MEGATRON_ADAPTOR_COMMIT=15582addff3f3d4680e350826fa70d012b475509 | ||
| ARG TRANSFORMER_ENGINE_NPU_COMMIT=d743c83d060d5edc48867ecb9e93ec80d81860e4 | ||
| ARG MBRIDGE_COMMIT=89eb10887887bc74853f89a4de258c0702932a1c | ||
| ARG SOC_VERSION="ascend910_9391" | ||
| ARG PIP_INDEX_URL="https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple" | ||
|
|
@@ -28,11 +30,12 @@ ENV SOC_VERSION=$SOC_VERSION \ | |
| HCCL_NPU_SOCKET_PORT_RANGE=61000-61050 \ | ||
| PYTORCH_NPU_ALLOC_CONF=expandable_segments:True \ | ||
| HYDRA_FULL_ERROR=1 \ | ||
| PYTHONPATH=/root/Megatron-Bridge/src:/root/Megatron-LM:/root/vime | ||
| PYTHONPATH=/root/Megatron-Bridge/src:/root/Megatron-LM:/root/MegatronAdaptor:/root/TransformerEngineNPU:/root/vime | ||
|
|
||
| # PATCH MAINTENANCE: keep patch COPY/apply operations and | ||
| # docker/npu_patch/series.conf synchronized. | ||
| COPY docker/npu_patch /opt/npu_patch | ||
| COPY docker/patch/latest/megatron.patch /opt/vime_patch/megatron.patch | ||
|
|
||
| RUN git config --global http.sslVerify false | ||
|
|
||
|
|
@@ -75,24 +78,40 @@ RUN git clone --branch bridge https://github.com/radixark/Megatron-Bridge.git \ | |
| RUN git clone https://gitcode.com/Ascend/MindSpeed.git /root/MindSpeed && \ | ||
| git -C /root/MindSpeed checkout "${MINDSPEED_COMMIT}" | ||
|
|
||
| RUN git clone https://gitcode.com/Ascend/MegatronAdaptor.git /root/MegatronAdaptor && \ | ||
| git -C /root/MegatronAdaptor checkout "${MEGATRON_ADAPTOR_COMMIT}" && \ | ||
| git clone https://gitcode.com/Ascend/TransformerEngineNPU.git /root/TransformerEngineNPU && \ | ||
| git -C /root/TransformerEngineNPU checkout "${TRANSFORMER_ENGINE_NPU_COMMIT}" | ||
|
|
||
| RUN git clone https://github.com/ISEEKYAN/mbridge.git /root/mbridge && \ | ||
| git -C /root/mbridge checkout "${MBRIDGE_COMMIT}" | ||
|
|
||
| # Apply NPU training-stack patches from the build-context snapshot. | ||
| RUN git -C /root/Megatron-LM apply --whitespace=nowarn \ | ||
| /opt/npu_patch/megatron_comm.patch && \ | ||
| # The NPU Megatron patch is based on the common Vime Megatron patch, so the | ||
| # common patch must be applied first. | ||
| RUN git -C /root/Megatron-LM apply --check --whitespace=nowarn \ | ||
| /opt/vime_patch/megatron.patch && \ | ||
| git -C /root/Megatron-LM apply --whitespace=nowarn \ | ||
| /opt/vime_patch/megatron.patch && \ | ||
| git -C /root/Megatron-LM apply --check --whitespace=nowarn \ | ||
| /opt/npu_patch/megatron.patch && \ | ||
| git -C /root/Megatron-LM apply --whitespace=nowarn \ | ||
| /opt/npu_patch/megatron.patch && \ | ||
| git -C /root/Megatron-Bridge apply --check --whitespace=nowarn \ | ||
| /opt/npu_patch/megatron-bridge.patch && \ | ||
| git -C /root/Megatron-Bridge apply --whitespace=nowarn \ | ||
| /opt/npu_patch/megatron-bridge.patch && \ | ||
| git -C /root/MindSpeed apply --whitespace=nowarn \ | ||
| /opt/npu_patch/mindspeed.patch | ||
|
|
||
| # Megatron-Bridge is used directly from PYTHONPATH. Installing its package | ||
| # metadata would pull CUDA-only dependencies into the Ascend environment. | ||
| RUN pip install --no-build-isolation "nvidia-modelopt[torch]>=0.37.0" && \ | ||
| RUN pip install --constraint /tmp/vime-npu-constraints.txt --no-build-isolation \ | ||
| "nvidia-modelopt==0.46.0" "nvdlfw-inspect==0.2.2" && \ | ||
| pip install --no-deps --no-build-isolation -e /root/mbridge && \ | ||
| pip install --no-deps --no-build-isolation -e /root/Megatron-LM && \ | ||
| pip install --no-deps --no-build-isolation -e /root/TransformerEngineNPU && \ | ||
| pip install --no-deps --no-build-isolation -e /root/MegatronAdaptor && \ | ||
| pip install --no-deps --no-build-isolation -e /root/MindSpeed | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. NPU CI (#193) likely fails during image build: Dockerfile.npu removes the MindSpeed clone but still runs pip install -e /root/MindSpeed.
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Offline weight updates depend on MindSpeed, so the MindSpeed code has been re-added. |
||
|
|
||
| # Defaults to the ascend branch for local builds. Release workflows should | ||
|
|
@@ -123,7 +142,7 @@ RUN git clone --depth 1 --branch 2026.6.0 \ | |
|
|
||
| # Minimal import check. | ||
| RUN source /usr/local/Ascend/ascend-toolkit/set_env.sh && \ | ||
| python3 -c 'import megatron, mindspeed, torch_memory_saver, vime, vllm, vllm_ascend;' | ||
| python3 -c 'import megatron, mindspeed, megatron_adaptor, transformer_engine, torch_memory_saver, vime, vllm, vllm_ascend;' | ||
|
|
||
| WORKDIR /root/vime | ||
| ENTRYPOINT [] | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -8,7 +8,9 @@ This guide provides instructions for installing Vime with NPU support, including | |
| | --------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | | ||
| | vime | main | [GitHub](https://github.com/vllm-project/vime/tree/main) | | ||
| | Megatron-Bridge | 3fd3768045422d0aa5c97e90a4e6c659aea9acb9 | [GitHub](https://github.com/radixark/Megatron-Bridge) | | ||
| | Megatron-LM | 3714d81d418c9f1bca4594fc35f9e8289f652862 | [GitHub](https://github.com/NVIDIA/Megatron-LM) | | ||
| | Megatron-LM | 1dcf0dafa884ad52ffb243625717a3471643e087 | [GitHub](https://github.com/NVIDIA/Megatron-LM) | | ||
| | MegatronAdaptor | main | [GitCode](https://gitcode.com/Ascend/MegatronAdaptor) | | ||
| | TransformerEngineNPU | main | [GitCode](https://gitcode.com/Ascend/TransformerEngineNPU) | | ||
| | MindSpeed | fc63de5c48426dd019c3b3f39e65f5bdf56e4086 | [GitCode](https://gitcode.com/Ascend/MindSpeed) | | ||
| | HDK | 25.3.RC1 | [Ascend](https://www.hiascend.com/hardware/firmware-drivers/commercial?product=7\&model=33) | | ||
| | CANN | 9.0.0 | [Ascend](https://www.hiascend.com/developer/download/community/result?module=cann\&cann=9.0.0\&product=7\&model=33) | | ||
|
|
@@ -31,7 +33,6 @@ git clone --branch ascend https://github.com/vllm-project/vime.git "${WORKSPACE} | |
| export PATCH_DIR="${WORKSPACE}/vime/docker/npu_patch" | ||
| ``` | ||
|
|
||
|
|
||
| #### 1. Megatron-Bridge | ||
|
|
||
| Used via `PYTHONPATH` (no editable install); it requires `nvidia-modelopt`. | ||
|
|
@@ -48,21 +49,29 @@ git -C "${WORKSPACE}/Megatron-Bridge" apply --whitespace=nowarn "${PATCH_DIR}/me | |
| pip install --no-build-isolation "nvidia-modelopt[torch]>=0.37.0" | ||
| ``` | ||
|
|
||
|
|
||
| #### 2. Megatron-LM | ||
|
|
||
| ```bash | ||
| export MEGATRON_COMMIT=3714d81d418c9f1bca4594fc35f9e8289f652862 | ||
| export MEGATRON_COMMIT=1dcf0dafa884ad52ffb243625717a3471643e087 | ||
| git clone https://github.com/NVIDIA/Megatron-LM.git "${WORKSPACE}/Megatron-LM" | ||
| git -C "${WORKSPACE}/Megatron-LM" checkout "${MEGATRON_COMMIT}" | ||
|
|
||
| git -C "${WORKSPACE}/Megatron-LM" apply --whitespace=nowarn "${PATCH_DIR}/megatron_comm.patch" | ||
| git -C "${WORKSPACE}/Megatron-LM" apply --whitespace=nowarn "${WORKSPACE}/vime/docker/patch/latest/megatron.patch" | ||
| git -C "${WORKSPACE}/Megatron-LM" apply --whitespace=nowarn "${PATCH_DIR}/megatron.patch" | ||
|
|
||
| pip install --no-deps --no-build-isolation -e "${WORKSPACE}/Megatron-LM" | ||
| ``` | ||
|
|
||
| #### 3. MindSpeed | ||
| #### 3. MegatronAdaptor and TransformerEngineNPU | ||
|
|
||
| The NPU training stack now uses the two source repositories directly. The mainline Megatron patch is applied first; `docker/npu_patch/megatron.patch` contains only the NPU-specific changes rebased onto that mainline patch: | ||
|
|
||
| pip install --no-deps --no-build-isolation -e ${WORKSPACE}/MegatronAdaptor | ||
| pip install --no-deps --no-build-isolation -e ${WORKSPACE}/TransformerEngineNPU | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Do we need to clone these two repos or not? I think this step could use more detail.
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. need install |
||
|
|
||
| Do not install the CUDA TransformerEngine package in the same environment. | ||
|
|
||
| #### 4. MegatronAdaptor and TransformerEngineNPU | ||
|
|
||
| ```bash | ||
| export MINDSPEED_COMMIT=fc63de5c48426dd019c3b3f39e65f5bdf56e4086 | ||
|
|
@@ -74,8 +83,7 @@ git -C "${WORKSPACE}/MindSpeed" apply --whitespace=nowarn "${PATCH_DIR}/mindspee | |
| pip install --no-deps --no-build-isolation -e "${WORKSPACE}/MindSpeed" | ||
| ``` | ||
|
|
||
|
|
||
| #### 4. Vime | ||
| #### 5. Vime | ||
|
|
||
| ```bash | ||
| pip install -r "${WORKSPACE}/vime/requirements.txt" | ||
|
|
@@ -98,7 +106,6 @@ pip install --no-deps output/torch_memory_saver-0.0.8-cp312-cp312-linux_aarch64. | |
|
|
||
| #### 5. Install vLLM and vLLM Ascend | ||
|
|
||
|
|
||
| ```bash | ||
| export VLLM_COMMIT=9090368b650896bf5fc990c921df7eb4c20355a5 | ||
|
|
||
|
|
@@ -126,4 +133,3 @@ pip install torch-npu==2.10.0 | |
| pip install torchvision==0.25.0 | ||
| pip install numpy==1.26.4 | ||
| ``` | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Could you please remove the megatron_comm.patch and mindspeed.patch entries from docker/npu_patch/series.conf? This keeps the declared patch stack consistent with the Dockerfile and available patch files.