| Cases | Node require | Description |
|---|---|---|
| one_device_per_thread | 1 | One Device(1 GPU) per Process or Thread |
| multi_devices_per_thread | 1 | Multiple Devices(more than one GPU) per Process or Thread |
| nonblocking_double_streams | 1 | One rank has two communicators. |
| nccl_with_mpi | 1 | Run with Open MPI |
| node_server/node_client | 2 | Using socket for init |
Clone this git lib to your local env, such as /home/xky/
Requirements:
- CUDA
- NVIDIA NCCL (optimized for NVLink)
- Open-MPI (option)
Recommend using docker images:
docker pull nvcr.io/nvidia/pytorch:24.07-py3If there is docker-ce, run docker:
sudo docker run --net=host --gpus=all -it -e UID=root --ipc host --shm-size="32g" \
-v /home/xky/:/home/xky \
-u 0 \
--name=nccl2 nvcr.io/nvidia/pytorch:24.07-py3 bashOthers:
docker run \
--runtime=nvidia \
--privileged \
--device /dev/nvidia0:/dev/nvidia0 \
--device /dev/nvidia1:/dev/nvidia1 \
--device /dev/nvidia2:/dev/nvidia2 \
--device /dev/nvidia3:/dev/nvidia3 \
--device /dev/nvidia4:/dev/nvidia4 \
--device /dev/nvidia5:/dev/nvidia5 \
--device /dev/nvidia6:/dev/nvidia6 \
--device /dev/nvidia7:/dev/nvidia7 \
--device /dev/nvidiactl:/dev/nvidiactl \
--device /dev/nvidia-uvm:/dev/nvidia-uvm \
--device /dev/nvidia-uvm-tools:/dev/nvidia-uvm-tools \
--device /dev/infiniband:/dev/infiniband \
-v /usr/local/bin/:/usr/local/bin/ \
-v /opt/cloud/cce/nvidia/:/usr/local/nvidia/ \
-v /home/xky/:/home/xky \
--ipc host \
--net host \
-it \
-u root \
--name nccl_env \
nvcr.io/nvidia/pytorch:24.07-py3 bashEnter the git directory and run makefile
cd /home/xky/BasicCUDA/nccl/
makeIf there is MPI lib in env, could compile MPI case:
make mpi./multi_devices_per_thread
./one_devices_per_thread
./nonblocking_double_streamsSet DEBUG=1 would print some debug information.
Could change ranks number by set '--nranks'. e.g:
DEBUG=1 ./nonblocking_double_streams --nranks 8MPI case run:
mpirun -n 6 --allow-run-as-root ./nccl_with_mpiTwo nodes case: using socket connection for nccl init.
Server run in one:
./node_serverClient run in another one, e.g. Server IP: 10.10.1.1
./node_client --hostname 10.10.1.1Add some envs:
# server:
NCCL_DEBUG=INFO NCCL_NET_PLUGIN=none NCCL_IB_DISABLE=1 ./node_server --port 8066 --nranks 8
# client:
NCCL_DEBUG=INFO NCCL_NET_PLUGIN=none NCCL_IB_DISABLE=1 ./node_client --hostname 10.10.1.1 --port 8066 --nranks 8