Skip to content

refactor: organize models, backends, contracts and train and rollout engines #892

refactor: organize models, backends, contracts and train and rollout engines

refactor: organize models, backends, contracts and train and rollout engines #892

Workflow file for this run

name: GPU CI

Check warning on line 1 in .github/workflows/gpu-ci.yml

View workflow run for this annotation

GitHub Actions / GPU CI

Workflow execution policy warning (evaluate mode)

On November 2, 2026, GitHub will restrict `pull_request_target` on public repositories by default. To continue allowing the event trigger, configure an Actions policy. Learn more: https://gh.io/securely-using-pull_request_target#default-policy-for-pull_request_target
on:
pull_request_target:
branches: [ main, test ]
paths:
- 'csrc/**'
- 'rl_engine/**'
- 'tests/**'
- 'pyproject.toml'
- 'setup.py'
- 'requirements*.txt'
- 'requirements/**'
- 'build_tools/**'
- 'MANIFEST.in'
- 'docker/Dockerfile.cuda'
- 'ci/scripts/run_gpu_ci.sh'
- '.github/workflows/gpu-ci.yml'
types: [ opened, synchronize, reopened, labeled ]
concurrency:
group: gpu-ci-serial
cancel-in-progress: false
jobs:
gpu-tests:
if: contains(github.event.pull_request.labels.*.name, 'needs-gpu-ci')
runs-on: ubuntu-latest
timeout-minutes: 60
strategy:
fail-fast: false
matrix:
include:
- { gpu_id: "NVIDIA RTX A4000", target_sm: "8.6" } # SM86
- { gpu_id: "NVIDIA H100 80GB HBM3", target_sm: "9.0", force_sm90: "1" } # SM90
steps:
- name: Checkout secure orchestrator script from base branch
uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.base.sha }}
- name: Install runpodctl
run: |
wget -qO runpodctl https://github.com/runpod/runpodctl/releases/latest/download/runpodctl-linux-amd64
chmod +x runpodctl
sudo mv runpodctl /usr/local/bin/runpodctl
runpodctl version
- name: Configure runpodctl
run: runpodctl config --apiKey "${{ secrets.RUNPOD_API_KEY }}"
- name: Setup SSH key
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "${{ secrets.RUNPOD_SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
chmod 600 ~/.ssh/id_ed25519
ssh-keygen -y -f ~/.ssh/id_ed25519 > /dev/null && echo "key OK" || echo "key BROKEN"
# Follow-up (#191): to test multiple architectures, add a matrix and pass
# GPU_ID + TARGET_SM (+ KERNEL_ALIGN_FORCE_SM90 for Hopper) through to the script.
# run_gpu_ci.sh reads all three, normalizes TARGET_SM, asserts the pod matches it,
# and forwards KERNEL_ALIGN_FORCE_SM90 into the remote build:
# strategy:
# matrix:
# include:
# - { gpu_id: "NVIDIA RTX A4000", target_sm: "8.6" } # Ampere
# - { gpu_id: "NVIDIA A100 80GB PCIe", target_sm: "8.0" }
# - { gpu_id: "NVIDIA H100 PCIe", target_sm: "9.0", force_sm90: "1" } # build Hopper TMA/WGMMA kernels
# - { gpu_id: "NVIDIA B200", target_sm: "10.0" }
# Per-arch jobs must NOT fall back to a different-capability GPU: the script
# fails fast when the pod arch != requested TARGET_SM, so keep fallback within
# the same compute capability (or unset it for these jobs).
- name: Run GPU tests on RunPod
env:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
PR_REPO_URL: ${{ github.event.pull_request.head.repo.clone_url }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
GPU_ID: ${{ matrix.gpu_id }}
TARGET_SM: ${{ matrix.target_sm }}
KERNEL_ALIGN_FORCE_SM90: ${{ matrix.force_sm90 }}
run: bash ci/scripts/run_gpu_ci.sh