Repository navigation
fix(ci): pin runpodctl v2.9.0 and drop the deprecated config step (RUNPOD_API_KEY auth) - #470
fusheng-ji wants to merge 3 commits into
Conversation
ws1-gtest-gpu and ws1-chain-gpu fail at "Configure runpodctl" before any test runs. They install releases/latest (v2.14.0 since 2026-09-10), whose deprecated `config --apiKey` aborts with `Config File ".runpod.yaml" Not Found` when it is the first runpodctl call on a fresh HOME. It only succeeds once an earlier call has created ~/.runpod/config.toml, which is why gpu-ci, whose install step happens to run `runpodctl version` first, got past the same command. Drop the config step in all three workflows. runpodctl v2 reads RUNPOD_API_KEY from the environment, and every step that runs ci/run_gpu_ci.sh already exports it, including the cleanup trap that calls `pod remove`. Verified for v2.9.0 inside `unshare -rn`, from a fresh HOME after `runpodctl version`: without the variable `pod list` fails locally with no_credentials, with RUNPOD_API_KEY=dummy it gets past the credential check to a network_error, and config.toml keeps `apikey = ''` throughout. v2.14.0 behaves the same. Pin the download to v2.9.0, the version gpu-ci run 31455127310 exercised end to end (pod create with sold-out fallback, pod id parsing, pod get polling, SSH, cleanup), and verify it with sha256sum -c against the runpodctl-linux-amd64 line of checksums_2.9.0_sha256.txt (06e6f54957db79d5cd9f1909a7f1d365076826751ba2f5df65d75dde43a64148), which matches GitHub's asset digest. Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
|
Important Review skippedReview was skipped as selected files did not have any reviewable changes. ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (3)
Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 1 remain after this review. 📝 WalkthroughWalkthroughThree GPU workflows now download runpodctl v2.9.0 and verify its SHA-256 checksum before installation. They remove separate API-key configuration steps. Two workflows print the installed version. ChangesGPU workflow runpodctl installation
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The workflows install the verified v2.9.0 binary and provide its required API key to the GPU script. No concrete regression from these changes remains; the PR is mergeable after normal checks. Architecture SummaryArchitecture risk: 🔵 Low · up to The changed surface does not map to a changed system, dependency edge, entrypoint, or external dependency. Changed systems: None identified. Architecture concerns Review detailsBefore / after behavior
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
@coderabbitai review |
|
Flink-ddd
left a comment
There was a problem hiding this comment.
Could you run this on RunPod and share logs showing that pod creation, SSH connection, and pod deletion all work?
|
Thanks for the review. I don't have a RunPod account, and the GPU jobs can't run from a fork, so this needs a maintainer run with the org key. Here is what is available now and the run that would close the gap. 1. Existing evidence with the pinned binary. Run 31455127310 (GPU CI, 2026-08-11, success) used
What that run does not prove. In that run 2. The failure this PR fixes is widespread. Across the 111 forks there are 138 runs of the GPU workflows: 108 failed, 30 skipped, 0 succeeded. A sample of 25 forks shows every failed job stopping at 3. Request. Could a maintainer push this PR's head ( For reference, a maintainer can dispatch it with: |
Latest Status [2026-10-04]
2026-10-04: rebased onto
main(6d8b4bc) and signed off (DCO): one commit,a1736d1, with anidentical patch (
git patch-id). The results below were measured before the rebase.Ready for review. One commit, three workflow files, no script changes.
Line numbers refer to this branch (
a1736d1). The workflow files shift by a few linescompared with
main;ci/run_gpu_ci.shis unchanged.Summary
releases/latest(v2.14.0 since2026-09-10) and then run
runpodctl config --apiKey …as the firstrunpodctlcall ona fresh runner. That fails because no config file exists yet.
gpu-ci.ymlgets pastthe same command only because it runs
runpodctl versionfirst, which creates the file.(31455127310)
exercised end to end:
pod create, pod-id parsing,pod getpolling, SSH,pod remove. The download is verified withsha256sum -c, andrunpodctl versionis printed into the log.
Configure runpodctl. Every laterrunpodctlcall is insideci/run_gpu_ci.sh, and every step that runs it setsRUNPOD_API_KEYin itsenv:.That includes the
pod removein itsEXITtrap, which runs in the same shell.configsucceeded (ingpu-ci, which runsversionfirst), it also generated an SSH key pair and uploaded it to the RunPod account. The script never used it:
ssh/scppass no-iand use the key written by
Setup SSH key.Files
.github/workflows/ws1-gtest-gpu.yml:75-88);Configure runpodctlremoved; key at:98, script at:106.github/workflows/ws1-chain-gpu.yml:93, script at:102.github/workflows/gpu-ci.yml:80, script at:86git diff --stat upstream/main(1968a87): 3 files, 29 insertions, 12 deletions.Test
Test results
Run locally against the pinned binary in a scratch directory. No pod was created, and
no request carrying a credential was sent to RunPod.
main'runpodctl config' is deprecated, thenerror saving config: Config File ".runpod.yaml" Not Found, exit 1ws1-gtest-gpuhistoryaction_required); its first runs (2026-08-19) already failed at this steprunpodctl-linux-amd64: OK; sha25606e6f549…a64148also matches GitHub's asset digestrunpodctl 2.9.0-c094cac, the same string run 31455127310 printedpod createflags present;pod get -o json;pod removeis an alias ofpod deleteRUNPOD_API_KEYno_credentials; dummy key:network_errorat DNS;config.tomlstaysapikey = ''; v2.14.0 identicalrun:block, executed in a scratch directory (script not attached): all three printOKand the version; a wrong checksum printsFAILEDand exits 1yaml.safe_loadpasses; noConfigure runpodctlleft; actionlint 1.7.12 clean (shellcheck was not installed, so actionlint's checks of therun:shell code did not run); pre-commit passesLocal reproduction of the root cause (v2.14.0, empty HOME)
gpu-ci run 34585491929
(2026-09-11) shows the same thing in CI: it ran
versionfirst and then got pastConfigure runpodctl. A later step (pod create) failed in that run for an unrelatedreason.
Notes
Risks this PR cannot verify
v2.9.0's output against today's API.
run_gpu_ci.shparsesrunpodctloutputwith
grep:"no longer any instances available"(:60,:76);:82,:84);:100,:101);"not found"(:37).Run 31455127310 (2026-08-11, this exact binary) exercised
:60/:76,:82,:100/:101and thepod removecall. The:84fallback (only reached when:82does not match) and the
:37branch were not exercised.Existing bug, not changed here. The pod-id fallback regex at
:84can take aJSON error
codesuch asunauthorizedas the pod id, and then poll it for up to10 minutes. A follow-up could check for a top-level
"error"key first.How this PR can be exercised before merge
ws1-gtest-gpu.yml:59andws1-chain-gpu.yml:46skip the GPU jobs at PRtime.
gpu-ci.ymlruns onpull_request_target(:4) and checks out its orchestrator fromthe base commit (
:40). Aneeds-gpu-cilabel would therefore runmain'sorchestrator, and its setup already passed before this PR.
mainafter merge:ws1-chain-gpu.ymlhas nopathsfilter, andws1-gtest-gpu.yml's push paths include the workflow file(
:47). A maintainer can also runworkflow_dispatchonws1-gtest-gpu.yml(:48)from a branch in
RL-Align/RL-Kernel.configruns uploaded may be worth pruning from the account.The generated public key carries the comment
runpodctl-ssh-key.Summary by CodeRabbit