Skip to content

Upgrade the MI355X default runtime to ROCm 10 - #123

Draft
irvineoy wants to merge 1 commit into
mainfrom
codex/mi355x-rocm10-runtime
Draft

irvineoy wants to merge 1 commit into
mainfrom
codex/mi355x-rocm10-runtime

Conversation

@irvineoy

@irvineoy irvineoy commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

The MI355X default still uses ROCm 7.2, while the ROCm 10 image changes Python, Triton, FlyDSL and SDK library layout. This change selects the digest-pinned SGLang 0.5.20 / ROCm 10 image and updates the runner, all six optional evaluation-tool runtimes, affected tasks and documentation together.

This PR remains a draft with two qualification blockers:

  • GPU ASan: the image builds and its packages pass checksum verification, but HIP safe and known-bug probes fault during runtime heap initialization; readiness then times out with other controls incomplete. Ordinary host-allocation GPU writes fail under both pinned old and new images on the tested hosts, while device, managed and pinned-host allocations pass. Qualification needs a working host configuration and a fresh complete control/candidate run.
  • FP8 fused-MoE: the existing numerical gate fails intermittently on both old and new runtimes (observed normalized error about 1.145% against a 1% limit). The original gate is retained and this task is not counted as passing. An upgrade-only regression has not been established.

Changes

  • Record the scoring image for ordinary runs and reject reuse of workspaces whose runtime identity does not match. Separate GEAK SDK caches by Python ABI and select one ROCm SDK library tree for workloads and rocprofv3.
  • Rebase all six tool runtimes onto the pinned scoring image. Keep the native rocJITsu builder independently pinned. Update Triton/FlyDSL AOT extraction, require independently attested FpSan safe/mismatch controls, and match ASan startup and candidate library environments. Runtime paths remain trusted sidecar evidence and cannot be supplied by task config.
  • Add task-local adapters for removed FlyDSL buffer/vector/type APIs, repair existing timing-evidence metadata, and protect reference state exposed during validation. Use a single-stage gfx950 MLA schedule to correct repeated decode failures under the new compiler. This fallback has a latency cost.
  • Update installation, compatibility, evaluation-tool and task documentation. Existing workload cases, numerical tolerances and timing parameters are retained; QK documentation now accurately describes the existing six BF16 workloads. Other GPU architecture defaults remain unchanged.

Validation

Hardware: real AMD Instinct MI355X GPUs, with Docker/Slurm runs on Linux. The new runtime is SGLang 0.5.20, ROCm SDK 10.0.0, Python 3.12, PyTorch 2.11, Triton 3.8 and FlyDSL 0.3.2, pinned to sha256:e20849665c105d389ef91d23c0dc73931aaa6f02056dd10e7b43e4f16c79df69.

  • All nine modified task packages have fresh framework-finalized task_validator PASS reports for their final task sources on MI355X.
  • Catalog initial actions and independent task runs cover 448 passing tasks out of 459 gfx950 selections. Ten need their dedicated vLLM runtime; the remaining fused-MoE task is unresolved. Catalog batches retain separate source snapshots, including early batches before the final SDK bootstrap fix, and do not establish full final-source validation of every task.
  • The MLA GPU regression test passed. A comparison using identical frozen inputs reproduced 104/120 failures for the original schedule on the new runtime and 0/120 with the fallback; both schedules passed on the old image. Diagnostic timing showed a latency cost, so the fallback is a correctness repair.
  • All five non-ASan tool startup controls passed. Real evaluator-manager candidate analysis passed for Triton FpSan, HIP-FpSan and rocJITsu. Waitcheck and ConSan retain their documented advisory boundaries.
  • python3 -m pytest -q tests/eval_tools tests/test_rocm_sdk_runtime.py: 296 passed on Linux with the final tool/runtime code.
  • The broader related FlyDSL CPU suite passed 4,863 checks, followed by focused checks for the final task repairs. make check-docker-runner and make check-perf-helpers passed. tests/test_mla_runtime_compat.py passed on MI355X.
  • Same-GPU old/new comparisons preserve source and runtime distinctions. The initial large MoE latency outlier did not recur in an old/new/new/old repeat.

Failed and incomplete controls remain failures or incomplete evidence. Runtime promotion remains unqualified, as recorded in the image lock and documentation. Raw reports and experiment artifacts are retained outside the repository.

Pin the new scoring image, adapt SDK loading and task dependencies, and
update all six evaluation-tool runtimes with matching documentation and
regression coverage. Keep qualification pending for GPU ASan and the
existing fused-MoE numerical failures.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant