Repository navigation
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
RL-Kernel needs a stable user switch as model coverage and upstream BI implementations grow. Selecting individual kernels by microbenchmark speed is insufficient: the training/rollout combination must preserve the numerical contract and improve the measured end-to-end workload.
This PR adds startup execution-plan selection behind
export RL_KERNEL_BI=1, based on main43f150f.User contract: the same switch for future models and upstream reuse
After a model's training, rollout and distributed paths are implemented and validated, and its profile, paired adapter and reference plan are registered, users continue to enable:
export RL_KERNEL_BI=1The same applies after an upstream-reuse PR is reviewed, merged and installed in the user's environment. At the next job startup, RL-Kernel identifies the model and runtime and selects the complete training/rollout plan with the lowest measured E2E time in the matching shipped, consistency-qualified comparison. Users do not need a separate switch per model or to choose the underlying provider manually.
Matching configuration includes model configuration, framework, GPU hardware, train/rollout topology, software/build/source identity, workload and numerical/runtime settings. Timings from another configuration or unrelated benchmark cohort do not qualify.
The initial catalog still contains only the Qwen3-8B reference recipes described below. Future model support and upstream kernel reuse are extension paths, not additional completed model implementations in this PR.
Runtime flow
The two installation branches below belong to one selected paired plan. They do not independently choose the fastest training and rollout kernels.
flowchart TD A["Read RL_KERNEL_BI and training configuration"] --> B{"BI enabled?"} B -->|No| C["Keep existing launcher behavior"] B -->|Yes| D["Identify model, hardware and software environment"] D --> E{"Supported model, reference plan and adapter?"} E -->|No| X["Report unmet conditions and stop"] E -->|Yes| F{"Valid E2E comparison for this exact context?"} F -->|Invalid| X F -->|Available| G["Select the fastest qualified complete plan"] F -->|Absent| H["Use the existing reference recipe"] G --> I["Seal the selected plan and distribute it to all workers"] H --> I I --> J{"Worker plan, configuration and runtime checks pass?"} J -->|No| X J -->|Yes| K["Install training-side implementation"] J -->|Yes| L["Install rollout-side implementation"] K --> M["Compile, capture graphs and run with the fixed plan"] L --> M M --> N["Record actual routes and validate plan agreement and E2E equality"] N --> O{"Validation passed?"} O -->|Yes| P["Report a validated run"] O -->|No| Q["Report failure; no unqualified route substitution"]Worker context checks precede framework adapter installation. Actual operator execution is established by runtime readbacks and E2E validation after execution; an installed hook alone is not execution evidence. Graph compilation/capture occurs where required by the selected framework recipe.
What changes
Initial scope
The shipped reference adapters cover Qwen3-8B BF16 with Vime + Megatron + vLLM on 8×H100 or 8×MI300X, training TP4/CP2 and rollout TP4/CP1. They reuse the existing strict numerical paths and explicitly pin CUDA no-split-K cuBLASLt or ROCm MFMA/chunked attention routing.
DSV4, Gemma, Qwen3-Next and H3 need their own completed paired adapters and evidence before being enabled. Direct unprepared
vllm serveis not supported by this initial launcher integration.No new performance results or candidate timing records are invented in this PR. Initially, the switch selects the shipped reference. A subsequent reviewed upstream-reuse PR can add a real candidate and an exact-context comparison; users keep the same flag after installing that release. Selection uses the installed catalog and remains fixed for the lifetime of each job.
Benchmark metadata is an admission check, not a proof of correctness: the linked WS1/WS2/E2E reports must be reviewed. The promise is the fastest qualified measured plan in the matching shipped comparison, not the globally fastest operator on every workload.
Validation completed
CPU-only environment: Python 3.12, PyTorch 2.14.1+cpu.
204 passed, 3 skipped.
Also passed: Black/isort checks, flake8 on changed Python files,
git diff --check,bash -nfor the ROCm launcher,mkdocs build --strict, CLI conflict/invalid-flag checks, and actual CPU runtime fingerprint collection.The uploaded Git tree was checked against the tested local tree and is identical.
GPU acceptance before merging
This is a draft because no CUDA/ROCm hardware was available. CPU tests do not certify GPU arithmetic, graph replay, integration with the real pinned framework stacks, or performance.
On the existing prepared reference environments:
See
docs/usage/model-bi-selection.mdfor details.