Skip to content

feat: select paired BI execution plans with RL_KERNEL_BI=1 - #477

Draft
Flink-ddd wants to merge 2 commits into
mainfrom
feat/model-bi-plan-selection
Draft

Flink-ddd wants to merge 2 commits into
mainfrom
feat/model-bi-plan-selection

Conversation

@Flink-ddd

@Flink-ddd Flink-ddd commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Why

RL-Kernel needs a stable user switch as model coverage and upstream BI implementations grow. Selecting individual kernels by microbenchmark speed is insufficient: the training/rollout combination must preserve the numerical contract and improve the measured end-to-end workload.

This PR adds startup execution-plan selection behind export RL_KERNEL_BI=1, based on main 43f150f.

User contract: the same switch for future models and upstream reuse

After a model's training, rollout and distributed paths are implemented and validated, and its profile, paired adapter and reference plan are registered, users continue to enable:

export RL_KERNEL_BI=1

The same applies after an upstream-reuse PR is reviewed, merged and installed in the user's environment. At the next job startup, RL-Kernel identifies the model and runtime and selects the complete training/rollout plan with the lowest measured E2E time in the matching shipped, consistency-qualified comparison. Users do not need a separate switch per model or to choose the underlying provider manually.

Contribution Required before automatic selection
New model, such as DSV4, Gemma, Qwen3-Next or H3 Model profile, complete paired framework adapter, supported-scope checks, reference plan and consistency evidence. Add comparable E2E timings to enable performance-based choice between plans.
Upstream kernel reuse Integrate the implementation into a complete paired plan; pass WS1, WS2 and nonempty E2E equality checks; register its plan digest and measured E2E time alongside the reference in the same comparison.
Publication Human review of implementation and evidence, merge, installation of the updated version, then a new job startup. Merging kernel code alone does not admit it to automatic performance selection.

Matching configuration includes model configuration, framework, GPU hardware, train/rollout topology, software/build/source identity, workload and numerical/runtime settings. Timings from another configuration or unrelated benchmark cohort do not qualify.

  • A matching comparison selects the fastest qualified complete plan; a tie retains the reference.
  • A supported stack with no matching candidate comparison uses its existing reference recipe.
  • A missing model/adapter, unsupported scope, invalid comparison or worker mismatch stops startup and reports the unmet condition.
  • Selection happens once at startup. Running jobs keep their selected plan through compilation, graph replay and weight updates. Existing strict numerical/fallback checks and E2E validators remain required.

The initial catalog still contains only the Qwen3-8B reference recipes described below. Future model support and upstream kernel reuse are extension paths, not additional completed model implementations in this PR.

Runtime flow

The two installation branches below belong to one selected paired plan. They do not independently choose the fastest training and rollout kernels.

flowchart TD
    A["Read RL_KERNEL_BI and training configuration"] --> B{"BI enabled?"}
    B -->|No| C["Keep existing launcher behavior"]
    B -->|Yes| D["Identify model, hardware and software environment"]
    D --> E{"Supported model, reference plan and adapter?"}
    E -->|No| X["Report unmet conditions and stop"]
    E -->|Yes| F{"Valid E2E comparison for this exact context?"}
    F -->|Invalid| X
    F -->|Available| G["Select the fastest qualified complete plan"]
    F -->|Absent| H["Use the existing reference recipe"]
    G --> I["Seal the selected plan and distribute it to all workers"]
    H --> I
    I --> J{"Worker plan, configuration and runtime checks pass?"}
    J -->|No| X
    J -->|Yes| K["Install training-side implementation"]
    J -->|Yes| L["Install rollout-side implementation"]
    K --> M["Compile, capture graphs and run with the fixed plan"]
    L --> M
    M --> N["Record actual routes and validate plan agreement and E2E equality"]
    N --> O{"Validation passed?"}
    O -->|Yes| P["Report a validated run"]
    O -->|No| Q["Report failure; no unqualified route substitution"]
Loading

Worker context checks precede framework adapter installation. Actual operator execution is established by runtime readbacks and E2E validation after execution; an installed hook alone is not execution evidence. Graph compilation/capture occurs where required by the selected framework recipe.

What changes

  • Add a model/config registry, immutable paired execution plans, framework/model adapters, runtime-context fingerprints, and reviewed E2E comparison records.
  • Select the fastest complete plan in a matching comparison. Require passing WS1/WS2, nonempty zero-mismatch/zero-difference E2E evidence, valid timings, and the incumbent reference in the same comparison. Retain the reference on ties or when no comparable candidate evidence exists.
  • Connect the CUDA and ROCm Qwen3 Vime launchers to preparation before Ray submission. Export the selected routing settings and a versioned plan envelope to every worker.
  • Dispatch Megatron/vLLM initialization through the selected paired adapter. Verify model config, package/build/source identity, routing settings and worker GPU compatibility. Frontend plugin discovery does not initialize CUDA.
  • Include plan/context identity in readbacks and validate agreement across training/rollout against the launch manifest. Partition ROCm graph caches by the selected plan/context.
  • Reject native/mixed-arm conflicts, unsupported models/configurations, missing preparation, changed worker identities, and attempts to replace an active plan. Keep existing strict runtime mismatch/fallback checks.
  • Document user commands, contributor registration interfaces, scope and GPU acceptance steps.

Initial scope

The shipped reference adapters cover Qwen3-8B BF16 with Vime + Megatron + vLLM on 8×H100 or 8×MI300X, training TP4/CP2 and rollout TP4/CP1. They reuse the existing strict numerical paths and explicitly pin CUDA no-split-K cuBLASLt or ROCm MFMA/chunked attention routing.

DSV4, Gemma, Qwen3-Next and H3 need their own completed paired adapters and evidence before being enabled. Direct unprepared vllm serve is not supported by this initial launcher integration.

No new performance results or candidate timing records are invented in this PR. Initially, the switch selects the shipped reference. A subsequent reviewed upstream-reuse PR can add a real candidate and an exact-context comparison; users keep the same flag after installing that release. Selection uses the installed catalog and remains fixed for the lifetime of each job.

Benchmark metadata is an admission check, not a proof of correctness: the linked WS1/WS2/E2E reports must be reviewed. The promise is the fastest qualified measured plan in the matching shipped comparison, not the globally fastest operator on every workload.

Validation completed

CPU-only environment: Python 3.12, PyTorch 2.14.1+cpu.

python -m pytest \
  tests/test_bi_plan_selection.py tests/test_bi_launcher_integration.py \
  tests/test_repro_launcher.py tests/test_repro_platform_alignment.py \
  tests/test_rocm_repro_profile.py tests/test_vime_tp4_example.py \
  tests/test_vime_rocm_attention_topology.py tests/test_vime_rocm_module_validation.py \
  tests/test_framework_runtime_adapters.py tests/test_framework_operator_integrations.py \
  tests/test_configurable_validation.py tests/test_pcp_validation.py -q

204 passed, 3 skipped.

Also passed: Black/isort checks, flake8 on changed Python files, git diff --check, bash -n for the ROCm launcher, mkdocs build --strict, CLI conflict/invalid-flag checks, and actual CPU runtime fingerprint collection.

The uploaded Git tree was checked against the tested local tree and is identical.

GPU acceptance before merging

This is a draft because no CUDA/ROCm hardware was available. CPU tests do not certify GPU arithmetic, graph replay, integration with the real pinned framework stacks, or performance.

On the existing prepared reference environments:

# 8×H100
export RL_KERNEL_BI=1
./rlk run --mode consistency --steps 2 --require-updates --wait

# 8×MI300X
./rlk run --profile examples/vime_rocm_attention_ablation/profiles/mi300x-qwen3-8b.json \
  --mode consistency --steps 8 --wait
  • Require matching plan/context identities in the manifest and all training/rollout readbacks.
  • Require nonempty E2E comparisons with mismatch count 0 and maximum absolute difference 0, successful weight refresh, graph capture/replay, and no fallback.
  • Repeat the existing 200-step workload with BI enabled versus the existing strict route with BI unset, using the same seeds/inputs/measurement protocol. Attach results and confirm no steady-state regression.
  • Any future upstream candidate needs its own complete WS1, WS2 and E2E consistency/performance comparison before selection is enabled.

See docs/usage/model-bi-selection.md for details.

@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant