Repository navigation
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
Signed-off-by: vensen <vensenmu@gmail.com>
Signed-off-by: vensen <vensenmu@gmail.com>
Signed-off-by: vensen <vensenmu@gmail.com>
Set OpenMP and MKL to one thread before PyTorch imports in the CPU unit-test job. The reordered cross-configuration suite can initialize a parent thread pool before forked smoke scorers run. Preserve the scoring timeout, numerical assertions, production routing, and GPU benchmark configuration. Signed-off-by: vensen <vensenmu@gmail.com>
Map PR types to code ownership, tests, configuration, benchmarks, and documentation. Explain model and hardware implementation status, train/rollout engine boundaries, numerical evidence, and DCO sign-offs. Add a dedicated Contributor Guide under Developer Guide while preserving all top-level navigation categories. Signed-off-by: vensen <vensenmu@gmail.com>
Render directory paths and inline terms as plain text. Escape angle-bracket placeholders so they remain visible, and preserve the command examples. Signed-off-by: vensen <vensenmu@gmail.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Model code, operator implementations, numerical contracts, framework adapters, experiments, and tests previously shared overlapping locations, making ownership unclear when adding a model or hardware platform. Starting from upstream
mainatbea224b, this PR reorganizes the repository on therefactorbranch into separate model, operator, backend, runtime, and integration layers.The layout follows vLLM's separation of model assembly, reusable layers, platforms, backends, engines, native sources, and tests, with explicit ownership for RL-Kernel's numerical contracts and train–rollout consistency validation.
Scope
Move existing implementations into
models,ops,contracts,reference,backends,runtime,distributed, andvalidation. Platform-specific attention and GEMM implementations belong to their hardware backend. Implementations shared across supported platforms live inbackends/shared/triton; ROCm-specific kernels remain underbackends/rocm, including those written in Triton.Split engines into
integrations/engines/rollout/vllmandintegrations/engines/train/{megatron,deepspeed}. Place engine-specific operator adapters with their engine, shared logic and state inintegrations/common, and VIME experiments, providers, and patches under its orchestrator integration.Extract the Qwen3 Dense model specification and operator inventory. Add empty packages for
deepseek_v4,gemma, andminimax_h3. These placeholders do not register backends or claim model support.Reorganize tests, benchmarks, configurations, tools, CI, native sources, requirements, documentation, and reports. Update Docker dependency paths, GitHub workflow paths, and CODEOWNERS. Move build logic into
build_tools, retain a thinsetup.py, and include contract and experiment resources in wheels and source distributions.Preserve compatibility for existing imports, scripts, and shell entry points. Legacy aliases share registries, caches, and state with their canonical implementations. Preserve backend IDs, dispatch priorities, numerical algorithms, precision thresholds, and strict-mode behavior. Keep WS1 workload and precision JSON files and historical evidence byte-for-byte unchanged.
Reserve runtime packages for future performance-based selection. This PR preserves existing routing and introduces no new fastest-path policy or unvalidated kernel fallback.
Add a dedicated English Contributor Guide under Developer Guide, mapping PR types to directory ownership, model/backend status, train/rollout integrations, tests, benchmarks, and numerical acceptance evidence. Keep all top-level navigation categories unchanged.
Repository layout
The tree below shows the canonical ownership structure. Braces group sibling directories or files. Existing compatibility paths remain available but are omitted from this overview. Reserved packages are extension points; directory presence alone does not establish model, engine, or hardware support.
Models own topology and operator requirements.
opsowns semantic entry points,contractsowns numerical requirements, andbackendsowns implementations. A common attention or GEMM implementation is shared across models; its location follows implementation ownership. Engine tensor layouts and hooks belong to the corresponding train or rollout adapter.WS1 operator validation belongs in
validation/operatorsand its matching tests; backend-specific tests live intests/backends. WS2 collective and partition tests live intests/distributed. Ablation, model, cross-configuration, and post-training integration checks have their own corresponding directories. Benchmarks contain executable measurement code, while reports contain captured results and historical evidence.Strict routing must establish numerical-contract eligibility, supported topology, precision, and validated provenance before performance can rank eligible implementations. The reserved selector and performance packages prepare that ownership boundary; existing dispatch behavior is retained in this PR.
Contributor Guide · Layout and ownership guide · Migration map · Local validation records
Local validation
Validation ran on a macOS ARM64 CPU environment. CUDA and ROCm GPU acceptance has not been run locally.
The existing baseline failures concern a missing CPU
_Cextension, a ROCm capability expectation, a missingget_loss_op, a GRPO negative control with the installed PyTorch version, and a macOS Gloo CLI timeout. Collection excluded the same four modules requiring unavailable Triton and one module referencing an AITER symbol missing from the baseline in both trees. The validation records document these exclusions; existing failures were not converted to skips or expected failures. An additional MyPy check with PyTorch type information also produced the same 14 existing errors on both trees.Dense model acceptance on CUDA and ROCm
Acceptance procedure and evidence checklist
CUDA and ROCm are accepted separately. Local CPU checks and historical GPU reports do not establish GPU acceptance for this PR. The PR remains a draft while the above GPU results are collected.
ROCm regression validation (rounds 2 and 3; warmup excluded)
Validated on the refactored tree with the source content now included in
33b7ea2b3bf.Environment/workload: 8x AMD MI300X (
gfx942), PyTorch2.12.0+rocm7.14.0a20260608, HIP7.14.60850, Triton3.7.0+gitb4e20bbe.rocm7.14.0a20260608, vLLM0.26.1.dev, Qwen3-8B, train TP4/CP2, rollout TP4, 3 rollouts x 8 samples, max response 6912, max tokens/GPU 4096, temperature 0.7, top-p 0.95, LR 5e-7, KL coefficient 0.01. Round 1 is treated as warmup; all tables below aggregate rounds 2 and 3 only.Consistency results — rounds 2 and 3 only
torch.equalR/R also reports
bitwise_equal=true,train_rollout_logprob_abs_diff=0, KL=0, 16/16 selected samples validated, no validation errors, and unchanged frozen sources during the run.Average performance — rounds 2 and 3 only
R/R is not faster in every sub-phase: after excluding warmup it wins 4/9 reported aggregate metrics (rollout time, effective rollout throughput, update weights, and end-to-end step), while reference/logp and training throughput remain regressions. The long-context round (all eight responses truncated at 6912 tokens) is the largest win: rollout 850.629 s -> 183.589 s.
Per-round details
Runtime path evidence
megatron.production.*; vLLM attention/FFN/logp usevllm.production.*. The native training logp marker is present. Aggregated runtime validation passed with zero errors.rlkernel.attention.deterministic.v1,rlkernel.ffn.qwen3.deterministic.v1, andrlkernel.linear_logp.bitwise.v1. Runtime validation observed 26,496 training attention/FFN calls, 552 training logp calls, 8,640 rollout attention calls, and 8 rollout FFN/logp calls.Regressions found and fixed during validation
Tests: 212 targeted ROCm/VIME/vLLM integration tests passed. Both 3-round E2E arms passed with zero validation errors; the reported comparison intentionally uses only rounds 2 and 3. Timing runs were collected only after removing unrelated platform GPU workloads; P/P had zero guard hits, and R/R had zero active injected processes before
rollout 0, with the conflicting Ray actor name reserved by an inert zero-resource actor for all timed rounds.