Skip to content

refactor: organize models, backends, contracts and train and rollout engines - #497

Draft
Flink-ddd wants to merge 7 commits into
RL-Align:mainfrom
Flink-ddd:refactor
Draft

Flink-ddd wants to merge 7 commits into
RL-Align:mainfrom
Flink-ddd:refactor

Conversation

@Flink-ddd

@Flink-ddd Flink-ddd commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Model code, operator implementations, numerical contracts, framework adapters, experiments, and tests previously shared overlapping locations, making ownership unclear when adding a model or hardware platform. Starting from upstream main at bea224b, this PR reorganizes the repository on the refactor branch into separate model, operator, backend, runtime, and integration layers.

The layout follows vLLM's separation of model assembly, reusable layers, platforms, backends, engines, native sources, and tests, with explicit ownership for RL-Kernel's numerical contracts and train–rollout consistency validation.

Scope

  • Move existing implementations into models, ops, contracts, reference, backends, runtime, distributed, and validation. Platform-specific attention and GEMM implementations belong to their hardware backend. Implementations shared across supported platforms live in backends/shared/triton; ROCm-specific kernels remain under backends/rocm, including those written in Triton.

  • Split engines into integrations/engines/rollout/vllm and integrations/engines/train/{megatron,deepspeed}. Place engine-specific operator adapters with their engine, shared logic and state in integrations/common, and VIME experiments, providers, and patches under its orchestrator integration.

  • Extract the Qwen3 Dense model specification and operator inventory. Add empty packages for deepseek_v4, gemma, and minimax_h3. These placeholders do not register backends or claim model support.

  • Reorganize tests, benchmarks, configurations, tools, CI, native sources, requirements, documentation, and reports. Update Docker dependency paths, GitHub workflow paths, and CODEOWNERS. Move build logic into build_tools, retain a thin setup.py, and include contract and experiment resources in wheels and source distributions.

  • Preserve compatibility for existing imports, scripts, and shell entry points. Legacy aliases share registries, caches, and state with their canonical implementations. Preserve backend IDs, dispatch priorities, numerical algorithms, precision thresholds, and strict-mode behavior. Keep WS1 workload and precision JSON files and historical evidence byte-for-byte unchanged.

  • Reserve runtime packages for future performance-based selection. This PR preserves existing routing and introduces no new fastest-path policy or unvalidated kernel fallback.

  • Add a dedicated English Contributor Guide under Developer Guide, mapping PR types to directory ownership, model/backend status, train/rollout integrations, tests, benchmarks, and numerical acceptance evidence. Keep all top-level navigation categories unchanged.

Repository layout

The tree below shows the canonical ownership structure. Braces group sibling directories or files. Existing compatibility paths remain available but are omitted from this overview. Reserved packages are extension points; directory presence alone does not establish model, engine, or hardware support.

RL-Kernel/
├── rl_engine/
│   ├── config/                         # Configuration and workload loaders
│   │   └── workloads/                  # Packaged, versioned WS1 workload identity
│   ├── models/
│   │   ├── qwen3/                      # Dense specification, operator inventory, assembly
│   │   ├── deepseek_v4/                 # Reserved package
│   │   ├── gemma/                       # Reserved package
│   │   └── minimax_h3/                  # Reserved package
│   ├── layers/                         # Reusable model layers; extension point
│   ├── ops/                            # Hardware-independent operator entry points
│   │   └── {attention,gemm,norm,activation,rope,embedding,moe,logprob,loss,sampling,packing,autograd}/
│   ├── contracts/
│   │   ├── operators/                  # Operator numerical contracts
│   │   ├── diagnostics/                # Ablation axes and reporting schema
│   │   └── profiles/{precision,invariance}/
│   ├── reference/                      # Mathematical references and diagnostic paths
│   ├── platforms/                      # Device discovery and capabilities
│   ├── backends/
│   │   ├── {cuda,rocm,musa,ascend,cpu}/
│   │   │   └── {attention,gemm,norm,activation,rope,embedding,moe,logprob,loss,sampling,packing}/
│   │   └── shared/triton/              # Shared implementations for supported devices
│   ├── runtime/
│   │   ├── registry.py                 # Existing contract-aware dispatch and priorities
│   │   ├── semantic_registry.py        # Capabilities, support checks, provenance
│   │   ├── operators.py                # Operator bridge and execution bindings
│   │   ├── policy.py                   # Strict and diagnostic execution policies
│   │   ├── plan.py                     # P/P, P/R, R/P, R/R execution selections
│   │   ├── executor.py
│   │   ├── provenance/                 # Stable, versioned operator identities
│   │   └── {selector,performance}/     # Reserved for measured selection policies
│   ├── distributed/
│   │   ├── algorithms/                 # Canonical arithmetic and collective ordering
│   │   └── transports/                 # RCCL and transport bindings
│   ├── integrations/
│   │   ├── common/                     # Shared adapters, weights, and state
│   │   ├── engines/
│   │   │   ├── rollout/
│   │   │   │   ├── interface.py
│   │   │   │   └── vllm/
│   │   │   └── train/
│   │   │       ├── contract.py
│   │   │       ├── megatron/
│   │   │       └── deepspeed/
│   │   └── orchestrators/
│   │       ├── vime/{providers,experiments,patches}/
│   │       └── {miles,areal}/           # Reserved integration points
│   ├── validation/
│   │   └── {operators,models,distributed,cross_config,ablation,reports,common,fixtures}/
│   ├── entrypoints/                    # Installed command-line entry points
│   └── utils/
├── csrc/
│   ├── bindings/                       # Python/native registration
│   ├── common/
│   └── {cuda,rocm,musa,ascend}/          # Native sources grouped by platform and operator
├── build_tools/                        # Extension selection, compiler flags, environment
├── configs/
│   ├── {models,hardware,policies,workloads,ablations,benchmarks,tuning,local}/
│   ├── suites/{ws1,ws2,e2e}/
│   ├── experiments/{cross_config,vime}/
│   └── integrations/
│       ├── engines/{rollout,train}/
│       └── orchestrators/vime/
├── tests/
│   ├── {contracts,reference,ops,layers,backends,platforms,runtime,models}/
│   ├── distributed/{collectives,transports,tp,cp,ep,dp}/
│   ├── integrations/
│   │   ├── common/
│   │   ├── engines/
│   │   │   ├── rollout/vllm/
│   │   │   └── train/{megatron,deepspeed}/
│   │   └── orchestrators/
│   ├── validation/{operators,models,cross_config,ablation,common}/
│   └── {benchmarks,e2e,entrypoints,build,helpers,data}/
├── benchmarks/
│   └── {common,operators,layers,backends,distributed,models,e2e,tuning,profiling}/
├── examples/{operators,post_training}/
├── bin/
│   ├── rlk                             # Thin checkout launcher
│   └── rlk-repro
├── tools/
│   └── {validation,weights,env,checks,benchmarking,migration}/
├── ci/
│   ├── run.py
│   └── {scripts,jobs,runners,providers}/
├── .github/
│   └── workflows/                      # CI triggers and security policy
├── docker/
│   ├── Dockerfile.cuda
│   ├── Dockerfile.rocm
│   └── Dockerfile.rocm_base
├── requirements/
│   ├── {runtime,test,dev,docs,benchmark,build}.txt
│   └── {integrations,constraints}/     # Reserved dependency organization
├── reports/{archive,experiments,releases}/
├── docs/
│   ├── {architecture,contracts,operators,integrations,validation,benchmarking}/
│   ├── {getting_started,usage,cli,api,contributing,community}/
│   └── {design,blog,archive,assets,mkdocs}/
├── artifacts/                          # Ignored local validation and build output
├── rlk                                 # Existing root launcher remains supported
├── requirements.txt                    # Compatibility dependency entry point
├── pyproject.toml                      # Package metadata, entry points, test configuration
├── setup.py                            # Thin build entry point
└── MANIFEST.in                         # Source distribution and packaged resources

Models own topology and operator requirements. ops owns semantic entry points, contracts owns numerical requirements, and backends owns implementations. A common attention or GEMM implementation is shared across models; its location follows implementation ownership. Engine tensor layouts and hooks belong to the corresponding train or rollout adapter.

WS1 operator validation belongs in validation/operators and its matching tests; backend-specific tests live in tests/backends. WS2 collective and partition tests live in tests/distributed. Ablation, model, cross-configuration, and post-training integration checks have their own corresponding directories. Benchmarks contain executable measurement code, while reports contain captured results and historical evidence.

Strict routing must establish numerical-contract eligibility, supported topology, precision, and validated provenance before performance can rank eligible implementations. The reserved selector and performance packages prepare that ownership boundary; existing dispatch behavior is retained in this PR.

Contributor Guide · Layout and ownership guide · Migration map · Local validation records

Local validation

Validation ran on a macOS ARM64 CPU environment. CUDA and ROCm GPU acceptance has not been run locally.

Check Result
Existing CPU-executable tests, using the same environment for the main baseline and refactor Both: 1832 passed / 1920 skipped / 5 failed, with the same failure identities
New layout, packaging, shared module state, and legacy CLI tests outside the checkout 28 passed / 1 skipped; the skip requires a native extension
Focused engine, VIME, WS1 path migration, and related checks 232 passed / 23 skipped
Native source and Python operator audit 35 native source files and 90 Python operator modules; computational logic unchanged, with differences limited to imports, comments, includes, and source-path resolution
Historical archives and numerical JSON files 74 archived files and two numerical JSON files remain byte-identical, including original line endings
pre-commit, strict MkDocs, and MyPy in the CI-equivalent environment Passed
Contributor Guide navigation and rendered content Strict documentation build and pre-commit passed; unchanged top-level navigation, one dedicated guide entry, 37 local links/anchors, and nine tables verified
Source distribution and wheel builds; installed-wheel imports, resources, and all entry points outside the checkout Passed

The existing baseline failures concern a missing CPU _C extension, a ROCm capability expectation, a missing get_loss_op, a GRPO negative control with the installed PyTorch version, and a macOS Gloo CLI timeout. Collection excluded the same four modules requiring unavailable Triton and one module referencing an AITER symbol missing from the baseline in both trees. The validation records document these exclusions; existing failures were not converted to skips or expected failures. An additional MyPy check with PyTorch type information also produced the same 14 existing errors on both trees.

Dense model acceptance on CUDA and ROCm

Acceptance procedure and evidence checklist

  • CUDA: Compare the main baseline and PR using the same environment, weights, workloads, and supported topologies across WS1, WS2, Dense model validation, and post-training integrations.
  • ROCm: Run the corresponding comparison over the scope already supported on ROCm.
  • Recompute train and rollout log probabilities independently. Require zero raw-bit mismatches after aligning logically active tokens. Do not reuse rollout scores, loosen numerical contracts, or rely on silent fallback to pass.
  • Record the actual backend, precision, topology, compiler and dependency versions, gradient/update coverage, and raw results.
  • After warmup, repeat measurements and report median latency, throughput, and memory differences. Performance differences are reported for reviewer acceptance; no automatic percentage regression threshold is imposed.

CUDA and ROCm are accepted separately. Local CPU checks and historical GPU reports do not establish GPU acceptance for this PR. The PR remains a draft while the above GPU results are collected.

ROCm regression validation (rounds 2 and 3; warmup excluded)

Validated on the refactored tree with the source content now included in 33b7ea2b3bf.

Environment/workload: 8x AMD MI300X (gfx942), PyTorch 2.12.0+rocm7.14.0a20260608, HIP 7.14.60850, Triton 3.7.0+gitb4e20bbe.rocm7.14.0a20260608, vLLM 0.26.1.dev, Qwen3-8B, train TP4/CP2, rollout TP4, 3 rollouts x 8 samples, max response 6912, max tokens/GPU 4096, temperature 0.7, top-p 0.95, LR 5e-7, KL coefficient 0.01. Round 1 is treated as warmup; all tables below aggregate rounds 2 and 3 only.

Consistency results — rounds 2 and 3 only

Configuration Mismatch Count Max |dlogp| torch.equal
P/P native 25,282 / 94,670 4.608247 false
R/R strict 0 / 96,166 0 true

R/R also reports bitwise_equal=true, train_rollout_logprob_abs_diff=0, KL=0, 16/16 selected samples validated, no validation errors, and unchanged frozen sources during the run.

Average performance — rounds 2 and 3 only

Metric P/P native R/R strict R/R relative to P/P
rollout time 482.659528 s 166.431715 s 65.52% faster
effective tokens/GPU/s 25.519699 35.936592 40.82% higher
update weights 6.256760 s 4.046481 s 35.33% faster
reference log probs 20.269574 s 34.024931 s 67.86% slower
log probs 21.250762 s 31.437890 s 47.94% slower
actor train 55.198876 s 73.521580 s 33.19% slower
train time 98.658031 s 139.699000 s 41.60% slower
actor train tok/s 885.434875 675.129916 23.75% lower
end-to-end step 595.777345 s 321.794659 s 45.99% faster

R/R is not faster in every sub-phase: after excluding warmup it wins 4/9 reported aggregate metrics (rollout time, effective rollout throughput, update weights, and end-to-end step), while reference/logp and training throughput remain regressions. The long-context round (all eight responses truncated at 6912 tokens) is the largest win: rollout 850.629 s -> 183.589 s.

Per-round details

Round P/P rollout R/R rollout P/P actor train R/R actor train P/P step R/R step
2 114.690 s 149.275 s 55.562 s 74.172 s 226.079 s 329.240 s
3 850.629 s 183.589 s 54.835 s 72.871 s 965.475 s 314.349 s

Runtime path evidence

  • P/P is native production code, with no fallback: Megatron attention/FFN use megatron.production.*; vLLM attention/FFN/logp use vllm.production.*. The native training logp marker is present. Aggregated runtime validation passed with zero errors.
  • R/R is RL-Kernel strict code, with no fallback: both Megatron and vLLM report rlkernel.attention.deterministic.v1, rlkernel.ffn.qwen3.deterministic.v1, and rlkernel.linear_logp.bitwise.v1. Runtime validation observed 26,496 training attention/FFN calls, 552 training logp calls, 8,640 rollout attention calls, and 8 rollout FFN/logp calls.
  • P/P and R/R used identical sealed model/data inputs, launcher content, VIME/Megatron/vLLM/RL-Kernel source content, and all common workload parameters; only the requested P/P vs R/R case selection differs.

Regressions found and fixed during validation

  • Dense CK strict attention incorrectly selected AITER's paged batch entry for batch sizes greater than one, breaking TP4/TP8 exactness. CK now retains the pinned one-row/one-KV-group schedule; the separately validated Triton dense schedule keeps its direct path.
  • Native vLLM full-graph execution completed FFN inside the compiled HIP graph, so Python hooks could not emit production execution evidence. A post-model sampler boundary now records native FFN completion without changing P/P computation.

Tests: 212 targeted ROCm/VIME/vLLM integration tests passed. Both 3-round E2E arms passed with zero validation errors; the reported comparison intentionally uses only rounds 2 and 3. Timing runs were collected only after removing unrelated platform GPU workloads; P/P had zero guard hits, and R/R had zero active injected processes before rollout 0, with the conflicting Ray actor name reserved by an inert zero-resource actor for all timed rounds.

@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@Flink-ddd Flink-ddd changed the title refactor: organize models, backends, contracts and train/rollout engines refactor: organize models, backends, contracts and train and rollout engines Oct 8, 2026
Signed-off-by: vensen <vensenmu@gmail.com>
Signed-off-by: vensen <vensenmu@gmail.com>
Set OpenMP and MKL to one thread before PyTorch imports in the CPU unit-test job. The reordered cross-configuration suite can initialize a parent thread pool before forked smoke scorers run. Preserve the scoring timeout, numerical assertions, production routing, and GPU benchmark configuration.

Signed-off-by: vensen <vensenmu@gmail.com>
@Flink-ddd Flink-ddd added type: refactor Code refactoring enhancement New feature or request component: kernels Tasks involving the development of CUDA and Triton underlying operators component: alignment Tasks involving RL loss functions such as DPO and GRPO, and mathematical alignment logic component: executors Tasks involving the interaction of vLLM inference and DeepSpeed ​​training endpoints. platform: cuda Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations) platform: rocm Specific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA) platform: triton Cross-platform Triton kernel related tasks Ascend priority: high Severe congestion issues require the highest priority for resolution. labels Oct 8, 2026
@Flink-ddd Flink-ddd added the MUSA label Oct 8, 2026
Flink-ddd and others added 3 commits October 8, 2026 19:24
Map PR types to code ownership, tests, configuration, benchmarks, and documentation. Explain model and hardware implementation status, train/rollout engine boundaries, numerical evidence, and DCO sign-offs. Add a dedicated Contributor Guide under Developer Guide while preserving all top-level navigation categories.

Signed-off-by: vensen <vensenmu@gmail.com>
Render directory paths and inline terms as plain text. Escape angle-bracket placeholders so they remain visible, and preserve the command examples.

Signed-off-by: vensen <vensenmu@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Ascend component: alignment Tasks involving RL loss functions such as DPO and GRPO, and mathematical alignment logic component: executors Tasks involving the interaction of vLLM inference and DeepSpeed ​​training endpoints. component: kernels Tasks involving the development of CUDA and Triton underlying operators enhancement New feature or request MUSA platform: cuda Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations) platform: rocm Specific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA) platform: triton Cross-platform Triton kernel related tasks priority: high Severe congestion issues require the highest priority for resolution. type: refactor Code refactoring

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants