diff --git a/.pre-commit-config.yaml b/.pre-commit-config.yaml index 128492896..6fb5fe947 100644 --- a/.pre-commit-config.yaml +++ b/.pre-commit-config.yaml @@ -7,7 +7,7 @@ repos: - id: check-yaml exclude: 'mkdocs\.yaml$' - id: check-added-large-files - exclude: 'docs/assets/logo\.png$' + exclude: 'docs/assets/.*' - repo: https://github.com/psf/black rev: 24.4.2 diff --git a/README.md b/README.md index d0a74ce3c..e9c7fedd8 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,29 @@ **RL-Kernel** is a high-performance, memory-efficient infrastructure for Reinforcement Learning post-training. It eliminates the memory and latency bottlenecks in Large Language Model alignment, This project targets AI infrastructure engineers, algorithm researchers, and enterprise-level large model alignment scenarios, providing specialized kernels for algorithms like **GRPO**, **PPO**, and **DPO**. + +--- + +## Our Core Philosophy + +**1. Operator-Level Train-Inference Consistency** +The biggest hidden barrier in large-scale RL is the subtle numerical divergence between rollout engines (e.g., vLLM) and training engines (e.g., Megatron/DeepSpeed). RL-Kernel provides mathematically rigorous, fused operators that lock down the computational graph. By guaranteeing absolute numerical consistency and deterministic reduction orders across the entire RL loop, we prevent reward hacking and distribution drift at the operator level. + +**2. Extreme Memory & Compute Efficiency** +We replace naive PyTorch paths—which suffer from $O(G \cdot L \cdot V)$ memory explosion—with specialized industrial-grade kernels (like `prefix_shared_attention` and `fused_logp`). This reduces VRAM consumption by up to 10x, unlocking massive batch sizes for GRPO workloads without triggering Out-Of-Memory (OOM) errors. + +--- + +## Global Architecture + +RL-Kernel sits strictly at the operator layer, acting as a non-intrusive bridge between high-level alignment orchestration (e.g., vime, slime) and foundational execution engines. We ensure maximum throughput and rigorous numerical parity without modifying upstream framework source code. + +
+
+