diff --git a/README.md b/README.md index 1cde9f794..f0e221b31 100644 --- a/README.md +++ b/README.md @@ -33,7 +33,7 @@ Compare RPent with reference methods on LIBERO, LIBERO-PRO, RoboCasa365 Target50 Codex / GPT-6 Astra / low / reasoning: **92.63% Overall (741/800)** across all eight LIBERO-PRO suites. See the [suite results and memory-batch explanation](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html#libero-pro-astra-memory), including the separately frozen Long and Spatial/Object/Goal memory batches. -[](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html) +[](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html) ## Who Should Consider Using RPent? diff --git a/README.zh-CN.md b/README.zh-CN.md index 6441d0d36..52061bd2e 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -33,7 +33,7 @@ Codex / GPT-6 Astra / low / reasoning 已完成全部八套 LIBERO-PRO,**Overall 92.63%(741/800)**。详见 [套件汇总与 memory 批次说明](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html#libero-pro-astra-memory),其中 Long 与 Spatial/Object/Goal 分别使用各自冻结的 memory 批次。 -[](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html) +[](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html) ## 适用用户 diff --git a/docs/source-en/rst_source/get_started/overview.rst b/docs/source-en/rst_source/get_started/overview.rst index b89c496f0..aa4dd34d6 100644 --- a/docs/source-en/rst_source/get_started/overview.rst +++ b/docs/source-en/rst_source/get_started/overview.rst @@ -32,13 +32,13 @@ Compare success rates on LIBERO, LIBERO-PRO, RoboCasa365 Target50, and RoboTwin C2R. Rankings apply to the methods and evaluation coverage shown; see :doc:`../leaderboard/index` for detailed results, configurations, and sources. -.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-light.png +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-light.png :alt: RPent Leaderboard :class: only-light :width: 100% :target: ../leaderboard/index.html -.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-dark.png +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-dark.png :alt: RPent Leaderboard :class: only-dark :width: 100% diff --git a/docs/source-en/rst_source/leaderboard/index.rst b/docs/source-en/rst_source/leaderboard/index.rst index a14e518d7..e4b48b3eb 100644 --- a/docs/source-en/rst_source/leaderboard/index.rst +++ b/docs/source-en/rst_source/leaderboard/index.rst @@ -11,7 +11,7 @@ RPent Leaderboard .. raw:: html - +
diff --git a/docs/source-en/rst_source/leaderboard/performance.rst b/docs/source-en/rst_source/leaderboard/performance.rst index b7089934b..a6d4b8baa 100644 --- a/docs/source-en/rst_source/leaderboard/performance.rst +++ b/docs/source-en/rst_source/leaderboard/performance.rst @@ -6,13 +6,13 @@ Performance .. raw:: htmlRPent Leaderboard
- - - - + + + +| Method | Success rate |
|---|---|
| Codex / GPT-6 Astra / low / reasoning[4] | 92.63% |
| Claude Code / Opus-4.7 / max.reasoning | 82.4% |
| Codex / GPT-5.5 / xhigh / reasoning | 75.13% |
| RPent Flash Mode[2] | 72.63% |
| Qwen3.6 27B / no-reasoning | 70.63% |
| ASPIRE[1] | 61.36% |
| π_RLinf | 50.0% |
| π0.5 | 11.0% |
| AtomVLA | 6.3% |
| X-VLA | 3.8% |
| MolmoAct | 1.5% |
| π0 | 0.3% |
* GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800). Qwen3.6 27B / no-reasoning: 70.63% (565/800).
| Method | Overall | Spatial Task | Spatial Swap | Object Task | Object Swap | Goal Task | Goal Swap | Long Task | Long Swap |
|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra[4] | 92.63% | 100% | 98% | 100% | 99% | 88% | 99% | 85% | 72% |
| Opus-4.7 / max.reasoning | 82.4% | — | — | — | — | — | — | — | — |
| GPT-5.5 | 75.13% | 81.0% | 69.0% | 94.0% | 91.0% | 75.0% | 66.0% | 70.00% | 55.00% |
| RPent Flash Mode / Molmo2-8B[2] | 72.63% | 79.00% | 70.00% | 86.00% | 93.00% | 74.00% | 65.00% | 60.00% | 54.00% |
| Qwen3.6 27B / no-reasoning | 70.63% | 82.00% | 78.00% | 83.00% | 84.00% | 68.00% | 68.00% | 61.00% | 41.00% |
| ASPIRE[1] | 61.36% | 60.0% | 51.0% | 95.0% | 98.0% | 45.0% | 81.0% | 38.3% | 22.6% |
| π_RLinf | 50.0% | 42.0% | 59.0% | 71.0% | 78.0% | 45.0% | 42.0% | 49.0% | 14.0% |
| π0.5 | 11.0% | 1.0% | 20.0% | 1.0% | 17.0% | 2.0% | 38.0% | 1.0% | 8.0% |
| AtomVLA | 6.3% | 1.0% | 16.0% | 0.0% | 10.0% | 11.0% | 2.0% | 9.0% | 1.0% |
| X-VLA | 3.8% | 0.0% | 0.0% | 8.0% | 2.0% | 9.0% | 1.0% | 10.0% | 0.0% |
| MolmoAct | 1.5% | 0.0% | 0.0% | 0.0% | 6.0% | 0.0% | 0.0% | 6.0% | 0.0% |
| π0 | 0.3% | 0.0% | 0.0% | 0.0% | 2.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| Cap-X | — | 14.0% | 12.0% | 18.0% | 22.0% | 17.0% | 26.0% | — | — |
| RPent / GPT-6 Motor Only[3] | — | — | — | — | — | — | — | 38.0% | — |
| RATS | — | 31.0% | 29.0% | 63.0% | 61.0% | 36.0% | 43.0% | — | — |
* [1] ASPIRE: Long Task and Long Swap use zero-shot transfer from the LIBERO-90 skill library.
* [2] RPent Flash Mode uses directly downloaded, officially released GPT-5.5 exploration memory; Molmo2-8B is used for visual localization. These results use the best-performing seed from s0-s9.
* [3] RPent / GPT-6 Motor Only uses motor-only control: it directly outputs end-effector pose increments and gripper commands through execute_action, without invoking VLA / primitives or loading memory.
* [4] GPT-6 Astra (memory-enabled configuration): Long Task/Swap and the other six suites use separate memory-file snapshots frozen after their respective exploration phases, with no updates during evaluation. Overall combines two non-overlapping batches: Long 157/200 plus the other suites 584/600, giving 741/800 (92.63%); the 800 episodes do not share a single memory snapshot.
| Method | Success rate |
|---|---|
| Codex / GPT-6 Astra / low / reasoning | 59.20% |
| Xiaomi-Robotics-1 | 57.4% |
| Codex / GPT-5.5 / xhigh / reasoning | 57.1% |
| Claude Code / Opus-4.7 / max.reasoning | 48.6% |
| WorldDreamer | 35.3% |
| RLDX-1 | 30.0% |
| π0.5 | 16.9% |
| π0 | 14.8% |
| Method | Overall | Atomic-Seen | Composite-Seen | Composite-Unseen |
|---|---|---|---|---|
| GPT-6 Astra | 59.20% | 87.78% | 43.75% | 42.50% |
| Xiaomi-Robotics-1 | 57.4% | 80.2% | 57.1% | 32.1% |
| GPT-5.5 | 57.1% | 92.0% | 61.0% | 13.8% |
| Opus-4.7 / max.reasoning | 48.6% | 79.4% | 47.5% | 15.0% |
| WorldDreamer | 35.3% | 66.3% | 26.7% | 9.0% |
| RLDX-1 | 30.0% | 60.0% | 21.3% | 5.0% |
| π0.5 | 16.9% | 39.6% | 7.1% | 1.2% |
| π0 | 14.8% | 34.6% | 6.1% | 1.1% |
| Method | Success rate |
|---|---|
| Codex / GPT-5.5 / xhigh / reasoning | 62.4% |
| Claude Code / Opus-4.7 / max.reasoning | 58.4% |
| LingBot-VLA | 50.4% |
| π0.5 | 47.9% |
| GR00T-N1.7 | 20.7% |
| StarVLA | 10.6% |
RPent Leaderboard
- - - - + + + +| Method | Mean time / episode (s) | Total output tokens |
|---|---|---|
| RPent / GPT-6 Astra Codex · low · reasoning | 412.14 | 3,364,938 |
| RPent / GPT-5.6 Sol Codex · xhigh · reasoning | 529.23 | 3,635,426 |
| RPent / GPT-5.6 Sol Codex · no-reasoning | 344.17 | 2,569,676 |
| RPent Flash Mode Molmo2-8B · visual localization | 60.19 | 0 |
| RPent / Qwen3.6 27B no-reasoning | 626.3 | 3,270,754 |
| Method | Mean time / episode (s) | Total output tokens |
|---|---|---|
| RPent / GPT-6 Astra Codex · low · reasoning | 1,168.95 | 5,894,436 |
| RPent / GPT-5.6 Sol Codex · xhigh · reasoning | 1,197.56 | 5,258,391 |
| RPent / GPT-5.6 Sol Codex · no-reasoning | 1,013.21 | 4,792,104 |
| Method | Mean time / episode (s) | Total output tokens |
|---|---|---|
| RPent / GPT-6 Astra Codex · low · reasoning | 1,078.8 | 2,763,641 |
| RPent / GPT-5.6 Sol Codex · xhigh · reasoning | 1,828.7 | 4,763,915 |
| RPent / GPT-5.6 Sol Codex · no-reasoning | 957.4 | 2,456,401 |
RPent 排行榜
- - - - + + + +| 方法 | 成功率 |
|---|---|
| Codex / GPT-6 Astra / low / reasoning[4] | 92.63% |
| Claude Code / Opus-4.7 / max.reasoning | 82.4% |
| Codex / GPT-5.5 / xhigh / reasoning | 75.13% |
| RPent Flash Mode[2] | 72.63% |
| Qwen3.6 27B / no-reasoning | 70.63% |
| ASPIRE[1] | 61.36% |
| π_RLinf | 50.0% |
| π0.5 | 11.0% |
| AtomVLA | 6.3% |
| X-VLA | 3.8% |
| MolmoAct | 1.5% |
| π0 | 0.3% |
* GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800). Qwen3.6 27B / no-reasoning: 70.63% (565/800).
| 方法 | 总体 | Spatial Task | Spatial Swap | Object Task | Object Swap | Goal Task | Goal Swap | Long Task | Long Swap |
|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra[4] | 92.63% | 100% | 98% | 100% | 99% | 88% | 99% | 85% | 72% |
| Opus-4.7 / max.reasoning | 82.4% | — | — | — | — | — | — | — | — |
| GPT-5.5 | 75.13% | 81.0% | 69.0% | 94.0% | 91.0% | 75.0% | 66.0% | 70.00% | 55.00% |
| RPent Flash Mode / Molmo2-8B[2] | 72.63% | 79.00% | 70.00% | 86.00% | 93.00% | 74.00% | 65.00% | 60.00% | 54.00% |
| Qwen3.6 27B / no-reasoning | 70.63% | 82.00% | 78.00% | 83.00% | 84.00% | 68.00% | 68.00% | 61.00% | 41.00% |
| ASPIRE[1] | 61.36% | 60.0% | 51.0% | 95.0% | 98.0% | 45.0% | 81.0% | 38.3% | 22.6% |
| π_RLinf | 50.0% | 42.0% | 59.0% | 71.0% | 78.0% | 45.0% | 42.0% | 49.0% | 14.0% |
| π0.5 | 11.0% | 1.0% | 20.0% | 1.0% | 17.0% | 2.0% | 38.0% | 1.0% | 8.0% |
| AtomVLA | 6.3% | 1.0% | 16.0% | 0.0% | 10.0% | 11.0% | 2.0% | 9.0% | 1.0% |
| X-VLA | 3.8% | 0.0% | 0.0% | 8.0% | 2.0% | 9.0% | 1.0% | 10.0% | 0.0% |
| MolmoAct | 1.5% | 0.0% | 0.0% | 0.0% | 6.0% | 0.0% | 0.0% | 6.0% | 0.0% |
| π0 | 0.3% | 0.0% | 0.0% | 0.0% | 2.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| Cap-X | — | 14.0% | 12.0% | 18.0% | 22.0% | 17.0% | 26.0% | — | — |
| RPent / GPT-6 Motor Only[3] | — | — | — | — | — | — | — | 38.0% | — |
| RATS | — | 31.0% | 29.0% | 63.0% | 61.0% | 36.0% | 43.0% | — | — |
* [1] ASPIRE:Long Task 和 Long Swap 使用 LIBERO-90 技能库进行 zero-shot 迁移。
* [2] RPent Flash Mode 使用直接下载的、官方公开的 GPT-5.5 explore memory;Molmo2-8B 用于视觉定位。该测试结果使用了 s0-s9 中表现最好的 seed。
* [3] RPent / GPT-6 Motor Only 为纯电机控制方法,通过 execute_action 直接输出末端位姿增量与夹爪指令,不调用 VLA / primitive,不加载 memory。
* [4] GPT-6 Astra(使用 memory 的配置):Long Task/Swap 与其余六套件使用各自探索后冻结的 memory 文件快照,评测期间不更新。Overall 合并两个不重叠批次:Long 157/200,加上其余套件 584/600,得到 741/800(92.63%);并非全部回合共享同一份 memory 快照。
| 方法 | 成功率 |
|---|---|
| Codex / GPT-6 Astra / low / reasoning | 59.20% |
| Xiaomi-Robotics-1 | 57.4% |
| Codex / GPT-5.5 / xhigh / reasoning | 57.1% |
| Claude Code / Opus-4.7 / max.reasoning | 48.6% |
| WorldDreamer | 35.3% |
| RLDX-1 | 30.0% |
| π0.5 | 16.9% |
| π0 | 14.8% |
| 方法 | 总体 | Atomic-Seen | Composite-Seen | Composite-Unseen |
|---|---|---|---|---|
| GPT-6 Astra | 59.20% | 87.78% | 43.75% | 42.50% |
| Xiaomi-Robotics-1 | 57.4% | 80.2% | 57.1% | 32.1% |
| GPT-5.5 | 57.1% | 92.0% | 61.0% | 13.8% |
| Opus-4.7 / max.reasoning | 48.6% | 79.4% | 47.5% | 15.0% |
| WorldDreamer | 35.3% | 66.3% | 26.7% | 9.0% |
| RLDX-1 | 30.0% | 60.0% | 21.3% | 5.0% |
| π0.5 | 16.9% | 39.6% | 7.1% | 1.2% |
| π0 | 14.8% | 34.6% | 6.1% | 1.1% |
| 方法 | 成功率 |
|---|---|
| Codex / GPT-5.5 / xhigh / reasoning | 62.4% |
| Claude Code / Opus-4.7 / max.reasoning | 58.4% |
| LingBot-VLA | 50.4% |
| π0.5 | 47.9% |
| GR00T-N1.7 | 20.7% |
| StarVLA | 10.6% |
RPent 排行榜
- - - - + + + +| 方法 | 平均每回合耗时(秒) | 总输出 token |
|---|---|---|
| RPent / GPT-6 Astra Codex · low · reasoning | 412.14 | 3,364,938 |
| RPent / GPT-5.6 Sol Codex · xhigh · reasoning | 529.23 | 3,635,426 |
| RPent / GPT-5.6 Sol Codex · no-reasoning | 344.17 | 2,569,676 |
| RPent Flash Mode Molmo2-8B · 视觉定位 | 60.19 | 0 |
| RPent / Qwen3.6 27B no-reasoning | 626.3 | 3,270,754 |
| 方法 | 平均每回合耗时(秒) | 总输出 token |
|---|---|---|
| RPent / GPT-6 Astra Codex · low · reasoning | 1,168.95 | 5,894,436 |
| RPent / GPT-5.6 Sol Codex · xhigh · reasoning | 1,197.56 | 5,258,391 |
| RPent / GPT-5.6 Sol Codex · no-reasoning | 1,013.21 | 4,792,104 |
| 方法 | 平均每回合耗时(秒) | 总输出 token |
|---|---|---|
| RPent / GPT-6 Astra Codex · low · reasoning | 1,078.8 | 2,763,641 |
| RPent / GPT-5.6 Sol Codex · xhigh · reasoning | 1,828.7 | 4,763,915 |
| RPent / GPT-5.6 Sol Codex · no-reasoning | 957.4 | 2,456,401 |