Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ Compare RPent with reference methods on LIBERO, LIBERO-PRO, RoboCasa365 Target50

Codex / GPT-6 Astra / low / reasoning: **92.63% Overall (741/800)** across all eight LIBERO-PRO suites. See the [suite results and memory-batch explanation](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html#libero-pro-astra-memory), including the separately frozen Long and Spatial/Object/Goal memory batches.

[![RPent Leaderboard](https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-light.png)](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html)
[![RPent Leaderboard](https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-light.png)](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html)


## Who Should Consider Using RPent?
Expand Down
2 changes: 1 addition & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@

Codex / GPT-6 Astra / low / reasoning 已完成全部八套 LIBERO-PRO,**Overall 92.63%(741/800)**。详见 [套件汇总与 memory 批次说明](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html#libero-pro-astra-memory),其中 Long 与 Spatial/Object/Goal 分别使用各自冻结的 memory 批次。

[![RPent 排行榜](https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-zh-light.png)](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html)
[![RPent 排行榜](https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-zh-light.png)](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html)


## 适用用户
Expand Down
4 changes: 2 additions & 2 deletions docs/source-en/rst_source/get_started/overview.rst
Original file line number Diff line number Diff line change
Expand Up @@ -32,13 +32,13 @@ Compare success rates on LIBERO, LIBERO-PRO, RoboCasa365 Target50, and RoboTwin
C2R. Rankings apply to the methods and evaluation coverage shown; see
:doc:`../leaderboard/index` for detailed results, configurations, and sources.

.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-light.png
.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-light.png
:alt: RPent Leaderboard
:class: only-light
:width: 100%
:target: ../leaderboard/index.html

.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-dark.png
.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-dark.png
:alt: RPent Leaderboard
:class: only-dark
:width: 100%
Expand Down
2 changes: 1 addition & 1 deletion docs/source-en/rst_source/leaderboard/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ RPent Leaderboard

.. raw:: html

<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/embed.js"></script>
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/embed.js"></script>
<div id="rpent-interactive-leaderboard" data-section="index"
data-performance-url="performance.html"
data-costs-url="time-token-costs.html"></div>
11 changes: 6 additions & 5 deletions docs/source-en/rst_source/leaderboard/performance.rst
Original file line number Diff line number Diff line change
Expand Up @@ -6,20 +6,21 @@ Performance
.. raw:: html

<p class="leaderboard-page-brand">RPent Leaderboard</p>
<link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/docs.css">
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/table-sort.js"></script>
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard.js"></script>
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/embed.js"></script>
<link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/docs.css">
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/table-sort.js"></script>
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard.js"></script>
<script defer src="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/embed.js"></script>
<div id="rpent-interactive-leaderboard" data-language="en" data-section="performance"
data-performance-url="performance.html" data-costs-url="time-token-costs.html"
data-results-url="https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/results.json">
data-results-url="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/results.json">
<div class="rpent-static-leaderboard">
<section id="native-performance" data-performance="true"><h2 class="page-section-title">Performance</h2><div class="section-layout"><details class="module-directory" open><summary>Environments</summary><nav class="module-nav" aria-label="Performance"><a href="#native-libero-pro">LIBERO-PRO</a><a href="#native-standard-libero">LIBERO</a><a href="#native-robocasa">RoboCasa365</a><a href="#native-robotwin">RoboTwin</a></nav></details><div class="section-results">
<section data-benchmark="libero-pro"><details data-module="libero-pro"><summary class="module-header"><h3>LIBERO-PRO</h3></summary><div id="native-libero-pro" class="module-content"><table><thead><tr><th>Method</th><th>Success rate</th></tr></thead><tbody><tr><td>Codex / GPT-6 Astra / low / reasoning<sup class="note-reference"><a href="#libero-pro-astra-memory" role="doc-noteref" aria-label="Evaluation note 4 for GPT-6 Astra">[4]</a></sup></td><td>92.63%</td></tr><tr><td>Claude Code / Opus-4.7 / max.reasoning</td><td>82.4%</td></tr><tr><td>Codex / GPT-5.5 / xhigh / reasoning</td><td>75.13%</td></tr><tr><td>RPent Flash Mode<sup class="note-reference"><a href="#libero-pro-note-2" role="doc-noteref" aria-label="Evaluation note 2 for RPent Flash Mode">[2]</a></sup></td><td>72.63%</td></tr><tr><td>Qwen3.6 27B / no-reasoning</td><td>70.63%</td></tr><tr><td>ASPIRE<sup class="note-reference"><a href="#libero-pro-note-1" role="doc-noteref" aria-label="Evaluation note 1 for ASPIRE">[1]</a></sup></td><td>61.36%</td></tr><tr><td>π_RLinf</td><td>50.0%</td></tr><tr><td>π0.5</td><td>11.0%</td></tr><tr><td>AtomVLA</td><td>6.3%</td></tr><tr><td>X-VLA</td><td>3.8%</td></tr><tr><td>MolmoAct</td><td>1.5%</td></tr><tr><td>π0</td><td>0.3%</td></tr></tbody></table><p>* GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800). Qwen3.6 27B / no-reasoning: 70.63% (565/800).</p><h4>All methods &amp; reported scores</h4><div class="table-wrap" tabindex="0"><table><thead><tr><th>Method</th><th>Overall</th><th>Spatial Task</th><th>Spatial Swap</th><th>Object Task</th><th>Object Swap</th><th>Goal Task</th><th>Goal Swap</th><th>Long Task</th><th>Long Swap</th></tr></thead><tbody><tr><td>GPT-6 Astra<sup class="note-reference"><a href="#libero-pro-astra-memory" role="doc-noteref" aria-label="Evaluation note 4 for GPT-6 Astra">[4]</a></sup></td><td>92.63%</td><td>100%</td><td>98%</td><td>100%</td><td>99%</td><td>88%</td><td>99%</td><td>85%</td><td>72%</td></tr><tr><td>Opus-4.7 / max.reasoning</td><td>82.4%</td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td></tr><tr><td>GPT-5.5</td><td>75.13%</td><td>81.0%</td><td>69.0%</td><td>94.0%</td><td>91.0%</td><td>75.0%</td><td>66.0%</td><td>70.00%</td><td>55.00%</td></tr><tr><td>RPent Flash Mode / Molmo2-8B<sup class="note-reference"><a href="#libero-pro-note-2" role="doc-noteref" aria-label="Evaluation note 2 for RPent Flash Mode">[2]</a></sup></td><td>72.63%</td><td>79.00%</td><td>70.00%</td><td>86.00%</td><td>93.00%</td><td>74.00%</td><td>65.00%</td><td>60.00%</td><td>54.00%</td></tr><tr><td>Qwen3.6 27B / no-reasoning</td><td>70.63%</td><td>82.00%</td><td>78.00%</td><td>83.00%</td><td>84.00%</td><td>68.00%</td><td>68.00%</td><td>61.00%</td><td>41.00%</td></tr><tr><td>ASPIRE<sup class="note-reference"><a href="#libero-pro-note-1" role="doc-noteref" aria-label="Evaluation note 1 for ASPIRE">[1]</a></sup></td><td>61.36%</td><td>60.0%</td><td>51.0%</td><td>95.0%</td><td>98.0%</td><td>45.0%</td><td>81.0%</td><td>38.3%</td><td>22.6%</td></tr><tr><td>π_RLinf</td><td>50.0%</td><td>42.0%</td><td>59.0%</td><td>71.0%</td><td>78.0%</td><td>45.0%</td><td>42.0%</td><td>49.0%</td><td>14.0%</td></tr><tr><td>π0.5</td><td>11.0%</td><td>1.0%</td><td>20.0%</td><td>1.0%</td><td>17.0%</td><td>2.0%</td><td>38.0%</td><td>1.0%</td><td>8.0%</td></tr><tr><td>AtomVLA</td><td>6.3%</td><td>1.0%</td><td>16.0%</td><td>0.0%</td><td>10.0%</td><td>11.0%</td><td>2.0%</td><td>9.0%</td><td>1.0%</td></tr><tr><td>X-VLA</td><td>3.8%</td><td>0.0%</td><td>0.0%</td><td>8.0%</td><td>2.0%</td><td>9.0%</td><td>1.0%</td><td>10.0%</td><td>0.0%</td></tr><tr><td>MolmoAct</td><td>1.5%</td><td>0.0%</td><td>0.0%</td><td>0.0%</td><td>6.0%</td><td>0.0%</td><td>0.0%</td><td>6.0%</td><td>0.0%</td></tr><tr><td>π0</td><td>0.3%</td><td>0.0%</td><td>0.0%</td><td>0.0%</td><td>2.0%</td><td>0.0%</td><td>0.0%</td><td>0.0%</td><td>0.0%</td></tr><tr><td>Cap-X</td><td>—</td><td>14.0%</td><td>12.0%</td><td>18.0%</td><td>22.0%</td><td>17.0%</td><td>26.0%</td><td>—</td><td>—</td></tr><tr><td>RPent / GPT-6 Motor Only<sup class="note-reference"><a href="#libero-pro-note-3" role="doc-noteref" aria-label="Evaluation note 3 for RPent / GPT-6 Motor Only">[3]</a></sup></td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td><td>—</td><td>38.0%</td><td>—</td></tr><tr><td>RATS</td><td>—</td><td>31.0%</td><td>29.0%</td><td>63.0%</td><td>61.0%</td><td>36.0%</td><td>43.0%</td><td>—</td><td>—</td></tr></tbody></table></div><p id="libero-pro-note-1" class="scope-note" tabindex="-1" role="note">* <span class="note-number">[1]</span> ASPIRE: Long Task and Long Swap use zero-shot transfer from the LIBERO-90 skill library.</p><p id="libero-pro-note-2" class="scope-note" tabindex="-1" role="note">* <span class="note-number">[2]</span> RPent Flash Mode uses directly downloaded, officially released GPT-5.5 exploration memory; Molmo2-8B is used for visual localization. These results use the best-performing seed from s0-s9.</p><p id="libero-pro-note-3" class="scope-note" tabindex="-1" role="note">* <span class="note-number">[3]</span> RPent / GPT-6 Motor Only uses motor-only control: it directly outputs end-effector pose increments and gripper commands through execute_action, without invoking VLA / primitives or loading memory.</p><p id="libero-pro-astra-memory" class="scope-note memory-context" tabindex="-1" role="note">* <span class="note-number">[4]</span> GPT-6 Astra (memory-enabled configuration): Long Task/Swap and the other six suites use separate memory-file snapshots frozen after their respective exploration phases, with no updates during evaluation. Overall combines two non-overlapping batches: Long 157/200 plus the other suites 584/600, giving 741/800 (92.63%); the 800 episodes do not share a single memory snapshot.</p></div></details></section>
<section data-benchmark="standard-libero"><details data-module="standard-libero"><summary class="module-header"><h3>LIBERO</h3></summary><div id="native-standard-libero" class="module-content"><table><thead><tr><th>Method</th><th>Success rate</th></tr></thead><tbody><tr><td>AtomVLA</td><td>97.0%</td></tr><tr><td>Claude Code / Opus-4.7 / max.reasoning</td><td>96.0%</td></tr><tr><td>π_RLinf</td><td>95.3%</td></tr><tr><td>π0</td><td>94.2%</td></tr><tr><td>NORA</td><td>79.5%</td></tr><tr><td>OpenVLA</td><td>76.5%</td></tr></tbody></table></div></details></section>
<section data-benchmark="robocasa"><details data-module="robocasa"><summary class="module-header"><h3>RoboCasa365 · Target50</h3></summary><div id="native-robocasa" class="module-content"><table><thead><tr><th>Method</th><th>Success rate</th></tr></thead><tbody><tr><td>Codex / GPT-6 Astra / low / reasoning</td><td>59.20%</td></tr><tr><td>Xiaomi-Robotics-1</td><td>57.4%</td></tr><tr><td>Codex / GPT-5.5 / xhigh / reasoning</td><td>57.1%</td></tr><tr><td>Claude Code / Opus-4.7 / max.reasoning</td><td>48.6%</td></tr><tr><td>WorldDreamer</td><td>35.3%</td></tr><tr><td>RLDX-1</td><td>30.0%</td></tr><tr><td>π0.5</td><td>16.9%</td></tr><tr><td>π0</td><td>14.8%</td></tr></tbody></table><h4>RoboCasa365 · Target50 · All methods &amp; reported scores</h4><div class="table-wrap" tabindex="0"><table><thead><tr><th>Method</th><th>Overall</th><th>Atomic-Seen</th><th>Composite-Seen</th><th>Composite-Unseen</th></tr></thead><tbody><tr><td>GPT-6 Astra</td><td>59.20%</td><td>87.78%</td><td>43.75%</td><td>42.50%</td></tr><tr><td>Xiaomi-Robotics-1</td><td>57.4%</td><td>80.2%</td><td>57.1%</td><td>32.1%</td></tr><tr><td>GPT-5.5</td><td>57.1%</td><td>92.0%</td><td>61.0%</td><td>13.8%</td></tr><tr><td>Opus-4.7 / max.reasoning</td><td>48.6%</td><td>79.4%</td><td>47.5%</td><td>15.0%</td></tr><tr><td>WorldDreamer</td><td>35.3%</td><td>66.3%</td><td>26.7%</td><td>9.0%</td></tr><tr><td>RLDX-1</td><td>30.0%</td><td>60.0%</td><td>21.3%</td><td>5.0%</td></tr><tr><td>π0.5</td><td>16.9%</td><td>39.6%</td><td>7.1%</td><td>1.2%</td></tr><tr><td>π0</td><td>14.8%</td><td>34.6%</td><td>6.1%</td><td>1.1%</td></tr></tbody></table></div></div></details></section>
<section data-benchmark="robotwin"><details data-module="robotwin"><summary class="module-header"><h3>RoboTwin</h3></summary><div id="native-robotwin" class="module-content"><table><thead><tr><th>Method</th><th>Success rate</th></tr></thead><tbody><tr><td>Codex / GPT-5.5 / xhigh / reasoning</td><td>62.4%</td></tr><tr><td>Claude Code / Opus-4.7 / max.reasoning</td><td>58.4%</td></tr><tr><td>LingBot-VLA</td><td>50.4%</td></tr><tr><td>π0.5</td><td>47.9%</td></tr><tr><td>GR00T-N1.7</td><td>20.7%</td></tr><tr><td>StarVLA</td><td>10.6%</td></tr></tbody></table></div></details></section>
</div></div></section>
<div class="leaderboard-footer"><p class="scope-note evaluation-notice">* All RPent method results shown in this leaderboard are from September 2026 evaluations.</p><span>RPent Leaderboard</span></div>

</div>
</div>
Loading
Loading