diff --git a/README.md b/README.md index 1cde9f794..f0e221b31 100644 --- a/README.md +++ b/README.md @@ -33,7 +33,7 @@ Compare RPent with reference methods on LIBERO, LIBERO-PRO, RoboCasa365 Target50 Codex / GPT-6 Astra / low / reasoning: **92.63% Overall (741/800)** across all eight LIBERO-PRO suites. See the [suite results and memory-batch explanation](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html#libero-pro-astra-memory), including the separately frozen Long and Spatial/Object/Goal memory batches. -[![RPent Leaderboard](https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-light.png)](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html) +[![RPent Leaderboard](https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-light.png)](https://rpent.readthedocs.io/en/latest/rst_source/leaderboard/index.html) ## Who Should Consider Using RPent? diff --git a/README.zh-CN.md b/README.zh-CN.md index 6441d0d36..52061bd2e 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -33,7 +33,7 @@ Codex / GPT-6 Astra / low / reasoning 已完成全部八套 LIBERO-PRO,**Overall 92.63%(741/800)**。详见 [套件汇总与 memory 批次说明](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html#libero-pro-astra-memory),其中 Long 与 Spatial/Object/Goal 分别使用各自冻结的 memory 批次。 -[![RPent 排行榜](https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-zh-light.png)](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html) +[![RPent 排行榜](https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-zh-light.png)](https://rpent.readthedocs.io/zh-cn/latest/rst_source/leaderboard/index.html) ## 适用用户 diff --git a/docs/source-en/rst_source/get_started/overview.rst b/docs/source-en/rst_source/get_started/overview.rst index b89c496f0..aa4dd34d6 100644 --- a/docs/source-en/rst_source/get_started/overview.rst +++ b/docs/source-en/rst_source/get_started/overview.rst @@ -32,13 +32,13 @@ Compare success rates on LIBERO, LIBERO-PRO, RoboCasa365 Target50, and RoboTwin C2R. Rankings apply to the methods and evaluation coverage shown; see :doc:`../leaderboard/index` for detailed results, configurations, and sources. -.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-light.png +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-light.png :alt: RPent Leaderboard :class: only-light :width: 100% :target: ../leaderboard/index.html -.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-en-dark.png +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-en-dark.png :alt: RPent Leaderboard :class: only-dark :width: 100% diff --git a/docs/source-en/rst_source/leaderboard/index.rst b/docs/source-en/rst_source/leaderboard/index.rst index a14e518d7..e4b48b3eb 100644 --- a/docs/source-en/rst_source/leaderboard/index.rst +++ b/docs/source-en/rst_source/leaderboard/index.rst @@ -11,7 +11,7 @@ RPent Leaderboard .. raw:: html - +
diff --git a/docs/source-en/rst_source/leaderboard/performance.rst b/docs/source-en/rst_source/leaderboard/performance.rst index b7089934b..a6d4b8baa 100644 --- a/docs/source-en/rst_source/leaderboard/performance.rst +++ b/docs/source-en/rst_source/leaderboard/performance.rst @@ -6,13 +6,13 @@ Performance .. raw:: html

RPent Leaderboard

- - - - + + + +
+ data-results-url="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/results.json">

Performance

Environments

LIBERO-PRO

MethodSuccess rate
Codex / GPT-6 Astra / low / reasoning[4]92.63%
Claude Code / Opus-4.7 / max.reasoning82.4%
Codex / GPT-5.5 / xhigh / reasoning75.13%
RPent Flash Mode[2]72.63%
Qwen3.6 27B / no-reasoning70.63%
ASPIRE[1]61.36%
π_RLinf50.0%
π0.511.0%
AtomVLA6.3%
X-VLA3.8%
MolmoAct1.5%
π00.3%

* GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800). Qwen3.6 27B / no-reasoning: 70.63% (565/800).

All methods & reported scores

MethodOverallSpatial TaskSpatial SwapObject TaskObject SwapGoal TaskGoal SwapLong TaskLong Swap
GPT-6 Astra[4]92.63%100%98%100%99%88%99%85%72%
Opus-4.7 / max.reasoning82.4%————————
GPT-5.575.13%81.0%69.0%94.0%91.0%75.0%66.0%70.00%55.00%
RPent Flash Mode / Molmo2-8B[2]72.63%79.00%70.00%86.00%93.00%74.00%65.00%60.00%54.00%
Qwen3.6 27B / no-reasoning70.63%82.00%78.00%83.00%84.00%68.00%68.00%61.00%41.00%
ASPIRE[1]61.36%60.0%51.0%95.0%98.0%45.0%81.0%38.3%22.6%
π_RLinf50.0%42.0%59.0%71.0%78.0%45.0%42.0%49.0%14.0%
π0.511.0%1.0%20.0%1.0%17.0%2.0%38.0%1.0%8.0%
AtomVLA6.3%1.0%16.0%0.0%10.0%11.0%2.0%9.0%1.0%
X-VLA3.8%0.0%0.0%8.0%2.0%9.0%1.0%10.0%0.0%
MolmoAct1.5%0.0%0.0%0.0%6.0%0.0%0.0%6.0%0.0%
π00.3%0.0%0.0%0.0%2.0%0.0%0.0%0.0%0.0%
Cap-X—14.0%12.0%18.0%22.0%17.0%26.0%——
RPent / GPT-6 Motor Only[3]———————38.0%—
RATS—31.0%29.0%63.0%61.0%36.0%43.0%——

* [1] ASPIRE: Long Task and Long Swap use zero-shot transfer from the LIBERO-90 skill library.

* [2] RPent Flash Mode uses directly downloaded, officially released GPT-5.5 exploration memory; Molmo2-8B is used for visual localization. These results use the best-performing seed from s0-s9.

* [3] RPent / GPT-6 Motor Only uses motor-only control: it directly outputs end-effector pose increments and gripper commands through execute_action, without invoking VLA / primitives or loading memory.

* [4] GPT-6 Astra (memory-enabled configuration): Long Task/Swap and the other six suites use separate memory-file snapshots frozen after their respective exploration phases, with no updates during evaluation. Overall combines two non-overlapping batches: Long 157/200 plus the other suites 584/600, giving 741/800 (92.63%); the 800 episodes do not share a single memory snapshot.

@@ -20,6 +20,7 @@ Performance

RoboCasa365 · Target50

MethodSuccess rate
Codex / GPT-6 Astra / low / reasoning59.20%
Xiaomi-Robotics-157.4%
Codex / GPT-5.5 / xhigh / reasoning57.1%
Claude Code / Opus-4.7 / max.reasoning48.6%
WorldDreamer35.3%
RLDX-130.0%
π0.516.9%
π014.8%

RoboCasa365 · Target50 · All methods & reported scores

MethodOverallAtomic-SeenComposite-SeenComposite-Unseen
GPT-6 Astra59.20%87.78%43.75%42.50%
Xiaomi-Robotics-157.4%80.2%57.1%32.1%
GPT-5.557.1%92.0%61.0%13.8%
Opus-4.7 / max.reasoning48.6%79.4%47.5%15.0%
WorldDreamer35.3%66.3%26.7%9.0%
RLDX-130.0%60.0%21.3%5.0%
π0.516.9%39.6%7.1%1.2%
π014.8%34.6%6.1%1.1%

RoboTwin

MethodSuccess rate
Codex / GPT-5.5 / xhigh / reasoning62.4%
Claude Code / Opus-4.7 / max.reasoning58.4%
LingBot-VLA50.4%
π0.547.9%
GR00T-N1.720.7%
StarVLA10.6%
+
diff --git a/docs/source-en/rst_source/leaderboard/time-token-costs.rst b/docs/source-en/rst_source/leaderboard/time-token-costs.rst index 19b30ca5b..9154c5660 100644 --- a/docs/source-en/rst_source/leaderboard/time-token-costs.rst +++ b/docs/source-en/rst_source/leaderboard/time-token-costs.rst @@ -6,14 +6,15 @@ Time & Token Costs .. raw:: html

RPent Leaderboard

- - - - + + + +
+ data-results-url="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/results.json">

Time & Token Costs

Environments

LIBERO-PRO

MethodMean time / episode (s)Total output tokens
RPent / GPT-6 Astra Codex · low · reasoning412.143,364,938
RPent / GPT-5.6 Sol Codex · xhigh · reasoning529.233,635,426
RPent / GPT-5.6 Sol Codex · no-reasoning344.172,569,676
RPent Flash Mode Molmo2-8B · visual localization60.190
RPent / Qwen3.6 27B no-reasoning626.33,270,754

RoboCasa365

MethodMean time / episode (s)Total output tokens
RPent / GPT-6 Astra Codex · low · reasoning1,168.955,894,436
RPent / GPT-5.6 Sol Codex · xhigh · reasoning1,197.565,258,391
RPent / GPT-5.6 Sol Codex · no-reasoning1,013.214,792,104

RoboTwin

MethodMean time / episode (s)Total output tokens
RPent / GPT-6 Astra Codex · low · reasoning1,078.82,763,641
RPent / GPT-5.6 Sol Codex · xhigh · reasoning1,828.74,763,915
RPent / GPT-5.6 Sol Codex · no-reasoning957.42,456,401
+
diff --git a/docs/source-zh/rst_source/get_started/overview.rst b/docs/source-zh/rst_source/get_started/overview.rst index 05596c5af..90e5c9d72 100644 --- a/docs/source-zh/rst_source/get_started/overview.rst +++ b/docs/source-zh/rst_source/get_started/overview.rst @@ -17,13 +17,13 @@ RPent 的三条核心设计原则是服务化、标准化和可组合(service- 对比 LIBERO、LIBERO-PRO、RoboCasa365 Target50 和 RoboTwin C2R 上的成功率。 排名仅限图中方法及评测范围;完整结果、模型配置和来源见 :doc:`../leaderboard/index`。 -.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-zh-light.png +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-zh-light.png :alt: RPent 排行榜 :class: only-light :width: 100% :target: ../leaderboard/index.html -.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@705bd44bfc8ad7586b76239de13167db35abcce7/rpent/benchmarks/leaderboard-zh-dark.png +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/leaderboard-zh-dark.png :alt: RPent 排行榜 :class: only-dark :width: 100% diff --git a/docs/source-zh/rst_source/leaderboard/index.rst b/docs/source-zh/rst_source/leaderboard/index.rst index 0639b83cc..638c1069b 100644 --- a/docs/source-zh/rst_source/leaderboard/index.rst +++ b/docs/source-zh/rst_source/leaderboard/index.rst @@ -11,7 +11,7 @@ RPent 排行榜 .. raw:: html - +
diff --git a/docs/source-zh/rst_source/leaderboard/performance.rst b/docs/source-zh/rst_source/leaderboard/performance.rst index 146d31411..20b6811e3 100644 --- a/docs/source-zh/rst_source/leaderboard/performance.rst +++ b/docs/source-zh/rst_source/leaderboard/performance.rst @@ -6,13 +6,13 @@ .. raw:: html

RPent 排行榜

- - - - + + + +
+ data-results-url="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/results.json">

评测成绩

环境

LIBERO-PRO

方法成功率
Codex / GPT-6 Astra / low / reasoning[4]92.63%
Claude Code / Opus-4.7 / max.reasoning82.4%
Codex / GPT-5.5 / xhigh / reasoning75.13%
RPent Flash Mode[2]72.63%
Qwen3.6 27B / no-reasoning70.63%
ASPIRE[1]61.36%
π_RLinf50.0%
π0.511.0%
AtomVLA6.3%
X-VLA3.8%
MolmoAct1.5%
π00.3%

* GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800). Qwen3.6 27B / no-reasoning: 70.63% (565/800).

完整方法与分项成绩

方法总体Spatial TaskSpatial SwapObject TaskObject SwapGoal TaskGoal SwapLong TaskLong Swap
GPT-6 Astra[4]92.63%100%98%100%99%88%99%85%72%
Opus-4.7 / max.reasoning82.4%————————
GPT-5.575.13%81.0%69.0%94.0%91.0%75.0%66.0%70.00%55.00%
RPent Flash Mode / Molmo2-8B[2]72.63%79.00%70.00%86.00%93.00%74.00%65.00%60.00%54.00%
Qwen3.6 27B / no-reasoning70.63%82.00%78.00%83.00%84.00%68.00%68.00%61.00%41.00%
ASPIRE[1]61.36%60.0%51.0%95.0%98.0%45.0%81.0%38.3%22.6%
π_RLinf50.0%42.0%59.0%71.0%78.0%45.0%42.0%49.0%14.0%
π0.511.0%1.0%20.0%1.0%17.0%2.0%38.0%1.0%8.0%
AtomVLA6.3%1.0%16.0%0.0%10.0%11.0%2.0%9.0%1.0%
X-VLA3.8%0.0%0.0%8.0%2.0%9.0%1.0%10.0%0.0%
MolmoAct1.5%0.0%0.0%0.0%6.0%0.0%0.0%6.0%0.0%
π00.3%0.0%0.0%0.0%2.0%0.0%0.0%0.0%0.0%
Cap-X—14.0%12.0%18.0%22.0%17.0%26.0%——
RPent / GPT-6 Motor Only[3]———————38.0%—
RATS—31.0%29.0%63.0%61.0%36.0%43.0%——

* [1] ASPIRE:Long Task 和 Long Swap 使用 LIBERO-90 技能库进行 zero-shot 迁移。

* [2] RPent Flash Mode 使用直接下载的、官方公开的 GPT-5.5 explore memory;Molmo2-8B 用于视觉定位。该测试结果使用了 s0-s9 中表现最好的 seed。

* [3] RPent / GPT-6 Motor Only 为纯电机控制方法,通过 execute_action 直接输出末端位姿增量与夹爪指令,不调用 VLA / primitive,不加载 memory。

* [4] GPT-6 Astra(使用 memory 的配置):Long Task/Swap 与其余六套件使用各自探索后冻结的 memory 文件快照,评测期间不更新。Overall 合并两个不重叠批次:Long 157/200,加上其余套件 584/600,得到 741/800(92.63%);并非全部回合共享同一份 memory 快照。

@@ -20,6 +20,7 @@

RoboCasa365 · Target50

方法成功率
Codex / GPT-6 Astra / low / reasoning59.20%
Xiaomi-Robotics-157.4%
Codex / GPT-5.5 / xhigh / reasoning57.1%
Claude Code / Opus-4.7 / max.reasoning48.6%
WorldDreamer35.3%
RLDX-130.0%
π0.516.9%
π014.8%

RoboCasa365 · Target50 · 完整方法与分项成绩

方法总体Atomic-SeenComposite-SeenComposite-Unseen
GPT-6 Astra59.20%87.78%43.75%42.50%
Xiaomi-Robotics-157.4%80.2%57.1%32.1%
GPT-5.557.1%92.0%61.0%13.8%
Opus-4.7 / max.reasoning48.6%79.4%47.5%15.0%
WorldDreamer35.3%66.3%26.7%9.0%
RLDX-130.0%60.0%21.3%5.0%
π0.516.9%39.6%7.1%1.2%
π014.8%34.6%6.1%1.1%

RoboTwin

方法成功率
Codex / GPT-5.5 / xhigh / reasoning62.4%
Claude Code / Opus-4.7 / max.reasoning58.4%
LingBot-VLA50.4%
π0.547.9%
GR00T-N1.720.7%
StarVLA10.6%
+
diff --git a/docs/source-zh/rst_source/leaderboard/time-token-costs.rst b/docs/source-zh/rst_source/leaderboard/time-token-costs.rst index b9111ca4e..862508ebc 100644 --- a/docs/source-zh/rst_source/leaderboard/time-token-costs.rst +++ b/docs/source-zh/rst_source/leaderboard/time-token-costs.rst @@ -6,14 +6,15 @@ .. raw:: html

RPent 排行榜

- - - - + + + +
+ data-results-url="https://cdn.jsdelivr.net/gh/RLinf/misc@756ab5ee0667f658332a34a72910df5e6448dfea/rpent/benchmarks/results.json">

耗时与 Token 开销

环境

LIBERO-PRO

方法平均每回合耗时(秒)总输出 token
RPent / GPT-6 Astra Codex · low · reasoning412.143,364,938
RPent / GPT-5.6 Sol Codex · xhigh · reasoning529.233,635,426
RPent / GPT-5.6 Sol Codex · no-reasoning344.172,569,676
RPent Flash Mode Molmo2-8B · 视觉定位60.190
RPent / Qwen3.6 27B no-reasoning626.33,270,754

RoboCasa365

方法平均每回合耗时(秒)总输出 token
RPent / GPT-6 Astra Codex · low · reasoning1,168.955,894,436
RPent / GPT-5.6 Sol Codex · xhigh · reasoning1,197.565,258,391
RPent / GPT-5.6 Sol Codex · no-reasoning1,013.214,792,104

RoboTwin

方法平均每回合耗时(秒)总输出 token
RPent / GPT-6 Astra Codex · low · reasoning1,078.82,763,641
RPent / GPT-5.6 Sol Codex · xhigh · reasoning1,828.74,763,915
RPent / GPT-5.6 Sol Codex · no-reasoning957.42,456,401
+