-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathindex.html
More file actions
454 lines (448 loc) · 29.8 KB
/
Copy pathindex.html
File metadata and controls
454 lines (448 loc) · 29.8 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<script>
(function () {
var lang = "en";
try { lang = localStorage.getItem("galbot-lang") || "en"; } catch (e) {}
var query = new URLSearchParams(location.search).get("lang");
if (query === "zh" || query === "en") lang = query;
document.documentElement.lang = lang === "zh" ? "zh-CN" : "en";
})();
</script>
<title>Galbot Open Algorithms</title>
<meta name="description" content="Public research code from Galaxy General Robotics for whole-body intelligence, manipulation, and navigation. Highlights: LATENT, Humanoid-GPT, and OpenTrack." />
<meta name="theme-color" content="#0a0a0f" />
<link rel="stylesheet" href="assets/site.css" />
</head>
<body id="top">
<a class="skip" href="#featured"><span class="en">Skip to highlights</span><span class="zh">跳到亮点</span></a>
<header class="top">
<div class="band">
<a class="brand" href="#top" aria-label="Galbot">
<img src="assets/galbot-wordmark.svg" alt="Galbot" />
</a>
<nav class="nav" aria-label="Page">
<a href="#featured"><span class="en">Highlight</span><span class="zh">亮点</span></a>
<a href="#motion"><span class="en">Whole-body Intelligence</span><span class="zh">全身智能</span></a>
<a href="#manipulation"><span class="en">Manipulation</span><span class="zh">操作</span></a>
<a href="#navigation"><span class="en">Navigation</span><span class="zh">导航</span></a>
</nav>
<div class="top-end">
<div class="top-links">
<a href="https://github.com/GalaxyGeneralRobotics">GitHub</a>
<a href="https://www.galbot.com/">galbot.com</a>
</div>
<div class="lang-switch" role="group" aria-label="Language">
<button type="button" data-set="en" aria-pressed="true">EN</button>
<button type="button" data-set="zh" aria-pressed="false">中文</button>
</div>
</div>
</div>
</header>
<main>
<section class="hero">
<div class="band">
<p class="eyebrow"><span class="en">Galaxy General Robotics</span><span class="zh">银河通用机器人</span></p>
<h1><span class="en">Open algorithms</span><span class="zh">开源算法</span></h1>
<p class="lead en">Galaxy General Robotics releases the research behind whole-body intelligence, manipulation, and navigation. Each entry below is a published result, with a paper, a project page, and code that can be reproduced.</p>
<p class="lead zh">银河通用公开全身智能、操作与导航方向的研究成果。以下每一项都包含论文、项目主页,以及可以复现的代码。</p>
<div class="hero-links">
<a class="button" href="https://github.com/GalaxyGeneralRobotics"><span class="en">GitHub</span><span class="zh">GitHub 组织</span></a>
<a class="button quiet" href="https://www.galbot.com/">galbot.com</a>
</div>
</div>
</section>
<section class="section" id="featured">
<div class="band">
<div class="section-head">
<div class="section-title"><span class="num">00</span><h2><span class="en">Highlight</span><span class="zh">亮点</span></h2></div>
<p class="en">Three results anchor whole-body intelligence: a general motion tracker, a tracker that holds under disturbance, and athletic tennis on a real humanoid.</p>
<p class="zh">三项结果撑起全身智能这条线:通用动作跟踪、扰动下仍然稳住的跟踪,以及真机上的网球对打。</p>
</div>
<div class="features">
<article class="feature">
<video controls playsinline muted loop autoplay preload="metadata" poster="assets/posters/latent.jpg" title="LATENT">
<source src="https://zzk273.github.io/LATENT/static/videos/teaser.mp4" type="video/mp4" />
</video>
<div class="body">
<p class="kicker">IROS 2026</p>
<h3>LATENT</h3>
<p class="paper en">Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data</p>
<p class="paper zh">从不完整的人类动作中学习人形机器人的网球技能</p>
<p class="en">Learns tennis from incomplete human swing fragments and keeps rallies going on a real humanoid.</p>
<p class="zh">用不完整的人类击球片段学习网球技能,在真机上完成持续对打。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/LATENT"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2603.12686"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://zzk273.github.io/LATENT/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</div>
</article>
<article class="feature">
<video controls playsinline muted loop autoplay preload="metadata" poster="assets/posters/humanoid-gpt.jpg" title="Humanoid-GPT">
<source src="https://qizekun.github.io/Humanoid-GPT/videos/home.mp4" type="video/mp4" />
</video>
<div class="body">
<p class="kicker">CVPR 2026</p>
<h3>Humanoid-GPT</h3>
<p class="paper en">Scaling Data and Structure for Zero-Shot Motion Tracking</p>
<p class="paper zh">扩展数据与模型结构,实现零样本全身动作跟踪</p>
<p class="en">A causal Transformer pretrained on about two billion retargeted frames tracks unseen whole-body motion without fine-tuning.</p>
<p class="zh">在约 20 亿帧重定向动作上预训练因果 Transformer,不经微调即可跟踪未见过的全身动作。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/Humanoid-GPT"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2606.03985"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://qizekun.github.io/Humanoid-GPT/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</div>
</article>
<article class="feature">
<video controls playsinline muted loop autoplay preload="metadata" poster="assets/posters/opentrack.jpg" title="Any2Track">
<source src="https://zzk273.github.io/Any2Track/static/videos/any2track_480p.mp4" type="video/mp4" />
</video>
<div class="body">
<p class="kicker">ICRA 2026</p>
<h3>Any2Track</h3>
<p class="paper en">Track Any Motions under Any Disturbances</p>
<p class="paper zh">在任意扰动下跟踪任意动作</p>
<p class="en">One policy tracks diverse, highly dynamic motion and stays stable across terrain, pushes, and changes in physical parameters.</p>
<p class="zh">一个策略跟踪多样、高动态的动作,并在地形、外力和物理参数变化下保持稳定。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/OpenTrack"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2509.13833"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://zzk273.github.io/Any2Track/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</div>
</article>
</div>
</div>
</section>
<section class="section" id="motion">
<div class="band">
<div class="section-head">
<div class="section-title"><span class="num">01</span><h2><span class="en">Whole-body Intelligence</span><span class="zh">全身智能</span></h2></div>
<p class="en">The three highlights sit here with teleoperation, traversal through clutter, a human-aligned tracking benchmark, and real-time co-speech gesture.</p>
<p class="zh">上述三项与遥操作、杂乱场景穿行、对齐人类判断的跟踪评测,以及实时共语手势列在一起。</p>
</div>
<div class="grid">
<article class="card">
<div class="card-top">
<h3>LATENT</h3>
<span class="badge">IROS 2026</span>
</div>
<p class="paper en">Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data</p>
<p class="paper zh">从不完整的人类动作中学习人形机器人的网球技能</p>
<p class="en">Learns tennis from incomplete human swing fragments and keeps rallies going on a real humanoid.</p>
<p class="zh">用不完整的人类击球片段学习网球技能,在真机上完成持续对打。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/LATENT"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2603.12686"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://zzk273.github.io/LATENT/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>Humanoid-GPT</h3>
<span class="badge">CVPR 2026</span>
</div>
<p class="paper en">Scaling Data and Structure for Zero-Shot Motion Tracking</p>
<p class="paper zh">扩展数据与模型结构,实现零样本全身动作跟踪</p>
<p class="en">A causal Transformer pretrained on about two billion retargeted frames tracks unseen whole-body motion without fine-tuning.</p>
<p class="zh">在约 20 亿帧重定向动作上预训练因果 Transformer,不经微调即可跟踪未见过的全身动作。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/Humanoid-GPT"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2606.03985"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://qizekun.github.io/Humanoid-GPT/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>Any2Track</h3>
<span class="badge">ICRA 2026</span>
</div>
<p class="paper en">Track Any Motions under Any Disturbances</p>
<p class="paper zh">在任意扰动下跟踪任意动作</p>
<p class="en">One policy tracks diverse, highly dynamic motion and stays stable across terrain, pushes, and changes in physical parameters.</p>
<p class="zh">一个策略跟踪多样、高动态的动作,并在地形、外力和物理参数变化下保持稳定。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/OpenTrack"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2509.13833"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://zzk273.github.io/Any2Track/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>OpenWBT</h3>
<span class="badge">RA-L 2025</span>
</div>
<p class="paper en">Whole-body teleoperation of Unitree G1 with Apple Vision Pro</p>
<p class="paper zh">用 Apple Vision Pro 遥操作 Unitree G1 的全身</p>
<p class="en">One person teleoperates the whole body with a Vision Pro, walking, squatting, bending, and grasping on the real robot and in simulation. The system builds on R2S2.</p>
<p class="zh">一个人用 Vision Pro 遥操作全身,在真机和仿真里完成行走、下蹲、弯腰和抓取。技术基础来自 R2S2。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/OpenWBT"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2505.10918"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://zzk273.github.io/R2S2/"><span class="en">Page</span><span class="zh">主页</span></a>
<a href="https://www.youtube.com/watch?v=EmWLJROMeB0"><span class="en">Video</span><span class="zh">视频</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>Click-and-Traverse</h3>
<span class="badge">RA-L 2026</span>
</div>
<p class="paper en">Collision-Free Humanoid Traversal in Cluttered Indoor Scenes</p>
<p class="paper zh">杂乱室内场景中的无碰撞人形穿行</p>
<p class="en">Traverses indoor scenes with obstacles on the floor, at the sides, and overhead, distilling several specialist policies into one generalist.</p>
<p class="zh">在地面、侧面和头顶同时有障碍的室内场景中穿行,并从多个专家策略蒸馏出通用策略。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/Click-and-Traverse"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2601.16035"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://axian12138.github.io/CAT/"><span class="en">Page</span><span class="zh">主页</span></a>
<a href="https://www.youtube.com/watch?v=blek__Qf0Vc"><span class="en">Video</span><span class="zh">视频</span></a>
<a href="https://www.bilibili.com/video/BV1aNr6BiEnL"><span class="en">Bilibili</span><span class="zh">哔哩哔哩</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>HumanTracker</h3>
<span class="badge">ECCV 2026</span>
</div>
<p class="paper en">Towards Comprehensive and Human-Aligned Motion Tracking Benchmark</p>
<p class="paper zh">面向全面、与人类判断对齐的动作跟踪评测</p>
<p class="en">About 153 hours of optical motion capture, plus a score aligned with human preference, for judging whether tracking looks natural.</p>
<p class="zh">约 153 小时光学动捕,加上和人类偏好对齐的轨迹评分,用来判断跟踪结果看起来是否自然。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/HumanTracker"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2608.13555"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://dairuliu.github.io/humantracker/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>RoboGesture</h3>
<span class="badge">ECCV 2026</span>
</div>
<p class="paper en">Real-Time Semantic-aligned Co-Speech Gestures for Humanoid Interaction</p>
<p class="paper zh">面向人形交互的实时语义对齐共语手势</p>
<p class="en">Generates semantically aligned co-speech gestures from audio in real time, and executes them on a humanoid without collisions.</p>
<p class="zh">从语音实时生成语义对齐的共语手势,并在人形机器人上做无碰撞执行。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/RoboGesture"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2608.28693"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://robogesture.github.io/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
</div>
</div>
</section>
<section class="section" id="manipulation">
<div class="band">
<div class="section-head">
<div class="section-title"><span class="num">02</span><h2><span class="en">Manipulation</span><span class="zh">操作</span></h2></div>
<p class="en">Grasping, latent action, spatial reasoning, and dexterous hands. The report on Astra as a policy belongs in this section.</p>
<p class="zh">抓取、潜动作、空间推理与灵巧手。把 Astra 当作策略的评测报告也归在这里。</p>
</div>
<div class="grid">
<article class="card">
<div class="card-top">
<h3>Astra Policy</h3>
<span class="badge"><span class="en">On this site</span><span class="zh">本站报告</span></span>
</div>
<p class="paper en">Exploring the Comprehensive Capabilities of GPT-6 Astra as Policies</p>
<p class="paper zh">探索 GPT-6 Astra 作为机器人策略的综合能力</p>
<p class="en">Evaluates Astra as an executable policy across manipulation, dexterous hands, visual navigation, humanoid control, and locomotion.</p>
<p class="zh">把 Astra 当作可执行策略,在操作、灵巧手、视觉导航、人形控制和运动上对照评测。</p>
<p class="links">
<a href="astra-policy/"><span class="en">Read the report</span><span class="zh">阅读报告</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>LDA-1B</h3>
<span class="badge">RSS 2026</span>
</div>
<p class="paper en">Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion</p>
<p class="paper zh">以通用具身数据扩展潜动力学动作模型</p>
<p class="en">Learns dynamics, future prediction, and action together on more than thirty thousand hours of mixed human and robot data.</p>
<p class="zh">在三万小时以上的人机异构数据上,同时学习动力学、未来预测和动作。</p>
<p class="links">
<a href="https://github.com/jiangranlv/LDA-1B"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2602.12215"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/LDA/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>GraspVLA</h3>
<span class="badge">CoRL 2025</span>
</div>
<p class="paper en">A Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data</p>
<p class="paper zh">在十亿帧合成动作上预训练的抓取基础模型</p>
<p class="en">Pretrains on a billion frames of synthetic grasps and transfers to open-vocabulary grasping without real-robot fine-tuning.</p>
<p class="zh">用十亿帧合成抓取轨迹预训练,不靠真实机器人数据微调,直接做开放词汇抓取。</p>
<p class="links">
<a href="https://github.com/PKU-EPIC/GraspVLA"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2505.03233"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/GraspVLA-web/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>SoFar</h3>
<span class="badge">NeurIPS 2025 Spotlight</span>
</div>
<p class="paper en">Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation</p>
<p class="paper zh">用语言描述的朝向连接空间推理与物体操作</p>
<p class="en">Connects spatial understanding to 6-DoF manipulation through language-described orientation, and includes the Open6DOR V2 evaluation.</p>
<p class="zh">用语言描述的朝向把空间理解和六自由度操作接起来,并包含 Open6DOR V2 评测。</p>
<p class="links">
<a href="https://github.com/qizekun/SoFar"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2502.13143"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://qizekun.github.io/sofar/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>UniDexGrasp++</h3>
<span class="badge">ICCV 2023</span>
</div>
<p class="paper en">Improving Dexterous Grasping Policy Learning via Geometry-aware Curriculum</p>
<p class="paper zh">以几何课程改进灵巧抓取策略的学习</p>
<p class="en">Learns dexterous grasping on a large object set with a geometry curriculum and generalist–specialist iteration. Oral and best-paper finalist at ICCV.</p>
<p class="zh">用几何课程和通用—专家迭代,把灵巧抓取策略学到大规模物体上。ICCV 口头报告,最佳论文提名。</p>
<p class="links">
<a href="https://github.com/PKU-EPIC/UniDexGrasp2"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2304.00464"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/UniDexGrasp++/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>DexGraspNet</h3>
<span class="badge">ICRA 2023</span>
</div>
<p class="paper en">A Large-Scale Robotic Dexterous Grasp Dataset for General Objects</p>
<p class="paper zh">面向通用物体的大规模机器人灵巧抓取数据集</p>
<p class="en">A large-scale dexterous grasp set and synthesis method for the Shadow Hand. Finalist for the ICRA outstanding manipulation paper.</p>
<p class="zh">面向 ShadowHand 的大规模灵巧抓取数据与合成方法。ICRA 操作方向杰出论文提名。</p>
<p class="links">
<a href="https://github.com/PKU-EPIC/DexGraspNet"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2210.02697"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/DexGraspNet/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
</div>
</div>
</section>
<section class="section" id="navigation">
<div class="band">
<div class="section-head">
<div class="section-title"><span class="num">03</span><h2><span class="en">Navigation</span><span class="zh">导航</span></h2></div>
<p class="en">Urban micromobility, video navigation across tasks, and visual tracking in the wild. Only work with released code or an open benchmark is included.</p>
<p class="zh">城市微出行、跨任务的视频导航,以及野外视觉跟踪。这里只收录已经放出代码或公开评测基准的工作。</p>
</div>
<div class="grid">
<article class="card">
<div class="card-top">
<h3>UrbanVLA</h3>
<span class="badge">ICRA 2026</span>
</div>
<p class="paper en">A Vision-Language-Action Model for Urban Micromobility</p>
<p class="paper zh">面向城市微出行的视觉-语言-动作模型</p>
<p class="en">Aligns a navigation route with onboard vision for long routes through city blocks.</p>
<p class="zh">把导航路线和车载视觉对齐,在城市街区里做长距离微出行。</p>
<p class="links">
<a href="https://github.com/GalaxyGeneralRobotics/UrbanVLA"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2510.23576"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/UrbanVLA-Web/"><span class="en">Page</span><span class="zh">主页</span></a>
<a href="https://www.youtube.com/watch?v=k98F77ugHQA"><span class="en">Video</span><span class="zh">视频</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>Uni-NaVid</h3>
<span class="badge">RSS 2025</span>
</div>
<p class="paper en">A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks</p>
<p class="paper zh">用视频视觉-语言-动作模型统一具身导航任务</p>
<p class="en">One video-action model covers language navigation, object search, and visual tracking.</p>
<p class="zh">一个视频动作模型覆盖语言导航、寻找物体和视觉跟踪。</p>
<p class="links">
<a href="https://github.com/jzhzhang/Uni-NaVid"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2412.06224"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/Uni-NaVid/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>TrackVLA</h3>
<span class="badge">CoRL 2025</span>
</div>
<p class="paper en">Embodied Visual Tracking in the Wild</p>
<p class="paper zh">野外具身视觉跟踪</p>
<p class="en">Recognizes and tracks a target in the wild. The release includes EVT-Bench and evaluation code.</p>
<p class="zh">在野外同时识别并跟踪目标。开放仓库包含 EVT-Bench 和评测代码。</p>
<p class="links">
<a href="https://github.com/ShaoanWang/TrackVLA"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2505.23189"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/TrackVLA-web/"><span class="en">Page</span><span class="zh">主页</span></a>
<a href="https://youtu.be/v51U3Nk-SK4"><span class="en">Video</span><span class="zh">视频</span></a>
</p>
</article>
<article class="card">
<div class="card-top">
<h3>NaVid</h3>
<span class="badge">RSS 2024</span>
</div>
<p class="paper en">Video-based VLM Plans the Next Step for Vision-and-Language Navigation</p>
<p class="paper zh">视频视觉语言模型规划语言导航的下一步</p>
<p class="en">Plans the next step of language navigation from a first-person video, without odometry or a map. The release is the VLN-CE evaluation code.</p>
<p class="zh">只看第一人称视频来规划语言导航的下一步,不依赖里程计和地图。开放的是 VLN-CE 评测代码。</p>
<p class="links">
<a href="https://github.com/jzhzhang/NaVid-VLN-CE"><span class="en">Code</span><span class="zh">代码</span></a>
<a href="https://arxiv.org/abs/2402.15852"><span class="en">Paper</span><span class="zh">论文</span></a>
<a href="https://pku-epic.github.io/NaVid/"><span class="en">Page</span><span class="zh">主页</span></a>
</p>
</article>
</div>
</div>
</section>
</main>
<footer class="foot">
<div class="band">
<p><span class="en">Galaxy General Robotics</span><span class="zh">北京银河通用机器人有限公司</span></p>
<p class="foot-links">
<a href="https://www.galbot.com/"><span class="en">Website</span><span class="zh">官网</span></a>
<a href="https://github.com/GalaxyGeneralRobotics">GitHub</a>
<a href="https://hughw19.github.io/"><span class="en">Related research</span><span class="zh">相关研究</span></a>
</p>
</div>
</footer>
<script>
(function () {
var titles = { en: "Galbot Open Algorithms", zh: "银河通用开源算法 · Galbot" };
function sync() {
var zh = document.documentElement.lang === "zh-CN";
document.title = zh ? titles.zh : titles.en;
document.querySelectorAll(".lang-switch button").forEach(function (button) {
var on = button.dataset.set === (zh ? "zh" : "en");
button.setAttribute("aria-pressed", on ? "true" : "false");
});
}
document.querySelectorAll(".lang-switch button").forEach(function (button) {
button.addEventListener("click", function () {
var lang = button.dataset.set;
document.documentElement.lang = lang === "zh" ? "zh-CN" : "en";
try { localStorage.setItem("galbot-lang", lang); } catch (e) {}
sync();
});
});
sync();
})();
</script>
</body>
</html>