Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
81 changes: 81 additions & 0 deletions docs/examples/inverted-pendulum.en.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# Inverted Pendulum

`examples/inverted_pendulum/` ships two PPO tasks for the classical cart-pole problem on flat
ground. The extension mirrors `examples/unitree/`: `ManagerBasedRlEnv` over Genesis, rsl_rl PPO,
a `BodyVelocitySensor` on the pole link, and the unified `genelab train` / `genelab play` CLI.

## Tasks

| Task id | Problem |
|---------|---------|
| `GeneLab-Inverted-Pendulum-v0` | Single inverted pole on a cart. |
| `GeneLab-Double-Inverted-Pendulum-v0` | Two stacked inverted poles on a cart. |

## Installation

The extension depends on the `rl` extra (rsl_rl). Pick the `torch-*` extra that matches the
hardware.

```bash
uv sync --extra rl --extra torch-cu128
uv pip install -e examples/inverted_pendulum

uv run genelab list tasks
# -> GeneLab-Inverted-Pendulum-v0
# -> GeneLab-Double-Inverted-Pendulum-v0
```

## Single inverted pendulum

```bash
uv run genelab train GeneLab-Inverted-Pendulum-v0 \
--num-envs 4096 --max-iterations 150

uv run genelab play GeneLab-Inverted-Pendulum-v0 \
--checkpoint logs/rsl_rl/inverted_pendulum_flat/<run>/model_150.pt --vis
```

`--checkpoint` makes `play` route through the RL runner with `--agent trained` by default.

## Double inverted pendulum

```bash
uv run genelab train GeneLab-Double-Inverted-Pendulum-v0 \
--num-envs 4096 --max-iterations 300

uv run genelab play GeneLab-Double-Inverted-Pendulum-v0 \
--checkpoint logs/rsl_rl/double_inverted_pendulum_flat/<run>/model_300.pt --vis
```

## Sensor and underactuation

Only the cart slide joint is PD-controlled. The pole hinges default to `kp=0, kv=0` so the
pendulum stays underactuated and the policy must learn balance through cart motion alone. A
`BodyVelocitySensor` attached to the top pole supplies a noisy angular-velocity observation
(corrupted with `Unoise` in the policy group, clean in the critic group).

## Interactive disturbance

Play mode launches a single environment (`num_envs=1`) and enables Genesis'
`MouseInteractionPlugin`. Left-click on the cart or pole and drag — a spring force pulls the
clicked link toward the cursor while the policy keeps balancing. Scroll wheel rotates the drag
plane around the surface normal. Release the button to remove the force.

!!! tip "Smoke-test budget"
A 5–10 iteration run with `--num-envs 64 --max-iterations 5` is enough to validate wiring
end-to-end. The reward signal will still be noisy at that scale; convergence requires the
150 / 300 iteration budgets above.

## Logs

Both tasks write to `logs/rsl_rl/<experiment>/<timestamp>_/` like the Unitree examples:

- `params/env.json` and `params/agent.json` — frozen configs at run time.
- `model_<iter>.pt` — checkpoints saved every `save_interval` iterations.
- TensorBoard event files alongside the checkpoints.

## See also

- [Unitree G1 quickstart](../getting-started/quickstart.md#unitree-g1)
- [Sensors](../concepts/sensors.md)
- [Play and Train CLI](../cli/play-train.md)
77 changes: 77 additions & 0 deletions docs/examples/inverted-pendulum.zh.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
# 倒立摆

`examples/inverted_pendulum/` 提供两个在平面上的 cart-pole 经典控制 PPO 任务。整条训练栈与
`examples/unitree/` 对齐:基于 Genesis 的 `ManagerBasedRlEnv`、rsl_rl PPO、挂在杆上的
`BodyVelocitySensor`,以及统一的 `genelab train` / `genelab play` CLI。

## 任务列表

| Task id | 问题 |
|---------|------|
| `GeneLab-Inverted-Pendulum-v0` | 小车 + 单杆的倒立摆。 |
| `GeneLab-Double-Inverted-Pendulum-v0` | 小车 + 串联双杆的倒立摆。 |

## 安装

扩展依赖 `rl` extra(rsl_rl)。`torch-*` extra 按硬件挑选。

```bash
uv sync --extra rl --extra torch-cu128
uv pip install -e examples/inverted_pendulum

uv run genelab list tasks
# -> GeneLab-Inverted-Pendulum-v0
# -> GeneLab-Double-Inverted-Pendulum-v0
```

## 单倒立摆

```bash
uv run genelab train GeneLab-Inverted-Pendulum-v0 \
--num-envs 4096 --max-iterations 150

uv run genelab play GeneLab-Inverted-Pendulum-v0 \
--checkpoint logs/rsl_rl/inverted_pendulum_flat/<run>/model_150.pt --vis
```

传入 `--checkpoint` 会让 `play` 自动经过 RL runner,并默认使用 `--agent trained`。

## 双倒立摆

```bash
uv run genelab train GeneLab-Double-Inverted-Pendulum-v0 \
--num-envs 4096 --max-iterations 300

uv run genelab play GeneLab-Double-Inverted-Pendulum-v0 \
--checkpoint logs/rsl_rl/double_inverted_pendulum_flat/<run>/model_300.pt --vis
```

## 传感器与欠驱动

只有小车的 slide 关节通过 PD 控制。两个 pole hinge 默认 `kp=0, kv=0`,保证整体处于欠驱动状态,
策略必须通过小车水平运动间接稳定杆。顶端 pole 上挂载的 `BodyVelocitySensor` 给出一路带噪声的
角速度观测:policy 观测组用 `Unoise` 做 corruption,critic 观测组直接读取干净值。

## 交互式扰动

`play` 默认只开 1 个环境(`num_envs=1`),并启用 Genesis 的 `MouseInteractionPlugin`。
左键点击 cart 或 pole 并拖动,会有一根弹簧把所点击的 link 拉向光标位置;策略仍然在背后试图
保持平衡。滚轮可绕表面法线旋转拖拽平面,松开左键即移除外力。

!!! tip "Smoke-test 预算"
使用 `--num-envs 64 --max-iterations 5` 跑 5–10 次迭代足以验证整条链路。此时 reward 信号
仍非常嘈杂,真正收敛需要上面给出的 150 / 300 次迭代预算。

## 日志

两个任务都把日志写到 `logs/rsl_rl/<experiment>/<timestamp>_/`,结构与 Unitree 示例一致:

- `params/env.json` 与 `params/agent.json` —— 运行时冻结的配置快照。
- `model_<iter>.pt` —— 按 `save_interval` 保存的 checkpoint。
- 同目录下的 TensorBoard 事件文件。

## See also

- [Unitree G1 快速开始](../getting-started/quickstart.md#unitree-g1)
- [传感器](../concepts/sensors.md)
- [play 与 train CLI](../cli/play-train.md)
10 changes: 10 additions & 0 deletions docs/examples/overview.en.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,16 @@
The repository ships several reference extensions under `examples/`. They double as integration
tests for the CLI and registry.

## inverted_pendulum

Two PPO cart-pole tasks built on the same `ManagerBasedRlEnv` + rsl_rl stack as the Unitree
example, sized to fit in a laptop training budget:

- **`GeneLab-Inverted-Pendulum-v0`** — single inverted pole on a cart.
- **`GeneLab-Double-Inverted-Pendulum-v0`** — two stacked inverted poles on a cart.

Source at `examples/inverted_pendulum/`; walkthrough at [Inverted Pendulum](inverted-pendulum.md).

## genelab_examples

The canonical in-tree extension, wiring two tasks:
Expand Down
10 changes: 10 additions & 0 deletions docs/examples/overview.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,16 @@

仓库在 `examples/` 下提供数个参考扩展,同时也是 CLI 与注册表的集成测试。

## inverted_pendulum

两个 PPO cart-pole 任务,训练栈与 Unitree 示例相同(`ManagerBasedRlEnv` + rsl_rl),训练预算
控制在单机能跑完的量级:

- **`GeneLab-Inverted-Pendulum-v0`** —— 小车 + 单杆倒立摆。
- **`GeneLab-Double-Inverted-Pendulum-v0`** —— 小车 + 串联双杆倒立摆。

源码位于 `examples/inverted_pendulum/`;完整流程见 [倒立摆](inverted-pendulum.md)。

## genelab_examples

仓库内的标准扩展,接通两个任务:
Expand Down
32 changes: 26 additions & 6 deletions examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,26 +5,36 @@ ship built-in tasks; example tasks are loaded like any other external project.

## Available Examples

- [Inverted Pendulum](inverted_pendulum/README.md): trainable single- and double-inverted-pendulum tasks that initialize Genesis while fully exercising `train` + `play`.
- [GeneLab Example Extension](genelab_examples/README.md): one Python project that registers the
Rubik's cube and Wuji hand tasks.
- [External Project](external_project/README.md): minimal standalone Python package that extends
GeneLab without editing `src/genelab/`.

The example extension registers these task IDs:
The bundled examples register these task IDs:

- `GeneLab-Inverted-Pendulum-v0`
- `GeneLab-Double-Inverted-Pendulum-v0`
- `GeneLab-Rubiks-Play-v0`
- `GeneLab-Wuji-Hand-Playback-v0`

List example tasks from the repository root without installing the example package:
List inverted-pendulum tasks from the repository root without installing the package:

```bash
PYTHONPATH=examples/inverted_pendulum/src uv run genelab --import genelab_inverted_pendulum.tasks list tasks
```

List the Genesis demo tasks the same way:

```bash
PYTHONPATH=examples/genelab_examples/src uv run genelab --import genelab_examples.tasks list tasks
```

Install the example extension once if you want `uv run genelab list tasks` to load it through entry
Install an example extension once if you want `uv run genelab list tasks` to load it through entry
points:

```bash
uv pip install -e examples/inverted_pendulum
uv pip install -e examples/genelab_examples
uv run genelab list tasks
```
Expand All @@ -43,21 +53,31 @@ override keys are converted to underscores.
Examples:

```bash
PYTHONPATH=examples/inverted_pendulum/src uv run genelab --import genelab_inverted_pendulum.tasks train GeneLab-Inverted-Pendulum-v0 --num-envs 4096 --max-iterations 150
PYTHONPATH=examples/inverted_pendulum/src uv run genelab --import genelab_inverted_pendulum.tasks play GeneLab-Inverted-Pendulum-v0 --checkpoint logs/rsl_rl/inverted_pendulum_flat/<run>/model_150.pt --vis
PYTHONPATH=examples/genelab_examples/src uv run genelab --import genelab_examples.tasks play GeneLab-Rubiks-Play-v0 --steps 5 --env.robot.cubie_size 0.04 --env.robot.gap 0.002
PYTHONPATH=examples/genelab_examples/src uv run genelab --import genelab_examples.tasks play GeneLab-Rubiks-Play-v0 --env.robot.welded true
PYTHONPATH=examples/genelab_examples/src uv run genelab --import genelab_examples.tasks play GeneLab-Wuji-Hand-Playback-v0 --env.reset_interval 0
```

`train` validates the task id and configuration path but currently reports that training is not
implemented:
`train` is implemented for the inverted-pendulum tasks. The Rubik's cube and Wuji hand demo tasks are
play-only and report that training is not implemented:

```bash
PYTHONPATH=examples/genelab_examples/src uv run genelab --import genelab_examples.tasks train GeneLab-Rubiks-Play-v0
```

## Smoke Tests

Run short headless smoke tests after the Genesis assets and cache have initialized:
Run the inverted-pendulum smoke tests first; a tiny rsl_rl run exercises the full Genesis +
PPO pipeline:

```bash
PYTHONPATH=examples/inverted_pendulum/src uv run genelab --import genelab_inverted_pendulum.tasks train GeneLab-Inverted-Pendulum-v0 --num-envs 64 --max-iterations 5
PYTHONPATH=examples/inverted_pendulum/src uv run genelab --import genelab_inverted_pendulum.tasks train GeneLab-Double-Inverted-Pendulum-v0 --num-envs 64 --max-iterations 5
```

Then run short headless smoke tests after the Genesis assets and cache have initialized:

```bash
PYTHONPATH=examples/genelab_examples/src uv run genelab --import genelab_examples.tasks play GeneLab-Rubiks-Play-v0 --steps 5
Expand Down
63 changes: 63 additions & 0 deletions examples/inverted_pendulum/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# Inverted pendulum examples

GeneLab extension that ships two PPO tasks for the classical cart-pole problem on flat ground:

- **`GeneLab-Inverted-Pendulum-v0`** — balance a single inverted pole on a cart.
- **`GeneLab-Double-Inverted-Pendulum-v0`** — stabilise two stacked inverted poles on a cart.

The training stack mirrors `examples/unitree/`: `ManagerBasedRlEnv` over Genesis, rsl_rl PPO,
`BodyVelocitySensor` for the noisy pole rate observation, and the unified `genelab train` /
`genelab play` CLI.

## Layout

```
examples/inverted_pendulum/
├── pyproject.toml
├── README.md
├── assets/ # cart-pole MJCFs
│ ├── inverted_pendulum.xml
│ └── double_inverted_pendulum.xml
└── src/genelab_inverted_pendulum/
├── tasks.py # registers both tasks
├── mdp.py # cart-pole-specific reward / termination / event terms
├── single/ # single-pendulum config (robot + env + PPO)
└── double/ # double-pendulum config (robot + env + PPO)
```

## Quickstart

```bash
# From the GeneLab repo root
uv sync --extra rl --extra torch-cu128 # pick whichever torch flavor fits your GPU
uv pip install -e examples/inverted_pendulum

uv run genelab list tasks
# -> GeneLab-Inverted-Pendulum-v0
# -> GeneLab-Double-Inverted-Pendulum-v0
```

### Single inverted pendulum

```bash
uv run genelab train GeneLab-Inverted-Pendulum-v0 --num_envs 4096 --max_iterations 150
uv run genelab play GeneLab-Inverted-Pendulum-v0 \
--checkpoint logs/rsl_rl/inverted_pendulum_flat/<run>/model_150.pt
```

### Double inverted pendulum

```bash
uv run genelab train GeneLab-Double-Inverted-Pendulum-v0 --num_envs 4096 --max_iterations 300
uv run genelab play GeneLab-Double-Inverted-Pendulum-v0 \
--checkpoint logs/rsl_rl/double_inverted_pendulum_flat/<run>/model_300.pt
```

## Notes

- Only the cart slide joint is PD-controlled. The pole hinges default to `kp=0, kv=0` so the
pendulum stays underactuated and the policy must learn balance via cart motion alone.
- The observation group corrupts joint position, joint velocity, and pole angular velocity with
`Unoise`. The critic group sees the same features without corruption.
- Logs land under `logs/rsl_rl/<experiment>/<timestamp>_/` exactly like the Unitree examples,
with `params/env.json`, `params/agent.json`, and `model_<iter>.pt` files.
30 changes: 30 additions & 0 deletions examples/inverted_pendulum/assets/double_inverted_pendulum.xml
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
<mujoco model="genelab_double_inverted_pendulum">
<option timestep="0.005" gravity="0 0 -9.81" integrator="implicitfast"/>
<default>
<joint armature="0" damping="0"/>
<geom contype="0" conaffinity="0"/>
</default>
<worldbody>
<body name="cart" pos="0 0 0.12">
<joint name="cart_slide" type="slide" axis="1 0 0" range="-3.0 3.0" damping="0.05"/>
<!-- contype/conaffinity bits picked so each geom is present in the rigid solver
(visible to the mouse-interaction raycaster) but never collides with the
plane or with sibling links. -->
<geom name="cart_box" type="box" size="0.17 0.11 0.05" mass="1.0"
rgba="0.10 0.45 0.85 1" contype="2" conaffinity="2"/>
<body name="pole_1" pos="0 0 0.05">
<joint name="pole_1_hinge" type="hinge" axis="0 1 0" range="-1.2 1.2" damping="0.002"/>
<geom name="pole_1_capsule" type="capsule" fromto="0 0 0 0 0 0.55" size="0.023"
mass="0.1" rgba="0.95 0.35 0.15 1" contype="4" conaffinity="4"/>
<body name="pole_2" pos="0 0 0.55">
<joint name="pole_2_hinge" type="hinge" axis="0 1 0" range="-1.5 1.5" damping="0.002"/>
<geom name="pole_2_capsule" type="capsule" fromto="0 0 0 0 0 0.45" size="0.020"
mass="0.08" rgba="0.95 0.65 0.12 1" contype="8" conaffinity="8"/>
</body>
</body>
</body>
</worldbody>
<actuator>
<motor name="cart_force" joint="cart_slide" gear="1"/>
</actuator>
</mujoco>
25 changes: 25 additions & 0 deletions examples/inverted_pendulum/assets/inverted_pendulum.xml
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
<mujoco model="genelab_inverted_pendulum">
<option timestep="0.005" gravity="0 0 -9.81" integrator="implicitfast"/>
<default>
<joint armature="0" damping="0"/>
<geom contype="0" conaffinity="0"/>
</default>
<worldbody>
<body name="cart" pos="0 0 0.12">
<joint name="cart_slide" type="slide" axis="1 0 0" range="-3.0 3.0" damping="0.05"/>
<!-- contype/conaffinity bits picked so the geom is present in the rigid solver
(visible to the mouse-interaction raycaster) but never participates in a
collision pair with the plane or the pole. -->
<geom name="cart_box" type="box" size="0.15 0.10 0.05" mass="1.0"
rgba="0.10 0.45 0.85 1" contype="2" conaffinity="2"/>
<body name="pole" pos="0 0 0.05">
<joint name="pole_hinge" type="hinge" axis="0 1 0" range="-1.2 1.2" damping="0.002"/>
<geom name="pole_capsule" type="capsule" fromto="0 0 0 0 0 0.80" size="0.025"
mass="0.1" rgba="0.95 0.35 0.15 1" contype="4" conaffinity="4"/>
</body>
</body>
</worldbody>
<actuator>
<motor name="cart_force" joint="cart_slide" gear="1"/>
</actuator>
</mujoco>
Loading