Project overview · Inference guide · Sample manifest
Explore 15 paired speech and motion clips on the bundled G1/BrainCo model.
Each sample_XX.npz has a matching sample_XX.wav, ready for a synchronized
MuJoCo video after environment setup. Rendering these saved motions needs no
neural checkpoints, cloud credentials or physical robot.
Run the commands below from the RoboGesture repository root.
.venv/bin/robogesture render examples/sample_01.npz \
--output outputs/sample_01.mp4 \
--preview outputs/sample_01.pngThe adjacent WAV is selected automatically. Open the MP4 to watch motion with speech, or inspect the PNG for a middle-frame preview. The default video is 960×720 at 30 FPS, with H.264 video and AAC audio.
for motion in examples/sample_*.npz; do
.venv/bin/robogesture render "$motion" \
--output "outputs/$(basename "${motion%.npz}").mp4"
doneThe 15 clips span 196 seconds of motion. Existing outputs are not overwritten by
default; choose a fresh output directory or add --overwrite if you intend to
replace an earlier render.
After downloading Mimi and retrieving the motion checkpoint, use any example WAV as a model input:
.venv/bin/robogesture infer examples/sample_06.wav \
--output outputs/new_motion.npz
.venv/bin/robogesture render outputs/new_motion.npz \
--output outputs/new_motion.mp4This runs inference again; it is not a replay or copy of the saved example NPZ. Keep the newly exported WAV beside the NPZ so its padded duration matches the new motion timeline.
The release includes 15 demonstration pairs, named sample_01 through sample_15.
manifest.json records their filenames, SHA-256 hashes, frame
counts and audio durations. These are demonstration assets, not an open training
dataset or a benchmark split.
All motions follow the 30 FPS qpos[T,60] replay contract. See the
format guide for column meanings and the distinction from
the unreleased training data. The verification record
covers rendering and full audio/video decoding of all 15 pairs.