Project overview · Models · Demo gallery · Robot deployment
Use this guide to turn a WAV into robot motion and a synchronized video, or to render an existing audio–motion pair. These offline workflows do not open a microphone, contact ASR/TTS services or send commands to a robot.
Install the environment using the quick start. Run the commands below from the repository root.
- Rendering: the environment and bundled robot assets are sufficient; no neural model weights are needed.
- WAV inference: retrieve the trained motion checkpoint with Git LFS and download Mimi. Qwen is only needed for dialogue.
git lfs pull
.venv/bin/python scripts/download_models.py --only mimi.venv/bin/robogesture infer examples/sample_06.wav \
--output outputs/generated.npzReplace the example path with your own WAV. The command produces:
generated.npz: 30 FPS motion and inference metadata.generated.wav: mono PCM16 audio at 24 kHz, aligned to the motion timeline.
The reader accepts different WAV sample rates and channel counts. It averages channels to mono and resamples to 24 kHz. Input must be nonempty and contain finite samples. No transcript or separate audio preprocessing command is needed.
The motion sampler generates a 30-frame block per nominal second. It pads a short final block, and the exported WAV is padded with silence to match the resulting motion duration. Use this exported WAV when rendering a newly generated NPZ.
The default output applies smoothing and collision filtering in a fixed-base scene with zero leg angles. Its collision metadata describes that offline scene, not measured G1 state or a guarantee about physical playback. Use the robot preflight before any real execution.
To export the original 41-joint model motion without smoothing/projection in the
exported qpos, use:
.venv/bin/robogesture infer examples/sample_06.wav \
--output outputs/generated_raw.npz --rawCUDA is selected by default. --device selects the PyTorch device; the tested
inference baseline is the RTX 4090 workstation described in the README.
.venv/bin/robogesture render outputs/generated.npz \
--output outputs/generated.mp4 \
--preview outputs/generated.pngThe renderer reads the adjacent WAV by default. Specify another paired WAV with
--wav when needed. Audio must be nonempty and no longer than the motion timeline;
shorter audio is padded with silence in the video.
| Option | Default / behavior |
|---|---|
--wav PATH |
Defaults to the input NPZ path with a .wav suffix |
--output PATH |
Required; must end in .mp4 |
--preview PATH |
Optionally saves the middle frame as a PNG |
--width, --height |
960 × 720; each must be an even integer of at least 64 |
--gl egl |
Default headless backend, used in local GPU acceptance |
--gl osmesa |
Alternative for a host configured with system OSMesa; not the tested backend |
--overwrite |
Explicitly permits replacement of existing output files |
Output is 30 FPS H.264 video with AAC audio. MuJoCo uses the repository's fixed-base
G1/BrainCo model and meshes; FFmpeg comes from imageio-ffmpeg inside .venv/.
If MUJOCO_GL is already set in your shell, it takes precedence over --gl.
For all 15 bundled examples, use the batch rendering command.
The renderer expects finite floating-point qpos[T,60] with at least one frame.
| Columns | Meaning in the fixed-base renderer |
|---|---|
0:7 |
Reserved root pose fields; unused by this renderer |
7:19 |
12 leg joints |
19:60 |
41 upper-body and hand joints |
The infer command writes fps=30. Input NPZ files without an fps field are
interpreted as 30 FPS. A declared rate of 60 FPS is rejected; the private training
format is not this replay contract.
Inference exports also include raw_motion41, motion30, audio_rate,
audio_samples and collision_filter. Default filtered exports additionally
record safe, corrected and held frame counts and minimum clearance. With --raw,
the exported qpos uses the original model output; the stored motion30 remains
the smoothed 30-channel representation, without collision filtering.
Code, checkpoints and robot assets resolve relative to the checkout. Public models
default to its sibling weights/ directory, independent of the current shell
directory. Your input and output arguments are relative to your shell directory
unless you provide absolute paths.
To use another public-model directory, set the same location for downloading and running:
export ROBOGESTURE_WEIGHTS_ROOT=/absolute/path/to/weights
.venv/bin/python scripts/download_models.py \
--weights-root "$ROBOGESTURE_WEIGHTS_ROOT"This directory should contain mimi/ and, for dialogue,
Qwen3-4B-Instruct-2507/. The trained checkpoint and LoRA adapter stay inside
checkpoints/. See the model guide for revisions and
integrity checks.
Inference and rendering protect existing outputs. Choose a fresh output name or
pass --overwrite when replacement is intentional. The inference output WAV must
never replace the input WAV, even with --overwrite.
Lightweight checks:
.venv/bin/python -m unittest discover -s tests -v
git lfs fsckAfter downloading both official models, run the complete offline verification:
.venv/bin/python scripts/check_install.py --require-cuda
.venv/bin/python scripts/verify_offline.py --output acceptance-outputThe acceptance directory must be empty. The runner checks real Qwen/LoRA and motion inference, renders every example plus a newly generated motion, and fully decodes the video and audio streams. It does not call cloud services or connect a robot. See the recorded results for the tested setup and limitations.