Skip to content

Latest commit

 

History

History
157 lines (116 loc) · 6.08 KB

File metadata and controls

157 lines (116 loc) · 6.08 KB

🚀 Inference & Visualization Guide

Project overview · Models · Demo gallery · Robot deployment

Use this guide to turn a WAV into robot motion and a synchronized video, or to render an existing audio–motion pair. These offline workflows do not open a microphone, contact ASR/TTS services or send commands to a robot.

Before you start

Install the environment using the quick start. Run the commands below from the repository root.

  • Rendering: the environment and bundled robot assets are sufficient; no neural model weights are needed.
  • WAV inference: retrieve the trained motion checkpoint with Git LFS and download Mimi. Qwen is only needed for dialogue.
git lfs pull
.venv/bin/python scripts/download_models.py --only mimi

Generate motion from speech

.venv/bin/robogesture infer examples/sample_06.wav \
  --output outputs/generated.npz

Replace the example path with your own WAV. The command produces:

  • generated.npz: 30 FPS motion and inference metadata.
  • generated.wav: mono PCM16 audio at 24 kHz, aligned to the motion timeline.

The reader accepts different WAV sample rates and channel counts. It averages channels to mono and resamples to 24 kHz. Input must be nonempty and contain finite samples. No transcript or separate audio preprocessing command is needed.

The motion sampler generates a 30-frame block per nominal second. It pads a short final block, and the exported WAV is padded with silence to match the resulting motion duration. Use this exported WAV when rendering a newly generated NPZ.

Motion processing

The default output applies smoothing and collision filtering in a fixed-base scene with zero leg angles. Its collision metadata describes that offline scene, not measured G1 state or a guarantee about physical playback. Use the robot preflight before any real execution.

To export the original 41-joint model motion without smoothing/projection in the exported qpos, use:

.venv/bin/robogesture infer examples/sample_06.wav \
  --output outputs/generated_raw.npz --raw

CUDA is selected by default. --device selects the PyTorch device; the tested inference baseline is the RTX 4090 workstation described in the README.

Render a synchronized video

.venv/bin/robogesture render outputs/generated.npz \
  --output outputs/generated.mp4 \
  --preview outputs/generated.png

The renderer reads the adjacent WAV by default. Specify another paired WAV with --wav when needed. Audio must be nonempty and no longer than the motion timeline; shorter audio is padded with silence in the video.

Option Default / behavior
--wav PATH Defaults to the input NPZ path with a .wav suffix
--output PATH Required; must end in .mp4
--preview PATH Optionally saves the middle frame as a PNG
--width, --height 960 × 720; each must be an even integer of at least 64
--gl egl Default headless backend, used in local GPU acceptance
--gl osmesa Alternative for a host configured with system OSMesa; not the tested backend
--overwrite Explicitly permits replacement of existing output files

Output is 30 FPS H.264 video with AAC audio. MuJoCo uses the repository's fixed-base G1/BrainCo model and meshes; FFmpeg comes from imageio-ffmpeg inside .venv/. If MUJOCO_GL is already set in your shell, it takes precedence over --gl.

For all 15 bundled examples, use the batch rendering command.

NPZ format

The renderer expects finite floating-point qpos[T,60] with at least one frame.

Columns Meaning in the fixed-base renderer
0:7 Reserved root pose fields; unused by this renderer
7:19 12 leg joints
19:60 41 upper-body and hand joints

The infer command writes fps=30. Input NPZ files without an fps field are interpreted as 30 FPS. A declared rate of 60 FPS is rejected; the private training format is not this replay contract.

Inference exports also include raw_motion41, motion30, audio_rate, audio_samples and collision_filter. Default filtered exports additionally record safe, corrected and held frame counts and minimum clearance. With --raw, the exported qpos uses the original model output; the stored motion30 remains the smoothed 30-channel representation, without collision filtering.

Model paths

Code, checkpoints and robot assets resolve relative to the checkout. Public models default to its sibling weights/ directory, independent of the current shell directory. Your input and output arguments are relative to your shell directory unless you provide absolute paths.

To use another public-model directory, set the same location for downloading and running:

export ROBOGESTURE_WEIGHTS_ROOT=/absolute/path/to/weights
.venv/bin/python scripts/download_models.py \
  --weights-root "$ROBOGESTURE_WEIGHTS_ROOT"

This directory should contain mimi/ and, for dialogue, Qwen3-4B-Instruct-2507/. The trained checkpoint and LoRA adapter stay inside checkpoints/. See the model guide for revisions and integrity checks.

Output protection & checks

Inference and rendering protect existing outputs. Choose a fresh output name or pass --overwrite when replacement is intentional. The inference output WAV must never replace the input WAV, even with --overwrite.

Lightweight checks:

.venv/bin/python -m unittest discover -s tests -v
git lfs fsck

After downloading both official models, run the complete offline verification:

.venv/bin/python scripts/check_install.py --require-cuda
.venv/bin/python scripts/verify_offline.py --output acceptance-output

The acceptance directory must be empty. The runner checks real Qwen/LoRA and motion inference, renders every example plus a newly generated motion, and fully decodes the video and audio streams. It does not call cloud services or connect a robot. See the recorded results for the tested setup and limitations.