Skip to content

Latest commit

 

History

History
142 lines (107 loc) · 5.79 KB

File metadata and controls

142 lines (107 loc) · 5.79 KB

🤖 English Interaction on G1

Project overview · Models · Offline inference · Verification

Bring the speech-to-gesture workflow to a Unitree G1 with BrainCo Revo2 hands. This guide connects the GPU workstation's English dialogue pipeline to the robot's audio and motion interfaces, using a field-validated deployment pipeline.

Each turn follows the same sequence: English ASR → Qwen + LoRA → TTS → complete motion generation and collision filtering → synchronized playback. The stream entry point is interactive, but playback is whole-utterance, not incremental.

Hardware and runtime configuration

The verified robot configuration is: GPU DDS interface enp4s0; G1 interface eth2; DDS domain 0; FSM 501; mode_pr=0, mode_machine=5; BrainCo Revo2 hands; an explicitly configured USB PulseAudio sink. Validation does not cover other hardware configurations.

The GPU workstation uses the full Python 3.10 environment. The G1 audio service only needs this checkout's robogesture.config and robogesture.robot_audio, which support Python 3.8. The robot does not need neural checkpoints or CUDA.

1. Set up the robot environment

On the G1, create a separate environment inside its own checkout:

sudo apt-get install build-essential cmake git python3.8-dev python3.8-venv libgl1 libglib2.0-0
cd RoboGesture
bash scripts/setup_robot.sh

Unlike the x86-64 workstation wheel, CycloneDDS 0.10.2 has no Python 3.8 ARM64 wheel. The installer therefore builds official CycloneDDS 0.10.2 at revision 9995905bce6c4cf9f740d6438bbf7fcfd1c83dfd into this checkout's .venv/lib, then builds the Python binding against it. run_robot_audio.py selects that environment's DDS library rather than a system-level DDS installation. This follows the upstream source-installation procedure.

The install finishes with a read-only import/IDL serialization check:

.venv/bin/python scripts/check_robot_install.py

2. Prepare the robot audio service

The device requires PulseAudio (pacat, pactl) and system-level brainco_hand.service. The latter supplies the left/right DDS state topics and must be enabled for cold boot. These vendor robot drivers/firmware, like the NVIDIA driver on the workstation, are host prerequisites. This repository does not replace the low-level motor or hand firmware. Hand control uses the DDS interfaces provided by the robot services.

Configure the audio sink using the exact name returned by pactl list short sinks:

mkdir -p ~/.config
# Create ~/.config/robogesture.env with:
# ROBOGESTURE_ROBOT_AUDIO_SINK=the-actual-sink-name
chmod 600 ~/.config/robogesture.env

systemd/robogesture-robot-audio.service assumes a robot checkout at ~/workplace2/RoboGesture. Adjust its two path entries for another location. Then install the user service:

mkdir -p ~/.config/systemd/user
cp systemd/robogesture-robot-audio.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now robogesture-robot-audio.service
sudo systemctl enable --now brainco_hand.service

3. Configure the GPU workstation

Complete the workstation setup and download both Mimi and Qwen using the model guide.

Recipients supply their own Volcengine ASR/TTS service credentials. Create an ignored .env from .env.example, fill in your credentials, then export them in the shell used to start the application:

cp .env.example .env
chmod 600 .env
# Fill in ROBOGESTURE_VOLC_APP_ID and ROBOGESTURE_VOLC_ACCESS_KEY before continuing.
set -a
. ./.env
set +a

Keep the populated .env private. The supplied configuration specifies ASR/TTS resource identifiers and an English voice. Speak English: the input gate requires Latin letters and excludes CJK; it is not a general language detector.

4. Preflight, then interact

Start with the disconnected entry point:

.venv/bin/robogesture stream

The initial command robogesture stream does not connect devices. A live --execute --preflight needs G1 state and an audio-service handshake, so it is not an offline simulation. Actual playback also requires --confirm SEND_MOTION, an operator, clear surroundings and an immediately accessible emergency stop.

After preparing the services and confirming the on-site conditions:

.venv/bin/robogesture stream --execute --preflight --listen-seconds 30
# Enable actual playback only after a successful preflight and operator approval.
.venv/bin/robogesture stream --execute --confirm SEND_MOTION --listen-seconds 120

Preflight requires live state and the audio-service handshake, but does not send motion or audio playback commands. The microphone is closed while the robot answers.

Runtime behavior and validation scope

The control implementation includes collision checks, 30 Hz pacing, FSM gating, listening-pose restore and SDK release. Do not remove those behaviors as part of installation or device configuration.

standing.py is available as a one-transition-at-a-time FSM utility. run_gpu.py and run_robot_audio.py are direct-script entry points. No robot services are installed or started by the GPU setup script or automated acceptance checks.

RoboGesture's local acceptance covers offline models, rendering and software-only robot-environment checks. Automated checks do not operate a physical G1 or test recipient cloud credentials; see the verification record.

The audio topic names (rt/audiomotion/audio/*) and IDL type names (audiomotion.msg.RobotAudio*) are wire-protocol identifiers, not Python imports. The GPU client and robot audio service must use matching topic and type names.