Reachy Mini · Voice Agent

Milo

A real-time conversational voice agent running on a Reachy Mini desktop robot. The robot handles audio capture and playback; a cloud speech pipeline handles understanding and response. This page documents how the system is put together.

End-to-end pipeline

One conversational turn, from microphone to speaker.

ON DEVICE — Reachy Mini CLOUD — Cloud Run 4-mic array 16 kHz capture Endpointing RMS gate → Smart Turn ONNX Speaker streamed 16 kHz playback Barge-in monitor concurrent thread STT Google LLM Gemini 2.5 Flash TTS Chirp3-HD 24 kHz PCM response stream chunked transfer · aborted on barge-in WAV upload → ← streamed back

Knowing when the user has finished speaking

The hardest part of a voice agent is not transcription — it is deciding when to reply. Cutting in too early truncates the user; waiting too long feels broken.

A naive silence timer fails on natural speech, because people pause mid-thought. Milo uses two stages: a cheap signal test to notice a pause, then a learned model to decide whether that pause actually ended the turn.

1

RMS gate cheap · every 20 ms

Root-mean-square energy per audio chunk, compared against a threshold. Speech starts when energy crosses it; a silence timer starts when it drops. Runs continuously, costs almost nothing. Utterances shorter than the minimum duration are discarded outright as noise.

2

Smart Turn v3 ONNX · on trigger only

Once the silence timer expires, the buffered audio is run through a Smart Turn ONNX model that returns the probability the turn is semantically complete. Above threshold, Milo replies. Below it, the silence timer resets and Milo keeps listening — the user was just thinking.

Why this split: the model only runs when the cheap test already suspects the turn ended, so inference happens once per pause rather than continuously. Whisper feature extraction plus an ONNX session on the robot's CPU is far too expensive to run every 20 ms — but perfectly affordable once.

Failure behaviour: if the model fails to load or throws, prediction returns 1.0 and the system falls back to plain silence-timer VAD. Endpointing degrades in quality but never hangs — the agent always eventually replies.

Latency: streaming instead of waiting

Perceived responsiveness is dominated by time-to-first-word, not total response time.

The obvious implementation downloads the entire TTS response, resamples it, then plays it — so the user waits for the whole reply to arrive before hearing anything. Milo instead consumes the HTTP response incrementally: it accumulates roughly half a second of PCM, resamples that chunk, pushes it to the speaker, and repeats. Playback begins while the rest of the sentence is still in flight.

16 kHz mic capture
→
24 kHz TTS output
→
16 kHz speaker playback

The speaker cannot consume the TTS sample rate directly, so every chunk is resampled 24 kHz → 16 kHz in flight rather than in one pass at the end.

Interrupting the robot mid-sentence

A voice agent that cannot be interrupted feels like a recording, not a conversation.

While audio is playing, a background thread keeps the microphone open and watches for sustained speech energy. It requires several consecutive loud chunks before firing, with a quiet-period reset, so a door closing or a single transient does not count as an interruption.

When it does fire, the effect is immediate: the in-flight HTTP response is closed, the resample-and-push loop breaks, and playback stops. Milo returns to listening rather than finishing a reply nobody is hearing any more.

Hardware constraint: this requires the robot SDK to support recording and playback concurrently. If that call fails, the monitor thread exits cleanly through its finally block and barge-in is silently disabled for that turn — the conversation continues uninterrupted rather than crashing.

Splitting work between robot and cloud

Default is cloud-heavy; the split is configurable.

By default the robot is deliberately thin — it records, uploads, streams back, and plays. All transcription, reasoning, and synthesis happen in the cloud, which keeps the on-device footprint small and the models swappable without touching the robot.

Transcription can optionally move on-device instead, running faster-whisper with int8 quantization on CPU. That trades accuracy and CPU headroom for lower round-trip latency and the ability to keep raw audio local. Endpointing already runs on-device in both configurations.

StageDefaultAlternative
EndpointingOn device — ONNX—
TranscriptionCloud STTOn device — faster-whisper int8
ReasoningCloud LLM—
SynthesisCloud TTS—

Running it on a Reachy Mini

The robot app is the thin client. It needs a speech backend to talk to, and credentials to reach it — both supplied through environment variables.

1 · Install on the robot

From the Reachy Mini dashboard at http://reachy-mini.local:8000, choose Install from Hugging Face and search for this Space. Once installed, the app appears as a tile you can start and stop. To run it from source instead, pip install -e . inside the package directory and launch with python -m reachy_mini_milo.

2 · Point it at a backend

This is the step that matters most. MILO_SERVICE_URL defaults to the author's own private Cloud Run deployment, which you will not be able to call. Set it to a speech service of your own that exposes a /speech2speech endpoint — accepting a WAV upload and returning a 24 kHz PCM stream.

3 · Handle authentication

If your backend is publicly reachable, set MILO_GCP_AUTH=false and you are done. If it is a private Cloud Run service, the app fetches a GCP identity token and refreshes it before the hour is out. It resolves credentials in this order, using the first that works:

OrderVariableUse case
1MILO_GCP_TOKEN A pre-fetched token — handy for one-off runs over SSH
2MILO_SA_KEY_JSON Whole service account key as a string — best for remote robots
3MILO_SA_KEY_PATH Path to a key file already copied onto the robot
4gcloud CLI Fallback, if the CLI is installed and authenticated

4 · Tune for your room (optional)

The defaults suit a reasonably quiet space. If Milo triggers on background noise or cuts you off mid-sentence, the endpointing variables below are the ones to reach for.

VariableDefaultEffect
MILO_VAD_RMS_THRESHOLD0.015 Raise in a noisy room so ambient sound stops registering as speech
MILO_VAD_SILENCE_DURATION0.7 Seconds of quiet before the turn model is consulted
MILO_VAD_MIN_SPEECH_DURATION0.4 Utterances shorter than this are dropped as noise
MILO_USE_SMART_TURNtrue Set false to use the plain silence timer alone
MILO_SMART_TURN_THRESHOLD0.5 Lower it to reply sooner, raise it to wait through longer pauses
MILO_BARGE_INtrue Set false if your hardware cannot record and play at once
MILO_BARGE_IN_RMS_THRESHOLD0.03 Raise it if Milo interrupts itself on its own playback
MILO_USE_LOCAL_STTfalse Run faster-whisper on device instead of sending audio to cloud STT
MILO_STT_MODELtiny Local model size — tiny, base, or small

One at a time. Do not run the dashboard tile and an SSH session together — both grab the microphone and speaker, and you get doubled audio.