Reachy Mini · Voice Agent
A real-time conversational voice agent running on a Reachy Mini desktop robot. The robot handles audio capture and playback; a cloud speech pipeline handles understanding and response. This page documents how the system is put together.
One conversational turn, from microphone to speaker.
The hardest part of a voice agent is not transcription — it is deciding when to reply. Cutting in too early truncates the user; waiting too long feels broken.
A naive silence timer fails on natural speech, because people pause mid-thought. Milo uses two stages: a cheap signal test to notice a pause, then a learned model to decide whether that pause actually ended the turn.
Root-mean-square energy per audio chunk, compared against a threshold. Speech starts when energy crosses it; a silence timer starts when it drops. Runs continuously, costs almost nothing. Utterances shorter than the minimum duration are discarded outright as noise.
Once the silence timer expires, the buffered audio is run through a Smart Turn ONNX model that returns the probability the turn is semantically complete. Above threshold, Milo replies. Below it, the silence timer resets and Milo keeps listening — the user was just thinking.
Why this split: the model only runs when the cheap test already suspects the turn ended, so inference happens once per pause rather than continuously. Whisper feature extraction plus an ONNX session on the robot's CPU is far too expensive to run every 20 ms — but perfectly affordable once.
Failure behaviour: if the model fails to load or throws, prediction
returns 1.0 and the system falls back to plain silence-timer VAD.
Endpointing degrades in quality but never hangs — the agent always eventually
replies.
Perceived responsiveness is dominated by time-to-first-word, not total response time.
The obvious implementation downloads the entire TTS response, resamples it, then plays it — so the user waits for the whole reply to arrive before hearing anything. Milo instead consumes the HTTP response incrementally: it accumulates roughly half a second of PCM, resamples that chunk, pushes it to the speaker, and repeats. Playback begins while the rest of the sentence is still in flight.
The speaker cannot consume the TTS sample rate directly, so every chunk is resampled 24 kHz → 16 kHz in flight rather than in one pass at the end.
A voice agent that cannot be interrupted feels like a recording, not a conversation.
While audio is playing, a background thread keeps the microphone open and watches for sustained speech energy. It requires several consecutive loud chunks before firing, with a quiet-period reset, so a door closing or a single transient does not count as an interruption.
When it does fire, the effect is immediate: the in-flight HTTP response is closed, the resample-and-push loop breaks, and playback stops. Milo returns to listening rather than finishing a reply nobody is hearing any more.
Hardware constraint: this requires the robot SDK to support recording
and playback concurrently. If that call fails, the monitor thread exits cleanly
through its finally block and barge-in is silently disabled for that
turn — the conversation continues uninterrupted rather than crashing.
Default is cloud-heavy; the split is configurable.
By default the robot is deliberately thin — it records, uploads, streams back, and plays. All transcription, reasoning, and synthesis happen in the cloud, which keeps the on-device footprint small and the models swappable without touching the robot.
Transcription can optionally move on-device instead, running faster-whisper with int8 quantization on CPU. That trades accuracy and CPU headroom for lower round-trip latency and the ability to keep raw audio local. Endpointing already runs on-device in both configurations.
| Stage | Default | Alternative |
|---|---|---|
| Endpointing | On device — ONNX | — |
| Transcription | Cloud STT | On device — faster-whisper int8 |
| Reasoning | Cloud LLM | — |
| Synthesis | Cloud TTS | — |
The robot app is the thin client. It needs a speech backend to talk to, and credentials to reach it — both supplied through environment variables.
From the Reachy Mini dashboard at http://reachy-mini.local:8000,
choose Install from Hugging Face and search for this Space.
Once installed, the app appears as a tile you can start and stop. To run it
from source instead, pip install -e . inside the package
directory and launch with python -m reachy_mini_milo.
This is the step that matters most. MILO_SERVICE_URL defaults to
the author's own private Cloud Run deployment, which you will not be able to
call. Set it to a speech service of your own that exposes a
/speech2speech endpoint — accepting a WAV upload and returning a
24 kHz PCM stream.
If your backend is publicly reachable, set MILO_GCP_AUTH=false and
you are done. If it is a private Cloud Run service, the app fetches a GCP
identity token and refreshes it before the hour is out. It resolves credentials
in this order, using the first that works:
| Order | Variable | Use case |
|---|---|---|
| 1 | MILO_GCP_TOKEN |
A pre-fetched token — handy for one-off runs over SSH |
| 2 | MILO_SA_KEY_JSON |
Whole service account key as a string — best for remote robots |
| 3 | MILO_SA_KEY_PATH |
Path to a key file already copied onto the robot |
| 4 | gcloud CLI | Fallback, if the CLI is installed and authenticated |
The defaults suit a reasonably quiet space. If Milo triggers on background noise or cuts you off mid-sentence, the endpointing variables below are the ones to reach for.
| Variable | Default | Effect |
|---|---|---|
MILO_VAD_RMS_THRESHOLD | 0.015 | Raise in a noisy room so ambient sound stops registering as speech |
MILO_VAD_SILENCE_DURATION | 0.7 | Seconds of quiet before the turn model is consulted |
MILO_VAD_MIN_SPEECH_DURATION | 0.4 | Utterances shorter than this are dropped as noise |
MILO_USE_SMART_TURN | true | Set false to use the plain silence timer alone |
MILO_SMART_TURN_THRESHOLD | 0.5 | Lower it to reply sooner, raise it to wait through longer pauses |
MILO_BARGE_IN | true | Set false if your hardware cannot record and play at once |
MILO_BARGE_IN_RMS_THRESHOLD | 0.03 | Raise it if Milo interrupts itself on its own playback |
MILO_USE_LOCAL_STT | false | Run faster-whisper on device instead of sending audio to cloud STT |
MILO_STT_MODEL | tiny | Local model size — tiny, base, or small |
One at a time. Do not run the dashboard tile and an SSH session together — both grab the microphone and speaker, and you get doubled audio.