CoolFace
Apppublic

Enricx/reachy_mini_gpt_live

sourceHugging Faceupdated 14d agoView on Hugging Face
4likes
App README

Reachy Mini GPT-Live πŸŽ™οΈπŸ€–

 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚  REACHY MINI Β· GPT-LIVE            Montagehandleiding / Guide β”‚
 β”‚                                                              β”‚
 β”‚   πŸ€–  +  πŸŽ™οΈ  +  πŸ”‘  β†’  πŸ’¬  πŸ•Ί  πŸ“ž                           β”‚
 β”‚                                                              β”‚
 β”‚   ⏱ 10 min      🧰 geen gereedschap / no tools               β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ“¦ In de doos / In the box

πŸ€– Reachy Mini (Wireless of Lite) met draaiende daemonπŸ”‘ OpenAI API-key met tegoed (gpt-live-1)
🌐 Wifi voor de robotπŸ“± Een browser (telefoon of laptop)
☎️ Optioneel: Twilio-account + ngrok-account (voor bellen)

πŸ”§ Montage / Assembly

 β‘  INSTALLEREN                      β‘‘ STARTEN
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ Dashboard          β”‚             β”‚ Dashboard          β”‚
 β”‚ http://reachy-mini β”‚             β”‚  β–Ά Reachy Mini     β”‚
 β”‚        .local:8000 β”‚   ──────▢   β”‚    GPT-Live        β”‚
 β”‚                    β”‚             β”‚                    β”‚
 β”‚ App store β–Έ [–]    β”‚             β”‚   βš™οΈ (open UI)     β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
   Reachy Mini GPT-Live               poort 8042

 β‘’ API-KEY                          β‘£ PRATEN
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ Instellingen β–Έ     β”‚             β”‚                    β”‚
 β”‚  OpenAI            β”‚             β”‚   πŸ—£οΈ ─────▢ πŸ€–     β”‚
 β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚   ──────▢   β”‚        ◀───── πŸ”Š    β”‚
 β”‚  β”‚ sk-proj-…    β”‚  β”‚             β”‚                    β”‚
 β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚             β”‚  "Hoi Reachy!"     β”‚
 β”‚      [Opslaan]     β”‚             β”‚                    β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

 β‘€ PERSOONLIJKHEID (optioneel)      β‘₯ BELLEN (optioneel)
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ Profiel β–Έ          β”‚             β”‚ Instellingen β–Έ     β”‚
 β”‚  β—‹ default         β”‚             β”‚  Bellen Β· Twilio   β”‚
 β”‚  β—‹ mars_rover      β”‚             β”‚  Bellen Β· webhook  β”‚
 β”‚  β—‹ butler          β”‚   ──────▢   β”‚   & tunnel (ngrok) β”‚
 β”‚  β—‹ storyteller     β”‚             β”‚                    β”‚
 β”‚  [Activeren]       β”‚             β”‚ OpenAI β–Έ Webhooks  β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
StapNLEN
β‘ Open het dashboard van je robot en installeer Reachy Mini GPT-Live uit de app store (of via de URL van deze Space).Open your robot's dashboard and install Reachy Mini GPT-Live from the app store (or from this Space's URL).
β‘‘Start de app en klik op βš™οΈ: de pagina opent op poort 8042.Start the app and click βš™οΈ: the page opens on port 8042.
β‘’Instellingen β†’ OpenAI: plak je API-key, Opslaan. De key blijft op de robot.Settings β†’ OpenAI: paste your API key, Save. The key stays on the robot.
β‘£Klik Start gesprek (of zet automatisch starten aan) en praat. Vraag om een dansje, een emotie, wat hij ziet, het weer, de tijd.Click Start conversation (or enable autostart) and talk. Ask for a dance, an emotion, what it sees, the weather, the time.
β‘€Kies een persoonlijkheid, of maak er zelf een met eigen stem, begroeting en tools.Pick a personality, or create your own with its own voice, greeting and tools.
β‘₯Voor bellen: Twilio-gegevens, ngrok-authtoken + vast domein, OpenAI Project ID en webhook-secret invullen. Zeg dan "bel …".For phone calls: enter Twilio details, ngrok authtoken + static domain, OpenAI Project ID and webhook secret. Then say "call …".
 ⚠️  Reachy praat alleen met tegoed op je OpenAI-account.     ⚠️  Stop = sessie netjes afsluiten (usage definitief).
 πŸ’‘  Tik op een antenne: robot wordt wakker en de app start.  πŸ’‘  Alles werkt ook op je telefoon (zelfde URL).

Talk to your Reachy Mini. The app opens a full-duplex OpenAI GPT-Live voice session over WebRTC and wires it to the robot's own microphone array and speaker. While you talk, Reachy looks at you (face tracking), turns toward your voice (microphone direction of arrival), wobbles its head in sync with its own voice and moves its antennas.

The GPT-Live backend can call robot tools: dance, play an emotion, move or sweep the head, follow your face, look through the camera, remember and forget facts about you, tell the time and the weather, search the web, and go to sleep. Personality profiles (voice, greeting, instructions, enabled tools) are selectable and editable in the web UI, and long-term memory persists on the robot between conversations.

A small web page (port 8042) shows the live transcript (including tool calls), a start/stop button, the profiles, the memory and the settings. The OpenAI API key stays on the machine that runs the app; the browser never sees it.

Snel starten (NL)

  1. 1.Installeer de app via het Reachy Mini dashboard (of pip install op een Lite, zie onder).
  2. 2.Start de app in het dashboard en open de app-pagina (βš™οΈ, poort 8042).
  3. 3.Plak je OpenAI API-key bij Instellingen en klik Opslaan.
  4. 4.Klik Start gesprek (of zet automatisch starten aan) en praat tegen Reachy.
  5. 5.Stop sluit de sessie netjes af: de app wacht op session.closed zodat het gebruik definitief is.

Requirements

  • β€”A Reachy Mini (Wireless or Lite) with a running daemon (reachy-mini β‰₯ 1.10).
  • β€”An OpenAI API key with access to gpt-live-1 (project-scoped key).
  • β€”Python β‰₯ 3.11. Extra Python dependencies: aiortc, av, aiohttp (wheels exist for Linux aarch64, so the Wireless robot installs them fine).

Install

From the dashboard (Wireless or Lite): open http://reachy-mini.local:8000 (Wireless) or http://localhost:8000 (Lite), find Reachy Mini GPT-Live in the app store and click Install, or install this Space by URL:

bash
curl -X POST http://<HOST>:8000/api/apps/install \
  -H "Content-Type: application/json" \
  -d '{"url": "https://huggingface.co/spaces/<user>/reachy_mini_gpt_live"}'

Manually on a Lite (daemon on your laptop, same Python env as reachy-mini):

bash
uv pip install git+https://huggingface.co/spaces/<user>/reachy_mini_gpt_live
# or, from a checkout:
uv pip install -e /path/to/reachy_mini_gpt_live

Manually on a Wireless robot without internet on the robot:

bash
scp -r reachy_mini_gpt_live pollen@reachy-mini.local:/tmp/
ssh pollen@reachy-mini.local "/venvs/apps_venv/bin/pip install /tmp/reachy_mini_gpt_live"

Configure the API key

Pick one (later entries win):

  1. 1.Settings page (recommended): start the app, open port 8042, paste the key. It is written to ~/.config/reachy_mini_gpt_live/.env (mode 600) on the machine running the app.
  2. 2.A .env file next to the installed package (see .env.example).
  3. 3.Environment variables of the daemon process (OPENAI_API_KEY=...).

All options are listed in .env.example. Everything except the key can also be changed from the settings page.

Run

  • β€”Dashboard: start Reachy Mini GPT-Live, then open its page (βš™οΈ) at http://reachy-mini.local:8042 (Wireless) or http://localhost:8042 (Lite).
  • β€”Terminal (daemon must be running): python -m reachy_mini_gpt_live.main
  • β€”Make it the default experience: reachy-mini-daemon --startup-app reachy_mini_gpt_live

With autostart on (default) the session starts as soon as the app starts, so you can simply launch the app from the dashboard and talk.

How it works

 Reachy microphone ──► RobotMicTrack ──► aiortc PeerConnection ──► OpenAI GPT-Live
 (16 kHz, SDK media)    (20 ms frames)        β”‚  β–²  oai-events data channel
                                               β–Ό  β”‚
 Reachy speaker ◄─── SpeakerSink ◄──── remote audio track
 (push_audio_sample,   (resample 16 kHz)
  head wobbler taps it)
                                        Bridge ──► WebSocket ──► browser UI (transcript, status)
                                          β”‚
                                        RobotBehavior (wake-up, face tracking, DoA look, antennas)

The GPT-Live contract as implemented (reachy_mini_gpt_live/live_client.py):

  • β€”The app (server side, holds the key) creates the session with POST /v1/live/sessions and body {"session": <session_config.json>, "transport": {"type": "webrtc", "sdp": <offer>}} using bearer authentication, then reads transport.sdp (answer) and the opaque session.id.
  • β€”Remote audio arrives on the WebRTC audio track; JSON events on the oai-events data channel.
  • β€”The client waits for session.started and never sends startup configuration again.
  • β€”session.input_transcript.delta / session.output_transcript.delta are shown with their exact delta text and grouped by start_ms / end_ms (see transcript.py), not by turns.
  • β€”Stop sends session.close and keeps audio and transport alive until session.closed (with usage.seconds and reason). A timeout or disconnect before that is reported as incomplete finalization in the UI and the log.

session_config.json is sent verbatim:

json
{
  "model": "gpt-live-1",
  "audio": { "output": { "voice": "marin" } },
  "delegation": {
    "type": "responses",
    "responses": {
      "parallel_tool_calls": false,
      "model": "gpt-5.6-terra",
      "reasoning": { "effort": "medium" },
      "tools": [ { "type": "web_search" } ]
    }
  }
}

At start the active profile, its tools and the remembered facts are merged into this base (instructions, voice, delegation.responses.instructions and delegation.responses.tools); session_config.json itself is never modified. The greeting is requested right after session.started with session.instructions.append.

Memory works two ways: the voice model delegates "remember this / call me X / I like …" to the backend (tool remember), and after every conversation the app distils stable facts from the transcript into memory (setting auto memory, on by default; "Extract facts" button in the Memory tab). Facts are injected into the next conversation.

Releasing the robot: while the app runs it holds the robot's app lock, so the dashboard and Reachy Mini Control show the power button disabled. Use the header button Sleep & stop app (or say "ga slapen"): the session closes, the robot goes to its sleep pose and the app stops. Because the app is the robot's startup app, it starts again on the next wake-up or antenna touch.

Tools

Tools are declared as function tools of the Responses delegation (delegation.responses.tools, next to web_search). When the backend calls one, the app executes it on the robot and returns the result with response.item.create + response.create; the Live model then tells you what happened.

ToolWhat it does
dance / stop_dancePlay a move from the dances library (random if unspecified) / stop
play_emotion / stop_emotionPlay a move + sound from the emotions library / stop
move_headLook left, right, up, down or front
sweep_lookSweep the head from left to right
head_trackingFace tracking on/off
cameraTake a picture and describe it (OpenAI vision model, default: the delegation model)
go_to_sleepSay goodbye, sleep pose, close the session and stop the app
remember / forgetLong-term facts about you, stored in memory.json next to the config
get_timeDate and time (robot's timezone or an IANA zone)
get_weatherCurrent weather and today's forecast (Open-Meteo, no key needed)
phone_call / end_phone_callCall a contact or number with a goal; hang up
phone_scenarioCall using a saved scenario (goal template with placeholders)
answer_phone_questionPass your decision to the assistant on the phone
add_contact / list_contactsManage the phone book

Which tools a session gets is decided by the active profile's default_tools and by availability (camera enabled, move libraries loaded). Both move libraries are Hugging Face datasets and are downloaded once in the background when the app starts.

Profiles

A profile is one profile.md (TOML metadata between +++ lines, then the instructions as Markdown), the same format as Pollen's conversation app:

markdown
+++
schema_version = 1
voice = "cedar"
greeting = "Greet the user in one sentence, in character."
default_tools = ["dance", "play_emotion", "camera", "remember", "forget"]
+++

### IDENTITY
You are Reachy Mini, ...

Built-in profiles: default, mars_rover, butler, storyteller. Your own profiles are saved under ~/.config/reachy_mini_gpt_live/profiles/<name>/profile.md on the machine running the app (create, edit and activate them in the UI). The profile's instructions become the Live instructions; its voice sets audio.output.voice; the backend prompt (delegation.responses.instructions) gets the tool guidance, a summary of the persona and the remembered facts.

Phone calls (Twilio β†’ OpenAI SIP)

Say "bel Pizza Mama in Dronten en reserveer een tafel voor twee om zeven uur" and Reachy's assistant makes the call: it introduces itself as your digital assistant (an AI), pursues the goal, checks with you when a real decision is needed ("een moment, ik leg het even voor", with 2–4 options you answer by voice or with a button in the UI), hangs up, and reports back: agreed / agreed with a change / not possible, plus exceptions.

How it works (outbound calls must be provider-owned, per the GPT-Live SIP guide):

  1. 1.phone_call asks Twilio to dial the number. When it is answered, the called party joins a Twilio conference and the app places a second Twilio leg to sip:{project}@sip.api.openai.com;transport=tls;secure=true?X-Reachy-Call={id} that joins the same conference. With Call route set to direct the app uses the older <Dial><Sip> bridge, without keys.
  2. 2.OpenAI posts live.transport.incoming to https://<public-url>/api/phone/webhook (Standard-Webhooks signature verified); the app accepts the session with call-specific instructions and a small Responses delegation with two tools: ask_owner and hang_up.
  3. 3.The sideband WebSocket (/live/sessions/{id}/attach) streams the transcript; owner questions are relayed to the robot conversation (session.commentary.append) and answered with the answer_phone_question tool or POST /api/phone/answer.
  4. 4.After hang-up the transcript is summarised into {status, summary, exceptions} and spoken by Reachy.

Call scenarios: reusable goals with {placeholders} (table reservation, appointment, insurance claim follow-up, order status, complaint, and a few tongue-in-cheek ones such as negotiating an extra pay step with your manager), each in Dutch, English, German, French and Italian. Pick one in the Calls tab, choose the language, fill in the fields, call; or say "call the hairdresser with the appointment scenario" and Reachy asks for the missing details. The assistant then speaks that language on the call (introduction, "one moment, let me check", goodbye). Your own scenarios are saved on the robot (scenarios.json).

Languages: the web UI itself is available in the same five languages (switch in the header; defaults to your browser language).

Keypad and phone menus (DTMF): in real calls GPT-Live said menu choices out loud instead of pressing keys, so Twilio plays the tones. To press keys the app briefly redirects the called party's leg to <Play digits> and straight back into the conference, while the assistant stays connected. Keys can come from three places:

  • β€”Keypad in the Calls tab, as soon as the call is answered. You can also type 0–9, * and # on your keyboard.
  • β€”The call assistant's backend tool press_keys. The assistant is told never to say a menu choice out loud and to press the key that fits the goal.
  • β€”Menu keys on a scenario or a call, pressed right after answering. WWW1 waits three seconds, then presses 1.

Why a call ended: every call records who ended it (you, the assistant, the maximum duration, the other party, OpenAI or the connection to the assistant), OpenAI's close reason and a timeline of the call legs, keys and tool calls. Open Timeline under a call in the Calls tab.

Recording and live listening: with record calls on, Twilio forks the call audio to the app (<Start><Stream> over the tunnel, independent of the SIP leg). The app writes a stereo WAV (left = the person called, right = the assistant) to recordings/ next to the config, shows a player per call in the UI, and streams the mixed audio live to your browser ("🎧 live meeluisteren"). Tell people you record if your local law requires it.

Sharing and cleaning up recordings: every recording has Download, Share and Delete buttons. Share creates a secret link (valid 7 days, revoke it any time with Stop sharing) to a small page with the audio player, the result and the transcript. With the tunnel running the link works for anyone you send it to (copy, WhatsApp or e-mail); without it, only inside your own network. Delete removes the audio, the transcript and any share links. To save space automatically, set Delete recordings after (days): older calls are then removed after each call, at app start and as soon as you save the setting. The UI shows how much space the recordings take. On a free ngrok account, people who open a share link first see ngrok's warning page ("You are about to visit…") and click Visit Site.

What the tunnel exposes: only the OpenAI webhook, Twilio's media stream and share links (/share/…). The web UI, its API and live listening answer 404 through the public URL and stay reachable on your own network only (http://reachy-mini.local:8042).

Setup:

  1. 1.Twilio: a (non-trial) account with a voice-capable number. Enter Account SID, Auth Token and the number in the app settings.
  2. 2.Public URL for the webhook: the robot is behind your router, so the app opens an ngrok tunnel itself (through pyngrok, no manual install on the robot). Create a free ngrok account, copy your authtoken and claim your free static domain (YOUR-NAME.ngrok-free.app), and enter both in the settings. The tunnel starts with the app and the public URL / webhook URL appear automatically. Make the app the robot's startup app (dashboard, or PUT /api/apps/startup-app) so it, and the tunnel, come back after a reboot or an antenna touch.
  3. 3.OpenAI dashboard (same project as the API key): copy the Project ID (Settings β†’ General) and create a webhook (Settings β†’ Webhooks) for the event live.transport.incoming pointing at https://YOUR-NAME.ngrok-free.app/api/phone/webhook; paste the webhook secret (whsec_…) in the settings.
  4. 4.Set your name (used in the introduction), the call language and voice.
  5. 5.Add contacts (name + number) in the UI, or let the assistant look businesses up with web search. Test with the Bel nu form.

Costs: Twilio per-minute rates for the outbound leg plus the SIP leg, and GPT-Live voice time for the call session. The assistant always discloses that it is an AI; keep it that way.

Settings

SettingEnv varDefaultMeaning
AutostartGPT_LIVE_AUTOSTARTtrueStart a session when the app starts
Face trackingGPT_LIVE_HEAD_TRACKINGtrueDaemon-side head tracking of the closest face
Look at voiceGPT_LIVE_DOA_LOOKtrueTurn toward the speaker when no face is tracked
AntennasGPT_LIVE_ANTENNAStrueAntenna animation (listening / speaking / idle)
ProfileGPT_LIVE_PROFILEdefaultActive personality profile
GreetingGPT_LIVE_GREETINGtrueAsk the model to greet you right after session.started
Camera toolGPT_LIVE_CAMERAtrueAllow the camera tool
Vision modelGPT_LIVE_VISION_MODEL(delegation model)Model used to describe camera pictures
Preload movesGPT_LIVE_PRELOAD_MOVEStrueDownload the move libraries at app start
TwilioTWILIO_ACCOUNT_SID, TWILIO_AUTH_TOKEN, TWILIO_FROM_NUMBEROutbound calling
OpenAI SIPOPENAI_PROJECT_ID, OPENAI_WEBHOOK_SECRET, GPT_LIVE_PUBLIC_URLWebhook + SIP target
OwnerGPT_LIVE_OWNER_NAME, GPT_LIVE_CALL_LANGUAGE, GPT_LIVE_CALL_VOICE, GPT_LIVE_CALL_MAX_MINnl, marin, 6Call persona
Call routeGPT_LIVE_CALL_MODEconferenceconference: keypad and menu keys via Twilio; direct: plain <Dial><Sip> bridge
SIP host/schemeGPT_LIVE_SIP_HOST, GPT_LIVE_SIP_SCHEMEsip.api.openai.com, sipsip-eu.api.openai.com for EU residency; sips for SRTP
ngrokNGROK_AUTHTOKEN, NGROK_DOMAINIn-app tunnel for the webhook (sets the public URL)
Record callsGPT_LIVE_RECORD_CALLStrueStereo WAV recording + live listening via a Twilio media stream
Keep recordingsGPT_LIVE_KEEP_RECORDINGS_DAYS0Delete recordings older than N days automatically (0 = keep)
Speaker gainGPT_LIVE_OUTPUT_GAIN1.6Software gain on the robot speaker (0.2–4, soft-clipped)
Auto memoryGPT_LIVE_AUTO_MEMORYtrueDistil facts from each conversation into memory
Silence timeoutGPT_LIVE_IDLE_TIMEOUT_MIN10Close the session after N minutes without transcript activity (0 = never)
Finalize timeoutGPT_LIVE_FINALIZE_TIMEOUT_S10How long to wait for session.closed
API base URLOPENAI_BASE_URLhttps://api.openai.com/v1Override for proxies / tests

Web API (port 8042)

GET /api/status, GET /api/config, POST /api/config, POST /api/start, POST /api/stop, POST /api/mute {"muted": true}, GET /api/transcript, GET /api/profiles, POST /api/profiles, POST /api/profiles/select, DELETE /api/profiles/{name}, GET /api/tools, POST /api/moves/reload, GET /api/memory, DELETE /api/memory[/{id}], GET /api/phone, POST /api/phone/call, POST /api/phone/hangup, POST /api/phone/answer, POST /api/phone/keys {"keys": "1#"}, POST /api/phone/webhook (OpenAI), GET /api/phone/recordings, GET /api/phone/recordings/{id}.wav, DELETE /api/phone/recordings/{id}, POST /api/phone/recordings/{id}/share?days=7, DELETE /api/phone/shares/{token}, the public share page GET /share/{token} (+ .wav), GET/POST/DELETE /api/contacts, and a WebSocket at /ws streaming status, transcript.delta, tool.call, tool.result, delegation, event and finalization messages.

Tests

The test-suite runs without a robot and without OpenAI: a fake GPT-Live server (tests/fake_live_server.py, aiortc + aiohttp) answers the SDP offer, echoes the microphone audio back, emits scripted transcript events and honours session.close. Fake robot media generates a tone and records what would be played on the speaker.

bash
uv pip install -e ".[dev]"
python -m pytest

Troubleshooting

  • β€”"No OPENAI_API_KEY configured" – open the app page (port 8042) and paste the key.
  • β€”HTTP 401/403 from /v1/live/sessions – wrong key, or the project has no GPT-Live access.
  • β€”HTTP 500 "Internal Server Error" from /v1/live/sessions – seen when the OpenAI account has no credits left (a plain Responses call then reports credit_balance_exhausted). Add credits at https://platform.openai.com/settings/organization/billing/ and try again. python -m tests.probe_live tests the endpoint with several config/SDP variants and prints the OpenAI request-id.
  • β€”"No session.started within 20s" – the WebRTC connection could not be established; check that the robot can reach the internet (UDP) and try again.
  • β€”Reachy hears itself – lower the speaker volume in the dashboard; the daemon's echo canceller and GPT-Live's turn handling cope with normal levels.
  • β€”Incomplete finalization – the session was closed but session.closed never arrived (timeout/disconnect). The conversation is over, but the final usage could not be confirmed.
  • β€”Tunnel "error" in the phone card – check the ngrok authtoken/domain; the first start downloads the ngrok binary on the robot, which needs internet access and a few seconds.
  • β€”Port 8042 already in use (e.g. Reachy Mini Control on a laptop) – start the daemon/app with GPT_LIVE_UI_PORT=8047 and open that port instead.
  • β€”Logs: Lite β†’ the terminal running the daemon; Wireless β†’ ssh pollen@reachy-mini.local "sudo journalctl -u reachy-mini-daemon -f".

Credits

Built with the Reachy Mini SDK by Pollen Robotics and the OpenAI GPT-Live API.