CoolFace
Apppublic

Jacid23/Lyon_chatbox

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes
App README

Lyon Chatbox

A modular voice conversation app for the Reachy Mini robot. Cheap to run, versatile, customizable, and fun.

ChatBox uses a cascade pipeline — ASR → LLM → TTS — where each stage is a swappable provider. Mix cloud APIs and local models, tweak the personality, add live reactions to keywords, and let the robot dance.

[image]

How it works

🎤 Microphone
    │
    ▼
┌─────────┐     ┌─────────┐     ┌─────────┐     🔊 Speaker
│   ASR   │ ──▶ │   LLM   │ ──▶ │   TTS   │ ──▶ 🤖 Robot moves
└─────────┘     └─────────┘     └─────────┘
 Speech to       Reasoning       Text to
 text            + tool calls    speech

Voice Activity Detection (VAD) segments microphone input. The ASR provider transcribes it, the LLM generates a response (with optional tool calls for head movements, dances, emotions, camera…), and the TTS provider speaks it back — all while the robot animates.

Installation

[!IMPORTANT] Before using this app, install Reachy Mini's SDK.

<details open> <summary><b>Using uv (recommended)</b></summary>

bash
# Create venv
uv venv --python python3.12 .venv
source .venv/bin/activate

# Install with cascade pipeline
uv sync --extra cascade

Install additional providers as needed:

bash
uv sync --extra cascade_parakeet_progressive  # Parakeet ASR (Apple Silicon)
uv sync --extra cascade_kokoro                # Kokoro TTS (local)
uv sync --extra cascade_gemini                # Gemini LLM
uv sync --extra cascade_elevenlabs            # ElevenLabs TTS
uv sync --extra cascade_deepgram              # Deepgram ASR
uv sync --extra cascade_parakeet              # Parakeet batch/RNNT ASR (Apple Silicon)
uv sync --extra cascade_voxtral_mlx           # Voxtral ASR (Apple Silicon)
uv sync --extra cascade_nemotron              # NeMo ASR (CUDA)
uv sync --extra cascade_gradium               # Gradium TTS (Python ≥3.12)
uv sync --extra cascade_all                   # All cascade providers

Vision and head-tracking extras:

bash
uv sync --extra yolo_vision          # YOLO head-tracking
uv sync --extra mediapipe_vision     # MediaPipe head-tracking
uv sync --extra local_vision         # Local VLM (SmolVLM2)
uv sync --extra all_vision           # All vision features

Combine extras freely:

bash
uv sync --extra cascade --extra cascade_kokoro --extra cascade_gemini --extra yolo_vision --group dev
Tip: Use uv sync --frozen to install from the lockfile without re-resolving.

</details>

<details> <summary><b>Using pip</b></summary>

bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[cascade]"

# Add providers
pip install -e ".[cascade_kokoro]"
pip install -e ".[cascade_gemini]"
# etc.

</details>

Optional extras reference

ExtraPurposeNotes
cascadeBase cascade pipeline (sounddevice, librosa, PyYAML)Required for cascade mode
cascade_parakeet_progressiveParakeet MLX progressive ASRApple Silicon only
cascade_kokoroKokoro local TTSAny PyTorch platform
cascade_geminiGoogle Gemini LLMRequires GEMINI_API_KEY
cascade_deepgramDeepgram streaming ASRRequires DEEPGRAM_API_KEY
cascade_elevenlabsElevenLabs TTSRequires ELEVENLABS_API_KEY
cascade_parakeetParakeet batch/RNNT ASRApple Silicon only
cascade_voxtral_mlxVoxtral Mini 4B multilingual ASRApple Silicon only
cascade_nemotronNeMo Parakeet/Nemotron ASRCUDA GPU required
cascade_silero_vadSilero VAD (torch-based)Optional VAD backend
cascade_gradiumGradium streaming TTSPython ≥3.12, requires GRADIUM_API_KEY
cascade_allAll cascade providersExcludes gradium (Python ≥3.12)
local_visionLocal VLM (SmolVLM2) via PyTorch/TransformersGPU recommended
yolo_visionYOLO head-trackingCPU or GPU
mediapipe_visionMediaPipe head-trackingCPU
all_visionAll vision extras

Configuration

Environment variables

Open the Gradio Settings panel to choose providers and save keys or local endpoint settings. Gradio cascade requires explicit ASR, LLM, and TTS choices because every runtime is different. The app writes those values to .env. You can also copy .env.example to .env and fill them manually:

If no .env is found near the working directory, the app falls back to ~/.config/settings/.env. Keys stored there survive app reinstalls and updates.

For preconfigured installs, copy config/examples/lyon_chatbox_cascade_setup.yaml to config/lyon_chatbox.yaml or lyon_chatbox.yaml. You can also point LYON_CHATBOX_CONFIG_FILE to any YAML file. That file chooses providers and can include non-secret endpoint/model settings while cascade.yaml remains the provider-definition catalog. API keys still belong in .env or real environment variables.

VariableUsed by
OPENAI_API_KEYWhisper ASR, OpenAI Realtime ASR, GPT LLMs, OpenAI TTS
GEMINI_API_KEYGemini LLMs
DEEPGRAM_API_KEYDeepgram ASR
ELEVENLABS_API_KEYElevenLabs TTS
GRADIUM_API_KEYGradium TTS
HF_HOMECache dir for Hugging Face models (default: ./cache)
HF_TOKENHugging Face access for gated models
LYON_CHATBOX_CONFIG_FILEOptional path to a YAML setup config
CASCADE_ASR_PROVIDERASR provider override
CASCADE_LLM_PROVIDERLLM provider override
CASCADE_TTS_PROVIDERTTS provider override
CASCADE_LLM_BASE_URLLocal OpenAI-compatible LLM endpoint
CASCADE_LLM_MODELLocal/OpenAI-compatible LLM model name
CASCADE_LLM_API_KEYOptional local LLM token/API key
CASCADE_ASR_MODELASR model override
CASCADE_TTS_BASE_URLLocal OpenAI-compatible TTS endpoint
CASCADE_TTS_MODELTTS model override
CASCADE_TTS_API_KEYOptional local TTS token/API key
CASCADE_TTS_VOICETTS voice override
LYON_CHATBOX_CUSTOM_PROFILESelect a profile (default: default)
LYON_CHATBOX_EXTERNAL_PROFILES_DIRECTORYPath to external profiles
LYON_CHATBOX_EXTERNAL_TOOLS_DIRECTORYPath to external tool modules
AUTOLOAD_EXTERNAL_TOOLSSet 1 to auto-load all external tools

cascade.yaml

The cascade.yaml file at the project root controls which providers are used and their settings. Structure:

yaml
asr:
  provider: parakeet_mlx_progressive   # Which ASR to use
  cloud_providers:
    whisper_openai:
      module: whisper_openai
      class: WhisperOpenAIASR
      streaming: false
      requires: [OPENAI_API_KEY]
      model: whisper-1
  local_providers:
    parakeet_mlx_progressive:
      module: parakeet_mlx_progressive
      class: ParakeetMLXProgressiveASR
      streaming: true
      hardware: apple_silicon
      # Provider-specific settings
      model: mlx-community/parakeet-tdt-0.6b-v3
      precision: float16

llm:
  provider: gemini-2.5-flash-lite      # Which LLM to use
  temperature: 1.0
  cloud_providers: { ... }
  local_providers: { ... }

tts:
  provider: kokoro                     # Which TTS to use
  trim_silence: true
  cloud_providers: { ... }
  local_providers: { ... }

Change providers from the Gradio Settings panel, cascade.yaml, .env, or the CLI with --asr-provider, --llm-provider, or --tts-provider.

ASR providers
ProviderLocationStreamingHardwareInstall extraAPI key
parakeet_mlx_progressiveLocalYesApple Siliconcascade_parakeet_progressive
voxtral_mlxLocalYesApple Siliconcascade_voxtral_mlx
parakeet_nemo_progressiveLocalYesCUDAcascade_nemotron
nemotronLocalYesCUDAcascade_nemotron
deepgramCloudYescascade_deepgramDEEPGRAM_API_KEY
openai_realtime_asrCloudYesOPENAI_API_KEY
whisper_openaiCloudNo (batch)OPENAI_API_KEY
LLM providers
ProviderLocationModelInstall extraAPI key
gemini-2.5-flash-liteCloudgemini-2.5-flash-litecascade_geminiGEMINI_API_KEY
gemini-3.1-flash-liteCloudgemini-3.1-flash-lite-previewcascade_geminiGEMINI_API_KEY
gpt-4o-miniCloudgpt-4o-miniOPENAI_API_KEY
gpt-5.2-chatCloudgpt-5.2-chat-latestOPENAI_API_KEY
local_openai_compatibleLocalUser-selectedOptional
TTS providers
ProviderLocationHardwareInstall extraAPI key
kokoroLocalAny (PyTorch)cascade_kokoro
local_openai_compatible_ttsLocal/networkOpenAI-compatible serverOptional
tts_openaiCloudOPENAI_API_KEY
elevenlabsCloudcascade_elevenlabsELEVENLABS_API_KEY
gradiumCloudcascade_gradiumGRADIUM_API_KEY

Running the app

[!TIP] Make sure the Reachy Mini daemon is running before launching. See Reachy Mini's SDK for setup.
bash
lyon-chatbox --gradio

CLI options

OptionDefaultDescription
--gradioFalseLaunch Gradio web UI at http://127.0.0.1:7860/. Without this, runs in console mode with VAD.
--asr-provider NAMEfrom yamlOverride ASR provider (e.g. deepgram, whisper_openai)
--llm-provider NAMEfrom yamlOverride LLM provider (e.g. gpt-4o-mini, gemini-2.5-flash-lite)
--tts-provider NAMEfrom yamlOverride TTS provider (e.g. kokoro, elevenlabs)
--head-tracker {yolo,mediapipe}NoneEnable head-tracking. Requires the matching vision extra.
--no-cameraFalseRun without camera.
--local-visionFalseUse local VLM (SmolVLM2) instead of cloud vision. Requires local_vision extra.
--autotest [FILE]Run automated testing with text utterances (default file: cascade/autotest.txt).
--realtimeFalseUse OpenAI realtime audio-to-audio API instead of cascade.
--robot-name NAMENoneConnect to a specific robot when multiple daemons run on the same subnet.
--debugFalseVerbose logging.

Examples

bash
# Gradio UI with YOLO head-tracking
lyon-chatbox --gradio --head-tracker yolo

# Console mode (no UI, VAD-based)
lyon-chatbox

# Override providers from CLI
lyon-chatbox --gradio --asr-provider deepgram --llm-provider gpt-4o-mini

# Audio-only (no camera)
lyon-chatbox --gradio --no-camera

# Automated testing
lyon-chatbox --autotest

Profiles

Profiles define the robot's personality: what it says, how it sounds, and which tools it can use.

Each profile is a folder under src/lyon_chatbox/profiles/<name>/ containing:

FileRequiredPurpose
instructions.txtYesSystem prompt — the robot's personality and behavior rules
tools.txtRecommendedEnabled tools, one per line. Falls back to default/tools.txt if missing.
voice.txtOptionalTTS voice name override (single line)
reactions.yamlOptionalLive reaction triggers (see Live reactions)
*.pyOptionalCustom tool implementations

Selecting a profile

  • Environment variable: LYON_CHATBOX_CUSTOM_PROFILE=pirate in .env
  • Gradio UI: Open the "Personality" accordion to switch profiles, edit instructions, or create new ones.

Template placeholders in instructions.txt

Reuse shared prompt snippets by referencing files under src/lyon_chatbox/prompts/:

[passion_for_lobster_jokes]
[identities/witty_identity]

Locked profile mode

Set LOCKED_PROFILE in src/lyon_chatbox/config.py to lock the app to a single profile. The Gradio UI shows "(locked)" and disables profile editing. Useful for dedicated app variants.

Live reactions

Reactions let the robot respond to keywords or entities in real time — while the user is still speaking — without waiting for the LLM.

Define them in profiles/<name>/reactions.yaml:

yaml
- name: music_excitement
  callback: excited_about_music
  trigger:
    words: [music, guitar, piano, drum, violin]

- name: food_reaction
  callback: react_to_food_entity
  trigger:
    entities: [food]
  repeatable: true

- name: groovy_dance
  callback: do_groovy_dance
  trigger:
    all:
      - words: [danc*]
      - words: [groov*]

- name: reachy_name
  callback: react_to_name
  trigger:
    words: [reachy, richie, reechy]
  params:
    emotion: helpful1

Trigger types

TypeDescriptionExample
wordsKeyword/glob match on transcript[music, danc*, "grand piano"]
entitiesNamed entity recognition via GLiNER[food, person, location]
allBoolean AND — all sub-triggers must matchSee groovy_dance above

Behavior

  • Reactions fire at most once per conversation turn by default.
  • Set repeatable: true to allow multiple triggers per turn (deduplicated by entity text for entity triggers).
  • params are passed as **kwargs to the callback.

Callback signature

Each callback value maps to a Python module in the profile folder exporting an async function:

python
async def my_callback(deps: ToolDependencies, match: TriggerMatch, **kwargs) -> None:
    ...

match.words contains matched keywords; match.entities contains matched entities (with .text, .label, .confidence).

LLM tools

Tools the LLM can call during a conversation:

ToolAction
speakSynthesize and play a speech segment (used internally by the pipeline)
danceQueue a dance from the dances library
stop_danceClear queued dances
play_emotionPlay a recorded emotion clip from pollen-robotics/reachy-mini-emotions-library
stop_emotionClear queued emotions
move_headMove head to a named position (left/right/up/down/front)
see_image_through_cameraCapture a camera frame and send it to the LLM for analysis
describe_camera_imageDescribe the current camera view
head_trackingEnable/disable head-tracking (requires --head-tracker)
task_statusCheck status of a background task
task_cancelCancel a running background task
do_nothingExplicitly remain idle

Profiles can also define custom tools (e.g. turn_left, turn_right, center_position in the default profile).

Advanced features

<details> <summary><b>External profiles and tools</b></summary>

Store profiles and tools outside the source tree:

text
external_content/
├── external_profiles/
│   └── my_profile/
│       ├── instructions.txt
│       ├── tools.txt
│       └── voice.txt
└── external_tools/
    └── my_custom_tool.py

Set in .env:

env
LYON_CHATBOX_CUSTOM_PROFILE=my_profile
LYON_CHATBOX_EXTERNAL_PROFILES_DIRECTORY=./external_content/external_profiles
LYON_CHATBOX_EXTERNAL_TOOLS_DIRECTORY=./external_content/external_tools
  • Default mode: tools.txt must list every tool explicitly. Names resolve against built-in tools first, then external tools.
  • Auto-load mode (AUTOLOAD_EXTERNAL_TOOLS=1): all *.py modules in the external tools directory are loaded automatically.

</details>

<details> <summary><b>Multiple robots on the same subnet</b></summary>

bash
lyon-chatbox --robot-name <name>

<name> must match the daemon's --robot-name value.

</details>

<details> <summary><b>Autotest mode</b></summary>

Run the full pipeline with synthetic text utterances instead of a microphone:

bash
lyon-chatbox --autotest
lyon-chatbox --autotest my_test_script.txt

Each line in the test file is treated as a user utterance. Useful for end-to-end testing without audio hardware.

</details>

<details> <summary><b>OpenAI Realtime mode (legacy)</b></summary>

The original audio-to-audio mode using OpenAI's realtime API is still available:

bash
lyon-chatbox --realtime --gradio

This bypasses the cascade pipeline entirely. Requires OPENAI_API_KEY.

</details>

Contributing

We welcome bug fixes, features, profiles, and documentation improvements. Please review our contribution guide for branch conventions, quality checks, and PR workflow.

Quick start:

  • Fork and clone the repo
  • Follow the installation steps (include the dev dependency group)
  • Run contributor checks listed in CONTRIBUTING.md

License

Apache 2.0