abrarfahim/moshi-tool-audio
Moshi Tool-Calling — Audio-Grounded Dataset Audio-grounded data teaching Moshi / PersonaPlex to emit tool-call special tokens in its inner monologue when it hears a request — and to stay quiet otherwise (listening/idle frames are trained to PAD). Each row is a code tensor codes[17, T] at 12.5 Hz: rows stream content 0 text monologue PAD while listening/idle, `< 1:9 Moshi audio silence 9:17 user audio the spoken question (edge-tts), Mimi-encoded mask=1 marks… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/moshi-tool-audio.
Upload README.md with huggingface_hub
Upload dataset
Upload dataset
Upload README.md with huggingface_hub
Upload dataset
Upload dataset
Upload tool_audio.pt with huggingface_hub
initial commit
