CoolFace
Datasetpublic

abrarfahim/moshi-tool-audio

Moshi Tool-Calling — Audio-Grounded Dataset Audio-grounded data teaching Moshi / PersonaPlex to emit tool-call special tokens in its inner monologue when it hears a request — and to stay quiet otherwise (listening/idle frames are trained to PAD). Each row is a code tensor codes[17, T] at 12.5 Hz: rows stream content 0 text monologue PAD while listening/idle, `< 1:9 Moshi audio silence 9:17 user audio the spoken question (edge-tts), Mimi-encoded mask=1 marks… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/moshi-tool-audio.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes31downloads
8 commits on main
93c6e733mo ago

Upload README.md with huggingface_hub

abrarfahim
a8c98bc3mo ago

Upload dataset

abrarfahim
b214e543mo ago

Upload dataset

abrarfahim
2710d5a3mo ago

Upload README.md with huggingface_hub

abrarfahim
e1397ed3mo ago

Upload dataset

abrarfahim
36596593mo ago

Upload dataset

abrarfahim
e7b101f3mo ago

Upload tool_audio.pt with huggingface_hub

abrarfahim
1dbe06f3mo ago

initial commit

abrarfahim