datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moshi-on-policy-prompts-3kBased on https://huggingface.co/datasets/yhytoto12/behavior-sd
moshi-on-policy-dpo-margin3mwanamke_moshi
Swahili Moshi Fine-tuning Dataset (mwanamke)
Overview
This dataset prepares Swahili conversational audio for fine-tuning
Moshika (the female-voice
Moshi variant) using
kyutai-labs/moshi-finetune.
It builds on the stereo, speaker-separated audio chunks from
rlabz/qsuperposition_mwanamke
and adds the .jsonl index and per-file .json transcripts that
moshi-finetune requires for training.
Source Data
Origin: rlabz/qsuperposition_mwanamke — stereo… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/mwanamke_moshi.moshi-tool-audio
Moshi Tool-Calling — Audio-Grounded Dataset
Audio-grounded data teaching Moshi / PersonaPlex to emit tool-call special
tokens in its inner monologue when it hears a request — and to stay quiet
otherwise (listening/idle frames are trained to PAD).
Each row is a code tensor codes[17, T] at 12.5 Hz:
rows
stream
content
0
text monologue
PAD while listening/idle, `<
1:9
Moshi audio
silence
9:17
user audio
the spoken question (edge-tts), Mimi-encoded
mask=1 marks… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/moshi-tool-audio.moshi-paired-prefsmall-german-medical-dialogue-dataset-for-moshi
Small german dialogue dataset
This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office.
Dataset Details
Dataset Description
500 made up phonecalls that were first created with AI as text.
The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added.
The audio files are formatted like this:
Stereo with split channels:
Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.moshi_chunk3sunbird-moshi-development-eval-v1eval_moshi_group_1moshi-continuationmoshi-on-policy-dpo-v17-partialmoshi-on-policy-dpo-v20-kyutai-smokemoshi-on-policy-dpo-v20-kyutai-alignedmoshi_chunk23moshi-on-policy-dpo-9kmoshi_chunk10moshi-lt-data
Moshi-LT Training Data
Synthetic Lithuanian audio data for training Moshi voice agents.
Contents
Split
Files
Duration
Format
Monologues
14899
unknownh
Mono WAV 24kHz + JSON timestamps
Dialogues
1703
unknownh
Stereo WAV 24kHz + JSON timestamps
Total
16602
unknownh
Generation
TTS engine: Google Chirp 3: HD (28 Lithuanian voices)
Monologues: Wikipedia + CulturaX Lithuanian text
Dialogues: LLM-generated scripts (Gemini 2.0 Flash) —… See the full description on the dataset page: https://huggingface.co/datasets/isLucid/moshi-lt-data.moshi-on-policy-dpo-tts-v18-fullprompt-smokemoshi_chunk0moshi-on-policy-dpo-9k-v17moshi-on-policy-dpo-tts-v18moshi-on-policy-dpo-tts-v18-fullpromptmoshi-on-policy-dpo-v20-kyutai-aligned-smokemoshi_dialogue
