datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mwanamke_moshi
Swahili Moshi Fine-tuning Dataset (mwanamke)
Overview
This dataset prepares Swahili conversational audio for fine-tuning
Moshika (the female-voice
Moshi variant) using
kyutai-labs/moshi-finetune.
It builds on the stereo, speaker-separated audio chunks from
rlabz/qsuperposition_mwanamke
and adds the .jsonl index and per-file .json transcripts that
moshi-finetune requires for training.
Source Data
Origin: rlabz/qsuperposition_mwanamke — stereo… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/mwanamke_moshi.pre-moshi-madammoshi-on-policy-prompts-3kBased on https://huggingface.co/datasets/yhytoto12/behavior-sd
moshi-on-policy-dpo-margin3moshi-zh-domain_wavfiledaily-life-hacks-grocery-nutrition-per-dollar
Daily Life Hacks: Grocery Nutrition per Dollar
This dataset contains two public CSV files ranking grocery foods by nutrition per dollar:
fiber-per-dollar-2026.csv — fiber per dollar
protein-per-dollar-2026.csv — protein per dollar
Methodology and disclosure
Nutrition values come from USDA FoodData Central, and prices are US prices from July 2026. This is an independent study and is not endorsed by USDA.
Read the companion articles:
Cheapest high-fiber foods… See the full description on the dataset page: https://huggingface.co/datasets/moshiko123/daily-life-hacks-grocery-nutrition-per-dollar.xx_beat_arkit_moshi_2025_07_20_30fps_attngreenobe-bd-resultszeroeggs_moshi_2025_06_06_30fps_conv4moshi-tool-audio
Moshi Tool-Calling — Audio-Grounded Dataset
Audio-grounded data teaching Moshi / PersonaPlex to emit tool-call special
tokens in its inner monologue when it hears a request — and to stay quiet
otherwise (listening/idle frames are trained to PAD).
Each row is a code tensor codes[17, T] at 12.5 Hz:
rows
stream
content
0
text monologue
PAD while listening/idle, `<
1:9
Moshi audio
silence
9:17
user audio
the spoken question (edge-tts), Mimi-encoded
mask=1 marks… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/moshi-tool-audio.moshi-paired-prefmoshimoshi_ai_girls_zh_captionedsmall-german-medical-dialogue-dataset-for-moshi
Small german dialogue dataset
This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office.
Dataset Details
Dataset Description
500 made up phonecalls that were first created with AI as text.
The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added.
The audio files are formatted like this:
Stereo with split channels:
Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.moshi_chunk3sunbird-moshi-development-eval-v1eval_moshi_group_1zeroeggs_moshi_2025_05_29moshi-on-policy-dpo-v17-partialmoshi-continuationmoshi-on-policy-dpo-v20-kyutai-smokemoshi_tts_dataset_dummymoshi-on-policy-dpo-v20-kyutai-alignedmoshi_chunk23moshi-on-policy-dpo-9kmoshi-on-policy-dpo-tts-v18-fullprompt-smokezeroeggs_moshi_2025_05_26_validmoshi-lt-data
Moshi-LT Training Data
Synthetic Lithuanian audio data for training Moshi voice agents.
Contents
Split
Files
Duration
Format
Monologues
14899
unknownh
Mono WAV 24kHz + JSON timestamps
Dialogues
1703
unknownh
Stereo WAV 24kHz + JSON timestamps
Total
16602
unknownh
Generation
TTS engine: Google Chirp 3: HD (28 Lithuanian voices)
Monologues: Wikipedia + CulturaX Lithuanian text
Dialogues: LLM-generated scripts (Gemini 2.0 Flash) —… See the full description on the dataset page: https://huggingface.co/datasets/isLucid/moshi-lt-data.xx_moshi_2025_07_09_30fps_conv4moshi_chunk10moshimoshi_ai_videos
