datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ljspeech-mimi-codes
LJSpeech — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens for the
LJSpeech corpus — 13,100 English utterances
from a single female speaker reading public-domain audiobook passages (~24 hours).
This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech
release; these codes are designed to be loaded alongside it for training Mimi-based speech
models without paying the ~1 hour of GPU extraction cost.
Schema
One row per utterance:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.librispeech-mimi-codes
LibriSpeech — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens for the
LibriSpeech corpus — multi-speaker English audiobook readings
from the LibriVox project.
This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech
mirrors (e.g. openslr/librispeech_asr);
these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models.
Schema
One row per utterance:
Column
Type
Notes
id
string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.mls-mimi-codes
Multilingual LibriSpeech (MLS) — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens
for Multilingual LibriSpeech —
LibriVox audiobooks in 7 non-English languages.
English is intentionally excluded. For English Mimi codes, use:
shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits)
shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native)
shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents)
shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.VoiceAssistant-400K-mimi
VoiceAssistant-400K with Mimi Tokens
This dataset is a processed version of gpt-omni/VoiceAssistant-400K with audio codec conversions.
Processing
Each sample has been processed to add:
answer_audio: Decoded audio waveform from SNAC tokens (24kHz Audio feature)
answer_mimi: Re-encoded audio using Kyutai's Mimi codec (32 codebooks)
Columns
Column
Type
Description
split_name
string
Original split name
index
int
Sample index
round
int
Conversation… See the full description on the dataset page: https://huggingface.co/datasets/Muvels/VoiceAssistant-400K-mimi.expresso-mimi-codes
Expresso — Mimi Codes (k = 32)
Pre-extracted Kyutai Mimi tokens (all 32 codebooks) for both the read and conversational subsets of Expresso. Source audio + transcripts live in shangeth/expresso; this dataset publishes the discrete-token version for training Mimi-based speech models without re-extracting.
⚠️ License: CC-BY-NC-4.0 — non-commercial use only.
Why Expresso for Wren?
Expresso is the most directly relevant dataset for speech disentanglement research — the… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso-mimi-codes.mimir-auto-clean
Mimir WIC Audio 24k
Audio chunks derived from public YouTube videos on https://www.youtube.com/@mimir-winter-is-coming.
Transcripts come from YouTube auto-generated captions only; no Whisper or other re-transcription is used.
Each parquet shard corresponds to one source video and embeds 24 kHz mono WAV audio bytes.
mimir-auto
Mimir WIC Audio 24k
Audio chunks derived from public YouTube videos on https://www.youtube.com/@mimir-winter-is-coming.
Transcripts come from YouTube auto-generated captions only; no Whisper or other re-transcription is used.
Each parquet shard corresponds to one source video and embeds 24 kHz mono WAV audio bytes.
mimir-auto-clean-deduple
Mimir WIC Audio 24k
Audio chunks derived from public YouTube videos on https://www.youtube.com/@mimir-winter-is-coming.
Transcripts come from YouTube auto-generated captions only; no Whisper or other re-transcription is used.
Each parquet shard corresponds to one source video and embeds 24 kHz mono WAV audio bytes.
