CoolFace
Datasetpublic

shangeth/ljspeech-mimi-codes

LJSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LJSpeech corpus — 13,100 English utterances from a single female speaker reading public-domain audiobook passages (~24 hours). This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech release; these codes are designed to be loaded alongside it for training Mimi-based speech models without paying the ~1 hour of GPU extraction cost. Schema One row per… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
0likes71downloads
7 commits on main
59019155mo ago

Add Links section (extraction code, Wren project, TTS models)

shangeth
d44092c5mo ago

Update citation to Wren 2026

shangeth
04405bd5mo ago

Add dataset card

shangeth
5eb848d5mo ago

Upload LJSpeech Mimi codes

shangeth
01eecaf7mo ago

Add files using upload-large-folder tool

shangeth
c0b15d37mo ago

Add dataset card

shangeth
28910557mo ago

initial commit

shangeth