shangeth/ljspeech-mimi-codes
LJSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LJSpeech corpus — 13,100 English utterances from a single female speaker reading public-domain audiobook passages (~24 hours). This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech release; these codes are designed to be loaded alongside it for training Mimi-based speech models without paying the ~1 hour of GPU extraction cost. Schema One row per… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.
Add Links section (extraction code, Wren project, TTS models)
Update citation to Wren 2026
Add dataset card
Upload LJSpeech Mimi codes
Add files using upload-large-folder tool
Add dataset card
initial commit
