CoolFace
Datasetpublic

shangeth/expresso-mimi-codes

Expresso — Mimi Codes (k = 32) Pre-extracted Kyutai Mimi tokens (all 32 codebooks) for both the read and conversational subsets of Expresso. Source audio + transcripts live in shangeth/expresso; this dataset publishes the discrete-token version for training Mimi-based speech models without re-extracting. ⚠️ License: CC-BY-NC-4.0 — non-commercial use only. Why Expresso for Wren? Expresso is the most directly relevant dataset for speech disentanglement research —… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso-mimi-codes.

sourceHugging Facecc-by-nc-4.0updated 5mo agoView on Hugging Face
1likes32downloads
settings

This repository belongs to shangeth on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameexpresso-mimi-codes
visibilitypublic
licencecc-by-nc-4.0
gatedno
ownershangeth
Account settings
shangeth/expresso-mimi-codes · CoolFace