datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_22_0_toki_pona_parquet
Common Voice 22.0 - Toki Pona Subset!
My own Parquet conversion of Toki Pona's subset of Fsicoli's reupload of Common Voice 22 so we don't have to downgrade to Datasets 3.6 anymore!
Why?
Because the original dataset required Hugging Face Datasets 3.6 or older because it has Python code and it's in TAR shards.
This is in Parquet and works with any recent version of Hugging Face Datasets!
Details
Dataset Structure
DatasetDict({… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/common_voice_22_0_toki_pona_parquet.emilia-token
Emilia EN Pocket Mimi continuous latents
This gated repository contains the English Emilia training data representation
used by the LatentTTS experiments in this project. Audio was encoded offline
with the continuous Gaussian Mimi speech VAE used by Pocket TTS. The files are
intended to let an authorized researcher reproduce latent-domain training
without encoding the source audio again.
Access and licensing
This is a derived representation of the
official Emilia… See the full description on the dataset page: https://huggingface.co/datasets/Gong1212/emilia-token.
