CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes236 downloads11d agoHugging Face02neuphonic /emilia-yodas-english-neucodecgated Dataset Card for NeuCodec Emilia-YODAS Dataset Summary The NeuCodec Emilia-YODAS dataset is an English-language dataset containing >30M audio samples (>78k hours), taken from the English-language subset of Emilia-YODAS and compressed with NeuCodec. Usage import torch from datasets import load_dataset from neucodec import NeuCodec # load dataset and model dataset = load_dataset("neuphonic/emilia-yodas-english-neucodec", split="train"… See the full description on the dataset page: https://huggingface.co/datasets/neuphonic/emilia-yodas-english-neucodec.tabular10M<n<100M17 likes195 downloads1y agoHugging Face03Steveeeeeeen /yodas-granary-it-neucodec-10s-20s Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-10s-20s.tabular100K<n<1M0 likes186 downloads4mo agoHugging Face04Steveeeeeeen /yodas-granary-it-neucodec-150k Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-150k.tabular100K<n<1M0 likes94 downloads5mo agoHugging Face05Steveeeeeeen /yodas-granary-it-neucodec-300k-5s30s Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast", "asr" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-300k-5s30s.tabular100K<n<1M0 likes90 downloads4mo agoHugging Face06rendchevi /emilia-yodas-english-neucodec-VJKL-250ktabular100K<n<1M0 likes12 downloads5mo agoHugging Face07rendchevi /emilia-yodas-english-neucodec-VJKL-250k-preptabular10K<n<100K0 likes12 downloads5mo agoHugging Face08Steveeeeeeen /granary-it-voxpopuli-neucodec Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from nvidia/Granary / it_voxpopuli and uploaded as resumable Parquet shards. { "source_dataset": "nvidia/Granary", "source_config": "it_voxpopuli", "source_split": "asr", "validation_source_split": "asr", "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config", "source_split"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/granary-it-voxpopuli-neucodec.tabularn<1K0 likes6 downloads5mo agoHugging Face09rendchevi /emilia-yodas-english-neucodec-VFLTYtabular100K<n<1M0 likes5 downloads6mo agoHugging Face10rendchevi /emilia-yodas-english-neucodec-VJKL-25k-preptabular10K<n<100K0 likes5 downloads5mo agoHugging Face11Steveeeeeeen /cml-tts-italian-neucodec CML TTS Italian NeuCodec Encoded Italian CML TTS dataset with NeuCodec speech tokens and speaker metadata. { "source_encoded_dataset": "Steveeeeeeen/cml-tts-italian-neucodec", "speaker_metadata_source": "ylacombe/cml-tts/italian", "columns": [ "text", "codes", "speaker_id", "cml_split", "cml_row_idx", "duration", "num_words" ], "summary": { "splits": { "train": { "rows": 35337, "missing": 0, "speakers": 59… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/cml-tts-italian-neucodec.tabular10K<n<100K0 likes4 downloads5mo agoHugging Face12TeeZee /emilia-yodas-english-neucodec-2000tabular1K<n<10K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.