CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01deepdml /cv22-neucodec Dataset Statistics The following table summarizes the number of examples for each config_name, the corresponding language, and each split. config_name language train_examples validation_examples test_examples other_examples af Afrikaans 139 125 131 306 am Amharic 523 248 252 579 ar Arabic 28,531 10,503 10,500 41,364 as Assamese 952 485 379 2,557 az Azerbaijani 157 78 95 529 be Belarusian 347,672 15,879 15,880 17,002 bg Bulgarian 4,952 2,932 3,354 1,787 bn… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/cv22-neucodec.text10M<n<100M0 likes552 downloads6mo agoHugging Face02malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes236 downloads11d agoHugging Face03deepdml /fleurs-neucodec Dataset Dataset Statistics This table shows the number of examples per language configuration and split. config_name train_examples validation_examples test_examples af_za 1.032 198 264 am_et 3.163 223 516 ar_eg 2.104 295 428 as_in 2.812 418 984 ast_es 2.511 398 946 az_az 2.665 400923 be_by 2.433 408 967 bg_bg 2.973 395 658 bn_in 3.006 402 920 bs_ba 3.091 400 925 ca_es 2.300 404 940 ceb_ph 3.261 225 541 ckb_iq 3.040 386 922 cmn_hans_cn… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/fleurs-neucodec.text100K<n<1M0 likes217 downloads6mo agoHugging Face04neuphonic /emilia-yodas-english-neucodecgated Dataset Card for NeuCodec Emilia-YODAS Dataset Summary The NeuCodec Emilia-YODAS dataset is an English-language dataset containing >30M audio samples (>78k hours), taken from the English-language subset of Emilia-YODAS and compressed with NeuCodec. Usage import torch from datasets import load_dataset from neucodec import NeuCodec # load dataset and model dataset = load_dataset("neuphonic/emilia-yodas-english-neucodec", split="train"… See the full description on the dataset page: https://huggingface.co/datasets/neuphonic/emilia-yodas-english-neucodec.tabular10M<n<100M17 likes195 downloads1y agoHugging Face05Vyvo-Research /emilia-yodas-en-neucodec10M<n<100M0 likes187 downloads10mo agoHugging Face06deepdml /cv17-neucodec Dataset Dataset Overview This dataset contains Common Voice speech data encoded into neural codec representations. Each sample includes: audio_path duration codes sentence language client_id The dataset is organized by language configuration and split into train, validation, and test sets when available. Dataset Statistics The following table summarizes the number of examples for each config_name and split. Dataset Statistics The following… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/cv17-neucodec.text1M<n<10M0 likes186 downloads6mo agoHugging Face07Steveeeeeeen /yodas-granary-it-neucodec-10s-20s Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-10s-20s.tabular100K<n<1M0 likes186 downloads4mo agoHugging Face08Steveeeeeeen /yodas-granary-it-neucodec-150k Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-150k.tabular100K<n<1M0 likes94 downloads5mo agoHugging Face09Steveeeeeeen /yodas-granary-it-neucodec-300k-5s30s Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast", "asr" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-300k-5s30s.tabular100K<n<1M0 likes90 downloads4mo agoHugging Face10cheryltian /emovoice-neucodectext10K<n<100K2 likes20 downloads10mo agoHugging Face11TeeZee /common_voice_17-pl-speakers-v2.0-neucodecaudio10K<n<100K0 likes15 downloads11mo agoHugging Face12ik /asante-twi-neucodec-encodedtext10K<n<100K0 likes15 downloads6mo agoHugging Face13arch-stanton-1 /common-voice-urdu-male-neucodectext1K<n<10K0 likes15 downloads2mo agoHugging Face14NsuMILab /neucodec_slr_susttext10K<n<100K0 likes13 downloads7mo agoHugging Face15rendchevi /emilia-yodas-english-neucodec-VJKL-250ktabular100K<n<1M0 likes12 downloads5mo agoHugging Face16rendchevi /emilia-yodas-english-neucodec-VJKL-250k-preptabular10K<n<100K0 likes12 downloads5mo agoHugging Face17jy1095 /fluers-token-neucodec0 likes10 downloads3mo agoHugging Face18arch-stanton-1 /agri-male-16kHz-neucodectextn<1K0 likes10 downloads2mo agoHugging Face19TeeZee /nEMO-speakers-v2.0-neucodecaudio1K<n<10K0 likes9 downloads11mo agoHugging Face20TeeZee /fleurs-pl-v1.0-neucodecaudio1K<n<10K0 likes9 downloads11mo agoHugging Face21arch-stanton-1 /UrduTTSDataset-16khz-neucodectext1K<n<10K0 likes8 downloads2mo agoHugging Face22BarryFutureman /maya-distill-data-neucodectext10K<n<100K0 likes7 downloads10mo agoHugging Face23harikc456 /waxal-lug-neucodectext1K<n<10K0 likes7 downloads7mo agoHugging Face24mimba /plt-neucodecgatedtext100K<n<1M0 likes7 downloads3mo agoHugging Face25BarryFutureman /EmoDB-neucodectext10K<n<100K0 likes6 downloads10mo agoHugging Face26Steveeeeeeen /granary-it-voxpopuli-neucodec Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from nvidia/Granary / it_voxpopuli and uploaded as resumable Parquet shards. { "source_dataset": "nvidia/Granary", "source_config": "it_voxpopuli", "source_split": "asr", "validation_source_split": "asr", "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config", "source_split"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/granary-it-voxpopuli-neucodec.tabularn<1K0 likes6 downloads5mo agoHugging Face27mimba /nnh-neucodecgatedtext1K<n<10K0 likes6 downloads3mo agoHugging Face28mimba /plt-clone-neucodecgatedtext10K<n<100K0 likes6 downloads3mo agoHugging Face29rendchevi /emilia-yodas-english-neucodec-VFLTYtabular100K<n<1M0 likes5 downloads5mo agoHugging Face30rendchevi /emilia-yodas-english-neucodec-VJKL-25k-preptabular10K<n<100K0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.