CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sirekist98 /tokenise_spanish_datasetDataset tokenizado para TTS en español. Incluye audio, transcripción, emoción, y códigos SNAC. audio100K<n<1M0 likes749 downloads1y agoHugging Face02omartariq612 /everyayah-with-tajweed-tokensaudio100K<n<1M2 likes214 downloads2y agoHugging Face03ittailup /spanish_tokenizedaudio10K<n<100K1 likes137 downloads2y agoHugging Face04ittailup /english_tokenizedaudio100K<n<1M0 likes136 downloads2y agoHugging Face05Alvor /audio-tokenizer-demo Audio Tokenizer Demo Dataset Audio samples processed by various neural audio tokenizers for comparison. Tokenizers Column Tokens/sec Codebooks cosyvoice2 25 1 glm4voice 12.5 1 mimoaudio 6.25 8 neucodec 50 1 wavtokenizer 40 1 xcodec2 50 1 unicodec 75 1 bicodec 50 1 (Semantic) + 1 (Global Fixed 35 tokens) flexicodec <100 1(FSQ)+7(RVQ) varstok <40 1 MOSS 400 32 H-Codec-1.5 <200 4(Acoustic) + 4(Semantic) audion<1K0 likes132 downloads7mo agoHugging Face06NGC404 /russian_sample-dataset-tokenised_Orpheus_TTSaudion<1K7 likes97 downloads1y agoHugging Face07Trelis /libritts-bpe-tokens libritts-bpe-tokens To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services. GPT-2 BPE tokens of LibriTTS-R text_normalized transcripts. Each utterance is terminated with the EOS token (50256). Tokens are in column token_ids (list[int]), vocab=50,257. Splits Mirrors the source LibriTTS-R splits (filtered by parler-tts; total ≈ 538 h): split utterances hours train.clean.100 ~32 k ~53 h train.clean.360 ~112 k ~218 h train.other.500… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-bpe-tokens.text100K<n<1M0 likes92 downloads4mo agoHugging Face08JerryAGENDD /JLSpeech_tokenizedThis dataset uses focal codec to tokenize audio in JL speech dataset. audio10K<n<100K1 likes78 downloads2y agoHugging Face09MihaiPopa-1 /common_voice_22_0_toki_pona_parquet Common Voice 22.0 - Toki Pona Subset! My own Parquet conversion of Toki Pona's subset of Fsicoli's reupload of Common Voice 22 so we don't have to downgrade to Datasets 3.6 anymore! Why? Because the original dataset required Hugging Face Datasets 3.6 or older because it has Python code and it's in TAR shards. This is in Parquet and works with any recent version of Hugging Face Datasets! Details Dataset Structure DatasetDict({… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/common_voice_22_0_toki_pona_parquet.audioautomatic-speech-recognition10K<n<100K0 likes67 downloads29d agoHugging Face10Trelis /libritts-snac-tokens libritts-snac-tokens To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services. LibriTTS-R encoded with hubertsiuzdak/snac_24khz (hierarchical RVQ, 3 levels at 12 / 24 / 48 fps, 4,096 entries each). Orpheus-style interleave per 1/12-sec audio frame: [L0[t], L1[2t], L1[2t+1], L2[4t], L2[4t+1], L2[4t+2], L2[4t+3]]. 7 tokens per audio frame, 84 fps flat. Offset vocab 12,288: L0 in [0, 4096), L1 in [4096, 8192), L2 in [8192, 12288). Decode with level = token //… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-snac-tokens.text100K<n<1M0 likes61 downloads4mo agoHugging Face11fguryel /va_tokenized Model Card for Model ID Model Details Model Description Developed by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Model type: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Finetuned from model [optional]: [More Information Needed] Model Sources [optional] Repository: [More Information Needed] Paper… See the full description on the dataset page: https://huggingface.co/datasets/fguryel/va_tokenized.audio100K<n<1M0 likes56 downloads1y agoHugging Face12omartariq612 /everyayah-mapped-to-tajweed-tokensaudio100K<n<1M0 likes53 downloads2y agoHugging Face13asahi417 /experiment-audio-tokenizeraudion<1K0 likes44 downloads2y agoHugging Face14Lwasinam /ljspeech-tokens-v2audio10K<n<100K0 likes42 downloads2y agoHugging Face15instinct-org /miscellaneous_yt_chunked_tokenizedgated miscellaneous_yt_chunked_48k_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes39 downloads4mo agoHugging Face16macabdul9 /librispeech-hubert-discrete-tokens Dataset Card for "librispeech-hubert-discrete-tokens" More Information needed audio10K<n<100K0 likes38 downloads3y agoHugging Face17Gong1212 /emilia-tokengated Emilia EN Pocket Mimi continuous latents This gated repository contains the English Emilia training data representation used by the LatentTTS experiments in this project. Audio was encoded offline with the continuous Gaussian Mimi speech VAE used by Pocket TTS. The files are intended to let an authorized researcher reproduce latent-domain training without encoding the source audio again. Access and licensing This is a derived representation of the official Emilia… See the full description on the dataset page: https://huggingface.co/datasets/Gong1212/emilia-token.audiotext-to-speech0 likes30 downloads17d agoHugging Face18macabdul9 /fleurs-hubert-discrete-tokens Dataset Card for "fleurs-hubert-discrete-tokens" More Information needed audio1K<n<10K0 likes29 downloads3y agoHugging Face19rocky730 /audio-tokenizer-demo Audio Tokenizer Demo Dataset Audio samples processed by various neural audio tokenizers for comparison. Tokenizers Column Tokens/sec Codebooks cosyvoice2 25 1 glm4voice 12.5 1 mimoaudio 6.25 8 neucodec 50 1 wavtokenizer 40 1 xcodec2 50 1 unicodec 75 1 bicodec 50 1 (Semantic) + 1 (Global Fixed 35 tokens) flexicodec <100 1(FSQ)+7(RVQ) varstok <40 1 MOSS 400 32 audion<1K0 likes27 downloads7mo agoHugging Face20Martingkc /dMel_tokenized_lj_speech_c80_sr16_hop400audio10K<n<100K0 likes26 downloads1y agoHugging Face21syarief-mulyadi /IWSE-InstructionBasedSpeechEdit-llasa_tokenizeaudion<1K0 likes22 downloads6mo agoHugging Face22LeeHarrold /musiccaps-mot-tokens MusicCaps Pre-Encoded Tokens for Mixture-of-Transformers (MoT) Dataset Description This dataset contains pre-encoded audio tokens from the MusicCaps dataset, processed through Meta's MusicGen EnCodec tokenizer for use in Mixture-of-Transformers (MoT) training. Dataset Summary 5,233 music clips encoded as discrete tokens 4 codebook layers from MusicGen's EnCodec ~500 tokens per 10-second clip Compressed from ~12GB audio to 82MB tokens Ready for multimodal… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/musiccaps-mot-tokens.tabulartext-to-audio1K<n<10K0 likes21 downloads10mo agoHugging Face23instinct-org /omni_chunked_tokenizedgated omni_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_tokenized.tabulartext-to-speech1K<n<10K0 likes20 downloads4mo agoHugging Face24milamarcheva /morphemically_tokenised_english_cdsaudion<1K0 likes19 downloads1y agoHugging Face25instinct-org /cv_chunked_tokenizedgated cv_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.tabulartext-to-speech10K<n<100K0 likes19 downloads1mo agoHugging Face26instinct-org /yt4_chunked_tokenizedgated yt4_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes19 downloads4mo agoHugging Face27instinct-org /audiobook_chunked_tokenizedgated audiobook_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes19 downloads4mo agoHugging Face28instinct-org /espeech_podcasts_chunked_tokenizedgated espeech_podcasts_chunked_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes19 downloads4mo agoHugging Face29instinct-org /yt3_chunked_tokenizedgated yt3_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face30instinct-org /yt2_chunked_tokenizedgated yt2_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.