CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01whyismydininghallonfire /orig-plus-asr-tamil-clean orig-plus-asr-tamil-clean Combined ASR dataset built from: albagon/til26-asr-split (orig rows) whyismydininghallonfire/asr-tamil-clean (asr_tamil_clean rows) Audio paths are namespaced under each split to avoid filename collisions: audio/orig/... audio/asr_tamil_clean/... Each row keeps key, audio, transcript, and language, with an added source_dataset field. Counts: train: 3595 orig + 891 asr_tamil_clean = 4486 validation: 899 orig + 224 asr_tamil_clean = 1123 audio1K<n<10K0 likes216 downloads4mo agoHugging Face02instinct-org /miscellaneous_yt_chunked_tokenizedgated miscellaneous_yt_chunked_48k_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes39 downloads4mo agoHugging Face03instinct-org /omni_chunked_tokenizedgated omni_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_tokenized.tabulartext-to-speech1K<n<10K0 likes20 downloads4mo agoHugging Face04instinct-org /cv_chunked_tokenizedgated cv_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.tabulartext-to-speech10K<n<100K0 likes19 downloads1mo agoHugging Face05instinct-org /yt4_chunked_tokenizedgated yt4_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes19 downloads4mo agoHugging Face06instinct-org /audiobook_chunked_tokenizedgated audiobook_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes19 downloads4mo agoHugging Face07instinct-org /espeech_podcasts_chunked_tokenizedgated espeech_podcasts_chunked_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes19 downloads4mo agoHugging Face08instinct-org /yt3_chunked_tokenizedgated yt3_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face09instinct-org /yt2_chunked_tokenizedgated yt2_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face10instinct-org /yt1_chunked_tokenizedgated yt1_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes16 downloads4mo agoHugging Face11instinct-org /yt_chunked_tokenizedgated yt_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes16 downloads4mo agoHugging Face12oridror /all-ceos-audio-caption-combined oridror/all-ceos-audio-caption-combined MYD multi-speaker Hebrew TTS training set for Orpheus-3B. Sourced from 6-ai/training/datasets/multimodal/audio_caption/ — Hebrew transcripts paired with synthetic edge-tts audio (Hila female, Avri male). License: Apache-2.0. Schema Each row in train.jsonl: { "audio_path": "<ceo>__<NNNNN>_<hash>.mp3", "transcript": "<utf-8 hebrew>", "speaker_id": "hila | avri", "ceo": "adel | rafael | antonio | lavi | noa"… See the full description on the dataset page: https://huggingface.co/datasets/oridror/all-ceos-audio-caption-combined.audion<1K0 likes15 downloads5mo agoHugging Face13instinct-org /default_voices_chunked_tokenizedgated default_voices_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face14instinct-org /tbp_chunked_tokenizedgated tbp_chunked_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face15instinct-org /audio_youtube_chunked_tokenizedgated audio_youtube_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face16instinct-org /education_chunked_tokenizedgated education_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes13 downloads4mo agoHugging Face17instinct-org /zy_chunked_tokenizedgated zy_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training shards… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes12 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.