CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instinct-org /multilingual-tts-voice-datasetgated Multilingual TTS Voice Dataset Multilingual speech and structured voice-control data for text-to-speech research and training. The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually. Configurations audio: utterances with embedded audio. prompt_specs: structured text and delivery specifications. clone_pairs: same-speaker reference and target pairs. voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.audiotext-to-speech100K<n<1M0 likes325 downloads14h agoHugging Face02Instinct-AI /Xerxes-Instruct-700K Dataset Card for "Xerxes-Instruct-700K" Description Xerxes, named after a Persian King renowned for his wisdom and strategic prowess, is an amalgamation of four distinct datasets. This dataset has been curated to cater to the burgeoning needs of natural language processing tasks, particularly in the domain of conversation modeling and comprehension. The dataset encompasses conversations sourced from a variety of sources, ranging from generative models to real-world… See the full description on the dataset page: https://huggingface.co/datasets/Instinct-AI/Xerxes-Instruct-700K.texttext-classification100K<n<1M10 likes132 downloads2y agoHugging Face03continuedev /instinct-data Instinct Dataset This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub. text1K<n<10K32 likes112 downloads1y agoHugging Face04instinct-org /miscellaneous_yt_chunked_tokenizedgated miscellaneous_yt_chunked_48k_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes39 downloads4mo agoHugging Face05instinct1912 /multilingual-tts-selected-speechgated Multilingual TTS Selected Speech This dataset contains text-and-audio pairs selected for TTS training. The initial audio/uzbek split has 2,184 utterances. Other languages and additional rows can be added as separate, append-only shards. Each row retains the complete structured text specification alongside audio and the OmniVoice fields utt, text, spk, and instruct. tts_tagged_text preserves the local delivery and vocal-event controls separately from the clean text field. For… See the full description on the dataset page: https://huggingface.co/datasets/instinct1912/multilingual-tts-selected-speech.audiotext-to-speech10K<n<100K0 likes29 downloads4h agoHugging Face06instinct-org /omni_chunked_tokenizedgated omni_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_tokenized.tabulartext-to-speech1K<n<10K0 likes20 downloads4mo agoHugging Face07instinct-org /cv_chunked_tokenizedgated cv_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.tabulartext-to-speech10K<n<100K0 likes19 downloads1mo agoHugging Face08instinct-org /yt4_chunked_tokenizedgated yt4_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes19 downloads4mo agoHugging Face09instinct-org /audiobook_chunked_tokenizedgated audiobook_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes19 downloads4mo agoHugging Face10instinct-org /espeech_podcasts_chunked_tokenizedgated espeech_podcasts_chunked_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes19 downloads4mo agoHugging Face11instinct-org /yt3_chunked_tokenizedgated yt3_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face12instinct-org /yt2_chunked_tokenizedgated yt2_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face13instinct-org /yt1_chunked_tokenizedgated yt1_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes16 downloads4mo agoHugging Face14instinct-org /yt_chunked_tokenizedgated yt_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes16 downloads4mo agoHugging Face15instinct-org /tbp_chunked_speech_restorisedgated tbp_chunked_speech_restorised This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes15 downloads4mo agoHugging Face16instinct-org /default_voices_chunked_tokenizedgated default_voices_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face17instinct-org /tbp_chunked_tokenizedgated tbp_chunked_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face18instinct-org /audio_youtube_chunked_tokenizedgated audio_youtube_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face19instinct-org /education_chunked_tokenizedgated education_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes13 downloads4mo agoHugging Face20instinct-org /zy_chunked_tokenizedgated zy_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training shards… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes12 downloads4mo agoHugging Face21vnixxa31 /instinct-data_nep_zeta-v0131text1K<n<10K1 likes10 downloads6mo agoHugging Face22instinct-org /cv_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/cv_chunked Aligned dataset: instinct-org/cv_chunked_nfa_aligned Rows: 71097 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.textautomatic-speech-recognition10K<n<100K0 likes10 downloads1mo agoHugging Face23instinct-org /stt_dataset_uzbekgated stt_dataset_uzbek This is a gated Uzbek speech-to-text dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/stt_dataset_uzbek.audioautomatic-speech-recognitionn<1K0 likes9 downloads4mo agoHugging Face24instinct-org /all15_speaker_deduped_tts_train_clone_pairs_rawgated All-15 Speaker-Deduped TTS Train Clone Pairs Raw This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline. Contents data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows. data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.tabulartext-to-speech10K<n<100K0 likes9 downloads4mo agoHugging Face25instinct-org /cv_chunkedgated cv_chunked This is a gated Uzbek chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads1mo agoHugging Face26instinct-org /yt4_unchunkedgated yt4_unchunked This is a gated Russian unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_unchunked.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads4mo agoHugging Face27sullivanUCSD /InstinctSCOPE-RL-datatext10K<n<100K0 likes8 downloads5mo agoHugging Face28instinct-org /default_voices_chunked_speech_restorisedgated default_voices_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes7 downloads4mo agoHugging Face29instinct-org /omni_chunked_speech_restorisedgated omni_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.audioautomatic-speech-recognition1K<n<10K0 likes7 downloads4mo agoHugging Face30instinct-org /yt3_chunked_speech_restorisedgated yt3_chunked_speech_restorised_48k This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.