CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instinct-org /multilingual-tts-voice-datasetgated Multilingual TTS Voice Dataset Multilingual speech and structured voice-control data for text-to-speech research and training. The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually. Configurations audio: utterances with embedded audio. prompt_specs: structured text and delivery specifications. clone_pairs: same-speaker reference and target pairs. voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.audiotext-to-speech100K<n<1M0 likes325 downloads1d agoHugging Face02Instinct-AI /Xerxes-Instruct-700K Dataset Card for "Xerxes-Instruct-700K" Description Xerxes, named after a Persian King renowned for his wisdom and strategic prowess, is an amalgamation of four distinct datasets. This dataset has been curated to cater to the burgeoning needs of natural language processing tasks, particularly in the domain of conversation modeling and comprehension. The dataset encompasses conversations sourced from a variety of sources, ranging from generative models to real-world… See the full description on the dataset page: https://huggingface.co/datasets/Instinct-AI/Xerxes-Instruct-700K.texttext-classification100K<n<1M10 likes132 downloads2y agoHugging Face03continuedev /instinct-data Instinct Dataset This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub. text1K<n<10K32 likes112 downloads1y agoHugging Face04instinct1912 /multilingual-tts-selected-speechgated Multilingual TTS Selected Speech This dataset contains text-and-audio pairs selected for TTS training. The initial audio/uzbek split has 2,184 utterances. Other languages and additional rows can be added as separate, append-only shards. Each row retains the complete structured text specification alongside audio and the OmniVoice fields utt, text, spk, and instruct. tts_tagged_text preserves the local delivery and vocal-event controls separately from the clean text field. For… See the full description on the dataset page: https://huggingface.co/datasets/instinct1912/multilingual-tts-selected-speech.audiotext-to-speech10K<n<100K0 likes29 downloads23h agoHugging Face05instinct-org /tbp_chunked_speech_restorisedgated tbp_chunked_speech_restorised This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes15 downloads4mo agoHugging Face06vnixxa31 /instinct-data_nep_zeta-v0131text1K<n<10K1 likes10 downloads6mo agoHugging Face07instinct-org /cv_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/cv_chunked Aligned dataset: instinct-org/cv_chunked_nfa_aligned Rows: 71097 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.textautomatic-speech-recognition10K<n<100K0 likes10 downloads1mo agoHugging Face08instinct-org /stt_dataset_uzbekgated stt_dataset_uzbek This is a gated Uzbek speech-to-text dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/stt_dataset_uzbek.audioautomatic-speech-recognitionn<1K0 likes9 downloads4mo agoHugging Face09instinct-org /cv_chunkedgated cv_chunked This is a gated Uzbek chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads1mo agoHugging Face10instinct-org /yt4_unchunkedgated yt4_unchunked This is a gated Russian unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_unchunked.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads4mo agoHugging Face11sullivanUCSD /InstinctSCOPE-RL-datatext10K<n<100K0 likes8 downloads5mo agoHugging Face12instinct-org /default_voices_chunked_speech_restorisedgated default_voices_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes7 downloads4mo agoHugging Face13instinct-org /omni_chunked_speech_restorisedgated omni_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.audioautomatic-speech-recognition1K<n<10K0 likes7 downloads4mo agoHugging Face14instinct-org /yt3_chunked_speech_restorisedgated yt3_chunked_speech_restorised_48k This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes7 downloads4mo agoHugging Face15instinct-org /cv_chunked_speech_restorisedgated cv_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads1mo agoHugging Face16instinct-org /default_voices_unchunkedgated default_voices_unchunked This is a gated Uzbek unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_unchunked.audioautomatic-speech-recognition1K<n<10K0 likes6 downloads4mo agoHugging Face17instinct-org /audiobook_chunked_speech_restorisedgated audiobook_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.audioautomatic-speech-recognition1M<n<10M0 likes6 downloads4mo agoHugging Face18instinct-org /yt2_chunked_speech_restorisedgated yt2_chunked_speech_restorised_48k This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M1 likes6 downloads4mo agoHugging Face19instinct-org /pods_chunkedgated Pods Chunked VAD Transcribed This public, manually gated dataset contains VAD-produced speech chunks paired with transcripts. Important: this release is not Sidonized yet. It has only completed VAD/chunk preparation and transcription. It has not gone through the later restoration, alignment, profiling, or final filtering stages. Contents Rows: 17966 Audio duration: 299.42 hours Language: Uzbek Access: public repository with manual gated access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/pods_chunked.tabularautomatic-speech-recognition10K<n<100K0 likes6 downloads4mo agoHugging Face20instinct-org /zy_unchunkedgated zy_unchunked This is a gated Uzbek unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_unchunked.audioautomatic-speech-recognition10K<n<100K0 likes5 downloads4mo agoHugging Face21instinct-org /audiobook_unchunkedgated audiobook_unchunked This is a gated Uzbek unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_unchunked.tabularautomatic-speech-recognition1M<n<10M0 likes5 downloads4mo agoHugging Face22instinct-org /audio_youtube_unchunkedgated audio_youtube_unchunked This is a gated Uzbek unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_unchunked.audioautomatic-speech-recognition1K<n<10K0 likes5 downloads4mo agoHugging Face23instinct-org /default_voices_chunkedgated default_voices_chunked This is a gated Uzbek chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked.audioautomatic-speech-recognition100K<n<1M0 likes5 downloads4mo agoHugging Face24instinct-org /tbp_chunkedgated tbp_chunked This is a gated Russian chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked.audioautomatic-speech-recognition100K<n<1M0 likes5 downloads4mo agoHugging Face25instinct-org /yt1_chunked_speech_restorisedgated yt1_chunked_speech_restorised_48k This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes5 downloads4mo agoHugging Face26instinct-org /audio_youtube_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/audio_youtube_chunked Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned Rows: 559484 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes5 downloads4mo agoHugging Face27instinct-org /default_voices_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/default_voices_chunked Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned Rows: 134236 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes5 downloads4mo agoHugging Face28instinct-org /miscellaneous_yt_chunked_speech_restorised_nfa_alignedgated miscellaneous_yt_chunked_speech_restorised_nfa_aligned Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised. Contents Parquet shards: 130 Rows: 528,187 Approx hours: 863.88 Audio column: audio with embedded FLAC bytes Transcript column: transcription Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes5 downloads4mo agoHugging Face29instinct-org /miscellaneous_yt_unchunkedgated miscellaneous_yt_unchunked This is a gated Uzbek unchunked source speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_unchunked.audioautomatic-speech-recognition10K<n<100K0 likes4 downloads4mo agoHugging Face30instinct-org /omni_chunkedgated omni_chunked This is a gated Uzbek chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked.audioautomatic-speech-recognition1K<n<10K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.