datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.Xerxes-Instruct-700K
Dataset Card for "Xerxes-Instruct-700K"
Description
Xerxes, named after a Persian King renowned for his wisdom and strategic prowess, is an amalgamation of four distinct datasets. This dataset has been curated to cater to the burgeoning needs of natural language processing tasks, particularly in the domain of conversation modeling and comprehension.
The dataset encompasses conversations sourced from a variety of sources, ranging from generative models to real-world… See the full description on the dataset page: https://huggingface.co/datasets/Instinct-AI/Xerxes-Instruct-700K.instinct-data
Instinct Dataset
This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub.
multilingual-tts-selected-speech
Multilingual TTS Selected Speech
This dataset contains text-and-audio pairs selected for TTS training. The initial
audio/uzbek split has 2,184 utterances. Other languages and additional
rows can be added as separate, append-only shards.
Each row retains the complete structured text specification alongside audio and
the OmniVoice fields utt, text, spk, and instruct. tts_tagged_text preserves
the local delivery and vocal-event controls separately from the clean text field.
For… See the full description on the dataset page: https://huggingface.co/datasets/instinct1912/multilingual-tts-selected-speech.tbp_chunked_speech_restorised
tbp_chunked_speech_restorised
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.instinct-data_nep_zeta-v0131cv_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/cv_chunked
Aligned dataset: instinct-org/cv_chunked_nfa_aligned
Rows: 71097 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.stt_dataset_uzbek
stt_dataset_uzbek
This is a gated Uzbek speech-to-text dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/stt_dataset_uzbek.cv_chunked
cv_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.yt4_unchunked
yt4_unchunked
This is a gated Russian unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_unchunked.InstinctSCOPE-RL-datadefault_voices_chunked_speech_restorised
default_voices_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.omni_chunked_speech_restorised
omni_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.yt3_chunked_speech_restorised
yt3_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised.cv_chunked_speech_restorised
cv_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.default_voices_unchunked
default_voices_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_unchunked.audiobook_chunked_speech_restorised
audiobook_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.yt2_chunked_speech_restorised
yt2_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised.pods_chunked
Pods Chunked VAD Transcribed
This public, manually gated dataset contains VAD-produced speech chunks paired with transcripts.
Important: this release is not Sidonized yet. It has only completed VAD/chunk preparation and transcription. It has not gone through the later restoration, alignment, profiling, or final filtering stages.
Contents
Rows: 17966
Audio duration: 299.42 hours
Language: Uzbek
Access: public repository with manual gated access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/pods_chunked.zy_unchunked
zy_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_unchunked.audiobook_unchunked
audiobook_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_unchunked.audio_youtube_unchunked
audio_youtube_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_unchunked.default_voices_chunked
default_voices_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked.tbp_chunked
tbp_chunked
This is a gated Russian chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked.yt1_chunked_speech_restorised
yt1_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised.audio_youtube_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audio_youtube_chunked
Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned
Rows: 559484 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.default_voices_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/default_voices_chunked
Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned
Rows: 134236 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.miscellaneous_yt_chunked_speech_restorised_nfa_aligned
miscellaneous_yt_chunked_speech_restorised_nfa_aligned
Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised.
Contents
Parquet shards: 130
Rows: 528,187
Approx hours: 863.88
Audio column: audio with embedded FLAC bytes
Transcript column: transcription
Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments
Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.miscellaneous_yt_unchunked
miscellaneous_yt_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_unchunked.omni_chunked
omni_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked.
