datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
instinct-stt-corruption-v1
Instinct STT Corruption v1
Fixed, reproducible offline corruption waveforms used for the robust-training stages of
instinct1912/instinct-stt-1.
Contents
467,862 paired clean/corrupted training rows.
800 hours of corrupted speech, balanced at 200 hours each for English, Russian,
Uzbek, and Kazakh.
234 deterministic waveform shards.
Hash-bound manifests, per-shard receipts, and the final completion receipt.
Canonical corruption topologies covering… See the full description on the dataset page: https://huggingface.co/datasets/instinct1912/instinct-stt-corruption-v1.tbp_chunked_speech_restorised
tbp_chunked_speech_restorised
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.speech-data-catalog
Instinct Speech Data Catalog
This is the machine-readable entry point for the instinct-org speech data estate. It contains
metadata only; source audio remains in its existing manually gated repositories.
Inventory
Data products: 113
Public/manual-gated data products: 109
Private/manual-gated exceptions: 4
Hub-reported storage: 7.81 TiB
Purpose
Purpose
Datasets
mixed
1
stt
56
tts
55
unknown
1
Lifecycle state… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/speech-data-catalog.stt_dataset_uzbek
stt_dataset_uzbek
This is a gated Uzbek speech-to-text dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/stt_dataset_uzbek.cv_chunked
cv_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.yt4_unchunked
yt4_unchunked
This is a gated Russian unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_unchunked.default_voices_chunked_speech_restorised
default_voices_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.omni_chunked_speech_restorised
omni_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.yt3_chunked_speech_restorised
yt3_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised.cv_chunked_speech_restorised
cv_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.default_voices_unchunked
default_voices_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_unchunked.default_voices_chunked_encoded
default_voices_chunked_encoded
This is a gated Uzbek encoded derived speech artifact from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_encoded.audiobook_chunked_speech_restorised
audiobook_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.yt2_chunked_speech_restorised
yt2_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised.pods
Pods
Uzbek audio dataset.
Each row contains:
audio: the audio file loaded by AudioFolder from file_name
source: the original audio filename
collection: the source folder name
language: uz
zy_unchunked
zy_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_unchunked.audiobook_unchunked
audiobook_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_unchunked.audio_youtube_unchunked
audio_youtube_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_unchunked.default_voices_chunked
default_voices_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked.tbp_chunked
tbp_chunked
This is a gated Russian chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked.yt1_chunked_speech_restorised
yt1_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised.uf_hd
UF_HD
Uzbek audio dataset staged from local audio folders.
Each row contains:
audio: the audio file loaded by AudioFolder from file_name
source: the original audio filename
collection: the source folder name
language: uz
Source folders:
kashtan films
Nevo Films
uzbek_films_hd
Total audio files: 2245
miscellaneous_yt_unchunked
miscellaneous_yt_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_unchunked.omni_chunked
omni_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked.yt_unchunked
yt_unchunked
This is a gated Russian unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_unchunked.yt2_unchunked
yt2_unchunked
This is a gated Russian unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_unchunked.audio_youtube_chunked_speech_restorised
audio_youtube_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_speech_restorised.zy_chunked_speech_restorised
zy_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_speech_restorised.yt4_chunked_speech_restorised
yt4_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised.education_chunked
education_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked.
