datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uzbek-STT-Dataset-780h
Uzbek STT Dataset (~780 hours)
A large Uzbek speech-to-text dataset for training and fine-tuning automatic
speech recognition (ASR) models such as Whisper.
Dataset summary
Language
Uzbek (uz)
Examples
122,464
Total audio
~780 hours
Clip length
up to 30 seconds each
Columns
audio, transcription
Audio
embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded
Split
single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.uzbek-speech-corpus
Uzbek Speech Corpus
Dataset Summary
The Uzbek speech corpus (USC) has been developed in collaboration between ISSAI and the Image and Speech Processing Laboratory in the Department of Computer Systems of the Tashkent University of Information Technologies. The USC comprises 958 different speakers with a total of 105 hours of transcribed audio recordings. To ensure high quality, the USC has been manually checked by native speakers. The USC is primarily designed for… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uzbek-speech-corpus.uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/DavronSherbaev/uzbekvoice-filtered.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.it_youtube_uzbek_speech_dataset
IT Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/islomov/it_youtube_uzbek_speech_dataset.podcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/ai4uz/uzbekvoice-filtered.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/news_youtube_uzbek_speech_dataset.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/news_youtube_uzbek_speech_dataset.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/news_youtube_uzbek_speech_dataset.uzbek-asr-curated-701h
Uzbek ASR Curated Dataset (701 hours)
A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation.
Dataset Description
Language
Uzbek (Latin script with okina ʻ)
Total utterances
337,920
Total duration
~701 hours
Audio format
16 kHz mono WAV (PCM_16)
Manifest format
NeMo JSONL
Splits
train (94%) / val (3%) / test (3%)
Splits
Split
Utterances
Hours
Train
317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.uzbek_speech_datauzbekvoice-podcast
UzbekVoice Podcast — O'zbek Podkast Ovoz Dataseti
Bitta YouTube podkast epizodidan tuzilgan kichik o'zbekcha nutq dataseti,
UzbekVoice
uslubida: qisqa audio bo'laklar mos matn (transkript) bilan, har biri
taxminan 10–20 soniyalik segmentlarga bo'lingan.
Manba
Video: Laganda semichka, Sotmayman! SEMICHKACHI BEKZOD | TEZKOR PODKAST
Kanal: QASHQADARYO_TEZKOR24
Havola: https://youtu.be/TsJdrXKYppg
Til: O'zbekcha (jonli/sheva nutqi, Qashqadaryo shevasi)
Ushbu dataset… See the full description on the dataset page: https://huggingface.co/datasets/Fdev071/uzbekvoice-podcast.uzbek-asr-train-manifests
Uzbek ASR Training Manifests
The exact training, validation and test splits behind
rustam1221/uzbek-asr-gigaam:
974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized,
and split by speaker.
No audio is copied. Each row is a pointer — a parquet file plus a row
index in the upstream dataset — and the training dataloader decodes the audio
when the batch is built. That keeps the whole corpus definition at 200 MB
instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.it_youtube_uzbek_speech_dataset
IT Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/it_youtube_uzbek_speech_dataset.uzbek_stt_datapodcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.it_youtube_uzbek_speech_dataset
IT Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/it_youtube_uzbek_speech_dataset.uzbekvoice-filtered2This is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/Jurabek/uzbekvoice-filtered2.podcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.stt_dataset_uzbek
stt_dataset_uzbek
This is a gated Uzbek speech-to-text dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/stt_dataset_uzbek.
