uzbek
Datasets
All datasets matching “uzbek”Uzbek-STT-Dataset-780h
Uzbek STT Dataset (~780 hours)
A large Uzbek speech-to-text dataset for training and fine-tuning automatic
speech recognition (ASR) models such as Whisper.
Dataset summary
Language
Uzbek (uz)
Examples
122,464
Total audio
~780 hours
Clip length
up to 30 seconds each
Columns
audio, transcription
Audio
embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded
Split
single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.uzbek-multi-speaker-35hnews_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.uzbek-speech-corpus
Uzbek Speech Corpus
Dataset Summary
The Uzbek speech corpus (USC) has been developed in collaboration between ISSAI and the Image and Speech Processing Laboratory in the Department of Computer Systems of the Tashkent University of Information Technologies. The USC comprises 958 different speakers with a total of 105 hours of transcribed audio recordings. To ensure high quality, the USC has been manually checked by native speakers. The USC is primarily designed for… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uzbek-speech-corpus.podcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/DavronSherbaev/uzbekvoice-filtered.
