datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.shkolkovo-bobr.video-webinars-audio
shkolkovo-bobr.video-webinars-audio
Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars are parts of free online school exams training courses made by Shkolkovo.
Language: Russian, includes some webinars on English
Dataset structure:
mp3 files in format ID.mp3, where ID is webinar ID. You can check original webinar with url like bobr.video/watch/ID. Some webinars may contain multiple speakers and music.
txt file in format… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/shkolkovo-bobr.video-webinars-audio.YouTube_Video_Transkriptleri_TR
Dataset Summary
This dataset consists of nearly 5 hours of video from over 40 Creative Commons-licensed videos on YouTube. The videos contain the voices of more than 100 different people. The audio files have been resampled to 16 kHz. The videos have been divided into chunks of up to 25 seconds. This dataset is intended for developing Turkish STT (Speech-to-Text) models.
Datasets Preparetion
The audio files and transcript data were scraped from YouTube. The scraped… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/YouTube_Video_Transkriptleri_TR.err-video-news-transcribed
Transcribed ERR Video News Dataset
This dataset contains transcriptions of video news stories from Estonian National Brroacasting (https://www.err.ee/). There are around 40K stories with a total duration of around 4000 hours.
Transcriptions are generated automatically using speech recognition (gemini-3-flash-preview). Contextual biasing was used to improve ASR quality, using the textual news story about the same topic.
The WER of the transcriptions is around 5% on the average.
The… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/err-video-news-transcribed.yubao_videos
YuBao: A New Chinese Dialect Speech Benchmark
Paper | Code
This repository contains the video metadata for the YuBao (語保) dataset, as presented in the paper "Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects".
YuBao is a comprehensive collection from the Chinese Language Resources Protection Project, featuring speech, dialect transcripts, phonetic (IPA) transcriptions, and Mandarin translations for parallel items (1,000 characters, 1,200 words, and 50 sentences)… See the full description on the dataset page: https://huggingface.co/datasets/kalbin/yubao_videos.zamai-pashto-video
ZamAI Pashto Video
ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation.
Dataset Summary
The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows.
Languages
Pashto
Dari
English
Modalities
Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.audio-video-conversation-4000h
Audio-Video Conversational Dataset
4,000 hours of synchronized speech and video of natural conversations: face movement, mouth motion, gestures, turn-taking, emotion, laughter, and interruptions across 20+ languages.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real samples.… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/audio-video-conversation-4000h.swahili-video-text
Swahili Video-Text Dataset
Automatically processed Swahili video clips and transcriptions.
audio-video-conversation-4000h
Audio-Video Conversational Dataset
4,000 hours of synchronized speech and video of natural conversations: face movement, mouth motion, gestures, turn-taking, emotion, laughter, and interruptions across 20+ languages.
This repository is a specification and preview listing. The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real samples.
Overview
The Audio-Video Conversational Dataset is a 4… See the full description on the dataset page: https://huggingface.co/datasets/DatoricAI/audio-video-conversation-4000h.video-dataset
📚 Russian Storytelling Video Dataset (700 participants)
This dataset contains full-body videos of 700 native Russian speakers engaged in unscripted storytelling. Participants freely tell personal stories, express a wide range of emotions, and naturally use facial expressions and hand gestures. Each video captures authentic human behavior in high resolution with high-quality audio.
📊 Sample
📺 Preview Video (10 Participants)
To get a quick impression of… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/video-dataset.videos
