datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.Real-TurnTurk
Real-TurnTurk
English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.YouTube-Commons-nl-audio
YouTube Commons NL Audio
This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions,
all under a CC BY 4.0 license.
It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB.
Source
The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.shofo-talking-head-en
Shofo Talking Head Dataset (English)
A curated dataset of ~10,000 talking-head videos (~186 hours, mean ~67s/clip), filtered for clean
single-speaker framing and paired with time-aligned transcripts. Built and released by Shofo.
This dataset is short-form social video, designed to support modern avatar, lip-sync, dubbing, and TTS work.
Every clip is curated by a multi-stage pipeline (face/framing analysis, on-screen-text detection,
object-occlusion detection, voice/face matching… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-talking-head-en.TeleEgo-Source
TeleEgo-Source
Source Videos and Time-Aligned Transcripts for TeleEgo
Official source release forTeleEgo: Benchmarking Egocentric AI Assistants in the Wild
Overview
TeleEgo is a multimodal benchmark for evaluating egocentric AI assistants in realistic, long-duration settings. It contains recordings from five participants over three days and covers four broad themes: Work & Study, Lifestyle &… See the full description on the dataset page: https://huggingface.co/datasets/H-oliday/TeleEgo-Source.darija_yt_2026
darija_yt_2026
Partition upload generated automatically.
Namespace: ohsn
Repo: ohsn/darija_yt_2026
Video count: 3511
Duration hours: 1565.31
This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline.
Sudan-MM
Sudan-MM: A Multimodal Dataset of Sudanese Arabic
Sudan-MM is the first publicly available multimodal dataset for Sudanese Arabic (السودانية), a low-resource dialect with no prior paired image-caption, video-caption, or voice-caption data. It was produced through a competitive shared task held in 2025, where five teams collected and annotated media depicting everyday Sudanese life.
Each item in the dataset pairs a visual or video recording with:
a written caption in Modern Standard… See the full description on the dataset page: https://huggingface.co/datasets/IndabaXSudan/Sudan-MM.APAC-Egocentric-Residential-Voiceover
APAC Egocentric Residential (with Voiceover)
Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track.
This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video.
Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.ai-researcher-roadmap-media
AI Researcher Roadmap Media
Optional video and subtitle assets for the
AI Researcher Roadmap
application.
Repository layout
manifest.json: file sizes and SHA-256 checksums used by the application.
videos/<stem>.mp4: lecture video.
subs/<stem>.<language>.vtt: subtitle tracks.
subs/<stem>.asr.<language>.vtt: ASR-generated subtitle tracks.
The application downloads only the selected lecture and its subtitle tracks.
Files are cached locally and can be played offline… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/ai-researcher-roadmap-media.CABankSakura
CABank Japanese Sakura Corpus
Susanne Miyata
Department of Medical Sciences
Aichi Shukotoku University
smiyata@asu.aasa.ac.jp
website: https://ca.talkbank.org/access/Sakura.html
Important
This data set is a copy from the original one located at https://ca.talkbank.org/access/Sakura.html.
Details
Participants: 31
Type of Study: xxx
Location: Japan
Media type: audio
DOI: doi:10.21415/T5M90R
Citation information
Some citation here.
In… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/CABankSakura.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.aiza-malagasy-administrative-asr
AIZA — Malagasy/French Code-Switched Administrative Speech (Benchmark Set)
10 short, self-recorded, consented audio clips of Malagasy/French code-switched
questions about Malagasy administrative procedures (lost ID card, birth
certificate, passport, land title, etc.) — created as the benchmark audio for
AIZA, a submission to the
Sahara CodeSwitch Africa Main Challenge.
Why this dataset exists
Neither Sahara's published language list nor the official… See the full description on the dataset page: https://huggingface.co/datasets/Arozo/aiza-malagasy-administrative-asr.Indic-subtitler-audio_evals
Indic_audio_evals
As part of this project. We are evaluating our performance of various ASR models as well
in a benchmarking dataset, we have created in various languages. This benchmarking dataset
is more alligned to real-world use-cases rather than having any academic datasets.
About Dataset
Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals
This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.zamai-pashto-video
ZamAI Pashto Video
ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation.
Dataset Summary
The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows.
Languages
Pashto
Dari
English
Modalities
Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya.
The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts.
The transcripts are:
Generated by ASR models (for the purpose of benchmarking)
Manual transcripts
Time stamps
Manual translations
This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/AgricultureVideosTranscript.2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/alj68/2M-Flores-ASL.ibibio-efik-speech-corpus-sample
Scuba Voice Dataset: Ibibio & Efik Sample (1 hour)
Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria.
This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.indic-audio-dialog-sample
Dataset Card for Indic Dialog Sample Dataset
Dataset Details
Dataset Description
The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.afrivoice-swahili-agriculture-subset
Dataset Card for the image text and voice dataset
Dataset Description
Subset of Afrivoice dataset:
approx. 100 hours of train (sampled, stratified)
Full dev + test from original repo (DigitalUmuganda/Afrivoice_Swahili)
Audio: .webm
Includes images + transcriptions from original repo (DigitalUmuganda/Afrivoice_Swahili)
Includes metadata csv
License
CC-BY-4.0 (derived from DigitalUmuganda/Afrivoice_Swahili)
Notes
Train split sampled using… See the full description on the dataset page: https://huggingface.co/datasets/egirma/afrivoice-swahili-agriculture-subset.portuguese-speech-recognition-dataset
Portuguese Telephone Dialogues Dataset - 10 Hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/portuguese-speech-recognition-dataset.swahili-video-text
Swahili Video-Text Dataset
Automatically processed Swahili video clips and transcriptions.
AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya.
The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts.
The transcripts are:
Generated by ASR models (for the purpose of benchmarking)
Manual transcripts
Time stamps
Manual translations
This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.mosla
Overview
The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one
of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction,
and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human
annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.anime-2024
Anime Video Dataset 2024
Overview
This is an anime video dataset curated specifically for multimedia research purposes.
It is sourced from anime series released in 2024 and includes both a complete combined file and four seasonal subsets (Winter, Spring, Summer, and Fall).
Data Details
Video Stream
Codec: H.264 (High Profile)
Format: YUV 4:2:0 (Progressive)
Resolution: 640×360 (SAR 1:1, DAR 16:9)
Frame Rate: 23.98 FPS
Audio Stream
Codec: AAC (LC)
Sample… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/anime-2024.
