datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.VCDB-Core-AudioVideo
VCDB Core Audio-Video Retrieval
This repository packages synchronized video and extracted audio from the
528-video core set of VCDB as a symmetric video+audio-to-video+audio
retrieval task for MTEB/MOEB. The separate 100,000-video background collection
is not included.
Terms and provenance
The source dataset is provided by Fudan University for research purposes
only. The source authors and Fudan University make no warranties about the
dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.shkolkovo-bobr.video-webinars-audio
shkolkovo-bobr.video-webinars-audio
Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars are parts of free online school exams training courses made by Shkolkovo.
Language: Russian, includes some webinars on English
Dataset structure:
mp3 files in format ID.mp3, where ID is webinar ID. You can check original webinar with url like bobr.video/watch/ID. Some webinars may contain multiple speakers and music.
txt file in format… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/shkolkovo-bobr.video-webinars-audio.mead_hdtf_400_merge_video_audio_frames_onlyvideo-dataset-audio_dataset
Video Dataset - audio_dataset
Dataset Description
This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Combined from tasks: task06, task07, task08
Dataset Structure
frames/ — extracted frames (first frame from each segment)
segments/ — video clips for each annotation interval
annotations/ — original JSON annotation
transcriptions/ — transcription files… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-audio_dataset.audio-video-conversation-4000h
Audio-Video Conversational Dataset
4,000 hours of synchronized speech and video of natural conversations: face movement, mouth motion, gestures, turn-taking, emotion, laughter, and interruptions across 20+ languages.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real samples.… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/audio-video-conversation-4000h.audio-video-conversation-4000h
Audio-Video Conversational Dataset
4,000 hours of synchronized speech and video of natural conversations: face movement, mouth motion, gestures, turn-taking, emotion, laughter, and interruptions across 20+ languages.
This repository is a specification and preview listing. The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real samples.
Overview
The Audio-Video Conversational Dataset is a 4… See the full description on the dataset page: https://huggingface.co/datasets/DatoricAI/audio-video-conversation-4000h.Video-Audio-MMEHDTF_audio_videomead_hdtf_400_merge_video_audio_preprocessvideos_for_audioVFHQ-Video-AudioThis dataset follows the same license as the original VFHQ dataset. Please ensure you have the necessary permissions to use VFHQ data.
The data is organized into the following folders:
data: The video and audio data.
Ditto_videos_hdtf_400_audio_60s_chunks_all_preprocess
