CoolFace
20 results

audio-video

elmoghany /Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text Dataset Overview A collection of 27 domains (“topics”) and 3100 question-answer pair. Each topic comes with average 117 QA pairs.Every QA entry comes with: references: one or more source files the answer is extracted from time with each reference comes the starting and ending time the answer is extracted from the reference video_files: the video files where the answer can be found (future) video title & description from metadata.csv File structure You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.question-answering1K<n<10K2 likes2.2k downloads1y agoHugging Faceyatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads6mo agoHugging FaceQybera /pkl-video-audio0 likes740 downloads1y agoHugging Facepranitchawla /VCDB-Core-AudioVideo VCDB Core Audio-Video Retrieval This repository packages synchronized video and extracted audio from the 528-video core set of VCDB as a symmetric video+audio-to-video+audio retrieval task for MTEB/MOEB. The separate 100,000-video background collection is not included. Terms and provenance The source dataset is provided by Fudan University for research purposes only. The source authors and Fudan University make no warranties about the dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.audioother10K<n<100K0 likes253 downloads27d agoHugging Facemorinaoki /news-media-audio-video-curated News Media Audio Video Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for News Media work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material loader.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/morinaoki/news-media-audio-video-curated.0 likes218 downloads18d agoHugging FaceZeroAgency /shkolkovo-bobr.video-webinars-audio shkolkovo-bobr.video-webinars-audio Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars are parts of free online school exams training courses made by Shkolkovo. Language: Russian, includes some webinars on English Dataset structure: mp3 files in format ID.mp3, where ID is webinar ID. You can check original webinar with url like bobr.video/watch/ID. Some webinars may contain multiple speakers and music. txt file in format… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/shkolkovo-bobr.video-webinars-audio.audioautomatic-speech-recognition100K<n<1M6 likes214 downloads1y agoHugging Face