datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video-scissors-sessions
Coding agent session traces for kaofelix/video-scissors-sessions
This dataset contains redacted coding agent session traces collected while working on git@github.com:kaofelix/video-scissors.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/kaofelix/video-scissors-sessions.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.omniVDCBench
🎬 OmniVDCBench
A bilingual video-domain comprehension benchmark for evaluating fine-grained event understanding, cultural-scene reasoning, and multiple-choice video QA
Dataset Card · Video Question Answering · Fine-grained Event Reasoning · Intangible Cultural Heritage · Chinese / English
Figure 1. Illustration of the OmniVDCBench evaluation setting. The example combines video evidence, multiple-choice questions, model predictions, and category-level… See the full description on the dataset page: https://huggingface.co/datasets/videobenchlab/omniVDCBench.VideoScienceBench
VideoScienceBench
A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon).
Dataset Summary
Attribute
Value
Examples
160
Domains
Physics, Chemistry
Format
JSONL (prompt + expected phenomenon + vid)
Data Creation Pipeline
Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample shows what a training-ready video understanding distillation dataset can look like.
Why this exists
Most teams evaluating outside data vendors want to know one thing first:
What does the delivered data actually look like?
This sample is designed to answer that question.
It demonstrates how raw video clips can be converted into structured, model-ready supervision for:
video understanding
multimodal SFT… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/video-understanding-distillation-sample.snfa-youtube-videodaten
SNFA YouTube-Videodaten
Ein strukturierter Datensatz mit veröffentlichten YouTube-Videos der SNF Academy und zugehörigen Inhalten aus den Bereichen Fitness, Ernährung, Coaching, Mindset, Personal Training und Ausbildung.
Datensatzübersicht
1'423 eindeutige Videos
1'423 eindeutige YouTube-Video-IDs
1'065 Videos mit Beschreibung
472'190 erfasste Views
Veröffentlichungszeitraum: 6. September 2013 bis 10. Mai 2026
Datenprüfung: 17. Juli 2026
Sprache: überwiegend… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-youtube-videodaten.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample demonstrates what a training-ready video understanding / multimodal distillation dataset can look like.
Intended purpose
This dataset is not a production corpus. It is a schema demonstration for potential partners evaluating SuperviseLab's delivery approach.
What it shows
clip-level metadata
short and long captions
OCR text
transcript
speaker attribution
structured JSON targets
distillation-ready… See the full description on the dataset page: https://huggingface.co/datasets/metavi/video-understanding-distillation-sample.VideoGameFRqwen3-omni-pairwise-video-train
Qwen3-Omni Pairwise Video Inference / Evaluation
Pairwise audio-video preference evaluation data for Qwen3-Omni models.
Each sample compares two generated videos (with audio) against a text caption and human/Gemini labels.
Source path on cluster: /inspire/hdd/project/autoregressive-video-generation/public/hym/data/final_train
Upload snapshot: 2026-06-12 10:46 UTC
Repository layout
Contents of final_infer are uploaded to the dataset repo root:
.cache/
ovi_davinci/… See the full description on the dataset page: https://huggingface.co/datasets/YinmingHuang/qwen3-omni-pairwise-video-train.
