datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QAEgo4D-MC-testThis benchmark was collected by QAEgo4D and updated by GroundVQA.
We conducted some processing for the experiments presented in our paper ReKV.
OmniVideo-Test
OmniVideo-Test
Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos/: Raw video files.
test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options.
OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.Video_Reality_Test
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:
real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/kolerk/Video_Reality_Test.GEN3C-Testing-Example
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
CVPR 2025 (Highlight)
Xuanchi Ren*,
Tianchang Shen*
Jiahui Huang,
Huan Ling,
Yifan Lu,
Merlin Nimier-David,
Thomas Müller,
Alexander Keller,
Sanja Fidler,
Jun Gao
* indicates equal contribution
Paper, Project Page
Abstract: We present GEN3C, a generative video model with precise Camera Control and
temporal 3D Consistency. Prior video models already generate realistic videos,
but they tend to leverage… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GEN3C-Testing-Example.video-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pyVideo_Reality_Test
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/ziweix/Video_Reality_Test.Video_Reality_Test
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/Dii2/Video_Reality_Test.Video_Reality_Test
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:
real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/videoasmrbench/Video_Reality_Test.Test-Dataset
AXIS held-out 20 — a cross-embodiment few-shot adaptation benchmark
20 tasks never seen in pretraining, 20 demonstrations each, LoRA adaptation, rollout evaluation.
This release is the SPECIFICATION and the INDEX, not the demonstration data. It is published
first and on purpose: everything here is what you need to render the benchmark on your own
embodiment, and none of it depends on our video encoding being finished.
Status: the task list is CANDIDATES. The learnability gate… See the full description on the dataset page: https://huggingface.co/datasets/axisrobotics/Test-Dataset.test_pipelineEgoPlan_test
EgoPlan-Test
Videos: We trim videos in EgoPlan test set by the start_frame and end_frame
Text: we add "video_path" key to the original json file.
korean-sign-word-classifier-mediapipe-test-100
KSL Test-100 Evaluation Dataset
This dataset contains 100 isolated Korean sign-language word videos used for the Test-100 evaluation of the MediaPipe keypoint word classifier.
Contents
videos/: MP4 isolated Korean sign-language word clips.
metadata.csv: Video-level labels, model predictions, and evaluation fields.
evaluation/: Full evaluation reports and machine-readable result files.
Evaluation Summary
Model:… See the full description on the dataset page: https://huggingface.co/datasets/Seoyoung07/korean-sign-word-classifier-mediapipe-test-100.E2E_20K_test
E2E 20K Synthetic Direction
A large-scale synthetic video benchmark for evaluating VideoLLMs' directional reasoning.20,000 videos of colored geometric shapes moving in four cardinal directions, with automatically generated MCQ annotations.
Directions
Direction
Description
up
Object moves toward the top of the frame
down
Object moves toward the bottom of the frame
left
Object moves toward the left of the frame
right
Object moves toward the right of the… See the full description on the dataset page: https://huggingface.co/datasets/KHUjongseo/E2E_20K_test.laser-vibrations-test
Laser Vibrations
Dataset of laser speckle vibration recordings used to locate objects hidden inside a cardboard box.
A 10×10 grid of lasers shines on the side of a box containing an object; as loudspeakers excite the box,
the speckle patterns shift in proportion to the local surface vibration. Per-sample metadata is in
data/metadata.jsonl; full signal data and media files live in per-sample subdirectories.
Dataset Viewer Columns
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/eturok-weizmann/laser-vibrations-test.yb_testvdub-test-assetstest-query
