datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DL3DV-Evaluation
DL3DV Testing Split Download Instructions
This repo contains all 55 scenes for evaluation. Note: it is an independent dataset, and none of its scenes overlap with those in DL3DV-10K. Have a galance on the preview page: https://dl3dv-10k.github.io/DL3DV-Testing-Split-Preview/.
Download
As the whole benchmark dataset is ~500G, a python script to download and untar files.
Environment Setup
The download script relies on huggingface hub, tqdm. You can download by… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Evaluation.SyncWorld-Evaluation
SyncWorld-Evaluation
Evaluation sets for yyuncong/SyncWorld, a robot world
model that predicts future video from past frames plus end-effector actions.
Paper: SyncWorld: Visual Calibration Enables World Models as Zero-Shot SimulatorsProject page: https://umass-embodied-agi.github.io/SyncWorld/
set
source
tasks
episodes
camera views
evaluation_maniskill
ManiSkill
5
50
base_camera_rgb, render_camera_rgb
evaluation_libero
LIBERO
5
50
agentview_rgb, side_rgb
Each… See the full description on the dataset page: https://huggingface.co/datasets/yyuncong/SyncWorld-Evaluation.bovine-embryo-video-evaluation
BovEmbryo: Bovine Embryo Time-Lapse Video Evaluation Dataset
Dataset Summary
BovEmbryo is a bovine embryo time-lapse video dataset for evaluating full-development biological video understanding models. The dataset contains 344 annotated videos and supports three evaluation settings: six-class IVF/SCNT outcome classification, three-class outcome-only classification, and IVF-to-SCNT / SCNT-to-IVF embryo-origin generalization.
Each video is an exported time-lapse imaging… See the full description on the dataset page: https://huggingface.co/datasets/embryo-video-eval/bovine-embryo-video-evaluation.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.evaluationVideo-Evaluationface-anonymization-public-evaluation-v58
Origin Data Lab — Face Anonymization Public Evaluation V58
Privacy processing for real-world video datasets
Automated face anonymization combined with targeted Human QA for autonomous driving, robotics, computer vision, urban mobility, and AI data teams.
This repository presents a public engineering evaluation of Origin Data Lab's face anonymization pipeline in dense, low-light urban traffic conditions.
Evaluate Your Own Video
Need privacy processing for traffic… See the full description on the dataset page: https://huggingface.co/datasets/origindatalab/face-anonymization-public-evaluation-v58.wenavigate-ppo-evaluation-v2evaluation-dataset
DeepSafe Evaluation Dataset
Evaluation set for DeepSafe,
a deepfake detection benchmark.
Tiers
Tier
Samples
Generators
Size
Use
master_eval_small/
198
116
1.7 GB
smoke test, under 2 min
master_eval/
15,454
411
10 GB
the standard benchmark
master_eval_full/
45,954
411
25 GB
complete set
Medium tier composition: 9,954 image, 3,500 audio, 2,000 video.
from huggingface_hub import snapshot_download
snapshot_download("deepsafe/evaluation-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/deepsafe/evaluation-dataset.speed-evaluation-kittivla-evaluationrelease-dataseteval_my_smolvla_evaluationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 14000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hangVLA/eval_my_smolvla_evaluation.eval_my_smolvla_evaluation2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 6,
"total_frames": 7357,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hangVLA/eval_my_smolvla_evaluation2.Evaluation_Videos_Lehome_ChallengeEvaluation Result
Train and Evaluate, pt 1
https://hackmd.io/@NYTCEE/rJjOk_9c-g
Train and Evaluate, pt 2
https://hackmd.io/@NYTCEE/Bks7KxioWl
vla-evaluation-v3egxo-household-egocentric-video-evaluation
EGXO Household Egocentric Video Dataset
EGXO maintains a continuously growing first-party household egocentric video catalogue. This repository documents the current commercial gold-standard 111-video, 10-hour evaluation release across 49 task families; it does not represent the size of the current catalogue.
Collection provenance
This release was collected through licensed GIG Rewards collection programs operated with telco partners. It is first-party inventory… See the full description on the dataset page: https://huggingface.co/datasets/egxodata/egxo-household-egocentric-video-evaluation.challenge-dataset
Challenge Dataset
Download
Run this once to fetch the parquet files and videos, and save them in a load_from_disk-compatible format:
from huggingface_hub import snapshot_download
from datasets import load_dataset
LOCAL_DIR = "challenge-dataset"
# 1. Download everything (parquet and video folders) into one local directory
snapshot_download(
repo_id="analogy-evaluation/challenge-dataset",
repo_type="dataset",
local_dir=LOCAL_DIR,
)
# 2. Load the splits from… See the full description on the dataset page: https://huggingface.co/datasets/analogy-evaluation/challenge-dataset.WorldHOI-evaluation-dataVLA_finetune_evaluationevaluation_outputsevaluation1eval_my_smolvla_evaluationevaluation5evaluation_dpdmd_dmd_oursvla-evaluation-v4evaluationEvaluation-fktevaluation_videosevaluation2
