datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DL3DV-Evaluation
DL3DV Testing Split Download Instructions
This repo contains all 55 scenes for evaluation. Note: it is an independent dataset, and none of its scenes overlap with those in DL3DV-10K. Have a galance on the preview page: https://dl3dv-10k.github.io/DL3DV-Testing-Split-Preview/.
Download
As the whole benchmark dataset is ~500G, a python script to download and untar files.
Environment Setup
The download script relies on huggingface hub, tqdm. You can download by… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Evaluation.longvideo_eval_videos
Long-RL: Scaling RL to Long Sequences (Evaluation Dataset - for research only)
Data Distribution
We strategically construct a high-quality dataset with CoT annotations for long video reasoning, named LongVideo-Reason. Leveraging a powerful VLM (NVILA-8B) and a leading open-source reasoning LLM, we develop a dataset comprising 52K high-quality Question-Reasoning-Answer pairs for long videos. We use 18K high-quality samples for Long-CoT-SFT to initialize… See the full description on the dataset page: https://huggingface.co/datasets/LongVideo-Reason/longvideo_eval_videos.WSYue-ASR-eval
WSYue-ASR-eval: Cantonese ASR Benchmark
To address the unique linguistic characteristics of Cantonese in speech recognition, we propose WSYue-ASR-eval, a benchmark specifically designed for evaluating Cantonese ASR systems. It is tailored to assess model performance across diverse lengths, domains, and linguistic phenomena of Cantonese speech.
The test set annotations are provided by Beijing AISHELL Technology Co., Ltd.
Key features:
Annotated through multiple rounds of manual… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/WSYue-ASR-eval.Semi-Truths-Evalset
Semi-Truths: The Evaluation Sample
Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation?
To address these questions, we introduce Semi-Truths, featuring 27,600 real images, 245,300 masks, and 850,200 AI-augmented images featuring varying degrees of targeted and localized edits, created using diverse augmentation methods… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths-Evalset.BridgeVLA_COLOSSEUM_EVAL_DATAarxiv: https://arxiv.org/abs/2506.07961
PixArt-Eval-30KDA-2-Evaluation
DA2: Depth Anything in Any Direction
DA2 predicts dense, scale-invariant distance from a single 360° panorama in an end-to-end manner, with remarkable geometric fidelity and strong zero-shot generalization.
🎮 Usage
Please see here.
🎓 Citation
If you find these datasets useful, please consider citing 🌹:
@article{li2025depth,
title={DA$^{2}$: Depth Anything in Any Direction},
author={Li, Haodong and Zheng, Wangguangdong and He, Jing and Liu, Yuhao and… See the full description on the dataset page: https://huggingface.co/datasets/haodongli/DA-2-Evaluation.Lyra-EvalNurec_eval
Nurec_eval — NRE Closed-Loop Eval Set with HD Map
59 OOD-longtail scenes for AlpaSim closed-loop evaluation, exported from NVIDIA NRE 26.02 and augmented with HD-map artifacts (lane graph, road boundaries, crosswalks) so models can be evaluated with full map context — not just geometry.
What's inside each pai_<uuid>.usdz
File
Purpose
default.usda
USDZ entry point
checkpoint.ckpt
Neural reconstruction (NRE-26.02 gsplat) weights
parsed_config.yaml
NRE… See the full description on the dataset page: https://huggingface.co/datasets/luuuulinnnn/Nurec_eval.BridgeVLA_RLBench_EVAL_DATAarxiv: https://arxiv.org/abs/2506.07961
Kling-Audio-Eval-cacheaidealab-videojp-eval
AIdeaLab VideoJP 評価再現用データ
はじめに
このリポジトリはAIdeaLab VideoJPのFVDを測定するためのデータを
集めました。再現手順を次のとおりに示します。
評価方法
まず、評価用ライブラリをダウンロードします。
git clone https://github.com/JunyaoHu/common_metrics_on_video_quality
ダウンロードできたら、ライブラリのインストール手順を踏んで、インストールします。
インストールしたら、同じディレクトリに次のファイルをコピーしてください
evaluate_videos.py
videos.tar
gen_ja.tar
コピーできたら、videos.tarとgen_ja.tarを展開します。
tar xf videos.tar
tar xf gen_ja.tar
最後にevaluate_videos.pyを実行すると、FVDが表示されるはずです。
おまけ: 評価用映像の作り方… See the full description on the dataset page: https://huggingface.co/datasets/aidealab/aidealab-videojp-eval.twoframe-eval-artifacts-20260505
TwoFrame Eval Artifacts 2026-05-05
Generated image artifacts for TwoFrame image-editing evaluation. The large image payload is stored as tar archives under archives/ to avoid uploading tens of thousands of loose PNG files.
Layout
archives/single_ref.tar: all complete single-reference outputs.
archives/multiref_part*.tar: complete K=2/K=3 multi-reference runs, sharded by run name.
archives/metadata.tar: README, index, manifests, and metrics as an archive.… See the full description on the dataset page: https://huggingface.co/datasets/wyhhey/twoframe-eval-artifacts-20260505.t2v-gen-evalOffline_Evaluationsyntra-testing-evals-v2cls-evaluation-datasetobject_images
Overview
This dataset contains 2D rendered images generated from 3D assets originally sourced from Objaverse and Objathor.The purpose of this dataset is to provide a large-scale collection of photo-realistic renderings for research on vision, multimodal learning, and text-to-3D understanding.
Following prior works such as Diffusion4D and Stable-Zero123,we release only the rendered 2D images (not the original 3D assets) to facilitate efficient experimentation while preserving the… See the full description on the dataset page: https://huggingface.co/datasets/LEGO-Eval/object_images.see-2-sound-evalWe sample images from Laion400M and the web to construct this small evaluation set.
ori_coco_evalrelo_coco_evalmembership-task-evalPAPO-EvalABot-PhysWorld_eval_640_480_HYPIR_vace_skeleton_3250stepWSYue-ASR-eval
WSYue-ASR-eval: Cantonese ASR Benchmark
To address the unique linguistic characteristics of Cantonese in speech recognition, we propose WSYue-ASR-eval, a benchmark specifically designed for evaluating Cantonese ASR systems. It is tailored to assess model performance across diverse lengths, domains, and linguistic phenomena of Cantonese speech.
The test set annotations are provided by Beijing AISHELL Technology Co., Ltd.
Key features:
Annotated through multiple rounds of manual… See the full description on the dataset page: https://huggingface.co/datasets/xuhaorran/WSYue-ASR-eval.Offline_Evaluationeval_benchmark_vlm_hallumembership-evalGaussianAnything-evalcontinuation-task-eval
