datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaianet-testtestdataInter-Edit-Test
Inter-Edit-Test
Official test benchmark release for the CVPR 2026 paper:
Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing
This repository hosts the public release of Inter-Edit-Test, a human-annotated benchmark for the Interactive Instruction-based Image Editing (I^3E) task.
Each sample contains:
a source image,
a coarse user-style interaction mask,
a concise editing instruction,
and a ground-truth edited image.
To simplify large-scale distribution on… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Test.testbnew_dataset_testdroid_testgaia-testwds_wilds-fmow_testwds_wilds-iwildcam_testegoxtreme-test
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
📖 Dataset Information
EgoXtreme is a novel large-scale dataset designed for robust egocentric 6D object pose estimation under extreme environmental conditions. It was captured at 30 fps using Aria glasses, providing high-resolution 1408 x 1408 raw fisheye RGB images. The dataset features 15 participants performing diverse interactions with 13 different objects… See the full description on the dataset page: https://huggingface.co/datasets/taegyoun88/egoxtreme-test.test_repo_1test_1225UCF-Crime_TEST_SEThf download backseollgi/UCF-Crime_TEST_SET --repo-type dataset --local-dir .
tar -xzvf UCF-Crime_TEST_SET.tar.gz
wds_fgvc_aircraft_test
FGVC-Aircraft (Test set only)
Original paper: Fine-Grained Visual Classification of Aircraft
Homepage: https://www.robots.ox.ac.uk/~vgg/data/fgvc-aircraft/
Bibtex:
@techreport{maji13fine-grained,
title = {Fine-Grained Visual Classification of Aircraft},
author = {S. Maji and J. Kannala and E. Rahtu
and M. Blaschko and A. Vedaldi},
year = {2013},
archivePrefix = {arXiv},
eprint = {1306.5151},
primaryClass = "cs-cv"… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_fgvc_aircraft_test.test-testGarments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.test_raw_video_data
ShareGPTVideo Raw Videos for Testing data
All dataset and models can be found at ShareGPTVideo.
Contents:
In case of need, this contains raw videos corresponding to test frames in
Test video frames
X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.wds_wilds-camelyon17_testPMV-test-targzwds_country211_test
Country-211 (Test set only)
Original paper: Learning Transferable Visual Models From Natural Language Supervision
Homepage: https://github.com/openai/CLIP/blob/main/data/country211.md
Derived from YFCC100M: https://multimediacommons.wordpress.com/yfcc100m-core-dataset/
Bibtex:
@article{DBLP:journals/corr/abs-2103-00020,
author = {Alec Radford and
Jong Wook Kim and
Chris Hallacy and
Aditya Ramesh and
Gabriel Goh and… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_country211_test.DressCode-Testtest_image_onlytestwds_vtab-clevr_closest_object_distance_test
CLEVR Closest Object Distance Webdataset (Test set only)
Original paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Homepage: https://cs.stanford.edu/people/jcjohns/clevr/
Bibtex:
@article{DBLP:journals/corr/JohnsonHMFZG16,
author = {Justin Johnson and
Bharath Hariharan and
Laurens van der Maaten and
Li Fei{-}Fei and
C. Lawrence Zitnick and
Ross B. Girshick}… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_vtab-clevr_closest_object_distance_test.wds_vtab-clevr_count_all_test
CLEVR Count All Webdataset (Test set only)
Original paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Homepage: https://cs.stanford.edu/people/jcjohns/clevr/
Bibtex:
@article{DBLP:journals/corr/JohnsonHMFZG16,
author = {Justin Johnson and
Bharath Hariharan and
Laurens van der Maaten and
Li Fei{-}Fei and
C. Lawrence Zitnick and
Ross B. Girshick},
title =… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_vtab-clevr_count_all_test.multilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.test_demoDicFace-test_datasetdeepfake_ecg_full_train_validation_test
