datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UrbanVerse-Training-Scenes
UrbanVerse Training Scenes (Urban Cousins)
A collection of ready-to-simulate urban 3D scenes in OpenUSD
for NVIDIA Isaac Sim / Isaac Lab, released by the
VAIL-UCLA lab. Each scene is a self-contained
USD stage with all of its materials and textures, so it can be opened and
simulated directly.
The scenes are generated with UrbanVerse — Scaling Urban Simulation by
Watching City-Tour Videos (Liu et al., ICLR 2026,
arXiv:2510.15018,
project page) — whose UrbanVerse-Gen
pipeline… See the full description on the dataset page: https://huggingface.co/datasets/UCLA-VAIL/UrbanVerse-Training-Scenes.ImageNetV2vaihingen-cr
Paper
This dataset is released as part of our ECCV 2026 paper:
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
Paper page: https://huggingface.co/papers/2607.02471
arXiv: https://arxiv.org/abs/2607.02471
Code: https://github.com/wzy6055/GACR
Idis
Dataset Card for Idis
Dataset Description
Paper Information
Dataset Structure
Dataset Usage
License
Citation
Dataset Description
Idis (Images with distractors) is a VQA benchmark suite for studying how distractors affect the test-time scaling of
reasoning vision-language models. Starting from two base datasets, we add distractors while keeping the target and the
answer unchanged, and vary them along three axes: modality (visual and linguistic), number (1 to 4)… See the full description on the dataset page: https://huggingface.co/datasets/Vail-2000/Idis.RS_Image_Segmentation_Vaihingenmegaunscene
Emergent Extreme-View Geometry in 3D Foundation Models
Yiwen Zhang¹ Joseph Tung² Ruojin Cai³ David Fouhey² Hadar Averbuch-Elor¹
¹Cornell University ²New York University ³Kempner Institute, Harvard University
MegaUnScene Benchmark
Overview
MegaUnScene is a dataset of Internet scenes unseen by existing 3DFMs for benchmarking. There are three test splits split across two evaluation tasks:
Relative Pose Estimation: UnScenePairs and UnScenePairs-t… See the full description on the dataset page: https://huggingface.co/datasets/cornell-vailab/megaunscene.mesh-snomed-entity-alignment-15k
MeSH-SNOMED Entity Alignment 15K
MeSH-SNOMED Entity Alignment 15K is a biomedical heterogeneous knowledge graph alignment benchmark for cross-ontology matching between MeSH and SNOMED CT. It is designed to evaluate entity alignment systems under realistic large-graph conditions, where gold-aligned concepts are embedded in much larger biomedical graphs containing many structurally relevant but non-aligned background entities. This release is intended for the accompanying EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/vaibhavalakshmiravideshik/mesh-snomed-entity-alignment-15k.VAIPE_PILLfinetune-data-for-vision-based-llms5AirQualty_imageConvTrain_Full_DatasetAirQualty_imageConv_2image-parquet2
My Dataset Card
Description…
pill-captionVAIPE_Padaption-chest-xray-pneumonia-labels
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-chest_xray_pneumonia_labels
This dataset contains labeled samples for chest X-ray image classification, distinguishing between normal cases and those with pneumonia. Each entry provides a diagnostic label indicating the presence of pneumonia or a normal lung condition. The data is structured as pairs with a single completion field holding the categorical diagnosis.… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavipadmanabhan/adaption-chest-xray-pneumonia-labels.osworld_tasks_filesfinetune-data-for-vision-based-llms3image-parquet4
My Dataset Card
Description…
finetune-data-for-vision-based-llmsparquetdataset
My Dataset Card
Description…
finetune-data-for-vision-llmsfinetune-data-for-vision-llms5image-parquet1
My Dataset Card
Description…
image-parquet23
My Dataset Card
Description…
toy-shapes-datasetThis dataset is used in the work "Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability" Paper
For more details, please refer to the GitHub repository.
indian_food_imagesfinetune-data-for-vision-based-llms2Project_Three_Datasetparquet7
