datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vdrop-crossview-8k
Cross-View Visual Thinking — 8K Training Set
Training data for “How and What to Imagine? Visual Thinking in Unified Multimodal Models
for Cross-View Spatial Reasoning”.
This repository contains the training set only (single train split). Evaluation
benchmarks used in the paper are third-party and are not distributed here.
Given two egocentric views of the same indoor scene with partially overlapping fields of view,
each sample asks a multiple-choice cross-view spatial… See the full description on the dataset page: https://huggingface.co/datasets/QianYangMILA/vdrop-crossview-8k.CrossViewUrbanTrafficDataset
Cross-View Urban Traffic Dataset
Dataset Summary
The Cross-View Urban Traffic Dataset (CVUTD) is a benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections in Regensburg, Germany.
The dataset is designed to support two linked tasks:
Cross-view identity matching between street-view and drone-view object tracks
Ego-to-BEV prediction using aerial supervision… See the full description on the dataset page: https://huggingface.co/datasets/prakharbh/CrossViewUrbanTrafficDataset.LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset
CrossView Prompt Dataset
The training dataset behind the
CrossView Prompt IC-LoRA
for LTX-Video 2.3 — a "virtual second camera" adapter that re-renders a scene
from a new viewpoint described by a short prompt.
Each sample is a pair of static-camera clips of the same scene (a reference
view and a target view) plus a camera-delta caption describing how the
target camera differs from the reference.
Contents
clips/<scene>/<cam>.mp4 # 504 unique clips, native… See the full description on the dataset page: https://huggingface.co/datasets/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset.CrossViewUrbanTrafficSubsetDataset
Cross-View Urban Traffic Subset Dataset
This repository provides a representative sample of the Cross-View Urban Traffic Dataset (CVUTD) for reviewer inspection and qualitative verification.
The subset is intended to let reviewers:
inspect the raw and annotated data format,
verify annotation quality,
understand the cross-view correspondence structure,
and assess the benchmark design without downloading the full dataset.
The full dataset is hosted separately and is available at… See the full description on the dataset page: https://huggingface.co/datasets/prakharbh/CrossViewUrbanTrafficSubsetDataset.arkitscenes-crossview-inpaint
Carlhahaha/arkitscenes-crossview-inpaint
Cross-view pairs generated from inpainting results (Topomap; meta root: /mnt/NAS/data/jz4725/topomap).
Splits
Uploaded splits: test, train, validation
Schema (columns)
Image_a: Original image of sample A (datasets.Image)
Image_b: Original image of sample B (datasets.Image)
Inpaint_b: Inpainted image of sample B (datasets.Image)
mask_b: Mask used for inpainting on B (datasets.Image)
point_b: Normalized centroid of B's… See the full description on the dataset page: https://huggingface.co/datasets/Carlhahaha/arkitscenes-crossview-inpaint.test-crossviewminival-crossview
