cross-view
vdrop-crossview-8k
Cross-View Visual Thinking — 8K Training Set
Training data for “How and What to Imagine? Visual Thinking in Unified Multimodal Models
for Cross-View Spatial Reasoning”.
This repository contains the training set only (single train split). Evaluation
benchmarks used in the paper are third-party and are not distributed here.
Given two egocentric views of the same indoor scene with partially overlapping fields of view,
each sample asks a multiple-choice cross-view spatial… See the full description on the dataset page: https://huggingface.co/datasets/QianYangMILA/vdrop-crossview-8k.CrossViewUrbanTrafficDataset
Cross-View Urban Traffic Dataset
Dataset Summary
The Cross-View Urban Traffic Dataset (CVUTD) is a benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections in Regensburg, Germany.
The dataset is designed to support two linked tasks:
Cross-view identity matching between street-view and drone-view object tracks
Ego-to-BEV prediction using aerial supervision… See the full description on the dataset page: https://huggingface.co/datasets/prakharbh/CrossViewUrbanTrafficDataset.LIBERO-CrossView-Pairs
LIBERO-CrossView-Pairs
LIBERO-CrossView-Pairs is a same-state paired-view dataset for training camera-robust vision-language-action policies on LIBERO. Each row contains two scene-camera observations of the exact same simulator state: a nominal LIBERO scene view and one camera-perturbed view. The paired images share the same robot state, language instruction, action target, episode index, frame index, and MuJoCo state; only the scene-camera extrinsics differ.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Bingqiii/LIBERO-CrossView-Pairs.LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset
CrossView Prompt Dataset
The training dataset behind the
CrossView Prompt IC-LoRA
for LTX-Video 2.3 — a "virtual second camera" adapter that re-renders a scene
from a new viewpoint described by a short prompt.
Each sample is a pair of static-camera clips of the same scene (a reference
view and a target view) plus a camera-delta caption describing how the
target camera differs from the reference.
Contents
clips/<scene>/<cam>.mp4 # 504 unique clips, native… See the full description on the dataset page: https://huggingface.co/datasets/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset.CrossViewQA_imgCross-view_Wireless_MIMO_Dataset
