yrlyrl/spatial-mmcot-vdrop
Spatial MMCoT v1 · vdrop VDrop cross-view (QianYangMILA/vdrop-crossview-8k, arXiv:2605.27310), Infinigen Indoors rooms. Two egocentric views of one room with partly overlapping fields of view are the input (input_image_0 = cam0, input_image_1 = cam1); the target is upstream's ground-truth panorama render, a wide panorama of the room that spans both views. The questions are multiple choice and the trained answer is the option letter. The upstream release ships no reasoning text… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-vdrop.
This repository belongs to yrlyrl on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
