CoolFace
Datasetpublic

FALCON-VLA/CALVIN-3D_PCD-ABCD_D

| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026) Zhengshen Zhang   Hao Li   Yalun Dai   Zhengbang Zhu   Lei Zhou   Chenchen Liu   Dong Wang   Francis E. H. Tay   Sijin Chen   Ziwei Liu   Yuxiao Liu*†   Xinghang Li*   Pan Zhou*   *Corresponding Author  †Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
2likes348downloads
Dataset Card

<div align="center">

<h1>| <em>FALCON</em> | From Spatial to Actions:<br>Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)</h1>

<a href="https://arxiv.org/abs/2510.17439" target="blank"> <img alt="arXiv" src="https://img.shields.io/badge/arXiv-FALCON-red?logo=arxiv" height="25" /> </a> <a href="https://falcon-vla.github.io/" target="blank"> <img alt="Website" src="https://img.shields.io/badge/🌎Website-falcon.io-blue.svg" height="25" /> </a> <a href="https://github.com/FALCON-VLA/FALCON" target="blank"> <img alt="GitHub Code: FALCON" src="https://img.shields.io/badge/Code-FALCON-181717?logo=github&logoColor=white" height="25" /> </a> <a href="https://huggingface.co/papers/2510.17439" target="blank"> <img alt="HF Paper: FALCON" src="https://img.shields.io/badge/%F0%9F%A4%97%20Paper-FALCON-ffc107?color=ffc107&logoColor=white" height="25" /> </a> <a href="https://huggingface.co/FALCON-VLA/FALCON-series" target="blank"> <img alt="HF Model: FALCON" src="https://img.shields.io/badge/%F0%9F%A4%97%20Model-FALCON-ffc107?color=ffc107&logoColor=white" height="25" /> </a> <br> <a href="https://www.python.org/" target="blank"> <img alt="Python 3.8" src="https://img.shields.io/badge/Python-%3E=3.8-blue" height="25" /> </a> <a href="https://pytorch.org/" target="blank"> <img alt="PyTorch" src="https://img.shields.io/badge/PyTorch-%3E=2.1-orange" height="25" /> </a>

</div>

<div align="center"> <br> <div style="text-align: center;"> <a href="https://scholar.google.com/citations?user=8nrJ1vsAAAAJ&hl=en" target="blank">Zhengshen Zhang</a> &emsp; <a href="https://scholar.google.com/citations?user=4dokjDoAAAAJ&hl=zh-CN" target="blank">Hao Li</a> &emsp; <a href="https://scholar.google.com/citations?user=6XyNVowAAAAJ&hl=en" target="blank">Yalun Dai</a> &emsp; <a href="https://scholar.google.com/citations?user=ozatRA0AAAAJ&hl=zh-CN" target="blank">Zhengbang Zhu</a> &emsp; <a href="https://scholar.google.com/citations?user=VhToj4wAAAAJ&hl=zh-CN" target="blank">Lei Zhou</a> &emsp; <br> <a href="https://sg.linkedin.com/in/liu-chenchen" target="blank">Chenchen Liu</a> &emsp; <a href="" target="blank">Dong Wang</a> &emsp; <a href="https://scholar.google.com/citations?user=mfH9UFIAAAAJ&hl=en" target="blank">Francis E. H. Tay</a> &emsp; <a href="https://ch3cook-fdu.github.io/" target="blank">Sijin Chen</a> &emsp; <br> <a href="https://liuziwei7.github.io/" target="blank">Ziwei Liu</a> &emsp; <a href="https://scholar.google.com/citations?user=i8wNtSgAAAAJ&hl=en" target="blank">Yuxiao Liu</a><sup>*</sup><sup>&dagger;</sup> &emsp; <a href="https://scholar.google.com/citations?user=laOWyTQAAAAJ&hl=zh-CN" target="blank">Xinghang Li</a><sup></sup> &emsp; <a href="https://panzhous.github.io/" target="_blank">Pan Zhou</a><sup></sup> &emsp; <br> <p style="text-align: center; margin-bottom: 0;"> <span class="author-note"><sup>*</sup>Corresponding Author</span>&emsp; <span class="author-note"><sup>&dagger;</sup>Project Lead</span> </p> <br> <p style="text-align: center;"> ByteDance Seed <br> National University of Singapore &emsp; Nanyang Technological University <br> Tsinghua University &emsp; Singapore Management University</p> </div> </div>

🚀 Introduction

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head of a VLA model, enabling robust spatial understanding and SOTA performance across diverse manipulation tasks without disrupting vision-language alignment. See our paper at here.

Dataset Card for CALVIN Dataset_ABCD-D in Point Cloud Format

Dataset Details

This dataset repo contains preprocessed point cloud data in world coordinate system from the static camera and wrist camera under the ABCD-D setting (training/validation) of the Calvin dataset.

NOTE: The related camera extrinsic parameters are deprecated. For accurate static camera parameters, please refer to this repo.

Dataset Sources

<!-- Provide the basic links for the dataset. -->

📦 Usage

<!-- Address questions around how the dataset is intended to be used. -->

We provide an efficient way to download the dataset in shards, merge them into a single archive, and then extract the final data.

  1. 1.Prepare a folder for downloaded parts
bash
mkdir -p downloaded_parts/
  1. 1.Download the dataset shards

You can use either the official Hugging Face CLI or the hfd.sh script:

Option A: Hugging Face CLI

bash
huggingface-cli download FALCON-VLA/CALVIN-3D_PCD-ABCD_D --repo-type dataset

Option B: High-speed download with hfd.sh (recommended)

bash
# Install dependencies
sudo apt-get install aria2 git-lfs -y

# Download the helper script
wget https://hf-mirror.com/hfd/hfd.sh
chmod +x hfd.sh

# Download the dataset shards
./hfd.sh FALCON-VLA/CALVIN-3D_PCD-ABCD_D --dataset --tool aria2c -x 8 -j 5 --include "*.tar.gz"
  1. 1.Merge the downloaded shards

Make sure all shard files are placed under downloaded_parts/, then merge them into a single tarball:

bash
cat downloaded_parts/packaged_ABCD_D*.tar.gz > packaged_ABCD_D.tar.gz
  1. 1.Extract the merged archive
bash
tar -xzvf ./packaged_ABCD_D.tar.gz -C ./

After extraction, the dataset files will be available in the current directory.

For point cloud data loading details, please refer to our released data class **DiskCalvinDataset3D**.

Dataset Structure

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->

After extraction, the dataset is organized as follows:

bash
packaged_ABCD_D/
├── training/
│   ├── episode_0000001.npz
│   ├── episode_0000002.npz
│   └── ...
└── validation/
    ├── episode_0000038.npz
    ├── episode_0000039.npz
    └── ...

Each .npz file corresponds to one episode and contains the following fields:

bash
static_pcd: point cloud from the static camera view, shape: (200, 200, 3)
static_rgb: RGB image from the static camera view, shape: (200, 200, 3)
gripper_pcd: point cloud from the wrist camera view, shape: (84, 84, 3)
gripper_rgb: RGB image from the wrist camera view, shape: (84, 84, 3)
annotation_id: episode annotation identifier, type: int

Note: static_cam_ex_mat and gripper_cam_ex_mat are deprecated and are no longer used in the current pipeline.

🤗 FAQs

If you encounter any issues, feel free to open an issue or reach out through discussions. We appreciate your feedback and contributions! 🚀

🖊️ Citation

If you find this project useful in your research, please consider cite:

BibTeX
@article{zhang2025spatial,
  title={From spatial to actions: Grounding vision-language-action model in spatial foundation priors},
  author={Zhang, Zhengshen and Li, Hao and Dai, Yalun and Zhu, Zhengbang and Zhou, Lei and Liu, Chenchen and Wang, Dong and Tay, Francis EH and Chen, Sijin and Liu, Ziwei and others},
  journal={arXiv preprint arXiv:2510.17439},
  year={2025}
}
BibTeX
@article{mees2022calvin,
  title={Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks},
  author={Mees, Oier and Hermann, Lukas and Rosete-Beas, Erick and Burgard, Wolfram},
  journal={IEEE Robotics and Automation Letters},
  volume={7},
  number={3},
  pages={7327--7334},
  year={2022},
  publisher={IEEE}
}

🪪 License

All datasets, as well as our codebase are released under the Apache-2.0 License.

❤️ Acknowledgement

FALCON is built with reference to the code of the following projects: RoboVLMs, Microsoft Kosmos-2, VGGT, and ManiUniCon. Thanks for their awesome work!