CoolFace
Datasetpublic

JimingYang/Touch-Vision-Language-Dataset

A Touch, Vision, and Language Dataset for Multimodal Alignment by Max (Letian) Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution). [Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation] This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.… See the full description on the dataset page: https://huggingface.co/datasets/JimingYang/Touch-Vision-Language-Dataset.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes188downloads
Dataset Card

A Touch, Vision, and Language Dataset for Multimodal Alignment

by <a href="https://max-fu.github.io">Max (Letian) Fu</a>, <a href="https://www.linkedin.com/in/gaurav-datta/">Gaurav Datta</a>, <a href="https://qingh097.github.io/">Huang Huang</a>, <a href="https://autolab.berkeley.edu/people">William Chung-Ho Panitch</a>, <a href="https://www.linkedin.com/in/jaimyn-drake/">Jaimyn Drake</a>, <a href="https://joeaortiz.github.io/">Joseph Ortiz</a>, <a href="https://www.mustafamukadam.com/">Mustafa Mukadam</a>, <a href="https://scholar.google.com/citations?user=p6DCMrQAAAAJ&hl=en">Mike Lambeta</a>, <a href="https://lasr.org/">Roberto Calandra</a>, <a href="https://goldberg.berkeley.edu">Ken Goldberg</a> at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution).

[Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation]

<p align="center"> <img src="img/splashfigurealt.png" width="800"> </p>

This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.

Instructions for Dataset

Due to the single file upload limit, we sharded the dataset into 8 zip files. To use the dataset, we first download them using the GUI or use git:

bash
# git lfs install (optional)
git clone git@hf.co:datasets/mlfu7/Touch-Vision-Language-Dataset

# There's an improved version of this dataset, which fixed the tactile/image folder swapping issue. 
git clone git@hf.co:datasets/yoorhim/TVL-revise

cd Touch-Vision-Language-Dataset
zip -s0 tvl_dataset_sharded.zip --out tvl_dataset.zip
unzip tvl_dataset.zip 

The structure of the dataset is as follows:

tvl_dataset
├── hct
│   ├── data1
│   │   ├── contact.json
│   │   ├── not_contact.json
│   │   ├── train.csv
│   │   ├── test.csv
│   │   ├── finetune.json
│   │   └── 0-1702507215.615537
│   │       ├── tactile
│   │       │   └── 165-0.025303125381469727.jpg
│   │       └── vision
│   │           └── 165-0.025303125381469727.jpg
│   ├── data2
│   │   ...
│   └── data3
│       ...
└── ssvtp
    ├── train.csv
    ├── test.csv
    ├── finetune.json
    ├── images_tac
    │   ├── image_0_tac.jpg
    │   ...
    ├── images_rgb
    │   ├── image_0_rgb.jpg
    │   ...
    └── text
        ├── labels_0.txt
        ...

Training and Inference

We provide the checkpoints of TVL tactile encoder and TVL-LLaMA here. Please refer to the official code release and the paper for more info.

Citation

Please give us a star 🌟 on Github to support us!

Please cite our work if you find our work inspiring or use our code in your work:

@inproceedings{
    fu2024a,
    title={A Touch, Vision, and Language Dataset for Multimodal Alignment},
    author={Letian Fu and Gaurav Datta and Huang Huang and William Chung-Ho Panitch and Jaimyn Drake and Joseph Ortiz and Mustafa Mukadam and Mike Lambeta and Roberto Calandra and Ken Goldberg},
    booktitle={Forty-first International Conference on Machine Learning},
    year={2024},
    url={https://openreview.net/forum?id=tFEOOH9eH0}
}