datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mantis-Instruct
Mantis-Instruct
Paper | Website | Github | Models | Demo
Summaries
Mantis-Instruct is a fully text-image interleaved multimodal instruction tuning dataset,
containing 721K examples from 14 subsets and covering multi-image skills including co-reference, reasoning, comparing, temporal understanding.
It's been used to train Mantis Model families
Mantis-Instruct has a total of 721K instances, consisting of 14 subsets to cover all the multi-image skills.
Among the… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Instruct.Cambrian10M_For_MantisWebsight_Mantis_DataDocmatix_For_MantisMantis-Eval
Overview
This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/.
Statistics
This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems.
Leaderboard
We list the current results as follows:
Models
Size
Mantis-Eval
LLaVA OneVision
72B
77.60
LLaVA OneVision
7B
64.20
GPT-4V
-
62.67… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Eval.cota-mantis
🌮 TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-Action
🌐 Website | 📑 Arxiv | 💻 Code| 🤗 Datasets
If you like our project or are interested in its updates, please star us :) Thank you! ⭐
Summary
TLDR: CoTA is a large-scale dataset of synthetic Chains-of-Thought-and-Action (CoTA) generated by multi-modal large language models.
Load data
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/cota-mantis.mantis_libero_lerobotThis dataset was created using LeRobot.
If you find our code or models useful in your work, please cite our paper:
@article{yang2025mantis,
title={Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight},
author={Yang, Yi and Li, Xueqi and Chen, Yiyang and Song, Jin and Wang, Yihan and Xiao, Zipeng and Su, Jiadi and Qiaoben, You and Liu, Pengfei and Deng, Zhijie},
journal={arXiv preprint arXiv:2511.16175},
year={2025}
}
mantis_datasetgello_hil_mantis
gello_hil_mantis
Pick up the plate and place it on mantis.
Split from aliy98/gello_hil_full_merged without modifying the original dataset.
Episodes: 120
Frames: 81357
Tasks:
'pick up the plate and place it on mantis'
vr_hil_mantis
vr_hil_mantis
Pick up the plate and place it on mantis.
Split from aliy98/vr_hil_full_merged without modifying the original dataset.
Episodes: 120
Frames: 85758
Tasks:
'pick up the plate and place it on mantis'
MIQA_evalprogram-cota-mantis
🌮 TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-Action
🌐 Website | 📑 Arxiv | 💻 Code| 🤗 Datasets
If you like our project or are interested in its updates, please star us :) Thank you! ⭐
Summary
TLDR: CoTA is a large-scale dataset of synthetic Chains-of-Thought-and-Action (CoTA) generated by programs.
Load data
from datasets import load_dataset
dataset = load_dataset("Salesforce/program-cota-mantis"… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/program-cota-mantis.mantis-translatedMIQA_samplegsma_prd_syntheticmantis-spacellavagsma_prd_synthetic_embeddingmantis_abs_joint_abs_joint_30hzThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5",
"total_episodes": 100,
"total_frames": 62815,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliy98/mantis_abs_joint_abs_joint_30hz.gsma_prd_synthetic_qamolmo2-mantis-instruct-llava_665k_multitelecom-questionsimnet1k_mantis_mantidmolmo2-mantis-instruct-starmantisMantisMantis-Evalmolmo2-mantis-instruct-nextqagsma_discover_synthetic_embeddingMantis_rumantis_rus_dataset
