datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MIT-Indoor-Scenesmit-adobe-fivek
Adobe FiveK
This is an upload of the Adobe FiveK dataset.
Note that I am not one of the authors of this dataset, if one of the authors would like to take ownership of this repository please reach out to me.
The data provided is not in the original format either.
Due to the massive size of the dataset >1TB I elected to convert all .tif and .dng files to a standard .webp with lossless compression.
Please refer to the dataset homepage for access to the uncompressed versions of the… See the full description on the dataset page: https://huggingface.co/datasets/logasja/mit-adobe-fivek.sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.rhizomorphic-networks-data
Rhizomorphic Networks — Data and Analysis Outputs
This dataset contains the experimental image data, segmentation outputs, and
downstream analysis results used for the quantitative analysis of
Armillaria gallica rhizomorphic networks.
The directory structure is organised according to the main stages of the analysis
pipeline:
Rhizomorphic Networks/
│
├── 01_raw inputs/
│ ├── control/
│ ├── furnace/
│ └── nutrient density/
│
├── 02_segmentation outputs/
│ ├── raw… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/rhizomorphic-networks-data.LADI-v2-dataset
Dataset Card for LADI-v2-dataset
Dataset Summary : v2
The LADI-v2 dataset is a set of aerial disaster images captured and labeled by the Civil Air Patrol (CAP). The images are geotagged (in their EXIF metadata). Each image has been labeled in triplicate by CAP volunteers trained in the FEMA damage assessment process for multi-label classification; where volunteers disagreed about the presence of a class, a majority vote was taken. The classes are:
bridges_any… See the full description on the dataset page: https://huggingface.co/datasets/MITLL/LADI-v2-dataset.toxo_mitoOCR-liboaccn-OPUS-MIT-5M-clean
Description
This dataset is a processed version of liboaccn/OPUS-MIT-5M to make it easier to use, particularly for a visual question answering task where answer is an OCR transcription.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than a path in an image column.We've also created a question column containing around 40 prompts based on via tutoiement, vouvoiement and imperative forms.
Note that this dataset contains only the… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/OCR-liboaccn-OPUS-MIT-5M-clean.mit-adobe-fivek
Adobe FiveK
This is an upload of the Adobe FiveK dataset.
Note that I am not one of the authors of this dataset, if one of the authors would like to take ownership of this repository please reach out to me.
The data provided is not in the original format either.
Due to the massive size of the dataset >1TB I elected to convert all .tif and .dng files to a standard .webp with lossless compression.
Please refer to the dataset homepage for access to the uncompressed versions of the… See the full description on the dataset page: https://huggingface.co/datasets/KlyaT/mit-adobe-fivek.acme-home-inbox
ACME Home Inbox
Public dataset: https://huggingface.co/datasets/Mitchins/acme-home-inbox
The v0.2.0 checkpoint contains the complete synthetic dataset plus the first
canonical OCR/vision deployment bake-off. Benchmark results are an auditable
research checkpoint, not a claim that any tested routing policy is ready for
unattended household use.
ACME Home Inbox is a fully synthetic, reproducible stress test for a practical
systems question:
Does graphical evidence change the… See the full description on the dataset page: https://huggingface.co/datasets/Mitchins/acme-home-inbox.jump-cp-0016-labelfree-MitoMitsuArtsvdquant-datasets
Quantization Library: DeepCompressor Inference Engine: Nunchaku
[Paper]
[Code]
[Website]
[Blog]
This is the sDCIdataset used in SVDQuant for benchmarking.
If you find this dataset useful or relevant to your research, please cite
@article{
li2024svdquant,
title={SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models},
author={Li*, Muyang and Lin*, Yujun and Zhang*, Zhekai and Cai, Tianle and Li, Xiuyu and Guo, Junxian and Xie, Enze and… See the full description on the dataset page: https://huggingface.co/datasets/mit-han-lab/svdquant-datasets.color-multi-fractal-db-1k
Dataset Card for Color Multi Fractal DB 1k
This is a pre-generated 1k classes, 1M images colored-multi-fractal-images dataset based on Improving Fractal Pre-training by Connor Anderson et al. and Multi-Fractal-Dataset by FYSignate1009.
We have changed some fractal parameters so that our ViT pretraining can converge. Modified parameters can be found on this repo.
You can pretrain vision transformers without worrying about dataset licensing for commercial use.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/color-multi-fractal-db-1k.vrm-color-concept-550k
VRM Color Concept 550K
Summary
This is a dataset to train anime-style text-to-image or any text and image multimodal models without copyright/licensing concerns.
All assets/materials utilized in this dataset are CC0 or properly licensed, and no pretrained models or any AI models are used to build this dataset.
Image, Metadata and Dataset License
All images, metadata in this dataset and the dataset itself are licensed under CC BY-NC 4.0 by ELAN MITSUA Project… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/vrm-color-concept-550k.art-museums-pd-440k
Art Museums PD 440K
Summary
This is a dataset to train text-to-image or any text and image multimodal models with minimized copyright/licensing concerns.
All images and texts in this dataset are orignally shared under CC0 or public domain, and no pretrained models or any AI models are used to build this dataset except for our ElanMT model to translate English captions to Japanese.
ElanMT model is trained solely on licensed corpus.
Data sources
Images and… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/art-museums-pd-440k.safe-commons-pd-3m
Safe Commons PD 3M
This is a balanced and safe-to-use public domain / CC0 images dataset.
All images and texts come from Wikimedia Commons and Wikidata with strict filtering.
Images license is either Public Domain or CC0 (varies by image).
Texts license is either CC0 or CC BY-SA (varies by caption source).
No synthetic data (AI generated images or captions) is in the dataset.
To build this dataset, we tried to avoid any knowledge leaks from existing pre-trained models at the… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/safe-commons-pd-3m.car-parts-segmentation-yolo
AutoInspect - Car Parts Dataset (Ultralytics YOLO segmentation)
YOLO-версия датасета с сегментацией деталей авто Car Parts Dataset.
Часть проекта AutoInspect (pipeline: view classification → car parts segmentation → damage segmentation).
Основан на датасете от HITL. Для парных деталей были прставлены тэги side (left/right) при помощи Supervisely App. Список таких деталей:
Headlight
Tail-light
Mirror
Front-window
Back-window
Front-door
Back-door
Front-wheel
Back-wheel
Fender… See the full description on the dataset page: https://huggingface.co/datasets/mitbersh/car-parts-segmentation-yolo.pixcell
PixCell Dataset
PixCell is a verified image-to-code curriculum for reconstructing photonic
geometry as primitive-only GDSFactory Python programs. Each model row contains
a high-visibility component image, its physical footprint, and an audited
program.
Core curriculum gallery |
Depth representation overview
Configurations
Configuration
Rows
Purpose
depth
4,560
Default corpus for supervised training and parameter recovery
core
738
Compact L0 to L4… See the full description on the dataset page: https://huggingface.co/datasets/qpaig-mit/pixcell.oakink2-vitra-streaming-v1
OakInk2 → VITRA Stage-1 (complete audited release)
This repository contains all 627 physical OakInk2-TaMF sequences converted to
VITRA Stage-1. Every sequence source pair is pinned to
kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact
frame identity, converted across the four calibrated views, checked by geometry
and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded,
and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.MITO_DatasetBNS-SSM-AFRAME
BNS-SSM-AFRAME
BNS waveform datasets for training/validating regression models (aframe).
Large HDF5 files are split into parts (Hugging Face's 50GB file limit).
Reassemble before use:
cat end_o3_ratesandpops_bns.hdf5.part-* > end_o3_ratesandpops_bns.hdf5
cat end_o3_ratesandpops_bns_uniform_chirp.hdf5.part-* > end_o3_ratesandpops_bns_uniform_chirp.hdf5
cat diagnostic.hdf5.part-* > diagnostic.hdf5
Layout
train/ — raw polarization waveforms… See the full description on the dataset page: https://huggingface.co/datasets/kyoon-mit/BNS-SSM-AFRAME.mitosis-gjs3g
Dataset Card for mitosis-gjs3g
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
mitosis-gjs3g
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/mitosis-gjs3g.shrec_empathic
Social Human Robot Embodied Conversation (SHREC) Dataset: Empathic Subset (RSS 2026)
The SHREC Empathic subset contains real-world human-robot interaction video data from Shen et al. (2024), collected over a month-long deployment of social robots in participants’ homes, as participants engage in natural, empathic storytelling interactions with AI agents.
Authors: Dong Won Lee, Yubin Kim, Sooyeon Jeong, Denison Guvenoz, Parker Malachowsky, Louis-Philippe Morency, Cynthia… See the full description on the dataset page: https://huggingface.co/datasets/MIT-personal-robots/shrec_empathic.MitoEM
Leaderboard: mitoem.grand-challenge.org
license: mit
task_categories:
- image-segmentation
language:
- en
pretty_name: MitoEM
size_categories:
- 1B<n<10B
mitomit-indoor-outdoorshrec_wellness_dorm
Social Human Robot Embodied Conversation (SHREC) Dataset: Wellness Dorm Subset (RSS 2026)
The SHREC Wellness Dorm subset contains longitudinal, real-world human-robot interaction video data data from Jeong et al. (2020), where a robotic positive psychology coach was deployed in MIT student dormitories. Participants engaged in daily wellbeing sessions with the robot over the course of 1–4 weeks.
Authors: Dong Won Lee, Yubin Kim, Sooyeon Jeong, Denison Guvenoz, Parker… See the full description on the dataset page: https://huggingface.co/datasets/MIT-personal-robots/shrec_wellness_dorm.mitosis-generations
Mitosis Generations
Every decomposition run of the Mitosis Space is
logged here: the input image, the RGBA layers produced from it, and the settings used.
⚠️ Public log
The Space is public and every run is recorded in this public dataset, including images uploaded
by visitors. Do not submit anything that cannot be shared publicly.
Layout
data/
metadata.jsonl # one JSON object per run
<run_id>/
input.png # exactly what… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/mitosis-generations.MetaMaterialsDiscoveryImages
Overview
Images used for image-to-simulator algorithm.
Citation
@misc{buehler2026metamaterialslaboratoryswarm,
title={Artificial intelligence agents autonomously build computational laboratories that reveal design principles of hierarchical metamaterial failure},
author={Markus J. Buehler},
year={2026},
eprint={xxxx.yyyyy},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/xxxx.yyyyy},
}
leaf-images
LeafGAN: Nature-inspired architected materials using unsupervised deep learning
Reference: Shen, S.C., Buehler, M.J. Nature-inspired Architected materials using unsupervised deep learning. Communications Enginering, 2022, DOI: https://www.nature.com/articles/s44172-022-00037-0
Abstract: Nature-inspired material design is driven by superior properties found in natural architected materials and enabled by recent developments in additive manufacturing and machine learning. Existing… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/leaf-images.
