datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SekaiMINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.fine-t2i
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv]
by Xu Ma, Yitian Zhang,
Qihua Dong, Yun Fu
Northeastern Univeristy
Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient).
🆕 What's New
[2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️
[2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.MINT-1T-PDF-CC-2023-14
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.mls_sidon
MLS-Sidon
Overview
This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
The dataset is provided in WebDataset format for efficient large-scale training.
Source: Multilingual LibriSpeech
Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese
Format: WebDataset (.tar shards)
License: CC-BY-4.0
Dataset Structure
Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.MINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.Mobile-O-Post-Train
Mobile-O Post-Training Data
Unified Multimodal Post-Training · ~105K Quadruplet Samples
📌 Overview
This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.
📊 Dataset Format
Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.MINT-1T-PDF-CC-2023-50
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.FTP-1-Dataset
FTP-1-Dataset
FTP-1-Dataset contains heterogeneous tactile manipulation data for FTP-1 pretraining.
This release currently includes 18 dataset archives:
FreeTacMan
MotionTrans
RDP
RDP_Bimanual
RH20TCfg5Franka
RH20TCfg6ATIAxia
RH20TCfg7Tactile
Unit
Unit_Bimanual
VLA_touch
ViTaMIn
VisuoTactile_D-WHEEL
VisuoTactile_QINGLOONG
exUMI
sharpa
Each dataset directory contains either a single <dataset>.tar file or split parts named <dataset>.tar.part-*. For split archives, concatenate… See the full description on the dataset page: https://huggingface.co/datasets/MJJJJ1064/FTP-1-Dataset.minty-astro-ph
MINT-1T ArXiv Astro-ph
An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers).
Overview
Papers
~845k
Total size
~804 GB
Format
WebDataset tar shards
Shards
287 (astro-ph-00000.tar to astro-ph-00286.tar)
Shard size
~3 GB each
Source
MINT-1T (Awadalla et al., 2024)
Data Format
Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.audiosnippetsmedpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.Clotho-Moment
Clotho-Moment
This repository provides wav files used in Language-based Audio Moment Retrieval.
Each sample includes long audio containing some audio events with the temporal and textual annotation.
Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/
Code: https://github.com/line/lighthouse
Split
Train
train/train-{000..715}.tar
37930 audio samples
Valid
valid/valid-{000..108}.tar
5741 audio samples
Test
test/test-{000..142}.tar
7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.marianne_pdf_7melee-ranked-replays
Melee Ranked Replays
Anonymized Slippi ranked replays (platinum+) from Super Smash Bros. Melee,
sharded by character and rank pair. Built for behavior-cloning and other
replay-driven ML work on Melee — notably MIMIC.
Contents
Raw .slp files grouped into tarballs by (character, rank_pair, source_archive),
organized into per-character folders:
{CHAR}/
{CHAR}_{rank_pair}_a{N}.tar.gz
metadata/
metadata_a{N}.json
Characters (25): BOWSER, CPTFALCON, DK, DOC, FALCO… See the full description on the dataset page: https://huggingface.co/datasets/erickfm/melee-ranked-replays.audiosnippets_small_with_detailed_annotationaudiosnippets_small_with_detailed_annotation2Mobile-O-Pre-Train
Mobile-O Pre-Training Data
Cross-Modal Alignment · 9M Text-Image Pairs
📌 Overview
This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.
📊 Dataset Composition
Source
Samples
Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.marianne_pdf_9S1-MMAlignS1-MMAlign
A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset
S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers.
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.marianne_pdf_5marianne_pdf_3captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
ms1mv3-wds
MS-Celeb-1M (v3)
This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper.
There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.Molmo2-ER-RoboPoint
Molmo2-ER · wentao-yuan/robopoint-data
1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: wentao-yuan/robopoint-data
Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721)
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.mls_hq_urgent_track1majestrino-data
