datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
witMINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.MINT-1T-PDF-CC-2023-14
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.Multimodal-Mind2Web
Dataset Summary
Multimodal-Mind2Web is the multimodal version of Mind2Web, a dataset for developing and evaluating generalist agents
for the web that can follow language instructions to complete complex tasks on any website. In this dataset, we align each HTML document in the dataset with
its corresponding webpage screenshot image from the Mind2Web raw dump. This multimodal version addresses the inconvenience of loading images from the ~300GB Mind2Web Raw Dump.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web.Unsafe2Safeimagery_mindbridgeMINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.2026-01-29T15-21-58plus00-00_gdpvalmini-imagenet
Dataset Description
A mini version of ImageNet-1k with 100 of 1000 classes present.
Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions.
Data Splits
Train
50000 samples from ImageNet-1k train split
Validation
10000 samples from ImageNet-1k train split
Test
5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/timm/mini-imagenet.XHRBench
XHRBench
Ultra-High-Resolution Remote Sensing Understanding and Reasoning
🤗 Hugging Face ·
🤖 ModelScope ·
📄 Paper ·
💻 Code
English | 中文
📚 Introduction
XHRBench evaluates fine-grained perception and complex reasoning in multimodal large language models using ultra-high-resolution remote-sensing imagery. This repository retains the name XHRBench and belongs to the same RSHR benchmark project as RSHR-Bench, with a… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/XHRBench.MINT-1T-PDF-CC-2023-50
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.scannet_mini_val_set_suiteminuszero-indian-autonomous-driving-dataset-v2
INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving
Overview
INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving.
This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.minty-astro-ph
MINT-1T ArXiv Astro-ph
An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers).
Overview
Papers
~845k
Total size
~804 GB
Format
WebDataset tar shards
Shards
287 (astro-ph-00000.tar to astro-ph-00286.tar)
Shard size
~3 GB each
Source
MINT-1T (Awadalla et al., 2024)
Data Format
Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.reachy-mini-wall-data
Reachy Mini — wall data (public)
posts.json for the Reachy Mini community wall: the AI-filtered posts shown publicly,
aggregated from Bluesky, YouTube, LinkedIn, TikTok, X and Reddit by the social-wall pipeline.
Fetch it directly (CORS-enabled) from any static site:
const url = "https://huggingface.co/datasets/pollen-robotics/reachy-mini-wall-data/resolve/main/posts.json";
const posts = await (await fetch(url)).json();
Each item: id, platform, author, handle, avatar, text… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-wall-data.MindCraftMindCraft2mini-VTAB
Mini-VTAB
A collection of VTAB (Visual Task Adaptation Benchmark) datasets. We sampled 1K training samples and 1K testing samples for each task.
Tasks
datasets = [
"caltech101",
"cifar10",
"cifar100",
"dtd",
"flowers",
"pets",
"sun397",
"svhn",
"pcam",
"eurosat",
"resisc45",
"diabetic_retinopathy",
"clevr_count_all",
"clevr_closest_object_distance",
"dmlab",
"dsprites_label_x_position"… See the full description on the dataset page: https://huggingface.co/datasets/antofuller/mini-VTAB.Robo4D-200k
Paper
Project Page
Model Repo
TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.minimax-h3-soup
MiniMax H3 Soup
Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090.
What is included
Original benchmark workflow graph (source_prompt.json), manifest, and result table.
Every one-second MP4 from the original C1-C11 benchmark grid and its Euler
repeat sweep. The separate
duration experiments are intentionally not included.
Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a
static render-time summary chart.… See the full description on the dataset page: https://huggingface.co/datasets/badincite/minimax-h3-soup.PolaRGB
PolarFree: Polarization-based Reflection-Free Imaging
Dataset Overview
PolarFree is a high-quality dataset designed for polarization-based reflection removal tasks, as introduced in the CVPR 2025 paper "PolarFree: Polarization-based Reflection-Free Imaging". The dataset aims to support tasks such as image reflection removal and image enhancement, particularly suitable for training and evaluating polarization-based image reflection removal models.
Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mingde/PolaRGB.open-imagesPittadsDB-AdsPics
Pitt Ads Dataset (PittAdsDB)
This folder holds the University of Pittsburgh Ads dataset artifacts we can download publicly, and a helper script to fetch images once access is granted.
What is included (downloaded now)
docs/
readme_images.txt (dataset readme for images)
readme_videos.txt (dataset readme for videos)
image_annotations/
image/QA_Action.json
image/QA_Combined_Action_Reason.json
image/QA_Reason.json
image/Sentiments.json
image/Sentiments_List.txt… See the full description on the dataset page: https://huggingface.co/datasets/Mindykkyan/PittadsDB-AdsPics.VisualProbe_train2026-01-29T07-51-20plus00-00_gdpvalMindCube_lmmseval
MindCube LMMs Eval Dataset
This dataset is formatted for use with lmms-eval framework.
Dataset Schema
Column
Type
Description
id
string
Unique identifier for each sample (format: {split}_{scene_id}_{question_id})
category
list[string]
Category labels (e.g., ["perpendicular", "P-O", "meanwhile", "self"])
type
string
Question type (e.g., "1_frame", "2_frame", "3_frame", "general")
meta_info
list[list[string]]
Metadata about scene objects and their spatial… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/MindCube_lmmseval.BDD100K-laneMET-Bench-Minecraft
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
Vanya Cohen and Raymond Mooney · ICML 2026
Paper · Publication page · Load the dataset · Citation
Domains: Chess · Shell Game · Minecraft
MET-Bench evaluates entity state tracking across text and image modalities. This repository contains the Minecraft domain.
Minecraft
Minecraft is a state prediction task involving partial observations, dynamic… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/MET-Bench-Minecraft.
