datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot.
T2I-ImageNet-Normalsdxl_images_easy_prompts-artists-seed1Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.waifu-preprocessed-datasetart-museums-pd-440k
Art Museums PD 440K
Summary
This is a dataset to train text-to-image or any text and image multimodal models with minimized copyright/licensing concerns.
All images and texts in this dataset are orignally shared under CC0 or public domain, and no pretrained models or any AI models are used to build this dataset except for our ElanMT model to translate English captions to Japanese.
ElanMT model is trained solely on licensed corpus.
Data sources
Images and… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/art-museums-pd-440k.computer-agent-arena
Computer Agent Arena: Evaluating Computer-Use Agents via Crowdsourcing from Real Users
Dataset Description
Computer Agent Arena is a comprehensive evaluation platform for multi-modal AI agents, particularly focusing on computer use and GUI interaction tasks. This dataset contains real interaction trajectories from various state-of-the-art AI agents performing complex computer tasks in controlled environments.
The dataset includes:
4,641 agent trajectories across diverse… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/computer-agent-arena.T2I-ImageNet-CutMixinfinigen-articulated
Infinigen-Articulated Assets
Formerly Infinigen-Sim
Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.Lowlight-Smartphone-Dataset
[WACV'26] Low-light Smartphone Dataset (LSD)
This is the official dataset proposed in our paper titled "Illuminating Darkness: Learning to Enhance Low-light Images In-the-Wild"
📄 Paper: arXiv💻 Code: GitHub - LSD-TFFormer
Overview
We introduce LSD, the largest in-the-wild Single-Shot Low-Light Image Enhancement (SLLIE) dataset to date.
Dataset Structure
This repository contains the following training data files:
patch_DLL_gtPatch.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/ARM4588/Lowlight-Smartphone-Dataset.VoT-video-latent-archiveffhq_captioned_1024
ffhq_captioned_1024
A captioned bucketed-shards export of gaunernst/ffhq-1024-wds.
This export contains 70,000 square face and portrait images from FFHQ, stored as
JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded
from the original dataset, deterministically converted to RGB, and re-encoded as
high-quality JPEG (quality=95, adaptive subsampling). Captions were generated
with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.HFGaussiansdxl_images_sb_prompts-multi_artist-seed1Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata… See the full description on the dataset page: https://huggingface.co/datasets/bbrangeo/Cowpea-Architecture-XML.dtasettarsphere-encoder-fid-artifacts
Sphere Encoder FID Evaluation Artifacts
This repository contains the evaluation artifacts for the paper Image Generation with a Sphere Encoder.
Project Page | GitHub Repository
These artifacts include data statistic files (fid_stats) and reference images (fid_refs) used to calculate Fréchet Inception Distance (FID) for generative models across several datasets, including CIFAR-10, ImageNet, Animal Faces, and Oxford Flowers.
Workspace Setup
Download the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/sphere-encoder-fid-artifacts.12HZ-SegmentationGarments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.dtasettar23cc12_imagenet21k_recap_hq_bucketed
cc12_imagenet21k_recap_hq_bucketed
Title: cc12_imagenet21k_recap_hq_bucketed
Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have
been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral.
To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.imagenet_22k_512_bucketable
ImageNet-22k 512-Bucketable Captioned Subset
This dataset is a pre-bucketed, captioned subset of timm/imagenet-22k-wds.
It is intended for text-to-image training and similar workflows that want images already grouped into aspect-ratio buckets near a 512-base training resolution. Images were kept only if they could fit one of the target buckets without upsampling after deterministic resize and crop.
Summary
Source: timm/imagenet-22k-wds (fall11 ImageNet-22k WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/imagenet_22k_512_bucketable.ArTbg_photo_concepts_bucketed_512
bg_photo_concepts_bucketed_512
Title: bg_photo_concepts_bucketed_512
Description: A recaptioned, self contained, bucketed and ready to train with version of https://huggingface.co/datasets/bghira/photo-concept-bucket, exported at 512^2 ish resolution buckets.
I lost the tracking data of which version of Gemini this was captioned with, likely 2.0 flash or 2.5 flash. The captions are on the long and datailed side and sometimes slightly redundant, but overall high quality.… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/bg_photo_concepts_bucketed_512.LAION_Aesthetics_1024_bucketed_512
LAION Aesthetics 1024 Bucketed 512 Captioned
This is a captioned bucketed-shards export of images from limingcv/LAION_Aesthetics_1024.
Images were filtered and resized/cropped into SDXL-style aspect-ratio buckets at a 512 base resolution, without upsampling. The export contains 382,144 images across 397 uncompressed WebDataset-style tar shards.
The .txt files now contain model-generated captions, not the original LAION web-scrape alt text or surrounding page text. Captions were… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_512.mjnj_tarArmchairclassical_paintings_bucketed_1024
Classical Paintings Captioned
A curated dataset of 7,131 classical paintings by 42 artists spanning the Baroque period through the 19th century, each with a descriptive plain-language caption (100--150 words). Intended for fine-tuning text-to-image models.
Artists (42)
Aelbert Cuyp, Albert Bierstadt, Anders Zorn, Anthony van Dyck, Artemisia Gentileschi, Caravaggio, Diego Velazquez, Frans Hals, Frederic Edwin Church, Georges de La Tour, Gerard ter Borch, Gerrit Dou… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/classical_paintings_bucketed_1024.
