datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot.
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.Nuplan-OccupancyCS-Arxiv-PDFs-08-25T2I-ImageNet-NormalCounterStrike-1K-360-wds
CounterStrike-1K — 360p WebDataset shards
This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets.
360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code.
Quickstart
Start a fresh uv project and add the loader:
mkdir cs1k-demo && cd cs1k-demo
uv init
uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.sdxl_images_easy_prompts-artists-seed1ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.waifu-preprocessed-datasetart-museums-pd-440k
Art Museums PD 440K
Summary
This is a dataset to train text-to-image or any text and image multimodal models with minimized copyright/licensing concerns.
All images and texts in this dataset are orignally shared under CC0 or public domain, and no pretrained models or any AI models are used to build this dataset except for our ElanMT model to translate English captions to Japanese.
ElanMT model is trained solely on licensed corpus.
Data sources
Images and… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/art-museums-pd-440k.computer-agent-arena
Computer Agent Arena: Evaluating Computer-Use Agents via Crowdsourcing from Real Users
Dataset Description
Computer Agent Arena is a comprehensive evaluation platform for multi-modal AI agents, particularly focusing on computer use and GUI interaction tasks. This dataset contains real interaction trajectories from various state-of-the-art AI agents performing complex computer tasks in controlled environments.
The dataset includes:
4,641 agent trajectories across diverse… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/computer-agent-arena.T2I-ImageNet-CutMixinfinigen-articulated
Infinigen-Articulated Assets
Formerly Infinigen-Sim
Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.Lowlight-Smartphone-Dataset
[WACV'26] Low-light Smartphone Dataset (LSD)
This is the official dataset proposed in our paper titled "Illuminating Darkness: Learning to Enhance Low-light Images In-the-Wild"
📄 Paper: arXiv💻 Code: GitHub - LSD-TFFormer
Overview
We introduce LSD, the largest in-the-wild Single-Shot Low-Light Image Enhancement (SLLIE) dataset to date.
Dataset Structure
This repository contains the following training data files:
patch_DLL_gtPatch.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/ARM4588/Lowlight-Smartphone-Dataset.VoT-video-latent-archiveffhq_captioned_1024
ffhq_captioned_1024
A captioned bucketed-shards export of gaunernst/ffhq-1024-wds.
This export contains 70,000 square face and portrait images from FFHQ, stored as
JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded
from the original dataset, deterministically converted to RGB, and re-encoded as
high-quality JPEG (quality=95, adaptive subsampling). Captions were generated
with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.HFGaussiansdxl_images_sb_prompts-multi_artist-seed1music-fingerprint-dataset
Neural Audio Fingerprint Dataset
(c) 2021 by Sungkyun Chang
https://github.com/mimbres/neural-audio-fp
This dataset includes all music sources, background noise and impulse-reponses
(IR) samples that have been used in the work ["Neural Audio Fingerprint for
High-specific Audio Retrieval based on Contrastive Learning"]
(https://arxiv.org/abs/2010.11910).
Format:
16-bit PCM Mono WAV, Sampling rate 8000 Hz
Description:
/
fingerprint_dataset_icassp2021/… See the full description on the dataset page: https://huggingface.co/datasets/arch-raven/music-fingerprint-dataset.physical-ai-bench-artifactsCowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata… See the full description on the dataset page: https://huggingface.co/datasets/bbrangeo/Cowpea-Architecture-XML.dtasettarsphere-encoder-fid-artifacts
Sphere Encoder FID Evaluation Artifacts
This repository contains the evaluation artifacts for the paper Image Generation with a Sphere Encoder.
Project Page | GitHub Repository
These artifacts include data statistic files (fid_stats) and reference images (fid_refs) used to calculate Fréchet Inception Distance (FID) for generative models across several datasets, including CIFAR-10, ImageNet, Animal Faces, and Oxford Flowers.
Workspace Setup
Download the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/sphere-encoder-fid-artifacts.12HZ-SegmentationGarments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.dtasettar23
