datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.MLS-Bench-Tasks
MLS-Bench Tasks
MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.Minecraft-Skins-20M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 19,973,928 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
id: A randomly generated UUID for each skin entry. These UUIDs are not linked to any external APIs or services (such as Mojang's player UUIDs) and serve solely as… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/Minecraft-Skins-20M.Minecraft-Skins-Captioned-1M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical.
image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.rocky_mountain_snowpack
Rocky Mountain Snowpack Dataset
The Rocky Mountain Snowpack dataset contains 4,040 preprocessed samples of snowpack imagery collected in the Colorado Rocky Mountains across the 2024–2025 and 2025–2026 winter seasons, from 7 snowpits dug between January 2025 and February 2026.Each sample segment of snow includes three types of images:
Magnified crystal images (close-up snow snow crystal profile photography)
Snowpack profile images (non-magnified snow crystal profiles… See the full description on the dataset page: https://huggingface.co/datasets/RMDig/rocky_mountain_snowpack.UrbanPersona-120K-Interpretive
UrbanPersona-120K-Interpretive
Annotation corpora and analysis outputs for "Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation" (EMNLP 2026 Workshop Pandora). Two
multimodal LLMs, Qwen3-VL-8B and Gemma4 E4B, annotate the same 50 PerceptSent urban scenes as the
same 1,200 demographic personas at T = 0.1, 60,000 persona × image attempts per model and 120,000
in all, next to their no-persona ablations, a greedy T = 0 decoding… See the full description on the dataset page: https://huggingface.co/datasets/MInDS-lab-UTFPR/UrbanPersona-120K-Interpretive.OpenGameArt-Mixed-Licenses
Dataset Card for OpenGameArt-Mixed-Licenses
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.vsr-sample-500
VSR Sample 500
This repository is a derivative sample of the Visual Spatial Reasoning (VSR) dataset. It contains 500 records and 483 unique COCO images in one train split. It is not the complete VSR corpus and is not a replacement for the upstream dataset.
The records were sampled without replacement from the upstream random-train split with deterministic seed 20260905. The sample preserves the source fields and values; the image field points to the bundled local file at… See the full description on the dataset page: https://huggingface.co/datasets/maujim/vsr-sample-500.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting instances, which follows the COCO format… See the full description on the dataset page: https://huggingface.co/datasets/eelianafang/MOUNT-Cattle.UrbanPersona-60K
UrbanPersona-60K: persona-conditioned urban sentiment annotations
Every annotation produced for "Stable Behavior, Limited Variation: Persona Validity in LLM
Agents for Urban Sentiment Perception" (arXiv:2604.28048): 60,000 attempts in
which 1,200 demographically distinct LLM personas each judged the same 50 urban scenes, plus the
two no-persona ablations the paper measures them against, the seed profiles that produced the
personas, and the full analysis outputs.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/MInDS-lab-UTFPR/UrbanPersona-60K.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.cnn_muffins
CNN Muffins
A compact dog-versus-muffin image-classification dataset built around the
well-known visual confusion between Chihuahua faces and blueberry muffins.
Dataset structure
Split
Dogs
Muffins
Total
Train
319
161
480
Validation
36
18
54
Hard-16 benchmark
8
8
16
The hard-16 benchmark is isolated from train and validation. The JSONL files
use repository-relative image paths:
The benchmark labels follow the original 4x4 checkerboard layout… See the full description on the dataset page: https://huggingface.co/datasets/VatsaDev/cnn_muffins.Meme-Sanity
Dataset Card for Meme-Sanity
Meme-Sanity is an extended multimodal dataset designed to improve hate speech detection in memes through counterfactual data augmentation. It contains 2,479 neutralized memes generated by isolating and rewriting the hateful component (text or image) using a large language–vision model pipeline. The dataset helps reduce spurious correlations and supports more robust, trustworthy, and context-sensitive hate classification.
Please note that all examples in… See the full description on the dataset page: https://huggingface.co/datasets/sahajps/Meme-Sanity.MM-Bench-E-CommerceThis is the HuggingFace repository of the paper named MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding in WSDM 2026 (oral).
In this paper, we argue that generative Multimodal Large Language Models (MLLMs) hold significant potential for improving product representation learning.
We propose the first generative MLLM-based model named MOON for product representation learning.
Furthermore, we contruct and publish a large-scale real-world… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/MM-Bench-E-Commerce.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.MONITRS
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
Dataset Description
Paper: NeurIPS 2025 (Spotlight)
Contact: revankar@cs.cornell.edu
MONITRS contains ~10,000 FEMA disaster events with temporal Sentinel-2 satellite imagery, natural language captions from news articles, geotagged locations, and question-answer pairs for disaster monitoring research.
Supported Tasks
Event classification
Temporal grounding
Location grounding
Visual… See the full description on the dataset page: https://huggingface.co/datasets/ShreelekhaR/MONITRS.Light-RAG-Marketing-Assets-Agent
🖼️ Light RAG Marketing Assets Agent — Pre-ingested Data
Pre-ingested LightRAG knowledge graph and vector data from 420 marketing images
analyzed with Gemini Vision API (gemini-3.5-flash) and processed through GPT-4o
for entity extraction and relationship mapping.
GitHub repo: 0xrphl/Light-RAG-Marketing-Assets-Agent
📊 Dataset Statistics
Metric
Value
Source images
420 (JPG/PNG/WebP)
Text chunks
2,095 (5 per image: core, visual, people/setting… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/Light-RAG-Marketing-Assets-Agent.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting… See the full description on the dataset page: https://huggingface.co/datasets/y1665065879/MOUNT-Cattle.minecraft-skins-1.1m-deduped-64x64-2.0
Minecraft Skins 1.1M Deduped (64x64 Edition) 2.0!
Minecraft Skins 1.1M Deduped 1.5 but it's tagged.
Format is just a 76.5 MB JSONL file and a 6.1 MB zipped JSONL file (as a JSONZ file)
Tools used
PIL Image (Python) and Google Colab (T4 GPU tier, but it didn't use the GPU at all!)
How it was made
Loaded Minecraft Skins 1.1M Deduped 1.5,
Tagged using a simple system where it looks for colors and complexity,
Output is given in a 6.1 MB ZIP archive or a 76.5 MB… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/minecraft-skins-1.1m-deduped-64x64-2.0.MUMU-Eval-6000
MUMU Eval 6000
This repository contains the 6,000-image source-data evaluation set used for
the Florence-2 and LFM2.5-VL-450M baselines in the MUMU evaluation repository.
It is an independently prepared research split, not an official MUMU Challenge
release.
Splits
Split
Images
Ground truth in manifest
validation
1,000
Yes
test
5,000
Yes
The split contains 2,001 Task A samples, 2,000 Task B samples, and 1,999 Task C
samples. All 6,000 image… See the full description on the dataset page: https://huggingface.co/datasets/JinyuLiu/MUMU-Eval-6000.movies_CLIP_ViT-L14
🎬 Movie Frame & Caption Dataset
📖 Introduction
This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted.
This dataset can be used for tasks such as:
Video understanding
Multimodal learning (image + text)
Image captioning
Vision-language retrieval
📂 Data… See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting… See the full description on the dataset page: https://huggingface.co/datasets/chenziyue-cattle/MOUNT-Cattle.MaskImageNet
MaskImageNet 460K Dataset Card
Dataset details
please checkout our paper
Paper or resources for more information:
[Project] [Paper] [Code]
CXR-CounterFact
CXR-CounterFact (CCF) Dataset
We are pioneers in introducing counterfactual cause into reinforced custom-tuning of MLLMs, we are deeply aware of the scarcity of counterfactual CoT in downstream tasks, especially in the highly professional medical field. Thus, our aspiration is for the model to adeptly acclimate to the concept drift by itself, acquiring abundant knowledge with more and more data, but not exhibiting bias.
In this context, a more realistic training dataset for… See the full description on the dataset page: https://huggingface.co/datasets/MiaoMiaoYang/CXR-CounterFact.mfg010-sample
MFG-010 — Manufacturing Defects Dataset (Sample)
A schema-identical preview of MFG-010, the XpertSystems.ai synthetic
defect events with visual-inspection ML metadata dataset for AOI
(Automated Optical Inspection) ML training, FMEA RPN modeling,
Ishikawa root cause classification, CAPA workflow simulation, and
defect-cohort quality engineering research. The full product covers
10,000-100,000 records. This sample is HF-sized at 3,000 records.
Built by XpertSystems.ai — Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/xpertsystems/mfg010-sample.ThyroidOpenDB
OpenThyroidDB
A multicenter, multitask thyroid ultrasound data resource supporting the ThyroidXAgent evidence workflow.
OpenThyroidDB integrates curated public thyroid ultrasound resources and institutionally governed clinical cohorts into a full-spectrum resource for segmentation, benign–malignant classification, report generation, and malignant-lesion stratification across diverse scanners and acquisition settings.
This Hugging Face repository (MedXAgent/OpenThyroidDB) is the… See the full description on the dataset page: https://huggingface.co/datasets/MedXAgent/ThyroidOpenDB.micah-boswell-photography
Micah Boswell photography on Unsplash
Metadata for the most viewed photographs by Micah Boswell (@micahboswell on Unsplash): Dallas at night, neon, wet streets, doors in Peru and Cuba. As of 2026-09-21 the full portfolio is 104 photographs with 53,380,492 views and 362,001 downloads, which places him in the top 10 percent of Unsplash contributors.
photos.jsonl holds one record per photograph: title, Unsplash page, hotlinkable image URL (Unsplash CDN), dimensions, date, likes… See the full description on the dataset page: https://huggingface.co/datasets/socraticstatic/micah-boswell-photography.aiconf-butterfly-learn-for-model-to-markupMaster dataset for the next markup stage of butterfly detection.
Sampling rules:
500 butterfly images
500 negative images
negative classes are distributed uniformly across: bee, beetle, flower, shrub
Files:
master_dataset.json
master_dataset.tsv
Columns:
photo_id
image
entity
selection_group
target_label
needs_bbox_markup
source_split
hard (if present in source)
photo_url (if present in source)
taxon (if present in source)
Source dataset: vsevolod-nv/aiconf-butterfly-detection-all… See the full description on the dataset page: https://huggingface.co/datasets/vsevolod-nv/aiconf-butterfly-learn-for-model-to-markup.oceanguard-marine-debris-eval-1000
OceanGuard AI — Marine Debris Evaluation Hold-out (annotations only)
The held-out evaluation split used to report the LoRA adapter delta in the
OceanGuard AI Kaggle Gemma 4 Good Hackathon submission
(Global Resilience track + Unsloth bonus track).
Important — this repository contains only the annotations and metadata.
The 1 000 underwater / coastal images are not redistributed here. They
come from three pre-existing third-party datasets, each with its own
license. Reviewers and… See the full description on the dataset page: https://huggingface.co/datasets/asferrer/oceanguard-marine-debris-eval-1000.Pumpkin-Maturity-Grading-Dataset
Pumpkin Maturity Grading Dataset
The current agricultural industry faces challenges in quality control of crops, especially in judging the maturity of pumpkins. Traditional methods often rely on manual identification, which is inefficient and prone to errors. Although existing image recognition technology has made some progress, there is a lack of high-quality datasets specifically targeted at pumpkin maturity. This dataset aims to improve the accuracy of maturity assessment for… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pumpkin-Maturity-Grading-Dataset.
