datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.rlbenchfail_val_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.rule34xyz
Dataset Card for rule34.xyz
Dataset Summary
This dataset contains information about image files from rule34.xyz, a booru-style imageboard. The dataset includes metadata for 590,983 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset metadata is… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34xyz.rocky_mountain_snowpack
Rocky Mountain Snowpack Dataset
The Rocky Mountain Snowpack dataset contains 4,040 preprocessed samples of snowpack imagery collected in the Colorado Rocky Mountains across the 2024–2025 and 2025–2026 winter seasons, from 7 snowpits dug between January 2025 and February 2026.Each sample segment of snow includes three types of images:
Magnified crystal images (close-up snow snow crystal profile photography)
Snowpack profile images (non-magnified snow crystal profiles… See the full description on the dataset page: https://huggingface.co/datasets/RMDig/rocky_mountain_snowpack.REOBench
Folder/File Descriptions
AID/AID_train.zip: Contains all AID images in the training set.
AID/AID_test.zip: Contains images in the test set under perturbation.
AID/AID_JSON/: Contains JSON files for zero-shot evaluation of LLM-based models.
Potsdam/Potsdam_Images_trian.zip: Contains all Potsdam images in the training set.
Potsdam/Potsdam_Anns_trian.zip: Contains annotations for images in the training set.
Potsdam/Potsdam_Images_test.zip: Contains Potsdam test images under… See the full description on the dataset page: https://huggingface.co/datasets/xiang709/REOBench.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.rule34lol-images-part2
Dataset Card for rule34lol-images-part2
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 77,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files (except the last archive). This is Part 2 of 2 for the complete rule34lol-images dataset. Part 1 can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part2.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.radread-public-results
RadRead — public results
Rollout-level results for RadRead, a benchmark of frontier models reading 150
radiographs. Every row is one graded model read: 5 saved rollouts per study
per model, scored by a deterministic grader (no judge model).
A read passes only when every required checklist finding, lesion box (the grader's
IoU / centre / containment test), lexical diagnosis check and action-set membership
check match the reference rubric. No partial credit inside a study;… See the full description on the dataset page: https://huggingface.co/datasets/tirandazdylan/radread-public-results.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.Light-RAG-Marketing-Assets-Agent
🖼️ Light RAG Marketing Assets Agent — Pre-ingested Data
Pre-ingested LightRAG knowledge graph and vector data from 420 marketing images
analyzed with Gemini Vision API (gemini-3.5-flash) and processed through GPT-4o
for entity extraction and relationship mapping.
GitHub repo: 0xrphl/Light-RAG-Marketing-Assets-Agent
📊 Dataset Statistics
Metric
Value
Source images
420 (JPG/PNG/WebP)
Text chunks
2,095 (5 per image: core, visual, people/setting… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/Light-RAG-Marketing-Assets-Agent.Latent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.rule34lol-images-part1
Dataset Card for rule34lol-images-part1
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 196,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. This is Part 1 of 2 for the complete rule34lol-images dataset. Part 2 can be found here.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part1.ReImageNet
ReImageNet
ReImageNet is a complete multilabel reannotation with localization of the ImageNet-1K validation
set (ILSVRC2012). A team of 7 trained in-house annotators reviewed all
50,000 validation images through an iterative annotation process, producing
per-image bounding boxes with class labels and annotation attributes,
correcting and extending the original single-label ground truth. Class
names and definitions were revised where the original WordNet-based names
no… See the full description on the dataset page: https://huggingface.co/datasets/vrg-prague/ReImageNet.nwpu-resisc45
NWPU-RESISC45
Train/val split manifests extracted from GeoChat_Instruct
for fine-tuning vision-language models on satellite scene classification.
Classes
45 scene categories (e.g. airplane, baseball diamond, beach, bridge, …)
Splits
Split
Samples
train
~25 200
val
~4 725
Format
Each JSON file is a list of objects with fields:
image (relative path), question, answer (class label).
lanternfly_research_dataset
Lantern Fly Research Dataset
This dataset contains human-verified spotted lanternfly sightings collected through the Lantern Fly Tracker app. Each entry includes high-quality photos, precise geolocation data, and comprehensive metadata for ecological research.
🎯 Purpose
This dataset supports:
Ecological research on spotted lanternfly distribution and spread patterns
Machine learning model training with verified, high-quality data
Temporal and spatial analysis of… See the full description on the dataset page: https://huggingface.co/datasets/rlogh/lanternfly_research_dataset.Raspberry-Variety-Classification-Dataset
Raspberry Variety Classification Dataset
The current agricultural industry faces challenges in managing the diversity of crop types, especially in raspberry cultivation and variety identification. Existing solutions often rely on manual identification, which is inefficient and prone to errors. This dataset aims to address the issue of low accuracy in variety classification by providing high-quality raspberry images, meeting the needs of intelligent agriculture. The dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Raspberry-Variety-Classification-Dataset.Rose
Dataset Card for "monetjoe/cv_backbones"
This repository consolidates the collection of backbone networks for pre-trained computer vision models available on the PyTorch official website. It mainly includes various Convolutional Neural Networks (CNNs) and Vision Transformer models pre-trained on the ImageNet1K dataset. The entire collection is divided into two subsets, V1 and V2, encompassing multiple classic and advanced versions of visual models. These pre-trained backbone… See the full description on the dataset page: https://huggingface.co/datasets/Limitless063/Rose.aidm-dogs-vs-cats-results
aidm-dogs-vs-cats-results
The experiment record of a dogs-vs-cats image-classification study, with CIFAR-10 and CIFAR-10-LT transfer and class-imbalance ablations. This repo holds the run registry, the splits, the report tables and figures, and the per-run predicted probabilities. It holds no images and no model weights; the checkpoints are in the companion model repo.
Generated by scripts/90_publish_hf.py on 2026-09-22 19:17 UTC. Every count, fingerprint and metric below was… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/aidm-dogs-vs-cats-results.Papaya-Tree-Recognition-Dataset
Papaya Tree Recognition Dataset
The current agricultural sector faces issues of low efficiency in crop recognition and management, especially against the backdrop of the growing development of smart agriculture. Traditional manual recognition methods are unable to meet the rapidly changing needs. Existing solutions often rely on image data from a single environment, lacking diversity and universality, which leads to poor model generalization. The Papaya Tree Recognition Dataset aims… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Papaya-Tree-Recognition-Dataset.Field-Pumpkin-Recognition-Dataset
Field Pumpkin Recognition Dataset
The current agriculture industry faces challenges such as low efficiency in crop recognition and difficulty in pest monitoring. Existing solutions often rely on manual identification, which is inefficient and prone to errors. This dataset aims to support the training of machine learning models by providing high-quality pumpkin image data, enhancing the accuracy and speed of pumpkin recognition. Data is primarily collected in the field using… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Field-Pumpkin-Recognition-Dataset.robot2robotIdentification
Robot2RobotIdentification
A dataset for machine-to-machine visual awareness.
Supported by A19Lab, Inc.
Dataset Description
Robot2RobotIdentification is a vision dataset designed to help drones, UGVs, and autonomous robots detect and recognize each other in real-world environments.
As autonomous machines become more common in skies, streets, and industrial spaces, reliable machine-to-machine perception is essential for safety, coordination, and navigation. This… See the full description on the dataset page: https://huggingface.co/datasets/A19Lab/robot2robotIdentification.
