datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.Icarus-dataset
Icarus
A unified multi-modal curriculum dataset for evolutionary neural architecture search. Every row is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. Support is non-empty for every task. Encoders read the Field descriptor (axes, value_type, n_classes, value_range, mask); mask is True where a value is padding/ignored. meta.class_names, when… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/Icarus-dataset.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.nesteo-prototype
NestEO: Modular and Hierarchical EO Dataset Framework
NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO.
Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.handwritten-digit-dataset
Handwritten Digit Dataset
This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability.
Dataset Details
The images are pre-processed to match the standard machine learning format for digit recognition:
Dimensions: 28x28 pixels.
Format: Grayscale (single channel).
Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.rlbenchfail_val_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.Ransomware_PE_Header_Feature_Dataset
Dataset Card for Ransomware PE Header Feature Dataset
Dataset Description
Dataset Summary
This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header.
Supported Tasks
Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.trade_vision_dataset
TradeVision: Hierarchical Physical Business & Multimodal Retail Provenance Dataset
This dataset is continuously seeded from OpenStreetMap, matched to Google Place IDs, harvested for temporal store photos, and enriched with zero-shot computer vision using Hugging Face Hub native pipelines.
Dataset Structure
The dataset is partitioned into three relational subsets loadable via Hugging Face datasets:
from datasets import load_dataset
# 1. Load Canonical Businesses… See the full description on the dataset page: https://huggingface.co/datasets/drksci/trade_vision_dataset.Guardian-FailCoT-OOD-datasets
Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks
This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026):
UR5-Fail — our newly collected three-view real-robot benchmark.
RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023).
RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.tripmatch-ai-dataset
TripMatch AI Dataset
A reproducible multimodal dataset for the TripMatch AI Final Project. It contains
10,000 synthetic text trip plans with a raw idea generated for every row by the
pretrained Hugging Face model google/flan-t5-small, plus 5,000 real street-view images
retained as extra multimodal work. The two configurations are separate so Dataset
Viewer can load each schema correctly.
Dataset statistics
Configuration
Rows
Main fields
Intended task… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-dataset.bdv2fail_train_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.ur5fail_train_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.bdv2fail_val_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_val_dataset.forestllava-dataset
Forest-LLaVA Multimodal Tree-Species Dataset
Forest-LLaVA is a multimodal remote-sensing dataset for tree-species
recognition and structured vision-language research. Each record is indexed by
a numeric sample_id from the US subset of GlobalGeoTree and is linked to a
60 m × 60 m patch from NAIP Optical, Sentinel-2 MSI and Sentinel-1 SAR data,
four-level taxonomic labels and geographic/environmental records.
The repository contains the complete image archives for the… See the full description on the dataset page: https://huggingface.co/datasets/minute1028/forestllava-dataset.ur5fail_val_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_val_dataset.bdv2fail_test_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.house_kg_full_dataset_frames
house.kg — Kyrgyzstan Real Estate, over time
Sale and rental listings scraped from house.kg, the largest
real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are
English; values are kept in the original language (Russian), exactly as the site
renders them.
Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions.
Subsets
subset
rows
description
listings
25,264
one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.dataset-registry-v0
Kámárí Dataset Registry (v0)
Provenance for the Kámárí datasets used to train the age model and build the
Kámárí-Safe Open benchmark. It holds
manifests (paths, hashes, labels, skin-tone band, quality), the dataset source and licence
registry, EDA reports, and the data-quality report. No raw images are redistributed.
Built from open, license-checked face datasets with an auto label-quality gate, MTCNN face crops, and
ITA skin-tone banding. v0 composition: 825,129 candidate rows… See the full description on the dataset page: https://huggingface.co/datasets/Shinzmann/dataset-registry-v0.violence-nonviolence-dataset
Violence vs Non-Violence Dataset
This dataset contains annotated interaction data for detecting violent vs non-violent human interactions.The data is extracted from video frames and includes bounding boxes, pose keypoints, motion features, and violence indicators for pairs of interacting persons.
Files
violence_data.csv → Frames labeled as violent interactions
non_violence_data.csv → Frames labeled as non-violent interactions
Each CSV contains structured features at… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/violence-nonviolence-dataset.InSpect
InSpect
InSpect is a curated natural history collection dataset for visual insect specimen understanding. It contains digitized insect specimen images with aligned crops, hierarchical taxonomy, label-derived structured metadata, and fine-grained anatomical part annotations.
Files
specimen_benchmark_metadata.csv: main metadata table. Each row corresponds to one specimen image and includes split information, taxonomic labels, image/crop paths, and label-derived… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset/InSpect.deepfake-detection-dataset-v2
Deepfake Detection Dataset V2
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images
CAM visualization images
CAM overlay images
Comparison images
Labels (real/fake)
Confidence scores
Image captions
Technical and non-technical… See the full description on the dataset page: https://huggingface.co/datasets/saakshigupta/deepfake-detection-dataset-v2.amt-airframe-handbook-dataset
AMT Airframe Handbook Dataset
A comprehensive dataset extracted from the FAA Aviation Maintenance Technician (AMT) Airframe Handbook, containing text content and rendered page images suitable for training vision-language models.

Overview
This dataset was created using the doc-parser-engine - a production-grade document parsing engine with HuggingFace integration. The source document is the FAA… See the full description on the dataset page: https://huggingface.co/datasets/Remixonwin/amt-airframe-handbook-dataset.cutclean-datasets
CutClean: balanced datasets
Data accompanying:
Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise, Enzo Tartaglione.
CutClean: Neural Network Pruning for Privacy-Preserving Inference.
Pattern Recognition — Proceedings of the 28th International Conference on Pattern
Recognition (ICPR 2026), Lyon, France. Lecture Notes in Computer Science, Springer
Nature Switzerland, pp. 450–465.
doi:10.1007/978-3-032-31452-9_30
Code: https://github.com/MaglioloLeonardo/CutClean
The… See the full description on the dataset page: https://huggingface.co/datasets/imDalton/cutclean-datasets.jwst-quality-analysis-dataset
JWST Quality Analysis Dataset
Overview
This dataset contains comprehensive quality analysis for 2,709 JWST (James Webb Space Telescope) NIRCam images from the MAST archive. Each image has been automatically analyzed for quality metrics, artifact detection, and noise characteristics.
Dataset Information
Size: 2,709 images
Format: JSONL (JSON Lines)
Source: JWST NIRCam observations from MAST
Targets: M16, NGC 3132, NGC 3324, SMACS 0723, Stephan's Quintet… See the full description on the dataset page: https://huggingface.co/datasets/norbertm/jwst-quality-analysis-dataset.minicar-dataset
🏎️ MiniCar Autonomous Driving Dataset
自動運転ミニカー用のトレーニングデータセット
概要
このデータセットには以下が含まれます:
カメラ画像
センサーデータ(IMU等)
アノテーション(ステアリング角度、スロットル)
データ構造
minicar-dataset/
├── train/
│ ├── images/ # カメラ画像 (JPG/PNG)
│ ├── sensors/ # センサーデータ (CSV)
│ └── annotations.csv # ラベルデータ
├── test/
│ └── ...
└── README.md
使い方
from datasets import load_dataset
dataset = load_dataset("Romihi50/minicar-dataset")
# トレーニングデータ
for sample in… See the full description on the dataset page: https://huggingface.co/datasets/Romihi50/minicar-dataset.oil_WTI_image_dataset_v1
FINANCE_PROJECT – Plot Image Dataset (Parquet Shards)
This dataset contains Matplotlib-rendered PNG plot images generated from sliding windows over financial time series features. Each row is one image sample plus 1–8 day trend labels derived from CL=F_Close, where labels are defined at the last index of the plotted window.
Repository Layout
The dataset is stored as Hugging Face–style sharded Parquet files:
train-00000-of-00210.parquet
train-00001-of-00210.parquet
…… See the full description on the dataset page: https://huggingface.co/datasets/mrzdu/oil_WTI_image_dataset_v1.
