datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/sunghong/CADS-dataset.Sparkle
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou
📦 Dataset
Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper.
The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.FISH_spots
FISH_spots Dataset
The manually verified in situ hybridization fluorescence images and point coordinate dataset.
This dataset contains images and annotations for the task of single-molecule fluorescence in situ hybridization (FISH) spot detection, supporting 2D, 3D, and simulated noisy data. The structure is designed for deep learning model development, training, and evaluation.
Directory Structure
FISH_spots/
├── 2d/
│ ├── csv/
│ ├── image/
│ ├── image_raw/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/GangCaoLab/FISH_spots.skin-cancer-ham10000-datasetsherlock
Sherlock
Naturalistic fMRI dataset: 16 subjects watched ~50 minutes of Sherlock across
two scanning runs (Part1, Part2) and then verbally recalled the narrative in
the scanner. TR = 1.5 s.
This repo mirrors the fmriprep-preprocessed dataset originally distributed via
DataLad at https://gin.g-node.org/ljchang/Sherlock. fmriprep version 1.2.6-1.
Layout
derivatives/fmriprep/sub-XX/
anat/ func/ figures/ log/
onsets/
Sherlock_Crop_Onsets.csv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/sherlock.robocurate-synth100
synth100 — 100 generated clips for validating Pre-Contact Level Filtering
100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so
Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs
travels with the package, in the space it is consumed in.
Read section 1 before using this. The single most important fact about this data is not in the
file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.E2AM_ResNet50
E2AM Ablation Results: ResNet-50
Energy-aware training ablation study for ResNet-50 across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet.
Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8).
Status: 45 completed variants, 0 partial.
Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_ResNet50.plant-disease-trainvindr-cxr-testsetparadetox
ParaDetox: Text Detoxification with Parallel Data (English)
This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.AniGen-Sample-Dataset
AniGen Sample Data
This directory is a compact example subset of the AniGen training dataset.
What Is Included
10 examples
10 unique raw assets
Full cross-modal files for each example
A subset metadata.csv with 10 rows
The retained directory layout follows the core structure of the reference test set:
raw/
renders/
renders_cond/
skeleton/
voxels/
features/
metadata.csv
statistics.txt
latents/ (encoded by the trained slat auto-encoder)
ss_latents/ (encoded by the… See the full description on the dataset page: https://huggingface.co/datasets/VAST-AI/AniGen-Sample-Dataset.agiqa-3kDataset in paper [IEEE TCSVT2023] [Agiqa-3k: An open database for ai-generated image quality assessment](https://huggingface.co/papers/2306.04717)
Code: https://github.com/lcysyzxdxc/AGIQA-3k-Database
@ARTICLE{10262331,
author={Li, Chunyi and Zhang, Zicheng and Wu, Haoning and Sun, Wei and Min, Xiongkuo and Liu, Xiaohong and Zhai, Guangtao and Lin, Weisi},
journal={IEEE Transactions on Circuits and Systems for Video Technology},
title={AGIQA-3K: An Open Database for AI-Generated Image… See the full description on the dataset page: https://huggingface.co/datasets/strawhat/agiqa-3k.PlantInquiryVQA
PlantInquiryVQA — Thinking Like a Botanist
Benchmark and framework for multi-turn, intent-driven visual question answering in plant pathology.
Accepted at ACL 2026 Findings.
Overview
PlantInquiryVQA formalises diagnostic reasoning in plant pathology as a Chain-of-Inquiry (CoI) — an ordered sequence of visually-grounded questions that adapts to the plant's severity and the expert's epistemic intent (Diagnosis / Prognosis / Management).
The benchmark evaluates… See the full description on the dataset page: https://huggingface.co/datasets/SyedNazmusSakib/PlantInquiryVQA.EditVerseBench
EditVerse
This repository contains the instruction-based video editing evaluation benchmark for EditVerseBench in paper "EditVerse: A Unified Framework for Editing and Generation via In-Context Learning".
Xuan Ju12, Tianyu Wang1, Yuqian Zhou1, He Zhang1, Qing Liu1, Nanxuan Zhao1, Zhifei Zhang1, Yijun Li1, Yuanhao Cai3, Shaoteng Liu1, Daniil Pakhomov1, Zhe Lin1, Soo Ye Kim1*, Qiang Xu2*
1Adobe Research 2The Chinese University of Hong Kong 3Johns Hopkins University *Corresponding… See the full description on the dataset page: https://huggingface.co/datasets/sooyek/EditVerseBench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.JL1-CUP-2024-Second-Format
JL1 CUP 2024 — Second-track format for semantic change detection
Bi-temporal 256×256 RGB patches with per-pixel semantic maps at times T1/T2 and a binary change map, aligned with the data split described in the literature for the JL1 cropland change-detection benchmark (Second Track / JL1-Second style layout).
Source
Resource
URL
JL1 Mall contest information
contest page
JL1 data / resources portal
resrepo
Data are provided by the JL1 / Jilin-1 ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/JL1-CUP-2024-Second-Format.StereoMIS_processedchange_eye_face_head_person
Overview
This data set contains manually curated, high quality images that can be used to
train image editing AI models like
FLUX.1 Kontext
to be able to take an input image and a reference image to create a target
image that is looking like the input image but with one of those parts replaced:
eyes
face
head
person (input image cloths are kept)
person (reference image cloths are kept)
Typical prompts for this editing could then be:
Change the eyes, keeping the rest of the image… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/change_eye_face_head_person.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
Accepted to CVPRW 2026.
GitHub | CVF Open Access | arXiv | Model
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.mac-app-store-apps-metadata
Dataset Card for Macappstore Applications Metadata
📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications Metadata sourced by the public API.
Curated by: MacPaw Way Ltd.
Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.Minecraft-winter-shaders-Lite
❄️ Minecraft Winter Shaders (Lite)
10,000+ winter Minecraft screenshots with Photon Shaders.
For generative models, style transfer, and high-quality vision tasks.
Structure
Single folder: winter/
Format: 960x540 JPEG
License
Dataset: CC BY NC 4.0.Shader: Photon Shaders by SixthSurge (commercial use of screenshots explicitly allowed).See PHOTON_SHADERS_LICENSE.txt.
Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ShafinSI/Obstacle-Detection-Dataset-YOLO.SeismicX-Cont-mini
SeismicX-Cont Mini Two-Hour Subset
This folder is the compact, Zenodo-archived two-hour mini release for
SeismicX-Cont. It is designed for quick download, tutorial use, software smoke
tests, and checking that the HDF5, annotation, SQLite, dataloader, picker, and
validation workflow all fit together before using the full 14-day data product.
Zenodo record: https://zenodo.org/records/21331024
DOI: https://doi.org/10.5281/zenodo.21331024
Hugging Face record:… See the full description on the dataset page: https://huggingface.co/datasets/cangyeone/SeismicX-Cont-mini.
