datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imglsdolympiad-math-contest-llama3-78ktrellis500k-github-archives-7InteriorVerseCaptions of InteriorVerse's RGB images extracted with microsoft/Florence-2-large.
olympiad-math-stepwise-solutions-llama3-20kThe MATH dataset is a collection of 20,300 problems from AMC and AIME competitions covering algebra, number theory, geometry, and precalculus problems and solution sets.
Problems and solutions are formatted in LATEX.
Step-by-step solutions and insight sections have been added in order to use a chain of thought to clarify the problem and solution.
BoxFusionFormulaBank-28K
FormulaBank-28K
FormulaBank-28K is a deterministic procedural-audio corpus for audio representation pre-training. It contains 28,000 mono clips at 16 kHz, each exactly 10.24 seconds long.
Configuration
Version: C125-I224-v1
Formula classes: 125
Renderings per class: 224
Total clips: 28,000
Audio format: lossless 24-bit FLAC
Source: frozen AudioPG-Atomic-H7-C224-R0-Clean FormulaBank manifest
Each formula class specifies an acoustic rendering rule. Each rendering… See the full description on the dataset page: https://huggingface.co/datasets/KevinJustin/FormulaBank-28K.s1-dataset
S1 Dataset
Access to the dataset files is gated. Approved users must sign in to their own
Hugging Face account before downloading.
Upload status
Last updated: 2026-08-25 02:58 UTC
Expected: 27,577 HDF5 files
Available: 27,577 HDF5 files
Overall progress: 100.00%
Complete sequences: 83 / 83
Remaining: 0 files (approximately 0.00 GiB)
Sequence
Status
Available
Expected
Missing
Progress
-
Complete
-
-
0
100.00%
All sequences not shown in the table… See the full description on the dataset page: https://huggingface.co/datasets/kevinlad/s1-dataset.Who_and_When
Who&When: #1 Benchmark for MAS automated failure attribution.
184 annotated failure tasks collected from
Algorithm-generated agentic systems built using CaptainAgent,
Hand-crafted systems such as Magnetic-One.
Fine-grained annotations for each failure, including:
The failure-responsible agent (who failed),
The decisive error step (when the critical error occurred),
A natural language explanation of the failure.
The dataset covers a wide range of realistic multi-agent scenarios… See the full description on the dataset page: https://huggingface.co/datasets/Kevin355/Who_and_When.libero_sft_policiesgdpval-gpt5
GDPval with GPT-5 Execution Results
This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks.
🎯 Dataset Overview
This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables.
📊 Key Statistics
Total tasks: 220
Tasks with AI deliverables: 87 (39.5%)
Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5.kev-suitesCUHK-X_Small_Model_Track
CUHK-X — Small Model Track
Multimodal human action recognition (classification).
Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes).
Repository layout
.
├── Training/
│ ├── class_mapping.csv # action_id <-> action_name (40 classes)
│ └── data/
│ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip
│ → HAR/data/<modality>/<action>/<user>/<trial>/<files>
└── Testing/
├── data/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.gdpval-gpt5-fork
GDPval Fork Dataset with GPT-5 Results
🏆 A comprehensive evaluation dataset featuring GPT-5 execution results on real-world professional tasks
This is an enhanced fork of the original OpenAI GDPval dataset with complete GPT-5 execution results, including actual deliverable files created by the AI model.
📊 Dataset Overview
Metric
Value
Total Tasks
220
AI-Completed Tasks
87 (39.5%)
Deliverable Files
492+ professional documents
Occupations
44
Industry… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5-fork.nuscenes-qa-mini
NuScenes-QA-mini Dataset
TL;DR:
This dataset is used for multimodal question-answering tasks in autonomous driving scenarios. We created this dataset based on nuScenes-QA dataset for evaluation in our paper Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI. The samples are divided into day and night scenes.
scene
# train samples
# validation samples
day
2,229
2,229
night
659
659
Each sample contains… See the full description on the dataset page: https://huggingface.co/datasets/KevinNotSmile/nuscenes-qa-mini.complet4r_preprocessed_dynamicreplicadiffpackDiffPack is the bigcode/commitpack dataset except diff'd between the old and new data.
olympiad-math-contest-llama3-20k
AMC/AIME Mathematics Problem and Solution Dataset
Dataset Details
Dataset Name: AMC/AIME Mathematics Problem and Solution Dataset
Version: 1.0
Release Date: 2024-06-1
Authors: Kevin Amiri
Intended Use
Primary Use: The dataset is created and intended for research and an AI Mathematical Olympiad Kaggle competition.
Intended Users: Researchers in AI & mathematics or science.
Dataset Composition
Number of Examples: 20,300 problems and solution sets… See the full description on the dataset page: https://huggingface.co/datasets/kevin009/olympiad-math-contest-llama3-20k.ScreenSpot3d-front-code
3D-Front-Code
RoomScript v4 Blender object programs, room-layout renders, code-only wall
architecture, and asset Blender artifacts derived from 3D-FRONT scene evidence.
Contents
13,917 object assets (reference and v4/best) in data/assets/*.tar
21,202 rooms (v4/best and code-only v4_wall/best) in data/rooms/*.tar
searchable JSONL indexes under metadata/
Each tar contains multiple samples while preserving the original
data/front_object_code/by_asset/... or… See the full description on the dataset page: https://huggingface.co/datasets/KevinFan111/3d-front-code.SLFCog-ResearchByteDance_Synthetic_Videos
Dataset Name
CGI synthetic videos generated in paper "Synthetic Video Enhances Physical Fidelity in Video Synthesis" (https://simulation.seaweed.video/)
Dataset Overview
Number of samples: [uploading...]
Annotations: [tags, captions]
License: [apache-2.0]
Citation: @article{zhao2025synthetic,
title={Synthetic Video Enhances Physical Fidelity in Video Synthesis},
author={Zhao, Qi and Ni, Xingyu and Wang, Ziyu and Cheng, Feng and Yang, Ziyan and Jiang, Lu and Wang, Bohan}… See the full description on the dataset page: https://huggingface.co/datasets/kevinzzz8866/ByteDance_Synthetic_Videos.typebert
Dataset Card for "typebert"
More Information needed
hssd-annotations
hssd-annotations
English | 中文
A standalone, locally-stored, zero-dependency Python API to search and
retrieve HSSD assets and their full annotation set — for downstream scene
generation (e.g. SceneSmith) and articulation/clearance research.
Every annotation family is merged into one per-asset record keyed by the HSSD
asset id. The library ships in post-replacement form: each asset is linked
to its articulated realization (official HSSD articulated, or a PartNet-Mobility… See the full description on the dataset page: https://huggingface.co/datasets/P-Kevin/hssd-annotations.ManyRefactors4Cstereo4d-lefteye-perspective
Dataset Summary
This dataset contains the left-eye rectified perspective views from the Stereo4D dataset (Paper). Each video is generated using the rectify.py script, which processes VR180 stereo videos to produce 512×512 video clips with a 60° field of view perspective camera. This dataset is intended to be used alongside the Stereo4D dataset annotations which can be found here.
This dataset is provided as-is for non-commercial research purposes only.
Download
git clone… See the full description on the dataset page: https://huggingface.co/datasets/KevinMathew/stereo4d-lefteye-perspective.llm-graph-poisoning-data
Generation-Time Poisoning of LLM-Generated Social Networks
This dataset contains synthetic personas, LLM-generated social graphs, cached
text embeddings, and evaluation metrics for clean generation and three
generation-time attack families. All names and profiles are synthetic and do
not represent real people.
Dataset variants
Variant
Nodes
Generator
Graph seeds per condition
Attack rates
p50
50
Qwen3-Max
10
10%, 20%, 30%, 40%, 50%
p200
200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.t2i-finegrain
t2i-finegrain Dataset
This dataset evaluates text-to-image (T2I) diffusion models using a benchmark of prompts designed to elicit specific failure modes. Human labels allow for T2I benchmarking evaluations.
Contents
10,587 total image–metadata entries
750+ prompts
11 failure mode categories
27 specific failure modes
14 total models evaluated:
5 models (with human ground truths):
SD3-XL
SD3-M
SD3.5-Large
SD3.5-Medium
Flux
9 models:
Flux-Kontext – 760 images… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/t2i-finegrain.
