datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WxC-Bench
Dataset Card for WxC-Bench
WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks.
Dataset Details
WxC-Bench contains datasets for six key tasks:
Nonlocal Parameterization of Gravity Wave Momentum Flux
Prediction of Aviation Turbulence
Identifying Weather Analogs
Generation of Natural Language Weather Forecasts
Long-Term Precipitation Forecasting
Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.pane-binding-functions-attributionIMPACT
IMPACT v1.1
IMPACT is a synchronized five-view RGB-D dataset and benchmark for multi-granularity human procedural action understanding in industrial assembly. It contains 112 trials from 13 participants and 39.5 video hours across one egocentric and four exocentric views.
Project page
Benchmark code and task protocols
Google Drive mirror
Release Update
July 2026, v1.1. All 560 TAS-B annotation files were revalidated, and 117 of 560 RGB videos (20.9%) were… See the full description on the dataset page: https://huggingface.co/datasets/KratosWen/IMPACT.paper-impact-dataopm-ehri-datadesign-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.notch-beam-2d-impact
NotchBeam2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/notch-beam-2d-impact",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/finebooks/bhl-impact-gt.IMPACTImpactMesh-Fire
ImpactMesh-Fire
ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab.
It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS.
The technical report is released soon. You find the flood subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood.
Features
Multimodal: SAR, optical, DEM
Multitemporal:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire.sentry-impact-risk
NASA Sentry: Earth Impact Risk Assessment
Credit: NASA/Johns Hopkins APL
Part of a dataset collection on Hugging Face.
Dataset description
Near-Earth objects with non-zero Earth impact probability from NASA JPL Sentry system.
The Sentry system, operated by NASA's Center for Near-Earth Object Studies (CNEOS) at the Jet Propulsion Laboratory, continuously monitors the most current asteroid catalog for possibilities of future Earth impact. Objects are listed… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/sentry-impact-risk.Genshin-Impact-Novel-Video
taylor-impact-2d
Taylor2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/taylor-impact-2d",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/taylor-impact-2d", repo_type="dataset")… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/taylor-impact-2d.bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/bhl-impact-gt.ImpactMesh-Flood
ImpactMesh-Flood
ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab.
It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS.
The technical report is released soon. You find the wildfire subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire.
Features
Multimodal: SAR, optical, DEM… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood.genshin-impact-voices
Genshin Impact — Voice Lines (Multi-Language)
An archive of character voice data extracted from Genshin Impact (原神), repackaged as Parquet shards per audio language.
Dataset Summary
Field
Value
Game
Genshin Impact (原神)
Publisher
HoYoverse / miHoYo Co., Ltd.
Languages
中文 (zh), 日本語 (ja), English (en), 한국어 (ko)
Game version
6.3
Source format
WAV + sidecar transcripts (.lab / .txt)
Distribution format
Apache Parquet (zstd), ~500 MiB audio per shard… See the full description on the dataset page: https://huggingface.co/datasets/ultemica/genshin-impact-voices.reward-projection-goal-generalisation-vlmsynthrad2023-impact-registration
🧭 SynthRAD2023 IMPACT Registrations (BSpline Transforms)
This repository provides Elastix B-spline transformation parameter files generated using the IMPACT method on the SynthRAD2023 dataset.
Each file corresponds to a non-rigid registration between a reference CT and another modality (MRI or CBCT), aligned into CT space using Elastix with the IMPACT similarity metric.
Task 1: 317 transforms (43 excluded cases)
Task 2: 289 transforms (69 excluded cases)
🚀 Overview… See the full description on the dataset page: https://huggingface.co/datasets/VBoussot/synthrad2023-impact-registration.denoising-impact-evaluation-dataset
Denoising Impact Evaluation Dataset
Dataset Description
The ekacare/denoising-impact-evaluation-dataset is a comprehensive benchmark dataset designed to evaluate the effects of speech enhancement on automatic speech recognition (ASR) systems in medical speech contexts. It includes paired noisy and denoised audio subsets under controlled acoustic conditions to support systematic analysis of denoising performance.
Source Data
Base Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/denoising-impact-evaluation-dataset.python4-leetcode-eft
Python4 LeetCode AFT (v2)
Execution-validated behavioral fine-tuning demonstrations for a controlled
study of Python 4, a fictional programming language executed by the Boa
interpreter. Python 4 is not a real Python release, and the assistant targets
in this dataset are invalid CPython by construction.
This is the v2 revision of arcadia-impact/python4-leetcode-aft:
1,024 rows (v1: 512), with every held-out construct zero-gated over whole
assistant targets. It supersedes the v1… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/python4-leetcode-eft.scimt-dispatch-charter-250m-v1ShareGPT-Genshin-Impact-Human-Gpt
honkai_impact_3rd_chinese_dialogue_corpus
崩坏三游戏剧情语料
总计 92,421 句剧情对白(带有角色标签)+旁白,从崩坏3的“主线1黄昏、少女、战舰”到“主线第二部03间章:一个梦游者的苦痛”
本数据集从 honkai_impact_3rd_game_playthrough 视频数据集出发,经过 AI pipeline 最终获取结构化的文本剧情语料。
AI pipeline 概述如下:
分P下载视频(使用 BBDown 下载 BiliBili崩三剧情视频)
视频帧分割(每1秒取一帧画面)
逐帧 OCR 检测文本(使用 Paddle-OCR)
逐帧 VLM 结构化解析(使用 MiniCPM-V-2_6,输入为帧图像 + OCR结果,输出为结构化 JSON)
基于规则的后处理
规范化 VLM 输出(e.g., 去噪、排除格式有问题的输出)
中间帧的信息去重与归并(e.g.… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/honkai_impact_3rd_chinese_dialogue_corpus.secret-traits
Secret Traits — eval & training prompts
Prompt datasets for the secret-traits
mini-eval, which scores RM-bias model organisms on two axes: whether they
exhibit 6 reward-model-bias behaviours, and whether they reveal those hidden
behaviours under 4 interrogation attacks.
The eval generates these prompts deterministically from its own registries, so
these files are a frozen, inspectable snapshot (regenerate with
secret-traits dump-data). All English, all synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/secret-traits.nasa-smd-IR-benchmark
NASA-IR benchmark
NASA SMD and IBM Research developed a domain-specific information retrieval benchmark, NASA-IR, spanning almost 500 question-answer pairs related to the Earth science, planetary science, heliophysics, astrophysics, and biological physical sciences domains. Specifically, we sampled a set of 166 paragraphs from AGU, AMS, ADS, PMC, and PubMed and manually annotated with 3 questions that are answerable from each of these paragraphs, resulting in 498 questions. We used… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-IR-benchmark.question-consistency-datasets
Question Consistency — concept datasets
Item pools for the question-consistency
preference/judgement-elicitation harness (forced-choice pairwise comparisons → Thurstonian fit →
consistency metrics). Each config is a flat list of items in a single item (string) column.
config
rows
what
items_500
500
500-concept sentiment/judgement pool
items_2000
2000
2000-concept pool (large-scale runs)
curated_concepts
250
curated rich multi-word concepts spanning categories… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/question-consistency-datasets.impact-psnc-polish-ocr
IMPACT-PSNC Polish OCR Diverse Subset
Compact, provenance-preserving subset of the Polish IMPACT ground truth released by the Poznan Supercomputing and Networking Center (PSNC). It is intended for OCR experiments on diverse historical Polish printed material.
This subset contains:
89 full-page images from 30 source collections;
599 text-region crops derived from PAGE XML polygons;
the 89 corresponding original PAGE XML files;
page and region transcriptions;
document-level train… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/impact-psnc-polish-ocr.synthrad2025-impact-registration
🧭 SynthRAD2025 IMPACT Registrations (BSpline Transforms)
This repository provides Elastix B-spline transformation parameter files generated using the IMPACT method on the SynthRAD2025 dataset.
Each file corresponds to a non-rigid registration between a reference CT and another modality (MRI or CBCT), aligned into CT space using Elastix with the IMPACT similarity metric.
Task 1: 411 transforms (102 excluded cases)
Task 2: 638 transforms (136 excluded cases)
Excluded: All cases… See the full description on the dataset page: https://huggingface.co/datasets/VBoussot/synthrad2025-impact-registration.python4-gemma3-27b-eft-v2-logsscimt-dispatch-graft-dose-v1
