datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opm-ehri-datadesign-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.notch-beam-2d-impact
NotchBeam2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/notch-beam-2d-impact",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/finebooks/bhl-impact-gt.sentry-impact-risk
NASA Sentry: Earth Impact Risk Assessment
Credit: NASA/Johns Hopkins APL
Part of a dataset collection on Hugging Face.
Dataset description
Near-Earth objects with non-zero Earth impact probability from NASA JPL Sentry system.
The Sentry system, operated by NASA's Center for Near-Earth Object Studies (CNEOS) at the Jet Propulsion Laboratory, continuously monitors the most current asteroid catalog for possibilities of future Earth impact. Objects are listed… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/sentry-impact-risk.taylor-impact-2d
Taylor2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/taylor-impact-2d",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/taylor-impact-2d", repo_type="dataset")… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/taylor-impact-2d.bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/bhl-impact-gt.denoising-impact-evaluation-dataset
Denoising Impact Evaluation Dataset
Dataset Description
The ekacare/denoising-impact-evaluation-dataset is a comprehensive benchmark dataset designed to evaluate the effects of speech enhancement on automatic speech recognition (ASR) systems in medical speech contexts. It includes paired noisy and denoised audio subsets under controlled acoustic conditions to support systematic analysis of denoising performance.
Source Data
Base Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/denoising-impact-evaluation-dataset.synthrad2023-impact-registration
🧭 SynthRAD2023 IMPACT Registrations (BSpline Transforms)
This repository provides Elastix B-spline transformation parameter files generated using the IMPACT method on the SynthRAD2023 dataset.
Each file corresponds to a non-rigid registration between a reference CT and another modality (MRI or CBCT), aligned into CT space using Elastix with the IMPACT similarity metric.
Task 1: 317 transforms (43 excluded cases)
Task 2: 289 transforms (69 excluded cases)
🚀 Overview… See the full description on the dataset page: https://huggingface.co/datasets/VBoussot/synthrad2023-impact-registration.impact-psnc-polish-ocr
IMPACT-PSNC Polish OCR Diverse Subset
Compact, provenance-preserving subset of the Polish IMPACT ground truth released by the Poznan Supercomputing and Networking Center (PSNC). It is intended for OCR experiments on diverse historical Polish printed material.
This subset contains:
89 full-page images from 30 source collections;
599 text-region crops derived from PAGE XML polygons;
the 89 corresponding original PAGE XML files;
page and region transcriptions;
document-level train… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/impact-psnc-polish-ocr.ShareGPT-Genshin-Impact-Human-Gpt
question-consistency-datasets
Question Consistency — concept datasets
Item pools for the question-consistency
preference/judgement-elicitation harness (forced-choice pairwise comparisons → Thurstonian fit →
consistency metrics). Each config is a flat list of items in a single item (string) column.
config
rows
what
items_500
500
500-concept sentiment/judgement pool
items_2000
2000
2000-concept pool (large-scale runs)
curated_concepts
250
curated rich multi-word concepts spanning categories… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/question-consistency-datasets.video-dataset-genshin-impact-landscape-organized
Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning.
synthrad2025-impact-registration
🧭 SynthRAD2025 IMPACT Registrations (BSpline Transforms)
This repository provides Elastix B-spline transformation parameter files generated using the IMPACT method on the SynthRAD2025 dataset.
Each file corresponds to a non-rigid registration between a reference CT and another modality (MRI or CBCT), aligned into CT space using Elastix with the IMPACT similarity metric.
Task 1: 411 transforms (102 excluded cases)
Task 2: 638 transforms (136 excluded cases)
Excluded: All cases… See the full description on the dataset page: https://huggingface.co/datasets/VBoussot/synthrad2025-impact-registration.cve-impacts
Image generated by DALL-E.
CVE_KeyPhrases
CVE_KeyPhrases is a dataset of published CVEs with the Key Risk Phrases (for Impact, Weakness, Attack) extracted.
It is released under license": "cc-by-sa-4.0"
Please see the BSides Dublin 2024 presentation video and deck.
The dataset includes:
~230K published CVEs (excluding those marked Rejected) i.e. all CVEs up to April 3 2024 NVD Published date.
The CVE ID, Description text, and Key Risk Phrases
As of April 2024… See the full description on the dataset page: https://huggingface.co/datasets/yahoo-inc/cve-impacts.NASA_IMPACT_TCB_OBJ
Transverse Cirrus Bands (TCB) Dataset
Dataset Overview
This dataset contains manually annotated satellite imagery of Transverse Cirrus Bands (TCBs), a type of cloud formation often associated with atmospheric turbulence. The dataset is formatted for object detection tasks using the YOLO and COCO annotation formats, making it suitable for training deep learning models for automated TCB detection.
Data Collection
Source: NASA-IMPACT Data Share
Satellite Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/viknesh1211/NASA_IMPACT_TCB_OBJ.genshin_impact_CHONGYUN_Paints_UNDO_Sketch_UnChunked
Genshin-Impact-Portrait-with-Tags-Filtered-IID-Gender-SP
name_mapping = {
'芭芭拉': 'BARBARA',
'柯莱': 'COLLEI',
'雷电将军': 'RAIDEN SHOGUN',
'云堇': 'YUN JIN',
'八重神子': 'YAE MIKO',
'妮露': 'NILOU',
'绮良良': 'KIRARA',
'砂糖': 'SUCROSE',
'珐露珊': 'FARUZAN',
'琳妮特': 'LYNETTE',
'纳西妲': 'NAHIDA',
'诺艾尔': 'NOELLE',
'凝光': 'NINGGUANG',
'鹿野院平藏': 'HEIZOU',
'琴': 'JEAN',
'枫原万叶': 'KAEDEHARA KAZUHA',
'芙宁娜': 'FURINA',
'艾尔海森': 'ALHAITHAM',
'甘雨': 'GANYU',
'凯亚': 'KAEYA',
'荒泷一斗': 'ARATAKI ITTO',
'优菈':… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin-Impact-Portrait-with-Tags-Filtered-IID-Gender-SP.Genshin_Impact_Yae_Miko_MMD_Video_Dataset_Captioned
In the style of Yae Miko , The video opens with a darkened scene where the details are not clearly visible. As the video progresses, the lighting improves, revealing a character dressed in traditional Japanese attire, standing on a stone pathway. The character is holding what appears to be a scroll or a piece of paper. Surrounding the character are several lanterns with intricate designs, casting a warm glow on the pathway and the character's clothing. In the background, there is a… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin_Impact_Yae_Miko_MMD_Video_Dataset_Captioned.honkai_impact_3rd_game_playthrough
Game Playthrough
最终解析出的语料在 honkai_impact_3rd_chinese_dialogue_corpus。
See honkai_impact_3rd_chinese_dialogue_corpus for final parsed result!
Description (English)
This is a collection of playthrough videos of Honkai Impact 3rd from Hoyoverse, along with efforts to build a Chinese text corpus (with OCR and MLLM-based parsing).
The language setting is Chinese.
All credits to the source author from BiliBili
The dataset contains the following contents:
Videos: The… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/honkai_impact_3rd_game_playthrough.Genshin_Impact_RaidenShogun_Voice_koreanimpactshift-us-county-day-extreme-weather
ImpactShift-US
ImpactShift-US is a delay-aware county-day benchmark built from real U.S.
government data for extreme-weather research under future-time and geographic
distribution shift. It separates NOAA nClimGrid-Daily county-average weather,
county-coded NOAA Storm Events and reported impacts, and CDC/ATSDR SVI 2022
county context.
The primary access layer contains 12,483,926 county-days for 3,107 GEOIDs over
2015–2025. Frozen time and geography roles support retrospective… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/impactshift-us-county-day-extreme-weather.qwen3-14b-conditional-lora-animal-20260625IMPACTS
I.M.P.A.C.T.S
Innovative Mimicry Patterns for Astrobiological Conditions and Terrestrial Shifts
Designed for Cross-Discipline/Interconnected Critical Thinking, Nuanced Understanding, Diverse Role Playing, and Innovative Problem Solving
I.M.P.A.C.T.S is a unique dataset created to empower large language models (LLMs) to explore and generate novel insights across the realms of biomimicry, climate change scenarios, and astrobiology. By intertwining detailed examples… See the full description on the dataset page: https://huggingface.co/datasets/Severian/IMPACTS.Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Captionimport os
import uuid
import re
import numpy as np
import pandas as pd
from datasets import load_dataset, Dataset
from PIL import Image
import toml
from tqdm import tqdm
from IPython import display
# 1. 加载数据集
ds = load_dataset("svjack/Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Caption")
# 2. 移除 image 列并转换为 Pandas DataFramedf = ds["train"].remove_columns(["image"]).to_pandas()
# 3. 定义字典
new_dict = {
'砂糖': 'SUCROSE', '五郎': 'GOROU', '雷电将军': 'RAIDEN SHOGUN', '七七': 'QIQI'… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Caption.video-dataset-genshin-impact-ep-character-organized
Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning.
premier-league-first-goal-impact-2025-26
What Is the First Goal Worth? — 2025/26 Premier League
Match-level data behind a 5DollarFootballAPI study of how the first confirmed goal changed
Bet365's normalized in-play win probabilities during the 2025/26 Premier League season.
Across 347 usable matches, the median within-match increase in the scoring team's normalized
win probability was 23.1 percentage points (bootstrap 95% CI: 22.4–24.2). The median
first goal after minute 75 moved the probability by 62.2 points… See the full description on the dataset page: https://huggingface.co/datasets/5dollarfootballapi/premier-league-first-goal-impact-2025-26.impact-craters
Impact Craters
Credit: NASA/Apollo 8
Part of a dataset collection on Hugging Face.
Dataset description
Comprehensive database of impact craters across the solar system, sourced from Wikidata.
Impact craters are among the most widespread geological features in the solar system, formed when asteroids, comets, or meteoroids collide with a planetary surface. They are critical windows into a body's geological history: crater size-frequency distributions reveal… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/impact-craters.genshin-impact-outfits
Genshin Impact Outfit
This is a collection of Genshin Impact character outfits (both wish and in-game version), with outfit description and detailed appearance, parsed from Fandom Wiki.
nasa-science-code-benchmark-v0.1.1
NASA Code Retrieval Benchmark v0.1.1
This repository is an updated version of the NASA Code Retrieval Benchmark. It provides a code retrieval benchmark based on code from 7 programming languages sourced from NASA's GitHub repositories.
What's New in v0.1.1?
v0.1.1 introduces a hierarchical structure and official Hugging Face dataset configurations. This allows you to evaluate models specifically by language or by query category without data redundancy in the file system.… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.1.
