datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/finebooks/bhl-impact-gt.libero-kinex-v0.6.0-vs-kinex-v0.6.0-improve-vs-codexEvaluation website · Data format
kinex (v0.6.0) vs kinex (v0.6.0-improve) vs codex · LIBERO Long · GPT-6 Astra / medium
45 planned episodes. Counts below are derived from the episode index.
Variant
Native successes / valid
Normal successes / valid
Interrupted
Missing
kinex (v0.6.0)
9/15
9/15
0
0
kinex (v0.6.0-improve)
9/15
9/15
0
0
codex
6/15
6/15
0
0
Native-valid counts include interrupted executions. Normal counts also require a finished execution and valid… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/libero-kinex-v0.6.0-vs-kinex-v0.6.0-improve-vs-codex.bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/bhl-impact-gt.improving-ca-lora-artifacts
Improving CA-LoRA — CA measurements and generated images
The two large evidence sets behind the reproduction-and-extension study of CA-LoRA
(Concept-Aware LoRA) on SDXL / Cityscapes:
measurements/ — the raw head-granularity concept-attribution tensors for a 13-timestep
sweep (backs Table 2 and Figure 1 of the report).
generated/ — every image that was scored for the main results (backs Table 3 of the
report).
The adapters that produced the images live in… See the full description on the dataset page: https://huggingface.co/datasets/chs35/improving-ca-lora-artifacts.ImageNet-C-impulse_noise-severity_5Sombench-IMP-Segmentation
SomBench Benchmark: Irregular Mare Patch (IMP) Segmentation
Science theme: Volcanic history
Task: Binary semantic segmentation
Dataset Summary
A binary semantic-segmentation benchmark for irregular mare patches
(IMPs): rare, morphologically subtle features interpreted as unusually young
volcanic landforms. Each sample is an LROC NAC image tile paired with a
binary IMP mask (IMP vs. background). The set is derived from published IMP
polygon annotations, framed as a… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-IMP-Segmentation.impact-psnc-polish-ocr
IMPACT-PSNC Polish OCR Diverse Subset
Compact, provenance-preserving subset of the Polish IMPACT ground truth released by the Poznan Supercomputing and Networking Center (PSNC). It is intended for OCR experiments on diverse historical Polish printed material.
This subset contains:
89 full-page images from 30 source collections;
599 text-region crops derived from PAGE XML polygons;
the 89 corresponding original PAGE XML files;
page and region transcriptions;
document-level train… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/impact-psnc-polish-ocr.multi-edit-clipd-improved
Multi-Edit Image Pairs - CLIP-D Improved Instructions
This dataset is a refined version of multi-instruct-image-editing/multi-edit-image-pairs where instructions have been selectively replaced with Qwen2-VL-7B generated instructions when CLIP Directional Similarity (CLIP-D) improved.
Dataset Statistics
Total Samples: 1000
Instructions Replaced: 526 (52.6%)
Original Instructions Kept: 474 (47.4%)
Improvements
CLIP-D (Directional Similarity)… See the full description on the dataset page: https://huggingface.co/datasets/engy58/multi-edit-clipd-improved.improved_aesthetics_4.5plus-ultra-hrversion https://git-lfs.github.com/spec/v1
oid sha256:98b45ea81164d1e1a1dd82255207053b15cd6c69d922a1c5cf3387ce604d4b74
size 28
Group_15_impedance_imagesguided_genshin_impact_official_server_recordings_01
原神 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game: 原神 (Genshin Impact)
Collection: guided (精数据)
Subset: official_server (官服(非私服))
Recordings: 566
Planned bytes: 4440851482194
Layout: recordings//
Parquet files are intentionally excluded.
metaworld_spatial_cardinal_15cm_strict_50deg_close_high_approach_improved_mwThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "metaworld",
"total_episodes": 800,
"total_frames": 97649,
"total_tasks": 8,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 24,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/metaworld_spatial_cardinal_15cm_strict_50deg_close_high_approach_improved_mw.Genshin-Impact-Portrait-with-Tags-Filtered-IID-Gender-SP
name_mapping = {
'芭芭拉': 'BARBARA',
'柯莱': 'COLLEI',
'雷电将军': 'RAIDEN SHOGUN',
'云堇': 'YUN JIN',
'八重神子': 'YAE MIKO',
'妮露': 'NILOU',
'绮良良': 'KIRARA',
'砂糖': 'SUCROSE',
'珐露珊': 'FARUZAN',
'琳妮特': 'LYNETTE',
'纳西妲': 'NAHIDA',
'诺艾尔': 'NOELLE',
'凝光': 'NINGGUANG',
'鹿野院平藏': 'HEIZOU',
'琴': 'JEAN',
'枫原万叶': 'KAEDEHARA KAZUHA',
'芙宁娜': 'FURINA',
'艾尔海森': 'ALHAITHAM',
'甘雨': 'GANYU',
'凯亚': 'KAEYA',
'荒泷一斗': 'ARATAKI ITTO',
'优菈':… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin-Impact-Portrait-with-Tags-Filtered-IID-Gender-SP.frakturline-testset
Fraktur/Other Text-Line — Test Set
A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman).
Developed for the Impresso digital humanities project.
Dataset Details
Property
Value
Task
Binary image classification
Classes
fraktur, other
Images per class
1 000
Total images
2 000
Image format
WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.dental-implant-surgery-sample
Dental Implant Surgery — Multimodal Annotated Video (Sample Case)
A public sample from one complete All-on-4 full-arch mandibular dental implant
surgery: two synchronised camera angles, the operating surgeon narrating while
he works, and six layers of structured clinical annotation (L0–L5) tied frame by
frame to what he said.
This is a showcase slice, not the whole case. What is here is enough to judge
the structure, the annotation quality and the honesty of the documentation.… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.f1_top100_improved_gif_framesNASA_IMPACT_TCB_OBJ
Transverse Cirrus Bands (TCB) Dataset
Dataset Overview
This dataset contains manually annotated satellite imagery of Transverse Cirrus Bands (TCBs), a type of cloud formation often associated with atmospheric turbulence. The dataset is formatted for object detection tasks using the YOLO and COCO annotation formats, making it suitable for training deep learning models for automated TCB detection.
Data Collection
Source: NASA-IMPACT Data Share
Satellite Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/viknesh1211/NASA_IMPACT_TCB_OBJ.Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Captionimport os
import uuid
import re
import numpy as np
import pandas as pd
from datasets import load_dataset, Dataset
from PIL import Image
import toml
from tqdm import tqdm
from IPython import display
# 1. 加载数据集
ds = load_dataset("svjack/Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Caption")
# 2. 移除 image 列并转换为 Pandas DataFramedf = ds["train"].remove_columns(["image"]).to_pandas()
# 3. 定义字典
new_dict = {
'砂糖': 'SUCROSE', '五郎': 'GOROU', '雷电将军': 'RAIDEN SHOGUN', '七七': 'QIQI'… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Caption.Impressions
Dataset Card for "Impressions"
Overview
The Impressions dataset is a multimodal benchmark that consists of 4,100 unique annotations and over 1,375 image-caption pairs from the photography domain. Each annotation explores (1) the aesthetic impactfulness of a photograph, (2) image descriptions in which pragmatic inferences are welcome, (3) emotions/thoughts/beliefs that the photograph may inspire, and (4) the aesthetic elements that elicited the expressed impression.
EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Impressions.corruption-impulse_noise
Corruption Dataset: Impulse_Noise
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using impulse_noise corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Impulse_Noise
Adds… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-impulse_noise.imprecision-bench
imprecision-bench
A multimodal benchmark for evaluating whether LLMs calibrate linguistic precision to pragmatic context, paired with 475 human productions and a peer-reviewed RSA baseline (r² ≈ 0.97).
This dataset accompanies the paper:
Modeling (Im)precision in Context
Roland Mühlenbernd, Stephanie Solt
Linguistics Vanguard, 2022
[Paper] · [Source Data] · [Companion Repo]
Notebook
notebook.ipynb — guided walkthrough: data loading, sample evaluation (1 row… See the full description on the dataset page: https://huggingface.co/datasets/RolandM/imprecision-bench.improved_aesthetics_6.5plusridgelora-stage2-imposebase-train160-50k-20260824
Stage-2 ControlNet retraining with the frozen IMPOSE base
This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE
checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM
t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices).
An initial Stage-1-from-scratch job was stopped at step 575 after correcting
the scope. It produced no scheduled checkpoint and is not used in any result;
its log is retained only as an audit trail.
Frozen IMPOSE… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824.genshin-impact-outfits
Genshin Impact Outfit
This is a collection of Genshin Impact character outfits (both wish and in-game version), with outfit description and detailed appearance, parsed from Fandom Wiki.
africa-synth-disability-visual-impairment-low-vision-all
Visual Impairment & Low Vision Services (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-disability-visual-impairment-low-vision-all.multi-edit-clipd-improved-batch0001ImpText-Bench
ImpText-Bench Metadata
This directory contains the benchmark JSONL manifests. The image files are
distributed through the Hugging Face dataset
Riversideli/ImpText-Bench.
Expected image layout:
images/
white/<id>.png
black/<id>.png
The full manifest has 1,630 records:
1,141 benign samples.
489 implicit-text samples across physical deformation, visual camouflage, and
cognitive suggestion categories.
The taxonomy follows the ImpText-Bench definition:
Primary category… See the full description on the dataset page: https://huggingface.co/datasets/Riversideli/ImpText-Bench.multimodal-genshin-impact
Genshin Impact Fandom Wiki Multimodal Dataset
Github repo here
Description
This dataset is a comprehensive collection of 22,162 fandom wiki pages for the popular game Genshin Impact.
The dataset includes markdown-formatted English content from the wiki, featuring interleaved text, as well as image, video, and audio file links. Additionally, the associated multimodal files (images, videos, and audio) have been downloaded and organized to facilitate the multimodal dataset… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/multimodal-genshin-impact.implicit-fast-ki-iar-heatmapsUme-Impressionism
Ume - Impressionism dataset
I provide the images used to create my LoRA : https://huggingface.co/UmeAiRT/FLUX.1-dev-LoRA-Impressionism
