datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
book2-lite-cleanedwmt-mqm-error-spans
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context in a form of error spans. Moreover, it contains some hallucinations used in the training of XCOMET models.
Please note that this is not an official release of the data and the original data can be found here.
The data is organised into 8 columns:
src: input text
mt: translation
ref: reference translation
annotations: List… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-error-spans.anime_pretraining_2Errors_Additive_Manufacturing_Plattform_Cam
Errors_Additive_Manufacturing_Plattform_Cam
3D Printing Nozzle Camera – YOLO Object Detection Dataset
This Repository is part of the Project: Künstliche Intelligenz zur Automatiserten Fehlerkorrektur in der Additiven Fertigung(Förderkennzeichen: 16IS23050B).
This dataset contains images captured from a camera positioned to capture the whole plattform of a 3D printer.
The task is object detection of both regular print elements and typical printing defects.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DasKunststoffZentrumSKZ/Errors_Additive_Manufacturing_Plattform_Cam.Qwen3.5_RL_ErrorCase
Qwen3.5 RL:错误案例与视频定位诊断
v3部分共660题;另新增V4 RL Step3000选帧诊断100题。每题含QA、完整原始输出及可见帧拼图。Dataset Viewer中,default为前60题,video_grounding为v3新增600题,v4_rl_step3000为V4新增100题。
序号
内容
入口
001–060
原三个主实验bench案例
第001题
061–560
RL训练视频500题:训练视觉处理下的新输出
第061题
561–660
ASR-Bench视频100题:复用既有评测输出
第561题
V4-001–100
V4 RL Step3000:VSI/ASR选帧与bbox诊断
V4诊断首页
本次新增的测试内容
新增600题为在看结果前固定的诊断抽样,包含成功与失败,不是600个错误案例。未加入88题附加对照,避免重复。
两部分均为最终Qwen3.5-9B RL… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/Qwen3.5_RL_ErrorCase.Spotlight-VideoGen-Errors
Spotlight Dataset
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
Aditya Chinchure, Sahithya Ravi, Pushkar Shukla, Vered Shwartz, Leonid Sigal
🎉 Accepted to ECCV 2026
🌐 Project Page
Summary
Spotlight is a benchmark for evaluating whether Vision Language Models (VLMs) can precisely
localize and explain errors in AI-generated videos. It contains 600 videos generated by
three state-of-the-art Text-to-Video (T2V) models —… See the full description on the dataset page: https://huggingface.co/datasets/UBC-ViL/Spotlight-VideoGen-Errors.sync_bigjob_8_finalised_processed_with_error_handling_from_51th_splitErrors_Additive_Manufacturing_Nozzle_Cam
Errors_Additive_Manufacturing_Nozzle_Cam
3D Printing Nozzle Camera – YOLO Object Detection Dataset
This Repository is part of the Project: Künstliche Intelligenz zur Automatiserten Fehlerkorrektur in der Additiven Fertigung(Förderkennzeichen: 16IS23050B).
This dataset contains images captured from a camera positioned directly next to the nozzle of a 3D printer.
The task is object detection of both regular print elements and typical printing defects.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DasKunststoffZentrumSKZ/Errors_Additive_Manufacturing_Nozzle_Cam.jupyter-errors-dataset
Dataset Summary
The presented dataset contains 10000 Jupyter notebooks,
each of which contains at least one error. In addition to the notebook content,
the dataset also provides information about the repository where the notebook is stored.
This information can help restore the environment if needed.
Getting Started
This dataset is organized such that it can be naively loaded via the Hugging Face datasets library. We recommend using streaming due to the large size… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/jupyter-errors-dataset.MusicLM99-GEO-Errors
99 Errors in GEO
Why organisations become invisible, misrepresented or unsupported in AI answers
GEO means Generative Engine Optimization. This six-language companion book turns 99 recurring representation failures into auditable warnings. Each warning records the evidence needed, a correction protocol, a revalidation question and a machine-readable rule.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz · NobleJackal ·… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/99-GEO-Errors.GBO-99-Errors
99 Mistakes in GBO
Why AI agents choose badly, exceed their authority and fail to stop
GBO means Generative Behavior Optimization: designing and governing what AI agents are allowed to do. This six-language companion to NOMOS GBO examines 99 failure patterns, each with a scenario, potential harm, detection signal, appropriate behaviour, machine rule and audit question.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz ·… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/GBO-99-Errors.quran-recitation-errors
Examples
Loading dataset:
from datasets import load_dataset
ds = load_dataset('sobolev210/quran-recitation-errors',)
print(ds["train"][0])
blimp-single-errorDataset for probing model preferences for linguistically acceptable sentences.
Generated by introducing automatic corruptions into sentences from Wikipedia, based on UniMorph minimal tag pairs.
More info coming soon!
@misc{glocker2025growmergescalingstrategies,
title={Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation},
author={Kevin Glocker and Kätriin Kukk and Romina Oji and Marcel Bollmann and Marco Kuhlmann and Jenny Kunz},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/blimp-single-error.icelandic-blimp-single-erroropen-smtp-error-dataset
Open SMTP Error Dataset
Open SMTP Error Dataset is an English-language, machine-readable reference
package for SMTP enhanced-status knowledge and cautious operational
classification. Version 1.1.0 contains 92 stable knowledge records and a
separate, auditable catalog of 120 classification rules.
The package is a reference artifact, not a live provider-policy feed. It helps
with observability, support, parser testing, and bounded delivery operations;
it does not establish… See the full description on the dataset page: https://huggingface.co/datasets/blazalek/open-smtp-error-dataset.tabular-errors-v1
TabFix multilingual table error pairs — version 2.0
This release keeps 18 business error categories and separates executable deterministic detection from two residual neural categories: text.encoding and text.spelling. The same repository and family-disjoint splits are retained.
Split
Records
Open-vocabulary views
train
27948
3260
validation
17127
1844
test
32776
3540
The seven string columns remain id, split, family_id, clean_xml, corrupt_xml, errors… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/tabular-errors-v1.german-blimp-single-errorappliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.errorsErrorAnalysis
AnchorSR Error Analysis
500道题,严格沿 failure_cases_500.json 的文件顺序排列,每题对照三个SFT模型。
推荐从第001题开始,点击“下一题”逐题阅读。
也可使用本页上方 Dataset Viewer,一行就是一题:图片、问题、标准答案、三个模型完整输出。
Q-Spatial 150题,SpatialRGPT 175题,VSI 175题(仅尺寸、距离,不含面积)。
全部500题:每个模型均提供原图和标注图,共3000张图;O编号和帧号来自模型声明。
原图与标注图使用相同源帧和拼图顺序。无效框/帧号或无声明会注明,不补造;此时标注页可能没有框。
原图指未添加模型框的网页展示副本,经过等比例缩放与JPEG编码,并非原始文件字节;SpatialRGPT原有区域标记保留。
视频只展示可绘制对象涉及帧,无有效框时展示第1帧,非完整视频。原图/标注图使用相同帧。
原生输出完整保留,包括循环、截断和格式错误;未修改答案或重新评分。
所有模型均为SFT,不是baseline。至少一个模型在该题失败,其他模型可能答对。… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/ErrorAnalysis.red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.toulmin_errors
Reasoning Rubrics — Toulmin-Typed Error Localization Benchmark
A multi-domain benchmark for studying typed reasoning errors in LLMs and AI
scientific reasoning agents. Errors are labeled along four Toulmin
argumentation dimensions: Grounds (premises/facts), Warrant
(inferential step), Qualifier (scope/certainty), Rebuttal
(competing evidence).
The benchmark has two parts:
Typed external benchmarks. Existing reasoning-error benchmarks
relabeled with Toulmin dimensions on top of the… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/toulmin_errors.hvac-error-codes
HVAC Bench Error Code Dataset
Error codes from HVAC equipment sold in the United States, United Kingdom, and Europe, as published by HVAC Bench. Each record gives the manufacturer, the code, the product family the definition applies to, a plain-language meaning, the checks an owner can safely make, the point at which a technician is needed, and the date the definition was last checked against manufacturer documentation. Codes are specific to a product family and are not… See the full description on the dataset page: https://huggingface.co/datasets/mukarram360/hvac-error-codes.Long-instructionsswedish-blimp-single-errorUPBench-Error-verified-v2
UPBench-Error-verified-v2
LCZZZZ/UPBench-Error → generation_error 子集,经两轮人工核验后保留的
1,219 条样本。每条样本视觉上看不出明显的低级生成缺陷。
筛选过程
步骤
剩余
原始 generation_error 样本
5,761
剔除 is_gui=true(GUI-World / egoproactive 屏幕录制)
3,390
第一轮:逐条过目 error_clip.mp4
good 1,420 / bad 1,970
第二轮:对第一轮 good 再过一遍
good 1,219 / bad 201
第二轮刷掉了第一轮 14.2% 的样本,最终保留率 1,219 / 3,390 = 36.0%。
内容
metadata/manifest-verified.jsonl 1,219 条,原 manifest 全部 25 个字段逐字保留,… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified-v2.kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models,
like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected
from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections.
The source videos are documented in the source.txt file.
Usage
from datasets import load_dataset
dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train")
print(dataset)
agentic-error-judge-v1
Agentic Root-Cause Judge Set — v1 (legacy)
LLM-judged root-cause labels for a 20K sample of agent tool-calling traces from
Agent-Ark/Toucan-1.5M,
using the B1–B8 agentic error taxonomy (as opposed to the hallucination-content
taxonomy used by the distill-reasoning/bert-spans datasets in this collection).
This is the first, flat-schema run — superseded by agentic-error-judge-v2
in this collection, which uses an updated multi-span-per-trace schema and covers
more traces. Kept here… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/agentic-error-judge-v1.assay-cupel-pose-error-v2-r7-mandatory-sandbox-rl
