datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wmt-mqm-error-spans
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context in a form of error spans. Moreover, it contains some hallucinations used in the training of XCOMET models.
Please note that this is not an official release of the data and the original data can be found here.
The data is organised into 8 columns:
src: input text
mt: translation
ref: reference translation
annotations: List… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-error-spans.99-GEO-Errors
99 Errors in GEO
Why organisations become invisible, misrepresented or unsupported in AI answers
GEO means Generative Engine Optimization. This six-language companion book turns 99 recurring representation failures into auditable warnings. Each warning records the evidence needed, a correction protocol, a revalidation question and a machine-readable rule.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz · NobleJackal ·… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/99-GEO-Errors.GBO-99-Errors
99 Mistakes in GBO
Why AI agents choose badly, exceed their authority and fail to stop
GBO means Generative Behavior Optimization: designing and governing what AI agents are allowed to do. This six-language companion to NOMOS GBO examines 99 failure patterns, each with a scenario, potential harm, detection signal, appropriate behaviour, machine rule and audit question.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz ·… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/GBO-99-Errors.red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.toulmin_errors
Reasoning Rubrics — Toulmin-Typed Error Localization Benchmark
A multi-domain benchmark for studying typed reasoning errors in LLMs and AI
scientific reasoning agents. Errors are labeled along four Toulmin
argumentation dimensions: Grounds (premises/facts), Warrant
(inferential step), Qualifier (scope/certainty), Rebuttal
(competing evidence).
The benchmark has two parts:
Typed external benchmarks. Existing reasoning-error benchmarks
relabeled with Toulmin dimensions on top of the… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/toulmin_errors.Long-instructionsUPBench-Error-verified-v2
UPBench-Error-verified-v2
LCZZZZ/UPBench-Error → generation_error 子集,经两轮人工核验后保留的
1,219 条样本。每条样本视觉上看不出明显的低级生成缺陷。
筛选过程
步骤
剩余
原始 generation_error 样本
5,761
剔除 is_gui=true(GUI-World / egoproactive 屏幕录制)
3,390
第一轮:逐条过目 error_clip.mp4
good 1,420 / bad 1,970
第二轮:对第一轮 good 再过一遍
good 1,219 / bad 201
第二轮刷掉了第一轮 14.2% 的样本,最终保留率 1,219 / 3,390 = 36.0%。
内容
metadata/manifest-verified.jsonl 1,219 条,原 manifest 全部 25 个字段逐字保留,… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified-v2.agentic-error-judge-v1
Agentic Root-Cause Judge Set — v1 (legacy)
LLM-judged root-cause labels for a 20K sample of agent tool-calling traces from
Agent-Ark/Toucan-1.5M,
using the B1–B8 agentic error taxonomy (as opposed to the hallucination-content
taxonomy used by the distill-reasoning/bert-spans datasets in this collection).
This is the first, flat-schema run — superseded by agentic-error-judge-v2
in this collection, which uses an updated multi-span-per-trace schema and covers
more traces. Kept here… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/agentic-error-judge-v1.assay-cupel-pose-error-v2-r7-mandatory-sandbox-rlalbanian-error-augmentation
Albanian Controlled Error Augmentation Dataset
Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation.
Each row is an (incorrect → correct) pair with an explicit error_type label.
Error types
error_type
Description
missing_diacritic
Missing ë / ç
c_q_confusion
Confusion between ç / q / c
digraph_reduction
Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.assay-cupel-pose-error-v2-r6-sandbox-harness-rlUPBench-Error-verified
UPBench-Error-verified
人工核验过的 LCZZZZ/UPBench-Error → generation_error 子集。
保留其中看不出明显低级生成缺陷的样本。
这是怎么筛出来的
从原数据集 5,761 条 generation_error 样本出发:
步骤
剩余
原始样本
5,761
剔除 is_gui=true(GUI-World / egoproactive 屏幕录制)
3,390
逐条人工过目 error_clip.mp4
3,390 全部看完
判定为无明显生成缺陷(good)
1,420
判定为有缺陷(bad)
1,970
内容
metadata/manifest-verified.jsonl 1,420 条,原 manifest 字段逐字保留,
另加 human_verification 字段… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified.zhtw-sentence-error-correction
中文錯字糾正資料集
由規則與字典自維基百科產生的錯誤糾正資料集。
包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。
資料集使用函式庫: p208p2002/zh-mistake-text-gen
子集
alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。
beta: 50%錯誤,50%不變。單句中僅有一個錯誤。
gamma: 100%錯誤。單句中可能有多個錯誤。
slm-reasoning-baseline-error-taxonomy
SLM Reasoning Research — Baseline failure taxonomy
Part of the SLM Reasoning Research project.
All 631 errors from Qwen3-0.6B-Base's zero-shot, zero-training GSM8K baseline (full 1,319-example test
set, 52.16% accuracy), classified into an 11-category failure taxonomy using GPT-OSS-20B as judge
(verified by hand against a 30-example sample first).
misread_semantics dominates at 55.5% (350/631) — far more than arithmetic_slip (10%) — meaning
this model's core weakness is… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-baseline-error-taxonomy.urdu-asr-error-correction-data
Urdu ASR Generative Error Correction Dataset
This dataset contains paired training and testing data for post-ASR error correction in Urdu.
Dataset Details
Language: Urdu (ur)
Task: ASR Error Correction
License: CC BY-NC 4.0
Dataset Structure
The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold).
train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.python-runtime-verified-error-correction
Python Runtime-Verified Error Correction Dataset 🐍⚡
Overview
Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution.
Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.repro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
api-error-retryability-dataset
API Error Retryability Classification
This dataset classifies common API error messages and HTTP status codes into two categories: 'retryable' or 'permanent'. Developers can use this to build robust error handling logic, distinguishing between transient issues that can be retried automatically and permanent failures that require user intervention or code changes.
30 rows · category: reliability · licence: CC0-1.0 (public domain)
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/SharkSkin/api-error-retryability-dataset.ErrorBench
ErrorBench: Fine-Grained Error Analysis of Multi-Family LLMs in Data-to-Text Generation
Dataset Summary
ErrorBench is a human-annotated, span-level benchmark for analyzing generation errors in Large Language Models (LLMs) for Data-to-Text (D2T) generation. The dataset consists of sentences generated from structured DBpedia triples and annotated with fine-grained span-level error labels across 10 error categories.
The dataset was introduced in our IJCNN 2026 paper… See the full description on the dataset page: https://huggingface.co/datasets/soumyaBharadwaj/ErrorBench.api-error-retry-classification
API Error Retryability Classification
This dataset classifies common API error messages and HTTP status codes into 'retryable' or 'permanent' categories. It helps developers implement robust error handling strategies, distinguishing between transient issues that warrant retries and fundamental problems requiring immediate client-side correction.
25 rows · category: reliability · licence: CC0-1.0 (public domain)
Usage
import api_error_retryability_classifier as… See the full description on the dataset page: https://huggingface.co/datasets/SharkSkin/api-error-retry-classification.RAG-Error-Critic-100K
RAG-Critic: Leveraging Automated Critic-Guided Agentic Workflow for Retrieval Augmented Generation
Guanting Dong,
Jiajie Jin,
Xiaoxi Li,
Yutao Zhu,
Zhicheng Dou ✉;
Ji-rong Wen
Github Page,
Gaoling School of Artificial Intelligence, Renmin University of China.
✉ Corresponding Author
Test_v3appliance-fault-error-codes
Appliance fault and error codes
Canonical, always-current version: https://referencesource.org/appliance-fault-error-codes/
Machine-readable: https://referencesource.org/appliance-fault-error-codes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-04
Stale after: 2027-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 170
Error/fault codes displayed by major home appliances (dishwashers, washing… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/appliance-fault-error-codes.dfm11-folketingets-dokumenter-error-correction
DFM11 Folketingets Dokumenter Error Correction
This dataset is the fully audited DFM11 replacement for
schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.
Every retained input was generated from its target using 1-8 declared
synthetic OCR substitutions. Deterministic text-quality filtering was followed
by a task-aware Gemma 4 audit of all 2,548,956 surviving
rows; 63,109 audit rejections were removed and
2,485,847 rows remain.
Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.jenkins_errors
Jenkins Error Outputs Dataset
A large synthetic dataset of Jenkins error outputs, generated for research, machine learning, and error analysis purposes. This dataset contains 100,000 entries, each simulating a real-world Jenkins error log with metadata.
Dataset Structure
Format: JSON Lines (.jsonl)
Fields:
timestamp: ISO8601 timestamp of the error
job_name: Jenkins job name
build_number: Jenkins build number
error_message: Jenkins error message
node: Jenkins node name… See the full description on the dataset page: https://huggingface.co/datasets/Snaseem2026/jenkins_errors.Bulgarian-Text-Errorsrepro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization
A Tight Theory of Error Feedback Algorithms in Distributed Optimization
Reproduction of ICML 2026 paper (OpenReview: dyRD6lBH8K)
Tags
trackio
trackio-logbook
open-experiment
icml2026-repro
paper-dyRD6lBH8K
KORMo-VLM-Dataset-Nemotron-Errorerror-voice-corpus
CatQualia error voice corpus — how a system reports its own failures
4,859 rows · 3,381,687 bytes · JSON Lines, one object per line.
What this is
Failure reports in the system's own voice, paired with the non-voice-aligned alternative. Useful for training a model to describe its errors plainly rather than deflecting them.
Schema
Fields of the first record, read from the file in this repository:
Field
Type
ts
int
grade
str… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/error-voice-corpus.Chat-Error_Pure-dove-sharegpt
