datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grammar-correction
grammar-correction
Dataset Summary
The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset,
derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction.
It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models.
Dataset Structure
Train set: 100 000 entries
Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.qualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.InterCode-Corrections
Dataset Card for InterCode-Corrections
This is a manually corrected version of the InterCode-Bash dataset, providing natural language prompts and Bash commands for the task of machine translation.
Dataset Details
Dataset Description
This dataset contains corrections for errors in the InterCode-Bash dataset. corrections.csv contains annotations for each error. final.csv contains the updated dataset with the corrections applied. The corrected dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/westenfelder/InterCode-Corrections.vn-spell-correction-eval-real
vn-spell-correction-eval-real
Out-of-distribution evaluation corpus for Vietnamese spell-correction
models — 150 hand-curated (noisy, clean) pairs sampled from real
VN error sources, not generated by nom.text.noise.
This is the test set we use to verify a spell-correction model
generalises beyond its own synthetic training distribution. A model
that scores 95 % on nom-vn's synthetic eval grid and 60 % on this
set is overfit to the noise generator.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.bias-correction-palestine-protocol
Dataset Card for LLM Bias Correction (Palestine/Israel Context)
This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel.
Dataset Structure
The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.zhtw-sentence-error-correction
中文錯字糾正資料集
由規則與字典自維基百科產生的錯誤糾正資料集。
包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。
資料集使用函式庫: p208p2002/zh-mistake-text-gen
子集
alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。
beta: 50%錯誤,50%不變。單句中僅有一個錯誤。
gamma: 100%錯誤。單句中可能有多個錯誤。
gemma4-onpolicy-student-corrections
Gemma 4 12B FrontierDistill - On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill).
Every example in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-student-corrections.ko-finance-asr-corrections
ko-finance-asr-corrections
Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube.
210 pairs
mined from 2,391 videos of auto-captions
across 47 channels
totalling 1,080.1 hours
Each pair carries how often the term was mangled and how often it was said correctly, plus
verification provenance.
한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답
표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다.
What makes it different
No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.urdu-asr-error-correction-data
Urdu ASR Generative Error Correction Dataset
This dataset contains paired training and testing data for post-ASR error correction in Urdu.
Dataset Details
Language: Urdu (ur)
Task: ASR Error Correction
License: CC BY-NC 4.0
Dataset Structure
The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold).
train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.python-runtime-verified-error-correction
Python Runtime-Verified Error Correction Dataset 🐍⚡
Overview
Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution.
Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.ocr-correction
OCR (Optical Character Recognition) Correction Dataset
This dataset comprises OCR-corrected text samples from English books and newspapers sourced from the Internet Archive. It provides pairs of raw OCR text and their AI-corrected versions, designed for OCR correction tasks.
Dataset Structure
Data Instances
Each instance contains:
input: Raw OCR text with errors
output: Corrected text
Example:
{
"input": "\n\n(ii) The income of Tarai and Bhabar… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ocr-correction.text-correction_collection
Human Samples
These samples contains contains human-written sentences produced during language learning practice, combined with AI-based grammatical verification and correction. The original sentences were written by language learners who often did not know whether their sentences were correct or incorrect. These authentic learner inputs capture a wide range of natural mistakes, such as spelling, syntax, word choice, and structure errors.
Synthetic Samples
These… See the full description on the dataset page: https://huggingface.co/datasets/marcelone/text-correction_collection.ego-mistake-corrections
Ego Mistake Corrections Benchmark (Ego-MC-Bench)
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type counts in annotations_release.json:… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-mistake-corrections.gemma4-onpolicy-50topics-2000-corrections
Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.text-correction-validationvn-spell-correction-train
nrl-ai/vn-spell-correction-train
459,478 (noisy, clean) Vietnamese training pairs for fine-tuning a
seq2seq spell-correction model. Each row:
{"input": "<noisy>", "target": "<clean>"}
Both fields are NFC-normalized.
How it was built
Clean side: same 500K register-balanced mix as
nrl-ai/vn-diacritic-train —
350K Vietnamese Wikipedia (CC-BY-SA-4.0,
hirine/wikipedia-vietnamese-1M296K-dataset) + 150K NFC-fixed
Vietnamese news (CC-BY-4.0, tmnam20/Vietnamese-News-dedup).… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-train.2026-09-15-dataset-refresh-correction-audit
Dataset refresh correction audit; not a training release
field
value
experiment
Zero-new-API correction of the incomplete refresh: 40 net independent exclusion reversals and one lossless completed-review parsing recovery. Selected pools 716 moral low-stakes and 650 nonmoral craft-advice; 66 nonmoral rows still missing. Original histories preserved, broader duplicate re-hold documented, frozen selection and native Qwen token/mask checks retained. Four saved-answer… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-correction-audit.pleias-post-ocr-correction-chonkie-aligned-en
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.dfm11-folketingets-dokumenter-error-correction
DFM11 Folketingets Dokumenter Error Correction
This dataset is the fully audited DFM11 replacement for
schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.
Every retained input was generated from its target using 1-8 declared
synthetic OCR substitutions. Deterministic text-quality filtering was followed
by a task-aware Gemma 4 audit of all 2,548,956 surviving
rows; 63,109 audit rejections were removed and
2,485,847 rows remain.
Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.Bird_explained_corrections
Dataset Card for Dataset Name
This dataset is truncated
vn-spell-correction-eval
nrl-ai/vn-spell-correction-eval
Vietnamese spell-correction evaluation grid: 4 source registers × 2
noise levels = 8 splits, 2,098 (noisy, clean) sentence pairs total.
Each pair is {"input": "<noisy>", "target": "<clean>"}. Both sides are
NFC-normalized. The clean target is the same sentence used as the
target in nrl-ai/vn-diacritic-eval —
spell correction is a strict superset of diacritic restoration, so we
reuse the same registers-balanced corpus.
Splits
Two noise… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval.pleias-post-ocr-correction-chonkie-aligned-fr
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.babble-corrections
babble — corrections
Training data for babble: a ~3M parameter
byte-level transformer that started from random weights and has only ever
learned from people correcting it in Discord.
There is no pretraining corpus. There is no scraped chat history. Every row here
is somebody deliberately teaching a small confused model to talk.
How a row happens
Someone @mentions the bot.
The bot replies with whatever its current weights produce. Early on this is
noise, and it is… See the full description on the dataset page: https://huggingface.co/datasets/kowo-co/babble-corrections.spell-correction
Spell-Check Dataset
This dataset consists of pairs of misspelled words and their corresponding correctly spelled words, designed for training and evaluating character-level spelling correction models. It is particularly useful for tasks such as:
Spelling correction
Character-level sequence-to-sequence modeling
Error detection and correction in text
Each data point in the dataset contains:
misspelled: A misspelled version of a word.
correct: The corrected spelling of the word.… See the full description on the dataset page: https://huggingface.co/datasets/torinriley/spell-correction.gemma4-onpolicy-50topics-corrections
Gemma 4 FrontierDistill - Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 FrontierDistill Project.
This dataset contains 1,000 authentic on-policy student failure corrections collected live from gemma-4-12b-it-qat-frontierdistill across 50 distinct… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-corrections.RPEval_correctionThe dataset is an enhanced version of https://github.com/yelboudouri/RPEval.git
The ratio of yes to no in the task type DECISION has been balanced from 5866:213 to 3067:3012.
RPEval: Role-Playing Evaluation for Large Language Models
This repository contains code and data referenced in: "Role-Playing Evaluation for Large Language Models".
Large Language Models (LLMs) demonstrate a notable capacity for adopting personas and engaging in role-playing. However,
evaluating… See the full description on the dataset page: https://huggingface.co/datasets/jiangchengchengNLP/RPEval_correction.indonesian-grammar-correction
Koreksi Tata Bahasa Indonesia ✍️
Kumpulan 126 pasangan kalimat (asli → koreksi) untuk grammar error correction bahasa Indonesia.
Kenapa dataset ini ada?
Grammar error correction (GEC) untuk bahasa Indonesia belum ada di HF — padahal model GEC global (271 downloads) jadi salah satu kategori paling dicari. Dataset ini isi gap itu: dari ejaan ("Dimana" → "Di mana"), ragam ("gw udah" → "saya sudah"), pleonasme ("para siswa-siswa" → "para siswa"), sampai huruf kapital.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-grammar-correction.
