CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /grammar-correction grammar-correction Dataset Summary The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset, derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction. It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models. Dataset Structure Train set: 100 000 entries Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.texttext-classification100K<n<1M11 likes389 downloads2y agoHugging Face02qualcomm /qualcomm-interactive-cooking-dataset-ego-mistake-corrections Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.textvideo-text-to-textn<1K1 likes318 downloads5mo agoHugging Face03greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes136 downloads2mo agoHugging Face04google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes130 downloads3y agoHugging Face05westenfelder /InterCode-Corrections Dataset Card for InterCode-Corrections This is a manually corrected version of the InterCode-Bash dataset, providing natural language prompts and Bash commands for the task of machine translation. Dataset Details Dataset Description This dataset contains corrections for errors in the InterCode-Bash dataset. corrections.csv contains annotations for each error. final.csv contains the updated dataset with the corrections applied. The corrected dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/westenfelder/InterCode-Corrections.texttranslationn<1K0 likes123 downloads1y agoHugging Face06nrl-ai /vn-spell-correction-eval-real vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.texttext-generationn<1K0 likes82 downloads5mo agoHugging Face07tech-equity-collective /bias-correction-palestine-protocol Dataset Card for LLM Bias Correction (Palestine/Israel Context) This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel. Dataset Structure The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.texttext-generationn<1K0 likes71 downloads17d agoHugging Face08sbussiso /synthetic-self-correction-and-thinking-samples Self Correction and Thinking A seed library for training language models to reason with self-correction. Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant. The structure at a glance graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.imagetext-generation1K<n<10K0 likes67 downloads1mo agoHugging Face09p208p2002 /zhtw-sentence-error-correction 中文錯字糾正資料集 由規則與字典自維基百科產生的錯誤糾正資料集。 包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。 資料集使用函式庫: p208p2002/zh-mistake-text-gen 子集 alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。 beta: 50%錯誤,50%不變。單句中僅有一個錯誤。 gamma: 100%錯誤。單句中可能有多個錯誤。 text100K<n<1M5 likes63 downloads3y agoHugging Face10True2456 /gemma4-onpolicy-student-corrections Gemma 4 12B FrontierDistill - On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill). Every example in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-student-corrections.texttext-generation1K<n<10K0 likes60 downloads2mo agoHugging Face11woongstar /ko-finance-asr-corrections ko-finance-asr-corrections Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube. 210 pairs mined from 2,391 videos of auto-captions across 47 channels totalling 1,080.1 hours Each pair carries how often the term was mangled and how often it was said correctly, plus verification provenance. 한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답 표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다. What makes it different No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.tabulartext-generationn<1K0 likes54 downloads11d agoHugging Face12sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes53 downloads6d agoHugging Face13SyntheticLogic-Labs /python-runtime-verified-error-correction Python Runtime-Verified Error Correction Dataset 🐍⚡ Overview Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution. Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.texttext-generation1K<n<10K0 likes51 downloads9mo agoHugging Face14agentlans /ocr-correction OCR (Optical Character Recognition) Correction Dataset This dataset comprises OCR-corrected text samples from English books and newspapers sourced from the Internet Archive. It provides pairs of raw OCR text and their AI-corrected versions, designed for OCR correction tasks. Dataset Structure Data Instances Each instance contains: input: Raw OCR text with errors output: Corrected text Example: { "input": "\n\n(ii) The income of Tarai and Bhabar… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ocr-correction.texttext-generation10K<n<100K1 likes50 downloads2y agoHugging Face15marcelone /text-correction_collection Human Samples These samples contains contains human-written sentences produced during language learning practice, combined with AI-based grammatical verification and correction. The original sentences were written by language learners who often did not know whether their sentences were correct or incorrect. These authentic learner inputs capture a wide range of natural mistakes, such as spelling, syntax, word choice, and structure errors. Synthetic Samples These… See the full description on the dataset page: https://huggingface.co/datasets/marcelone/text-correction_collection.texttext-generation1K<n<10K0 likes47 downloads10mo agoHugging Face16neuripsedtracksub /ego-mistake-corrections Ego Mistake Corrections Benchmark (Ego-MC-Bench) Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type counts in annotations_release.json:… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-mistake-corrections.textvideo-text-to-textn<1K0 likes47 downloads5mo agoHugging Face17True2456 /gemma4-onpolicy-50topics-2000-corrections Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.texttext-generation1K<n<10K0 likes46 downloads2mo agoHugging Face18protonx-models /text-correction-validationtext100K<n<1M11 likes45 downloads10mo agoHugging Face19nrl-ai /vn-spell-correction-train nrl-ai/vn-spell-correction-train 459,478 (noisy, clean) Vietnamese training pairs for fine-tuning a seq2seq spell-correction model. Each row: {"input": "<noisy>", "target": "<clean>"} Both fields are NFC-normalized. How it was built Clean side: same 500K register-balanced mix as nrl-ai/vn-diacritic-train — 350K Vietnamese Wikipedia (CC-BY-SA-4.0, hirine/wikipedia-vietnamese-1M296K-dataset) + 150K NFC-fixed Vietnamese news (CC-BY-4.0, tmnam20/Vietnamese-News-dedup).… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-train.texttext-generation100K<n<1M0 likes42 downloads5mo agoHugging Face20dougalldeepmind /2026-09-15-dataset-refresh-correction-audit Dataset refresh correction audit; not a training release field value experiment Zero-new-API correction of the incomplete refresh: 40 net independent exclusion reversals and one lossless completed-review parsing recovery. Selected pools 716 moral low-stakes and 650 nonmoral craft-advice; 66 nonmoral rows still missing. Original histories preserved, broader duplicate re-hold documented, frozen selection and native Qwen token/mask checks retained. Four saved-answer… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-correction-audit.tabularn<1K0 likes42 downloads7d agoHugging Face21emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-en PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face22schneiderkamplab /dfm11-folketingets-dokumenter-error-correction DFM11 Folketingets Dokumenter Error Correction This dataset is the fully audited DFM11 replacement for schneiderkamplab/dfm10-folketingets-dokumenter-error-correction. Every retained input was generated from its target using 1-8 declared synthetic OCR substitutions. Deterministic text-quality filtering was followed by a task-aware Gemma 4 audit of all 2,548,956 surviving rows; 63,109 audit rejections were removed and 2,485,847 rows remain. Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.text1M<n<10M0 likes39 downloads17d agoHugging Face23arthurdubrou /Bird_explained_corrections Dataset Card for Dataset Name This dataset is truncated textn<1K0 likes29 downloads3y agoHugging Face24nrl-ai /vn-spell-correction-eval nrl-ai/vn-spell-correction-eval Vietnamese spell-correction evaluation grid: 4 source registers × 2 noise levels = 8 splits, 2,098 (noisy, clean) sentence pairs total. Each pair is {"input": "<noisy>", "target": "<clean>"}. Both sides are NFC-normalized. The clean target is the same sentence used as the target in nrl-ai/vn-diacritic-eval — spell correction is a strict superset of diacritic restoration, so we reuse the same registers-balanced corpus. Splits Two noise… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval.texttext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face25emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-fr PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face26kowo-co /babble-corrections babble — corrections Training data for babble: a ~3M parameter byte-level transformer that started from random weights and has only ever learned from people correcting it in Discord. There is no pretraining corpus. There is no scraped chat history. Every row here is somebody deliberately teaching a small confused model to talk. How a row happens Someone @mentions the bot. The bot replies with whatever its current weights produce. Early on this is noise, and it is… See the full description on the dataset page: https://huggingface.co/datasets/kowo-co/babble-corrections.texttext-generationn<1K0 likes26 downloads1mo agoHugging Face27torinriley /spell-correction Spell-Check Dataset This dataset consists of pairs of misspelled words and their corresponding correctly spelled words, designed for training and evaluating character-level spelling correction models. It is particularly useful for tasks such as: Spelling correction Character-level sequence-to-sequence modeling Error detection and correction in text Each data point in the dataset contains: misspelled: A misspelled version of a word. correct: The corrected spelling of the word.… See the full description on the dataset page: https://huggingface.co/datasets/torinriley/spell-correction.text10K<n<100K2 likes25 downloads2y agoHugging Face28True2456 /gemma4-onpolicy-50topics-corrections Gemma 4 FrontierDistill - Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 FrontierDistill Project. This dataset contains 1,000 authentic on-policy student failure corrections collected live from gemma-4-12b-it-qat-frontierdistill across 50 distinct… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-corrections.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face29jiangchengchengNLP /RPEval_correctionThe dataset is an enhanced version of https://github.com/yelboudouri/RPEval.git The ratio of yes to no in the task type DECISION has been balanced from 5866:213 to 3067:3012. RPEval: Role-Playing Evaluation for Large Language Models This repository contains code and data referenced in: "Role-Playing Evaluation for Large Language Models". Large Language Models (LLMs) demonstrate a notable capacity for adopting personas and engaging in role-playing. However, evaluating… See the full description on the dataset page: https://huggingface.co/datasets/jiangchengchengNLP/RPEval_correction.textquestion-answering1K<n<10K0 likes19 downloads10mo agoHugging Face30LorthGyu /indonesian-grammar-correction Koreksi Tata Bahasa Indonesia ✍️ Kumpulan 126 pasangan kalimat (asli → koreksi) untuk grammar error correction bahasa Indonesia. Kenapa dataset ini ada? Grammar error correction (GEC) untuk bahasa Indonesia belum ada di HF — padahal model GEC global (271 downloads) jadi salah satu kategori paling dicari. Dataset ini isi gap itu: dari ejaan ("Dimana" → "Di mana"), ragam ("gw udah" → "saya sudah"), pleonasme ("para siswa-siswa" → "para siswa"), sampai huruf kapital.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-grammar-correction.texttext-generationn<1K0 likes19 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.