datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text_error_correction文本纠错的相关数据
text_error_correction文本纠错的相关数据
red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models,
like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected
from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections.
The source videos are documented in the source.txt file.
Usage
from datasets import load_dataset
dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train")
print(dataset)
zhtw-sentence-error-correction
中文錯字糾正資料集
由規則與字典自維基百科產生的錯誤糾正資料集。
包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。
資料集使用函式庫: p208p2002/zh-mistake-text-gen
子集
alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。
beta: 50%錯誤,50%不變。單句中僅有一個錯誤。
gamma: 100%錯誤。單句中可能有多個錯誤。
urdu-asr-error-correction-data
Urdu ASR Generative Error Correction Dataset
This dataset contains paired training and testing data for post-ASR error correction in Urdu.
Dataset Details
Language: Urdu (ur)
Task: ASR Error Correction
License: CC BY-NC 4.0
Dataset Structure
The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold).
train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.dfm10-folketingets-dokumenter-error-correction
dfm10-folketingets-dokumenter-error-correction
Audited folketingets-dokumenter-error-correction tasks derived from Folketing documents.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 13
Rows: 3,105,440
Category: Danish transformation
Upstream material
Rigsarkivet handover 14004 / Folketinget
Processing
The complete generated task… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.vietnamese-error-correction-corpus
Data Summary
The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets.
• Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text.
• Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs.
• Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.python-runtime-verified-error-correction
Python Runtime-Verified Error Correction Dataset 🐍⚡
Overview
Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution.
Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.Grammar_Error_Correctiongrammatical_error_correctionquantum-error-correction-failure-v0.1
quantum-error-correction-failure-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum error correction regimes.
Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction.
The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency.
Core stability idea
Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.dfm11-folketingets-dokumenter-error-correction
DFM11 Folketingets Dokumenter Error Correction
This dataset is the fully audited DFM11 replacement for
schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.
Every retained input was generated from its target using 1-8 declared
synthetic OCR substitutions. Deterministic text-quality filtering was followed
by a task-aware Gemma 4 audit of all 2,548,956 surviving
rows; 63,109 audit rejections were removed and
2,485,847 rows remain.
Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.bak_ocr_error_correction_2022
Dataset Card for "bak_ocr_error_correction_2022"
More Information needed
grammatical-error-correctionvi-error-correction-v2vi-error-correction-2.0error_correction_model_dataset_raw
Dataset Card for "error_correction_model_dataset"
More Information needed
error-correction-viquantum_error_correction_telemetryrobot-error-correction-tr-v1
Robot Error Correction TR v1
This dataset focuses on failure detection and corrective behavior in embodied AI systems.
Unlike standard instruction datasets, each sample represents:
an incorrect real-world outcome
a corrective decision
The goal is improving humanoid robot autonomy and reliability in real environments.
Capabilities trained:
self-correction
safety awareness
environment feedback handling
recovery planning
synthetic-error-generated-spelling-correction-dataset-100knepali_grammatical_error_correctionturkish-sft-error-correction-10k
kilicai/turkish-sft-error-correction-10k
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('kilicai/turkish-sft-error-correction-10k')
turkish-sft-error_correction_20k
kilicai/turkish-sft-error_correction_20k
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('kilicai/turkish-sft-error_correction_20k')
advanced-math-error-correctionvi-error-correctionvi-error-correction-super-smallvi-error-correction-super-small-upper-3vi-error-correction-2
