datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.zhtw-sentence-error-correction
中文錯字糾正資料集
由規則與字典自維基百科產生的錯誤糾正資料集。
包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。
資料集使用函式庫: p208p2002/zh-mistake-text-gen
子集
alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。
beta: 50%錯誤,50%不變。單句中僅有一個錯誤。
gamma: 100%錯誤。單句中可能有多個錯誤。
urdu-asr-error-correction-data
Urdu ASR Generative Error Correction Dataset
This dataset contains paired training and testing data for post-ASR error correction in Urdu.
Dataset Details
Language: Urdu (ur)
Task: ASR Error Correction
License: CC BY-NC 4.0
Dataset Structure
The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold).
train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.python-runtime-verified-error-correction
Python Runtime-Verified Error Correction Dataset 🐍⚡
Overview
Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution.
Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.dfm11-folketingets-dokumenter-error-correction
DFM11 Folketingets Dokumenter Error Correction
This dataset is the fully audited DFM11 replacement for
schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.
Every retained input was generated from its target using 1-8 declared
synthetic OCR substitutions. Deterministic text-quality filtering was followed
by a task-aware Gemma 4 audit of all 2,548,956 surviving
rows; 63,109 audit rejections were removed and
2,485,847 rows remain.
Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.robot-error-correction-tr-v1
Robot Error Correction TR v1
This dataset focuses on failure detection and corrective behavior in embodied AI systems.
Unlike standard instruction datasets, each sample represents:
an incorrect real-world outcome
a corrective decision
The goal is improving humanoid robot autonomy and reliability in real environments.
Capabilities trained:
self-correction
safety awareness
environment feedback handling
recovery planning
vi-error-correctionvi-error-correction-super-smallvi-error-correction-super-small-upper-3vi-error-correction-2vi-error-correction-6vi-error-correction-super-small-uppervi-error-correction-super-small-upper-2vi-error-correction-super-small-trashvi-error-correction-11vi-error-correction-4vi-error-correction-5vi-error-correction-9vi-error-correction-small-10
