CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cyberfish /text_error_correction文本纠错的相关数据 1 likes3.7k downloads5y agoHugging Face028uBob /text_error_correction文本纠错的相关数据 0 likes259 downloads2mo agoHugging Face03google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes130 downloads3y agoHugging Face04muzaffercky /kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models, like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections. The source videos are documented in the source.txt file. Usage from datasets import load_dataset dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train") print(dataset) textn<1K1 likes86 downloads1y agoHugging Face05p208p2002 /zhtw-sentence-error-correction 中文錯字糾正資料集 由規則與字典自維基百科產生的錯誤糾正資料集。 包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。 資料集使用函式庫: p208p2002/zh-mistake-text-gen 子集 alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。 beta: 50%錯誤,50%不變。單句中僅有一個錯誤。 gamma: 100%錯誤。單句中可能有多個錯誤。 text100K<n<1M5 likes63 downloads3y agoHugging Face06sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes53 downloads6d agoHugging Face07schneiderkamplab /dfm10-folketingets-dokumenter-error-correction dfm10-folketingets-dokumenter-error-correction Audited folketingets-dokumenter-error-correction tasks derived from Folketing documents. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 13 Rows: 3,105,440 Category: Danish transformation Upstream material Rigsarkivet handover 14004 / Folketinget Processing The complete generated task… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.0 likes52 downloads23d agoHugging Face08yammdd /vietnamese-error-correction-corpus Data Summary The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets. • Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text. • Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs. • Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.text10K<n<100K0 likes51 downloads2mo agoHugging Face09SyntheticLogic-Labs /python-runtime-verified-error-correction Python Runtime-Verified Error Correction Dataset 🐍⚡ Overview Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution. Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.texttext-generation1K<n<10K0 likes51 downloads9mo agoHugging Face10sarayusapa /Grammar_Error_Correctiontext100K<n<1M0 likes45 downloads1y agoHugging Face11Karthi02 /grammatical_error_correction0 likes44 downloads1y agoHugging Face12ClarusC64 /quantum-error-correction-failure-v0.1 quantum-error-correction-failure-v0.1 What this dataset does This dataset evaluates whether models can detect instability in quantum error correction regimes. Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction. The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency. Core stability idea Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.tabulartabular-classificationn<1K0 likes41 downloads5mo agoHugging Face13schneiderkamplab /dfm11-folketingets-dokumenter-error-correction DFM11 Folketingets Dokumenter Error Correction This dataset is the fully audited DFM11 replacement for schneiderkamplab/dfm10-folketingets-dokumenter-error-correction. Every retained input was generated from its target using 1-8 declared synthetic OCR substitutions. Deterministic text-quality filtering was followed by a task-aware Gemma 4 audit of all 2,548,956 surviving rows; 63,109 audit rejections were removed and 2,485,847 rows remain. Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.text1M<n<10M0 likes39 downloads17d agoHugging Face14slone /bak_ocr_error_correction_2022 Dataset Card for "bak_ocr_error_correction_2022" More Information needed text10K<n<100K0 likes33 downloads3y agoHugging Face15Karthi02 /grammatical-error-correction0 likes32 downloads1y agoHugging Face16bmd1905 /vi-error-correction-v2text100K<n<1M3 likes27 downloads2y agoHugging Face17bmd1905 /vi-error-correction-2.0text1M<n<10M1 likes24 downloads2y agoHugging Face18shahidul034 /error_correction_model_dataset_raw Dataset Card for "error_correction_model_dataset" More Information needed text1M<n<10M0 likes15 downloads4y agoHugging Face19bmd1905 /error-correction-vitext100K<n<1M8 likes14 downloads4y agoHugging Face20Shoriful025 /quantum_error_correction_telemetrytabularn<1K0 likes14 downloads8mo agoHugging Face21vosap52 /robot-error-correction-tr-v1 Robot Error Correction TR v1 This dataset focuses on failure detection and corrective behavior in embodied AI systems. Unlike standard instruction datasets, each sample represents: an incorrect real-world outcome a corrective decision The goal is improving humanoid robot autonomy and reliability in real environments. Capabilities trained: self-correction safety awareness environment feedback handling recovery planning textroboticsn<1K0 likes12 downloads7mo agoHugging Face22SPEAK-PP /synthetic-error-generated-spelling-correction-dataset-100ktext10K<n<100K0 likes12 downloads6mo agoHugging Face23sumitaryal /nepali_grammatical_error_correctiontext1M<n<10M0 likes11 downloads2y agoHugging Face24kilicai /turkish-sft-error-correction-10k kilicai/turkish-sft-error-correction-10k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-error-correction-10k') text10K<n<100K1 likes11 downloads4mo agoHugging Face25kilicai /turkish-sft-error_correction_20k kilicai/turkish-sft-error_correction_20k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-error_correction_20k') text10K<n<100K0 likes11 downloads4mo agoHugging Face26reverendish /advanced-math-error-correctiontext100K<n<1M1 likes7 downloads4mo agoHugging Face27duyle2408 /vi-error-correctiontext1K<n<10K0 likes4 downloads1y agoHugging Face28duyle2408 /vi-error-correction-super-smalltextn<1K0 likes3 downloads1y agoHugging Face29duyle2408 /vi-error-correction-super-small-upper-3text1K<n<10K0 likes3 downloads1y agoHugging Face30duyle2408 /vi-error-correction-2text10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.