CoolFace
15 results

post-correction

PleIAs /Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay. Description All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.tabular10K<n<100K135 likes2.1k downloads1y agoHugging Faceacomquest /sanskrit-ocr-post-correction\ A Benchmark and Dataset for Post-OCR text correction in Sanskrit. This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script. It includes: - Train/Validation/Test splits with OCR text and corrected ground truth - An out-of-domain test set of 500 sentences - Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyayatext-classification100K<n<1M2 likes272 downloads1y agoHugging FaceAbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes120 downloads3mo agoHugging FaceGenSEC-LLM /SLT-Task1-Post-ASR-Text-Correction Dataset Name: Pilot dataset for Multi-domain ASR corrections Description This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains. It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0 Structure Data Split The dataset is divided into training and test splits: Training Data: 281,082 entries Approximately 6,255… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task1-Post-ASR-Text-Correction.text100K<n<1M2 likes114 downloads2y agoHugging Facejeanflop /post-ocr-correction Synthetic OCR Correction Dataset This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.text1M<n<10M2 likes78 downloads2y agoHugging FaceFrancophonIA /ICDAR_2019_Competition_Post-OCR_Text_Correction [!NOTE] Dataset origin: https://zenodo.org/records/3515403 Corpus for the ICDAR2019 Competition on Post-OCR Text Correction (October 2019) => Website: http://l3i.univ-larochelle.fr/ICDAR2019PostOCR Description: The corpus accounts for 22M OCRed characters along with the corresponding Gold Standard (GS). The documents come from different digital collections available, among others, at the National Library of France (BnF) and the British Library (BL). The corresponding GS… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/ICDAR_2019_Competition_Post-OCR_Text_Correction.0 likes77 downloads1y agoHugging Face