post-correction
Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias.
Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay.
Description
All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.sanskrit-ocr-post-correction\
A Benchmark and Dataset for Post-OCR text correction in Sanskrit.
This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script.
It includes:
- Train/Validation/Test splits with OCR text and corrected ground truth
- An out-of-domain test set of 500 sentences
- Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and GoladhyayaIndic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.SLT-Task1-Post-ASR-Text-Correction
Dataset Name: Pilot dataset for Multi-domain ASR corrections
Description
This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains.
It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0
Structure
Data Split
The dataset is divided into training and test splits:
Training Data: 281,082 entries
Approximately 6,255… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task1-Post-ASR-Text-Correction.post-ocr-correction
Synthetic OCR Correction Dataset
This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction.
Description
To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.ICDAR_2019_Competition_Post-OCR_Text_Correction
[!NOTE]
Dataset origin: https://zenodo.org/records/3515403
Corpus for the ICDAR2019 Competition on Post-OCR Text Correction (October 2019)
=> Website: http://l3i.univ-larochelle.fr/ICDAR2019PostOCR
Description:
The corpus accounts for 22M OCRed characters along with the corresponding Gold Standard (GS). The documents come from different digital collections available, among others, at the National Library of France (BnF) and the British Library (BL). The corresponding GS… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/ICDAR_2019_Competition_Post-OCR_Text_Correction.
