CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay. Description All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.tabular10K<n<100K135 likes2.1k downloads1y agoHugging Face02acomquest /sanskrit-ocr-post-correction\ A Benchmark and Dataset for Post-OCR text correction in Sanskrit. This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script. It includes: - Train/Validation/Test splits with OCR text and corrected ground truth - An out-of-domain test set of 500 sentences - Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyayatext-classification100K<n<1M2 likes272 downloads1y agoHugging Face03AbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes120 downloads3mo agoHugging Face04GenSEC-LLM /SLT-Task1-Post-ASR-Text-Correction Dataset Name: Pilot dataset for Multi-domain ASR corrections Description This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains. It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0 Structure Data Split The dataset is divided into training and test splits: Training Data: 281,082 entries Approximately 6,255… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task1-Post-ASR-Text-Correction.text100K<n<1M2 likes114 downloads2y agoHugging Face05jeanflop /post-ocr-correction Synthetic OCR Correction Dataset This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.text1M<n<10M2 likes78 downloads2y agoHugging Face06FrancophonIA /ICDAR_2019_Competition_Post-OCR_Text_Correction [!NOTE] Dataset origin: https://zenodo.org/records/3515403 Corpus for the ICDAR2019 Competition on Post-OCR Text Correction (October 2019) => Website: http://l3i.univ-larochelle.fr/ICDAR2019PostOCR Description: The corpus accounts for 22M OCRed characters along with the corresponding Gold Standard (GS). The documents come from different digital collections available, among others, at the National Library of France (BnF) and the British Library (BL). The corresponding GS… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/ICDAR_2019_Competition_Post-OCR_Text_Correction.0 likes77 downloads1y agoHugging Face07jeanflop /post_ocr_correction2 Synthetic OCR Correction Dataset This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post_ocr_correction2.text1M<n<10M1 likes44 downloads2y agoHugging Face08emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-en PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face09buddhist-nlp /translation-postcorrectiontext1K<n<10K0 likes27 downloads2y agoHugging Face10emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-fr PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face11buddhist-nlp /sentence-alignment-merged-postcorrection Dataset Card for "sentence-alignment-merged-postcorrection" More Information needed text100K<n<1M0 likes21 downloads2y agoHugging Face12floriandebaene /EmDComF_OCR_post-correctionOCR-to-gold parallel dataset for OCR post-correcting early modern Dutch, originating from EmDComF. text10K<n<100K2 likes21 downloads1y agoHugging Face13chronbmm /sanskrit-ocr-postcorrectiontext100K<n<1M1 likes20 downloads2y agoHugging Face14LIACC /Emakhuwa-Portuguese-OCR-post-correctionBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.imagetranslationn<1K0 likes17 downloads2y agoHugging Face15emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-de0 likes16 downloads3mo agoHugging Face16jeanflop /post_ocr_correction-512 Synthetic OCR Correction Dataset <700 input length This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 1,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post_ocr_correction-512.text1M<n<10M0 likes13 downloads2y agoHugging Face17Alijeff1214 /Ocr_Post_Correctionimagen<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.