CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tinixai /ocr_annual_financials 📋 TiniX Vietnam OCR Annual Financial Statements (2015–2025) 📌 Overview TiniX Vietnam OCR Annual Financial Statements là bộ dữ liệu tài liệu tài chính tiếng Việt quy mô lớn, bao gồm: File PDF gốc của báo cáo tài chính thường niên Văn bản OCR được trích xuất tương ứng từ từng tài liệu Dữ liệu bao phủ các doanh nghiệp niêm yết tại Việt Nam trong giai đoạn 2015–2025, được thu thập và xử lý bởi TiniX AI. Bộ dữ liệu hiện bao gồm: 18.231 báo cáo tài chính thường… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/ocr_annual_financials.documentdocument-question-answering10K<n<100K26 likes3.3k downloads4mo agoHugging Face02SBB /sbb-dc-ocr Dataset Card for Berlin State Library OCR data Dataset Summary The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945. At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages. For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012). Supported Tasks and Leaderboards language-modeling: this dataset has the potential to be used… See the full description on the dataset page: https://huggingface.co/datasets/SBB/sbb-dc-ocr.textfill-mask1M<n<10M8 likes1.7k downloads4y agoHugging Face03vduydong /ocr_annual_financials 📋 TiniX Vietnam OCR Annual Financial Statements (2015–2025) 📌 Overview TiniX Vietnam OCR Annual Financial Statements là bộ dữ liệu văn bản OCR từ báo cáo tài chính thường niên của các doanh nghiệp niêm yết tại Việt Nam trong giai đoạn 2015–2025. Dữ liệu được thu thập và xử lý bởi TiniX AI bao gồm 18.231 báo cáo tiếng Việt tương ứng với 1491 mã chứng khoán khác nhau. 🧩 Data Description Mỗi file văn bản chứa toàn bộ nội dung OCR của một báo cáo tài chính.… See the full description on the dataset page: https://huggingface.co/datasets/vduydong/ocr_annual_financials.textdocument-question-answering10K<n<100K1 likes855 downloads4mo agoHugging Face04AdaMLLab /nyu-aco-ocr-full nyu-aco-ocr-full Scanned Arabic books from the Arabic Collections Online (ACO) archive, OCR'd page by page with the dots.ocr vision-language model served through vLLM. The archive spans seven partner collections: NYU, Princeton, Cornell, Columbia, AUB, AUC, and UAE National Archives. Each PDF page is rendered at 200 DPI, OCR'd individually, and the pages of a book are joined into one markdown document separated by \n\n---\n\n. Each row is one complete book. Schema… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/nyu-aco-ocr-full.texttext-generation10K<n<100K0 likes465 downloads1mo agoHugging Face05thuongdz /ocr_annual_financials 📋 TiniX Vietnam OCR Annual Financial Statements (2015–2025) 📌 Overview TiniX Vietnam OCR Annual Financial Statements là bộ dữ liệu văn bản OCR từ báo cáo tài chính thường niên của các doanh nghiệp niêm yết tại Việt Nam trong giai đoạn 2015–2025. Dữ liệu được thu thập và xử lý bởi TiniX AI bao gồm 18.231 báo cáo tiếng Việt tương ứng với 1491 mã chứng khoán khác nhau. 🧩 Data Description Mỗi file văn bản chứa toàn bộ nội dung OCR của một báo cáo… See the full description on the dataset page: https://huggingface.co/datasets/thuongdz/ocr_annual_financials.documentdocument-question-answering10K<n<100K1 likes423 downloads4mo agoHugging Face06racineai /ocr-pdf-degraded OCR-PDF-Degraded Dataset Overview This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments. Purpose Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.imagetext-generation10K<n<100K3 likes382 downloads1y agoHugging Face07nubuwwat /khatme-nubuwwat-ocr-dataset Khatme-Nubuwwat Urdu OCR Corpus This is a structure-aware, fully OCR'd text dataset of around 215 Urdu Khatme Nubuwat books/volumes (approximately 86,557 pages of text) The majority of the books were in Urdu Nastaliq font, with Arabic Naskh and English text present minimally as well. The text dataset is paried with source-page scans. The books are composed of Nastaliq prose with heavy references to Quran and Hadith. Effort was made to ensure that the OCR pipeline transcribed the… See the full description on the dataset page: https://huggingface.co/datasets/nubuwwat/khatme-nubuwwat-ocr-dataset.imagetext-generation10K<n<100K2 likes129 downloads1mo agoHugging Face08AbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes116 downloads3mo agoHugging Face09beatsprom /multimodal-vision-ocr-document-parsing-2026 📐 Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) This repository provides the official 100-sample production teaser of the Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights Vision-Language Models (Qwen2-VL, Pixtral-12B, Llama-3.2-Vision, ColPali) on dense document parsing, normalized spatial bounding boxes (<box>[ymin, xmin, ymax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-ocr-document-parsing-2026.textimage-to-textn<1K0 likes88 downloads21d agoHugging Face10fatihburakkaragoz /old-nogay-turkish-ocr-corpus Old Nogay Turkish OCR Corpus This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis. We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.tabulartext-generationn<1K0 likes77 downloads5mo agoHugging Face11christopherxzyx /ocr_financials_statements_2020_2025 📋 Vietnam Annual Financial Statements (2020–2025) (PDF & OCR) 📌 Overview OCR Vietnam Annual Financial Statements (2020–2025) là bộ dữ liệu báo cáo tài chính thường niên của các doanh nghiệp niêm yết tại Việt Nam. Đây là phiên bản được chọn lọc và tinh chỉnh từ bộ dữ liệu gốc TiniX Vietnam OCR Annual Financial Statements do TiniX AI thu thập. Điểm khác biệt của bộ dữ liệu này: Tập trung chuyên sâu vào giai đoạn mới nhất: 2020–2025. Cung cấp song song cả định… See the full description on the dataset page: https://huggingface.co/datasets/christopherxzyx/ocr_financials_statements_2020_2025.documentdocument-question-answering10K<n<100K0 likes62 downloads28d agoHugging Face12LR-AI-Labs /vi-OCR_VQA Dataset Card for "vi-OCR-VQA" imagevisual-question-answering10K<n<100K7 likes55 downloads2y agoHugging Face13AbhishekBhandari /Devanagari-OCR-ICL-Benchmark Devanagari Post-OCR Correction Benchmark A benchmark for evaluating post-OCR correction systems on Hindi and Marathi text rendered in Devanagari script, accompanying the paper "Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction" (Bhandari and Harit, 2026). Summary 20,000 evaluation pairs (10k Hindi, 10k Marathi) of (OCR-corrupted, ground-truth) sentences ~6,600 shot-bank pairs (~3.3k per language) for in-context example… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Devanagari-OCR-ICL-Benchmark.texttext-generation10K<n<100K1 likes53 downloads3mo agoHugging Face14Hemanth-thunder /ocr-data-tnpsc Tamil Public Domain Books (Tamil) The dataset comprises over 30 school textbooks and certain TNPSC (Tamil Nadu Public Service Commission) materials in Tamil medium, presumed to be in the public domain. texttext-generation1K<n<10K2 likes49 downloads2y agoHugging Face15Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads15d agoHugging Face16lkevincc0 /Akkadian-OCR-Corpus Akkadian OCR Corpus OCR-extracted corpus from Old Assyrian cuneiform tablet publications and Akkadian linguistic resources, designed for LLM pretraining on ancient Mesopotamian languages. Dataset Description This dataset contains OCR-processed text from academic publications on Old Assyrian studies, including cuneiform tablet transliterations, translations, and the Chicago Assyrian Dictionary (CAD). Splits Split Description Pages Characters OCR Quality… See the full description on the dataset page: https://huggingface.co/datasets/lkevincc0/Akkadian-OCR-Corpus.tabulartext-generation100K<n<1M0 likes42 downloads8mo agoHugging Face17emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-en PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.texttext-generation100K<n<1M0 likes39 downloads3mo agoHugging Face18agentlans /ocr-correction OCR (Optical Character Recognition) Correction Dataset This dataset comprises OCR-corrected text samples from English books and newspapers sourced from the Internet Archive. It provides pairs of raw OCR text and their AI-corrected versions, designed for OCR correction tasks. Dataset Structure Data Instances Each instance contains: input: Raw OCR text with errors output: Corrected text Example: { "input": "\n\n(ii) The income of Tarai and Bhabar… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ocr-correction.texttext-generation10K<n<100K1 likes36 downloads2y agoHugging Face19sarvamai /indic-ocr-bench Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material. The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.imageimage-to-text1K<n<10K2 likes36 downloads2d agoHugging Face20yifeihu /ACL-23-Paper-OCR-Markdown ACL 2023 Paper in Markdown after OCR This dataset contains 2150 papers from Association for Computational Linguistics (ACL) 2023: Long Papers (912 papers) Short Papers (185 papers) System Demonstrations (59 paper) Student Research Workshop (35 papers) Industry Track (77 papers) Tutorial Abstracts (7 papers) Findings (902 papers) This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project. OCR process The… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/ACL-23-Paper-OCR-Markdown.textsummarization1K<n<10K19 likes35 downloads2y agoHugging Face21fatihburakkaragoz /anadolu-ocr-corpus Anadolu OCR Corpus Anadolu OCR Corpus is an OpenCR export of OCR text and document metadata for 52 historical Ottoman Turkish, Turkish, and Arabic-containing PDF sources. The dataset is provided in two Hugging Face configs: pages: one row per source page, including page-level OCR text, metadata, validation status, script direction, language detection, hashes, and split labels. documents: one row per source document, including document-level concatenated text, markdown, aggregate… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/anadolu-ocr-corpus.tabulartext-generation10K<n<100K3 likes27 downloads4mo agoHugging Face22public-records-research /epstractor-ocr epstractor-ocr This dataset contains OCR-extracted text from documents processed using DeepSeek-OCR. Dataset Summary Total Records: 56,774 Total Source Files: 56,774 Format: Parquet (Snappy compression) Dataset Structure Data Fields source_file: Original source filename (without page numbers or extensions) page_number: Page number for multi-page documents (null for single-page images) split: Dataset split (train/test/validation) text:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-ocr.tabulartext-generation10K<n<100K0 likes26 downloads10mo agoHugging Face23FatimahEmadEldin /arabic-ocr-jokes Extracted Arabic Jokes Dataset Dataset Description This dataset contains Arabic jokes extracted and cleaned from various PDF booklets using OCR and LLM-based post-processing. It was created to support research in Arabic Humor Generation. Data Fields Joke_Text: The full text of the joke in Arabic. Source Extracted from public domain joke books. texttext-generationn<1K0 likes26 downloads8mo agoHugging Face24emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-fr PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.texttext-generation10K<n<100K0 likes25 downloads3mo agoHugging Face25OumarDicko /Bourse_Casablanca_Rapports_Financiers_Bruts_OCR Bourse Casablanca : Rapports Financiers Bruts (OCR) Présentation Ce dataset contient l'extraction OCR intégrale des Rapports Annuels et Financiers de 58 sociétés cotées à la Bourse de Casablanca (BVC). L'archive couvre l'intégralité de l'historique disponible sur le site officiel de la Bourse de Casablanca pour chaque émetteur, transformant des documents PDF complexes en données HTML exploitables. Chiffres clés Périmètre : 58 sociétés de la cote marocaine.… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bourse_Casablanca_Rapports_Financiers_Bruts_OCR.textfeature-extraction1K<n<10K0 likes24 downloads8mo agoHugging Face26aimgo /Latin-OCR-Artifacts Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG. OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations. If you use this in your work, please cite: @misc{mccarthy2025LACOROCR, author = {McCarthy, A. M.}, title = {{Latin OCR Artifacts}}, year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.tabulartext-generation100K<n<1M0 likes22 downloads8mo agoHugging Face27tensorlake /OCRBenchV2-DocParsing-UpdatedGTOCRBenchV2-DocParsing-UpdatedGT is an improved ground truth for the document parsing subset of OCRBench V2, created and verified by Tensorlake. This version is used in Tensorlake’s OCR and document understanding benchmark comparisons. The dataset is intended for evaluation and research purposes only.For the original benchmark and other subsets, please refer to OCRBench V2 . texttext-generationn<1K0 likes20 downloads11mo agoHugging Face28titou4ng /smolified-ocr-data-extract-and-compare 🤏 smolified-ocr-data-extract-and-compare Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extract-and-compare. 📦 Asset Details Origin: Smolify Foundry (Job ID: 790dd5fa) Records: 21835 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. texttext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face29LIACC /Emakhuwa-Portuguese-OCR-post-correctionBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.imagetranslationn<1K0 likes17 downloads2y agoHugging Face30JingweiNi /ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513 GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps. Summary Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 Source rows: 500 K2-Think Codeforces traces Source manifest-selected Qwen3.5 labels: 10,000 steps GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.tabulartext-generationn<1K0 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.