datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ocr_annual_financials
📋 TiniX Vietnam OCR Annual Financial Statements (2015–2025)
📌 Overview
TiniX Vietnam OCR Annual Financial Statements là bộ dữ liệu tài liệu tài chính tiếng Việt quy mô lớn, bao gồm:
File PDF gốc của báo cáo tài chính thường niên
Văn bản OCR được trích xuất tương ứng từ từng tài liệu
Dữ liệu bao phủ các doanh nghiệp niêm yết tại Việt Nam trong giai đoạn 2015–2025, được thu thập và xử lý bởi TiniX AI.
Bộ dữ liệu hiện bao gồm:
18.231 báo cáo tài chính thường… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/ocr_annual_financials.sbb-dc-ocr
Dataset Card for Berlin State Library OCR data
Dataset Summary
The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.
At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.
For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).
Supported Tasks and Leaderboards
language-modeling: this dataset has the potential to be used… See the full description on the dataset page: https://huggingface.co/datasets/SBB/sbb-dc-ocr.ocr_annual_financials
📋 TiniX Vietnam OCR Annual Financial Statements (2015–2025)
📌 Overview
TiniX Vietnam OCR Annual Financial Statements là bộ dữ liệu văn bản OCR từ báo cáo tài chính thường niên của các doanh nghiệp niêm yết tại Việt Nam trong giai đoạn 2015–2025. Dữ liệu được thu thập và xử lý bởi TiniX AI bao gồm 18.231 báo cáo tiếng Việt tương ứng với 1491 mã chứng khoán khác nhau.
🧩 Data Description
Mỗi file văn bản chứa toàn bộ nội dung OCR của một báo cáo tài chính.… See the full description on the dataset page: https://huggingface.co/datasets/vduydong/ocr_annual_financials.nyu-aco-ocr-full
nyu-aco-ocr-full
Scanned Arabic books from the Arabic Collections Online (ACO) archive, OCR'd page by page with the dots.ocr vision-language model served through vLLM. The archive spans seven partner collections: NYU, Princeton, Cornell, Columbia, AUB, AUC, and UAE National Archives.
Each PDF page is rendered at 200 DPI, OCR'd individually, and the pages of a book are joined into one markdown document separated by \n\n---\n\n. Each row is one complete book.
Schema… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/nyu-aco-ocr-full.ocr_annual_financials
📋 TiniX Vietnam OCR Annual Financial Statements (2015–2025)
📌 Overview
TiniX Vietnam OCR Annual Financial Statements là bộ dữ liệu văn bản OCR từ báo cáo tài chính thường niên của các doanh nghiệp niêm yết tại Việt Nam trong giai đoạn 2015–2025. Dữ liệu được thu thập và xử lý bởi TiniX AI bao gồm 18.231 báo cáo tiếng Việt tương ứng với 1491 mã chứng khoán khác nhau.
🧩 Data Description
Mỗi file văn bản chứa toàn bộ nội dung OCR của một báo cáo… See the full description on the dataset page: https://huggingface.co/datasets/thuongdz/ocr_annual_financials.ocr-pdf-degraded
OCR-PDF-Degraded Dataset
Overview
This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments.
Purpose
Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.khatme-nubuwwat-ocr-dataset
Khatme-Nubuwwat Urdu OCR Corpus
This is a structure-aware, fully OCR'd text dataset of around 215 Urdu Khatme Nubuwat books/volumes (approximately 86,557 pages of text) The majority of the books were in Urdu Nastaliq font, with Arabic Naskh and English text present minimally as well. The text dataset is paried with source-page scans.
The books are composed of Nastaliq prose with heavy references to Quran and Hadith. Effort was made to ensure that the OCR pipeline transcribed the… See the full description on the dataset page: https://huggingface.co/datasets/nubuwwat/khatme-nubuwwat-ocr-dataset.Indic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.multimodal-vision-ocr-document-parsing-2026
📐 Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026)
This repository provides the official 100-sample production teaser of the Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights Vision-Language Models (Qwen2-VL, Pixtral-12B, Llama-3.2-Vision, ColPali) on dense document parsing, normalized spatial bounding boxes (<box>[ymin, xmin, ymax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-ocr-document-parsing-2026.old-nogay-turkish-ocr-corpus
Old Nogay Turkish OCR Corpus
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.
We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.ocr_financials_statements_2020_2025
📋 Vietnam Annual Financial Statements (2020–2025) (PDF & OCR)
📌 Overview
OCR Vietnam Annual Financial Statements (2020–2025) là bộ dữ liệu báo cáo tài chính thường niên của các doanh nghiệp niêm yết tại Việt Nam. Đây là phiên bản được chọn lọc và tinh chỉnh từ bộ dữ liệu gốc TiniX Vietnam OCR Annual Financial Statements do TiniX AI thu thập.
Điểm khác biệt của bộ dữ liệu này:
Tập trung chuyên sâu vào giai đoạn mới nhất: 2020–2025.
Cung cấp song song cả định… See the full description on the dataset page: https://huggingface.co/datasets/christopherxzyx/ocr_financials_statements_2020_2025.vi-OCR_VQA
Dataset Card for "vi-OCR-VQA"
Devanagari-OCR-ICL-Benchmark
Devanagari Post-OCR Correction Benchmark
A benchmark for evaluating post-OCR correction systems on Hindi and Marathi
text rendered in Devanagari script, accompanying the paper
"Evaluating In-Context Learning and Retrieval Strategies for Devanagari
Post-OCR Correction" (Bhandari and Harit, 2026).
Summary
20,000 evaluation pairs (10k Hindi, 10k Marathi) of (OCR-corrupted, ground-truth) sentences
~6,600 shot-bank pairs (~3.3k per language) for in-context example… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Devanagari-OCR-ICL-Benchmark.ocr-data-tnpsc
Tamil Public Domain Books (Tamil)
The dataset comprises over 30 school textbooks and certain TNPSC (Tamil Nadu Public Service Commission) materials in Tamil medium, presumed to be in the public domain.
Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.Akkadian-OCR-Corpus
Akkadian OCR Corpus
OCR-extracted corpus from Old Assyrian cuneiform tablet publications and Akkadian linguistic resources, designed for LLM pretraining on ancient Mesopotamian languages.
Dataset Description
This dataset contains OCR-processed text from academic publications on Old Assyrian studies, including cuneiform tablet transliterations, translations, and the Chicago Assyrian Dictionary (CAD).
Splits
Split
Description
Pages
Characters
OCR Quality… See the full description on the dataset page: https://huggingface.co/datasets/lkevincc0/Akkadian-OCR-Corpus.pleias-post-ocr-correction-chonkie-aligned-en
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.ocr-correction
OCR (Optical Character Recognition) Correction Dataset
This dataset comprises OCR-corrected text samples from English books and newspapers sourced from the Internet Archive. It provides pairs of raw OCR text and their AI-corrected versions, designed for OCR correction tasks.
Dataset Structure
Data Instances
Each instance contains:
input: Raw OCR text with errors
output: Corrected text
Example:
{
"input": "\n\n(ii) The income of Tarai and Bhabar… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ocr-correction.indic-ocr-bench
Sarvam Indic OCR Bench
Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.
The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.ACL-23-Paper-OCR-Markdown
ACL 2023 Paper in Markdown after OCR
This dataset contains 2150 papers from Association for Computational Linguistics (ACL) 2023:
Long Papers (912 papers)
Short Papers (185 papers)
System Demonstrations (59 paper)
Student Research Workshop (35 papers)
Industry Track (77 papers)
Tutorial Abstracts (7 papers)
Findings (902 papers)
This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project.
OCR process
The… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/ACL-23-Paper-OCR-Markdown.anadolu-ocr-corpus
Anadolu OCR Corpus
Anadolu OCR Corpus is an OpenCR export of OCR text and document metadata for 52 historical Ottoman Turkish, Turkish, and Arabic-containing PDF sources. The dataset is provided in two Hugging Face configs:
pages: one row per source page, including page-level OCR text, metadata, validation status, script direction, language detection, hashes, and split labels.
documents: one row per source document, including document-level concatenated text, markdown, aggregate… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/anadolu-ocr-corpus.epstractor-ocr
epstractor-ocr
This dataset contains OCR-extracted text from documents processed using DeepSeek-OCR.
Dataset Summary
Total Records: 56,774
Total Source Files: 56,774
Format: Parquet (Snappy compression)
Dataset Structure
Data Fields
source_file: Original source filename (without page numbers or extensions)
page_number: Page number for multi-page documents (null for single-page images)
split: Dataset split (train/test/validation)
text:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-ocr.arabic-ocr-jokes
Extracted Arabic Jokes Dataset
Dataset Description
This dataset contains Arabic jokes extracted and cleaned from various PDF booklets using OCR and LLM-based post-processing. It was created to support research in Arabic Humor Generation.
Data Fields
Joke_Text: The full text of the joke in Arabic.
Source
Extracted from public domain joke books.
pleias-post-ocr-correction-chonkie-aligned-fr
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.Bourse_Casablanca_Rapports_Financiers_Bruts_OCR
Bourse Casablanca : Rapports Financiers Bruts (OCR)
Présentation
Ce dataset contient l'extraction OCR intégrale des Rapports Annuels et Financiers de 58 sociétés cotées à la Bourse de Casablanca (BVC).
L'archive couvre l'intégralité de l'historique disponible sur le site officiel de la Bourse de Casablanca pour chaque émetteur, transformant des documents PDF complexes en données HTML exploitables.
Chiffres clés
Périmètre : 58 sociétés de la cote marocaine.… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bourse_Casablanca_Rapports_Financiers_Bruts_OCR.Latin-OCR-Artifacts
Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG.
OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations.
If you use this in your work, please cite:
@misc{mccarthy2025LACOROCR,
author = {McCarthy, A. M.},
title = {{Latin OCR Artifacts}},
year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.OCRBenchV2-DocParsing-UpdatedGTOCRBenchV2-DocParsing-UpdatedGT is an improved ground truth for the document parsing subset of OCRBench V2, created and verified by Tensorlake.
This version is used in Tensorlake’s OCR and document understanding benchmark comparisons.
The dataset is intended for evaluation and research purposes only.For the original benchmark and other subsets, please refer to OCRBench V2
.
smolified-ocr-data-extract-and-compare
🤏 smolified-ocr-data-extract-and-compare
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extract-and-compare.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 790dd5fa)
Records: 21835
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
Emakhuwa-Portuguese-OCR-post-correctionBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513
GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps
This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps.
Summary
Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513
Source rows: 500 K2-Think Codeforces traces
Source manifest-selected Qwen3.5 labels: 10,000 steps
GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.
