datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Openpdf-Analysis-Recognition
Openpdf-Analysis-Recognition
The Openpdf-Analysis-Recognition dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks.
Attribute
Value
Task
Image-to-Text
Modality
Image
Format
ImageFolder… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Analysis-Recognition.Openpdf-MultiReceipt-1K
Openpdf-MultiReceipt-1K
Openpdf-MultiReceipt-1K is a dataset consisting of over 1,000 receipt documents in PDF format. This dataset is designed for use in image-to-text and document understanding tasks, particularly Optical Character Recognition (OCR), receipt parsing, and layout analysis.
Notes
No text annotations or metadata are provided — only the raw PDFs.
Ideal for tasks requiring raw document inputs like PDF-to-Text pipelines.
Dataset Summary
Size: 1… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-MultiReceipt-1K.Openpdf-Blank-v2.0
Openpdf-Blank-v2.0
Openpdf-Blank-v2.0 is a small dataset containing blank or near-blank PDF image samples. This dataset is primarily designed to help train and evaluate document processing models, especially in tasks like:
Identifying and filtering blank or noise-filled documents.
Preprocessing stages for OCR pipelines.
Receipt/document classification tasks.
Dataset Structure
Modality: Image
Languages: English (if applicable)
Size: Less than 1,000 samples
License:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Blank-v2.0.Openpdf-Blank-v2.0-Sample
Openpdf-Blank-v2.0-Sample
Openpdf-Blank-v2.0-Sample is a sample dataset of blank or near-blank invoice and receipt documents. It contains 255 high-resolution scanned images extracted and cleaned from document PDFs. This dataset is intended to support training and evaluation of OCR, document classification, and layout-based filtering models where blank or structurally minimal pages must be identified and processed.
Dataset Summary
Format: Parquet (auto-converted)… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Blank-v2.0-Sample.
