datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDocVQA-Corpus-ExtendedOpenDoc-Pdf-Preview
OpenDoc-Pdf-Preview
OpenDoc-Pdf-Preview is a compact visual preview dataset containing 6,000 high-resolution document images extracted from PDFs. This dataset is designed for Image-to-Text tasks such as document OCR pretraining, layout understanding, and multimodal document analysis.
Dataset Summary
Modality: Image-to-Text
Content Type: PDF-based document previews
Number of Samples: 6,000
Language: English
Format: Parquet
Split: train only
Size: 606 MB
License: Apache… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDoc-Pdf-Preview.OpenDocVQA-Corpus
Dataset Card for OpenDocVQA
This is a training and evaluation corpus data file for VDocRAG, a new RAG framework that can directly understand diverse real-world documents purely from visual features.
Dataset Description
OpenDocVQA is the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats.
Supported Tasks and Leaderboards
Given a large collection of document images and a question… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/OpenDocVQA-Corpus.Opendoc2-Analysis-Recognition
Opendoc2-Analysis-Recognition Dataset
Overview
The Opendoc2-Analysis-Recognition dataset is a collection of data designed for tasks involving image analysis and recognition. It is suitable for various machine learning tasks, including image-to-text conversion, text classification, and image feature extraction.
Dataset Details
Modalities: Likely includes images and associated labels (specific modalities can be confirmed on the dataset's page).
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Opendoc2-Analysis-Recognition.OpenDoc-Null-6K
OpenDoc-Null-6K
The OpenDoc-Null-6K dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks.
Attribute
Value
Task
Image-to-Text
Modality
Image
Format
ImageFolder
Language
English
License… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDoc-Null-6K.Opendoc1-Analysis-Recognition
Opendoc1-Analysis-Recognition Dataset
Overview
The Opendoc1-Analysis-Recognition dataset is designed for tasks involving image-to-text, text classification, and image feature extraction. It contains images paired with class labels, making it suitable for vision-language tasks.
Dataset Details
Modalities: Image
Languages: English
Size: Approximately 1,000 samples (n=1K)
Tags: image, analysis, vision-language
License: Apache 2.0
Tasks
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Opendoc1-Analysis-Recognition.
