datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docvqa-single-page-questions
Dataset Card for DocVQA Dataset
Dataset Summary
DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images.
Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information.
Usage
This dataset can be used with current releases of Hugging Face datasets library.
Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.kangaroo_math_mc_questionsroco2-question-id-dataset
ROCOv2: Radiology Object in COntext version 2
Introduction
ROCOv2 is a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PMC Open Access Subset. It is an updated version of the ROCO dataset, adding 35,705 new images and improving concept extraction and filtering.
Dataset Overview
The ROCOv2 dataset contains 79,789 radiological images, each with a corresponding caption and medical concepts. The… See the full description on the dataset page: https://huggingface.co/datasets/Jiiwonn/roco2-question-id-dataset.dhivehi-vrd-batch-1-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
Art-Vision-Question-Answering-Dataset
Art Vision Question Answering Dataset
🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering.
Dataset Overview
This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks.
✨ Key Features
🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer
💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.dhivehi-vrd-batch-3-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-vrd-batch-6-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
video-game-question-answeringrocov2-questions-radiologyrvl-cdip-questionnaire⚠️ This only a subpart of the original dataset, containing only questionnaire.
The RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images. The images are sized so their largest dimension does not exceed 1000 pixels.
For questions and comments please contact Adam Harley (aharley@scs.ryerson.ca).
The full… See the full description on the dataset page: https://huggingface.co/datasets/chainyo/rvl-cdip-questionnaire.dhivehi-vrd-batch-2-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
visual-question-answering-cocomovie-frames-questionsGaoKao-questionsdhivehi-vrd-batch-4-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-vrd-batch-5-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
docvqa-single-page-questions-answer-ocr
DocVQA with Answer Localization
This dataset provides answer-localization annotations produced by our pipeline on top of the DocVQA dataset.
Usage
from datasets import load_dataset
# Load the dataset with answer OCR annotations
ds = load_dataset("indrehus/docvqa-single-page-questions-answer-ocr", split="validation")
# Get a single sample
sample = ds[0]
# Available fields in each sample:
print("Image:", sample["image"]) # PIL.Image
print("Question:"… See the full description on the dataset page: https://huggingface.co/datasets/indrehus/docvqa-single-page-questions-answer-ocr.translated_visual_puzzles_with_questiondhivehi-vrd-batch-2-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
question-answervsr_random_with_questionsroco2-question-dataset-validationindian-exam-questionsvsr_zeroshot_with_questionsdhivehi-vrd-batch-7-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
alevel-exam-questionsdriving-benchmark-questionsAP-exam-questionstranslated_mmiq_dataset_with_questioncauldron-question
