datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QueST-PartNetMobility-SAPIEN
QueST: PartNet-Mobility SAPIEN Simulation Dataset
This dataset accompanies the paper:
QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon TrackingMayank Anand, Mohammad Saqlain, Kyan Mahajan, Priya Shukla, G.C Nandi, Andrew MelnikCAO Workshop at ICLR 2026
What Is This Dataset?
Synchronized RGB-D simulation sequences rendered in SAPIEN from PartNet-Mobility articulated objects, designed to stress-test long-horizon point… See the full description on the dataset page: https://huggingface.co/datasets/AnandMayank/QueST-PartNetMobility-SAPIEN.docvqa-single-page-questions
Dataset Card for DocVQA Dataset
Dataset Summary
DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images.
Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information.
Usage
This dataset can be used with current releases of Hugging Face datasets library.
Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.questFish2024
Dataset Card for QUEST Fish 2024
Images collected by teachers during a QUEST workshop. In 2024, the images were of fish collected from bodies of water near Princeton University.
Dataset Details
Dataset Structure
/dataset/
<folder>/
<img_id 1>.png
<img_id 2>.png
...
<img_id n>.png
...
<img_id 1>.png
<img_id 2>.png
...
<img_id n>.png
fieldData2024.csv
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/questFish2024.kangaroo_math_mc_questionsjee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.roco2-question-id-dataset
ROCOv2: Radiology Object in COntext version 2
Introduction
ROCOv2 is a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PMC Open Access Subset. It is an updated version of the ROCO dataset, adding 35,705 new images and improving concept extraction and filtering.
Dataset Overview
The ROCOv2 dataset contains 79,789 radiological images, each with a corresponding caption and medical concepts. The… See the full description on the dataset page: https://huggingface.co/datasets/Jiiwonn/roco2-question-id-dataset.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.dhivehi-vrd-batch-1-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
Art-Vision-Question-Answering-Dataset
Art Vision Question Answering Dataset
🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering.
Dataset Overview
This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks.
✨ Key Features
🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer
💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.jee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/soughed/jee-main-questions.QuestRoomScanTurkish-medical-visual-question-answering-LLaVa-dataset
Türkçe Radyoloji Görüntüleme Veri Seti - data_RAD
data_RAD veri seti, radyoloji görüntüleri üzerinde görsel soru-cevaplama (VQA) araştırmaları yapmak amacıyla Türkçeye çevrilmiş ve LLaVa mimarisiyle uyumlu hale getirilmiştir. Bu veri seti, tıbbi görüntü analizi ve yapay zeka destekli radyoloji uygulamalarını geliştirmek için kullanılabilir.
Veri Seti İçeriği
Toplam Görüntü Sayısı: 316
Veri Yapısı: DatasetDict({ train: Dataset({ features: ['image'], num_rows: 316 }) })
Özellikler:… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/Turkish-medical-visual-question-answering-LLaVa-dataset.dhivehi-vrd-batch-3-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
SeeTRUE-Feedback
Dataset Card for SeeTRUE-Feedback
Dataset Description
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Fields
Data Splits
Dataset Creation
Licensing Information
Citation Information
Dataset Description
The SeeTRUE-Feedback dataset is a diverse benchmark for the meta-evaluation of image-text matching/alignment feedback. It aims to overcome limitations in current benchmarks, which primarily focus on predicting a matching score between 0-1.… See the full description on the dataset page: https://huggingface.co/datasets/mismatch-quest/SeeTRUE-Feedback.quest-bowling-ball-yolo
Quest Bowling Ball YOLO Dataset
This dataset contains YOLO-format bounding-box annotations for detecting a bowling ball in Quest mixed-reality bowling footage.
The dataset was created for a live Quest-to-laptop bowling replay pipeline. YOLO is used only to find the first reliable ball seed; SAM2 then takes over for mask tracking and trajectory reconstruction.
Dataset Summary
Task: single-class object detection
Class: bowling_ball
Format: YOLO detection labels
Images: 801… See the full description on the dataset page: https://huggingface.co/datasets/sri299792458/quest-bowling-ball-yolo.dhivehi-vrd-batch-6-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
2-000-Physics-Chemistry-And-Biology-Questions-With-Image-Explanations-Dataset
2,000 Physics, Chemistry, and Biology Questions with Image Explanations Dataset
This collection contains 2,000 high-quality physics, chemistry, and biology questions, featuring a core modality of original images paired with text explanations. The data covers various formats, including multiple-choice, fill-in-the-blanks, experimental, and calculation questions. Each record provides the original question image, precise OCR-extracted text, and detailed step-by-step textual… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/2-000-Physics-Chemistry-And-Biology-Questions-With-Image-Explanations-Dataset.video-game-question-answeringrocov2-questions-radiologydhivehi-vrd-batch-2-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
rvl-cdip-questionnaire⚠️ This only a subpart of the original dataset, containing only questionnaire.
The RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images. The images are sized so their largest dimension does not exceed 1000 pixels.
For questions and comments please contact Adam Harley (aharley@scs.ryerson.ca).
The full… See the full description on the dataset page: https://huggingface.co/datasets/chainyo/rvl-cdip-questionnaire.Questrovisual-question-answering-cocodhivehi-vrd-batch-5-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-vrd-batch-4-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
bonus-one-question-pipeline
MIS 752 bonus: the whole pipeline in five cells
Optional · 50 extra credit points · about 20 minutes. Also the notebook to run when Hugging Face
is not working for you: it tests each piece and prints a report you can send Dr. Young.
Get the notebook and submit here: https://unlv.instructure.com/courses/214078/assignments/2856035
(download Bonus_One_Question_Pipeline.ipynb from this repo or from WebCampus, then in Colab: File → Upload notebook).
Five cells, top to bottom. Each… See the full description on the dataset page: https://huggingface.co/datasets/MIS-752/bonus-one-question-pipeline.translated_visual_puzzles_with_questiondocvqa-single-page-questions-answer-ocr
DocVQA with Answer Localization
This dataset provides answer-localization annotations produced by our pipeline on top of the DocVQA dataset.
Usage
from datasets import load_dataset
# Load the dataset with answer OCR annotations
ds = load_dataset("indrehus/docvqa-single-page-questions-answer-ocr", split="validation")
# Get a single sample
sample = ds[0]
# Available fields in each sample:
print("Image:", sample["image"]) # PIL.Image
print("Question:"… See the full description on the dataset page: https://huggingface.co/datasets/indrehus/docvqa-single-page-questions-answer-ocr.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-advanced-questions.quest-under-capricorn
Dataset Card for "quest-under-capricorn"
TODO: upload blip2 captions
update readme with tSNE
UPDATE README
change to darc-ai-quc
More Information needed
