datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pubtables-qa
PubTables-QA: A Multi-Page Document Table QA Benchmark
PubTables-QA is a benchmark for evaluating vision-language models on document-level table question answering over multi-page scientific papers. Questions require understanding tables that span multiple pages, cross-referencing multiple tables, and jointly reasoning over tables and surrounding text.
Dataset Summary
Count
QA pairs
2,106
Documents
276
Page images
4,151
Avg. pages per document
~15… See the full description on the dataset page: https://huggingface.co/datasets/pubpub/pubtables-qa.mimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.motif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.xbd-damage-qa
xBD Damage Inventory QA Dataset
Overview
A question-answering dataset for building damage assessment
from post-disaster satellite imagery, built on the
xBD / xView2 dataset.
Dataset Statistics
300 scenes selected across 5 damage buckets
1,656 QA samples covering 8 question templates
100% validation pass rate
Disaster types: Wind, Flooding, Fire, Tsunami, Volcano, Earthquake
QA Templates
Template
Description
Difficulty
XBD-Q1
Full damage… See the full description on the dataset page: https://huggingface.co/datasets/DakshJ27/xbd-damage-qa.OpenUAV-QA
✨OpenUAV-QA✨
OpenUAV-QA is a large-scale multiple-choice question-answering benchmark for UAV (drone) navigation decision-making, built upon the TravelUAV (OpenUAV) dataset. It transforms raw UAV flight trajectories into structured, text-polished 4-option QA pairs that test a multimodal model's ability to reason about path planning, action sequences, and spatial dynamics from first-person drone video footage.
The dataset covers 22 distinct simulated… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/OpenUAV-QA.JavaError-QA
JErrRAG-Eval-800
JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record.
This Hugging Face repository contains:
java_error_qa_v2/: the canonical public benchmark package
paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles
SHA256SUMS.txt: release-side hash anchors referenced by the paper
Dataset Summary
Total records: 800
Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.vcr-qa-llavabrazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.adaption-practical-life-advice-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-practical_life_advice_qa
This dataset consists of prompt-completion pairs offering practical advice on diverse personal and professional topics such as workplace etiquette, career changes, financial decisions, and wealth management. The responses provide concise, actionable guidance tailored to specific user scenarios involving economic constraints or life transitions. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/Mariia1234/adaption-practical-life-advice-qa.reddit-finance-qa-json
Dataset Overview
This repository contains files used in the fine-tuning and retrieval-augmented generation (RAG) system built on Reddit finance data. Check out the Github repo to use this data here
reddit_finance_qa.jsonl
This is a JSON Lines (jsonl) file containing cleaned and deduplicated Reddit question-answer (QA) pairs from finance-related subreddits such as:
r/personalfinance
r/investing
r/wallstreetbets
r/cryptocurrency
r/stocks
Format (One… See the full description on the dataset page: https://huggingface.co/datasets/egupta/reddit-finance-qa-json.
