datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Target-QA
🎯 Target-QA: The First QA Dataset Benchmarking Target Priorization Based on DepMap
📑 Dataset Summary
Target-QA is derived from the DepMap multi-omics and CRISPR screening cohorts, harmonized via BioMedGraphica.It enables multi-modal reasoning by combining numeric evidence, topological knowledge and language context for CRISPR target prioritization.
This dataset supports the training and benchmarking of… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/Target-QA.Data-Analytics-Digital-Marketing-Project-Management-QA_DBBhagavad-Gita-QA
Bhagavad-Gita-QA-Multilingual
Dataset Summary
Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework.
This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita-QA.UIS-QA
UIS-QA: A Benchmark for Unindexed Information Seeking
Figure 1. UIS problem. Standard agents (bottom) rely on indexed information and often fail or hallucinate; UIS-capable agents (top) use additional tools to excavate unindexed information and solve UIS tasks.
If .figs do not load, see the paper.
🔔 News
[2026.03.10] 🎉 We release the UIS-QA dataset and the paper (ICLR 2026, arXiv) today!
📋 Dataset Description
Homepage
Paper… See the full description on the dataset page: https://huggingface.co/datasets/UIS-Digger/UIS-QA.patient-doctor-qa-tr-321179
Patient Doctor Q&A TR 321179 Veri Kümesi
Patient Doctor Q&A TR 321179 veri kümesi, Patient Doctor Q&A TR 19583, Patient Doctor Q&A TR 167732, Patient Doctor Q&A TR 5695 ve Patient Doctor Q&A TR 95588 veri kümelerinin birleştirilmiş ve karıştırılmış halidir.
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 2 sütun içerir: Soru, Cevap.Dil: Türkçe.
Potansiyel Kullanım Alanları:
Tıbbi araştırmalar
Doğal Dil… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-321179.agrillm-qa-eval-800
Dataset Card for agrillm-qa-eval-800
agrillm-qa-eval-800 is a high-quality evaluation dataset focused on agricultural knowledge and reasoning. The dataset was assembled by ai71 in partnership with leading organizations and partners across the agricultural sector such as CGIAR, ECHO, Digital Green, Embrapa, FAO, the World Bank, IFAD, the Gates Foundation, KALRO, KIADPAI, the Extension Foundation, and additional contributors across the agricultural domain.
It is intended as an open… See the full description on the dataset page: https://huggingface.co/datasets/AI71ai/agrillm-qa-eval-800.physiotherapy-evidence-qa
🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus
Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology.
This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.agronomy-qa-agriculture
Agronomy QA Pairs — Agriculture Instruction Dataset
A concise, English-language question-and-answer dataset covering practical agriculture:
crop management and planting, soil health and fertility, irrigation, and pest and disease
control, with a smaller amount of livestock content. Curated for instruction fine-tuning
of language models in the agriculture domain.
Rows
1,914 question/answer pairs
Unique questions
1,914 — no duplicate questions
Unique answers
1,870… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/agronomy-qa-agriculture.market-news-qa
Market News QA — Market-Analysis Instruction Dataset
Concise question-and-answer pairs for market analysis and financial news
interpretation: classifying news by market area, reading sentiment, identifying
who a story matters to, and answering forward-looking questions from earnings calls.
Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category).
Rows
10,011
Distinct answers
8,469 (85%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.chinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.human_curated_qa_dataset
Human Curated QA Dataset
DigiGreen/human_curated_qa_dataset is a human-verified question-answer dataset designed to support research and development in natural language question answering and agriculture-focused conversational AI.
This dataset contains realistic, domain-relevant QA pairs that were manually curated to ensure accurate and contextually rich answers. It can be used to benchmark models for QA generation.
📌 Dataset Overview
Name: Human Curated QA Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_curated_qa_dataset.patient-doctor-qa-tr-95588
Patient Doctor Q&A TR 95588 Veri Seti
Patient Doctor Q&A TR 95588 veri seti, chat_doctor veri setinin Türkçeye çevrilmiş halidir.
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 3 sütun içerir: Talimat, Soru, Cevap.
Dil: Türkçe.
Potansiyel Kullanım Alanları:
Tıbbi araştırmalar
Doğal Dil İşleme (NLP)
Tıbbi eğitim
Sınırlamalar:
Veri gizliliği endişeleri
Yanıt kalitesinde değişkenlik
Potansiyel… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-95588.QASpanish
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Description
This dataset contains a collection of questions and answers related to topics commonly taught in Spanish high schools. The questions are categorized into different subjects, such as math, science, history, and literature. The answers are provided in Spanish and are intended to be informative and comprehensive.… See the full description on the dataset page: https://huggingface.co/datasets/alivi/QASpanish.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.HC3-textgen-qa
HC3-textgen-qa
the Hello-SimpleAI/HC3 reformatted for textgen
special tokens for question/answer, see dataset preview
math-code-qa
Math & Code QA — Instruction Dataset
Worked mathematical solutions and short code answers, built for the
Adaption Labs AutoScientist Challenge (Math & Code category).
Rows
5,200
Math
3,600
Code
1,600
Distinct answers
5,199 (100%)
Duplicate questions
none
Nulls
none
Question length
median 27 words
Answer length
median 58 words (max 89)
License
CC-BY-4.0
What makes the math rows unusual
Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead.
文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。
推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量
经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式
chat格式为chatml
QASpanis
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Description
This dataset contains a collection of questions and answers related to topics commonly taught in Spanish high schools. The questions are categorized into different subjects, such as math, science, history, and literature. The answers are provided in Spanish and are intended to be informative and comprehensive.… See the full description on the dataset page: https://huggingface.co/datasets/Dontallo/QASpanis.math-code-qa-v2
Math & Code QA v2 — Instruction Dataset
Worked mathematical solutions and short code answers, spanning arithmetic word
problems through to algebra, geometry and combinatorics.
Built for the Adaption Labs AutoScientist Challenge (Math & Code category).
The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the
held-out Math category evaluation.
Rows
5,297 (4,197 math, 1,100 code)
Distinct answers
5,297 (100%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.interview_QA_sample_set
Huy Interview Instruction Dataset
Dataset Description
This is an instruction-answer dataset for fine-tuning conversational AI models to answer interview-style questions based on a personal CV/profile.
The dataset has two columns: instruction and answer.
The dataset contains 5,100 instruction-answer pairs.
Data Creation
This dataset was created using a GenAI-assisted pipeline. A personal CV/profile was provided as source material, and GenAI was used… See the full description on the dataset page: https://huggingface.co/datasets/dinhxuanhuy/interview_QA_sample_set.naija-pidgin-health-qa-rivers-2026PakLegalAid-Lawyer-QA
PakLegalAid – Lawyer-Written Legal QA Dataset
A dataset of 78 real legal questions answered by practicing lawyers, created for the PakLegalAid project.
Why this exists
After fine-tuning LLaMA 3.2 on a bookish legal corpus, the model gave academic, textbook-style answers. To fix this, we had real lawyers write the responses — making the model sound like an actual legal professional, not a law textbook.
Dataset
Column
Description
Query
Legal question… See the full description on the dataset page: https://huggingface.co/datasets/heyIamUmair/PakLegalAid-Lawyer-QA.FL_QA_GERQuestion-Answer style Dataset containing 3069 different questions regarding the Principality of Liechtenstein.
Contains 1409 questions in the legal domain and 1660 questions in the historical / cultural domain.
The questions are generated using OpenAI ChatGPT 4, asking ChatGPT to produce question-answer pairs for the text content given. The text content is based on the documents and articles used in the FL_Legal_GER and FL_History_GER textual datasets.
Dataset is made available in German… See the full description on the dataset page: https://huggingface.co/datasets/JoeUnili/FL_QA_GER.patient-doctor-qa-tr-167732
Patient Doctor Q&A TR 167732 Veri Seti
Patient Doctor Q&A TR 167732 veri seti, doktorsitesi veri setinin temizlenmiş halidir.
Temizlenen Kısımlar
Telefon numaraları
Mail adresleri
Açık adresler
Bağlam bağımlı cümleler
Noktalama işaretleri (gereksiz kullanımlar kaldırıldı, eksik kısımlar dolduruldu)
Büyük/Küçük harfler
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 4 sütun içerir: Ünvan, Alan, Soru, Cevap.… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-167732.IT_QA-QGnews-qa-summarization-73Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.film_qa_pairs_datasetQASpanish
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Description
This dataset contains a collection of questions and answers related to topics commonly taught in Spanish high schools. The questions are categorized into different subjects, such as math, science, history, and literature. The answers are provided in Spanish and are intended to be informative and comprehensive.… See the full description on the dataset page: https://huggingface.co/datasets/Dontallo/QASpanish.bash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
