datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-qa-datasets
all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets
The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset
Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples
truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/lavita/medical-qa-datasets.egms-qa-dataset
EGMS-QA Dataset
Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables,
and natural-language QA records for 10,000 overlapping 7 km tiles. This card
describes the available data, file formats, and download options.
Data access
Data needed
Files to download
Details
Published QA records
train.jsonl, validation.jsonl, test.jsonl
QA loading example
Encoder inputs
Source tiles, metadata
Encoder data
Translator inputs
Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.qa_zre
Dataset Card for QaZre
Dataset Summary
A dataset reducing relation extraction to simple reading comprehension questions
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 516.06 MB
Size of the generated dataset: 2.09 GB
Total amount of disk used: 2.60 GB
An example of 'validation' looks as follows.
{… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/qa_zre.proto_qa
Dataset Card for [Dataset Name]
Dataset Summary
This dataset is for studying computational models trained to reason about prototypical situations. It is anticipated that still would not lead to usage in a downstream task, but as a way of studying the knowledge (and biases) of prototypical situations already contained in pre-trained models. The data it is partially based on (Family Feud).
Using deterministic filtering a sampling from a larger set of all transcriptions was… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/proto_qa.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.financial-qa-dataset
financial-qa-dataset
This dataset consists of Question-Answer_Context Pairs. It also consists of metadata for filtering the records.
Repo Structure
financial-qa-dataset
├── financial-qa-dataset.csv
├── metadata.csv
├── notebooks
│ |── loading_dataset.ipynb
│ |── Loading_dataset_huggingface.ipynb
│ |── basic_rag_langchain_vertexai.ipynb
│ |── basic_rag_with_evaluation.ipynb
|
├── data
|── Statements
|── Reports… See the full description on the dataset page: https://huggingface.co/datasets/adityarane/financial-qa-dataset.disfl_qa
Dataset Card for DISFL-QA: A Benchmark Dataset for Understanding Disfluencies in Question Answering
Dataset Summary
Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages. Disfl-QA builds upon the SQuAD-v2 (Rajpurkar et al., 2018) dataset, where each question in the dev set is annotated to add a contextual disfluency using the paragraph as a source of distractors.
The final dataset… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/disfl_qa.pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.ko_QA_datasetmaywell/korean_textbooks 의 dataset을 Q&A 형식으로 재구성한 dataset입니다.
medical-qa-datasets
all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets
The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset
Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples
truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/ZuoXiaojia/medical-qa-datasets.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.multi_re_qaMultiReQA contains the sentence boundary annotation from eight publicly available QA datasets including SearchQA, TriviaQA, HotpotQA, NaturalQuestions, SQuAD, BioASQ, RelationExtraction, and TextbookQA. Five of these datasets, including SearchQA, TriviaQA, HotpotQA, NaturalQuestions, SQuAD, contain both training and test data, and three, including BioASQ, RelationExtraction, TextbookQA, contain only the test dataamateur-radio-qa-dataset
📻 Amateur Radio & Electronics QA Dataset (SFT / DPO / Chat)
This dataset is a comprehensive, production-grade bilingual (English and Turkish) corpus dedicated to Amateur Radio (Ham Radio), RF Engineering, Software Defined Radio (SDR), Signal Processing (DSP), Antennas, and Telecommunications Electronics.
Generated and verified using the Elektor Universal Dataset Generator Pipeline (Phase 1-4) with strict LLM-as-a-Judge 5D quality filtering and Google LangExtract… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/amateur-radio-qa-dataset.korean-legal-qa-dataset
Korean Legal QA Dataset (한국어 법률 QA 데이터셋)
법령 조문과 판례를 연결고리로 묶은 한국어 법률 질의응답 데이터셋입니다.
베제처 국가법령정보센터의 공개 법령·판례 데이터를 바탕으로 구축되었습니다.
형식 (Format)
각 항목은 질문(question), 정답(answer), 검색문서(retrieved_documents) 3가지 필드로 구성됩니다:
{
"id": "qa_conn_01_a",
"question": "[민사소송법 제80조] 관련하여, 대법원 1948.04.07 선고 '부동산소유권이전등기' 사건에서 법원은 어떤 판단을 내렸는가?",
"answer": "【사건명】부동산소유권이전등기 (4281민상362)\n【판시사항】...\n【판결요지】...",
"retrieved_documents": [
{
"type": "statute",
"law_name": "민사소송법"… See the full description on the dataset page: https://huggingface.co/datasets/ggh5454/korean-legal-qa-dataset.turkish-qa-multi-dialog-dataset
Turkish QA & Multi-Dialog Dataset
Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir:
Yaklaşık 19.000 adet soru-cevap (QA) örneği
Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri
Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur.
Veri İçeriği
QA Bölümü (~19K)
SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri
Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.devops-qa-dataset
DevOps Q&A Dataset v1.0
Overview
High-quality dataset of 25,670 DevOps technical examples collected from GitHub repositories, Stack Exchange, and official documentation.
Statistics
Total examples: 25,670
Average quality score: ~0.82
Unique (deduplicated): High accuracy via MD5
Categories: Docker, Kubernetes, CI/CD, Cloud, Linux, Terraform, Ansible
Sources: StackExchange/HuggingFace (70%), GitHub Repositories (29%), Official Documentation (~1%)
Use… See the full description on the dataset page: https://huggingface.co/datasets/Skilln/devops-qa-dataset.ifc-bim-qa-dataset
IFC BIM Question-Answering Dataset
A comprehensive question-answering dataset for Building Information Modeling (BIM) and Industry Foundation Classes (IFC) domain knowledge.
Dataset Summary
This dataset contains 13,485 question-answer pairs covering comprehensive BIM domain knowledge:
IFC Schema Knowledge: Entities, constraints, functions, and global rules
IFC Documentation: Specifications, concepts, geometry, and processes
Professional Certification: BIM practices… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-qa-dataset.Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.qa-dataset-k1000
QA Dataset K1000 — The First Drop of Ink
Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026).
Paper · Full text (v2) · Hugging Face paper page
The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.Mistral_Trivia-QA_Dataset
Mistral Trivia QA Dataset
The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation.
Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.MoroccanHistory-QA-Datasetcareer-guidance-qa-dataset
Dataset Card for Career Guidance Dataset
Dataset Overview
This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.customer_qa_dataset
Enterprise Customer Support Q&A Benchmark (500k Complete Dataset)
A comprehensive enterprise-grade customer support dialogue dataset designed for fine-tuning Large Language Models (LLMs), training conversational AI agents, preference optimization (DPO/RLHF), tool-use function calling, RAG, intent classification, and sentiment analysis.
The dataset spans 500,700 structured records covering single-turn & multi-turn customer care interactions across 4 major industries (E-commerce… See the full description on the dataset page: https://huggingface.co/datasets/Saif7800/customer_qa_dataset.turkish_law_qa_dataset
Not: Bu veri seti orijinal olarak OrionCAF tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: OrionCAF/turkish_law_qa_dataset
🔗 Derleyen Platform: VeriPazarı
📚 Türkçe Hukuk Soru-Cevap Veri Seti (Turkish Law QA Dataset)
Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18.300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish_law_qa_dataset.faithfulness-qa-dataset
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
Overview
Faithfulness-QA is a large-scale dataset of 99,094 question-answer pairs designed to train and evaluate the faithfulness of Retrieval-Augmented Generation (RAG) models to retrieved context.
The core idea is counterfactual entity substitution: for each QA sample, we replace the answer-bearing entity in the context with a type-consistent alternative… See the full description on the dataset page: https://huggingface.co/datasets/Laurie/faithfulness-qa-dataset.comprehensive-qa-dataset
Comprehensive Question Answering Dataset
A large-scale, diverse collection of question answering datasets combined into a unified format for training and evaluating QA models. This dataset contains over 160,000 question-answer pairs from three popular QA benchmarks.
Dataset Summary
This comprehensive dataset combines three popular question answering datasets into a single, unified format:
SQuAD 2.0 (Stanford Question Answering Dataset) - Context passages from Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/Successmove/comprehensive-qa-dataset.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.bilingual-coding-qa-dataset
🌐 Bilingual Coding Q&A Dataset
📊 Dataset Description
A comprehensive bilingual (English-Hindi) dataset containing 25,151 high-quality question-answer pairsfocused on programming concepts, particularly Python, machine learning, and AI. This dataset was used to fine-tune coding assistant models and contains over 7 million tokens of training data.
Dataset Statistics
Metric
Value
Total Examples
25,151 Q&A pairs
Total Lines
250,320+… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/bilingual-coding-qa-dataset.QA_LAW_Egyptian_dataset
Egyptian Arabic Legal QA Dataset
Dataset Description
This dataset contains 3,725 question-answer pairs in Egyptian Arabic focused on legal topics. The dataset serves as a valuable resource for developing Arabic natural language processing models, particularly for legal domain applications in Egyptian Arabic dialect.
Key Features
Language: Egyptian Arabic (العامية المصرية)
Domain: Legal/Law (القانون)
Size: 3,725 examples
Topics: 745 unique legal categories… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/QA_LAW_Egyptian_dataset.turkish_law_qa_dataset
📚 Turkish Law QA Dataset (Türkçe Hukuk Soru-Cevap Veri Seti)
Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18,300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.
Bu veri seti, özellikle hukuk alanında uzmanlaşmış Büyük Dil Modellerini (LLM) ince ayarlamak (fine-tuning), RAG (Retrieval-Augmented Generation) sistemlerinin performansını test etmek ve Türk hukuk sistemine hakim… See the full description on the dataset page: https://huggingface.co/datasets/OrionCAF/turkish_law_qa_dataset.
