datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Knowledge-QA-SingleTurn-Dataset
Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン)
概要
本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。
生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom)
データの説明
項目
内容
件数
約7,000件
形式
JSONL(1行1JSON)
言語
日本語
ターン数
1ターン(質問1 + 回答1)
ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.egms-qa-dataset
EGMS-QA Dataset
Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables,
and natural-language QA records for 10,000 overlapping 7 km tiles. This card
describes the available data, file formats, and download options.
Data access
Data needed
Files to download
Details
Published QA records
train.jsonl, validation.jsonl, test.jsonl
QA loading example
Encoder inputs
Source tiles, metadata
Encoder data
Translator inputs
Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.hr-policies-qa-dataset
📚 HR Policies Q&A Dataset
🔎 Overview
This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for:
🤖 LLM fine-tuning
💬 HR & compliance chatbots
🏢 Enterprise policy automation
By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can:
✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/hr-policies-qa-dataset.pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.NEET_2021_QA_Datasetturkish-qa-multi-dialog-dataset
Turkish QA & Multi-Dialog Dataset
Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir:
Yaklaşık 19.000 adet soru-cevap (QA) örneği
Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri
Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur.
Veri İçeriği
QA Bölümü (~19K)
SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri
Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.Knowledge-QA-MultiTurn-Dataset
Knowledge QA Multi-turn Dataset(知識質問データセット・マルチターン)
概要
本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形・フォローアップ質問を生成、Kimi K2.5で回答を生成した 3ターンのマルチターン知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom)
データの説明
項目
内容
件数
約3,000件
形式
JSONL(1行1JSON)
言語
日本語
ターン数
3ターン(質問3 + 回答3)
ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-MultiTurn-Dataset.qa-dataset-k1000
QA Dataset K1000 — The First Drop of Ink
Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026).
Paper · Full text (v2) · Hugging Face paper page
The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.faithfulness-qa-dataset
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
Overview
Faithfulness-QA is a large-scale dataset of 99,094 question-answer pairs designed to train and evaluate the faithfulness of Retrieval-Augmented Generation (RAG) models to retrieved context.
The core idea is counterfactual entity substitution: for each QA sample, we replace the answer-bearing entity in the context with a type-consistent alternative… See the full description on the dataset page: https://huggingface.co/datasets/Laurie/faithfulness-qa-dataset.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.QA-dataset-ersynthetic-neurology-QA-datasetqkg_qa_dataset
Data Card: qkg_qa_dataset
Summary
qkg_qa_dataset is a cleaned biomedical multiple-choice QA dataset released with the QKG project. The current release contains 2,788 question-answer records in a compact QA-only schema.
Each record includes:
a stable sample_key
the question text
answer choices, when exported separately
the normalized gold answer label
the original answer key, when available
The main release file is:
top_samples_filtered.jsonl
Earlier versions of this… See the full description on the dataset page: https://huggingface.co/datasets/HKAI-Sci/qkg_qa_dataset.Machine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.cybersec-qa-dataset-zh
Cybersecurity QA Dataset (zh) · 中文网络安全技术问答数据集
面向 防御与安全教育 的中文网络安全技术问答数据集,适用于 LLM 指令微调(SFT)。
21,799 条纯技术问答,零国家归因、零地缘内容,附可复现质检流水线与 CI 校验。
数据概览
总条数:21,799(149 批)
格式:JSONL,每行 {"user": ..., "assistant": ...}
平均答案长度:约 1,231 字,结构化分层(原理 → 攻击面 → 检测 → 缓解)
主题分布(按问题关键词约略归类)
主题
条数
二进制 / 漏洞利用
5186
Web 安全
4858
其他 / 综合
2730
密码学
1648
蓝队 / DFIR / 检测
1622
AD 域 / 内网 / 后渗透
1278
网络协议攻防
1151
云原生 / 容器
1133
移动 / IoT / 固件
837
恶意软件 / 逆向分析
764… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/cybersec-qa-dataset-zh.thai-qa-rag-answer-dataset
Thai QA RAG Answer Synthesis Dataset
Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined
Rows: 9999 rows.
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.medical-qa-datasetqa_legal_dataset_trainRadFAQs_Radiology_Imaging_Health_QA_Dataset
RadFAQs — Radiology & Imaging Health Q&A Dataset (Nepali)
Overview
This dataset (radfaqs_health_qa_nepali_all_nepali.jsonl) is a large, single-source collection of 1,235 instruction-following conversation pairs in Nepali, entirely focused on radiology and medical imaging — CT scans, MRI, X-rays, ultrasound, PET scans, bone density (DEXA) tests, echocardiograms, angiography, radiotherapy/brachytherapy, and related procedures. Each record is a single-turn human↔gpt… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/RadFAQs_Radiology_Imaging_Health_QA_Dataset.Rust_Master_QA_Dataset
Dataset Card for Rust_Master_QA_Dataset
Rust QA Dataset
Dataset Details
Dataset Description
Rust QA Dataset including Questions from:
General Programming
Types
Ownership and Moves
References
Expressions
Error Handling
Crates and Modules
Structs
Enums and Patterns
Traits and Generics
Closures
Iterators
Collections
Strings and Text
Input and Output
Concurrency
Asynchronous Programming
Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset
Nepal CRS Company FAQ — Contraception & Family Planning Nepali Q&A Dataset
1. Overview
This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about contraception and family planning methods — oral pills, emergency contraceptive pills, injectables (e.g. DMPA), implants, IUDs, and condoms. The content originates from Nepal CRS Company, a well-established Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset.Nepali_Commercial_Bank_Financial_Indicators_QA_Dataset
Nepali Commercial Bank Financial Indicators — Grounded QA Dataset
A Nepali-language, grounded single-metric question–answering dataset built from the quarterly Key Financial Indicators of Commercial Banks published for Nepal's commercial banking sector. Each example is a single-turn human↔assistant conversation (ShareGPT / Hermes style) in which a question about one specific bank, one specific quarter, and one specific financial metric is answered with the exact value drawn from… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_Commercial_Bank_Financial_Indicators_QA_Dataset.thai-qa-multiturn-answer-dataset
Thai QA Multiturns Answer Synthesis Dataset
Rows: 11,992 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"}
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.Open-Domain-Oral-Disease-QA-Dataset
Open-Domain-Oral-Disease-QA-Dataset
Dataset Details
Dataset Description
This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in the domain of oral disease.
We currently offer a suite of evaluation datasets encompassing models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B. More data is under reviewed. This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/Lines/Open-Domain-Oral-Disease-QA-Dataset.ais-qa-datasetFight_Vitiligo_FAQ_Nepali_Health_QA_Dataset
Fight Vitiligo FAQ — Nepali Health Q&A Dataset
1. Overview
This dataset is a pure Nepali (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about vitiligo (भिटिलिगो) — a real, non-synthetic health-education FAQ set. Unlike template-generated MCQ datasets, this one contains naturally written, open-ended, explanatory Q&A pairs authored around a single health topic: vitiligo and related skin health.
Each… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Fight_Vitiligo_FAQ_Nepali_Health_QA_Dataset.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset-jsonl.legal-qa-datasetchatgpt_4o_medical_qa_dataset_50000
