datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.turkish-qa-multi-dialog-dataset
Turkish QA & Multi-Dialog Dataset
Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir:
Yaklaşık 19.000 adet soru-cevap (QA) örneği
Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri
Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur.
Veri İçeriği
QA Bölümü (~19K)
SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri
Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.faithfulness-qa-dataset
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
Overview
Faithfulness-QA is a large-scale dataset of 99,094 question-answer pairs designed to train and evaluate the faithfulness of Retrieval-Augmented Generation (RAG) models to retrieved context.
The core idea is counterfactual entity substitution: for each QA sample, we replace the answer-bearing entity in the context with a type-consistent alternative… See the full description on the dataset page: https://huggingface.co/datasets/Laurie/faithfulness-qa-dataset.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.cybersec-qa-dataset-zh
Cybersecurity QA Dataset (zh) · 中文网络安全技术问答数据集
面向 防御与安全教育 的中文网络安全技术问答数据集,适用于 LLM 指令微调(SFT)。
21,799 条纯技术问答,零国家归因、零地缘内容,附可复现质检流水线与 CI 校验。
数据概览
总条数:21,799(149 批)
格式:JSONL,每行 {"user": ..., "assistant": ...}
平均答案长度:约 1,231 字,结构化分层(原理 → 攻击面 → 检测 → 缓解)
主题分布(按问题关键词约略归类)
主题
条数
二进制 / 漏洞利用
5186
Web 安全
4858
其他 / 综合
2730
密码学
1648
蓝队 / DFIR / 检测
1622
AD 域 / 内网 / 后渗透
1278
网络协议攻防
1151
云原生 / 容器
1133
移动 / IoT / 固件
837
恶意软件 / 逆向分析
764… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/cybersec-qa-dataset-zh.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset-jsonl.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/jmvalder/healthcare-qa-dataset-jsonl.us-law-qa-dataset
US Law Q&A Dataset v1.0
673 high-quality question-answer pairs on US law — the perfect dataset for SFT, RAG, and LLM-as-a-Judge.
This is a fully cleaned and verified corpus covering all major areas of American law: constitutional, criminal, civil, contract, property, corporate, evidence, labor, intellectual property, antitrust, maritime law, and many more.
Judge Score: 9.4/10 (evaluated by Grok, built by xAI).
📸 Data Preview
Actual rows from the dataset (646-673)… See the full description on the dataset page: https://huggingface.co/datasets/Gregorgramtoshi/us-law-qa-dataset.arabic-qa-dataset-sigir2024
Arabic QA Dataset | 10,000 Instruction-Tuning Pairs
High-quality Arabic question-answering dataset designed for fine-tuning and instruction-tuning LLMs on Arabic language tasks.
Dataset Details
Size
10,000 entries
Format
JSONL (instruction/input/output)
Language
Arabic
License
MIT
Provider
AlTal Datamining
Schema
Field
Type
Description
instruction
string
The question or task in Arabic
input
string
Additional context… See the full description on the dataset page: https://huggingface.co/datasets/bobez999/arabic-qa-dataset-sigir2024.MineSafety-QA-Dataset
矿山安全领域 QA 数据集
基于中国矿山安全法规构建的问答对数据集,用于 QLoRA 领域微调。
数据来源
《煤矿安全规程》(2025)
《金属非金属矿山安全规程》(2020)
数据规模
原始生成:7874 条
AI 质量评估过滤后:7265 条
数据格式
Alpaca 格式,包含 <think> 推理链:
{
"instruction": "问题",
"input": "",
"output": "<think>\n推理过程...\n</think>\n\n正式回答...",
"system": "你是一位精通中国矿山安全法律法规的资深专家..."
}
构建流程
PDF 规程文档经 MinerU 转为 Markdown
Easy Dataset 自动分块、提取问题、生成答案(DeepSeek-R1-0528-Qwen3-8B)
AI 自动评分(满分 5 分),过滤 3.5 分以下的低质量 QA 对… See the full description on the dataset page: https://huggingface.co/datasets/FateDefier/MineSafety-QA-Dataset.mikrotik-routeros-qa-dataset
MikroTik RouterOS Q&A Dataset
The first public structured Q&A dataset for MikroTik RouterOS fine-tuning. 1,672 instruction/response pairs grounded in the official MikroTik Confluence documentation, covering 215 distinct documentation pages.
Built to fine-tune zilalzihar/mikrotik-routeros-v01-GGUF, and released so others can train their own RouterOS-specialized models without re-doing the grounding work.
Files
File
Records
Purpose
qa-pairs.jsonl
1,672
Full… See the full description on the dataset page: https://huggingface.co/datasets/zilalzihar/mikrotik-routeros-qa-dataset.iu-bd-qa-datasetunseen-aptitude-qa-dataset
Unseen Aptitude QA Dataset
This dataset contains categorized quantitative and logical aptitude questions explicitly structured for campus placement preparation (e.g., TCS, Wipro, Infosys). It is formatted using the standard ChatML / OpenAI Messages schema, making it natively compatible with fine-tuning models like SmolLM2-1.7B.
Dataset Structure
Each data sample contains a messages array featuring a structured system persona, metadata-enriched user questions, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/unseen-aptitude-qa-dataset.gensyn-qa-dataset
Gensyn Protocol QA Dataset
A comprehensive Question-Answer dataset about the Gensyn Protocol, its products, technical architecture, and ecosystem.
Dataset Description
This dataset contains 65 curated question-answer pairs covering all aspects of the Gensyn decentralized machine learning protocol. It is designed to help developers, researchers, and community members understand Gensyn's technology and participate in the ecosystem.
What is Gensyn?
Gensyn is a… See the full description on the dataset page: https://huggingface.co/datasets/planzb/gensyn-qa-dataset.JinJinLeDao_QA_Dataset
JinJinLeDao QA Dataset
Dataset Description
Repository: https://github.com/tech-podcasts/JinJinLeDao_QA_Dataset
HuggingFace: https://huggingface.co/datasets/wavpub/JinJinLeDao_QA_Dataset
Dataset Summary
The dataset contains over 18,000 Chinese question-answer pairs extracted from 281 episodes of the Chinese podcast "JinJinLeDao". The subtitles were extracted using the OpenAI Whisper transcription tool, and the question-answer pairs were generated using GPT-3.5… See the full description on the dataset page: https://huggingface.co/datasets/wavpub/JinJinLeDao_QA_Dataset.QA-Dataset-Generator
RAG Scientific QA Dataset (Generated)
Dataset Description
This dataset contains 711 high-quality Question-Answering pairs synthetically generated from ArXiv scientific papers. It is specifically designed to fine-tune Large Language Models (LLMs) for Retrieval-Augmented Generation (RAG) tasks.
Source Data: 200 ArXiv papers (Computer Science: AI, CL, LG, IR).
Generation Method: Generated using gpt-4o-mini with strict rules to prevent hallucination.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/xunnhi/QA-Dataset-Generator.reddit-qa-dataset
Reddit QA Dataset
Dataset Summary
The Reddit QA Dataset by EnDevSols is a large-scale, conversational dataset curated from diverse community discussions on Reddit. It is designed to capture the authentic, natural flow of human dialogue, making it an excellent resource for training Large Language Models (LLMs) to handle casual, unstructured queries and complex, multi-turn community interactions.
By leveraging highly upvoted answers and community-vetted knowledge, this… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/reddit-qa-dataset.
