datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.amateur-radio-qa-dataset
📻 Amateur Radio & Electronics QA Dataset (SFT / DPO / Chat)
This dataset is a comprehensive, production-grade bilingual (English and Turkish) corpus dedicated to Amateur Radio (Ham Radio), RF Engineering, Software Defined Radio (SDR), Signal Processing (DSP), Antennas, and Telecommunications Electronics.
Generated and verified using the Elektor Universal Dataset Generator Pipeline (Phase 1-4) with strict LLM-as-a-Judge 5D quality filtering and Google LangExtract… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/amateur-radio-qa-dataset.turkish-qa-multi-dialog-dataset
Turkish QA & Multi-Dialog Dataset
Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir:
Yaklaşık 19.000 adet soru-cevap (QA) örneği
Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri
Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur.
Veri İçeriği
QA Bölümü (~19K)
SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri
Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.ifc-bim-qa-dataset
IFC BIM Question-Answering Dataset
A comprehensive question-answering dataset for Building Information Modeling (BIM) and Industry Foundation Classes (IFC) domain knowledge.
Dataset Summary
This dataset contains 13,485 question-answer pairs covering comprehensive BIM domain knowledge:
IFC Schema Knowledge: Entities, constraints, functions, and global rules
IFC Documentation: Specifications, concepts, geometry, and processes
Professional Certification: BIM practices… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-qa-dataset.Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.customer_qa_dataset
Enterprise Customer Support Q&A Benchmark (500k Complete Dataset)
A comprehensive enterprise-grade customer support dialogue dataset designed for fine-tuning Large Language Models (LLMs), training conversational AI agents, preference optimization (DPO/RLHF), tool-use function calling, RAG, intent classification, and sentiment analysis.
The dataset spans 500,700 structured records covering single-turn & multi-turn customer care interactions across 4 major industries (E-commerce… See the full description on the dataset page: https://huggingface.co/datasets/Saif7800/customer_qa_dataset.turkish_law_qa_dataset
Not: Bu veri seti orijinal olarak OrionCAF tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: OrionCAF/turkish_law_qa_dataset
🔗 Derleyen Platform: VeriPazarı
📚 Türkçe Hukuk Soru-Cevap Veri Seti (Turkish Law QA Dataset)
Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18.300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish_law_qa_dataset.faithfulness-qa-dataset
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
Overview
Faithfulness-QA is a large-scale dataset of 99,094 question-answer pairs designed to train and evaluate the faithfulness of Retrieval-Augmented Generation (RAG) models to retrieved context.
The core idea is counterfactual entity substitution: for each QA sample, we replace the answer-bearing entity in the context with a type-consistent alternative… See the full description on the dataset page: https://huggingface.co/datasets/Laurie/faithfulness-qa-dataset.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.bilingual-coding-qa-dataset
🌐 Bilingual Coding Q&A Dataset
📊 Dataset Description
A comprehensive bilingual (English-Hindi) dataset containing 25,151 high-quality question-answer pairsfocused on programming concepts, particularly Python, machine learning, and AI. This dataset was used to fine-tune coding assistant models and contains over 7 million tokens of training data.
Dataset Statistics
Metric
Value
Total Examples
25,151 Q&A pairs
Total Lines
250,320+… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/bilingual-coding-qa-dataset.actuary-enough-qa-dataset
👋 Connect with me on LinkedIn!
Manuel Caccone - Actuarial Data Scientist & Open Source Educator
Let's discuss actuarial science, AI, and open source projects!
🎯 Actuary Enough - Actuarial Question Simplification Dataset
🚩 Dataset Description
The Actuary Enough Dataset contains examples of complex actuarial and insurance questions that have been simplified and rephrased to improve clarity and accessibility. This dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuary-enough-qa-dataset.turkish_law_qa_dataset
📚 Turkish Law QA Dataset (Türkçe Hukuk Soru-Cevap Veri Seti)
Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18,300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.
Bu veri seti, özellikle hukuk alanında uzmanlaşmış Büyük Dil Modellerini (LLM) ince ayarlamak (fine-tuning), RAG (Retrieval-Augmented Generation) sistemlerinin performansını test etmek ve Türk hukuk sistemine hakim… See the full description on the dataset page: https://huggingface.co/datasets/OrionCAF/turkish_law_qa_dataset.cybersec-qa-dataset-zh
Cybersecurity QA Dataset (zh) · 中文网络安全技术问答数据集
面向 防御与安全教育 的中文网络安全技术问答数据集,适用于 LLM 指令微调(SFT)。
21,799 条纯技术问答,零国家归因、零地缘内容,附可复现质检流水线与 CI 校验。
数据概览
总条数:21,799(149 批)
格式:JSONL,每行 {"user": ..., "assistant": ...}
平均答案长度:约 1,231 字,结构化分层(原理 → 攻击面 → 检测 → 缓解)
主题分布(按问题关键词约略归类)
主题
条数
二进制 / 漏洞利用
5186
Web 安全
4858
其他 / 综合
2730
密码学
1648
蓝队 / DFIR / 检测
1622
AD 域 / 内网 / 后渗透
1278
网络协议攻防
1151
云原生 / 容器
1133
移动 / IoT / 固件
837
恶意软件 / 逆向分析
764… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/cybersec-qa-dataset-zh.banking-finance-qa-dataset
Banking & Finance QA Dataset
This dataset is a custom instruction-response dataset created for banking, finance, AML/KYC, compliance, and regulatory question answering.
It was built as part of a 3-project GenAI portfolio:
Banking RAG Assistant
Banking & Finance QA Dataset
Banking Finance QLoRA Fine-Tuned Model
Dataset Summary
Total samples: 3,002
Format: Alpaca-style instruction / input / output
Language: English
Domain: Banking and Finance
Primary use case:… See the full description on the dataset page: https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset.kaoyan-qa-dataset
CN-Grad-Consult-Dataset (高等教育考研咨询数据集)
📢 特别说明 (Availability Note)
⚠️ 关于数据完整性:
由于 GitHub 文件大小限制,本仓库目前仅直接提供 CPT (连续预训练) 部分的数据文件。
SFT (指令微调) 部分包含高质量问答对,如需获取完整版 SFT 数据集,请查看底部的 📧 联系方式。
📖 Dataset Summary (数据集简介)
本数据集是面向考研与高校招生教育领域的垂直语料库,内容覆盖招生目录、政策解读、院校概况、历年分数线、考试结构解析与考生备考经验等。
旨在帮助开发者构建教育领域的垂直大模型,适用于以下场景:
**教育领域大模型微调 (SFT)**:提升模型在考研咨询场景下的对话能力。
检索增强生成 (RAG) 知识库构建:提供高质量的外部知识源。
**领域适应连续预训练 (CPT)**:注入教育领域的专业背景知识。
数据总量:约 398.36 MB (全量)… See the full description on the dataset page: https://huggingface.co/datasets/L7July/kaoyan-qa-dataset.k-ifrs-qa-dataset
K-IFRS QA Dataset
한국채택국제회계기준(K-IFRS) 기반의 오픈소스 QA 데이터셋입니다.LLM 파인튜닝(SFT), RAG 시스템 구축, 회계 도메인 벤치마크 평가 등 다양한 목적에 활용할 수 있도록 설계되었습니다.
기준 연도: 본 데이터셋은 2026년 5월 24일자 K-IFRS 기준으로 작성되었습니다.회계기준은 지속적으로 개정되므로, 사용 시 기준 연도를 반드시 확인하시기 바랍니다.
데이터셋 개요
항목
내용
총 데이터 수
34,418개
Train 분할
약 30,976개 (90%)
Validation 분할
약 3,442개 (10%)
언어
한국어
형식
Instruction-Input-Output (Alpaca 형식)
기준
K-IFRS (2026년 5월 24일자)
라이선스
CC BY-NC-SA 4.0
데이터 구조 (Data Fields)
각 데이터는… See the full description on the dataset page: https://huggingface.co/datasets/sonsdf/k-ifrs-qa-dataset.human_curated_qa_dataset
Human Curated QA Dataset
DigiGreen/human_curated_qa_dataset is a human-verified question-answer dataset designed to support research and development in natural language question answering and agriculture-focused conversational AI.
This dataset contains realistic, domain-relevant QA pairs that were manually curated to ensure accurate and contextually rich answers. It can be used to benchmark models for QA generation.
📌 Dataset Overview
Name: Human Curated QA Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_curated_qa_dataset.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset-jsonl.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.buddism-qa-dataset
Buddhism Question-Answer Dataset
A comprehensive Vietnamese-English Buddhism question-answering dataset created by merging and processing multiple Buddhism-related datasets.
Dataset Description
This dataset combines two high-quality Buddhism question-answer datasets to create a unified resource for training and evaluating models on Buddhism-related knowledge. The dataset contains questions and answers in both Vietnamese and English, making it suitable for multilingual… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddism-qa-dataset.African-History-QA-Dataset
Dataset Name
African History Dataset
Dataset Structure
Data Fields
question: Questions about African History
answer: Answers to Questions.
Data Splits
Training: 2114 examples
Validation: 200 examples
Testing: 100 examples
Usage
from datasets import load_dataset
dataset = load_dataset("DannyAI/African-History-QA-Dataset")
Citation Information
If you use this dataset, please cite:
@dataset{
Ihenacho2026African_History_Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DannyAI/African-History-QA-Dataset.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/jmvalder/healthcare-qa-dataset-jsonl.us-law-qa-dataset
US Law Q&A Dataset v1.0
673 high-quality question-answer pairs on US law — the perfect dataset for SFT, RAG, and LLM-as-a-Judge.
This is a fully cleaned and verified corpus covering all major areas of American law: constitutional, criminal, civil, contract, property, corporate, evidence, labor, intellectual property, antitrust, maritime law, and many more.
Judge Score: 9.4/10 (evaluated by Grok, built by xAI).
📸 Data Preview
Actual rows from the dataset (646-673)… See the full description on the dataset page: https://huggingface.co/datasets/Gregorgramtoshi/us-law-qa-dataset.turkish_law_qa_dataset
📚 Turkish Law QA Dataset (Türkçe Hukuk Soru-Cevap Veri Seti)
Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18,300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.
Bu veri seti, özellikle hukuk alanında uzmanlaşmış Büyük Dil Modellerini (LLM) ince ayarlamak (fine-tuning), RAG (Retrieval-Augmented Generation) sistemlerinin performansını test etmek ve Türk hukuk sistemine… See the full description on the dataset page: https://huggingface.co/datasets/M3hm3t32/turkish_law_qa_dataset.bioinformatics-qa-dataset
Bioinformatics QA Dataset
A curated question-answer dataset for bioinformatics and computational biology model training.
Summary
Total examples: 5880
Unique topics: 65
Columns: id, topic, question, answer
Source format: CSV converted to Hugging Face Dataset
Intended Use
This dataset is intended for:
Instruction tuning and domain adaptation for biomedical and bioinformatics LLMs
QA benchmarking in life-science terminology
Prompt-response training pipelines… See the full description on the dataset page: https://huggingface.co/datasets/yashm/bioinformatics-qa-dataset.lib3m_qa_dataset_v2
Libertarian Large Language Model QA Dataset (Lib3M QAD) — v2.0.0
Large-scale synthetic Question–Answer dataset distilled from a curated corpus of
libertarian books and magazines. Designed for instruction-tuning / fine-tuning
language models on Austrian economics and classical-liberal philosophy.
What's new in v2 vs v1
+89,321 QA pairs (426,846 total, up from 337,525)
Magazine content added (16.4% of pairs) — previously books only
Third generation model (Qwen 3.6 35B A3B) joins… See the full description on the dataset page: https://huggingface.co/datasets/lib3m/lib3m_qa_dataset_v2.healthcare-qa-dataset
Healthcare Q&A Dataset
This dataset contains 51 healthcare-related question-answer pairs designed for training conversational AI models in the medical domain.
Dataset Structure
Each entry contains:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{
"prompt": "What are the symptoms of diabetes?",
"completion": "Common symptoms of diabetes include increased thirst, frequent urination, unexplained weight loss… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset.arabic-qa-dataset-sigir2024
Arabic QA Dataset | 10,000 Instruction-Tuning Pairs
High-quality Arabic question-answering dataset designed for fine-tuning and instruction-tuning LLMs on Arabic language tasks.
Dataset Details
Size
10,000 entries
Format
JSONL (instruction/input/output)
Language
Arabic
License
MIT
Provider
AlTal Datamining
Schema
Field
Type
Description
instruction
string
The question or task in Arabic
input
string
Additional context… See the full description on the dataset page: https://huggingface.co/datasets/bobez999/arabic-qa-dataset-sigir2024.QA_Dataset_SOP_Akademik
