CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01paodigitalhub /pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub. The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language. The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.textquestion-answeringn<1K1 likes249 downloads4d agoHugging Face02momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes165 downloads25d agoHugging Face03onkanat /amateur-radio-qa-dataset 📻 Amateur Radio & Electronics QA Dataset (SFT / DPO / Chat) This dataset is a comprehensive, production-grade bilingual (English and Turkish) corpus dedicated to Amateur Radio (Ham Radio), RF Engineering, Software Defined Radio (SDR), Signal Processing (DSP), Antennas, and Telecommunications Electronics. Generated and verified using the Elektor Universal Dataset Generator Pipeline (Phase 1-4) with strict LLM-as-a-Judge 5D quality filtering and Google LangExtract… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/amateur-radio-qa-dataset.textquestion-answering100K<n<1M0 likes159 downloads21d agoHugging Face04sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes140 downloads10mo agoHugging Face05Dietmar2020 /ifc-bim-qa-dataset IFC BIM Question-Answering Dataset A comprehensive question-answering dataset for Building Information Modeling (BIM) and Industry Foundation Classes (IFC) domain knowledge. Dataset Summary This dataset contains 13,485 question-answer pairs covering comprehensive BIM domain knowledge: IFC Schema Knowledge: Entities, constraints, functions, and global rules IFC Documentation: Specifications, concepts, geometry, and processes Professional Certification: BIM practices… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-qa-dataset.textquestion-answering10K<n<100K7 likes111 downloads1y agoHugging Face06Laurie /faithfulness-qa-dataset Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models Overview Faithfulness-QA is a large-scale dataset of 99,094 question-answer pairs designed to train and evaluate the faithfulness of Retrieval-Augmented Generation (RAG) models to retrieved context. The core idea is counterfactual entity substitution: for each QA sample, we replace the answer-bearing entity in the context with a type-consistent alternative… See the full description on the dataset page: https://huggingface.co/datasets/Laurie/faithfulness-qa-dataset.textquestion-answering100K<n<1M0 likes95 downloads2mo agoHugging Face07Taklaxbr /turkish_law_qa_dataset Not: Bu veri seti orijinal olarak OrionCAF tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: OrionCAF/turkish_law_qa_dataset 🔗 Derleyen Platform: VeriPazarı 📚 Türkçe Hukuk Soru-Cevap Veri Seti (Turkish Law QA Dataset) Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18.300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish_law_qa_dataset.textquestion-answering10K<n<100K0 likes88 downloads4mo agoHugging Face08manuelcaccone /actuary-enough-qa-dataset 👋 Connect with me on LinkedIn! Manuel Caccone - Actuarial Data Scientist & Open Source Educator Let's discuss actuarial science, AI, and open source projects! 🎯 Actuary Enough - Actuarial Question Simplification Dataset 🚩 Dataset Description The Actuary Enough Dataset contains examples of complex actuarial and insurance questions that have been simplified and rephrased to improve clarity and accessibility. This dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuary-enough-qa-dataset.texttext-generation10K<n<100K0 likes66 downloads1y agoHugging Face09OrionCAF /turkish_law_qa_dataset 📚 Turkish Law QA Dataset (Türkçe Hukuk Soru-Cevap Veri Seti) Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18,300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir. Bu veri seti, özellikle hukuk alanında uzmanlaşmış Büyük Dil Modellerini (LLM) ince ayarlamak (fine-tuning), RAG (Retrieval-Augmented Generation) sistemlerinin performansını test etmek ve Türk hukuk sistemine hakim… See the full description on the dataset page: https://huggingface.co/datasets/OrionCAF/turkish_law_qa_dataset.textquestion-answering10K<n<100K5 likes66 downloads8mo agoHugging Face10hamidsalimi /Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1 Persian Civil Procedure QA Dataset Dataset Description این مجموعه‌داده شامل پرسش‌وپاسخ‌های حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است. هر نمونه شامل سه فیلد اصلی است: question: پرسش حقوقی answer: پاسخ پرسش evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است هدف مجموعه‌داده، فراهم‌کردن داده‌ای ساختاریافته برای آموزش، ارزیابی و توسعه مدل‌های زبانی فارسی در زمینه پرسش‌وپاسخ حقوقی است. Dataset Structure نمونه‌ای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.textquestion-answeringn<1K1 likes62 downloads2mo agoHugging Face11uninhibited-scholar /cybersec-qa-dataset-zh Cybersecurity QA Dataset (zh) · 中文网络安全技术问答数据集 面向 防御与安全教育 的中文网络安全技术问答数据集,适用于 LLM 指令微调(SFT)。 21,799 条纯技术问答,零国家归因、零地缘内容,附可复现质检流水线与 CI 校验。 数据概览 总条数:21,799(149 批) 格式:JSONL,每行 {"user": ..., "assistant": ...} 平均答案长度:约 1,231 字,结构化分层(原理 → 攻击面 → 检测 → 缓解) 主题分布(按问题关键词约略归类) 主题 条数 二进制 / 漏洞利用 5186 Web 安全 4858 其他 / 综合 2730 密码学 1648 蓝队 / DFIR / 检测 1622 AD 域 / 内网 / 后渗透 1278 网络协议攻防 1151 云原生 / 容器 1133 移动 / IoT / 固件 837 恶意软件 / 逆向分析 764… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/cybersec-qa-dataset-zh.texttext-generation10K<n<100K0 likes52 downloads3mo agoHugging Face12RakeshMadasani /banking-finance-qa-dataset Banking & Finance QA Dataset This dataset is a custom instruction-response dataset created for banking, finance, AML/KYC, compliance, and regulatory question answering. It was built as part of a 3-project GenAI portfolio: Banking RAG Assistant Banking & Finance QA Dataset Banking Finance QLoRA Fine-Tuned Model Dataset Summary Total samples: 3,002 Format: Alpaca-style instruction / input / output Language: English Domain: Banking and Finance Primary use case:… See the full description on the dataset page: https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset.texttext-generation1K<n<10K0 likes50 downloads6mo agoHugging Face13sonsdf /k-ifrs-qa-dataset K-IFRS QA Dataset 한국채택국제회계기준(K-IFRS) 기반의 오픈소스 QA 데이터셋입니다.LLM 파인튜닝(SFT), RAG 시스템 구축, 회계 도메인 벤치마크 평가 등 다양한 목적에 활용할 수 있도록 설계되었습니다. 기준 연도: 본 데이터셋은 2026년 5월 24일자 K-IFRS 기준으로 작성되었습니다.회계기준은 지속적으로 개정되므로, 사용 시 기준 연도를 반드시 확인하시기 바랍니다. 데이터셋 개요 항목 내용 총 데이터 수 34,418개 Train 분할 약 30,976개 (90%) Validation 분할 약 3,442개 (10%) 언어 한국어 형식 Instruction-Input-Output (Alpaca 형식) 기준 K-IFRS (2026년 5월 24일자) 라이선스 CC BY-NC-SA 4.0 데이터 구조 (Data Fields) 각 데이터는… See the full description on the dataset page: https://huggingface.co/datasets/sonsdf/k-ifrs-qa-dataset.texttext-generation10K<n<100K0 likes44 downloads4mo agoHugging Face14DigiGreen /human_curated_qa_dataset Human Curated QA Dataset DigiGreen/human_curated_qa_dataset is a human-verified question-answer dataset designed to support research and development in natural language question answering and agriculture-focused conversational AI. This dataset contains realistic, domain-relevant QA pairs that were manually curated to ensure accurate and contextually rich answers. It can be used to benchmark models for QA generation. 📌 Dataset Overview Name: Human Curated QA Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_curated_qa_dataset.texttext-generation1K<n<10K1 likes40 downloads5mo agoHugging Face15M3hm3t32 /turkish_law_qa_dataset 📚 Turkish Law QA Dataset (Türkçe Hukuk Soru-Cevap Veri Seti) Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18,300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir. Bu veri seti, özellikle hukuk alanında uzmanlaşmış Büyük Dil Modellerini (LLM) ince ayarlamak (fine-tuning), RAG (Retrieval-Augmented Generation) sistemlerinin performansını test etmek ve Türk hukuk sistemine… See the full description on the dataset page: https://huggingface.co/datasets/M3hm3t32/turkish_law_qa_dataset.textquestion-answering10K<n<100K0 likes39 downloads3mo agoHugging Face16DannyAI /African-History-QA-Dataset Dataset Name African History Dataset Dataset Structure Data Fields question: Questions about African History answer: Answers to Questions. Data Splits Training: 2114 examples Validation: 200 examples Testing: 100 examples Usage from datasets import load_dataset dataset = load_dataset("DannyAI/African-History-QA-Dataset") Citation Information If you use this dataset, please cite: @dataset{ Ihenacho2026African_History_Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DannyAI/African-History-QA-Dataset.textquestion-answering1K<n<10K2 likes38 downloads8mo agoHugging Face17crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes38 downloads6mo agoHugging Face18vanloc1808 /buddism-qa-dataset Buddhism Question-Answer Dataset A comprehensive Vietnamese-English Buddhism question-answering dataset created by merging and processing multiple Buddhism-related datasets. Dataset Description This dataset combines two high-quality Buddhism question-answer datasets to create a unified resource for training and evaluating models on Buddhism-related knowledge. The dataset contains questions and answers in both Vietnamese and English, making it suitable for multilingual… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddism-qa-dataset.textquestion-answering10K<n<100K0 likes37 downloads1y agoHugging Face19adrianf12 /healthcare-qa-dataset-jsonl Healthcare Q&A Dataset (JSONL Format) This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain. Dataset Structure The dataset is provided as a JSONL file where each line contains a JSON object with: prompt: A healthcare-related question completion: A detailed, informative answer Sample Entry {"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset-jsonl.textquestion-answeringn<1K0 likes36 downloads1y agoHugging Face20jmvalder /healthcare-qa-dataset-jsonl Healthcare Q&A Dataset (JSONL Format) This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain. Dataset Structure The dataset is provided as a JSONL file where each line contains a JSON object with: prompt: A healthcare-related question completion: A detailed, informative answer Sample Entry {"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/jmvalder/healthcare-qa-dataset-jsonl.textquestion-answeringn<1K0 likes36 downloads8mo agoHugging Face21yashm /bioinformatics-qa-dataset Bioinformatics QA Dataset A curated question-answer dataset for bioinformatics and computational biology model training. Summary Total examples: 5880 Unique topics: 65 Columns: id, topic, question, answer Source format: CSV converted to Hugging Face Dataset Intended Use This dataset is intended for: Instruction tuning and domain adaptation for biomedical and bioinformatics LLMs QA benchmarking in life-science terminology Prompt-response training pipelines… See the full description on the dataset page: https://huggingface.co/datasets/yashm/bioinformatics-qa-dataset.textquestion-answering1K<n<10K0 likes34 downloads5mo agoHugging Face22Gregorgramtoshi /us-law-qa-dataset US Law Q&A Dataset v1.0 673 high-quality question-answer pairs on US law — the perfect dataset for SFT, RAG, and LLM-as-a-Judge. This is a fully cleaned and verified corpus covering all major areas of American law: constitutional, criminal, civil, contract, property, corporate, evidence, labor, intellectual property, antitrust, maritime law, and many more. Judge Score: 9.4/10 (evaluated by Grok, built by xAI). 📸 Data Preview Actual rows from the dataset (646-673)… See the full description on the dataset page: https://huggingface.co/datasets/Gregorgramtoshi/us-law-qa-dataset.texttext-generationn<1K3 likes33 downloads6mo agoHugging Face23lib3m /lib3m_qa_dataset_v2 Libertarian Large Language Model QA Dataset (Lib3M QAD) — v2.0.0 Large-scale synthetic Question–Answer dataset distilled from a curated corpus of libertarian books and magazines. Designed for instruction-tuning / fine-tuning language models on Austrian economics and classical-liberal philosophy. What's new in v2 vs v1 +89,321 QA pairs (426,846 total, up from 337,525) Magazine content added (16.4% of pairs) — previously books only Third generation model (Qwen 3.6 35B A3B) joins… See the full description on the dataset page: https://huggingface.co/datasets/lib3m/lib3m_qa_dataset_v2.textquestion-answering100K<n<1M1 likes28 downloads5mo agoHugging Face24bobez999 /arabic-qa-dataset-sigir2024 Arabic QA Dataset | 10,000 Instruction-Tuning Pairs High-quality Arabic question-answering dataset designed for fine-tuning and instruction-tuning LLMs on Arabic language tasks. Dataset Details Size 10,000 entries Format JSONL (instruction/input/output) Language Arabic License MIT Provider AlTal Datamining Schema Field Type Description instruction string The question or task in Arabic input string Additional context… See the full description on the dataset page: https://huggingface.co/datasets/bobez999/arabic-qa-dataset-sigir2024.textquestion-answering10K<n<100K0 likes25 downloads6mo agoHugging Face25adrianf12 /healthcare-qa-dataset Healthcare Q&A Dataset This dataset contains 51 healthcare-related question-answer pairs designed for training conversational AI models in the medical domain. Dataset Structure Each entry contains: prompt: A healthcare-related question completion: A detailed, informative answer Sample Entry { "prompt": "What are the symptoms of diabetes?", "completion": "Common symptoms of diabetes include increased thirst, frequent urination, unexplained weight loss… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset.textquestion-answeringn<1K1 likes24 downloads1y agoHugging Face26ghaniashafiqa /QA_Dataset_SOP_Akademiktexttext-generationn<1K0 likes23 downloads2y agoHugging Face27david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes22 downloads4mo agoHugging Face28FateDefier /MineSafety-QA-Dataset 矿山安全领域 QA 数据集 基于中国矿山安全法规构建的问答对数据集,用于 QLoRA 领域微调。 数据来源 《煤矿安全规程》(2025) 《金属非金属矿山安全规程》(2020) 数据规模 原始生成:7874 条 AI 质量评估过滤后:7265 条 数据格式 Alpaca 格式,包含 <think> 推理链: { "instruction": "问题", "input": "", "output": "<think>\n推理过程...\n</think>\n\n正式回答...", "system": "你是一位精通中国矿山安全法律法规的资深专家..." } 构建流程 PDF 规程文档经 MinerU 转为 Markdown Easy Dataset 自动分块、提取问题、生成答案(DeepSeek-R1-0528-Qwen3-8B) AI 自动评分(满分 5 分),过滤 3.5 分以下的低质量 QA 对… See the full description on the dataset page: https://huggingface.co/datasets/FateDefier/MineSafety-QA-Dataset.textquestion-answering1K<n<10K0 likes21 downloads4mo agoHugging Face29K-University-AIED /level-aware-qa-dataset Construction of LLM-based Level-Aware QA Dataset for AI Tutor Development This dataset is a resource-based QA dataset of approximately 5,000 items generated using GPT-4o based on deep learning major lecture materials and foundational papers. It is a multi-modal dataset containing not only text but also visual information such as formulas and charts. The questions and answers are systematically organized according to the learner's comprehension level (High/Mid/Low). In particular… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/level-aware-qa-dataset.textquestion-answering1K<n<10K0 likes19 downloads8mo agoHugging Face30Prathamesh25 /aptitude-qa-dataset Aptitude QA Dataset This dataset contains foundational quantitative and logical reasoning questions used for fine-tuning compact language models on structured mathematical problem-solving tasks. It uses a structured chat format designed for training placement-focused conversational agents. Dataset Structure The data points are organized into standard conversational messages, establishing a robust training structure for Hugging Face SFTTrainer pipelines. JSON… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/aptitude-qa-dataset.texttext-generation1K<n<10K0 likes19 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.