CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01paodigitalhub /pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub. The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language. The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.textquestion-answeringn<1K1 likes247 downloads4d agoHugging Face02momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads25d agoHugging Face03sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes135 downloads10mo agoHugging Face04Laurie /faithfulness-qa-dataset Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models Overview Faithfulness-QA is a large-scale dataset of 99,094 question-answer pairs designed to train and evaluate the faithfulness of Retrieval-Augmented Generation (RAG) models to retrieved context. The core idea is counterfactual entity substitution: for each QA sample, we replace the answer-bearing entity in the context with a type-consistent alternative… See the full description on the dataset page: https://huggingface.co/datasets/Laurie/faithfulness-qa-dataset.textquestion-answering100K<n<1M0 likes83 downloads2mo agoHugging Face05hamidsalimi /Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1 Persian Civil Procedure QA Dataset Dataset Description این مجموعه‌داده شامل پرسش‌وپاسخ‌های حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است. هر نمونه شامل سه فیلد اصلی است: question: پرسش حقوقی answer: پاسخ پرسش evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است هدف مجموعه‌داده، فراهم‌کردن داده‌ای ساختاریافته برای آموزش، ارزیابی و توسعه مدل‌های زبانی فارسی در زمینه پرسش‌وپاسخ حقوقی است. Dataset Structure نمونه‌ای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.textquestion-answeringn<1K1 likes72 downloads2mo agoHugging Face06uninhibited-scholar /cybersec-qa-dataset-zh Cybersecurity QA Dataset (zh) · 中文网络安全技术问答数据集 面向 防御与安全教育 的中文网络安全技术问答数据集,适用于 LLM 指令微调(SFT)。 21,799 条纯技术问答,零国家归因、零地缘内容,附可复现质检流水线与 CI 校验。 数据概览 总条数:21,799(149 批) 格式:JSONL,每行 {"user": ..., "assistant": ...} 平均答案长度:约 1,231 字,结构化分层(原理 → 攻击面 → 检测 → 缓解) 主题分布(按问题关键词约略归类) 主题 条数 二进制 / 漏洞利用 5186 Web 安全 4858 其他 / 综合 2730 密码学 1648 蓝队 / DFIR / 检测 1622 AD 域 / 内网 / 后渗透 1278 网络协议攻防 1151 云原生 / 容器 1133 移动 / IoT / 固件 837 恶意软件 / 逆向分析 764… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/cybersec-qa-dataset-zh.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face07adrianf12 /healthcare-qa-dataset-jsonl Healthcare Q&A Dataset (JSONL Format) This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain. Dataset Structure The dataset is provided as a JSONL file where each line contains a JSON object with: prompt: A healthcare-related question completion: A detailed, informative answer Sample Entry {"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset-jsonl.textquestion-answeringn<1K0 likes38 downloads1y agoHugging Face08jmvalder /healthcare-qa-dataset-jsonl Healthcare Q&A Dataset (JSONL Format) This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain. Dataset Structure The dataset is provided as a JSONL file where each line contains a JSON object with: prompt: A healthcare-related question completion: A detailed, informative answer Sample Entry {"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/jmvalder/healthcare-qa-dataset-jsonl.textquestion-answeringn<1K0 likes33 downloads8mo agoHugging Face09Gregorgramtoshi /us-law-qa-dataset US Law Q&A Dataset v1.0 673 high-quality question-answer pairs on US law — the perfect dataset for SFT, RAG, and LLM-as-a-Judge. This is a fully cleaned and verified corpus covering all major areas of American law: constitutional, criminal, civil, contract, property, corporate, evidence, labor, intellectual property, antitrust, maritime law, and many more. Judge Score: 9.4/10 (evaluated by Grok, built by xAI). 📸 Data Preview Actual rows from the dataset (646-673)… See the full description on the dataset page: https://huggingface.co/datasets/Gregorgramtoshi/us-law-qa-dataset.texttext-generationn<1K3 likes32 downloads6mo agoHugging Face10bobez999 /arabic-qa-dataset-sigir2024 Arabic QA Dataset | 10,000 Instruction-Tuning Pairs High-quality Arabic question-answering dataset designed for fine-tuning and instruction-tuning LLMs on Arabic language tasks. Dataset Details Size 10,000 entries Format JSONL (instruction/input/output) Language Arabic License MIT Provider AlTal Datamining Schema Field Type Description instruction string The question or task in Arabic input string Additional context… See the full description on the dataset page: https://huggingface.co/datasets/bobez999/arabic-qa-dataset-sigir2024.textquestion-answering10K<n<100K0 likes25 downloads6mo agoHugging Face11FateDefier /MineSafety-QA-Dataset 矿山安全领域 QA 数据集 基于中国矿山安全法规构建的问答对数据集,用于 QLoRA 领域微调。 数据来源 《煤矿安全规程》(2025) 《金属非金属矿山安全规程》(2020) 数据规模 原始生成:7874 条 AI 质量评估过滤后:7265 条 数据格式 Alpaca 格式,包含 <think> 推理链: { "instruction": "问题", "input": "", "output": "<think>\n推理过程...\n</think>\n\n正式回答...", "system": "你是一位精通中国矿山安全法律法规的资深专家..." } 构建流程 PDF 规程文档经 MinerU 转为 Markdown Easy Dataset 自动分块、提取问题、生成答案(DeepSeek-R1-0528-Qwen3-8B) AI 自动评分(满分 5 分),过滤 3.5 分以下的低质量 QA 对… See the full description on the dataset page: https://huggingface.co/datasets/FateDefier/MineSafety-QA-Dataset.textquestion-answering1K<n<10K0 likes22 downloads4mo agoHugging Face12zilalzihar /mikrotik-routeros-qa-dataset MikroTik RouterOS Q&A Dataset The first public structured Q&A dataset for MikroTik RouterOS fine-tuning. 1,672 instruction/response pairs grounded in the official MikroTik Confluence documentation, covering 215 distinct documentation pages. Built to fine-tune zilalzihar/mikrotik-routeros-v01-GGUF, and released so others can train their own RouterOS-specialized models without re-doing the grounding work. Files File Records Purpose qa-pairs.jsonl 1,672 Full… See the full description on the dataset page: https://huggingface.co/datasets/zilalzihar/mikrotik-routeros-qa-dataset.texttext-generation1K<n<10K0 likes18 downloads4mo agoHugging Face13sofikulislam /iu-bd-qa-datasettexttext-generation1K<n<10K0 likes16 downloads1y agoHugging Face14Prathamesh25 /unseen-aptitude-qa-dataset Unseen Aptitude QA Dataset This dataset contains categorized quantitative and logical aptitude questions explicitly structured for campus placement preparation (e.g., TCS, Wipro, Infosys). It is formatted using the standard ChatML / OpenAI Messages schema, making it natively compatible with fine-tuning models like SmolLM2-1.7B. Dataset Structure Each data sample contains a messages array featuring a structured system persona, metadata-enriched user questions, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/unseen-aptitude-qa-dataset.texttext-generationn<1K0 likes14 downloads5mo agoHugging Face15planzb /gensyn-qa-dataset Gensyn Protocol QA Dataset A comprehensive Question-Answer dataset about the Gensyn Protocol, its products, technical architecture, and ecosystem. Dataset Description This dataset contains 65 curated question-answer pairs covering all aspects of the Gensyn decentralized machine learning protocol. It is designed to help developers, researchers, and community members understand Gensyn's technology and participate in the ecosystem. What is Gensyn? Gensyn is a… See the full description on the dataset page: https://huggingface.co/datasets/planzb/gensyn-qa-dataset.textquestion-answeringn<1K0 likes12 downloads10mo agoHugging Face16wavpub /JinJinLeDao_QA_Datasetgated JinJinLeDao QA Dataset Dataset Description Repository: https://github.com/tech-podcasts/JinJinLeDao_QA_Dataset HuggingFace: https://huggingface.co/datasets/wavpub/JinJinLeDao_QA_Dataset Dataset Summary The dataset contains over 18,000 Chinese question-answer pairs extracted from 281 episodes of the Chinese podcast "JinJinLeDao". The subtitles were extracted using the OpenAI Whisper transcription tool, and the question-answer pairs were generated using GPT-3.5… See the full description on the dataset page: https://huggingface.co/datasets/wavpub/JinJinLeDao_QA_Dataset.textquestion-answering10K<n<100K17 likes10 downloads3y agoHugging Face17xunnhi /QA-Dataset-Generator RAG Scientific QA Dataset (Generated) Dataset Description This dataset contains 711 high-quality Question-Answering pairs synthetically generated from ArXiv scientific papers. It is specifically designed to fine-tune Large Language Models (LLMs) for Retrieval-Augmented Generation (RAG) tasks. Source Data: 200 ArXiv papers (Computer Science: AI, CL, LG, IR). Generation Method: Generated using gpt-4o-mini with strict rules to prevent hallucination. Language:… See the full description on the dataset page: https://huggingface.co/datasets/xunnhi/QA-Dataset-Generator.textquestion-answering1K<n<10K1 likes10 downloads3mo agoHugging Face18EnDevSols /reddit-qa-dataset Reddit QA Dataset Dataset Summary The Reddit QA Dataset by EnDevSols is a large-scale, conversational dataset curated from diverse community discussions on Reddit. It is designed to capture the authentic, natural flow of human dialogue, making it an excellent resource for training Large Language Models (LLMs) to handle casual, unstructured queries and complex, multi-turn community interactions. By leveraging highly upvoted answers and community-vetted knowledge, this… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/reddit-qa-dataset.textquestion-answering1M<n<10M0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.