CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01truthfulqa /truthful_qa Dataset Card for truthful_qa Dataset Summary TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.textmultiple-choice1K<n<10K292 likes202k downloads3y agoHugging Face02Azzindani /Legal_Corpus_QA_SynDeepThink 🧠 Legal Corpus QA SynDeepThink Dataset This repository contains a high-intelligence Legal Question-and-Answer dataset, generated through an advanced Iterative and Recursive Thinking process. It bridges the gap between static legal corpora and the dynamic "check-and-recheck" nature of human legal expertise. 🏛️ 💡 The Concept: Iterative & Recursive Legal Logic While standard synthetic datasets are often generated in a single pass, Legal_Corpus_QA_SynDeepThink mimics the… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Legal_Corpus_QA_SynDeepThink.tabulartext-generation1K<n<10K1 likes11k downloads7mo agoHugging Face03rezaduty /cybersecurity-qa-v2 Cybersecurity Q&A Dataset v2 — 2.6M Examples A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics. 2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies. Statistics Source Examples Description NIST NVD CVE Database ~1,954,225 All CVEs (2002–2025): overview, severity, detection, remediation AlicanKiraz0/All-CVE-Records-Training-Dataset ~297,441 Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/rezaduty/cybersecurity-qa-v2.textquestion-answering1M<n<10M2 likes4.4k downloads4mo agoHugging Face04Azzindani /ID_Legal_QA_SynDeepThink 🧠 Indonesian Legal QA SynDeepThink Dataset This repository hosts a specialized Indonesian Legal QA dataset that incorporates a Deep Thinking Phase. It is engineered for researchers and developers focusing on high-level judicial reasoning and complex regulatory analysis. 🏛️ 💡 The Concept: Deep Thinking vs. Standard QA While standard models often provide "System 1" (snap) judgments, the SynDeepThink approach simulates "System 2" (slow, deliberate) thinking. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynDeepThink.tabulartext-generationn<1K1 likes4k downloads7mo agoHugging Face05yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.7k downloads6mo agoHugging Face06artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes3k downloads1mo agoHugging Face07Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face08Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes2.7k downloads9m agoHugging Face09KevinNotSmile /nuscenes-qa-mini NuScenes-QA-mini Dataset TL;DR: This dataset is used for multimodal question-answering tasks in autonomous driving scenarios. We created this dataset based on nuScenes-QA dataset for evaluation in our paper Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI. The samples are divided into day and night scenes. scene # train samples # validation samples day 2,229 2,229 night 659 659 Each sample contains… See the full description on the dataset page: https://huggingface.co/datasets/KevinNotSmile/nuscenes-qa-mini.textvisual-question-answering1K<n<10K4 likes1.3k downloads3y agoHugging Face10swe-qa /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.textquestion-answering1K<n<10K5 likes1.1k downloads2mo agoHugging Face11FreedomIntelligence /huatuo_encyclopedia_qa Dataset Card for Huatuo_encyclopedia_qa Dataset Summary This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen Health… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_encyclopedia_qa.texttext-generation100K<n<1M91 likes877 downloads3y agoHugging Face12bofenghuang /medical-qa-fr-v0.1 Medical QA (FR) v0.1 A French medical instruction-tuning dataset (~508K examples) compiled from three public medical QA / dialogue sources: ruslanmv/ai-medical-chatbot — 256,010 examples (config ai_medical_chatbot, default) lavita/medical-qa-datasets (all-processed config) — 230,041 examples (config medical_qa_datasets) FreedomIntelligence/Medical-R1-Distill-Data — 21,641 examples (config medical_r1_distill_data) Each source question was machine-translated into French, then a… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/medical-qa-fr-v0.1.textquestion-answering100K<n<1M0 likes848 downloads3mo agoHugging Face13emre570 /uscode_qacA compact question-answer set for the Prime Intellect U.S. legal evaluation environment. Each record pairs a natural-language question with an extractive answer and the source statute snippet drawn from the Cornell Law School Legal Information Institute (LII) U.S. Code site. Fields also include title_id, section_id, and section_url to support retrieval-style evaluations; the snippet lives in context and is used to build the search index rather than being passed directly to the model.… See the full description on the dataset page: https://huggingface.co/datasets/emre570/uscode_qac.textquestion-answeringn<1K0 likes843 downloads10mo agoHugging Face14zabir1996 /mimic-medical-imaging-qa MIMIC Medical Imaging QA Dataset 5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction. License The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.imagequestion-answering1K<n<10K3 likes796 downloads5mo agoHugging Face15isaacus /open-australian-legal-qa Open Australian Legal QA ‍⚖️ Open Australian Legal QA by Isaacus is the first open dataset of Australian legal questions and answers. Comprised of 2,124 questions and answers synthesised by gpt-4 from the Open Australian Legal Corpus, the largest open database of Australian law, the dataset is intended to facilitate the development of legal AI assistants in Australia. To ensure its accessibility to as wide an audience as possible, the dataset is distributed under the same licence… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-qa.textquestion-answering1K<n<10K23 likes630 downloads7mo agoHugging Face16open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes541 downloads1y agoHugging Face17Meddies /meddies-qa Meddies QA Vietnamese medical question-answer data reorganized into five domain groups, with chat-template QA rows and user-only question-bank configs for generation workflows. [!IMPORTANT] This is a research and training dataset, not clinical guidance. The answers are machine-generated and cleaned from a source medical corpus; review them before using them in products, clinical education, or patient-facing systems. Why this dataset Good… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-qa.textquestion-answering10M<n<100M1 likes520 downloads2mo agoHugging Face18lamini /earnings-calls-qa Lamini Earning Calls QA Dataset Description This dataset contains transcripts of earning calls for various companies, along with questions and answers related to the companies' financial performance and other relevant topics. Format The transcripts, questions, and answers are in the form of jsonlines files, with each json object in the file containing the transcript of an earning call for a single company. Data Pipeline Code The entire data pipeline… See the full description on the dataset page: https://huggingface.co/datasets/lamini/earnings-calls-qa.texttext-classification100K<n<1M54 likes488 downloads3y agoHugging Face19qinchuanhui /UDA-QA Dataset Card for Dataset Name [NIPS-2024] UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis (https://arxiv.org/abs/2406.15187) UDA (Unstructured Document Analysis) is a benchmark suite for Retrieval Augmented Generation (RAG) in real-world document analysis. Each entry in the UDA dataset is organized as a document-question-answer triplet, where a question is raised from the document, accompanied by a corresponding ground-truth answer. The… See the full description on the dataset page: https://huggingface.co/datasets/qinchuanhui/UDA-QA.textquestion-answering10K<n<100K6 likes488 downloads2y agoHugging Face20humanfia-lab /QAlg QAlg Humanize-Physic Formalizations and Proofs QAlg (Quantum Algorithms) is a blind benchmark for formalizing theorems in quantum algorithms. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 36 tasks cover quantum circuits, linear algebra, the quantum Fourier transform, Hamiltonian simulation, hidden subgroups, QSP/QSVT, and parameterized… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QAlg.texttext-generationn<1K0 likes477 downloads2mo agoHugging Face21nshportun /usa-immigration-law-qa USA Immigration Law Q&A Dataset A large-scale, source-grounded Q&A dataset covering U.S. immigration law and policy, built entirely from official government sources, open legal datasets, and curated community materials. Complete pipeline available at https://github.com/nshportun/usa-immigration and pre-print https://arxiv.org/abs/2605.30589 Dataset Contents Split Records Description train 16,065 Training Q&A pairs eval 993 Stratified held-out… See the full description on the dataset page: https://huggingface.co/datasets/nshportun/usa-immigration-law-qa.textquestion-answering10K<n<100K3 likes455 downloads4mo agoHugging Face22FreedomIntelligence /huatuo_consultation_qa Dataset Card for huatuo_consultation_qa Dataset Summary We collected data from a website for medical consultation , consisting of many online consultation records by medical experts. Each record is a QA pair: a patient raises a question and a medical doctor answers the question. The basic information of doctors (including name, hospital organization, and department) was recorded. We directly crawl patient’s questions and doctor’s answers as QA pairs, getting 32,708,346… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_consultation_qa.texttext-generation10M<n<100M16 likes449 downloads3y agoHugging Face23Voxel51 /fiftyone-qa-pairs-14k FiftyOne QA 14k Dataset Overview This dataset is derived from the FiftyOne Function Calling 14k dataset and is designed to train AI assistants to understand and answer questions about FiftyOne's functionality. Purpose Train AI models to understand FiftyOne-related queries Provide examples of FiftyOne syntax Map natural language questions to appropriate FiftyOne code snippets Demonstrate correct usage for FiftyOne functions and method calls This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-qa-pairs-14k.textquestion-answering10K<n<100K1 likes431 downloads1y agoHugging Face24FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes424 downloads3y agoHugging Face25alianassmaaa /ameli-assurance-maladie-qa Ameli Assurance Maladie - Question Answering Dataset Description Dataset de 100 paires question-réponse (QA) basé sur les publications officielles de l'Assurance Maladie française, extraites du site assurance-maladie.ameli.fr. Conçu pour évaluer des systèmes de RAG (Retrieval-Augmented Generation) sur des documents institutionnels français dans le domaine de la santé publique et de la protection sociale. Format du dataset { "question": "Quel article du… See the full description on the dataset page: https://huggingface.co/datasets/alianassmaaa/ameli-assurance-maladie-qa.documentquestion-answeringn<1K1 likes362 downloads5mo agoHugging Face26zai-org /webglm-qa WebGLM-QA Dataset Description WebGLM-QA is the dataset used to train the WebGLM generator module. It consists of 43,579 high-quality data samples for the train split, 1,000 for the validation split, and 400 for the test split. Refer to our paper for the data construction details. Dataset Structure To load the dataset, you can try the following code. from datasets import load_dataset load_dataset("THUDM/webglm-qa") DatasetDict({ train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/webglm-qa.texttext-generation10K<n<100K65 likes358 downloads3y agoHugging Face27MercanAI /turkce-sft-qa-3.7m 🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti 3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri setinden geldiğini taşır. English: A merged, row-level deduplicated and quality-filtered collection of 24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries its source dataset, source URL and original license. 🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.tabulartext-generation1M<n<10M0 likes341 downloads2mo agoHugging Face28fzmnm /TinyBooks-QA-Chinese本数据集已停止更新,请移步https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh TinyBooks-QA-Chinese Inspired by the (TinyStories)[https://arxiv.org/abs/2305.07759] paper, where a small language model exhibits strong capabilities when trained on high-quality, 🍼baby-friendly stories synthesized by AI, I present an AI-generated Encyclopedia suitable for kindergarten and grade school levels. This AI-synthesized dataset converts classical literature into a question-answer style curriculum with… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyBooks-QA-Chinese.texttext-generation1K<n<10K8 likes315 downloads2y agoHugging Face29HTJ008 /JavaError-QA JErrRAG-Eval-800 JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record. This Hugging Face repository contains: java_error_qa_v2/: the canonical public benchmark package paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles SHA256SUMS.txt: release-side hash anchors referenced by the paper Dataset Summary Total records: 800 Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.documentquestion-answeringn<1K2 likes303 downloads2mo agoHugging Face30lateesha-bhatia /sec-filings-qa-instruct SEC Filings Instruction-Tuning Dataset (Llama-3 Format) This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template. Dataset Details Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.textquestion-answering1K<n<10K0 likes273 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.