CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lavita /medical-qa-datasets all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/lavita/medical-qa-datasets.textquestion-answering1M<n<10M64 likes10k downloads3y agoHugging Face02DataPilot /Knowledge-QA-SingleTurn-Dataset Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。 生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約7,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1ターン(質問1 + 回答1) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.text1K<n<10K2 likes6.1k downloads6mo agoHugging Face03turing-motors /STRIDE-QA-Dataset STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks. Category Description Object-centric Spatial QA Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.imagevisual-question-answering100K<n<1M9 likes1.4k downloads8mo agoHugging Face04risenyard /egms-qa-dataset EGMS-QA Dataset Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables, and natural-language QA records for 10,000 overlapping 7 km tiles. This card describes the available data, file formats, and download options. Data access Data needed Files to download Details Published QA records train.jsonl, validation.jsonl, test.jsonl QA loading example Encoder inputs Source tiles, metadata Encoder data Translator inputs Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.textquestion-answering100K<n<1M1 likes1.1k downloads18d agoHugging Face05community-datasets /qa_zre Dataset Card for QaZre Dataset Summary A dataset reducing relation extraction to simple reading comprehension questions Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 516.06 MB Size of the generated dataset: 2.09 GB Total amount of disk used: 2.60 GB An example of 'validation' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/qa_zre.textquestion-answering1M<n<10M5 likes589 downloads2y agoHugging Face06community-datasets /proto_qa Dataset Card for [Dataset Name] Dataset Summary This dataset is for studying computational models trained to reason about prototypical situations. It is anticipated that still would not lead to usage in a downstream task, but as a way of studying the knowledge (and biases) of prototypical situations already contained in pre-trained models. The data it is partially based on (Family Feud). Using deterministic filtering a sampling from a larger set of all transcriptions was… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/proto_qa.textquestion-answering1K<n<10K2 likes545 downloads2y agoHugging Face07khengkok /annual-report-qa-datasettext1K<n<10K0 likes429 downloads10mo agoHugging Face08tolgadev /atpl_qa_dataset Q&A ATPL Dataset for Aviation Industry This dataset contains a collection of questions and answers related to the Airline Transport Pilot License (ATPL) exam. It is designed to assist in the preparation for the ATPL exams and can be used for fine-tuning large language models (LLMs) for the aviation industry. Dataset Details Name: JAA ATPL Question Bank Format: Excel (converted to Hugging Face Dataset) Content: Questions, multiple-choice answers, correct answers, and… See the full description on the dataset page: https://huggingface.co/datasets/tolgadev/atpl_qa_dataset.text1K<n<10K0 likes360 downloads2y agoHugging Face09jun-2018 /multimodal_qa_dataset_v1text10K<n<100K0 likes359 downloads11mo agoHugging Face10romiroll /logical-reasoning-qa-dataset Dataset Card for "logical-reasoning-qa-dataset" More Information needed textn<1K0 likes355 downloads1y agoHugging Face11google-research-datasets /disfl_qa Dataset Card for DISFL-QA: A Benchmark Dataset for Understanding Disfluencies in Question Answering Dataset Summary Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages. Disfl-QA builds upon the SQuAD-v2 (Rajpurkar et al., 2018) dataset, where each question in the dev set is annotated to add a contextual disfluency using the paragraph as a source of distractors. The final dataset… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/disfl_qa.textquestion-answering10K<n<100K7 likes319 downloads2y agoHugging Face12hellotayssir /FinQA_TAT-QA_financial_finetuning_dataset Dataset Summary This dataset provides a unified, flattened context / question / answer format for question answering over financial documents that combine tabular and textual data. It is built to support training and evaluating models on numerical and discrete reasoning tasks in the finance domain, drawing on the structure and style of established finance-QA benchmarks such as TAT-QA and FinQA. Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.text10K<n<100K1 likes299 downloads2mo agoHugging Face13Yana /ft-llm-2026-qa-dataset FT-LLM 2026 QA Dataset A Japanese visual-question-answering dataset used for Stage 1-2 visual instruction tuning of the COMPASS Vision-Language Model. Each sample contains a document or natural image together with one or more Japanese question–answer pairs, and is designed to give the VLM its instruction-following and VQA capabilities. Images are embedded in the dataset, so no external downloads are required. Part of the Compass collection. License Released under the… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-qa-dataset.imagevisual-question-answering100K<n<1M1 likes279 downloads5mo agoHugging Face14168mxie /mnemonic-qa-datasettext1K<n<10K0 likes269 downloads5mo agoHugging Face15strova-ai /hr-policies-qa-dataset 📚 HR Policies Q&A Dataset 🔎 Overview This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for: 🤖 LLM fine-tuning 💬 HR & compliance chatbots 🏢 Enterprise policy automation By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can: ✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/hr-policies-qa-dataset.textn<1K0 likes267 downloads1y agoHugging Face16paodigitalhub /pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub. The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language. The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.textquestion-answeringn<1K1 likes247 downloads4d agoHugging Face17abcasas /VIGIA-QA-datasetimagen<1K0 likes214 downloads2mo agoHugging Face18kikikara /ko_QA_datasetmaywell/korean_textbooks 의 dataset을 Q&A 형식으로 재구성한 dataset입니다. textquestion-answering100K<n<1M3 likes203 downloads2y agoHugging Face19ZuoXiaojia /medical-qa-datasets all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/ZuoXiaojia/medical-qa-datasets.textquestion-answering1M<n<10M0 likes196 downloads8mo agoHugging Face20jun-2018 /multimodal_qa_dataset_v4_traintabular10K<n<100K0 likes191 downloads10mo agoHugging Face21momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads25d agoHugging Face22jun-2018 /multimodal_qa_dataset_v3_traintabularn<1K0 likes162 downloads10mo agoHugging Face23onkanat /amateur-radio-qa-dataset 📻 Amateur Radio & Electronics QA Dataset (SFT / DPO / Chat) This dataset is a comprehensive, production-grade bilingual (English and Turkish) corpus dedicated to Amateur Radio (Ham Radio), RF Engineering, Software Defined Radio (SDR), Signal Processing (DSP), Antennas, and Telecommunications Electronics. Generated and verified using the Elektor Universal Dataset Generator Pipeline (Phase 1-4) with strict LLM-as-a-Judge 5D quality filtering and Google LangExtract… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/amateur-radio-qa-dataset.textquestion-answering100K<n<1M0 likes154 downloads21d agoHugging Face24roshansk23 /NEET_2021_QA_Datasettext1K<n<10K0 likes146 downloads2y agoHugging Face25sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes135 downloads10mo agoHugging Face26DataPilot /Knowledge-QA-MultiTurn-Dataset Knowledge QA Multi-turn Dataset(知識質問データセット・マルチターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形・フォローアップ質問を生成、Kimi K2.5で回答を生成した 3ターンのマルチターン知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約3,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 3ターン(質問3 + 回答3) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-MultiTurn-Dataset.text1K<n<10K2 likes135 downloads6mo agoHugging Face27kgrabko /JiRack-GammaCorpus-Fact-QA-Datasettext1M<n<10M0 likes129 downloads4mo agoHugging Face28jun-2018 /multimodal_qa_dataset_v2_image_focustext1K<n<10K0 likes111 downloads10mo agoHugging Face29Dietmar2020 /ifc-bim-qa-dataset IFC BIM Question-Answering Dataset A comprehensive question-answering dataset for Building Information Modeling (BIM) and Industry Foundation Classes (IFC) domain knowledge. Dataset Summary This dataset contains 13,485 question-answer pairs covering comprehensive BIM domain knowledge: IFC Schema Knowledge: Entities, constraints, functions, and global rules IFC Documentation: Specifications, concepts, geometry, and processes Professional Certification: BIM practices… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-qa-dataset.textquestion-answering10K<n<100K7 likes110 downloads1y agoHugging Face30lab-flair /qa-dataset-k1000 QA Dataset K1000 — The First Drop of Ink Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026). Paper · Full text (v2) · Hugging Face paper page The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.textquestion-answeringn<1K1 likes102 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.