CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ScaleAI /SWE-Atlas-QnAUpdate 03/30/2026: We released the dataset in harbor format in our official GitHub repo for SWE-Atlas. We recommend using the harbor scaffold with modal runtime sandboxes as the official way to run the benchmark. SWE-Atlas QnA Codebase QnA is the first benchmark in the SWE-Atlas suite. It evaluates AI agents on deep code comprehension — tracing execution paths, explaining architectural decisions, and answering deeply technical questions about production-grade software systems. 124… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SWE-Atlas-QnA.textn<1K21 likes701 downloads6mo agoHugging Face02ruh-ai /grafite-jee-mains-qna-no-imgtext10K<n<100K3 likes689 downloads1y agoHugging Face03sadeem-ai /arabic-qna Sadeem QnA: An Arabic QnA Dataset 🌍✨ Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike. About Sadeem QnA The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.textquestion-answering1K<n<10K4 likes392 downloads3y agoHugging Face04pgurazada1 /document-qna-chroma-anyscale-logstextn<1K0 likes363 downloads2y agoHugging Face05169Pi /Science-QnA Science-QnA The Science-QnA is a large-scale, high-quality science-focused dataset (~5.63M rows) curated using synthetic data generation through distillation techniques and select open-source resources. Designed to train and evaluate reasoning-capable models in science domains with emphasis on conceptual understanding, numerical problem-solving, and exam-style Q&A patterns across Physics, Chemistry, Biology, and Mathematics. Summary • Domain: Science, Physics… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/Science-QnA.texttext-generation1M<n<10M3 likes314 downloads7mo agoHugging Face06Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes234 downloads5mo agoHugging Face07tmquan /pbgdpl-vn-legal-qna pbgdpl.gov.vn — Vietnamese Legal Q&A · Hỏi đáp pháp luật 🇻🇳 Tóm tắt. Bản thu thập đầy đủ chuyên mục Hỏi đáp pháp luật của Cổng thông tin điện tử Phổ biến giáo dục pháp luật — cổng giáo dục pháp luật công khai do Bộ Tư pháp vận hành. Mỗi dòng là một cặp câu hỏi của công dân (Q) và trả lời chính thức (A), kèm chú thích nguồn, lĩnh vực pháp lý, ngày gửi, và đường dẫn về trang gốc. 🇬🇧 Summary. A complete crawl of the public Hỏi đáp pháp luật ("Legal Q&A") section of pbgdpl.gov.vn —… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/pbgdpl-vn-legal-qna.tabularquestion-answering1K<n<10K0 likes201 downloads4mo agoHugging Face08arjunth2001 /online_privacy_qnaOnline Privacy Policy QnA Dataset textn<1K5 likes129 downloads5y agoHugging Face09neifuisan /Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand! Why? Because I was just bored during my summer. textquestion-answeringn<1K50 likes119 downloads2y agoHugging Face100x22almostEvil /tatoeba-mt-qna-oa Dataset Card for multilingual tatoeba QnA translation with ~120K entries. Dataset Summary Contains Parquet of a list of instructions and translation articles on different languages. Each row consists of INSTRUCTION RESPONSE SOURCE (tatoeba) METADATA (json with language, text length, uuid, langs-pair). Original Dataset is avalible here: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt textquestion-answering100K<n<1M2 likes117 downloads3y agoHugging Face11mayankchugh-learning /document-qna-chroma-anyscale-logstextn<1K0 likes116 downloads2y agoHugging Face12BCCard /BCCard-Finance-Kor-QnAtext10K<n<100K17 likes113 downloads2y agoHugging Face13datavorous /jee-exam-qnatabular10K<n<100K2 likes108 downloads1y agoHugging Face14msamg /QnA_Descriptivetextn<1K0 likes101 downloads2y agoHugging Face150x22almostEvil /reasoning-gsm-qna-oa Dataset Card for GSM QnA reasoning with ~8.8K entries. Dataset Summary Contains Parquet of a list of instructions and answers. Each row consists of INSTRUCTION RESPONSE SOURCE METADATA (json with language). Original Datasets are available here: https://huggingface.co/datasets/gsm8k https://huggingface.co/datasets/reasoning-machines/gsm-hard textquestion-answering1K<n<10K8 likes91 downloads3y agoHugging Face16Qnancy /magicmotion MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance Quanhao Li*, Zhen Xing*, Rui Wang, Hui Zhang, Qi Dai, and Zuxuan Wu * equal contribution 💡 Abstract Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths. However, existing methods… See the full description on the dataset page: https://huggingface.co/datasets/Qnancy/magicmotion.imagen<1K0 likes91 downloads8mo agoHugging Face17tkxwaweru /medical_QnAtext10K<n<100K1 likes88 downloads2y agoHugging Face18LexiconShiftInnovations /Dental_QnA_Instructtext1K<n<10K0 likes75 downloads2y agoHugging Face19agufsamudra /alodokter-qna Dataset Question Answer Health Indonesian Dataset Summary The Question Answer Health Indonesian dataset contains +250,000 question-and-answer pairs related to health topics sourced from the Alodokter website. The dataset spans a collection period from July 2023 to September 2023 (approximately 2 months). It is designed to facilitate research and development in the fields of natural language processing (NLP), particularly for Indonesian language models, health information… See the full description on the dataset page: https://huggingface.co/datasets/agufsamudra/alodokter-qna.text100K<n<1M4 likes75 downloads2y agoHugging Face20lcw99 /wikipedia-korean-20240501-1million-qnatext100K<n<1M40 likes67 downloads2y agoHugging Face21mayankchugh-learning /streamlit-qna-chroma-anyscale-logstextn<1K0 likes64 downloads2y agoHugging Face22atitaarora /qdrant_doc_qnatextn<1K1 likes63 downloads2y agoHugging Face23pgurazada1 /tesla-qna-feedback-logs Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/pgurazada1/tesla-qna-feedback-logs.textn<1K0 likes62 downloads2y agoHugging Face24genloop /FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1 Dataset Card This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs. textquestion-answeringn<1K0 likes61 downloads2y agoHugging Face25daekeun-ml /GLAN-qna-kr-300k Korean GLAN (Generalized Instruction Tuning) Instructions Dataset GLAN-QnA-KR — a 303,581-row seedless, taxonomy-driven Korean instruction corpus. 📄 A technical report documenting the generation pipeline, duplication analysis, and a two-layer contamination audit is available on arXiv: arXiv:2607.20443. Please cite it if you use this dataset (Citation). What is GLAN? Catastrophic forgetting, also known as catastrophic interference, occurs during SLM/LLM… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/GLAN-qna-kr-300k.textquestion-answering100K<n<1M5 likes60 downloads2mo agoHugging Face26eagle0504 /warren-buffett-letters-qna-r1-enhanced-1998-2024 🧠 Warren Buffett Letters Q&A Dataset Pipeline This project extracts question-answer-reasoning triplets from Warren Buffett's annual shareholder letters using OCR and LLMs. The pipeline is modular and divided into the following stages: You can clone the repo here. 1. Setup Create a virtual environment and install dependencies using requirements.txt. 2. Data Curation (curate_data.py) Load a list of PDF URLs from the Berkshire Hathaway website. Use Mistral's… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/warren-buffett-letters-qna-r1-enhanced-1998-2024.textquestion-answering10K<n<100K2 likes58 downloads1y agoHugging Face27ghiffaryr /qna-japanesetext1M<n<10M0 likes55 downloads2y agoHugging Face28worldboss /lauki-qna Lauki Phones Q&A (chat) Supervised fine-tuning dataset of Lauki Phones customer-support Q&A pairs, converted from lauki_qna.jsonl into Hugging Face chat messages. Format Each row is a two-turn conversation: { "messages": [ {"content": "<question>", "role": "user"}, {"content": "<answer>", "role": "assistant"} ] } Load from datasets import load_dataset ds = load_dataset("worldboss/lauki-qna", split="train") print(ds[0]["messages"]) texttext-generationn<1K0 likes55 downloads24d agoHugging Face29Mr-Vicky-01 /Security-QnAtext1K<n<10K6 likes54 downloads2y agoHugging Face30ekacare /indian_protocols_based_clinical_QnA Indian Protocols-Based Clinical Q&A A rubric-graded evaluation dataset built from clinical guideline documents (Indian and international). Each sample is a realistic doctor-side query against a known protocol, paired with rubrics that grade (a) whether the system retrieved/identified the correct guideline content and (b) whether the final answer is clinically complete and safe. What this evaluates This dataset is built to stress-test clinical assistants on… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/indian_protocols_based_clinical_QnA.textquestion-answeringn<1K0 likes54 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.