datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Atlas-QnAUpdate 03/30/2026: We released the dataset in harbor format in our official GitHub repo for SWE-Atlas.
We recommend using the harbor scaffold with modal runtime sandboxes as the official way to run the benchmark.
SWE-Atlas QnA
Codebase QnA is the first benchmark in the SWE-Atlas suite. It evaluates AI agents on deep code comprehension — tracing execution paths, explaining architectural decisions, and answering deeply technical questions about production-grade software systems.
124… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SWE-Atlas-QnA.grafite-jee-mains-qna-no-imgarabic-qna
Sadeem QnA: An Arabic QnA Dataset 🌍✨
Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike.
About Sadeem QnA
The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.document-qna-chroma-anyscale-logsScience-QnA
Science-QnA
The Science-QnA is a large-scale, high-quality science-focused dataset (~5.63M rows) curated using synthetic data generation through distillation techniques and select open-source resources. Designed to train and evaluate reasoning-capable models in science domains with emphasis on conceptual understanding, numerical problem-solving, and exam-style Q&A patterns across Physics, Chemistry, Biology, and Mathematics.
Summary
• Domain: Science, Physics… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/Science-QnA.StackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.pbgdpl-vn-legal-qna
pbgdpl.gov.vn — Vietnamese Legal Q&A · Hỏi đáp pháp luật
🇻🇳 Tóm tắt. Bản thu thập đầy đủ chuyên mục Hỏi đáp pháp luật
của Cổng thông tin điện tử Phổ biến giáo dục pháp luật
— cổng giáo dục pháp luật công khai do Bộ Tư pháp vận hành. Mỗi
dòng là một cặp câu hỏi của công dân (Q) và trả lời chính
thức (A), kèm chú thích nguồn, lĩnh vực pháp lý, ngày gửi, và đường
dẫn về trang gốc.
🇬🇧 Summary. A complete crawl of the public Hỏi đáp pháp luật
("Legal Q&A") section of
pbgdpl.gov.vn —… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/pbgdpl-vn-legal-qna.online_privacy_qnaOnline Privacy Policy QnA Dataset
Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand!
Why? Because I was just bored during my summer.
tatoeba-mt-qna-oa
Dataset Card for multilingual tatoeba QnA translation with ~120K entries.
Dataset Summary
Contains Parquet of a list of instructions and translation articles on different languages.
Each row consists of
INSTRUCTION
RESPONSE
SOURCE (tatoeba)
METADATA (json with language, text length, uuid, langs-pair).
Original Dataset is avalible here:
https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt
document-qna-chroma-anyscale-logsBCCard-Finance-Kor-QnAjee-exam-qnaQnA_Descriptivereasoning-gsm-qna-oa
Dataset Card for GSM QnA reasoning with ~8.8K entries.
Dataset Summary
Contains Parquet of a list of instructions and answers.
Each row consists of
INSTRUCTION
RESPONSE
SOURCE
METADATA (json with language).
Original Datasets are available here:
https://huggingface.co/datasets/gsm8k
https://huggingface.co/datasets/reasoning-machines/gsm-hard
magicmotion
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
Quanhao Li*, Zhen Xing*, Rui Wang, Hui Zhang, Qi Dai, and Zuxuan Wu
* equal contribution
💡 Abstract
Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths.
However, existing methods… See the full description on the dataset page: https://huggingface.co/datasets/Qnancy/magicmotion.medical_QnADental_QnA_Instructalodokter-qna
Dataset Question Answer Health Indonesian
Dataset Summary
The Question Answer Health Indonesian dataset contains +250,000 question-and-answer pairs related to health topics sourced from the Alodokter website. The dataset spans a collection period from July 2023 to September 2023 (approximately 2 months). It is designed to facilitate research and development in the fields of natural language processing (NLP), particularly for Indonesian language models, health information… See the full description on the dataset page: https://huggingface.co/datasets/agufsamudra/alodokter-qna.wikipedia-korean-20240501-1million-qnastreamlit-qna-chroma-anyscale-logsqdrant_doc_qnatesla-qna-feedback-logs
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/pgurazada1/tesla-qna-feedback-logs.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
GLAN-qna-kr-300k
Korean GLAN (Generalized Instruction Tuning) Instructions Dataset
GLAN-QnA-KR — a 303,581-row seedless, taxonomy-driven Korean instruction corpus.
📄 A technical report documenting the generation pipeline, duplication analysis, and a two-layer
contamination audit is available on arXiv: arXiv:2607.20443.
Please cite it if you use this dataset (Citation).
What is GLAN?
Catastrophic forgetting, also known as catastrophic interference, occurs during SLM/LLM… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/GLAN-qna-kr-300k.warren-buffett-letters-qna-r1-enhanced-1998-2024
🧠 Warren Buffett Letters Q&A Dataset Pipeline
This project extracts question-answer-reasoning triplets from Warren Buffett's annual shareholder letters using OCR and LLMs. The pipeline is modular and divided into the following stages:
You can clone the repo here.
1. Setup
Create a virtual environment and install dependencies using requirements.txt.
2. Data Curation (curate_data.py)
Load a list of PDF URLs from the Berkshire Hathaway website.
Use Mistral's… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/warren-buffett-letters-qna-r1-enhanced-1998-2024.qna-japaneselauki-qna
Lauki Phones Q&A (chat)
Supervised fine-tuning dataset of Lauki Phones customer-support Q&A pairs,
converted from lauki_qna.jsonl into Hugging Face chat messages.
Format
Each row is a two-turn conversation:
{
"messages": [
{"content": "<question>", "role": "user"},
{"content": "<answer>", "role": "assistant"}
]
}
Load
from datasets import load_dataset
ds = load_dataset("worldboss/lauki-qna", split="train")
print(ds[0]["messages"])
Security-QnAindian_protocols_based_clinical_QnA
Indian Protocols-Based Clinical Q&A
A rubric-graded evaluation dataset built from clinical guideline documents (Indian and international). Each sample is a realistic doctor-side query against a known protocol, paired with rubrics that grade (a) whether the system retrieved/identified the correct guideline content and (b) whether the final answer is clinically complete and safe.
What this evaluates
This dataset is built to stress-test clinical assistants on… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/indian_protocols_based_clinical_QnA.
