datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.natural-questions-slim-short-answer
Natural Questions Slim Short Answer
This is a slim, flattened derived version of
google-research-datasets/natural_questions
for short-answer question answering experiments.
The conversion keeps examples with extractable short answers and removes the
original document HTML, token-level document spans, long answer candidates, and
yes/no-only examples. Each record is a simple question-answer pair. It is
intended for lightweight QA prompting and evaluation, not as a full replacement
for… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/natural-questions-slim-short-answer.quora-question-answer-datasetQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora.
Usage:
For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset
minecraft-question-answer-700k
minecraft-question-answer-700k
Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline.
about the dataset
rows - 694,814
tokens - 47,133,624
source - https://minecraft.wiki/
Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.IMO-AnswerBench-Verified
IMO AnswerBench Verified
IMO AnswerBench Verified is a human-expert-verified derivative of OpenEvals/IMO-AnswerBench, originally curated by the Google DeepMind Superhuman Reasoning team. Every record in the 400-problem benchmark was reviewed individually. The review identified and corrected 13 records while preserving the benchmark's balanced coverage of four major mathematical areas.
Dataset summary
Total records: 400
Verification method: record-by-record human… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/IMO-AnswerBench-Verified.chaii-hindi-and-tamil-question-answeringpsychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles.
The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.bodo-legal-question-answering-ai4bharat
Bodo Legal Question Answering Dataset
Overview
This dataset is a Bodo-language legal Question Answering (QA) resource
created for research in low-resource Natural Language Processing (NLP)
and legal language processing.
The supplied source files contain legal judgment contexts together with
multiple questions and answers. For Hugging Face compatibility and
question-answering model training, each question-answer pair has been
flattened into a separate JSONL example… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-ai4bharat.IELTs-Speaking-answer
Overview
This dataset consists of 2 json files named 'ielts_new.json' and 'ielts_old.json', which contain ielts questions and its corresponding answers for part 1 and part 2.
'ielts_new.json': new IELTs topics for 2024 September-December.
'ielts_old.json': remained IELTs topics for 2024 September-December.
Quality
Since the dataset is analysed and generated by ChatGPT based on my own pdf file, the answer may be incomplete(only part of the sentence is extracted, leading to… See the full description on the dataset page: https://huggingface.co/datasets/qwertyuiopasdfg/IELTs-Speaking-answer.aime-2026-fable-5-answers
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
aime-2026-formatted-fable — это обработанный и структурированный датасет на основе задач AIME 2026 из бенчмарка MathArena. Датасет сохранён в формате JSONL и помимо условий задач с финальными ответами содержит сгенерированные цепочки рассуждений (think) с ограничением объёма до 2048 токенов на пример.
Data Fields
Каждая запись в… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/aime-2026-fable-5-answers.bodo-legal-question-answering-iiith
Bodo Legal Question Answering Dataset — IIITH Translation
Overview
A Bodo-language legal Question Answering (QA) resource derived from
English legal judgments. Each example contains a judgment context, a
question, and its corresponding answer.
Data Provenance
Original Legal Source
The underlying English legal judgments were extracted from the publicly
accessible Gauhati High Court judgment repository:… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-iiith.Math-Question-Answerscas_verified_teacher_pool
SCAS Verified Teacher Answer Pool
This dataset provides an aligned, correctness-verified pool of
teacher-generated mathematical reasoning solutions for studying
student-centric data selection in distillation.
The release covers two source corpora, Hendrycks MATH and DeepScaleR. For each
corpus, we retain the subset of questions on which all nine selected teacher
models produce verified correct answers. Each retained question is paired with
nine alternative teacher solutions, one… See the full description on the dataset page: https://huggingface.co/datasets/Student-Centric-Answer-Sampling/scas_verified_teacher_pool.thai-qa-rag-answer-dataset
Thai QA RAG Answer Synthesis Dataset
Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined
Rows: 9999 rows.
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.crypto-sales-question-answersA dataset consisting of questions, answers, and cryptocurrency descriptions
thai-qa-multiturn-answer-dataset
Thai QA Multiturns Answer Synthesis Dataset
Rows: 11,992 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"}
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.einstein_answers
What would Einstein Say?
This dataset contains a set of questions and answers, mimicking Einstein's approach to answer general scientific and philosophical queries.
The data points have been generated synthetically, however the factual correctness of the data is ensured, not guaranteed whatsoever.
ru_sberquad_long_answersUPD 29.05.2023: Добавлены негативные примеры.
Датасет для ответов на вопросы по тексту.
Сгенерирован моделью Den4ikAI/FRED-T5-XL_instructor
Отличия от sberquad, xquad и т.д:
Ответы не односложные, развернутые, представляют несколько предложений
Не подходит для обучения энкодерных моделей!
Question-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.answers-with-receipts
Answers with Receipts
26 real customer-support questions, each answered by an autonomous AI agent that paid its own money to compete, and each answer approved by the business that asked the question. Every row carries the on-chain transaction that paid the agent.
The preference label in this dataset is backed by a payment, not a click.
Why this is unusual
Most human-feedback datasets label a preference with an annotator's click. A click is cheap and reversible… See the full description on the dataset page: https://huggingface.co/datasets/deskcrew/answers-with-receipts.Question-answeringsmall-ru
Dataset Card for Question Answering Russian Dataset
🧠 Quick Summary
Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях.
📚 Dataset Details
Curated by: @kurumikz
Language(s): Russian (ru)
License: CC-BY 4.0 — свободное использование с обязательным указанием автора
Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024,
title={Questions_Answers_In_Sinhala_Language},
author={Ayesha Kalpani},
year={2024},
url={},
}
Questions_Answers_In_Sinhala_Language
Dataset Description
A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala.
Dataset Details
License
This dataset is licensed under the MIT License.
Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.Pidgin_Question-English_Answer_Dataset
Pidgin Question - English Answer Dataset (Sample)
Data Card v1.0
Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/
Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.pcos-patient-assist-question-and-answer
Dataset Card for PCOS Patient Assist Question and Answer Dataset
Dataset Details
Dataset Description
The PCOS Patient Assist Question and Answer Dataset is a curated dataset of question–answer pairs designed to represent common questions asked by patients diagnosed with or concerned about Polycystic Ovary Syndrome (PCOS).
The dataset is structured to simulate real patient queries that arise during different stages of the PCOS journey, including diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-patient-assist-question-and-answer.pcos_question_answer_hindi
PCOS Hindi Lifestyle & Clinical Q&A Dataset
Dataset Details
Dataset Description
This dataset contains patient-facing conversational question–answer pairs in Hindi (Devanagari script) focused on Polycystic Ovary Syndrome (PCOS/PCOD).
The dataset is designed to support training and evaluation of healthcare conversational AI systems that provide lifestyle and general clinical guidance for women diagnosed with PCOS.
All conversations are structured in a chat format… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos_question_answer_hindi.fashion_questions_answersvimqa-generated-answers-pass1
Vi-MQA - Pass 1 Generated Answers & Evaluation
This repo contains the Pass 1 outputs and evaluation results for the Vi-MQA Dataset from the VMLU Benchmark Suite with a total of 4,762 records.
Folder Structure
1. Model Outputs (raw_outputs/)
Contains the formatted outputs from the 3 models evaluated in Pass 1:
results_pass1_gemma.jsonl (Gemma 4 31B IT)
results_pass1_llama.jsonl (Llama 4 Scout)
results_pass1_qwen.jsonl (Qwen3 32B)
2.… See the full description on the dataset page: https://huggingface.co/datasets/nygdon/vimqa-generated-answers-pass1.question_answering
Dataset Information
This Question Answering dataset is a reading comprehension resource derived from Persian Wikipedia. This crowd-sourced dataset contains over 9,000 entries, each of which can either be an unanswerable question or a question with one or more answers based on the provided context. Similar to the SQuAD2.0 dataset, the inclusion of unanswerable questions allows for the development of systems that "know they don't know the answer." Additionally, the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/azizmatin/question_answering.bva-answer-or-abstain
BVA Answer-or-Abstain QA
Question–answer pairs over public U.S. Board of Veterans' Appeals (BVA)
decisions, labeled for whether the document actually supports an answer.
Built to train and evaluate a model's ability to answer when grounded and
abstain when the answer is not in the document — i.e., to say "not in the
document" instead of hallucinating.
Rows: 3,014 (open sample)
Answerable: 2,032 (67%)
Unanswerable (abstain): 982 (33%)
Scale: this is a 3,000-row sample of a 31… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-answer-or-abstain.khmer_question_answerThe data collected from https://www.khsearch.com/ related to the general question-answering examination.
It used to train fine-tuned models from many LLMs, including LlaMa, Qwen, Mistral, and Gemma.
Under the research title "Fine-tuning for Question Answering in Low-Resource Languages: A Case Study on Khmer" conducted at ViLa Lab, Institute of Technology of Cambodia, Phnom Penh.
Lab Info: https://www.facebook.com/vilalabitc
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/kimleang123/khmer_question_answer.
