datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.natural_questions_cleanjee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.SQuAD-EN-Passage-to-Question
Dataset Card for SQuAD-EN-Passage-to-Question
Dataset Summary
SQuAD-EN-Passage-to-Question is a reformatted and reorganized version of the Stanford Question Answering Dataset (SQuAD). The dataset is designed for text generation and question generation research tasks.
In the original SQuAD dataset, each context passage is associated with multiple question-answer pairs stored as separate entries. In this modified version, all questions associated with the same context… See the full description on the dataset page: https://huggingface.co/datasets/Siam0703/SQuAD-EN-Passage-to-Question.Finance-Questions-Essay_and_Calculation-Chinese
Overview
Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks.
Stage 1: Data Collection & Standardization
Extract financial question samples from professional textbooks via Easy Dataset.
Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.minecraft-question-answer-700k
minecraft-question-answer-700k
Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline.
about the dataset
rows - 694,814
tokens - 47,133,624
source - https://minecraft.wiki/
Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles.
The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-advanced-questions.product-catalog-questions
Lamini Product Catalog QA Dataset
Description
This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle.
Format
The questions and product information are in the form of jsonlines file.
Data Pipeline Code
The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.Math-Question-AnswerSQuAD-BN-Passage-to-Question
Dataset Card for SQuAD-BN-Passage-to-Question
Dataset Summary
SQuAD-BN-Passage-to-Question is a reformatted and filtered version of the Bangla Question Answering dataset derived from csebuetnlp/squad_bn. The dataset is designed for Bangla text generation and question generation research tasks.
In the original dataset, each context passage is associated with multiple question-answer pairs stored as separate entries. In this modified version:
All questions associated with… See the full description on the dataset page: https://huggingface.co/datasets/Siam0703/SQuAD-BN-Passage-to-Question.Question-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.spm-synthetic-questions
Dataset Card: SPM Synthetic Questions
Dataset Summary
14,135 synthetic exam-style questions for Malaysia's SPM (Sijil Pelajaran
Malaysia) curriculum, Form 5, covering 10 subjects. Every item is
LLM-generated and aligned to the KSSM curriculum and the SPM examination
format. Each item belongs to one of three categories:
hots — Higher Order Thinking Skills (KBAT) questions
lazim — soalan lazim (commonly-asked question styles)
perangkap — soalan perangkap (trap… See the full description on the dataset page: https://huggingface.co/datasets/VixeroAI/spm-synthetic-questions.Question-answeringsmall-ru
Dataset Card for Question Answering Russian Dataset
🧠 Quick Summary
Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях.
📚 Dataset Details
Curated by: @kurumikz
Language(s): Russian (ru)
License: CC-BY 4.0 — свободное использование с обязательным указанием автора
Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.ridiculous_math_questions
Dataset Card for Ridiculous Math Questions
A Set of ridiculous math questions that you won't find a teacher to write!
Dataset Details
Dataset Description
This dataset is a list of math questions generated by large language a model.
Which model is used depends on the version:
v0.05 was written by a 20B model, specifically DaringMaid-20B-V1.1-6bpw-exl2.
Curated by: KaraKaraWitch
Funded by [optional]: N/A
Shared by [optional]: KaraKaraWitch
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/WitchesSocialStream/ridiculous_math_questions.Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024,
title={Questions_Answers_In_Sinhala_Language},
author={Ayesha Kalpani},
year={2024},
url={},
}
Questions_Answers_In_Sinhala_Language
Dataset Description
A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala.
Dataset Details
License
This dataset is licensed under the MIT License.
Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.BenchMAX_Question_Answering
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios.
The subtasks are similar to the subtasks in RULER.
The data is sourcing from UN Parallel Corpus and xquad.
The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.Pidgin_Question-English_Answer_Dataset
Pidgin Question - English Answer Dataset (Sample)
Data Card v1.0
Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/
Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.questions
📝 Overview
Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی.
🎯 Purpose
دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی:
Pashto instruction-tuning
Pashto question-answering
Pashto reasoning
Pashto dialogue modeling
Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.cultural-questions-dataset
Turkish General Knowledge & Trivia CoT Dataset (TR-GenK-CoT)
TR-GenK-CoT is a high-quality, synthetic, and carefully curated Turkish dataset designed for Instruction Tuning and Chain of Thought (CoT) reasoning. It contains exactly 500 completely unique, non-repetitive general knowledge and trivia conversations with rich step-by-step thinking processes.
The dataset is formatted using standard Chat Template formats (matching OpenAI/Hugging Face chat schemas) making it directly… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/cultural-questions-dataset.BYOD-open-ended-questions
BYOD fixed open-ended evaluation prompts
This anonymous review artifact contains the 100 fixed, ordered questions used
for the paper's open-ended generation evaluation. Every evaluated model receives
the same prompt text and ordering. The repository intentionally contains prompts
only; model generations and scores are not part of this dataset.
Schema
id: zero-based prompt index used by the evaluation harness.
prompt: the exact user question.
corpus-presse-question-algerienne-dpo-analyse-qualitative-608
Corpus Presse: Question algérienne — qualitative analysis-only DPO pairs (608)
This dataset contains 608 preference pairs for qualitative stance analysis of
Arabic and French historical newspaper articles about the Algerian question and
French colonial order.
Each training row has only:
prompt
chosen
rejected
The model-facing rows deliberately avoid numeric stance scores. The pairs are
intended to teach evidence weighting, source-frame interpretation,
quote/reporting… See the full description on the dataset page: https://huggingface.co/datasets/yakz1/corpus-presse-question-algerienne-dpo-analyse-qualitative-608.questionizer
Questionizer
A dataset of sentences and the questions that they answer.
Propositions were randomly sampled from agentlans/wikipedia-propositions
Then rewritten as questions using google/gemma-3-12b-it
Limitations
Some question-answer pairs sound unnatural
Lacks context when processing single sentences
MatLab-Questions-Answers
Dataset Card for Dataset Name
Matlab-Questions-Answers
Dataset Details
Contains 150+ matlab/octave related questions and answers.
Dataset Description
Contains 150+ matlab/octave related questions and answers ranging from Grade School to Graduate level mathematics.
Language(s) (NLP): English
License: Apache 2.0
Uses
Small Language Models on matlab/octave specific code generation
Evaluation of Language Models on Matlab related questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/EmAllen-TW/MatLab-Questions-Answers.pharmacy_rx_questions
Pharmacy_Rx_Questions (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Pharmacy Prescription Inquiries & Advisory Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/pharmacy_rx_questions.pashto-math-questions
Pashto Math Questions Dataset
This dataset contains a clean collection of mathematical word problems translated into Pashto and localized for regional context.
The original source data consisted of raw conversational Chinese math strings. Because many machine-translated math solutions contain logical and arithmetic inaccuracies, this dataset extracts and preserves only the question fields to provide a high-quality foundation for custom mathematical alignment, instruction tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math-questions.khmer_question_answerThe data collected from https://www.khsearch.com/ related to the general question-answering examination.
It used to train fine-tuned models from many LLMs, including LlaMa, Qwen, Mistral, and Gemma.
Under the research title "Fine-tuning for Question Answering in Low-Resource Languages: A Case Study on Khmer" conducted at ViLa Lab, Institute of Technology of Cambodia, Phnom Penh.
Lab Info: https://www.facebook.com/vilalabitc
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/kimleang123/khmer_question_answer.cognitive-question-generator-v1
Cognitive Question Generator Dataset
Dataset for fine-tuning an expert analysis and question generation model. Contains 5,637 prompt-response pairs capturing expert reasoning patterns for technology transactions and product counseling.
Dataset Description
This dataset was generated from the CognitiveTrainer platform's Mode 1 (Expert Analysis) system, capturing:
Initial scenario analysis
Claim validation with chain-of-trust
Multi-turn expert dialogue
Final synthesis… See the full description on the dataset page: https://huggingface.co/datasets/KevinKeller/cognitive-question-generator-v1.steve-jobs-question-and-answers
