datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.natural_questions_cleanjee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.minecraft-question-answer-700k
minecraft-question-answer-700k
Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline.
about the dataset
rows - 694,814
tokens - 47,133,624
source - https://minecraft.wiki/
Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.Finance-Questions-Essay_and_Calculation-Chinese
Overview
Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks.
Stage 1: Data Collection & Standardization
Extract financial question samples from professional textbooks via Easy Dataset.
Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.SQuAD-EN-Passage-to-Question
Dataset Card for SQuAD-EN-Passage-to-Question
Dataset Summary
SQuAD-EN-Passage-to-Question is a reformatted and reorganized version of the Stanford Question Answering Dataset (SQuAD). The dataset is designed for text generation and question generation research tasks.
In the original SQuAD dataset, each context passage is associated with multiple question-answer pairs stored as separate entries. In this modified version, all questions associated with the same context… See the full description on the dataset page: https://huggingface.co/datasets/Siam0703/SQuAD-EN-Passage-to-Question.psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles.
The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-advanced-questions.product-catalog-questions
Lamini Product Catalog QA Dataset
Description
This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle.
Format
The questions and product information are in the form of jsonlines file.
Data Pipeline Code
The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.Math-Question-AnswerSQuAD-BN-Passage-to-Question
Dataset Card for SQuAD-BN-Passage-to-Question
Dataset Summary
SQuAD-BN-Passage-to-Question is a reformatted and filtered version of the Bangla Question Answering dataset derived from csebuetnlp/squad_bn. The dataset is designed for Bangla text generation and question generation research tasks.
In the original dataset, each context passage is associated with multiple question-answer pairs stored as separate entries. In this modified version:
All questions associated with… See the full description on the dataset page: https://huggingface.co/datasets/Siam0703/SQuAD-BN-Passage-to-Question.Question-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.NPC-Quest-Dialogue
NPC-Quest-Dialogue
A successor to NPC-Dialogue_v2
Hunting for good quality NPC RP datasets are not an easy job, and that is why we took on the job to create high quality ones ourselves.
Our previous release of NPC-Dialogue_v2 had good feedback, so by using the same strategy we created NPC-Quest-Dialogue, a high quality dataset for NPC behavior in video games, specifially focusing on quest-related conversations.
Dataset Statistics
Total source quests: 1,994 RPG… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-Quest-Dialogue.questions
📝 Overview
Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی.
🎯 Purpose
دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی:
Pashto instruction-tuning
Pashto question-answering
Pashto reasoning
Pashto dialogue modeling
Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.Question-answeringsmall-ru
Dataset Card for Question Answering Russian Dataset
🧠 Quick Summary
Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях.
📚 Dataset Details
Curated by: @kurumikz
Language(s): Russian (ru)
License: CC-BY 4.0 — свободное использование с обязательным указанием автора
Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.spm-synthetic-questions
Dataset Card: SPM Synthetic Questions
Dataset Summary
14,135 synthetic exam-style questions for Malaysia's SPM (Sijil Pelajaran
Malaysia) curriculum, Form 5, covering 10 subjects. Every item is
LLM-generated and aligned to the KSSM curriculum and the SPM examination
format. Each item belongs to one of three categories:
hots — Higher Order Thinking Skills (KBAT) questions
lazim — soalan lazim (commonly-asked question styles)
perangkap — soalan perangkap (trap… See the full description on the dataset page: https://huggingface.co/datasets/VixeroAI/spm-synthetic-questions.EVM-QuestBench
EVM-QuestBench
EVM-QuestBench is an execution-grounded benchmark for evaluating whether large language models and AI agents can translate natural-language blockchain intent into executable transaction code that produces the intended EVM state transition.
Released with the ACL 2026 Long Paper by Pei Yang, Wanyi Chen, Ke Wang, Lynn Ai, Eric Yang, and Tianyu Shi, the benchmark contains 107 expert-authored tasks: 62 atomic tasks and 45 composite workflows. Unlike code-similarity… See the full description on the dataset page: https://huggingface.co/datasets/berryccc1/EVM-QuestBench.BenchMAX_Question_Answering
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios.
The subtasks are similar to the subtasks in RULER.
The data is sourcing from UN Parallel Corpus and xquad.
The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.cultural-questions-dataset
Turkish General Knowledge & Trivia CoT Dataset (TR-GenK-CoT)
TR-GenK-CoT is a high-quality, synthetic, and carefully curated Turkish dataset designed for Instruction Tuning and Chain of Thought (CoT) reasoning. It contains exactly 500 completely unique, non-repetitive general knowledge and trivia conversations with rich step-by-step thinking processes.
The dataset is formatted using standard Chat Template formats (matching OpenAI/Hugging Face chat schemas) making it directly… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/cultural-questions-dataset.Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024,
title={Questions_Answers_In_Sinhala_Language},
author={Ayesha Kalpani},
year={2024},
url={},
}
Questions_Answers_In_Sinhala_Language
Dataset Description
A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala.
Dataset Details
License
This dataset is licensed under the MIT License.
Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.ridiculous_math_questions
Dataset Card for Ridiculous Math Questions
A Set of ridiculous math questions that you won't find a teacher to write!
Dataset Details
Dataset Description
This dataset is a list of math questions generated by large language a model.
Which model is used depends on the version:
v0.05 was written by a 20B model, specifically DaringMaid-20B-V1.1-6bpw-exl2.
Curated by: KaraKaraWitch
Funded by [optional]: N/A
Shared by [optional]: KaraKaraWitch
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/WitchesSocialStream/ridiculous_math_questions.Pidgin_Question-English_Answer_Dataset
Pidgin Question - English Answer Dataset (Sample)
Data Card v1.0
Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/
Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.RolePlay-NPC-Quest
RolePlay-NPC-Quest
The newest RP dataset containing some high-quality dataset for Gemma4NPC.
A combination of pippa, NPC-Quest-Dialogue, and NPC-Dialogue_v2.
Shuffled to ensure proper training quality.
Note: There are some REALLY long conversations in PIPPA (upwards to 11493 chat turns), which will get truncated during training, those should be removed for training.
corpus-presse-question-algerienne-dpo-analyse-qualitative-608
Corpus Presse: Question algérienne — qualitative analysis-only DPO pairs (608)
This dataset contains 608 preference pairs for qualitative stance analysis of
Arabic and French historical newspaper articles about the Algerian question and
French colonial order.
Each training row has only:
prompt
chosen
rejected
The model-facing rows deliberately avoid numeric stance scores. The pairs are
intended to teach evidence weighting, source-frame interpretation,
quote/reporting… See the full description on the dataset page: https://huggingface.co/datasets/yakz1/corpus-presse-question-algerienne-dpo-analyse-qualitative-608.questionizer
Questionizer
A dataset of sentences and the questions that they answer.
Propositions were randomly sampled from agentlans/wikipedia-propositions
Then rewritten as questions using google/gemma-3-12b-it
Limitations
Some question-answer pairs sound unnatural
Lacks context when processing single sentences
pharmacy_rx_questions
Pharmacy_Rx_Questions (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Pharmacy Prescription Inquiries & Advisory Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/pharmacy_rx_questions.MatLab-Questions-Answers
Dataset Card for Dataset Name
Matlab-Questions-Answers
Dataset Details
Contains 150+ matlab/octave related questions and answers.
Dataset Description
Contains 150+ matlab/octave related questions and answers ranging from Grade School to Graduate level mathematics.
Language(s) (NLP): English
License: Apache 2.0
Uses
Small Language Models on matlab/octave specific code generation
Evaluation of Language Models on Matlab related questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/EmAllen-TW/MatLab-Questions-Answers.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.Gemma4NPC-Quest-Dataset
Dataset Card for Gemma4NPC Preference Dataset
Dataset Description
The Gemma4NPC Preference Dataset is a specialized text-generation and reinforcement learning dataset designed to train Large Language Models (LLMs) for use as Non-Playable Characters (NPCs) in video games.
Integrating LLMs into game engines requires models that can seamlessly blend creative roleplay with strict formatting requirements. This dataset addresses two primary training objectives:… See the full description on the dataset page: https://huggingface.co/datasets/spy5er/Gemma4NPC-Quest-Dataset.
