datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
answercarefully-dpo-ja-2026
AnswerCarefully-derived Japanese DPO data for LLM safety
本データセットは、llm-jp/AnswerCarefullyを参照して作成した日本語LLMの安全応答をDPOで学習するためのpreference datasetです。
利用条件
本データセットには、llm-jp/AnswerCarefullyと同じ利用規約を適用します。
利用者は、llm-jp/AnswerCarefullyと本データセットの両方で利用規約に同意する必要があります。
データ
train: 417件
validation: 44件
各行には次のフィールドが含まれます。
id: 本リリース内だけで使用するID
prompt: 元質問の意味と危険性を変えずに言い換えた質問
chosen: DPOで望ましい応答として扱う回答
rejected: DPOで望ましくない応答として扱う回答
category, harm_type, risk_area… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/answercarefully-dpo-ja-2026.text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.minecraft-question-answer-700k
minecraft-question-answer-700k
Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline.
about the dataset
rows - 694,814
tokens - 47,133,624
source - https://minecraft.wiki/
Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles.
The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.IELTs-Speaking-answer
Overview
This dataset consists of 2 json files named 'ielts_new.json' and 'ielts_old.json', which contain ielts questions and its corresponding answers for part 1 and part 2.
'ielts_new.json': new IELTs topics for 2024 September-December.
'ielts_old.json': remained IELTs topics for 2024 September-December.
Quality
Since the dataset is analysed and generated by ChatGPT based on my own pdf file, the answer may be incomplete(only part of the sentence is extracted, leading to… See the full description on the dataset page: https://huggingface.co/datasets/qwertyuiopasdfg/IELTs-Speaking-answer.Math-Question-AnswerHuman-Style-Answers
Human Style Answers
This Datasets contains question and answers on different topics in Human style. (For Chatbots training)
This Datasets is build using TOP AI like (GPT4, Claude3 , Command R+, etc.)
Dataset Details
Description
The Human Style Response Dataset is a rich collection of question-and-answer pairs, meticulously crafted in a human-like style. It serves as a valuable resource for training chatbots and conversational AI models. Let's dive into the… See the full description on the dataset page: https://huggingface.co/datasets/innova-ai/Human-Style-Answers.scas_verified_teacher_pool
SCAS Verified Teacher Answer Pool
This dataset provides an aligned, correctness-verified pool of
teacher-generated mathematical reasoning solutions for studying
student-centric data selection in distillation.
The release covers two source corpora, Hendrycks MATH and DeepScaleR. For each
corpus, we retain the subset of questions on which all nine selected teacher
models produce verified correct answers. Each retained question is paired with
nine alternative teacher solutions, one… See the full description on the dataset page: https://huggingface.co/datasets/Student-Centric-Answer-Sampling/scas_verified_teacher_pool.einstein_answers
What would Einstein Say?
This dataset contains a set of questions and answers, mimicking Einstein's approach to answer general scientific and philosophical queries.
The data points have been generated synthetically, however the factual correctness of the data is ensured, not guaranteed whatsoever.
BenchMAX_Question_Answering
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios.
The subtasks are similar to the subtasks in RULER.
The data is sourcing from UN Parallel Corpus and xquad.
The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.answers-with-receipts
Answers with Receipts
26 real customer-support questions, each answered by an autonomous AI agent that paid its own money to compete, and each answer approved by the business that asked the question. Every row carries the on-chain transaction that paid the agent.
The preference label in this dataset is backed by a payment, not a click.
Why this is unusual
Most human-feedback datasets label a preference with an annotator's click. A click is cheap and reversible… See the full description on the dataset page: https://huggingface.co/datasets/deskcrew/answers-with-receipts.Question-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.Question-answeringsmall-ru
Dataset Card for Question Answering Russian Dataset
🧠 Quick Summary
Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях.
📚 Dataset Details
Curated by: @kurumikz
Language(s): Russian (ru)
License: CC-BY 4.0 — свободное использование с обязательным указанием автора
Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.do_not_answer_zh_response
puwaer/do_not_answer_en_response
This dataset is based on LibrAI/do-not-answer translated into Chinese, with model answers added for both positive and negative examples.
It is a dataset intended for the performance evaluation of reward models.
For the positive examples, outputs from deepseek-ai/DeepSeek-V3.2-Exp are used.
For the prompts and negative examples, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 are used.… See the full description on the dataset page: https://huggingface.co/datasets/puwaer/do_not_answer_zh_response.Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024,
title={Questions_Answers_In_Sinhala_Language},
author={Ayesha Kalpani},
year={2024},
url={},
}
Questions_Answers_In_Sinhala_Language
Dataset Description
A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala.
Dataset Details
License
This dataset is licensed under the MIT License.
Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.Pidgin_Question-English_Answer_Dataset
Pidgin Question - English Answer Dataset (Sample)
Data Card v1.0
Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/
Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.do_not_answer_en_response
puwaer/do_not_answer_en_response
This dataset is based on LibrAI/do-not-answer, with model answers added for both positive and negative examples.
It is a dataset intended for the performance evaluation of reward models.
For the positive examples, outputs from deepseek-ai/DeepSeek-V3.2-Exp are used.
For the negative examples, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 are used.
このデータセットは、LibrAI/do-not-answerをもとに、正例、負例に対して模範回答を付与したデータセットです。
reward… See the full description on the dataset page: https://huggingface.co/datasets/puwaer/do_not_answer_en_response.boardgamebench-answer-dpo
BoardGameBench Answer DPO Dataset
This dataset contains the reviewed preference examples used for the DPO stage of the nemotron-boardgame-answer-lora-b4-safe-final adapter.
It is a compact pilot set of 10 BoardGameBench preference rows. Each row presents the same board-game decision prompt with a preferred answer and a plausible rejected answer. The preferred answer is selected from engine-guided move comparisons and includes the exact move label.
Format
The main… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/boardgamebench-answer-dpo.do_not_answer_jp_response
puwaer/do_not_answer_jp_response
This dataset is based on kunishou/do-not-answer-ja, with model answers added for both positive and negative examples.
It is a dataset intended for the performance evaluation of reward models.
For the positive examples, outputs from deepseek-ai/DeepSeek-V3.2-Exp are used.
For the negative examples, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 are used.
このデータセットは、kunishou/do-not-answer-jaをもとに、正例、負例に対して模範回答を付与したデータセットです。
reward… See the full description on the dataset page: https://huggingface.co/datasets/puwaer/do_not_answer_jp_response.khmer_question_answerThe data collected from https://www.khsearch.com/ related to the general question-answering examination.
It used to train fine-tuned models from many LLMs, including LlaMa, Qwen, Mistral, and Gemma.
Under the research title "Fine-tuning for Question Answering in Low-Resource Languages: A Case Study on Khmer" conducted at ViLa Lab, Institute of Technology of Cambodia, Phnom Penh.
Lab Info: https://www.facebook.com/vilalabitc
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/kimleang123/khmer_question_answer.MatLab-Questions-Answers
Dataset Card for Dataset Name
Matlab-Questions-Answers
Dataset Details
Contains 150+ matlab/octave related questions and answers.
Dataset Description
Contains 150+ matlab/octave related questions and answers ranging from Grade School to Graduate level mathematics.
Language(s) (NLP): English
License: Apache 2.0
Uses
Small Language Models on matlab/octave specific code generation
Evaluation of Language Models on Matlab related questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/EmAllen-TW/MatLab-Questions-Answers.grok_answer_mail_ru
Датасет ответов на Маил.ру
В этом датасете собраны ответы от Grok-3-latest (и немного chatgpt-4o-latest) на вопросы с Ответы Маил.ру
steve-jobs-question-and-answers
