datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AnswerCarefully
AnswerCarefully
概要
AnswerCarefullyは日本語LLM 出力の安全性・適切性に特化したインストラクションデータセットです。
このデータセットは、英語の要注意回答を集めた Do-Not-Answer データセット の包括的なカテゴリ分類に基づき、人手で質問・回答ともに日本語サンプルを集めたオリジナルのデータセットです。
データセットの詳細については、こちらをご覧ください。
Overview
AnswerCarefully is an instruction dataset specifically aimed at ensuring safety and appropriateness of LLM output in Japanese.
This dataset consists of original pairs of questions and reference (safe) responses based on the extensive safety taxonomy proposed in… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/AnswerCarefully.answercarefully-dpo-ja-2026
AnswerCarefully-derived Japanese DPO data for LLM safety
本データセットは、llm-jp/AnswerCarefullyを参照して作成した日本語LLMの安全応答をDPOで学習するためのpreference datasetです。
利用条件
本データセットには、llm-jp/AnswerCarefullyと同じ利用規約を適用します。
利用者は、llm-jp/AnswerCarefullyと本データセットの両方で利用規約に同意する必要があります。
データ
train: 417件
validation: 44件
各行には次のフィールドが含まれます。
id: 本リリース内だけで使用するID
prompt: 元質問の意味と危険性を変えずに言い換えた質問
chosen: DPOで望ましい応答として扱う回答
rejected: DPOで望ましくない応答として扱う回答
category, harm_type, risk_area… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/answercarefully-dpo-ja-2026.yahoo_answers_topics
Dataset Card for "Yahoo Answers Topics"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.hle_math_exact_match_no_image_int_answerhle_math_exact_match_no_image_int_answer_random128do-not-answer
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Overview
Do not answer is an open-source dataset to evaluate LLMs' safety mechanism at a low cost. The dataset is curated and filtered to consist only of prompts to which responsible language models do not answer.
Besides human annotations, Do not answer also implements model-based evaluation, where a 600M fine-tuned BERT-like evaluator achieves comparable results with human and GPT-4.
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/LibrAI/do-not-answer.alignment_faking_harm_answersfinal_answer_comps
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains all public final-answer competitions in MathArena. Thus, it includes: AIME 2025, HMMT 2025, CMIMC 2025, BRUMO 2025, and Apex 2025.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/final_answer_comps.CEH_question_answermedical-question-answering-datasetsNLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.pico-banana-smolvlm-format-with-rejected-answer
pico-banana-smolvlm-format-with-rejected-answer
Balanced image-level tampering detection dataset in SmolVLM-style format
with chosen/rejected answer pairs, derived from the pico-banana MCQ
pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training.
Dataset overview
Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a
rejected_answer field: the answer from the counterpart sample (same
edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.hendrycks_math_with_answersstackexchange-question-answering
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
reddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit.
The score for the question / answer is the upvote count (i.e. positive-negative upvotes).
Only questions / answers that have these properties were extracted:
min_score = 3
min_title_len = 20
min_body_len = 100
imo-answerbench
IMO-AnswerBench
Dataset Description
IMO-AnswerBench is a benchmark dataset for evaluating the mathematical reasoning capabilities of large language models. It consists of 400 challenging short-answer problems from the International Mathematical Olympiad (IMO) and other sources.
This dataset is part of the IMO-Bench suite, released by Google DeepMind in conjunction with their 2025 IMO gold medal achievement.
Supported Tasks and Leaderboards
The primary task… See the full description on the dataset page: https://huggingface.co/datasets/Hwilner/imo-answerbench.cybersecurity_full_question_answersgpqa-metadata-blind-answerscugnizz-v22-tool-answer
scugnizz-v22-tool-answer
Synthetic tool-loop data teaching grounded answers and correct refusal on empty/error tool results.
Format: Hermes/OpenAI-style messages plus tools.
enwiki
English Wikipedia as clean md
This dataset is a cleaned, structurally faithful approximation of the English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline.
It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset creation. The… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/enwiki.iheval-benign-answerssycophancy_eval_answerdo-not-answer-scenarios
Do-Not-Answer Scenarios
Refusal-evaluation scenarios derived from the Do-Not-Answer dataset, serialized as
giskard Scenario objects (one JSON object per line).
Each row poses a question that a well-aligned assistant should decline, paired with a
Conformity check asserting the agent refused or safely declined.
Format
Every line is a serialized giskard.checks.Scenario:
name — "Do not answer #<id>"
steps[].interacts[].inputs — the question to send to the agent under… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/do-not-answer-scenarios.simplewiki
Simple English Wikipedia as clean md
This dataset is a cleaned, structurally faithful approximation of the Simple English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline.
It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/simplewiki.multi-context-long-answer-datasetmodel_answer
Dataset Card for "livebench/model_answer"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/model_answer.
