CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llm-jp /AnswerCarefullygated AnswerCarefully 概要 AnswerCarefullyは日本語LLM 出力の安全性・適切性に特化したインストラクションデータセットです。 このデータセットは、英語の要注意回答を集めた Do-Not-Answer データセット の包括的なカテゴリ分類に基づき、人手で質問・回答ともに日本語サンプルを集めたオリジナルのデータセットです。 データセットの詳細については、こちらをご覧ください。 Overview AnswerCarefully is an instruction dataset specifically aimed at ensuring safety and appropriateness of LLM output in Japanese. This dataset consists of original pairs of questions and reference (safe) responses based on the extensive safety taxonomy proposed in… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/AnswerCarefully.text1K<n<10K179 likes27k downloads1mo agoHugging Face02community-datasets /yahoo_answers_topics Dataset Card for "Yahoo Answers Topics" Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.texttext-classification1M<n<10M63 likes8.4k downloads2y agoHugging Face03hbXNov /hle_math_exact_match_no_image_int_answerimagen<1K1 likes7.3k downloads2y agoHugging Face04hbXNov /hle_math_exact_match_no_image_int_answer_random128imagen<1K0 likes6.3k downloads2y agoHugging Face05LibrAI /do-not-answer Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs Overview Do not answer is an open-source dataset to evaluate LLMs' safety mechanism at a low cost. The dataset is curated and filtered to consist only of prompts to which responsible language models do not answer. Besides human annotations, Do not answer also implements model-based evaluation, where a 600M fine-tuned BERT-like evaluator achieves comparable results with human and GPT-4. Instruction… See the full description on the dataset page: https://huggingface.co/datasets/LibrAI/do-not-answer.tabulartext-generationn<1K57 likes3.6k downloads3y agoHugging Face06FabienRoger /alignment_faking_harm_answerstext1K<n<10K0 likes3.3k downloads1y agoHugging Face07MathArena /final_answer_comps Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains all public final-answer competitions in MathArena. Thus, it includes: AIME 2025, HMMT 2025, CMIMC 2025, BRUMO 2025, and Apex 2025. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/final_answer_comps.textn<1K0 likes2.4k downloads4mo agoHugging Face08tamdd18 /CEH_question_answertextn<1K0 likes1.9k downloads2y agoHugging Face09Malikeh1375 /medical-question-answering-datasetstextquestion-answering1M<n<10M85 likes1.7k downloads6mo agoHugging Face10aisingapore /NLU-Question-Answeringgated SEA Question Answering SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese. Supported Tasks and Leaderboards SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.texttext-generation1K<n<10K0 likes1.6k downloads9mo agoHugging Face11vanloc1808 /pico-banana-smolvlm-format-with-rejected-answer pico-banana-smolvlm-format-with-rejected-answer Balanced image-level tampering detection dataset in SmolVLM-style format with chosen/rejected answer pairs, derived from the pico-banana MCQ pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training. Dataset overview Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a rejected_answer field: the answer from the counterpart sample (same edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.image100K<n<1M1 likes1.1k downloads7mo agoHugging Face12Asap7772 /hendrycks_math_with_answerstext10K<n<100K1 likes624 downloads2y agoHugging Face13PrimeIntellect /stackexchange-question-answering SYNTHETIC-1 This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here text100K<n<1M16 likes605 downloads2y agoHugging Face14Hwilner /imo-answerbench IMO-AnswerBench Dataset Description IMO-AnswerBench is a benchmark dataset for evaluating the mathematical reasoning capabilities of large language models. It consists of 400 challenging short-answer problems from the International Mathematical Olympiad (IMO) and other sources. This dataset is part of the IMO-Bench suite, released by Google DeepMind in conjunction with their 2025 IMO gold medal achievement. Supported Tasks and Leaderboards The primary task… See the full description on the dataset page: https://huggingface.co/datasets/Hwilner/imo-answerbench.textn<1K5 likes512 downloads11mo agoHugging Face15answerdotai /enwiki English Wikipedia as clean md This dataset is a cleaned, structurally faithful approximation of the English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline. It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset creation. The… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/enwiki.tabular10M<n<100M5 likes412 downloads10d agoHugging Face16hirundo-io /iheval-benign-answerstextn<1K0 likes369 downloads8mo agoHugging Face17cfierro /sycophancy_eval_answertext1K<n<10K0 likes342 downloads1y agoHugging Face18answerdotai /simplewiki Simple English Wikipedia as clean md This dataset is a cleaned, structurally faithful approximation of the Simple English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline. It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/simplewiki.tabular100K<n<1M4 likes337 downloads11d agoHugging Face19nbtpj /multi-context-long-answer-datasettext1M<n<10M11 likes326 downloads4y agoHugging Face20livebench /model_answer Dataset Card for "livebench/model_answer" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/model_answer.text10K<n<100K1 likes309 downloads2y agoHugging Face21WenxingZhu /msmarco_answerai_colbert_small_embeddings MS MARCO ColBERT Embeddings Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1. Dataset Structure The dataset contains: data/corpus/: 177 parquet files with document embeddings data/queries/: 11 parquet files with query embeddings data/qrels/train.parquet: Relevance judgments (532,751 pairs) Usage from datasets import load_dataset # Load from directory (recommended for large datasets) corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.tabularfeature-extraction1M<n<10M0 likes308 downloads11mo agoHugging Face22sentence-transformers /yahoo-answers Dataset Card for Yahoo Answers This dataset is a collection of pairs containing titles, questions, and answers collected from Yahoo Answers. See the Yahoo Answers dataset for additional information. This dataset can be used directly with Sentence Transformers to train embedding models. Dataset Subsets title-question-answer-pair subset Columns: "question", "answer" Column types: str, str Examples:{ 'question': "why doesn't an optical mouse work on a glass… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/yahoo-answers.textfeature-extraction1M<n<10M12 likes291 downloads2y agoHugging Face23OpenEvals /IMO-AnswerBench IMO-AnswerBench Dataset Description IMO-AnswerBench is a benchmark dataset for evaluating the mathematical reasoning capabilities of large language models. It consists of 400 challenging short-answer problems from the International Mathematical Olympiad (IMO) and other sources. This dataset is part of the IMO-Bench suite, released by Google DeepMind in conjunction with their 2025 IMO gold medal achievement. Supported Tasks and Leaderboards The primary task… See the full description on the dataset page: https://huggingface.co/datasets/OpenEvals/IMO-AnswerBench.textn<1K2 likes285 downloads8mo agoHugging Face24Lots-of-LoRAs /task453_swag_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task453_swag_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task453_swag_answer_generation.texttext-generation1K<n<10K0 likes258 downloads2y agoHugging Face25abhayesian /answers-with-reasoning-mmlu-pro answers-with-reasoning-mmlu-pro Self-distillation SFT corpus: Qwen3-8B-Instruct's own correct chain-of-thought rollouts on MMLU-Pro multiple-choice questions (general-QA domain). Generation Source problems: TIGER-Lab/MMLU-Pro test split (12,032 multiple-choice questions across 14 subject categories). Sampling model: qwen/qwen3-8b via OpenRouter (providers: Alibaba, AtlasCloud) with reasoning enabled. Sampling parameters: temperature=0.6, top_p=0.95, max_tokens=8000.… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/answers-with-reasoning-mmlu-pro.texttext-generation1K<n<10K0 likes252 downloads5mo agoHugging Face26petkopetkov /medical-question-answering-splittext100K<n<1M0 likes243 downloads2y agoHugging Face27Lots-of-LoRAs /task067_abductivenli_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task067_abductivenli_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task067_abductivenli_answer_generation.texttext-generation1K<n<10K0 likes237 downloads2y agoHugging Face28mteb /yahoo_answers_topicstext1M<n<10M0 likes230 downloads1y agoHugging Face29copenlu /answerable_tydiqa Dataset Card for "answerable-tydiqa" Dataset Summary TyDi QA is a question answering dataset covering 11 typologically diverse languages. Answerable TyDi QA is an extension of the GoldP subtask of the original TyDi QA dataset to also include unanswertable questions. Dataset Structure The dataset contains a train and a validation set, with 116067 and 13325 examples, respectively. Access them with from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/answerable_tydiqa.textquestion-answering100K<n<1M10 likes221 downloads2y agoHugging Face30addy88 /nq-question-answeronlytext100K<n<1M1 likes220 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.