datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
casimedicos-exp
Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams
We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments
for the correct answer but also arguments to explain why the remaining possible answers are incorrect.
This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation.
The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.ExploreToM
Data sample for ExploreToM: Program-guided adversarial data generation for theory of mind reasoning
ExploreToM is the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation.
Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs.
Our A* search procedure aims to find… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ExploreToM.expert-rag-benchmarks
Expert RAG Benchmarks
A unified collection of four expert-level legal RAG benchmarks, exposed as six
named splits and three relational configurations: questions, documents, and
qrels.
The KCL split is named kcl_essay because Hugging Face split identifiers do not
permit hyphens; its source name remains kcl-essay.
Loading
from datasets import load_dataset
repo_id = "jinulee-v/expert-rag-benchmarks"
questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.Open-Omega-Explora-2.5M
Open-Omega-Explora-2.5M
Open-Omega-Explora-2.5M is a high-quality, large-scale reasoning dataset blending the strengths of both Open-Omega-Forge-1M and Open-Omega-Atom-1.5M. This unified dataset is crafted for advanced tasks in mathematics, coding, and science reasoning, featuring a robust majority of math-centric examples. Its construction ensures comprehensive coverage and balanced optimization for training, evaluation, and benchmarking in AI research, STEM education, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Explora-2.5M.ExploitDB_DataSet
🛡️ ExploitDB Cybersecurity Dataset
A comprehensive cybersecurity dataset containing 70,233 vulnerability records from ExploitDB, processed and optimized for machine learning and security research.
📊 Dataset Overview
This dataset provides structured information about cybersecurity vulnerabilities, exploits, and security advisories collected from ExploitDB - one of the world's largest exploit databases.
🎯 Key Statistics
Total Records: 70,233 vulnerability… See the full description on the dataset page: https://huggingface.co/datasets/Waiper/ExploitDB_DataSet.cibench-experiments
CIBench Experiments
Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era.
If a benchmark result cannot be replayed from its manifest alone, it did not happen.
Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.expertqa
Dataset Card for ExpertQA
Dataset Summary
We provide here the data accompanying the paper: ExpertQA: Expert-Curated Questions and Attributed Answers. The ExpertQA dataset contains 2177 examples from 32 different fields.
Supported Tasks
The main data contains 2177 examples that can be used to evaluate new methods for estimating factuality and attribution, while the lfqa_domain and lfqa_rand data can be used to evaluate long-form question answering systems.… See the full description on the dataset page: https://huggingface.co/datasets/cmalaviya/expertqa.pentesting-explanations
Pentesting Explanations - Adversarial Reasoning & Vulnerability Research
A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names.
The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/pentesting-explanations.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
balanced-copa-explanations
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.Causal-Intervention-Tests-For-Explanation-Faithfulness
Faithfulness via Causal Interventions — Evaluation Pipeline
Paper: What Does Answer Change Rate Actually Measure? A Specificity Audit
of Causal Intervention Tests for Explanation Faithfulness
Accepted at: EMNLP 2026 Workshop GroundLM, Budapest, Hungary (emnlp.org)
This pipeline implements the causal-intervention evaluation for LLM
explanation faithfulness described in the accompanying paper, including two
controls: a content-free specificity check and a decoding-noise floor.… See the full description on the dataset page: https://huggingface.co/datasets/durgesh-rao/Causal-Intervention-Tests-For-Explanation-Faithfulness.fine-grained-medical-reasoning
Dataset Card for Fine-Grained Medical Reasoning
Fine-grained medical reasoning QA dataset introduced in "Can LLMs Reason Like Doctors? Exploring the Limits of Large Language Models in Complex Medical Reasoning"
(Findings of EACL 2026). Manually annotated from the MedAgentsBench test_hard set,
it evaluates LLMs’ abduction, deduction, and induction capabilities, offering detailed insights into physician-like reasoning.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/expertailab/fine-grained-medical-reasoning.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.swahili-language-exposure
swahili-language-exposure
Dataset Summary
swahili-language-exposure is a large-scale Swahili (Kiswahili) corpus designed for language exposure and continued pretraining of language models.
Unlike instruction-tuning datasets, this dataset focuses on exposing models to natural Swahili usage across conversations, explanations, narratives, technical discussions, and mixed-domain text. The goal is to improve fluency, vocabulary coverage, syntax, and cultural grounding in… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-language-exposure.eli5_rlhf_explainlikeim5
ELI5 paired
This is a processed version of the eli5 dataset.
Compared to "eli5_rlhf", this dataset contains only QA pairs from the train split of the eli5 dataset and only from the subreddit explainlikeimfive.
Furthermore, the function
def get_question(example):
title = example["title"]
selftext = example["selftext"]
if selftext:
if selftext[-1] not in [".", "?", "!"]:
seperator = ". "
else:
seperator = " "
question = title… See the full description on the dataset page: https://huggingface.co/datasets/vincentmin/eli5_rlhf_explainlikeim5.pentesting-explanations
Pentesting Explanations - Adversarial Reasoning & Vulnerability Research
A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names.
The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-explanations.swahili-language-exposure-v2
Swahili Language Exposure
Large-scale Swahili corpus for continued pretraining and language exposure.
Maintained by NileAGI.
ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
ReLU
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
Picture: The pipeline of using open-weight LLMs to train/finetune over new information (Finetuned-LLM). Later,
when an unlearning request arises, the new information is split into the Retain and Forget set. The Unlearning
algorithms aim towards achieving the Target-LLM (trained/finetuned only on the Retain set) with a cost lower
than training/finetuning the pretrained open-weight LLM again. The spider plot… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/ReLU.qa-expert-multi-hop-qa-V1.0
Dataset Card for QA-Expert-multi-hop-qa-V1.0
This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering.
In total, this dataset contains 25.5k for training and 3.19k for evaluation.
You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0
The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.MM-TelecoBenchARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
Standard-Multimodal-Explanation
Dataset Card for Standard Multimodal Explanation (SME)
This is a dataset for Multimodal Explanation for Visual Question Answering (MEVQA).
Dataset Details
Dataset Description
This dataset contains questions, images, answers, and the multimodal explanations of the underlying reasoning process.
The explanations are in standard English with additional [BOX] for visual grounding.
Language(s) (NLP): English
License: apache-2.0
Modality:
Language… See the full description on the dataset page: https://huggingface.co/datasets/LivXue/Standard-Multimodal-Explanation.XXXXXL-chain-of-thought-expanded
XXXXXL Chain-of-Thought — Expanded
A 1,195-row expansion of wop/XXXXXL-chain-of-thought (11 rows → 1,195 rows), preserving its "Narrative Technicality" style: stream-of-consciousness inner monologue that performs explicit, low-level verification inside the prose rather than as a labeled checklist. The original <|thinking|> tag is re-cased to the standard <think>...</think> for drop-in interop with DeepSeek-R1 / Qwen-QwQ pipelines. The original 11 anchors are preserved as rows… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/XXXXXL-chain-of-thought-expanded.ethics_expression_preferencesReformat of the ethics dataset to be in DPO format, in this expression format
Examples:
subset commonsense
before 13910 {'label': 1, 'input': "I went to the principal's office to change my records before going to a different school."}
after 1513 {'prompt': 'I won my match by ', 'chosen': 'making a lot of good shots.', 'rejected': "moving my ball when the other player wasn't looking."}
subset utilitarianism
before 13737 {'baseline': 'I forgot to wear my virus-proof mask when I visited… See the full description on the dataset page: https://huggingface.co/datasets/wassname/ethics_expression_preferences.PolyFiQA-Expert
Dataset Card for PolyFiQA-Expert
Dataset Summary
PolyFiQA-Expert is a multilingual financial question-answering dataset designed to evaluate expert-level financial reasoning in low-resource and multilingual settings. Each instance consists of a task identifier, a query prompt, an associated financial question, and the correct answer.The Expert split emphasizes complex, high-level financial understanding, requiring deeper domain knowledge and nuanced reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/PolyFiQA-Expert.Privacy-Expert-Instructions
Dataset Card: Privacy-Expert-Instructions
This dataset contains 13000+ high-quality instruction-tuning pairs focused on Privacy and Data Protection. The data was curated from several StackExchange communities (Security, SuperUser, StackOverflow, etc.) and processed into a clean Alpacca-style format.
Dataset Summary
The primary goal of this dataset is to provide fine-tuning data for LLMs to understand and answer questions regarding:
Online Privacy: Tracking, anonymity… See the full description on the dataset page: https://huggingface.co/datasets/meeAtif/Privacy-Expert-Instructions.science-cot-dataset
ExpertData Science — Scientific Reasoning
Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean
Each record captures a complete experimental or theoretical reasoning chain:
Hypothesis → Methodology → Causal Chain → Validated Conclusion.
Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction.
This dataset is produced by the ExpertData-Factory pipeline
(Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.23f-expediente
Expediente 23-F: Declassified Documents from the 1981 Spanish Coup Attempt
A page-level multimodal dataset of declassified Spanish government documents related to the failed coup d'etat of February 23, 1981 (known as 23-F). Each row contains a rendered page image paired with its extracted text, sourced from official state archives.
Historical Context
On February 23, 1981, Lieutenant Colonel Antonio Tejero stormed the Spanish Congress of Deputies with armed Guardia Civil… See the full description on the dataset page: https://huggingface.co/datasets/JorgeAV/23f-expediente.ExploreToM
PNYX/ExploreToM
This is a enriched version of the facebook/ExploreToM.
This version is designed to be executed with the lm-eval package using the A-VERT methodology.
It contains the same data as the original dataset, but with additional columns to facilitate systematic evaluation of reasoning across different orders of theory of mind.
Note: In the A-VERT repository can be found the task definition in yaml format to be used with lm-evaluation-harness.
New Columns… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/ExploreToM.
