datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
gvg-explainable-game-similarity
GVG Explainable Game Similarity Dataset 2026 (v0.1.0)
50 human-reviewed pairs of similar PC games (82 games) from Game V Game. Each row says why the two games are similar, what differs, and who each suits, in English and Chinese, and cites the two official Steam store records the claims were checked against (with capture dates).
Most similarity data says two games are alike and stops. This one is small on purpose: every row was read by a person against the current Steam records… See the full description on the dataset page: https://huggingface.co/datasets/lette2/gvg-explainable-game-similarity.explainable-user-profilesExplained user profiles for 50 users based on IMDB Dataset created by GPT-4o.
explainDepression-social-media-xaiCleaned, balanced, and clinically annotated social media dataset
for explainable depression detection research.
Combined from real Twitter and Reddit posts, engineered with
DSM-5-aligned clinical lexicon features, and used to train a
DistilBERT model achieving 96.17% accuracy.
━━━━━━━━━━━━━━━━━━━━━━━━━━━
DATASET STATS
Total Rows → 40,770
Class Balance → 50% Depressed / 50% Not Depressed
Feature Columns → 12
Sources → Twitter + Reddit
━━━━━━━━━━━━━━━━━━━━━━━━━━━… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/explainDepression-social-media-xai.linguistic-echoes-explainable-nlp-stress-detectionautotrain-data-ai-explains-codeExplainableAI-emotions-DPO-ORPO-RLHF
Preference Dataset for Explainable Multi-Label Emotion Classification
This repository contains a preference dataset compiled to compare two model-generated responses for explaining multi-label emotion classifications on Tweets. The dataset is accompanied by human annotations indicating which response was preferred, based on a set of defined dimensions (clarity, correctness, helpfulness, and verbosity). The annotation guidelines are included to describe how these preference judgments… See the full description on the dataset page: https://huggingface.co/datasets/imhmdf/ExplainableAI-emotions-DPO-ORPO-RLHF.WikiSQL_explainedsarcasm-explain-5k
SarcasmExplain-5K
Dataset Description
SarcasmExplain-5K is a balanced dataset of 5,000 Reddit sarcasm
instances annotated with five complementary natural language explanation
types, generated via GPT-4 and validated through human evaluation.
Created by: Maliha Binte Mamun
Year: 2025
License: CC BY 4.0
GitHub: https://github.com/maliha-usui/sarcasm-explain-5k
🔒 Access
Complete 3 annotation forms to download:
👉… See the full description on the dataset page: https://huggingface.co/datasets/maliha/sarcasm-explain-5k.autotrain-data-ai-explains-codeEmo-Explainabilityts-detect-test-smell-gemini-explained11k_History_MCQAs_gen_ExplainExplainLikeIm5_Annotated_Dataset
LLM Argumentation Preference Dataset
Dataset created for the NLP Research Course 097920 (Technion).
Each example includes a user query and two responses annotated by 3 human annotators for preference, source identification etc.
🧩 Tasks
The dataset includes four main annotation tasks:
Preference Task – Which response is easier to understand?
Source Identification Task – Which response is written by a human or an AI?
Appeal to Expert Task – Does the response's writer… See the full description on the dataset page: https://huggingface.co/datasets/SlowSenik/ExplainLikeIm5_Annotated_Dataset.11k_History_MCQAs_gen_Explain
