datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mechanic-Interpretability-Research-Datamechanistic-interpretability-papers
Mechanistic Interpretability Papers — FineSet
A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.Emotional_Interpretability
Components
Dataset :
Emotional_perspectives : Response to a given context under 27 emotional lenses
Description
Synthetic dataset created to mimic emotional responses primarily made for alignment and interpretability research
More details will be listed on github soon
license: mit
llm-interpretability-v1simulation-interpretability-dataset
Qwen3.5-2B-Base Blind Spots Dataset
A curated dataset documenting systematic failure modes ("blind spots") discovered in Qwen/Qwen3.5-2B-Base through structured probing experiments.
Dataset Description
This dataset contains 12 carefully selected examples where Qwen3.5-2B-Base exhibits predictable, reproducible failures across three major categories:
Category
Examples
Key Finding
Authority-Induced Sycophancy
4
Model accepts false claims when framed with… See the full description on the dataset page: https://huggingface.co/datasets/Znreza/simulation-interpretability-dataset.llm-interpretability-v1-messages
