fineset-io/mechanistic-interpretability-papers
Mechanistic Interpretability Papers β FineSet A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. πΈ This is a dated snapshot β generated 2026-06-12. It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast β new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. β Why thisβ¦ See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.
Mechanistic Interpretability Papers β FineSet
A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar.
πΈ This is a dated snapshot β generated 2026-06-12. It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast β new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. β
Why this dataset
- Quality-scored:
quality_scorefloat (0β1), citation-normalized β filter out the noise - Papers with code: 133 flagged via
has_codeβ find reproducible work fast - Deduplicated: arXiv + Semantic Scholar cross-referenced, duplicate records merged
- Clean JSONL: 748 records, one per line, normalized fields β no encoding garbage
Dataset details
- Records: 748
- Date range: 2022β2026
- Snapshot date: 2026-06-12 (frozen β see note above)
- Sources: arXiv, Semantic Scholar (cross-referenced, duplicates merged)
- arXiv categories: cs.LG, cs.AI
- Quality scoring: citation-normalized, 0β1 (p50=0.119, p90=0.355)
- Format: JSONL, one record per line
Fields
Quality score methodology
quality_score = min(1.0, log10(citation_count + 1) / 4)
A citation-normalized heuristic: 0 for uncited papers, ~0.5 at 100 citations, ~0.75 at 1,000, 1.0 at 10,000+. Useful for filtering training data by impact.
π Want this on YOUR topic, updated daily?
This snapshot is frozen at 2026-06-12. The live FineSet pipeline keeps a dataset like this refreshed every day on whatever topic you describe β new papers in, dedup and quality scoring automatic, export as JSONL/Parquet or push straight to the Hub.
Try it now β it's live: β fineset.io β describe your research topic in plain English and get a fresh, quality-scored dataset in minutes. Free to start.
