CoolFace
Datasetpublic

fineset-io/mechanistic-interpretability-papers

Mechanistic Interpretability Papers β€” FineSet A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. πŸ“Έ This is a dated snapshot β€” generated 2026-06-12. It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast β€” new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
1likes22downloads
Dataset Card

Mechanistic Interpretability Papers β€” FineSet

A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar.

πŸ“Έ This is a dated snapshot β€” generated 2026-06-12. It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast β€” new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓

Why this dataset

  • β€”Quality-scored: quality_score float (0–1), citation-normalized β€” filter out the noise
  • β€”Papers with code: 133 flagged via has_code β€” find reproducible work fast
  • β€”Deduplicated: arXiv + Semantic Scholar cross-referenced, duplicate records merged
  • β€”Clean JSONL: 748 records, one per line, normalized fields β€” no encoding garbage

Dataset details

  • β€”Records: 748
  • β€”Date range: 2022–2026
  • β€”Snapshot date: 2026-06-12 (frozen β€” see note above)
  • β€”Sources: arXiv, Semantic Scholar (cross-referenced, duplicates merged)
  • β€”arXiv categories: cs.LG, cs.AI
  • β€”Quality scoring: citation-normalized, 0–1 (p50=0.119, p90=0.355)
  • β€”Format: JSONL, one record per line

Fields

FieldTypeDescription
idstringDeterministic SHA256 record id
sourceslistWhich sources contributed (arxiv, semantic_scholar)
titlestringPaper title
abstractstringFull abstract
authorslistAuthor names
categorieslistarXiv category codes
fieldsofstudylistSemantic Scholar field tags
published_datestringISO 8601 date
urlstringarXiv abstract URL
pdf_urlstring\nullOpen-access PDF if available
arxiv_idstring\nullarXiv identifier
doistring\nullDOI if available
citation_countintCitation count (Semantic Scholar)
influentialcitationcountintInfluential citations (Semantic Scholar)
has_codeboolCode repo detected in the arXiv comment
code_urlstring\nullGitHub URL if detected
venuestring\nullPublication venue
quality_scorefloat0–1, citation-normalized

Quality score methodology

quality_score = min(1.0, log10(citation_count + 1) / 4)

A citation-normalized heuristic: 0 for uncited papers, ~0.5 at 100 citations, ~0.75 at 1,000, 1.0 at 10,000+. Useful for filtering training data by impact.

πŸ‘‰ Want this on YOUR topic, updated daily?

This snapshot is frozen at 2026-06-12. The live FineSet pipeline keeps a dataset like this refreshed every day on whatever topic you describe β€” new papers in, dedup and quality scoring automatic, export as JSONL/Parquet or push straight to the Hub.

Try it now β€” it's live: β†’ fineset.io β€” describe your research topic in plain English and get a fresh, quality-scored dataset in minutes. Free to start.