yuz9yuz/NLPImpact
This dataset accompanies our ACL 2025 paper: Internal and External Impacts of Natural Language Processing Papers. We present a comprehensive dataset for analyzing both the internal (academic) and external (public) impacts of NLP papers published in top-tier conferences — ACL, EMNLP, and NAACL — between 1979 and 2024. The dataset supports a wide range of scientometric studies, including topic-level impact evaluation across patents, media, policy, and code repositories. Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/yuz9yuz/NLPImpact.
This dataset accompanies our ACL 2025 paper: Internal and External Impacts of Natural Language Processing Papers.
We present a comprehensive dataset for analyzing both the internal (academic) and external (public) impacts of NLP papers published in top-tier conferences — ACL, EMNLP, and NAACL — between 1979 and 2024. The dataset supports a wide range of scientometric studies, including topic-level impact evaluation across patents, media, policy, and code repositories.
Data Sources
Our dataset integrates signals from several open and restricted resources:
- ACL Anthology: NLP papers
- OpenAlex: Citation counts
- Reliance on Science: Patent-to-paper links.
- Papers with Code + GitHub API: Linking NLP papers to GitHub repositories, including stars and forks.
⚠️ Not publicly included:
Altmetric (media-to-paper links) and Overton (policy-document-to-paper links) are used in our analysis but are not released here due to data access restrictions. Approval from Altmetric and Overton needed to access these signals. Please refer to our paper for more details.
Dataset Format
Each record corresponds to one NLP paper and includes the following fields:
Citation
If you find this dataset useful, please cite the following paper:
@article{zhang2025internal,
title={Internal and External Impacts of Natural Language Processing Papers},
author={Zhang, Yu},
journal={arXiv preprint arXiv:2505.16061},
year={2025}
}