datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acl-anthology-md
ACL Anthology Markdown Corpus
A snapshot of the ACL Anthology consisting of bibliographic metadata for 120,034 papers and full-text markdown conversions of 114,484 papers (≈95% of the catalogue, the remainder are frontmatter, abstract-only entries, or papers without an available PDF).
This corpus is the document collection used by ACL-Verbatim, a hallucination-free question-answering system for NLP research papers built on top of VerbatimRAG.
Configurations
The… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/acl-anthology-md.acl-anthology-atomic-claims
ACL Anthology Atomic Contribution Claims (ACC)
346,010 atomic contribution claims extracted from the abstracts of 80,144 ACL Anthology papers — 423 venues, 1963–2026.
An atomic contribution claim (ACC) is a single self-contained sentence stating one concrete contribution of a paper: atomic (exactly one contribution-bearing proposition), decontextualized (pronouns resolved, meta-language removed), and falsifiable (a verifiable assertion). Unlike raw abstracts or keywords, ACCs… See the full description on the dataset page: https://huggingface.co/datasets/Hamyrappy/acl-anthology-atomic-claims.aclAnthology-9k-filtered
Description
This dataset contains filtered ACL papers between certain abstract length ranges. The data includes columns such as paper_name, year, venue, url, bibkey, and cite_acl.
We also have the filtered_dataset.jsonl that holds the main text info.
Note: Some records might be missing certain fields, especially bibkey or cite_acl. We plan to fill them via partial manual / fuzzy matching.
License
Materials prior to 2016: CC BY-NC-SA 3.0.
Materials from… See the full description on the dataset page: https://huggingface.co/datasets/yilmazzey/aclAnthology-9k-filtered.acl-anthology_sum
