CoolFace
Datasetpublic

bigbio/nlmchem

NLM-Chem corpus consists of 150 full-text articles from the PubMed Central Open Access dataset, comprising 67 different chemical journals, aiming to cover a general distribution of usage of chemical names in the biomedical literature. Articles were selected so that human annotation was most valuable (meaning that they were rich in bio-entities, and current state-of-the-art named entity recognition systems disagreed on bio-entity recognition.

sourceHugging Facecc0-1.0updated 4y agoView on Hugging Face
3likes51downloads

bigbio/nlmchem · main · files are served by the source, never re-hosted here