datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCIERCSciERCSCIERC (Luan et al., 2018) via "Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks" (Gururangan et al., 2020) reuploaded because of error encountered when trying to load zj88zj/SCIERC with the huggingfaces/datasets library.
scierc_processed_datasciERCsciercsci-ercscierc_aeco
Dataset Card for SciERC AECO dataset
Dataset Summary
The SciERC AECO dataset is an English-language dataset containing 1016 sentences from research papers in the AECO domain, annotated for scientific entities and relations based on the SciERC annotation schema.
Supported Tasks and Leaderboards
'NER': the dataset can be used to train a model to detect scientific entities according to the SciERC annotation schema
'Relation extraction': the dataset can be used… See the full description on the dataset page: https://huggingface.co/datasets/zavavan/scierc_aeco.SCIERCscierc_v1relabel_SciERC
Dataset Card for "relabel_SciERC"
More Information needed
SCIERC-PT
SCIERC-PT
Overview
SCIERC-PT is a European Portuguese (PT-PT) translation of the SCIERC dataset, created to support research on Scientific Information Extraction in Portuguese. The dataset provides translated scientific abstracts suitable for evaluating Named Entity Recognition (NER) and Relation Extraction (RE) models while preserving the original annotation schema.
The dataset was automatically translated and subsequently curated to improve alignment between the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SCIERC-PT.scierc_v2_processedmethod_only_scierc
