amalia-llm/SCIERC-PT
SCIERC-PT Overview SCIERC-PT is a European Portuguese (PT-PT) translation of the SCIERC dataset, created to support research on Scientific Information Extraction in Portuguese. The dataset provides translated scientific abstracts suitable for evaluating Named Entity Recognition (NER) and Relation Extraction (RE) models while preserving the original annotation schema. The dataset was automatically translated and subsequently curated to improve alignment between the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SCIERC-PT.
SCIERC-PT
Overview
SCIERC-PT is a European Portuguese (PT-PT) translation of the SCIERC dataset, created to support research on Scientific Information Extraction in Portuguese. The dataset provides translated scientific abstracts suitable for evaluating Named Entity Recognition (NER) and Relation Extraction (RE) models while preserving the original annotation schema.
The dataset was automatically translated and subsequently curated to improve alignment between the translated text and the original annotations.
Original Dataset
SCIERC-PT is derived from the SCIERC dataset:
- Paper: Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction (EMNLP 2018)
- Website: https://nlp.cs.washington.edu/sciIE/
Please refer to the original website and publication for further details regarding the original dataset construction and annotation methodology.
Dataset
The dataset contains computer science paper abstracts annotated with scientific entities and semantic relations.
Splits
Entity Types
Relation Types
Translation and Curation
The original SCIERC dataset was translated into European Portuguese (PT-PT) using GPT-5.1.
To improve annotation consistency, an automatic validation script was first used to identify candidate cases where annotated entity spans no longer matched the translated text. The flagged instances were then manually reviewed, correcting entity span alignment and, when necessary, making minimal adjustments to the translated text while preserving its original meaning.
The original annotation schema, including entity and relation types, was preserved throughout the process.
Although care was taken during curation, minor inconsistencies may still remain due to the inherent challenges of cross-lingual dataset adaptation.
Intended Use
SCIERC-PT is intended for research on:
- Scientific Information Extraction
- Named Entity Recognition
- Relation Extraction
- Scientific Knowledge Graph Construction
- Evaluation of multilingual and Portuguese Large Language Models
More Information
For more information about the dataset creation process and its application to Portuguese scientific information extraction, see our paper:
- Paper: Evaluating Generative Large Language Models for Portuguese Scientific Information Extraction (NSLP 2026)
- https://lrec.elra.info/lrec2026-ws-nslp-15
