CoolFace
Datasetpublic

amalia-llm/SCIERC-PT

SCIERC-PT Overview SCIERC-PT is a European Portuguese (PT-PT) translation of the SCIERC dataset, created to support research on Scientific Information Extraction in Portuguese. The dataset provides translated scientific abstracts suitable for evaluating Named Entity Recognition (NER) and Relation Extraction (RE) models while preserving the original annotation schema. The dataset was automatically translated and subsequently curated to improve alignment between the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SCIERC-PT.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes16downloads
Dataset Card

SCIERC-PT

Overview

SCIERC-PT is a European Portuguese (PT-PT) translation of the SCIERC dataset, created to support research on Scientific Information Extraction in Portuguese. The dataset provides translated scientific abstracts suitable for evaluating Named Entity Recognition (NER) and Relation Extraction (RE) models while preserving the original annotation schema.

The dataset was automatically translated and subsequently curated to improve alignment between the translated text and the original annotations.

Original Dataset

SCIERC-PT is derived from the SCIERC dataset:

  • —Paper: Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction (EMNLP 2018)
  • —Website: https://nlp.cs.washington.edu/sciIE/

Please refer to the original website and publication for further details regarding the original dataset construction and annotation methodology.

Dataset

The dataset contains computer science paper abstracts annotated with scientific entities and semantic relations.

Splits

SplitSamples
Train350
Dev50
Test100

Entity Types

EnglishPortuguese
TaskTarefa
MethodMétodo
Evaluation MetricMétrica de Avaliação
MaterialMaterial
Other Scientific TermOutros Termos Científicos
GenericGenérico

Relation Types

EnglishPortuguese
Used-forUsado-para
Feature-ofCaracterística-de
Hyponym-ofHipónimo-de
Part-ofParte-de
CompareCompara
ConjunctionConjunção
Evaluate-forAvaliado-para

Translation and Curation

The original SCIERC dataset was translated into European Portuguese (PT-PT) using GPT-5.1.

To improve annotation consistency, an automatic validation script was first used to identify candidate cases where annotated entity spans no longer matched the translated text. The flagged instances were then manually reviewed, correcting entity span alignment and, when necessary, making minimal adjustments to the translated text while preserving its original meaning.

The original annotation schema, including entity and relation types, was preserved throughout the process.

Although care was taken during curation, minor inconsistencies may still remain due to the inherent challenges of cross-lingual dataset adaptation.

Intended Use

SCIERC-PT is intended for research on:

  • —Scientific Information Extraction
  • —Named Entity Recognition
  • —Relation Extraction
  • —Scientific Knowledge Graph Construction
  • —Evaluation of multilingual and Portuguese Large Language Models

More Information

For more information about the dataset creation process and its application to Portuguese scientific information extraction, see our paper:

  • —Paper: Evaluating Generative Large Language Models for Portuguese Scientific Information Extraction (NSLP 2026)
  • —https://lrec.elra.info/lrec2026-ws-nslp-15