CoolFace
Datasetpublic

proxectonos/Galician_NER

Galician NER test Dataset created by combining four galician datasets for Named Entity Recognition, annotated according to the new standards for NER annotations: corNER: Updated version of the original corNER dataset, which was created by annotating for NER the corga dataset. LREC: Updated version to keep up with the new standards for NER annotations. PUD: Dataset created by annotating for NER the Galician PUD treebank. TreeGal: Dataset created by annotating for NER the TreeGal… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Galician_NER.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes50downloads
Dataset Card

Galician NER test

Dataset created by combining four galician datasets for Named Entity Recognition, annotated according to the new standards for NER annotations:

  • corNER: Updated version of the original corNER dataset, which was created by annotating for NER the corga dataset.
  • LREC: Updated version to keep up with the new standards for NER annotations.
  • PUD: Dataset created by annotating for NER the Galician PUD treebank.
  • TreeGal: Dataset created by annotating for NER the TreeGal treebank.

All datasets where manually annotated from random extractions of journalistic galician corpus. The labels identify their corresponding tokens as:

  • proper nouns in an initial position (B)
  • an internal position (I)
  • another grammatical element (O).

Proper nouns are classified according to the enamex standard notation:

  • PER (person)
  • ORG (organization)
  • LOC (location)
  • MISC (other).

Acknowledgments

This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA. Esta publicación del proyecto Desarrollo de Modelos ALIA está financiada por el Ministerio para la Transformación Digital y de la Función Pública y por el Plan de Recuperación, Transformación y Resiliencia – Financiado por la Unión Europea – NextGenerationEU.

Citations:

Marcos Garcia, Pablo Gamallo. Multilingual corpora with coreferential annotation of person entities.

Marcos Garcia, Iria Gayo, Isaac González López. Identificação e classificação de entidades mencionadas em galego.

Xulia Sánchez-Rodríguez, Albina Sarymsakova, Laura Castro, Marcos Garcia. Increasing manually annotated resources for Galician: the Parallel Universal Dependencies Treebank.

Marcos Garcia, Carlos Gómez-Rodríguez, and Miguel A Alonso. New treebank or repurposed? On the feasibility of cross-lingual parsing of romance languages with universal dependencies.