CoolFace
Datasetpublic

evalitahf/entity_recognition

Data for the NERMuD shared task (Evalita 2023) This data is the one used for the NERMuD shared task organized at Evalita 2023. The dataset contains the Wikinews, fiction, and De Gasperi subsets of KIND, where test data is used for development. Content of the dataset Split Sentences wn_train 10,912 wn_dev 2,594 wn_test 2,088 fic_train 11,423 fic_dev 1,051 fic_test 1,517 adg_train 5,147 adg_dev 1,122 adg_test 521 Set Sentences… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/entity_recognition.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
0likes105downloads
Dataset Card

Data for the NERMuD shared task (Evalita 2023)

This data is the one used for the NERMuD shared task organized at Evalita 2023.

The dataset contains the Wikinews, fiction, and De Gasperi subsets of KIND, where test data is used for development.

Content of the dataset

SplitSentences
wn_train10,912
wn_dev2,594
wn_test2,088
fic_train11,423
fic_dev1,051
fic_test1,517
adg_train5,147
adg_dev1,122
adg_test521
SetSentences
Total train27,482
Total dev4,767
Total test4,126
Total36,380

Publications

@inproceedings{evalita2023nermud,
    title={{NERMuD} at {EVALITA} 2023: Overview of the Named-Entities Recognition on Multi-Domain Documents Task},
    author={Palmero Aprosio, Alessio and Paccosi, Teresa},
    booktitle={Proceedings of the Eighth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2023)},
    publisher = {CEUR.org},
    year = {2023},
    month = {September},
    address = {Parma, Italy}
}