datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
named_entity_recognition_document_contextflan_combined_task1544_conll2002_named_entity_recognition_answer_generationtask960_ancora-ca-ner_named_entity_recognition
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task960_ancora-ca-ner_named_entity_recognition
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task960_ancora-ca-ner_named_entity_recognition.amharic-named-entity-recognition
Amharic Named Entity Recognition Dataset
This dataset can be used to train models for Named Entity Recognition.
Dataset Source
https://github.com/uhh-lt/ethiopicmodels/blob/master/am/data/NER/train.txt
Finetuned Models
The following transformer models were finetuned using this dataset. The reported precision, recall, and f1 metrics are macro averages.
Model
Size (# params)
Precision
Recall
F1
bert-medium-amharic
40.5M
0.64
0.73
0.68… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-named-entity-recognition.named-entity-recognition
Sinhala Named Entity Recognition
Sinhala Named Entity Recognition is a token-level named entity recognition dataset for Sinhala.
This repository is a re-upload of the original Sinhala NER dataset introduced by Manamini et al. (2016) in "Ananya - a Named-Entity-Recognition (NER) System for Sinhala Language" with proper train/ test splits. The dataset was subsequently included as the Named Entity Recognition (NER) task in the SINHALA-GLUE benchmark introduced in "Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/named-entity-recognition.named_entity_recognitionnlp.6.named_entity_recognition
Dataset Card for "nlp.6.named_entity_recognition"
More Information needed
APIS_OEBL__Named_Entity_RecognitionJSON file of 6,941 sentences of historical biographies, annotated with "PER" (Person), "ORG" (Organisation), "LOC" (Location).
source
The original data was extracted from the Austrian Biographical Lexicon (ÖBL) in the context of the Austrian Prosopographical Information System (APIS) project.
From there, samples were randomly pulled and annotated for Named Entity Recognition tasks, which form this dataset.
The texts concern numerous smaller biographies in the time period between… See the full description on the dataset page: https://huggingface.co/datasets/SteffRhes/APIS_OEBL__Named_Entity_Recognition.Named_entity_recognitionnamed_entity_recognitiontask1544_conll2002_named_entity_recognition_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1544_conll2002_named_entity_recognition_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1544_conll2002_named_entity_recognition_answer_generation.NamedEntityRecognition_SLUE-VoxPopulinamed_entity_recognitionflan_combined_task960_ancora-ca-ner_named_entity_recognition
