named-entity
named_entity_recognition_document_contexttask960_ancora-ca-ner_named_entity_recognition
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task960_ancora-ca-ner_named_entity_recognition
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task960_ancora-ca-ner_named_entity_recognition.cs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.flan_combined_task1544_conll2002_named_entity_recognition_answer_generationpioNER-Armenian-Named-Entity
pioNER - named entity annotated datasets
pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
Alongside the datasets, we release 50-, 100-, 200-, and 300-dimensional GloVe word embeddings trained on a collection of Armenian texts from Wikipedia, news, blogs, and encyclopedia.
Silver-standard dataset
The generated corpus is automatically extracted and annotated using Armenian Wikipedia. We used a modification of… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/pioNER-Armenian-Named-Entity.amharic-named-entity-recognition
Amharic Named Entity Recognition Dataset
This dataset can be used to train models for Named Entity Recognition.
Dataset Source
https://github.com/uhh-lt/ethiopicmodels/blob/master/am/data/NER/train.txt
Finetuned Models
The following transformer models were finetuned using this dataset. The reported precision, recall, and f1 metrics are macro averages.
Model
Size (# params)
Precision
Recall
F1
bert-medium-amharic
40.5M
0.64
0.73
0.68… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-named-entity-recognition.
