masakhaner
finetuned-xlm-r-masakhaner-swa-whole-word-phoneticxlm-roberta-base-masakhanerafroxlmr-large-ner-masakhaner-1.0_2.02-finetuned-xlm-r-masakhaner-swa-whole-word-phoneticbert-base-multilingual-cased-masakhanerdistilbert-base-multilingual-cased-masakhanerxlm-roberta-large-masakhanerpixel-base-finetuned-masakhaner-pcm
masakhaner2MasakhaNER 2.0 is the largest publicly available high-quality dataset for named entity recognition (NER) in 20 African languages.
Named entities are phrases that contain the names of persons, organizations, locations, times and quantities.
Example:
[PER Wolff] , currently a journalist in [LOC Argentina] , played with [PER Del Bosque] in the final years of the seventies in [ORG Real Madrid] .
MasakhaNER is a named entity dataset consisting of PER, ORG, LOC, and DATE entities annotated by Masakhane for 20 African languages:
- Bambara (bam)
- Ghomala (bbj)
- Ewe (ewe)
- Fon (fon)
- Hausa (hau)
- Igbo (ibo)
- Kinyarwanda (kin)
- Luganda (lug)
- Dholuo (luo)
- Mossi (mos)
- Chichewa (nya)
- Nigerian Pidgin
- chShona (sna)
- Kiswahili (swą)
- Setswana (tsn)
- Twi (twi)
- Wolof (wol)
- isiXhosa (xho)
- Yorùbá (yor)
- isiZulu (zul)
The train/validation/test sets are available for all the ten languages.
For more details see https://arxiv.org/abs/2103.11811masakhanerMasakhaNER is the first large publicly available high-quality dataset for named entity recognition (NER) in ten African languages.
Named entities are phrases that contain the names of persons, organizations, locations, times and quantities.
Example:
[PER Wolff] , currently a journalist in [LOC Argentina] , played with [PER Del Bosque] in the final years of the seventies in [ORG Real Madrid] .
MasakhaNER is a named entity dataset consisting of PER, ORG, LOC, and DATE entities annotated by Masakhane for ten African languages:
- Amharic
- Hausa
- Igbo
- Kinyarwanda
- Luganda
- Luo
- Nigerian-Pidgin
- Swahili
- Wolof
- Yoruba
The train/validation/test sets are available for all the ten languages.
For more details see https://arxiv.org/abs/2103.11811masakhanerV1masakhaner-x-parquetmasakhaner-xMasakhaNER-X is an aggregation of MasakhaNER 1.0 and MasakhaNER 2.0 datasets for 20 African languages. The dataset is not in CoNLL format. The input is the original raw text while the output is byte-level span annotations.
Example:
{"example_id": "test-00015916", "language": "pcm", "text": "By Bashir Ibrahim Hassan", "spans": [{"start_byte": 3, "limit_byte": 24, "label": "PER"}], "target": "PER: Bashir Ibrahim Hassan"}
MasakhaNER-X is a named entity dataset consisting of PER, ORG, LOC, and DATE entities annotated by Masakhane for twenty African languages:
- Amharic
- Ghomala
- Bambara
- Ewe
- Hausa
- Igbo
- Kinyarwanda
- Luganda
- Luo
- Mossi
- Chichewa
- Nigerian-Pidgin
- chiShona
- Swahili
- Setswana
- Twi
- Wolof
- Xhosa
- Yoruba
- Zulu
The train/validation/test sets are available for all the twenty languages.
For more details see https://aclanthology.org/2022.emnlp-main.298masakhaner2
MasakhaNER2
...
annotations_creators:
expert-generated
language:
bm
bbj
ee
fon
ha
ig
rw
lg
luo
mos
ny
pcm
sn
sw
tn
tw
wo
xh
yo
zu
language_creators:
expert-generated
license:
afl-3.0
multilinguality:
multilingual
pretty_name: masakhaner2.0
size_categories:
1K<n<10K
source_datasets:
original
tags:
ner
masakhaner
masakhane
task_categories:
token-classification
task_ids:
named-entity-recognition
Dataset Card for [Dataset Name]
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Vibrant8849/masakhaner2.
