CoolFace
Modelpublic

megantosh/flair-arabic-dialects-codeswitch-egy-lev

sourceHugging Faceapache-2.0updated 5y agoView on Hugging Face
0likes39downloads
Model Card

Arabic Flair + fastText Part-of-Speech tagging Model (Egyptian and Levant)

Pretrained Part-of-Speech tagging model built on a joint corpus written in Egyptian and Levantine (Jordanian, Lebanese, Palestinian, Syrian) dialects with code-switching of Egyptian Arabic and English. The model is trained using Flair (forward+backward)and fastText embeddings.

Pretraining Corpora:

This sequence labeling model was pretrained on three corpora jointly:

  1. 1.4 Dialects A Dialectal Arabic Datasets containing four dialects of Arabic, Egyptian (EGY), Levantine (LEV), Gulf (GLF), and Maghrebi (MGR). Each dataset consists of a set of 350 manually segmented and PoS tagged tweets.
  2. 2.UD South Levantine Arabic MADAR A Dataset with 100 manually-annotated sentences taken from the MADAR (Multi-Arabic Dialect Applications and Resources) project by Shorouq Zahra.
  3. 3.Parts of the Cairo Students Code-Switch (CSCS) corpus developed for "Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus" by Hamed et al.

Usage

python
from flair.data import Sentence
from flair.models import SequenceTagger
  
tagger = SequenceTagger.load("megantosh/flair-arabic-dialects-codeswitch-egy-lev")
sentence = Sentence('عمرو عادلي أستاذ للاقتصاد السياسي المساعد في الجامعة الأمريكية  بالقاهرة .')
tagger.predict(sentence)
for entity in sentence.get_spans('pos'):
    print(entity)

Due to the right-to-left in left-to-right context, some formatting errors might occur. and your code might appear like this, (link accessed on 2020-10-27)

<!--# Example

Tagset-->

Scores & Tagset

<details>

precisionrecallf1-scoresupport
INTJ0.81820.90000.857110
OUN0.90090.94020.9201435
NUM0.95240.83330.888924
ADJ0.87620.76030.8142121
ADP0.99030.96230.9761106
CCONJ0.96000.97300.966474
PROPN0.93330.93330.933315
ADV0.91350.80510.8559118
VERB0.88520.92310.9038117
PRON0.96200.94650.9542187
SCONJ0.85710.94740.900019
PART0.93500.97910.9565191
DET0.93480.91490.924747
PUNCT1.00001.00001.000035
AUX0.92860.98110.954153
MENTION0.92311.00000.960012
V0.85710.87800.867582
FUT-PART+V+PREP+PRON1.00000.00000.00001
PROG-PART+V+PRON+PREP+PRON0.00001.00000.00000
ADJ+NSUFF0.61110.84620.709726
NOUN+NSUFF0.81820.84380.830864
PREP+PRON0.95650.95650.956523
PUNC0.99411.00000.9971169
EOS1.00001.00001.000070
NOUN+PRON0.69860.85000.766960
V+PRON0.72580.80360.762756
PART+PRON1.00000.94740.973019
PROG-PART+V0.83330.93020.879143
DET+NOUN0.96251.00000.980977
NOUN+NSUFF+PRON0.90910.71430.800014
PROG-PART+V+PRON0.70830.94440.809518
PREP+NOUN+NSUFF0.66670.40000.5000 5
NOUN+NSUFF+NSUFF1.00000.00000.00003
CONJ0.97221.00000.985935
V+PRON+PRON0.63640.58330.608712
FOREIGN0.66670.66670.66673
PREP+NOUN0.63160.75000.685716
DET+NOUN+NSUFF0.90000.93100.915329
DET+ADJ+NSUFF1.00000.57140.72737
CONJ+PRON1.00000.87500.93338
NOUN+CASE0.00000.00000.00002
DET+ADJ1.00000.66670.80006
PREP1.00000.97180.985771
CONJ+FUT-PART+V0.00000.00000.00001
CONJ+V0.66670.75000.70598
FUT-PART1.00001.00001.00002
ADJ+PRON1.00000.00000.00008
CONJ+PREP+NOUN+PRON1.00000.00000.00001
CONJ+NOUN+PRON0.37501.00000.54553
PART+ADJ1.00000.00000.00001
PART+NOUN0.50001.00000.66671
CONJ+PREP+NOUN1.00000.00000.00001
CONJ+NOUN0.70000.77780.73689
URL1.00001.00001.00003
CONJ+FUT-PART1.00000.00000.00001
FUT-PART+V0.85710.60000.705910
PREP+NOUN+NSUFF+NSUFF1.00000.00000.00001
HASH1.00000.94120.969717
ADJ+PREP+PRON1.00000.00000.00003
PREP+NOUN+PRON0.00000.00000.00001
EMOT1.00000.88890.941218
CONJ+PREP1.00000.75000.85714
PREP+DET+NOUN+NSUFF1.00000.75000.85714
PRON+DET+NOUN+NSUFF0.00001.00000.00000
V+PREP+PRON1.00000.00000.00005
V+PRON+PREP+PRON0.00001.00000.00000
CONJ+NOUN+NSUFF0.50000.50000.50002
V+NEG-PART1.00000.00000.00002
PREP+DET+NOUN0.90911.00000.952410
PREP+V1.00000.00000.00002
CONJ+PART1.00000.77780.87509
CONJ+V+PRON1.00001.00001.00005
PROG-PART+V+PREP+PRON1.00000.50000.66672
PREP+NOUN+NSUFF+PRON1.00001.00001.00001
ADJ+CASE1.00000.00000.00001
PART+NOUN+PRON1.00001.00001.00001
PART+V1.00000.00000.00003
PART+V+PRON0.00001.00000.00000
FUT-PART+V+PRON0.00001.00000.00000
FUT-PART+V+PRON+PRON1.00000.00000.00001
CONJ+PREP+PRON1.00000.00000.00001
CONJ+V+PRON+PREP+PRON1.00000.00000.00001
CONJ+V+PREP+PRON0.00001.00000.00000
CONJ+DET+NOUN+NSUFF1.00000.00000.00001
CONJ+DET+NOUN0.66671.00000.80002
CONJ+PREP+DET+NOUN1.00001.00001.00001
PREP+PART1.00000.00000.00002
PART+V+PRON+NEG-PART0.33330.33330.33333
PART+V+NEG-PART0.33330.50000.40002
PART+PREP+NEG-PART1.00001.00001.00003
PART+PROG-PART+V+NEG-PART1.00000.33330.50003
PREP+DET+NOUN+NSUFF+PREP+PRON1.00000.00000.00001
PREP+PRON+DET+NOUN0.00001.00000.00000
PART+NSUFF1.00000.00000.00001
CONJ+PROG-PART+V+PRON1.00001.00001.00001
PART+PREP+PRON1.00000.00000.00001
CONJ+PART+PREP1.00000.00000.00001
NUM+NSUFF0.66670.66670.66673
CONJ+PART+V+PRON+NEG-PART1.00001.00001.00001
PART+NOUN+NEG-PART1.00001.00001.00001
CONJ+ADJ+NSUFF1.00000.00000.00001
PREP+ADJ1.00000.00000.00001
ADJ+NSUFF+PRON1.00000.00000.00002
CONJ+PROG-PART+V1.00000.00000.00001
CONJ+PART+PROG-PART+V+PREP+PRON+NEG-PART1.00000.00000.00001
CONJ+PART+PREP+PRON+NEG-PART0.00001.00000.00000
PREP+PART+PRON1.00000.00000.00001
CONJ+ADV+NSUFF1.00000.00000.00001
CONJ+ADV0.00001.00000.00000
PART+NOUN+PRON+NEG-PART0.00001.00000.00000
CONJ+ADJ1.00001.00001.00001

</details>

  • F-score (micro): 0.8974
  • F-score (macro): 0.5188
  • Accuracy (incl. no class): 0.901

Expand details below to show class scores for each tag. Note that tag compounds (a tag made for multiple agglutinated parts of speech) are considered as separate ones.

# Citation if you use this model, please consider citing [this work](https://www.researchgate.net/publication/358956953_Sequence_Labeling_Architectures_in_Diglossia_-_a_case_study_of_Arabic_and_its_dialects):

latex
@unpublished{MMHU21
author = "M. Megahed",
title = "Sequence Labeling Architectures in Diglossia",
year = {2021},
doi = "10.13140/RG.2.2.34961.10084"
url = {https://www.researchgate.net/publication/358956953_Sequence_Labeling_Architectures_in_Diglossia_-_a_case_study_of_Arabic_and_its_dialects}
}