CoolFace
Modelpublic

nikitast/lang-segmentation-roberta

sourceHugging Faceupdated 4y agoView on Hugging Face
4likes15downloads
Model Card

RoBERTa for Multilabel Language Segmentation

Training

RoBERTa fine-tuned on small parts of Open Subtitles, Oscar and Tatoeba datasets (~9k samples per language).

Implemented heuristic algorithm for multilingual training data creation with generation of target masks- https://github.com/n1kstep/lang-classifier

data sourcelanguage
open_subtitleska, he, en, de
oscarbe, kk, az, hu
tatoebaru, uk

Validation

The metrics obtained from validation on the another part of dataset (~1k samples per language).

Validation LossPrecisionRecallF1-ScoreAccuracy
0.0291720.9196230.9335860.9265520.991883