nikitast/lang-segmentation-roberta
415
RoBERTa for Multilabel Language Segmentation
Training
RoBERTa fine-tuned on small parts of Open Subtitles, Oscar and Tatoeba datasets (~9k samples per language).
Implemented heuristic algorithm for multilingual training data creation with generation of target masks- https://github.com/n1kstep/lang-classifier
Validation
The metrics obtained from validation on the another part of dataset (~1k samples per language).
