seara/rubert-tiny2-russian-sentiment
3527k
This is RuBERT-tiny2 model fine-tuned for _sentiment classification of short Russian texts. The task is a multi-class classification_ with the following labels:
0: neutral
1: positive
2: negativeLabel to Russian label:
neutral: нейтральный
positive: позитивный
negative: негативныйUsage
from transformers import pipeline
model = pipeline(model="seara/rubert-tiny2-russian-sentiment")
model("Привет, ты мне нравишься!")
# [{'label': 'positive', 'score': 0.9398769736289978}]Dataset
This model was trained on the union of the following datasets:
- Kaggle Russian News Dataset
- Linis Crowd 2015
- Linis Crowd 2016
- RuReviews
- RuSentiment
An overview of the training data can be found on S. Smetanin Github repository.
_Download links for all Russian sentiment datasets collected by Smetanin can be found in this [repository](https://github.com/searayeah/russian-sentiment-emotion-datasets)._
Training
Training were done in this project with this parameters:
tokenizer.max_length: 512
batch_size: 64
optimizer: adam
lr: 0.00001
weight_decay: 0
epochs: 5Train/validation/test splits are 80%/10%/10%.
