CoolFace
Modelpublic

gurkan08/bert-turkish-text-classification

sourceHugging Faceupdated 5y agoView on Hugging Face
3likes23downloads
Model Card

Turkish News Text Classification

Turkish text classification model obtained by fine-tuning the Turkish bert model (dbmdz/bert-base-turkish-cased)

Dataset

Dataset consists of 11 classes were obtained from https://www.trthaber.com/. The model was created using the most distinctive 6 classes.

Dataset can be accessed at https://github.com/gurkan08/datasets/tree/master/trt11category.

labeldict = { 'LABEL0': 'ekonomi', 'LABEL1': 'spor', 'LABEL2': 'saglik', 'LABEL3': 'kultursanat', 'LABEL4': 'bilimteknoloji', 'LABEL_5': 'egitim' }

70% of the data were used for training and 30% for testing.

train f1-weighted score = %97

test f1-weighted score = %94

Usage

from transformers import pipeline, AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.frompretrained("gurkan08/bert-turkish-text-classification") model = AutoModelForSequenceClassification.frompretrained("gurkan08/bert-turkish-text-classification")

nlp = pipeline("sentiment-analysis", model=model, tokenizer=tokenizer)

text = ["Süper Lig'in 6. haftasında Sivasspor ile Çaykur Rizespor karşı karşıya geldi...", "Son 24 saatte 69 kişi Kovid-19 nedeniyle yaşamını yitirdi, 1573 kişi iyileşti"]

out = nlp(text)

labeldict = { 'LABEL0': 'ekonomi', 'LABEL1': 'spor', 'LABEL2': 'saglik', 'LABEL3': 'kultursanat', 'LABEL4': 'bilimteknoloji', 'LABEL_5': 'egitim' }

results = [] for result in out: result['label'] = label_dict[result['label']] results.append(result) print(results)

# > [{'label': 'spor', 'score': 0.9992026090621948}, {'label': 'saglik', 'score': 0.9972177147865295}]