CoolFace
Modelpublic

Polygl0t/portuguese-bertimbau-large-edu-classifier

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes15downloads
Model Card

BERTimbau-large Edu Classifier

BERTimbau-large Edu Classifier is a BERT based model that can be used for judging the educational value of a given Portuguese text string. This model was trained on the Portuguese Educational Qwen Annotations dataset.

Details

For training, we added a classification head with a single regression output to neuralmind/bert-large-portuguese-cased. Only the classification head was trained, i.e., the rest of the model was frozen.

  • Dataset: Portuguese Educational Qwen Annotations
  • Language: Portuguese
  • Number of Training Epochs: 20
  • Batch size: 256
  • Optimizer: torch.optim.AdamW
  • Learning Rate: 3e-4
  • Eval Metric: f1-score

This repository has the source code used to train this model.

Evaluation Results

Confusion Matrix
**1****2****3****4****5**
1565815353220
21080566484870
317120124602300
40355906272
5001101
  • Precision: 0.6368
  • Recall: 0.5482
  • F1 Macro: 0.5731
  • Accuracy: 0.7205

Usage

Here's an example of how to use the BERTimbau-large Edu Classifier:

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

tokenizer = AutoTokenizer.from_pretrained("Polygl0t/portuguese-bertimbau-large-edu-classifier")
model = AutoModelForSequenceClassification.from_pretrained("Polygl0t/portuguese-bertimbau-large-edu-classifier")
model.to(device)

text = "Coloque aqui o seu texto ..."
encoded_input  =  tokenizer(text, return_tensors="pt", padding="longest", truncation=True).to(device)

with  torch.no_grad():
	model_output  =  model(**encoded_input)
	logits  =  model_output.logits.squeeze(-1).float().cpu().numpy()

# scores are produced in the range [0, 4]. To convert to the range [1, 5], we can simply add 1 to the score.
score = [x + 1 for x in logits.tolist()][0]

print({
 "text": text,
 "score": score,
 "int_score": [int(round(max(0, min(score, 4)))) + 1 for score in logits][0],
})

Cite as 🤗

latex
@misc{correa2026tucano2cool,
      title={{Tucano 2 Cool: Better Open Source LLMs for Portuguese}}, 
      author={Nicholas Kluge Corr{\^e}a and Aniket Sen and Shiza Fatimah and Sophia Falk and Lennard Landgraf and Julia Kastner and Lucie Flek},
      year={2026},
      eprint={2603.03543},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2603.03543}, 
}

Aknowlegments

Polyglot is a project funded by the Federal Ministry of Education and Research (BMBF) and the Ministry of Culture and Science of the State of North Rhine-Westphalia (MWK) as part of TRA Sustainable Futures (University of Bonn) and the Excellence Strategy of the federal and state governments.

We also gratefully acknowledge the granted access to the Marvin cluster hosted by University of Bonn along with the support provided by its High Performance Computing & Analytics Lab.

License

BERTimbau-large Edu Classifier is licensed under the Apache License, Version 2.0. For more details, see the LICENSE file.