gplsi/Aitana-FraudDetection-R-1.0
Aitana-FraudDetection-R-1.0
Description
This model is fine-tuned from BSC-LT/mRoBERTa for binary classification of phishing detection in English texts. It predicts whether a given SMS or email message belongs to the category of phishing or not phishing.
Dataset
The dataset used for fine-tuning contains SMS and email texts labeled as phishing or not phishing.
- Training set: 9,422 instances
- Test set: 2,357 instances
Training Parameters
- learning_rate: 2e-5
- numtrainepochs: 2
- perdevicetrainbatchsize: 8
- perdeviceevalbatchsize: 8
- overwriteoutputdir: true
- logging_strategy: steps
- logging_steps: 10
- seed: 852
- fp16: true
Results
Combined dataset (SMS + emails)
Confusion Matrix
- Accuracy: 0.9856
- Macro Avg F1: 0.9798 ---
Only Emails
Confusion Matrix
- Accuracy: 0.9776
- Macro Avg F1: 0.9723 ---
Only SMS
Confusion Matrix | | Pred Not Phishing | Pred Phishing | | --------------------- | ----------------- | ------------- | | True Not Phishing | 969 | 5 | | True Phishing | 6 | 215 |
- Accuracy: 0.9908
- Macro Avg F1: 0.9847 ---
Funding
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública, co-financed by the EU – NextGenerationEU, within the framework of the project Desarrollo de Modelos ALIA.
Reference
@misc{gplsi-mroberta-fraudephishing,
author = {Martínez-Murillo, Iván and Consuegra-Ayala, Juan Pablo and Bonora, Mar and Sepúlveda-Torres, Robiert},
title = {Aitana-FraudDetection-R-1.0: Fine-tuned model for phishing detection},
year = {2025},
howpublished = {\url{https://huggingface.co/gplsi/Aitana-FraudDetection-R-1.0}},
note = {Accessed: 2025-10-03}
}
