francesco-zatto/twitter-xlm-roberta-train-es-sexism-detector
06
XLM-RoBERTa Sexism Classifier (Trained: ES / Evaluated: EN)
This model is a fine-tuned version of cardiffnlp/twitter-xlm-roberta-base, trained for multi-class sexism detection on the EXIST 2023 Task 2 dataset.
Experiment Details: cross_lingual_transfer
This repository contains the Zero-Shot Cross-Lingual variant of our study.
- Training Data: The model was fine-tuned exclusively on the Spanish (ES) subset of the EXIST 2023 dataset.
- Evaluation Data: The model was evaluated exclusively on the English (EN) test set.
- Objective: This experiment measures the base model's zero-shot cross-lingual transfer capabilities—specifically, its ability to learn the conceptual nuances of sexism and reporting in Spanish and successfully map those concepts to English text without explicit English fine-tuning.
- Loss Function: Trained using a weighted Cross-Entropy loss to account for class imbalances in the Spanish training split.
Intended Use
Categorizes tweets into one of four sexist intentions:
-(Non-sexist)DIRECT(Directly sexist messages)JUDGEMENTAL(Messages condemning sexist behaviors)REPORTED(Messages reporting a sexist situation)
Preprocessing
Inputs must be preprocessed to match the CardiffNLP base model formatting:
- Replace user mentions (
@user) with the token@user - Replace URLs with the token
http
Evaluation Results (English Test Set)
- Macro F1: 0.4547
- Precision: 0.4596
- Recall: 0.4938
How to Use
from transformers import AutoTokenizer, AutoModelForSequenceClassification
repo_id = "francesco-zatto/twitter-xlm-roberta-train-es-sexism-detector"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)
inputs = tokenizer("Your cleaned tweet text here", return_tensors="pt")
outputs = model(**inputs)