CoolFace
Modelpublic

francesco-zatto/twitter-xlm-roberta-train-es-sexism-detector

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes6downloads
Model Card

XLM-RoBERTa Sexism Classifier (Trained: ES / Evaluated: EN)

This model is a fine-tuned version of cardiffnlp/twitter-xlm-roberta-base, trained for multi-class sexism detection on the EXIST 2023 Task 2 dataset.

Experiment Details: cross_lingual_transfer

This repository contains the Zero-Shot Cross-Lingual variant of our study.

  • —Training Data: The model was fine-tuned exclusively on the Spanish (ES) subset of the EXIST 2023 dataset.
  • —Evaluation Data: The model was evaluated exclusively on the English (EN) test set.
  • —Objective: This experiment measures the base model's zero-shot cross-lingual transfer capabilities—specifically, its ability to learn the conceptual nuances of sexism and reporting in Spanish and successfully map those concepts to English text without explicit English fine-tuning.
  • —Loss Function: Trained using a weighted Cross-Entropy loss to account for class imbalances in the Spanish training split.

Intended Use

Categorizes tweets into one of four sexist intentions:

  1. 1.- (Non-sexist)
  2. 2.DIRECT (Directly sexist messages)
  3. 3.JUDGEMENTAL (Messages condemning sexist behaviors)
  4. 4.REPORTED (Messages reporting a sexist situation)

Preprocessing

Inputs must be preprocessed to match the CardiffNLP base model formatting:

  • —Replace user mentions (@user) with the token @user
  • —Replace URLs with the token http

Evaluation Results (English Test Set)

  • —Macro F1: 0.4547
  • —Precision: 0.4596
  • —Recall: 0.4938

How to Use

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification

repo_id = "francesco-zatto/twitter-xlm-roberta-train-es-sexism-detector"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)

inputs = tokenizer("Your cleaned tweet text here", return_tensors="pt")
outputs = model(**inputs)