CoolFace
Modelpublic

MHCTDS/visage

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes6downloads
Model Card

Model Card for Model ID

Best model between those tested on the VISAGE paper. Made for violence severity classification for social media messages in Brazilian Portuguese, according to an adaption of the Quantification of Violence Scale (QOVS) Improved with upsampling and further using CV K-Fold

Model Details

Has a total of 5 classes: No violence, low violence, medium violence, high violence and very high violence

Model Description

  • Developed by: Matheus Henrique Cajueiro Tobias de Souza
  • Funded by : Grant E-26/210.936/2024 by CNPq, CAPES, and FAPERJ
  • Model type: BERT
  • Language(s) (NLP): Brazilian Portuguese
  • License: Apache 2.0
  • Finetuned from model: BERTimbau

Model Sources [optional]

  • Repository: https://github.com/ufrrj-labweb/visage
  • New Repository: https://github.com/ufrrj-labweb/Visage-V1.1.
  • Paper: https://sol.sbc.org.br/index.php/brasnam/article/view/36366/36153

Direct Use

Predict violence severity of social media messages on Brazilian Portuguese based on the Quantification of Violence Scale (QOVS)

Downstream Use [optional]

General violence prediction or improved social media violence prediction for Portuguese and maybe Spanish

Out-of-Scope Use

Things not mentioned in the use cases, other form of written texts outside of social media without fine tuning.

Bias, Risks, and Limitations

Most of the data labels were no violence, high violence or very high violence, so the model tends to have difficulties with differentiating low violence from null and medium from high.

We trained it based on social media posts collected on X for major violent events happening especially in Rio de Janeiro. If the model is not fine tuned for your data and is used on a region with very different writing patterns or slang, the model might have difficulties.

Recommendations

Always fine tune it for your data, the model is pretty light so it should always be feasible.

If you don't want to accidentally overtune it set a low learning rate, but that shouldn't be an issue with proper fine tuning techniques.

How to Get Started with the Model

Use the code below to get started with the model.

model = BertForSequenceClassification.from_pretrained("MHCTDS/visage")

The github includes a BERT class for model evaluation and training using CV K-Fold, but be warned that it is not optimal. I recommend just using your own functions like any other BERT model

Accelerate can be used to parallelize it on any computer, including multiple GPUs or NPU setups (I don't know why you would need multiple GPUs for a 0.5 Gigabyte model though).

Training Details

Training Data

Trained on social media posts collected on X for major violent events happening especially in Rio de Janeiro.

Over 2k posts were labeled by 2 students based on the Quantification of Violence Scale (QOVS).

1758 of those entries (non-duplicates) were used on a 70/30 training-validation split.

For more information consult the paper or github. We plan on releasing the data on a dataset paper in the future.

Training Procedure

After evaluation using 10 folds with CV K-Fold, the model was retrained using all of the data available with upsampling.

Preprocessing [optional]

The data was upsampled, with the aim of having all classes distributed in a roughly equal manner. Low violence class ws multiplied by 240 times, medium violence class by 71 times, high violence class by 13 times and very high violence by 3 times

Training Hyperparameters
  • Training regime: [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
Speeds, Sizes, Times

20k+ entries per minute with a rtx 5070 ti

Only 1 size available for now. Conversion and use on mobile devices should is possible.

Evaluation

Evaluation class used is available on the github.

Testing Data, Factors & Metrics

Testing Data

<!-- This should link to a Dataset Card if possible. -->

[More Information Needed]

Factors

<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->

[More Information Needed]

Metrics

F1, accuracy, recall and precision were used to evaluate to model over a 10 fold CV K-Fold process along with ROC AUC, AP and their standard deviations.

Results

The model was evaluated witha 10 fold CV K-Fold process, the following metrics are averages for all the folds.

The current upsampled model had 88% recall, 87% F1, 86% Precision, 88% Accuracy, 92% ROC AUC and 78% AP, 0.01 of standard deviation per fold for ROC AUC and 0.03 of standard deviation for AP

The past non-upsampled model had 88% recall, 86% F1, 84% Precision, 88% Accuracy, 92% ROC AUC and 78% AP, 0.01 of standard deviation per fold for ROC AUC and 0.03 of standard deviation for AP

For comparison, on our dataset, the results with a dummy model were: 68% recall, 55% F1, 46% Precision, 68% Accuracy, 80% ROC AUC and 53% APS

Summary

Model Examination [optional]

<!-- Relevant interpretability work for the model goes here -->

[More Information Needed]

Environmental Impact

<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

  • Hardware Type: Mac studio m2 max
  • Hours used: 100
  • Carbon Emitted: ~0.3kg | A negligible amount compared to a dishwasher in a year

Technical Specifications

Hardware

Needs at least 0.5G of ram

Software

Can run on any OS. Was tested on mps(Mac) and CUDA(Nvidia).

Citation

BibTeX:

@online{noauthorvistanodate, title = {Vista do {VISAGE}: Detection and automatic classification of urban violence through social media data}, url = {https://sol.sbc.org.br/index.php/brasnam/article/view/36366/36153}, urldate = {2025-08-03}, file = {Vista do VISAGE\: Detection and automatic classification of urban violence through social media data:/Users/mhctds/Zotero/storage/2KE57U9X/36153.html:text/html}, }

@inproceedings{souzavisage2025, title = {{VISAGE}: Detection and automatic classification of urban violence through social media data}, rights = {Copyright (c)}, url = {https://sol.sbc.org.br/index.php/brasnam/article/view/36366}, doi = {10.5753/brasnam.2025.8111}, shorttitle = {{VISAGE}}, abstract = {Urban violence remains a challenge for cities, requiring innovative approaches to provide actionable insights for authorities. This study presents {VISAGE}, a framework designed to detect and classify urban violence in social media textual data using an extension of the Quantification of Violence Scale ({QoVS}). Ten Machine learning models, including {BERT} and Random Forest, were evaluated on Twitter datasets from violent events in Brazil, with {BERT} achieving an F1-score of 0.86. Results demonstrate the feasibility of automating violence assessment from text. Limitations like dataset imbalance and labeling process are discussed with future work targeting real-time and multimodal analysis.}, eventtitle = {Brazilian Workshop on Social Network Analysis and Mining ({BraSNAM})}, pages = {26--39}, booktitle = {Brazilian Workshop on Social Network Analysis and Mining ({BraSNAM})}, publisher = {{SBC}}, author = {Souza, Matheus Henrique C. T. de and Silva, Eliel Roger da and França, Tiago Cruz de and Oliveira, Jonice}, urldate = {2025-08-03}, date = {2025-07-20}, langid = {english}, note = {{ISSN}: 2595-6094}, file = {Full Text PDF:/Users/mhctds/Zotero/storage/XNAUFBFX/Souza et al. - 2025 - VISAGE Detection and automatic classification of urban violence through social media data.pdf:application/pdf}, }

APA:

[More Information Needed]

Model Card Authors [optional]

Matheus Henrique Cajueiro Tobias de Souza, Eliel Roger da Silva, Tiago Cruz de França, Jonice Oliveira

Model Card Contact

1MHCTDS8@gmail.com, elielsilva@ufrj.br, tcruzfanca@ufrrj.br, jonice@dcc.ufrj.br