mteb/BrazilianToxicTweetsClassification
BrazilianToxicTweetsClassification An MTEB dataset Massive Text Embedding Benchmark ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity, sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia, Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification.
16.2k
1---2annotations_creators:3- expert-annotated4language:5- por6license: cc-by-sa-4.07multilinguality: monolingual8source_datasets:9- mteb/BrazilianToxicTweetsClassification10task_categories:11- text-classification12- sentiment-analysis13- sentiment-scoring14- sentiment-classification15- hate-speech-detection16task_ids:17- sentiment-analysis18- sentiment-scoring19- sentiment-classification20- hate-speech-detection21dataset_info:22 features:23 - name: text24 dtype: string25 - name: label26 list: string27 splits:28 - name: train29 num_bytes: 85696830 num_examples: 819231 - name: test32 num_bytes: 20791933 num_examples: 204834 download_size: 67650235 dataset_size: 106488736configs:37- config_name: default38 data_files:39 - split: train40 path: data/train-*41 - split: test42 path: data/test-*43tags:44- mteb45- text46---47<!-- adapted from https://github.com/huggingface/huggingface_hub/blob/v0.30.2/src/huggingface_hub/templates/datasetcard_template.md -->48 49<div align="center" style="padding: 40px 20px; background-color: white; border-radius: 12px; box-shadow: 0 2px 10px rgba(0, 0, 0, 0.05); max-width: 600px; margin: 0 auto;">50 <h1 style="font-size: 3.5rem; color: #1a1a1a; margin: 0 0 20px 0; letter-spacing: 2px; font-weight: 700;">BrazilianToxicTweetsClassification</h1>51 <div style="font-size: 1.5rem; color: #4a4a4a; margin-bottom: 5px; font-weight: 300;">An <a href="https://github.com/embeddings-benchmark/mteb" style="color: #2c5282; font-weight: 600; text-decoration: none;" onmouseover="this.style.textDecoration='underline'" onmouseout="this.style.textDecoration='none'">MTEB</a> dataset</div>52 <div style="font-size: 0.9rem; color: #2c5282; margin-top: 10px;">Massive Text Embedding Benchmark</div>53</div>54 55 56 ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from57 a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity,58 sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia,59 Xenophobia, Obscene, Insult, Misogyny and Racism.60 61 62| | |63|---------------|---------------------------------------------|64| Task category | t2c |65| Domains | Constructed, Written |66| Reference | https://paperswithcode.com/dataset/told-br |67 68Source datasets:69- [mteb/BrazilianToxicTweetsClassification](https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification)70 71 72## How to evaluate on this task73 74You can evaluate an embedding model on this dataset using the following code:75 76```python77import mteb78 79task = mteb.get_task("BrazilianToxicTweetsClassification")80evaluator = mteb.MTEB([task])81 82model = mteb.get_model(YOUR_MODEL)83evaluator.run(model)84```85 86<!-- Datasets want link to arxiv in readme to autolink dataset with paper -->87To learn more about how to run models on `mteb` task check out the [GitHub repository](https://github.com/embeddings-benchmark/mteb).88 89## Citation90 91If you use this dataset, please cite the dataset as well as [mteb](https://github.com/embeddings-benchmark/mteb), as this dataset likely includes additional processing as a part of the [MMTEB Contribution](https://github.com/embeddings-benchmark/mteb/tree/main/docs/mmteb).92 93```bibtex94 95@article{DBLP:journals/corr/abs-2010-04543,96 author = {Joao Augusto Leite and97Diego F. Silva and98Kalina Bontcheva and99Carolina Scarton},100 eprint = {2010.04543},101 eprinttype = {arXiv},102 journal = {CoRR},103 timestamp = {Tue, 15 Dec 2020 16:10:16 +0100},104 title = {Toxic Language Detection in Social Media for Brazilian Portuguese:105New Dataset and Multilingual Analysis},106 url = {https://arxiv.org/abs/2010.04543},107 volume = {abs/2010.04543},108 year = {2020},109}110 111 112@article{enevoldsen2025mmtebmassivemultilingualtext,113 title={MMTEB: Massive Multilingual Text Embedding Benchmark},114 author={Kenneth Enevoldsen and Isaac Chung and Imene Kerboua and Márton Kardos and Ashwin Mathur and David Stap and Jay Gala and Wissam Siblini and Dominik Krzemiński and Genta Indra Winata and Saba Sturua and Saiteja Utpala and Mathieu Ciancone and Marion Schaeffer and Gabriel Sequeira and Diganta Misra and Shreeya Dhakal and Jonathan Rystrøm and Roman Solomatin and Ömer Çağatan and Akash Kundu and Martin Bernstorff and Shitao Xiao and Akshita Sukhlecha and Bhavish Pahwa and Rafał Poświata and Kranthi Kiran GV and Shawon Ashraf and Daniel Auras and Björn Plüster and Jan Philipp Harries and Loïc Magne and Isabelle Mohr and Mariya Hendriksen and Dawei Zhu and Hippolyte Gisserot-Boukhlef and Tom Aarsen and Jan Kostkan and Konrad Wojtasik and Taemin Lee and Marek Šuppa and Crystina Zhang and Roberta Rocca and Mohammed Hamdy and Andrianos Michail and John Yang and Manuel Faysse and Aleksei Vatolin and Nandan Thakur and Manan Dey and Dipam Vasani and Pranjal Chitale and Simone Tedeschi and Nguyen Tai and Artem Snegirev and Michael Günther and Mengzhou Xia and Weijia Shi and Xing Han Lù and Jordan Clive and Gayatri Krishnakumar and Anna Maksimova and Silvan Wehrli and Maria Tikhonova and Henil Panchal and Aleksandr Abramov and Malte Ostendorff and Zheng Liu and Simon Clematide and Lester James Miranda and Alena Fenogenova and Guangyu Song and Ruqiya Bin Safi and Wen-Ding Li and Alessia Borghini and Federico Cassano and Hongjin Su and Jimmy Lin and Howard Yen and Lasse Hansen and Sara Hooker and Chenghao Xiao and Vaibhav Adlakha and Orion Weller and Siva Reddy and Niklas Muennighoff},115 publisher = {arXiv},116 journal={arXiv preprint arXiv:2502.13595},117 year={2025},118 url={https://arxiv.org/abs/2502.13595},119 doi = {10.48550/arXiv.2502.13595},120}121 122@article{muennighoff2022mteb,123 author = {Muennighoff, Niklas and Tazi, Nouamane and Magne, Loïc and Reimers, Nils},124 title = {MTEB: Massive Text Embedding Benchmark},125 publisher = {arXiv},126 journal={arXiv preprint arXiv:2210.07316},127 year = {2022}128 url = {https://arxiv.org/abs/2210.07316},129 doi = {10.48550/ARXIV.2210.07316},130}131```132 133# Dataset Statistics134<details>135 <summary> Dataset Statistics</summary>136 137The following code contains the descriptive statistics from the task. These can also be obtained using:138 139```python140import mteb141 142task = mteb.get_task("BrazilianToxicTweetsClassification")143 144desc_stats = task.metadata.descriptive_stats145```146 147```json148{149 "test": {150 "num_samples": 2048,151 "number_texts_intersect_with_train": 23,152 "text_statistics": {153 "total_text_length": 172708,154 "min_text_length": 5,155 "average_text_length": 84.330078125,156 "max_text_length": 304,157 "unique_texts": 2046158 },159 "image_statistics": null,160 "label_statistics": {161 "min_labels_per_text": 0,162 "average_label_per_text": 0.57958984375,163 "max_labels_per_text": 4,164 "unique_labels": 7,165 "labels": {166 "obscene": {167 "count": 653168 },169 "insult": {170 "count": 430171 },172 "misogyny": {173 "count": 46174 },175 "racism": {176 "count": 13177 },178 "xenophobia": {179 "count": 13180 },181 "homophobia": {182 "count": 32183 },184 "None": {185 "count": 1145186 }187 }188 }189 },190 "train": {191 "num_samples": 8192,192 "number_texts_intersect_with_train": null,193 "text_statistics": {194 "total_text_length": 714281,195 "min_text_length": 4,196 "average_text_length": 87.1925048828125,197 "max_text_length": 322,198 "unique_texts": 8172199 },200 "image_statistics": null,201 "label_statistics": {202 "min_labels_per_text": 0,203 "average_label_per_text": 0.5751953125,204 "max_labels_per_text": 4,205 "unique_labels": 7,206 "labels": {207 "None": {208 "count": 4580209 },210 "obscene": {211 "count": 2576212 },213 "insult": {214 "count": 1700215 },216 "homophobia": {217 "count": 139218 },219 "misogyny": {220 "count": 179221 },222 "racism": {223 "count": 54224 },225 "xenophobia": {226 "count": 64227 }228 }229 }230 }231}232```233 234</details>235 236---237*This dataset card was automatically generated using [MTEB](https://github.com/embeddings-benchmark/mteb)*