CoolFace
Datasetpublic

KoseiUemura/EmotionAnalysis

EmotionAnalysisPlus An MTEB dataset Massive Text Embedding Benchmark Multi-label emotion classification dataset for 28 languages released with the BRIGHTER project and SemEval-2025 Task 11. Task category t2c Domains Social, Written Reference https://github.com/emotion-analysis-project/SemEval2025-Task11 Source datasets: llama-lang-adapt/EmotionAnalysisFinal Dataset Preparation in MTEB This repository is a staging copy of… See the full description on the dataset page: https://huggingface.co/datasets/KoseiUemura/EmotionAnalysis.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes82downloads
Dataset Card

<!-- adapted from https://github.com/huggingface/huggingfacehub/blob/v0.30.2/src/huggingfacehub/templates/datasetcard_template.md -->

<div align="center" style="padding: 40px 20px; background-color: white; border-radius: 12px; box-shadow: 0 2px 10px rgba(0, 0, 0, 0.05); max-width: 600px; margin: 0 auto;"> <h1 style="font-size: 3.5rem; color: #1a1a1a; margin: 0 0 20px 0; letter-spacing: 2px; font-weight: 700;">EmotionAnalysisPlus</h1> <div style="font-size: 1.5rem; color: #4a4a4a; margin-bottom: 5px; font-weight: 300;">An <a href="https://github.com/embeddings-benchmark/mteb" style="color: #2c5282; font-weight: 600; text-decoration: none;" onmouseover="this.style.textDecoration='underline'" onmouseout="this.style.textDecoration='none'">MTEB</a> dataset</div> <div style="font-size: 0.9rem; color: #2c5282; margin-top: 10px;">Massive Text Embedding Benchmark</div> </div>

Multi-label emotion classification dataset for 28 languages released with the BRIGHTER project and SemEval-2025 Task 11.

Task categoryt2c
DomainsSocial, Written
Referencehttps://github.com/emotion-analysis-project/SemEval2025-Task11

Source datasets:

Dataset Preparation in MTEB

This repository is a staging copy of llama-lang-adapt/EmotionAnalysisFinal for MTEB. The intended long-term canonical benchmark copy is mteb/EmotionAnalysis.

Transformations

  • Converted per-emotion indicator columns into the MTEB multi-label format: label: list[int]
  • Preserved the task-facing subset names, including gaz and swh, while sourcing from the original Hub configs
  • Backfilled a train split where missing for benchmark compatibility before cleaning
  • Applied dataset cleaning before upload to reduce duplicates and copied-train overlap in the staging copy

Label Schema

  • 0: anger
  • 1: disgust
  • 2: fear
  • 3: joy
  • 4: sadness
  • 5: surprise

Splits and subsets

  • Language-specific configs from the benchmark task are preserved
  • The staged copy includes the cleaned train/validation/test-style splits used by EmotionAnalysisPlus

How to evaluate on this task

You can evaluate an embedding model on this dataset using the following code:

python
import mteb

task = mteb.get_task("EmotionAnalysisPlus")
evaluator = mteb.MTEB([task])

model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)

<!-- Datasets want link to arxiv in readme to autolink dataset with paper --> To learn more about how to run models on mteb task check out the GitHub repository.

Citation

If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a part of the MMTEB Contribution.

bibtex


@article{enevoldsen2025mmtebmassivemultilingualtext,
  title={MMTEB: Massive Multilingual Text Embedding Benchmark},
  author={Kenneth Enevoldsen and Isaac Chung and Imene Kerboua and Márton Kardos and Ashwin Mathur and David Stap and Jay Gala and Wissam Siblini and Dominik Krzemiński and Genta Indra Winata and Saba Sturua and Saiteja Utpala and Mathieu Ciancone and Marion Schaeffer and Gabriel Sequeira and Diganta Misra and Shreeya Dhakal and Jonathan Rystrøm and Roman Solomatin and Ömer Çağatan and Akash Kundu and Martin Bernstorff and Shitao Xiao and Akshita Sukhlecha and Bhavish Pahwa and Rafał Poświata and Kranthi Kiran GV and Shawon Ashraf and Daniel Auras and Björn Plüster and Jan Philipp Harries and Loïc Magne and Isabelle Mohr and Mariya Hendriksen and Dawei Zhu and Hippolyte Gisserot-Boukhlef and Tom Aarsen and Jan Kostkan and Konrad Wojtasik and Taemin Lee and Marek Šuppa and Crystina Zhang and Roberta Rocca and Mohammed Hamdy and Andrianos Michail and John Yang and Manuel Faysse and Aleksei Vatolin and Nandan Thakur and Manan Dey and Dipam Vasani and Pranjal Chitale and Simone Tedeschi and Nguyen Tai and Artem Snegirev and Michael Günther and Mengzhou Xia and Weijia Shi and Xing Han Lù and Jordan Clive and Gayatri Krishnakumar and Anna Maksimova and Silvan Wehrli and Maria Tikhonova and Henil Panchal and Aleksandr Abramov and Malte Ostendorff and Zheng Liu and Simon Clematide and Lester James Miranda and Alena Fenogenova and Guangyu Song and Ruqiya Bin Safi and Wen-Ding Li and Alessia Borghini and Federico Cassano and Hongjin Su and Jimmy Lin and Howard Yen and Lasse Hansen and Sara Hooker and Chenghao Xiao and Vaibhav Adlakha and Orion Weller and Siva Reddy and Niklas Muennighoff},
  publisher = {arXiv},
  journal={arXiv preprint arXiv:2502.13595},
  year={2025},
  url={https://arxiv.org/abs/2502.13595},
  doi = {10.48550/arXiv.2502.13595},
}

@article{muennighoff2022mteb,
  author = {Muennighoff, Niklas and Tazi, Nouamane and Magne, Loïc and Reimers, Nils},
  title = {MTEB: Massive Text Embedding Benchmark},
  publisher = {arXiv},
  journal={arXiv preprint arXiv:2210.07316},
  year = {2022}
  url = {https://arxiv.org/abs/2210.07316},
  doi = {10.48550/ARXIV.2210.07316},
}

Dataset Statistics

<details> <summary> Dataset Statistics</summary>

The following code contains the descriptive statistics from the task. These can also be obtained using:

python
import mteb

task = mteb.get_task("EmotionAnalysisPlus")

desc_stats = task.metadata.descriptive_stats
json
{
    "validation": {
        "num_samples": 9913,
        "number_texts_intersect_with_train": 9912,
        "text_statistics": {
            "total_text_length": 913579,
            "min_text_length": 6,
            "average_text_length": 92.15968929688287,
            "max_text_length": 2022,
            "unique_texts": 9912
        },
        "image_statistics": null,
        "audio_statistics": null,
        "label_statistics": {
            "min_labels_per_text": 0,
            "average_label_per_text": 0.9135478664380107,
            "max_labels_per_text": 5,
            "unique_labels": 7,
            "labels": {
                "3": {
                    "count": 1956
                },
                "None": {
                    "count": 2679
                },
                "2": {
                    "count": 721
                },
                "0": {
                    "count": 1509
                },
                "1": {
                    "count": 1584
                },
                "4": {
                    "count": 2091
                },
                "5": {
                    "count": 1195
                }
            }
        }
    },
    "test": {
        "num_samples": 43882,
        "number_texts_intersect_with_train": 17,
        "text_statistics": {
            "total_text_length": 4126105,
            "min_text_length": 3,
            "average_text_length": 94.02727769928444,
            "max_text_length": 3779,
            "unique_texts": 43848
        },
        "image_statistics": null,
        "audio_statistics": null,
        "label_statistics": {
            "min_labels_per_text": 0,
            "average_label_per_text": 1.0114169819060206,
            "max_labels_per_text": 6,
            "unique_labels": 7,
            "labels": {
                "3": {
                    "count": 9680
                },
                "None": {
                    "count": 10566
                },
                "4": {
                    "count": 9202
                },
                "2": {
                    "count": 4844
                },
                "0": {
                    "count": 7680
                },
                "1": {
                    "count": 6957
                },
                "5": {
                    "count": 6020
                }
            }
        }
    },
    "train": {
        "num_samples": 9913,
        "number_texts_intersect_with_train": null,
        "text_statistics": {
            "total_text_length": 913579,
            "min_text_length": 6,
            "average_text_length": 92.15968929688287,
            "max_text_length": 2022,
            "unique_texts": 9912
        },
        "image_statistics": null,
        "audio_statistics": null,
        "label_statistics": {
            "min_labels_per_text": 0,
            "average_label_per_text": 0.9135478664380107,
            "max_labels_per_text": 5,
            "unique_labels": 7,
            "labels": {
                "3": {
                    "count": 1956
                },
                "None": {
                    "count": 2679
                },
                "2": {
                    "count": 721
                },
                "0": {
                    "count": 1509
                },
                "1": {
                    "count": 1584
                },
                "4": {
                    "count": 2091
                },
                "5": {
                    "count": 1195
                }
            }
        }
    }
}

</details>


This dataset card was automatically generated using [MTEB](https://github.com/embeddings-benchmark/mteb)