CoolFace
Datasetpublic

vgaraujov/semeval-2025-task11-track-c

SemEval 2025 Task 11 - Track C Dataset This dataset contains the data for SemEval 2025 Task 11: Bridging the Gap in Text-Based Emotion Detection - Track C, organized as language-specific configurations. Dataset Description The dataset is a multi-language, multi-label emotion classification dataset with separate configurations for each language. Total languages: 30 standard ISO codes Total examples: 57254 Splits: dev, test (Track C has no train split)… See the full description on the dataset page: https://huggingface.co/datasets/vgaraujov/semeval-2025-task11-track-c.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes342downloads
Dataset Card

SemEval 2025 Task 11 - Track C Dataset

This dataset contains the data for SemEval 2025 Task 11: Bridging the Gap in Text-Based Emotion Detection - Track C, organized as language-specific configurations.

Dataset Description

The dataset is a multi-language, multi-label emotion classification dataset with separate configurations for each language.

  • —Total languages: 30 standard ISO codes
  • —Total examples: 57254
  • —Splits: dev, test (Track C has no train split)

Track Information

Track C has more languages than Track B, but does not include a training set. It only provides dev and test splits for each language.

Language Configurations

Each language is available as a separate configuration with the following statistics:

ISO CodeOriginal Code(s)Dev ExamplesTest ExamplesTotal
afafr9810651163
amamh59217742366
arary, arq36717142081
dedeu20026042804
eneng11627672883
esesp18416951879
hahau35610801436
hihin10010101110
idind1568511007
igibo47914441923
jvjav151837988
mrmar10010001100
omorm57417212295
pcmpcm62018702490
ptptbr, ptmz45730023459
roron12311191242
rurus19910001199
rwkin40712311638
sosom56616962262
susun1999261125
svswe20011881388
swswa55116562207
titir61418402454
tttat20010001200
ukukr24922342483
vmwvmw2587771035
xhxho68215942276
yoyor49715001997
zhchn20026422842
zuzul87520472922

Features

  • —id: Unique identifier for each example
  • —text: Text content to classify
  • —anger, disgust, fear, joy, sadness, surprise: Presence of emotion
  • —emotions: List of emotions present in the text

Usage

python
from datasets import load_dataset

# Load all data for a specific language
eng_dataset = load_dataset("YOUR_USERNAME/semeval-2025-task11-track-c", "eng")

# Or load a specific split for a language
eng_dev = load_dataset("YOUR_USERNAME/semeval-2025-task11-track-c", "eng", split="dev")

Citation

If you use this dataset, please cite the following papers:

@misc{{muhammad2025brighterbridginggaphumanannotated,
      title={{BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages}}, 
      author={{Shamsuddeen Hassan Muhammad and Nedjma Ousidhoum and Idris Abdulmumin and Jan Philip Wahle and Terry Ruas and Meriem Beloucif and Christine de Kock and Nirmal Surange and Daniela Teodorescu and Ibrahim Said Ahmad and David Ifeoluwa Adelani and Alham Fikri Aji and Felermino D. M. A. Ali and Ilseyar Alimova and Vladimir Araujo and Nikolay Babakov and Naomi Baes and Ana-Maria Bucur and Andiswa Bukula and Guanqun Cao and Rodrigo Tufiño and Rendi Chevi and Chiamaka Ijeoma Chukwuneke and Alexandra Ciobotaru and Daryna Dementieva and Murja Sani Gadanya and Robert Geislinger and Bela Gipp and Oumaima Hourrane and Oana Ignat and Falalu Ibrahim Lawan and Rooweither Mabuya and Rahmad Mahendra and Vukosi Marivate and Andrew Piper and Alexander Panchenko and Charles Henrique Porto Ferreira and Vitaly Protasov and Samuel Rutunda and Manish Shrivastava and Aura Cristina Udrea and Lilian Diana Awuor Wanzare and Sophie Wu and Florian Valentin Wunderlich and Hanif Muhammad Zhafran and Tianhui Zhang and Yi Zhou and Saif M. Mohammad}},
      year={{2025}},
      eprint={{2502.11926}},
      archivePrefix={{arXiv}},
      primaryClass={{cs.CL}},
      url={{https://arxiv.org/abs/2502.11926}}, 
}}
@misc{{muhammad2025semeval2025task11bridging,
      title={{SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Detection}}, 
      author={{Shamsuddeen Hassan Muhammad and Nedjma Ousidhoum and Idris Abdulmumin and Seid Muhie Yimam and Jan Philip Wahle and Terry Ruas and Meriem Beloucif and Christine De Kock and Tadesse Destaw Belay and Ibrahim Said Ahmad and Nirmal Surange and Daniela Teodorescu and David Ifeoluwa Adelani and Alham Fikri Aji and Felermino Ali and Vladimir Araujo and Abinew Ali Ayele and Oana Ignat and Alexander Panchenko and Yi Zhou and Saif M. Mohammad}},
      year={{2025}},
      eprint={{2503.07269}},
      archivePrefix={{arXiv}},
      primaryClass={{cs.CL}},
      url={{https://arxiv.org/abs/2503.07269}}, 
}}

License

This dataset is licensed under CC-BY 4.0.