CoolFace
Datasetpublic

saeedzou/yodas2-en-000-116-00000000-emotion-filtered

YODAS2 Emotion Dataset Pipeline This project builds an automatically labeled speech emotion dataset from the English portion of YODAS2. The pipeline combines speaker diarization, voice activity detection, speech segmentation, and predictions from five pretrained speech emotion recognition models. The resulting labels are filtered using model agreement and then downsampled to reduce the strong class imbalance in the source data. Source Data The source dataset is… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/yodas2-en-000-116-00000000-emotion-filtered.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes63downloads
Dataset Card

YODAS2 Emotion Dataset Pipeline

This project builds an automatically labeled speech emotion dataset from the English portion of YODAS2.

The pipeline combines speaker diarization, voice activity detection, speech segmentation, and predictions from five pretrained speech emotion recognition models. The resulting labels are filtered using model agreement and then downsampled to reduce the strong class imbalance in the source data.

Source Data

The source dataset is YODAS2, available on Hugging Face:

https://huggingface.co/datasets/espnet/yodas2

The English subsets used by the project are organized as:

  • —en000–en007
  • —en100–en129

For this processing run, the 00000000.tar.gz audio shard was downloaded from:

  • —en000–en007
  • —en100–en110

For example:

data/en000/audio/00000000.tar.gz

Processing Pipeline

Speaker Diarization

The downloaded audio was first processed with the community-1 diarization pipeline to identify individual speakers and produce speaker-specific speech regions.

Voice Activity Detection

The diarized regions were then processed with Silero VAD to identify active speech and remove non-speech portions.

Segment Processing

The VAD segments were cleaned using the following rules:

  • —Adjacent segments shorter than one second were merged.
  • —Segments shorter than one second were discarded.
  • —Adjacent segments longer than one second were not merged.

The decision not to merge longer segments was intentional. Longer segments can contain multiple emotional contexts, and merging them could make the subsequent emotion prediction less reliable.

The resulting segments were treated as individual utterances for emotion recognition.

Emotion Recognition

Each utterance was passed through five pretrained speech emotion recognition models:

  1. 1.emotion2vec+ Seed
  2. 2.emotion2vec+ Base
  3. 3.emotion2vec+ Large
  4. 4.MERaLiON-SER-v1
  5. 5.WavLM Large Categorical Emotion

The five predictions were combined using majority voting. An utterance was retained when at least four of the five models agreed on the emotion. Utterances without sufficient agreement were discarded.

The retained labels were:

  • —Angry
  • —Happy
  • —Neutral
  • —Sad
  • —Fearful
  • —Disgusted
  • —Surprised

Dataset Statistics

The initial segmentation produced 101,421 utterances. After majority-vote filtering, 63,290 utterances were retained and 38,131 were discarded.

Neutral was heavily overrepresented, with 56,691 utterances (~90%). The second-largest class was Happy, with 4,597 utterances, corresponding to approximately 14.17 hours of speech.

Neutral was therefore downsampled to match the size of the Happy class. No oversampling was performed for the minority classes.

The final dataset contains 11,196 utterances:

EmotionUtterances
Neutral4,597
Happy4,597
Angry1,166
Sad646
Surprised141
Fearful27
Disgusted22
Total11,196

[image]

The duration distribution of the utterances is shown below.

[image]

Why Multiple Emotion Models?

Five pretrained emotion recognition models were used to reduce dependence on any single model. The final label was determined by majority vote.

Empirically, a 4-model agreement threshold performed better than a 3-model threshold in our evaluation, so we used agreement from at least four of the five models for the final dataset.

Limitations

  • —Language: Although this dataset is derived from the English subset of YODAS2, we observed videos that appear to contain speech in languages other than English, likely due to errors in the upstream language metadata. No additional language identification or filtering was performed.
  • —Automatically generated labels: The emotion labels are pseudo-labels generated by pretrained SER models and have not been manually verified.
  • —Source data: YODAS2 is not an emotion-specific corpus, resulting in a strong natural bias toward Neutral speech.
  • —Class imbalance: Disgusted (22), Fearful (27), and Surprised (141) have very few examples.
  • —Model selection: Requiring agreement from four of five models may favor clearer emotional examples and exclude ambiguous ones.
  • —Segmentation: Keeping longer adjacent segments separate can result in utterances with limited emotional context.

Pipeline

text
YODAS2 English audio
        |
        v
Speaker diarization
        |
        v
Silero VAD
        |
        v
Segment cleanup
  - merge adjacent segments < 1s
  - discard segments < 1s
  - keep longer adjacent segments separate
        |
        v
5 speech emotion recognition models
        |
        v
Majority voting
        |
        v
Keep 7 emotion classes
        |
        v
Downsample Neutral
        |
        v
11,196 utterances

Project Status

The current results are based on 25 processed shards:

  • —en000–en007
  • —en100–en110

Each YODAS2 subset contains approximately 500 shards, with each shard containing around 38 minutes of audio. Across the 25 processed shards, the pipeline produced 14.17 hours of retained Happy speech.

Based on this processing rate, applying the same pipeline to the complete English YODAS2 dataset could yield approximately 1,076 hours of speech. This is an estimate based on the current processing results and may vary across shards.

Source Models

The emotion recognition stage currently uses:

Citation

If you use YODAS2 as part of your work, please cite the original YODAS2 publication and dataset.

@inproceedings{li2023yodas,
  title={Yodas: Youtube-Oriented Dataset for Audio and Speech},
  author={Li, Xinjian and Takamichi, Shinnosuke and Saeki, Takaaki and Chen, William and Shiota, Sayaka and Watanabe, Shinji},
  booktitle={2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)},
  pages={1--8},
  year={2023},
  organization={IEEE}
}

If you use this derived dataset, please cite this repository and specify the dataset version or commit used in your experiments.

@misc{zou2026yodas2emotionfiltered,
  author       = {Zouashkiani, Saeed},
  title        = {YODAS2 English Speech Emotion Recognition Filtered Dataset},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/saeedzou/yodas2-en-000-116-00000000-emotion-filtered}},
}