datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aaiEmoTa
EmoTa: A Tamil Emotional Speech Dataset
EmoTa is the first emotional speech dataset in Tamil, designed to reflect the
linguistic diversity of Sri Lankan Tamil speakers. It contains 936 recorded
utterances from 22 native Tamil speakers (11 male, 11 female), each articulating
19 semantically neutral sentences across five emotions: anger, happiness,
sadness, fear, and neutral.
🌐 Project page: https://aaivu.github.io/EmoTa/
💻 Code & loader: https://github.com/aaivu/EmoTa
📄 Paper… See the full description on the dataset page: https://huggingface.co/datasets/aaivu-labs/EmoTa.DEAR
DEAR Dataset
Dataset Summary
The Deep Evaluation of Audio Representations (DEAR) dataset is a benchmark designed to assess general-purpose audio foundation models on properties critical for hearable devices.
It comprises 1,158 mono audio tracks (30 s each), spatially mixing proprietary anechoic speech monologues with high-quality everyday acoustic scene recordings from the HOA‑SSR library.
DEAR enables controlled evaluation of:
Context (environment type:… See the full description on the dataset page: https://huggingface.co/datasets/HSLU-AAI/DEAR.tunisian-english-speech
