datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liepa-3
LIEPA-3 — Lithuanian Speech Corpus
Didysis lietuvių kalbos garsynas (LIEPA-3)
Dataset Summary
LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours,
~7.5 million audio files) built for automatic speech recognition (ASR),
text-to-speech (TTS) and linguistic research. It spans read, spontaneous,
phonetically-annotated and dialectal speech recorded under a wide range of
conditions (studio, dictaphone, radio, TV, telephone, audiobooks).
Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.liepa-2
Dataset Card for LIEPA-2
Dataset Summary
The LIEPA-2 dataset is a large-scale annotated speech corpus for the Lithuanian language, developed under the project "Development of Services Controlled by Lithuanian Speech" (LIEPA-2). It is a phonetically representative, structured collection of data (audio recordings and annotations) designed for scientific research in speech technologies and the development of electronic services.
Total Duration: 1000 hours
Access:… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-2.liepa-tts
LIEPA TTS Dataset
Dataset Summary
This dataset contains recovered and organized utterance-level audio from four human speakers recorded for the Vilnius University LIEPA speech-synthesis project. It includes 20,180 WAV recordings (about 12 hours), aligned text, and several stress representations.
The original LIEPA project produced the recordings, synthesis voices, and synthesis system. The current dataset presents those resources in a structured, stress-enriched… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-tts.meld_emotion_test@article{poria2018meld,
title={Meld: A multimodal multi-party dataset for emotion recognition in conversations},
author={Poria, Soujanya and Hazarika, Devamanyu and Majumder, Navonil and Naik, Gautam and Cambria, Erik and Mihalcea, Rada},
journal={arXiv preprint arXiv:1810.02508},
year={2018}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/meld_emotion_test.MELDm-meldWav2Vec_MELD_AudioMELD
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset meld-emo --model gpt4o_audio
python audio_evals/main.py --dataset meld-sentiment --model gpt4o_audio
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/MELD.MELD-Preprocessed
MELD Preprocessed for SER
This dataset is the manually preprocessed audio only version of MELD, only audio IDs, utterance transcriptions, dialogue IDs and Utterance IDs were extracted.
S. Poria, D. Hazarika, N. Majumder, G. Naik, R. Mihalcea,
E. Cambria. MELD: A Multimodal Multi-Party Dataset
for Emotion Recognition in Conversation. (2018)
Chen, S.Y., Hsu, C.C., Kuo, C.C. and Ku, L.W.
EmotionLines: An Emotion Corpus of Multi-Party
Conversations. arXiv preprint arXiv:1802.08379… See the full description on the dataset page: https://huggingface.co/datasets/Vano04/MELD-Preprocessed.MELD-splitsmeld_sentiment_test@article{poria2018meld,
title={Meld: A multimodal multi-party dataset for emotion recognition in conversations},
author={Poria, Soujanya and Hazarika, Devamanyu and Majumder, Navonil and Naik, Gautam and Cambria, Erik and Mihalcea, Rada},
journal={arXiv preprint arXiv:1810.02508},
year={2018}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/meld_sentiment_test.meld-acoustic-datasetMELD-audio-test
Dataset Card for "MELD-audio-test"
More Information needed
SpeechSentimentAnalysis_MELDliepa-asr
Liepa ASR Dataset
Lithuanian Automatic Speech Recognition (ASR) dataset from the LIEPA project (Lietuvių šnekos garsynas LIEPA) developed at Vilnius University.It provides a phonetically representative corpus for ASR and TTS research, capturing diverse speakers and recording styles.
Dataset Summary
The Liepa ASR dataset contains speech recordings and their transcriptions, designed for both speech recognition and speech synthesis research.
Total speakers: 376 (248… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-asr.HateSpeechDetection_Detoxy_VCTK_LJSpeech_CV_MELDMELDVibeCheckAudio_MELD_chunk_002_of_999MELD-processed
Dataset Card for "MELD-processed"
More Information needed
MELD-videos-absolute-pathsVibeCheckAudio_MELD_chunk_004_of_999VibeCheckRandomSamples_VibeCheckAudio_MELDMELD_audioMELDVibeCheckAudio_MELD_chunk_001_of_999VibeCheckAudio_MELD_chunk_003_of_999VibeCheckAudio_MELDmeld_subsetMELD_Audio_3LabelsMultimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
MELD contains the same dialogue instances available in EmotionLines, but it also encompasses audio and
visual modality along with text. MELD has more than 1400 dialogues and 13000 utterances from Friends TV series.
Multiple speakers participated in the dialogues. Each utterance in a dialogue has been labeled by any of these
seven emotions -- Anger, Disgust, Sadness, Joy, Neutral, Surprise and Fear. MELD also has sentiment (positive,
negative and neutral) annotation for each utterance.
This dataset is slightly modified, so that it concentrates on Emotion recognition in audio input only.MELD-PC-VA
