datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liepa-3
LIEPA-3 — Lithuanian Speech Corpus
Didysis lietuvių kalbos garsynas (LIEPA-3)
Dataset Summary
LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours,
~7.5 million audio files) built for automatic speech recognition (ASR),
text-to-speech (TTS) and linguistic research. It spans read, spontaneous,
phonetically-annotated and dialectal speech recorded under a wide range of
conditions (studio, dictaphone, radio, TV, telephone, audiobooks).
Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.liepa-2
Dataset Card for LIEPA-2
Dataset Summary
The LIEPA-2 dataset is a large-scale annotated speech corpus for the Lithuanian language, developed under the project "Development of Services Controlled by Lithuanian Speech" (LIEPA-2). It is a phonetically representative, structured collection of data (audio recordings and annotations) designed for scientific research in speech technologies and the development of electronic services.
Total Duration: 1000 hours
Access:… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-2.MELD-DS-448
Appendix: MELD-DS-448 Dataset Overview
Dataset Overview
MELD-DS-448 contains 26,166 malicious samples spanning 448 distinct malware families collected from April 2020 to August 2025. All samples are uniquely identified by SHA-256 hashes and include precise "First Seen" timestamps.
Family Distribution Characteristics: The dataset exhibits a typical long-tail distribution, with 35.7% singleton families (only 1 sample) and 64.7% small-scale families (≤5 samples).… See the full description on the dataset page: https://huggingface.co/datasets/MeldProject/MELD-DS-448.liepa-tts
LIEPA TTS Dataset
Dataset Summary
This dataset contains recovered and organized utterance-level audio from four human speakers recorded for the Vilnius University LIEPA speech-synthesis project. It includes 20,180 WAV recordings (about 12 hours), aligned text, and several stress representations.
The original LIEPA project produced the recordings, synthesis voices, and synthesis system. The current dataset presents those resources in a structured, stress-enriched… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-tts.meld-open
MELD Open
MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details.
Note: This version of MELD Open retains the original labels from its source datasets. For normalized labels, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open.meld_emotion_test@article{poria2018meld,
title={Meld: A multimodal multi-party dataset for emotion recognition in conversations},
author={Poria, Soujanya and Hazarika, Devamanyu and Majumder, Navonil and Naik, Gautam and Cambria, Erik and Mihalcea, Rada},
journal={arXiv preprint arXiv:1810.02508},
year={2018}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/meld_emotion_test.meld-open-normalized
MELD Open (Normalized)
MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details.
Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.MELDm-meldMELD-dataset
MELD — Mathematical Equivalence under Linguistic Diversity
MELD is a small, hand-curated evaluation benchmark for math-aware text embedding
models. It tests one specific capability: does the model recognize that two statements
describing the same mathematical fact are equivalent even when they are written in
the vocabulary, notation, and conventions of different mathematical subfields?
MELD was originally part of
uw-math-ai/Math2Vec-embedding-dataset
and is released here as a… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/MELD-dataset.lt-stressed-corpus
LT Stressed Corpus
Lithuanian sentences with word stress marks.
The dataset combines MATAS v1.0 and ALKSNIS v3.0. It is useful for speech technology, pronunciation work, language learning, and research on Lithuanian stress.
This release contains sentences where every Lithuanian word that needs stress has a selected stressed form. Sentences with Arabic or Roman numerals are left out because reading a number correctly depends on context and grammatical form.
Some Lithuanian words… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/lt-stressed-corpus.MELD-processed-v3-wavlm
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
MELD
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset meld-emo --model gpt4o_audio
python audio_evals/main.py --dataset meld-sentiment --model gpt4o_audio
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/MELD.MELD-Preprocessed
MELD Preprocessed for SER
This dataset is the manually preprocessed audio only version of MELD, only audio IDs, utterance transcriptions, dialogue IDs and Utterance IDs were extracted.
S. Poria, D. Hazarika, N. Majumder, G. Naik, R. Mihalcea,
E. Cambria. MELD: A Multimodal Multi-Party Dataset
for Emotion Recognition in Conversation. (2018)
Chen, S.Y., Hsu, C.C., Kuo, C.C. and Ku, L.W.
EmotionLines: An Emotion Corpus of Multi-Party
Conversations. arXiv preprint arXiv:1802.08379… See the full description on the dataset page: https://huggingface.co/datasets/Vano04/MELD-Preprocessed.MELD-emotion-detection-preprocessed
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me

Multimodal Emotion Recognition Dataset (Processed from MELD)
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/MELD-emotion-detection-preprocessed.MELD-processed-v4-opensmile25d
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
MELD-splitsMELD-processed-v5-emotion2vec
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
meld_sentiment_test@article{poria2018meld,
title={Meld: A multimodal multi-party dataset for emotion recognition in conversations},
author={Poria, Soujanya and Hazarika, Devamanyu and Majumder, Navonil and Naik, Gautam and Cambria, Erik and Mihalcea, Rada},
journal={arXiv preprint arXiv:1810.02508},
year={2018}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/meld_sentiment_test.meld-acoustic-datasetMELD-processed-v2
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
MELD-MPCA
Dataset Information
the dataset is a jsonl file containing each dialogue (context) per line.
Field
Amount
Dialogue (context/line)
1022
Diff User
260
each dialogue context messages of a conversation, with those informations:
user
content
emotion
type
n_turn
summary
traits
distanglement
ref_speaker
ref_utterance
tar_speaker
selected_speaker
The user of the message
The content of the message
The emotion of the user
Either a positive or negative emotion
The… See the full description on the dataset page: https://huggingface.co/datasets/neoluigi/MELD-MPCA.vlm-compositionality-embeddings
VLM Compositionality Embeddings
Pre-computed image and text embeddings for the thesis "From Euclidean to Hyperbolic Vision-Language Spaces: A Study of Attribute–Object Compositionality" by Meelad Dashti (Politecnico di Torino & University of Twente, 2026).
Code repository: github.com/MelDashti/hyperbolic-vlm-compositionality
Models
Model
Geometry
Architecture
Training Data
CLIP ViT-L/14
Spherical
ViT-L/14
WIT (400M+ pairs)
DINOv2 ViT-L/14
Spherical
ViT-L/14… See the full description on the dataset page: https://huggingface.co/datasets/Meldashti/vlm-compositionality-embeddings.turkish-medical-articles-rag
Turkish Medical Articles - RAG System & Vector Database
Bu proje, Türkçe tıbbi makaleler üzerinde çalışan bir Retrieval-Augmented Generation (RAG) sistemi ve Vektör Veritabanı uygulamasıdır. Proje kapsamında ham veriler Hugging Face üzerinden çekilmiş, semantik parçalama uygulanmış, magibu/embeddingmagibu-200m modeliyle vektörleştirilmiş, ChromaDB üzerinde saklanmış ve başlangıç eşiği ile dinamik eşik optimizasyonu adımlarını içeren 30 soruluk bir benchmark testiyle… See the full description on the dataset page: https://huggingface.co/datasets/meldakahramann/turkish-medical-articles-rag.MELD-audio-test
Dataset Card for "MELD-audio-test"
More Information needed
meld-sentiment-classificationSpeechSentimentAnalysis_MELDliepa-asr
Liepa ASR Dataset
Lithuanian Automatic Speech Recognition (ASR) dataset from the LIEPA project (Lietuvių šnekos garsynas LIEPA) developed at Vilnius University.It provides a phonetically representative corpus for ASR and TTS research, capturing diverse speakers and recording styles.
Dataset Summary
The Liepa ASR dataset contains speech recordings and their transcriptions, designed for both speech recognition and speech synthesis research.
Total speakers: 376 (248… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-asr.meld-dataset-instruct-Option_2meld-transcript-final
