meld
Datasets
All datasets matching “meld”MELDliepa-3
LIEPA-3 — Lithuanian Speech Corpus
Didysis lietuvių kalbos garsynas (LIEPA-3)
Dataset Summary
LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours,
~7.5 million audio files) built for automatic speech recognition (ASR),
text-to-speech (TTS) and linguistic research. It spans read, spontaneous,
phonetically-annotated and dialectal speech recorded under a wide range of
conditions (studio, dictaphone, radio, TV, telephone, audiobooks).
Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.MELDMELD-processed
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
liepa-2
Dataset Card for LIEPA-2
Dataset Summary
The LIEPA-2 dataset is a large-scale annotated speech corpus for the Lithuanian language, developed under the project "Development of Services Controlled by Lithuanian Speech" (LIEPA-2). It is a phonetically representative, structured collection of data (audio recordings and annotations) designed for scientific research in speech technologies and the development of electronic services.
Total Duration: 1000 hours
Access:… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-2.MELD-DS-448
Appendix: MELD-DS-448 Dataset Overview
Dataset Overview
MELD-DS-448 contains 26,166 malicious samples spanning 448 distinct malware families collected from April 2020 to August 2025. All samples are uniquely identified by SHA-256 hashes and include precise "First Seen" timestamps.
Family Distribution Characteristics: The dataset exhibits a typical long-tail distribution, with 35.7% singleton families (only 1 sample) and 64.7% small-scale families (≤5 samples).… See the full description on the dataset page: https://huggingface.co/datasets/MeldProject/MELD-DS-448.
