datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MELD-DS-448
Appendix: MELD-DS-448 Dataset Overview
Dataset Overview
MELD-DS-448 contains 26,166 malicious samples spanning 448 distinct malware families collected from April 2020 to August 2025. All samples are uniquely identified by SHA-256 hashes and include precise "First Seen" timestamps.
Family Distribution Characteristics: The dataset exhibits a typical long-tail distribution, with 35.7% singleton families (only 1 sample) and 64.7% small-scale families (≤5 samples).… See the full description on the dataset page: https://huggingface.co/datasets/MeldProject/MELD-DS-448.MELD-ST
MELD-ST: An Emotion-aware Speech Translation Dataset
Paper: https://arxiv.org/abs/2405.13233
Overview
This emotion-aware speech translation dataset is a multi-language dataset extracted from the TV show "Friends." It includes English, Japanese, and German subtitles along with corresponding timestamps. This dataset is designed for natural language processing tasks.
Contents
The dataset is partitioned into train, test, and development subsets to streamline… See the full description on the dataset page: https://huggingface.co/datasets/ku-nlp/MELD-ST.MELD_text10K_MELD_Plus_v1.0
Synthetic MELD-Plus (10K Patients)
Watch a demo
This dataset contains 10,000 synthetic patients inspired by the published MELD-Plus study (a collboration between Massachusetts General Hospital and IBM Research). Each row corresponds to a single admission, with demographics, labs, comorbidities, medications, derived scores (MELD, MELD-Na, MELD-Plus), and the binary outcome Death_Within_90_Days.
All data are artificially generated and contain no identifiable patient records.… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/10K_MELD_Plus_v1.0.1M_MELD_Plus_v1.0
Synthetic MELD-Plus (1M Patients)
This dataset contains 1,000,000 synthetic patients inspired by the published MELD-Plus study (a collboration between Massachusetts General Hospital and IBM Research). Each row corresponds to a single admission, with demographics, labs, comorbidities, medications, derived scores (MELD, MELD-Na, MELD-Plus), and the binary outcome Death_Within_90_Days.
All data are artificially generated and contain no identifiable patient records.… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/1M_MELD_Plus_v1.0.meld_clips_v3
目录结构
/
├── metadata.csv # 数据集元数据文件
├── videos/ # 视频文件目录
│ ├── dev_sample_2.mp4
│ ├── dev_sample_3.mp4
│ └── ...
└── audios/ # 音频文件目录
├── mfa_meld_dev_oov.txt # MFA对齐时的OOV(out-of-vocabulary)词汇
├── ost/ # 原始音轨(Original SoundTrack)
│ ├── dev_sample_2.wav
│ ├── dev_sample_3.wav
│ └── ...
├── vocals/ # 人声音轨(分离后)
│ ├── dev_sample_2_vocals.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BigfufuOuO/meld_clips_v3.100K_MELD_Plus_v1.0
Synthetic MELD-Plus (100K Patients)
Watch a demo
This dataset contains 100,000 synthetic patients inspired by the published MELD-Plus study (a collboration between Massachusetts General Hospital and IBM Research). Each row corresponds to a single admission, with demographics, labs, comorbidities, medications, derived scores (MELD, MELD-Na, MELD-Plus), and the binary outcome Death_Within_90_Days.
All data are artificially generated and contain no identifiable patient records.… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/100K_MELD_Plus_v1.0.meld
Dataset Card for "MELD_Text"
More Information needed
