datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HD-Hunyuan-30K-Distill-Dataverl-distill-datasetmultilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.Dataset_Distillation_ReproductionThis dataset is for keeping track on the reproduction of some Dataset Distillation methods. The format of each folder is like dataset/ipc{n}/class_name.
Reference:
SRe2L:https://github.com/VILA-Lab/SRe2L/tree/main/SRe2L
WMDD:https://github.com/Liu-Hy/WMDD
CVDD:https://github.com/Jiacheng8/CV-DD
G-VBSM:https://github.com/shaoshitong/G_VBSM_Dataset_Condensation
math128_R1_1.5B_VER-Qwen2.5-1.5B_data-distill_r1_qwen_1p5B_gpt_4o_verify_proc_all_train_1e-5distilled_with_prompts
