datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.Dataset_Distillation_ReproductionThis dataset is for keeping track on the reproduction of some Dataset Distillation methods. The format of each folder is like dataset/ipc{n}/class_name.
Reference:
SRe2L:https://github.com/VILA-Lab/SRe2L/tree/main/SRe2L
WMDD:https://github.com/Liu-Hy/WMDD
CVDD:https://github.com/Jiacheng8/CV-DD
G-VBSM:https://github.com/shaoshitong/G_VBSM_Dataset_Condensation
