datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
groupK_project_BroadcastflowCN
A Dataset of Longitudinal Acoustic Changes for Middle-Aged Voice Aging based on 15 Years of Voice Data from Program Host in Mainland China
Abstract
This dataset provides a unique 15-year longitudinal acoustic record tracking the vocal aging of a professional Mandarin-speaking television host from age 35 to 50 (2012–2026). Comprising high-quality audio extracted from broadcast interviews and programs, the collection captures the subtle physiological and prosodic… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/groupK_project_BroadcastflowCN.Annotated_Food_Vlog_Dataset_GroupL
Dataest Description
This project has constructed a multimodal corpus of language strategies for food exploration videos on Chinese social media. The dataset is centered around the videos of the well-known blogger "Diao Yueshe Shi Yu Ji", containing approximately 1,000 entries with a total of 90 minutes of transcribed video content. The dataset is stored in CSV format and meticulously records the original dialogue, synthetic text generated by large language models (LLMs), rhetorical… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/Annotated_Food_Vlog_Dataset_GroupL.
