CoolFace
Datasetpublic

eduhk-compling/GroupE_groupproject_DayoWong_StandUp_Comedy

Project Title DayoWong StandUp Comedy Last Cantonese word Annotation and Joke Structure Dataset Description This corpus annotates the last Cantonese word in every sentence of Dayo Wong's 1999 stand-up set (“黃子華 Dayo 1999 棟篤笑 拾下拾下”, https://www.youtube.com/watch?v=nSe3nhopbpg), as well as the structure of the jokes. This corpus excludes the last word if it is in English, using for relevant tags. The corpus is annotated with pitch(1-5), using the tone system of… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/GroupE_groupproject_DayoWong_StandUp_Comedy.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes166downloads
Dataset Card

Project Title

DayoWong StandUp Comedy Last Cantonese word Annotation and Joke Structure

Dataset Description

This corpus annotates the last Cantonese word in every sentence of Dayo Wong's 1999 stand-up set (“黃子華 Dayo 1999 棟篤笑 拾下拾下”, https://www.youtube.com/watch?v=nSe3nhopbpg), as well as the structure of the jokes. This corpus excludes the last word if it is in English, using </na> for relevant tags.

The corpus is annotated with pitch(1-5), using the tone system of numbers based on: Chao, Y.-R. (1947), Cantonese Primer (Cambridge, Mass.: Harvard University Press).

The other categories of tags are as follows:

tones \</rising> \</level> \</falling>: Which describes the movement of the voice pitch.

duration \</elongation> \</truncation> \</nochange>: Which describes the duration of the sound in the recording.

Structure of joke \</setup> \</misdirection> \</punchline> \</tag> \</transition>: Which describes the part of the joke the sentence is.

pic \</name>: Which is the person responsible for annotating the segment.

check1 \</initials1>: Which is the first person checking the segment.

check2 \</initials2>: Which is the second person checking the segment.