CoolFace
Datasetpublic

nanyang-technological-university-singapore/hkcancor

The Hong Kong Cantonese Corpus (HKCanCor) comprise transcribed conversations recorded between March 1997 and August 1998. It contains recordings of spontaneous speech (51 texts) and radio programmes (42 texts), which involve 2 to 4 speakers, with 1 text of monologue. In total, the corpus contains around 230,000 Chinese words. The text is word-segmented, annotated with part-of-speech (POS) tags and romanised Cantonese pronunciation. Romanisation scheme - Linguistic Society of Hong Kong (LSHK) POS scheme - Peita-Fujitsu-Renmin Ribao (PRF) corpus (Duan et al., 2000), with extended tags for Cantonese-specific phenomena added by Luke and Wang (see original paper for details).

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
17likes194downloads
settings

This repository belongs to nanyang-technological-university-singapore on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namehkcancor
visibilitypublic
licencecc-by-4.0
gatedno
ownernanyang-technological-university-singapore
Account settings
nanyang-technological-university-singapore/hkcancor · CoolFace