CoolFace
Datasetpublicgated

alvanlii/cantonese-radio

Cantonese Radio Pseudo-Transcription Dataset Contains 14k hours of audio sourced from Archive.org Columns order_index: Represents the order of the audio compared to those from the same filename link: Link of the original full audio transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.

sourceHugging Faceupdated 2y agoView on Hugging Face
25likes2.2kdownloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.