CoolFace
20 results

cantonese

alvanlii /cantonese-radiogated Cantonese Radio Pseudo-Transcription Dataset Contains 14k hours of audio sourced from Archive.org Columns order_index: Represents the order of the audio compared to those from the same filename link: Link of the original full audio transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.audioautomatic-speech-recognition1M<n<10M25 likes2.2k downloads2y agoHugging Faceziyou-li /cantonese_dailyaudio1K<n<10K4 likes1.5k downloads4y agoHugging Facejed351 /Cantonese_Common_Crawl_Filtered Cantonese Chinese C4 Dataset Summary Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant traditional Chinese dataset can be found here. This dataset contains data processed with CantoneseDetect. In CantoneseDetect, you can choose whether to include quotes (i.e. categorise the data as Cantonese even if Cantonese appeared only in quotes). And I found that a lot of entries came from Wikipedia and LIHKG. If you… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese_Common_Crawl_Filtered.text1M<n<10M4 likes1.1k downloads1y agoHugging Facealvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes804 downloads6mo agoHugging FaceScicom-intl /YouTube-Cantonese-Emilia YouTube Cantonese — Emilia 2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by running alvanlii/cantonese-youtube through the Emilia speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering). Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.tabularautomatic-speech-recognition1M<n<10M1 likes790 downloads1mo agoHugging FaceAlienKevin /sbs_cantonese SBS Cantonese Speech Corpus This speech corpus contains 435 hours of SBS Cantonese podcasts from Auguest 2022 to October 2023. There are 2,519 episodes and each episode is split into segments that are at most 10 seconds long. In total, there are 189,216 segments in this corpus. Here is a breakdown on the categories of episodes present in this dataset: Category SBS Channels Episodes news 中文新聞, 新聞簡報 622 business 寰宇金融 148 vaccine 疫苗快報 71 gardening 園藝趣談 58 tech 科技世界… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/sbs_cantonese.audio10K<n<100K7 likes789 downloads3y agoHugging Face