CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yvthyvq /KarenSo_CantoneseRecordings_Liujgoj KarenSo_CantoneseRecordings_Liujgoj 本數據集係基於開源粵語語音庫 kakiso/KarenSo_CantoneseRecordings 進行重構與正字、拼音重映射嘅溜歌粵語(Liujgoj)語音正字 Ground Truth 數據集。主要用於音標切分、粵語語音流形(Language Manifold)對齊、以及 Stage 2 SFT 翻譯與語言工程任務。 👥 致謝與上游數據集說明 (Acknowledgment & Upstream Source) 本數據集嘅原始音頻與文本來源於 Hugging Face 社群成員 kakiso 分享嘅項目: 原始數據集 (Original Dataset): kakiso/KarenSo_CantoneseRecordings 原始授權協議 (License): CC-BY-4.0 在此由衷感謝原創作者 Karen So 及其團隊錄製並無私分享高品質(44.1kHz / 16bit / Mono)嘅純淨粵語口語語料,為廣東話開源 AI… See the full description on the dataset page: https://huggingface.co/datasets/Yvthyvq/KarenSo_CantoneseRecordings_Liujgoj.audioautomatic-speech-recognition0 likes162 downloads3mo agoHugging Face02freococo /sagaw_karen_asrThis is the first public Sagaw Karen language ASR dataset in AI history. Sagaw Karen ASR This dataset contains audio recordings and aligned metadata in the Sagaw Karen language (ISO 639-3: ksw), a major Sgaw Karenic language spoken throughout southern and eastern Myanmar. The language is sometimes also referred to as Sgaw Karen or Sakaw Karen in English transliterations. All audio segments in this dataset were sourced from publicly available news broadcasts published by PVTV… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sagaw_karen_asr.audioautomatic-speech-recognition1K<n<10K0 likes39 downloads1y agoHugging Face03freococo /karenni_language_asr_audio RFA Karenni (Kayah) Language Voices This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. This dataset was created by freococo. The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.audioautomatic-speech-recognition1K<n<10K0 likes37 downloads1y agoHugging Face04eduhk-compling /KarenSo Dataset Description This dataset includes 50 sentences spoken in colloquial Hong Kong Cantonese (HKC), covering interrogatives and statements. Sentences are sourced from online Cantonese teaching materials and classical commercial slogans. It includes approxiametly 198 seconds of audio recorded by a female native speaker of HKC. The sampling rate is 44.1 kHz with 16-bit resolution. Transcription in Jyutping were also provided. Issues Encountered & Solution I… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/KarenSo.audion<1K0 likes34 downloads8mo agoHugging Face05freococo /western_poe_karen_asrThis is the first public Western Poe Karen language ASR dataset in AI history. Western Poe Karen ASR This dataset contains audio recordings and aligned transcriptions in the Western Poe Karen language (also known in linguistic literature as Western Pwo or Delta Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in the Ayeyarwady Delta region of Myanmar. Although linguists commonly refer to this language as Western Pwo Karen, the community and this project prefer the spelling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/western_poe_karen_asr.audioautomatic-speech-recognition1K<n<10K0 likes30 downloads1y agoHugging Face06freococo /eastern_poe_karen_asrThis is the first public Eastern Poe Karen language ASR dataset in AI history. Eastern Poe Karen ASR This dataset contains audio recordings and aligned metadata in the Eastern Poe Karen language (a regional variety of Eastern Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in Mon State and Kayin State in southeastern Myanmar. While linguistically described as Eastern Pwo Karen, the community and this project prefer the term Poe as a community-endorsed spelling. All audio… See the full description on the dataset page: https://huggingface.co/datasets/freococo/eastern_poe_karen_asr.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads1y agoHugging Face07collectivat /ladino-karen-TTS Ladino Text-to-Speech (TTS) Training Dataset Dataset Description This dataset contains a single-speaker speech corpus in Ladino (Judeo-Spanish) recorded by a native speaker from Istanbul. The corpus was created for training text-to-speech synthesis models for this endangered language. Dataset Statistics Speaker: Karen (native Ladino speaker) Recordings: 1987 segments Total Duration: ~3.3 hours Sampling Rate: 16 kHz Audio Format: WAV (16-bit, mono) Language:… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/ladino-karen-TTS.audiotext-to-speech1K<n<10K0 likes20 downloads11mo agoHugging Face08kakiso /KarenSo_CantoneseRecordings Dataset Description This dataset includes 50 sentences spoken in colloquial Hong Kong Cantonese (HKC), covering interrogatives and statements. Sentences are sourced from online Cantonese teaching materials and classical commercial slogans. It includes approxiametly 198 seconds of audio recorded by a female native speaker of HKC. The sampling rate is 44.1 kHz with 16-bit resolution. Transcription in Jyutping were also provided. Issues Encountered & Solution I… See the full description on the dataset page: https://huggingface.co/datasets/kakiso/KarenSo_CantoneseRecordings.audion<1K0 likes17 downloads5mo agoHugging Face09aravdash /karenTTSaudio1K<n<10K1 likes13 downloads1y agoHugging Face10archivartaunik /maksim-garetski-rodnae-karenne-andrei-kaliada Роднае карэнне Metadata Author: Максім Гарэцкі Title: Роднае карэнне Narrator: Андрэй Каляда Source Group: Аўдыёкнігі Source: Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders. Target maximum split size:… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/maksim-garetski-rodnae-karenne-andrei-kaliada.audion<1K0 likes4 downloads4mo agoHugging Face11archivartaunik /maksim-garetski-rodnae-karenne-valer-mazynski Роднае карэнне Metadata Author: Максім Гарэцкі Title: Роднае карэнне Narrator: Валер Мазынскі Source Group: Аўдыёкнігі Source: rutracker.org Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders. Target maximum… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/maksim-garetski-rodnae-karenne-valer-mazynski.audion<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.