CoolFace
15 results

afrikaans

andreoosthuizen /afrikaans-30s Afrikaans Speech Dataset for Whisper Fine-Tuning Dataset Card Dataset Summary This dataset consists of approximately 56 hours of Afrikaans speech extracted from church sermons, paired with cleaned and aligned transcripts. It is specifically prepared for fine-tuning multilingual ASR models like OpenAI's Whisper (particularly large-v3) on low-resource Afrikaans speech The audio is segmented into fixed 30-second chunks (with 3-second overlaps for context… See the full description on the dataset page: https://huggingface.co/datasets/andreoosthuizen/afrikaans-30s.audioautomatic-speech-recognition1K<n<10K0 likes288 downloads2mo agoHugging Facenwu-ctext /afrikaans_ner_corpus Dataset Card for Afrikaans Ner Corpus Dataset Summary The Afrikaans Ner Corpus is an Afrikaans dataset developed by The Centre for Text Technology (CTexT), North-West University, South Africa. The data is based on documents from the South African goverment domain and crawled from gov.za websites. It was created to support NER task for Afrikaans language. The dataset uses CoNLL shared task annotation standards. Supported Tasks and Leaderboards [More… See the full description on the dataset page: https://huggingface.co/datasets/nwu-ctext/afrikaans_ner_corpus.texttoken-classification1K<n<10K8 likes197 downloads3y agoHugging Facemichsethowusu /afrikaans-english-emotions-corpus Afrikaans-english Emotion Analysis Corpus Dataset Description This dataset contains emotion-labeled text data in Afrikaans-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-english-emotions-corpus.texttext-classification1M<n<10M0 likes152 downloads1y agoHugging Faceshunyalabs /afrikaans-speech-datasetaudio1K<n<10K1 likes111 downloads1y agoHugging Facevoice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes86 downloads2y agoHugging FaceNicheVault /nichevault-afrikaans-asr-sample NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.audioautomatic-speech-recognitionn<1K0 likes56 downloads2mo agoHugging Face