Afrikaans
afrikaans-30s
Afrikaans Speech Dataset for Whisper Fine-Tuning
Dataset Card
Dataset Summary
This dataset consists of approximately 56 hours of Afrikaans speech extracted from church sermons, paired with cleaned and aligned transcripts. It is specifically prepared for fine-tuning multilingual ASR models like OpenAI's Whisper (particularly large-v3) on low-resource Afrikaans speech
The audio is segmented into fixed 30-second chunks (with 3-second overlaps for context… See the full description on the dataset page: https://huggingface.co/datasets/andreoosthuizen/afrikaans-30s.afrikaans_ner_corpus
Dataset Card for Afrikaans Ner Corpus
Dataset Summary
The Afrikaans Ner Corpus is an Afrikaans dataset developed by The Centre for Text Technology (CTexT), North-West University, South Africa. The data is based on documents from the South African goverment domain and crawled from gov.za websites. It was created to support NER task for Afrikaans language. The dataset uses CoNLL shared task annotation standards.
Supported Tasks and Leaderboards
[More… See the full description on the dataset page: https://huggingface.co/datasets/nwu-ctext/afrikaans_ner_corpus.afrikaans-english-emotions-corpus
Afrikaans-english Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Afrikaans-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-english-emotions-corpus.afrikaans-speech-datasetopenslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.alpaca_afrikaans_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_afrikaans_taco.
