CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01andreoosthuizen /afrikaans-30s Afrikaans Speech Dataset for Whisper Fine-Tuning Dataset Card Dataset Summary This dataset consists of approximately 56 hours of Afrikaans speech extracted from church sermons, paired with cleaned and aligned transcripts. It is specifically prepared for fine-tuning multilingual ASR models like OpenAI's Whisper (particularly large-v3) on low-resource Afrikaans speech The audio is segmented into fixed 30-second chunks (with 3-second overlaps for context… See the full description on the dataset page: https://huggingface.co/datasets/andreoosthuizen/afrikaans-30s.audioautomatic-speech-recognition1K<n<10K0 likes261 downloads2mo agoHugging Face02nwu-ctext /afrikaans_ner_corpus Dataset Card for Afrikaans Ner Corpus Dataset Summary The Afrikaans Ner Corpus is an Afrikaans dataset developed by The Centre for Text Technology (CTexT), North-West University, South Africa. The data is based on documents from the South African goverment domain and crawled from gov.za websites. It was created to support NER task for Afrikaans language. The dataset uses CoNLL shared task annotation standards. Supported Tasks and Leaderboards [More… See the full description on the dataset page: https://huggingface.co/datasets/nwu-ctext/afrikaans_ner_corpus.texttoken-classification1K<n<10K8 likes184 downloads3y agoHugging Face03michsethowusu /afrikaans-english-emotions-corpus Afrikaans-english Emotion Analysis Corpus Dataset Description This dataset contains emotion-labeled text data in Afrikaans-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-english-emotions-corpus.texttext-classification1M<n<10M0 likes151 downloads1y agoHugging Face04shunyalabs /afrikaans-speech-datasetaudio1K<n<10K1 likes131 downloads1y agoHugging Face05voice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes92 downloads2y agoHugging Face06saillab /alpaca_afrikaans_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_afrikaans_taco.text10K<n<100K0 likes74 downloads2y agoHugging Face07NicheVault /nichevault-afrikaans-asr-sample NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.audioautomatic-speech-recognitionn<1K0 likes54 downloads2mo agoHugging Face08michsethowusu /afrikaans-english_sentence-pairstext10M<n<100M0 likes44 downloads1y agoHugging Face09Max5ive /nchlt_speech_afrikaans NCHLT Speech Corpus -- Afrikaans This is the Afrikaans language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): afr URI: https://hdl.handle.net/20.500.12185/280 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/nchlt_speech_afrikaans.audioautomatic-speech-recognition10K<n<100K0 likes38 downloads6mo agoHugging Face10okezieowen /afrispeech_afrikaansaudio1K<n<10K0 likes34 downloads1y agoHugging Face11michsethowusu /afrikaans-dinka_sentence-pairs Afrikaans-Dinka_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Dinka_Sentence-Pairs Number of Rows: 113795 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-dinka_sentence-pairs.text100K<n<1M0 likes27 downloads1y agoHugging Face12michsethowusu /english-afrikaans_sentence-pairs_mt560 English-Afrikaans Parallel Dataset This dataset contains parallel sentences in English and Afrikaans (South Africa). Dataset Information Language Pair: English ↔ Afrikaans Language Code: afr Country: South Africa Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-afrikaans_sentence-pairs_mt560.text1M<n<10M0 likes27 downloads1y agoHugging Face13Beijuka /Afrikaans_testaudio1K<n<10K1 likes23 downloads2y agoHugging Face14jojo-ai-mst /Roleplay-Afrikaans RolePlay-Afrikaans Roleplay-Afrikaans Dataset is a dataset for roleplaying in the Afrikaans language for Large Language Model. The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, it can be found at this… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Afrikaans.texttext-generation1K<n<10K2 likes20 downloads2y agoHugging Face15michsethowusu /afrikaans-sentiments-corpus Afrikaans Sentiment Corpus Dataset Description This dataset contains sentiment-labeled text data in Afrikaans for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources. Dataset Statistics Total samples: 1,500,000 Positive sentiment: 803944… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-sentiments-corpus.texttext-classification1M<n<10M0 likes20 downloads1y agoHugging Face16saillab /alpaca-afrikaans-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-afrikaans-cleaned.text10K<n<100K1 likes19 downloads2y agoHugging Face17michsethowusu /afrikaans-tswana_sentence-pairs Afrikaans-Tswana_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Tswana_Sentence-Pairs Number of Rows: 779259… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-tswana_sentence-pairs.text100K<n<1M0 likes19 downloads1y agoHugging Face18jacobmorrison /MATH-500-afrikaanstextn<1K0 likes19 downloads1y agoHugging Face19michsethowusu /afrikaans-wolof_sentence-pairs Afrikaans-Wolof_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Wolof_Sentence-Pairs Number of Rows: 237049 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-wolof_sentence-pairs.text100K<n<1M0 likes16 downloads1y agoHugging Face20michsethowusu /afrikaans-somali_sentence-pairs Afrikaans-Somali_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Somali_Sentence-Pairs Number of Rows: 1432536… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-somali_sentence-pairs.text1M<n<10M0 likes16 downloads1y agoHugging Face21michsethowusu /Code-170k-afrikaans Dataset Description Code-170k-afrikaans is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Afrikaans, making coding education accessible to Afrikaans speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Afrikaans language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-afrikaans.texttext-generation100K<n<1M0 likes16 downloads11mo agoHugging Face22michsethowusu /afrikaans-tigrinya_sentence-pairs Afrikaans-Tigrinya_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Tigrinya_Sentence-Pairs Number of Rows: 454333… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-tigrinya_sentence-pairs.text100K<n<1M0 likes15 downloads1y agoHugging Face23michsethowusu /afrikaans-ganda_sentence-pairs Afrikaans-Ganda_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Ganda_Sentence-Pairs Number of Rows: 477046 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-ganda_sentence-pairs.text100K<n<1M0 likes14 downloads1y agoHugging Face24michsethowusu /afrikaans-fulah_sentence-pairs Afrikaans-Fulah_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Fulah_Sentence-Pairs Number of Rows: 168995 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-fulah_sentence-pairs.text100K<n<1M0 likes14 downloads1y agoHugging Face25michsethowusu /afrikaans-bambara_sentence-pairs Afrikaans-Bambara_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Bambara_Sentence-Pairs Number of Rows: 121709… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-bambara_sentence-pairs.text100K<n<1M0 likes14 downloads1y agoHugging Face26ULAIS /Glossaries_Sample_Afrikaans_SMCgatedtextn<1K1 likes14 downloads5mo agoHugging Face27michsethowusu /afrikaans-kamba_sentence-pairs Afrikaans-Kamba_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Kamba_Sentence-Pairs Number of Rows: 99196 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-kamba_sentence-pairs.text10K<n<100K0 likes13 downloads1y agoHugging Face28jacobmorrison /alpacaeval-afrikaanstextn<1K0 likes13 downloads1y agoHugging Face29eugenetanjc /speech_accent_5_afrikaanstextn<1K0 likes12 downloads4y agoHugging Face30michsethowusu /afrikaans-akan_sentence-pairs Afrikaans-Akan_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Akan_Sentence-Pairs Number of Rows: 96786 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-akan_sentence-pairs.text10K<n<100K1 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.