CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajesujoba /yoruba_text_c3Yoruba Text C3 is the largest Yoruba texts collected and used to train FastText embeddings in the YorubaTwi Embedding paper: https://www.aclweb.org/anthology/2020.lrec-1.335/text-generation100K<n<1M3 likes190 downloads3y agoHugging Face02adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes102 downloads14d agoHugging Face03michsethowusu /Code-170k-yoruba Dataset Description Code-170k-yoruba is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Yoruba, making coding education accessible to Yoruba speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Yoruba language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-yoruba.texttext-generation100K<n<1M0 likes35 downloads11mo agoHugging Face04Lots-of-LoRAs /task612_yorubabbc_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task612_yorubabbc_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task612_yorubabbc_classification.texttext-generation1K<n<10K0 likes34 downloads2y agoHugging Face05saaga /yoruba-cultural-reasoning-blindspots# Yoruba Cultural Reasoning Blind Spots in Frontier Models ## Overview This dataset captures the "blind spots" of frontier base models when evaluating non-Western abstract reasoning, specifically focusing on Yoruba proverbs. It contains 10 diverse examples demonstrating "generative collapse" and "Cultural Hallucination." ## 1. Loading the Model This evaluation was conducted using a standard Google Colab environment with a T4 GPU. The model evaluated was `Qwen/Qwen2.5-1.5B`. It was loaded in… See the full description on the dataset page: https://huggingface.co/datasets/saaga/yoruba-cultural-reasoning-blindspots.texttext-generationn<1K1 likes33 downloads7mo agoHugging Face061nnocent /adtc-agri-yoruba-training-mix ADTC Agri-Yoruba Training Mix An English/Yoruba instruction tuning dataset for agricultural extension advisory, built for the Africa Deep Tech Challenge 2026 Laptop LLM track by Team Support Vector. Used to fine tune llama-3.2-3b-agri-yoruba. Composition 43,640 examples total, in chat message format (system, user, assistant): Source Rows Share naija_yoruba 27,464 63% ai71_agrillm 14,996 34% identity_control 800 2% behavior_control 240 1%… See the full description on the dataset page: https://huggingface.co/datasets/1nnocent/adtc-agri-yoruba-training-mix.texttext-generation10K<n<100K0 likes21 downloads1mo agoHugging Face07Enochid /foundry-y-yoruba-corpus Yorùbá News & Tasks — a contamination-proof, 100% human Yorùbá corpus A 19,997-row supervised fine-tuning corpus ({"prompt","completion"} JSONL) for adapting a language model to Yorùbá, built for the Adaption AutoScientist challenge (language track). Every row is human-written, from an approved, train-split-only source. No synthetic or LLM-generated text. SHA-256: a711692096719c0d11f8e9c3784061211733029bf32a20e36b7e920488f8cc0c Rows: 19,997 · Format: JSONL, prompt + completion… See the full description on the dataset page: https://huggingface.co/datasets/Enochid/foundry-y-yoruba-corpus.texttranslation10K<n<100K0 likes16 downloads3mo agoHugging Face08BabsDest /yoruba_pidgin_agriculture_datatexttext-generationn<1K1 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.