datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-170k-zulu
Dataset Description
Code-170k-zulu is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Zulu, making coding education accessible to Zulu speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Zulu language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-zulu.zulu-pretraining-datasetThis is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B
Books: Zulu-English Dictionary – A dictionary offering Zulu terms with English definitions, ideal for teaching basic word mappings.
Translation: South African Government Speeches – Official speeches in Zulu, which help the model understand structured Zulu sentences and phrases.
Transcription: Zulu Community Corpus – A collection of transcriptions, exposing the model to real-life conversational Zulu.
Document:… See the full description on the dataset page: https://huggingface.co/datasets/ChallengerSpaceShuttle/zulu-pretraining-dataset.Roleplay-Zulu
RolePlay-Zulu
Roleplay-Zulu Dataset is a dataset for roleplaying in the Zulu language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Zulu.
