datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-170k-mauritian-creole
Dataset Description
Code-170k-mauritian-creole is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Mauritian Creole, making coding education accessible to Mauritian Creole speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Mauritian Creole language - democratizing coding education
Multi-turn dialogues covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-mauritian-creole.Code-170k-seychellois-creole
Dataset Description
Code-170k-seychellois-creole is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Seychellois Creole, making coding education accessible to Seychellois Creole speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Seychellois Creole language - democratizing coding education
Multi-turn dialogues covering various… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-seychellois-creole.haitian-creole-synthetic-v1
Haitian Creole Synthetic Dataset v2
Dataset Summary
This is a synthetic dataset of Haitian Creole text samples across multiple domains, including legal, medical, education, business, technology, daily life, and culture. The dataset is designed for natural language processing tasks such as text classification, language modeling, and machine translation.
Dataset Splits
Split
Samples
Description
Train
908
Training set
Validation
113
Validation set… See the full description on the dataset page: https://huggingface.co/datasets/Vladimirht/haitian-creole-synthetic-v1.
