creole
Datasets
All datasets matching “creole”cmu_haitian_creole_speech###########################################################################
Language Technologies Institute
Carnegie Mellon University
Copyright (c) 2010
All Rights Reserved.
Permission is hereby granted, free of charge, to use and distribute
this data and its documentation without restriction, including
without limitation the rights to use, copy, modify, merge, publish,
distribute… See the full description on the dataset page: https://huggingface.co/datasets/jsbeaudry/cmu_haitian_creole_speech.haitian-creole-processed-extendeddcreole_rcCreoleRC is a subset created by the CreoleVal paper. Relation classification (RC) aims to identify semantic associations between entities within a text, essential for applications like knowledge base completion and question answering. The dataset is sourced from Wikipedia and manually annotated. CreoleRC contains 5 creoles, but SEACrowd is interested specifically in the Chavacano subset.english-seychelles-french-creole_sentence-pairs_mt560
English-Seychelles French Creole Parallel Dataset
This dataset contains parallel sentences in English and Seychelles French Creole (Seychelles).
Dataset Information
Language Pair: English ↔ Seychelles French Creole
Language Code: crs
Country: Seychelles
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-seychelles-french-creole_sentence-pairs_mt560.alpaca_haitian_creole_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_haitian_creole_taco.Code-170k-mauritian-creole
Dataset Description
Code-170k-mauritian-creole is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Mauritian Creole, making coding education accessible to Mauritian Creole speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Mauritian Creole language - democratizing coding education
Multi-turn dialogues covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-mauritian-creole.
