CoolFace
20 results

zulu

ZuluVision /RaCig-Datatext10M<n<100M2 likes1.2k downloads1y agoHugging Facestrombergnlp /zulu_stanceThis is a stance detection dataset in the Zulu language. The data is translated to Zulu by Zulu native speakers, from English source texts. Misinformation has become a major concern in recent last years given its spread across our information sources. In the past years, many NLP tasks have been introduced in this area, with some systems reaching good results on English language datasets. Existing AI based approaches for fighting misinformation in literature suggest automatic stance detection as an integral first step to success. Our paper aims at utilizing this progress made for English to transfers that knowledge into other languages, which is a non-trivial task due to the domain gap between English and the target languages. We propose a black-box non-intrusive method that utilizes techniques from Domain Adaptation to reduce the domain gap, without requiring any human expertise in the target language, by leveraging low-quality data in both a supervised and unsupervised manner. This allows us to rapidly achieve similar results for stance detection for the Zulu language, the target language in this work, as are found for English. We also provide a stance detection dataset in the Zulu language.texttext-classification1K<n<10K1 likes84 downloads4y agoHugging FaceGyimah3 /zulu-music-listening-clipsaudion<1K0 likes43 downloads1mo agoHugging Facemichsethowusu /Code-170k-zulu Dataset Description Code-170k-zulu is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Zulu, making coding education accessible to Zulu speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Zulu language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-zulu.texttext-generation100K<n<1M0 likes29 downloads11mo agoHugging Facesaillab /alpaca-zulu-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-zulu-cleaned.text10K<n<100K0 likes28 downloads2y agoHugging Faceokezieowen /afrispeech_zuluaudio1K<n<10K0 likes28 downloads1y agoHugging Face