Israel
Datasets
All datasets matching “Israel”ProverbEval
ProverbEval: Benchmark for Evaluating LLMs on Low-Resource Proverbs
This dataset accompanies the paper:"ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding"ArXiv:2411.05049v3
Dataset Summary
ProverbEval is a culturally grounded evaluation benchmark designed to assess the language understanding abilities of large language models (LLMs) in low-resource settings. It consists of tasks based on proverbs in five languages:
Amharic
Afaan… See the full description on the dataset page: https://huggingface.co/datasets/israel/ProverbEval.AfriGuard
AfriGuard: Safety Evaluation Data for African Languages
AfriGuard is a human-annotated safety dataset covering 10 African languages: Amharic, Hausa, Igbo, Oromo, Shona, Swahili, Twi, Wolof, Yoruba, and Zulu. Each example contains a culturally grounded prompt/response pair in English and the target language, labeled with a safety top category, a safe/unsafe label, and majority-vote annotations from three native-speaker annotators.
Splits
Each language config… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard.frontend_dpo
DPO JavaScript Dataset
This repository contains a modified and expanded version of a closed-source JavaScript dataset. The dataset has been adapted to fit the DPO (Dynamic Programming Object) format, making it compatible with the LLaMA-Factory project. The dataset includes a variety of JavaScript code snippets with optimizations and best practices, generated using closed-source tools and expanded by me.
License
This dataset is licensed under the Apache 2.0 License.… See the full description on the dataset page: https://huggingface.co/datasets/israellaguan/frontend_dpo.kwaiklear-sample-level-agent-trajectories-2.2Mwaxal-autolabled
Auot-Lableing Waxal unlabeled dataset on Best Multilingual Ethio-ASR models
@article{abdullah2026ethio,
title={Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages},
author={Abdullah, Badr M and Azime, Israel Abebe and Tonja, Atnafu Lambebo and Alabi, Jesujoba O and Alemu, Abel Mulat and Hagos, Eyob G and Balcha, Bontu Fufa and Nerea, Mulubrhan A and Yadeta, Debela Desalegn and Marilign, Dagnachew Mekonnen and others}… See the full description on the dataset page: https://huggingface.co/datasets/israel/waxal-autolabled.amharic-speech-expanded
