datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProverbEval
ProverbEval: Benchmark for Evaluating LLMs on Low-Resource Proverbs
This dataset accompanies the paper:"ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding"ArXiv:2411.05049v3
Dataset Summary
ProverbEval is a culturally grounded evaluation benchmark designed to assess the language understanding abilities of large language models (LLMs) in low-resource settings. It consists of tasks based on proverbs in five languages:
Amharic
Afaan… See the full description on the dataset page: https://huggingface.co/datasets/israel/ProverbEval.AfriGuard
AfriGuard: Safety Evaluation Data for African Languages
AfriGuard is a human-annotated safety dataset covering 10 African languages: Amharic, Hausa, Igbo, Oromo, Shona, Swahili, Twi, Wolof, Yoruba, and Zulu. Each example contains a culturally grounded prompt/response pair in English and the target language, labeled with a safety top category, a safe/unsafe label, and majority-vote annotations from three native-speaker annotators.
Splits
Each language config… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard.flores-parallelAmharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.Israel-Stock-Symbols-and-Metadata
Israel Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Israel.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Israel-Stock-Symbols-and-Metadata.ldYgx5w24IGOdMfAfriGuard-inst
AfriGuard-inst
Alpaca-style instruction-tuning data for safety-aligned fine-tuning, built from
the train/validation splits of
israel/AfriGuard and
israel/AfriGuard-XL.
Test splits are excluded and reserved for evaluation.
Construction
AfriGuard (10 languages): each row yields TWO items — one English
(prompt/response) and one native-language
(prompt_translated/response_translated).
AfriGuard-XL (34 split configs, English-only): each row yields one item.
Rows with… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard-inst.AmharicStoryQA6OHOLvM38r8Kgy1
MoleculeNet Benchmark (website)
MoleculeNet is a benchmark specially designed for testing machine learning methods of molecular properties. As we aim to facilitate the development of molecular machine learning method, this work curates a number of dataset collections, creates a suite of software that implements many known featurizations and previously proposed algorithms. All methods and datasets are integrated as parts of the open source DeepChem package(MIT license).
MoleculeNet… See the full description on the dataset page: https://huggingface.co/datasets/israel/6OHOLvM38r8Kgy1.AfriGuard-XL
AfriGuard-XL
AfriGuard-XL extends israel/AfriGuard
with culturally grounded safety prompts for many more African countries/regions.
Each scenario contributes 4 examples (2 safe / 2 unsafe).
Splits
Country configs not covered by AfriGuard provide train / validation / test
splits (~40% / 10% / 50%), following the AfriGuard split methodology:
Split at the scenario level (scenario_id) — every validation/test scenario
is fully unseen in train.
Scenario → split… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard-XL.flores_plusprojectaccept_or_denylocalizationstil-2024-main_datasetAmharicZefenAmharicSpellCheckAmharicQAIsraeli-Palestinian-Conflict
Dataset Card Creation Guide
Dataset Summary
The Israeli-Palestinian-Conflict dataset is an English-language dataset contains manually collected claims, regarding the Israel-Palestine conflict, annotated both objectively with multi-labels to categorize the content according to common themes in such arguments, and subjectively by their level of impact on a moderately informed citizen. The primary purpose of this dataset is to support Israeli public relations efforts at… See the full description on the dataset page: https://huggingface.co/datasets/avishagnevo/Israeli-Palestinian-Conflict.NEAQKba0G6fNVmIsemeval2007_task_14SNETstil-2024-human_expertsNEAQKba0G6fNVmI-new-2eval_data_alphabet_sortMezmurCompletionSNET_Archiveiabd-datasetExtraído de https://github.com/anthony-wang/BestPractices/tree/master/data.
Campos:
Formula (string)
T (float64): Temperatura (k)
CP (float64): Capacidad calorifica (J/mol K)
Act5_titanicdoc-exp
