datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProverbEval
ProverbEval: Benchmark for Evaluating LLMs on Low-Resource Proverbs
This dataset accompanies the paper:"ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding"ArXiv:2411.05049v3
Dataset Summary
ProverbEval is a culturally grounded evaluation benchmark designed to assess the language understanding abilities of large language models (LLMs) in low-resource settings. It consists of tasks based on proverbs in five languages:
Amharic
Afaan… See the full description on the dataset page: https://huggingface.co/datasets/israel/ProverbEval.AfriGuard
AfriGuard: Safety Evaluation Data for African Languages
AfriGuard is a human-annotated safety dataset covering 10 African languages: Amharic, Hausa, Igbo, Oromo, Shona, Swahili, Twi, Wolof, Yoruba, and Zulu. Each example contains a culturally grounded prompt/response pair in English and the target language, labeled with a safety top category, a safe/unsafe label, and majority-vote annotations from three native-speaker annotators.
Splits
Each language config… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard.frontend_dpo
DPO JavaScript Dataset
This repository contains a modified and expanded version of a closed-source JavaScript dataset. The dataset has been adapted to fit the DPO (Dynamic Programming Object) format, making it compatible with the LLaMA-Factory project. The dataset includes a variety of JavaScript code snippets with optimizations and best practices, generated using closed-source tools and expanded by me.
License
This dataset is licensed under the Apache 2.0 License.… See the full description on the dataset page: https://huggingface.co/datasets/israellaguan/frontend_dpo.kwaiklear-sample-level-agent-trajectories-2.2Mwaxal-autolabled
Auot-Lableing Waxal unlabeled dataset on Best Multilingual Ethio-ASR models
@article{abdullah2026ethio,
title={Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages},
author={Abdullah, Badr M and Azime, Israel Abebe and Tonja, Atnafu Lambebo and Alabi, Jesujoba O and Alemu, Abel Mulat and Hagos, Eyob G and Balcha, Bontu Fufa and Nerea, Mulubrhan A and Yadeta, Debela Desalegn and Marilign, Dagnachew Mekonnen and others}… See the full description on the dataset page: https://huggingface.co/datasets/israel/waxal-autolabled.amharic-speech-expandedipfs_israel_laws_ir
Israel legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_israel_laws (revision 86221a671797406ddfd71daf033378292578a86d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Israel prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_israel_laws_ir.israeli_law
Open Israeli Law (Hebrew Wikisource)
Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query.
Nearly 6,000 pages. Almost 100 million characters. One jsonl.
This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.med-safety-bench-reproduced
Med Safety Bench Dataset
This dataset is a collection of harmful medical requests and safe responses, originally sourced from the AI4LIFE-GROUP/med-safety-bench GitHub repository.
The dataset is intended for research purposes related to the safety of medical AI models.
Data Structure
The dataset contains the following columns:
harmful_medical_request: The harmful medical request.
safe_response: The safe response to the request.
source: Indicates whether the data… See the full description on the dataset page: https://huggingface.co/datasets/israel-adewuyi/med-safety-bench-reproduced.c4-10k-tokenized-gpt2kwaiklear-sample-level-agent-trajectories-350Kcase-law-israel
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/case-law-israel.kwaiklear-sample-level-agent-trajectories-170Kipfs_israel_laws
Israel In-Force National Laws (Knesset / Reshumot)
Research snapshot of official national legislation from Knesset OData + Reshumot PDFs (fs.knesset.gov.il).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-18
Coverage
in-force snapshot
Source
Knesset OData + Reshumot PDFs (fs.knesset.gov.il)
Collector
scrapers/collect_knesset_laws.py
Laws / instruments
995
Articles… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_israel_laws.israel-public-transitflores-parallelbank_of_israel
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/bank_of_israel
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Bank of Israel.
Label Interpretation
Stance Detection… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/bank_of_israel.israel-declaration-executable-companion-v1.0
Israel Declaration Executable Companion
Credit: This dataset and companion package were generated by DBbun LLC.
This repository contains an executable companion package generated from the Hebrew text of the Israeli Declaration of Independence (Megillat Ha'Atzmaut). The package includes structured documentation, simulator code, a simulation specification, synthetic output tables, generated figures, and summary metadata.
The goal is to demonstrate a document-to-executable-companion… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/israel-declaration-executable-companion-v1.0.Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.sdsdsdsdsldYgx5w24IGOdMfIsrael-Photos
Israel Photos Dataset
A collection of 369 photographs captured across Israel between 2024 and 2025, with LLM-generated captions and location annotations. The images are sourced from the photographer's Pexels gallery.
About This Collection
This dataset was deliberately curated to provide a diverse visual representation of Israel, encompassing:
Varied locations: From the historic streets of Jerusalem's Old City to Tel Aviv's urban landscape, desert vistas in the Negev, and… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Israel-Photos.Israel-Stock-Symbols-and-Metadata
Israel Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Israel.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Israel-Stock-Symbols-and-Metadata.Iran-Israel-War-2026
Iran-Israel War — OSINT Dataset
Open-source intelligence dataset tracking Iranian missile and drone attack waves against Israel and US/coalition targets across four operations in the "True Promise" series (2024–2026).
Dataset Description
53 attack waves across four Iranian military operations, each with 89 structured fields covering timing, weapons systems, targets, interception performance, casualties, and escalation indicators. Also includes international reactions data… See the full description on the dataset page: https://huggingface.co/datasets/brainrot8756/Iran-Israel-War-2026.wa-commerce-blindspot-eval
West African Micro-Commerce Blind Spot Eval
Author: Israel Olanrewaju Odeajo (israelkingz)Challenge: Fatima Fellowship — Technical Challenge (Blind Spots of Frontier Models)Model under test: Qwen/Qwen2.5-3B-Instruct (~3B open-weight)Scope: A 24-item diagnostic set (not a large-scale benchmark claim) + reproducible GPU eval + qualitative failure analysis
1. The blind spot (lived experience)
Title
Semantic inversion and high-stakes failure modes in… See the full description on the dataset page: https://huggingface.co/datasets/israelkingz/wa-commerce-blindspot-eval.AmazonReview
Dataset Card for "AmazonReview"
More Information needed
Iran-Israel-War-2026
Iran-Israel War — OSINT Dataset
Open-source intelligence dataset tracking Iranian missile and drone attack waves against Israel and US/coalition targets across four operations in the "True Promise" series (2024–2026).
Dataset Description
53 attack waves across four Iranian military operations, each with 89 structured fields covering timing, weapons systems, targets, interception performance, casualties, and escalation indicators. Also includes international reactions data… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Iran-Israel-War-2026.israeli-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Israel
The Synthetic Israel Passports Dataset compiles more than 1,000 AI-generated passport images intended for training OCR and computer vision models on identity documents. Each record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/israeli-passports.AfriGuard-inst
AfriGuard-inst
Alpaca-style instruction-tuning data for safety-aligned fine-tuning, built from
the train/validation splits of
israel/AfriGuard and
israel/AfriGuard-XL.
Test splits are excluded and reserved for evaluation.
Construction
AfriGuard (10 languages): each row yields TWO items — one English
(prompt/response) and one native-language
(prompt_translated/response_translated).
AfriGuard-XL (34 split configs, English-only): each row yields one item.
Rows with… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard-inst.my-quiz-eval
