datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chat-resultsmacula-hebrew-syntax
NuBerea MACULA Hebrew Syntax Trees (OT)
Full syntactic tree annotation of the Hebrew Bible from the MACULA Hebrew Linguistic Dataset, packaged as relational tables for computational biblical studies. The dataset covers word-level linguistic annotation (morphology, glosses, lexical semantics), sentence segmentation, and hierarchical syntactic structure (clauses and phrases with their roles and containment relations) over the Westminster Leningrad Codex base text.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-hebrew-syntax.hebrew-poetry
NuBerea Hebrew Poetry
Computational study of the two signature rhetorical devices of biblical Hebrew verse: chiasmus (mirrored, inverted structure) and semantic parallelism (paired lines that echo or contrast one another). The dataset covers the poetic corpus of the Hebrew Bible with detected and ranked chiastic passages, classified poetry verses, line-segmentation of verses into their paired cola, and scored parallelism judgments — resources for literary, rhetorical, and… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/hebrew-poetry.hebrew_this_world
Dataset Card for HebrewSentiment
Dataset Summary
HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license.
Data Annotation:
Supported Tasks and Leaderboards
Language modeling
Languages
Hebrew
Dataset Structure
csv file with "," delimeter
Data Instances
Sample:
{
"issue_num": 637,
"page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.Hebrew_VAD_lexicon
Hebrew VAD Lexicon
The Hebrew VAD Lexicon is an enhanced version of an automatically translated affective lexicon, originally derived from the English VAD lexicon created by Mohammad (2018) .
It provides valence, arousal, and dominance (VAD) scores for Hebrew words.
The lexicon was carefully curated by manually reviewing and correcting the automatic translations and enriching the dataset with additional linguistic information.
There are two versions of the dataset:
English-Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/GiliGold/Hebrew_VAD_lexicon.sefaria-hebrew
Dataset Card for "sefaria-hebrew"
this dataset contains jewish texts in hebrew from the sefaria project
hebrew-asr-vn
Hebrew ASR three-source training dataset
Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees.
Original data and transcripts are credited to ivrit.ai and its contributors.
Pinned revisions and preparation rules are in metadata/sources.json and
metadata/preparation-config.json. VoxKnesset is excluded by user decision.
Source/split
Clips
Hours
crowd-recital/test
1,557
1.071
crowd-recital/train
45,372
33.258
crowd-recital/validation
1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.hebrew-wikipedia-sentences-corpus
Hebrew Wikipedia Sentences Corpus
A corpus of 10,999,257 cleaned, deduplicated Hebrew sentences extracted from 366,610 Hebrew Wikipedia articles.
Dataset Description
This dataset contains Hebrew sentences extracted from Hebrew Wikipedia (crawled 2026-02). Each sentence has been cleaned, filtered for quality, and deduplicated. The dataset is intended for Hebrew NLP tasks including language modeling, text classification, NER, sentence similarity, and more.
Source… See the full description on the dataset page: https://huggingface.co/datasets/tomron87/hebrew-wikipedia-sentences-corpus.hebrew_suffix_verbal_forms
Suffixed Verbal Forms Detection Dataset for Modern Hebrew
Dataset Summary
This dataset contains annotated Hebrew sentences containing verbal forms that are ambiguous as to whether they include a pronominal suffix or not (e.g., the Hebrew word lamed-yod-mem-daled-vav can be understood as either "he taught him" or "they taught"). The goal of the dataset is to support tasks involving the identification and disambiguation of verbs with pronominal suffixes in Hebrew literature… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/hebrew_suffix_verbal_forms.hebrew-targum-vocalized
Vocalized Hebrew–Targum Parallel Corpus
Verse-aligned parallel corpus of Biblical Hebrew and Targumic Aramaic, covering 15119 verses of the Hebrew Bible.
Structure
field
type
description
book
int
Book number, 1–39 in standard Hebrew Bible order
book_name
string
English book name
chapter
int
Chapter
verse
int
Verse
hebrew
string
Masoretic Hebrew
targum
string
Targum Onkelos / Jonathan
split
verses
train
13646
validation
718… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/hebrew-targum-vocalized.hebrew-psychotechnique-MQA
Israeli Psychometric Exam (NITE) — Multiple-Choice QA
1570 multiple-choice questions extracted from 26 publicly released NITE
psychometric entrance exams (2019–2026).
Splits
split
rows
notes
verbal
927
Hebrew, RTL
english
620
English
quantitative
23
almost nothing survives filtering
Fields
question, options (4), answer (1-indexed into options)
section / part / number — position within the exam
source_pdf / page — provenance, for… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-psychotechnique-MQA.yam-peleg__Hebrew-Mistral-7B-details
Dataset Card for Evaluation run of yam-peleg/Hebrew-Mistral-7B
Dataset automatically created during the evaluation run of model yam-peleg/Hebrew-Mistral-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yam-peleg__Hebrew-Mistral-7B-details.judaic-aramaic-hebrew-lexicon-v1
Judaic Aramaic-Hebrew Lexicon v1
Auxiliary lexical-alignment data. Hebrew glosses are paired with
Aramaic headwords and vocalized aliases. Multiple senses share a group id and
splits are assigned by canonical headword to prevent leakage.
Provenance, contribution, and license
The Aramaic-Hebrew lexical content originates in Torat Emet; it was not authored by Otzaria. This project extracted and normalized the entries, constructed alignment records, grouped senses… See the full description on the dataset page: https://huggingface.co/datasets/ArieLLL123/judaic-aramaic-hebrew-lexicon-v1.HebrewRecipes
Dataset Card for Hebrew Recipes
A dataset of recipes scraped from the Israeli recipe websites.
Dataset Details
Dataset Description
The dataset contains recipes scraped from multiple Israeli recipe websites, including sugat.com and hashulchan.co.il. It includes structured JSON-LD data conforming to schema.org Recipe specifications, cleaned HTML from the printable recipe views, and various metadata for each recipe URL.
Curated by: [@Wissotsky]
Language: [Hebrew]… See the full description on the dataset page: https://huggingface.co/datasets/Wissotsky/HebrewRecipes.yam-peleg__Hebrew-Gemma-11B-Instructmodern_hebrew_bible_he
Modern Hebrew Bible
Description
The complete Bible in Modern Hebrew, including both the Old Testament (Tanakh) and the New Testament (Brit Hadashah). The Old Testament follows the Masoretic tradition while the New Testament is a translation into Modern Hebrew by various translators. This is the primary Bible used by Hebrew-speaking Christians and Messianic Jews in Israel today.
Source: Unbound Bible / Biola University electronic text.
License: Public Domain… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/modern_hebrew_bible_he.HebrewSearch-gemma-3-dataHebrewSearch-bge-datahebrew_masoretic_ot_hbo
Hebrew Masoretic Old Testament (כתבי הקודש)
Description
The complete Hebrew Masoretic Text of the Old Testament (Tanakh), based on the Leningrad Codex — the oldest complete manuscript of the Hebrew Bible, dating to 1008 CE. This text serves as the basis for the Biblia Hebraica Stuttgartensia (BHS) and represents the authoritative Masoretic tradition.
The text includes full vocalization (niqqud) and cantillation marks (te'amim), rendered in the square Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/hebrew_masoretic_ot_hbo.yam-peleg__Hebrew-Gemma-11B-Instruct-details
Dataset Card for Evaluation run of yam-peleg/Hebrew-Gemma-11B-Instruct
Dataset automatically created during the evaluation run of model yam-peleg/Hebrew-Gemma-11B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yam-peleg__Hebrew-Gemma-11B-Instruct-details.HebrewBible_HapaxLegomenon
📖 NLP Research Course 097920: Hapax Legomenon Dataset
A dataset created for the NLP Research Course 097920, focusing on Hapax Legomenon — words that appear only once in the entire Hebrew Bible.
This dataset is designed to study LLM understanding of rare words in context, comparing a Hebrew-specific LLM (dicta-il/dictalm2.0-instruct) with a general-purpose LLM (gemini-2.0-flash).
🎯 Tasks
We designed three annotation tasks to evaluate LLM outputs:
1️⃣ Preference… See the full description on the dataset page: https://huggingface.co/datasets/wrom/HebrewBible_HapaxLegomenon.yam-peleg__Hebrew-Mistral-7B-200K-details
Dataset Card for Evaluation run of yam-peleg/Hebrew-Mistral-7B-200K
Dataset automatically created during the evaluation run of model yam-peleg/Hebrew-Mistral-7B-200K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yam-peleg__Hebrew-Mistral-7B-200K-details.hebrew-wiki-entitiestrc-hebrew-no-special-markers
