datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hebrew_VAD_lexicon
Hebrew VAD Lexicon
The Hebrew VAD Lexicon is an enhanced version of an automatically translated affective lexicon, originally derived from the English VAD lexicon created by Mohammad (2018) .
It provides valence, arousal, and dominance (VAD) scores for Hebrew words.
The lexicon was carefully curated by manually reviewing and correcting the automatic translations and enriching the dataset with additional linguistic information.
There are two versions of the dataset:
English-Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/GiliGold/Hebrew_VAD_lexicon.hebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.trc-hebrewHebrew-Speech-Dataset
🎧 Hausa Speech Dataset
The Hausa Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that require diverse audio data and reliable voice data for multilingual model training. It contains 160 hours of recordings across 849 files, stored in MP3 and WAV formats, with a total size of 270 MB. This carefully engineered audio dataset ensures balanced representation with 48% female and 52% male speakers, covering an age range from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hebrew-Speech-Dataset.Hebrew_sentimenthebrew_sa
Sentiment Analysis Data for the Hebrew Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Amram et al. (2018).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{amram-etal-2018-representations,
title = "Representations and Architectures in Neural Sentiment Analysis for Morphologically Rich Languages: A Case Study from {M}odern {H}ebrew"… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/hebrew_sa.hebrew-trc-special-markershebrew-ocr-datasetRoleplay-Hebrew
RolePlay-Hebrew
Roleplay-Hebrew Dataset is a dataset for roleplaying in the Hebrew language for the Large Language Model.
The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Hebrew.HebrewBible_HapaxLegomenon
📖 NLP Research Course 097920: Hapax Legomenon Dataset
A dataset created for the NLP Research Course 097920, focusing on Hapax Legomenon — words that appear only once in the entire Hebrew Bible.
This dataset is designed to study LLM understanding of rare words in context, comparing a Hebrew-specific LLM (dicta-il/dictalm2.0-instruct) with a general-purpose LLM (gemini-2.0-flash).
🎯 Tasks
We designed three annotation tasks to evaluate LLM outputs:
1️⃣ Preference… See the full description on the dataset page: https://huggingface.co/datasets/wrom/HebrewBible_HapaxLegomenon.cleaned-hebrew-lexicon
Hebrew Cleaned Lexicon
מילון תדירויות של קורפוס טקסט עברי מנוקה.
מילים ייחודיות: 1,547,130
סה"כ מופעים: 294,312,712
נוקה ממרכאות בודדות ותווי רעש
קבצים
lexicon_cleaned.pkl – מילון בפורמט Pickle (Counter)
lexicon_cleaned.csv – מילון בפורמט CSV (word, frequency)
ai-questions-hebrewtrc-hebrew-no-special-markersHEBREW-MIL-CLEANHebrewLyricsUsed the great data set Norod78/HebrewLyricsDataet
Hebrew_lyrics_40kUsed the great data set Norod78/HebrewLyricsDataet
