datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hebrew_this_world
Dataset Card for HebrewSentiment
Dataset Summary
HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license.
Data Annotation:
Supported Tasks and Leaderboards
Language modeling
Languages
Hebrew
Dataset Structure
csv file with "," delimeter
Data Instances
Sample:
{
"issue_num": 637,
"page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.hebrew-speech-datasethebrew-translated-retrieval-datasetslatin-greek-hebrew-english-dataset
Proverbs: Ancient Languages Set
This repository contains a collection of 2,000 short phrases translated into three ancient languages: Ancient Latin, Ancient Greek, Biblical Hebrew, and English. The phrases cover a wide variety of contexts, providing insight into the linguistic, cultural, and philosophical landscapes of these ancient civilizations.
Overview
The "Proverbs: Ancient Languages Set" is a resource designed to help individuals explore and understand ancient… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin-greek-hebrew-english-dataset.hebrew-ocr-doctags-dataset_v2Hebrew-Speech-Dataset
🎧 Hausa Speech Dataset
The Hausa Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that require diverse audio data and reliable voice data for multilingual model training. It contains 160 hours of recordings across 849 files, stored in MP3 and WAV formats, with a total size of 270 MB. This carefully engineered audio dataset ensures balanced representation with 48% female and 52% male speakers, covering an age range from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hebrew-Speech-Dataset.fine-tune-hebrew-dataset-2
Dataset Card for "fine-tune-hebrew-dataset-2"
More Information needed
hebrew-ocr-datasetHebrew-Paraphrase-DatasetHebrew Paraphrase Dataset
This repository contains a high-quality paraphrase dataset in Hebrew, consisting of 9785 instances.
The dataset includes both paragraph-level (75%) and sentence-level (25%) paraphrases generated with the help of a large language model.
Among these, 300 instances have been manually validated as gold standard examples.
What Is a Paraphrase?
A paraphrase is a restatement of a text using different words and structures while preserving the original meaning.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HebArabNlpProject/Hebrew-Paraphrase-Dataset.hebrew-words-dataset
Dataset Card for "hebrew-words-dataset"
More Information needed
hebrew-tzfira-dataset
Ha-Tsfira OCR and POS-Tagged Dataset
Dataset Summary
This dataset contains OCR-processed and POS-tagged text from Ha-Tsfira (הצפירה), a Hebrew-language newspaper published in Poland from 1862 and then from 1874 to 1931. The dataset includes 50 newspaper issues that have been digitized, cleaned, and linguistically annotated.
Languages
Hebrew (he)
Dataset Structure
DatasetDict({
train: Dataset({
features: ['id', 'ocr_text', 'cleaned_text'… See the full description on the dataset page: https://huggingface.co/datasets/mbole/hebrew-tzfira-dataset.hebrew-ocr-doctags-datasethebrew_private_dataset
