CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /hebrew_this_world Dataset Card for HebrewSentiment Dataset Summary HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license. Data Annotation: Supported Tasks and Leaderboards Language modeling Languages Hebrew Dataset Structure csv file with "," delimeter Data Instances Sample: { "issue_num": 637, "page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.imagetext-generation1K<n<10K1 likes152 downloads2y agoHugging Face02shunyalabs /hebrew-speech-datasetaudio1K<n<10K0 likes54 downloads1y agoHugging Face03MitchMitchon /hebrew-translated-retrieval-datasetstext100K<n<1M0 likes54 downloads4mo agoHugging Face04Dddixyy /latin-greek-hebrew-english-dataset Proverbs: Ancient Languages Set This repository contains a collection of 2,000 short phrases translated into three ancient languages: Ancient Latin, Ancient Greek, Biblical Hebrew, and English. The phrases cover a wide variety of contexts, providing insight into the linguistic, cultural, and philosophical landscapes of these ancient civilizations. Overview The "Proverbs: Ancient Languages Set" is a resource designed to help individuals explore and understand ancient… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin-greek-hebrew-english-dataset.texttranslation1K<n<10K1 likes40 downloads2y agoHugging Face05youssefkhalil320 /hebrew-ocr-doctags-dataset_v2image10K<n<100K0 likes32 downloads1y agoHugging Face06Speech-data /Hebrew-Speech-Dataset 🎧 Hausa Speech Dataset The Hausa Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that require diverse audio data and reliable voice data for multilingual model training. It contains 160 hours of recordings across 849 files, stored in MP3 and WAV formats, with a total size of 270 MB. This carefully engineered audio dataset ensures balanced representation with 48% female and 52% male speakers, covering an age range from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hebrew-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes30 downloads6mo agoHugging Face07sabuhi1997 /fine-tune-hebrew-dataset-2 Dataset Card for "fine-tune-hebrew-dataset-2" More Information needed audion<1K0 likes14 downloads3y agoHugging Face08mbole /hebrew-ocr-datasettextn<1K0 likes13 downloads2y agoHugging Face09HebArabNlpProject /Hebrew-Paraphrase-DatasetHebrew Paraphrase Dataset This repository contains a high-quality paraphrase dataset in Hebrew, consisting of 9785 instances. The dataset includes both paragraph-level (75%) and sentence-level (25%) paraphrases generated with the help of a large language model. Among these, 300 instances have been manually validated as gold standard examples. What Is a Paraphrase? A paraphrase is a restatement of a text using different words and structures while preserving the original meaning. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HebArabNlpProject/Hebrew-Paraphrase-Dataset.text1K<n<10K3 likes11 downloads2y agoHugging Face10sivan22 /hebrew-words-dataset Dataset Card for "hebrew-words-dataset" More Information needed imagen<1K0 likes8 downloads3y agoHugging Face11mbole /hebrew-tzfira-dataset Ha-Tsfira OCR and POS-Tagged Dataset Dataset Summary This dataset contains OCR-processed and POS-tagged text from Ha-Tsfira (הצפירה), a Hebrew-language newspaper published in Poland from 1862 and then from 1874 to 1931. The dataset includes 50 newspaper issues that have been digitized, cleaned, and linguistically annotated. Languages Hebrew (he) Dataset Structure DatasetDict({ train: Dataset({ features: ['id', 'ocr_text', 'cleaned_text'… See the full description on the dataset page: https://huggingface.co/datasets/mbole/hebrew-tzfira-dataset.texttext-classificationn<1K0 likes8 downloads2y agoHugging Face12youssefkhalil320 /hebrew-ocr-doctags-datasetimage1K<n<10K0 likes7 downloads1y agoHugging Face13MitchMitchon /hebrew_private_datasetgatedtext10K<n<100K1 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.