datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hebrew_speech_coursera
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
{'audio': {'path':… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_coursera.hebrew_speech_campus
Data Description
Hebrew Speech Recognition dataset from Campus IL.
Data was scraped from the Campus website, which contains video lectures from various courses in Hebrew.Then subtitles were extracted from the videos and aligned with the audio.Subtitles that are not on Hebrew were removed (WIP: need to remove non-Hebrew audio as well, e.g. using simple classifier).Samples with duration less than 3 second were removed.Total duration of the dataset is 152 hours.Outliers in terms… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_campus.hebrew_speech_kan
Dataset Card for Dataset Name
Dataset Summary
Hebrew Dataset for ASR
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
{'audio': {'path': '/root/.cache/huggingface/datasets/downloads/extracted/8ce7402f6482c6053251d7f3000eec88668c994beb48b7ca7352e77ef810a0b6/train/e429593fede945c185897e378a5839f4198.wav',
'array': array([-0.00265503, -0.0018158… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_kan.hebrew-impairment-speech-v1
Hebrew Atypical Speech Dataset (Down Syndrome)
Dataset name: hebrew-impairment-speech-v1
Speaker: A single Hebrew speaker with Down syndrome.
Purpose: To advance research in personalized ASR for atypical speech, especially Hebrew.
Dataset Summary
This dataset contains 2307 Hebrew audio clips spoken by one individual with Down syndrome, accompanied by transcriptions.
Speech includes dysarthric, stuttered, and non-standard pronunciation and grammar.
The dataset aims to… See the full description on the dataset page: https://huggingface.co/datasets/akiva-skolnik/hebrew-impairment-speech-v1.English-Hebrew-Mixed-Sentences
English-Hebrew Mixed Sentences Dataset
A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel.
Overview
This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences.
Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.hebrew-speech-datasetqwen3-asr-hebrew-100kHebrew-Speech-Dataset
🎧 Hausa Speech Dataset
The Hausa Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that require diverse audio data and reliable voice data for multilingual model training. It contains 160 hours of recordings across 849 files, stored in MP3 and WAV formats, with a total size of 270 MB. This carefully engineered audio dataset ensures balanced representation with 48% female and 52% male speakers, covering an age range from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hebrew-Speech-Dataset.hebrew_speech_kan_nikudHebrew-talkhebrew_keyword_spotfine-tune-hebrew-dataset
Dataset Card for "fine-tune-hebrew-dataset"
More Information needed
fine-tune-hebrew-dataset-2
Dataset Card for "fine-tune-hebrew-dataset-2"
More Information needed
hebrew_kanhebrew_kan_sentence0hebrew_keywords1hebrew_keywords2hebrew_keywordsHebrew-talkhebrew_kan_sentence50000hebrew_kan_sentence90000hebrew_kan_sentence100000YodaLingua-Hebrew
YodaLingua-Hebrew
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Hebrew portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
45,877 audio–transcription pairs
Total duration
120.3 hours
Speakers
2,226 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Hebrew.hebrew_kan_sentence20000hebrew_kan_sentence60000hebrew_kan_sentence120000hebrew_kan_sentence130000hebrew_kan_sentence10000hebrew_kan_sentence70000hebrew_kan_sentence30000
