CoolFace
Datasetpublic

akiva-skolnik/hebrew-impairment-speech-v1

Hebrew Atypical Speech Dataset (Down Syndrome) Dataset name: hebrew-impairment-speech-v1 Speaker: A single Hebrew speaker with Down syndrome. Purpose: To advance research in personalized ASR for atypical speech, especially Hebrew. Dataset Summary This dataset contains 2307 Hebrew audio clips spoken by one individual with Down syndrome, accompanied by transcriptions. Speech includes dysarthric, stuttered, and non-standard pronunciation and grammar. The dataset aims… See the full description on the dataset page: https://huggingface.co/datasets/akiva-skolnik/hebrew-impairment-speech-v1.

sourceHugging Facecc-by-nc-4.0updated 9mo agoView on Hugging Face
1likes98downloads
Dataset Card

Hebrew Atypical Speech Dataset (Down Syndrome)

Dataset name: hebrew-impairment-speech-v1 Speaker: A single Hebrew speaker with Down syndrome. Purpose: To advance research in personalized ASR for atypical speech, especially Hebrew.

Dataset Summary

This dataset contains 2307 Hebrew audio clips spoken by one individual with Down syndrome, accompanied by transcriptions. Speech includes dysarthric, stuttered, and non-standard pronunciation and grammar.

The dataset aims to support research into:

  • —Personalized ASR
  • —Hebrew impaired-speech recognition
  • —Assistive communication systems

Approved for release by the speaker’s family for non-commercial use.

Data Structure

audio/folder/*.wav
dataset.jsonl  // {"file": "audio/0001.wav", "text": "…"}

Statistics

MetricValue
Clips2307
Total duration4655.7 sec (~1.29 hours)
Min0.286 sec
Max37.989 sec
Average2.02 sec
Median1.03 sec

Example: Whisper Fine-Tuning Notebook

See notebook.ipynb