datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval-whatsapp
Dataset Card for ivrit.ai Whatsapp Eval
Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed.
Dataset Details
Dataset Description
This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025.
The recordings are of volunteer native hebrew speakers using consumer devices in natural environments.
Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner.
Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.
