datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knesset-committees
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols.
We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.audio-v2This dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It wa released on April 20th, 2025.
You can find the full list of sources in this dataset under the dataset's sources.txt.
Paper: https://arxiv.org/abs/2307.08720
If you use our datasets, the following quote is preferable:
@misc{marmor2023ivritai,
title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development},
author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.audio-v2-transcripts
Overview
This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on May 18th, 2025.
You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt.
All files were transcribed using the process.py pipeline, performing:
Frame-level VAD
Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.crowd-recital
About
This dataset was created by crowd-sourced recording sessions in Hebrew as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital.eval-whatsapp
Dataset Card for ivrit.ai Whatsapp Eval
Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed.
Dataset Details
Dataset Description
This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025.
The recordings are of volunteer native hebrew speakers using consumer devices in natural environments.
Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner.
Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.knesset-plenums
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps.
We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts).
The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.crowd-recital-whisper-training
Dataset Card for ivrit.ai - Crowd Recital
Dataset Details
Dataset Description
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
- Full license: https://www.ivrit.ai/en/the-license/
- FAQs: https://www.ivrit.ai/en/license-faqs/
Dataset Structure
Data Fields
Each example in the dataset contains:
audio: An audio column containing:
bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.crowd-recital-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Recital - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~78h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.audio-v2-opusThis dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on April 20th, 2025.
You can find the full list of sources in this dataset under the dataset's sources.txt.
Paper: https://arxiv.org/abs/2307.08720
If you use our datasets, the following quote is preferable:
@misc{marmor2023ivritai,
title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development},
author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-opus.crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.crowd-transcribe-v4Note: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai.
crowd-transcribe-v4
This is ivrit.ai's 4th crowd-sourced transcribed dataset release.
It contains over 250 hours of volunteer-transcribed data, randomly selected from our audio-vad dataset of over 10,000 hours.
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
Full license:… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-transcribe-v4.eval-forced-alignment
Hebrew Forced Alignment Evaluation Dataset
Human-verified, word-level time-aligned Hebrew speech clips.
To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was
built. The system lets labelers fix the transcript and align each spoken word to the audio,
down to 1ms precision (though annotators typically work at ~10ms granularity).
The audio samples were gathered by randomly sampling from several of ivrit-ai's larger,
published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.
