datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crowd-transcribe-v5
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
Full license: https://www.ivrit.ai/en/the-license/
FAQs: https://www.ivrit.ai/en/license-faqs/
VoxKnesset
VoxKnesset
Voice recordings of Israeli politicians from Knesset proceedings, annotated with
speaker age and demographic metadata.
Dataset Summary
Total hours (longitudinal subset): 2,307
Plenary sessions: ~1,550
Unique speakers: 393 Members of Knesset
Language: Hebrew
Recording years: 2009–2025 (16 years)
Maximum span per speaker: 15 years
Median span per speaker: 3.4 years
Speakers with >10 years coverage: 47 (12%)
Age range: 28–81 years
Split
Samples… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/VoxKnesset.jbdknesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audio-vadivrit.ai is a database of Hebrew audio and text content.
audio-base contains the raw, unprocessed sources.
audio-vad contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset.
audio-transcripts contains transcriptions for each snippet in the audio-vad dataset.
The audio-base dataset contains data from the following sources:
Geekonomy (Podcast, https://geekonomy.net)
HaCongress (Podcast, https://hacongress.podbean.com/)
Idan Eretz's… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-vad.hebrew-handwriting-ocr-benchmark
Hebrew Handwriting OCR Benchmark
A small, human-verified benchmark for OCR / handwritten text recognition (HTR) on
modern Hebrew handwriting: 225 gold lines across 10 pages, one page per
writer, drawn from the transcriptor.ivrit.ai
volunteer transcription corpus.
This is a test set. There is no train split, by design — it exists to be held
out. It is deliberately small and clean rather than large and noisy: every line
was transcribed by at least two volunteers independently and… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/hebrew-handwriting-ocr-benchmark.audio-baseivrit.ai is a database of Hebrew audio and text content.
audio-base contains the raw, unprocessed sources.
audio-vad contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset.
v1 data is generated using silero-vad's default parameters.
v2 data is generated using min_speech_duration_ms=2000 (milliseconds), and max_speech_duration_s=30 (seconds).
audio-transcripts contains transcriptions for each snippet in the audio-vad dataset.
You… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-base.eval-whatsapp
Dataset Card for ivrit.ai Whatsapp Eval
Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed.
Dataset Details
Dataset Description
This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025.
The recordings are of volunteer native hebrew speakers using consumer devices in natural environments.
Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner.
Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.knesset-plenums
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps.
We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts).
The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.whisper-training
Dataset Card for "whisper-training"
Note: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai.
More Information needed
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
Full license: https://www.ivrit.ai/en/the-license/
FAQs: https://www.ivrit.ai/en/license-faqs/
eval-d1Note: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai.
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
Full license: https://www.ivrit.ai/en/the-license/
FAQs: https://www.ivrit.ai/en/license-faqs/
brain-teasers_ENcrowd-recital-whisper-training
Dataset Card for ivrit.ai - Crowd Recital
Dataset Details
Dataset Description
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
- Full license: https://www.ivrit.ai/en/the-license/
- FAQs: https://www.ivrit.ai/en/license-faqs/
Dataset Structure
Data Fields
Each example in the dataset contains:
audio: An audio column containing:
bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.tofu-translit-ivrit_lat-lebnani-hepburn
TOFU — Hebrish / Arabizi / Romaji (Latin-script transliterations)
Three Latin-script transliteration arms of the TOFU fictitious-author unlearning
benchmark, built for "Script, Not Syntax: Transliteration as a Blind Spot in
Multilingual Unlearning" (Tsir Cohen, Rubinstein, Spira — Trustworthy Machine
Learning, Tel Aviv University, 2026).
Why this exists
TOFU (Maini et al., 2024) asks factual questions about invented authors, so it's
answerable only from what a… See the full description on the dataset page: https://huggingface.co/datasets/Aya168/tofu-translit-ivrit_lat-lebnani-hepburn.crowd-recital-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Recital - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~78h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.medmcqa-instructionaudio-transcriptsivrit.ai is a database of Hebrew audio and text content.
audio-base contains the raw, unprocessed sources.
audio-vad contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset.
audio-transcripts contains transcriptions for each snippet in the audio-vad dataset.
The audio-base dataset contains data from the following sources:
Geekonomy (Podcast, https://geekonomy.net)
HaCongress (Podcast, https://hacongress.podbean.com/)
Idan Eretz's… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-transcripts.crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.medmcqa-conversationaudio-labeled
Dataset Card for "audio-labeled"
More Information needed
crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.medmcqa-benchmarkIVR-pilot-benchmark
IntentSpec Benchmark — Data Supplement
This archive contains the benchmark data used to compute Intent Violation
Rate (IVR) in the paper: 49 tasks, each derived from a HumanEval problem and
extended with an ambiguous/gold prompt pair and a decomposed set of
executable constraints.
Files
spec_pairs.jsonl
The benchmark itself — one JSON object per line, one line per task. This is
the file consumed directly by the evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/IVR-pilot-benchmark.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.crowd-transcribe-v4Note: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai.
crowd-transcribe-v4
This is ivrit.ai's 4th crowd-sourced transcribed dataset release.
It contains over 250 hours of volunteer-transcribed data, randomly selected from our audio-vad dataset of over 10,000 hours.
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
Full license:… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-transcribe-v4.eval-forced-alignment
Hebrew Forced Alignment Evaluation Dataset
Human-verified, word-level time-aligned Hebrew speech clips.
To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was
built. The system lets labelers fix the transcript and align each spoken word to the audio,
down to 1ms precision (though annotators typically work at ~10ms granularity).
The audio samples were gathered by randomly sampling from several of ivrit-ai's larger,
published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.ivrit-manifestsivrit-manifests-textsivrit-recordingsjpress-demo
