ivr
Datasets
All datasets matching “ivr”knesset-committees
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols.
We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.ShapeNetSDF
ShapeNetSDF
Signed Distance Field (SDF) point samples derived from
ShapeNet Core, for training and evaluating implicit
neural representations / neural fields on 3D shapes.
This dataset is shared as part of the CVPR 2026 paper Weight Space Representation Learning via Neural Field Adaptaion.
Code for producing this dataset is shared in the wsr.pytorch neural-field codebase.
Each shape is converted into a watertight manifold, normalized into the unit
cube [-1, 1]³, and sampled… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-IVRL/ShapeNetSDF.audio-v2This dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It wa released on April 20th, 2025.
You can find the full list of sources in this dataset under the dataset's sources.txt.
Paper: https://arxiv.org/abs/2307.08720
If you use our datasets, the following quote is preferable:
@misc{marmor2023ivritai,
title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development},
author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.crowd-transcribe-v5
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
Full license: https://www.ivrit.ai/en/the-license/
FAQs: https://www.ivrit.ai/en/license-faqs/
VoxKnesset
VoxKnesset
Voice recordings of Israeli politicians from Knesset proceedings, annotated with
speaker age and demographic metadata.
Dataset Summary
Total hours (longitudinal subset): 2,307
Plenary sessions: ~1,550
Unique speakers: 393 Members of Knesset
Language: Hebrew
Recording years: 2009–2025 (16 years)
Maximum span per speaker: 15 years
Median span per speaker: 3.4 years
Speakers with >10 years coverage: 47 (12%)
Age range: 28–81 years
Split
Samples… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/VoxKnesset.audio-v2-transcripts
Overview
This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on May 18th, 2025.
You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt.
All files were transcribed using the process.py pipeline, performing:
Frame-level VAD
Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.
