datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knesset-committees
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols.
We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.knesset-committees-speakers
Knesset Committees Speakers
An index that attaches a verified Knesset member identity, and through it
demographics, to the committee audio in
ivrit-ai/knesset-committees.
No audio is included. Each row names a span (session, start, end) of that
dataset's audio.m4a; filename follows the VoxKnesset convention
{speaker_id}_{session}_{start_ms}_{end_ms}.wav so the same tooling applies.
speaker_id is the Knesset's official PersonID -- the same id space as the
Knesset Corpus and… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/knesset-committees-speakers.
