datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.knesset-committees-speakers
Knesset Committees Speakers
An index that attaches a verified Knesset member identity, and through it
demographics, to the committee audio in
ivrit-ai/knesset-committees.
No audio is included. Each row names a span (session, start, end) of that
dataset's audio.m4a; filename follows the VoxKnesset convention
{speaker_id}_{session}_{start_ms}_{end_ms}.wav so the same tooling applies.
speaker_id is the Knesset's official PersonID -- the same id space as the
Knesset Corpus and… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/knesset-committees-speakers.knesset-plenums
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps.
We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts).
The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.
