Knesset
Datasets
All datasets matching “Knesset”KnessetCorpus
The Knesset (Israeli Parliament) Proceedings Corpus
💻 [Github Repo] •
📃 [Paper] •
📊 [ES kibana dashboard]
Dataset Description
An annotated corpus of Hebrew parliamentary proceedings containing over 35 million sentences from all the (plenary and committee) protocols held in the Israeli parliament
from 1992 to 2024.Sentences are annotated with various levels of linguistic information, including part-of-speech tags, morphological features, dependency… See the full description on the dataset page: https://huggingface.co/datasets/HaifaCLGroup/KnessetCorpus.knesset-committees
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols.
We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.knesset-committees-chunksknesset-data
🇮🇱 נתוני הכנסת הפתוחה
מאגר נתונים פתוח של הכנסת — ישירות מה-API הרשמי, בפורמט JSONL מחולק לקבצים.
מקור: OData API של הכנסת
רישיון: CC-BY-SA-4.0
תחזוקה: זה עלינו
כל קובץ JSONL מכיל שורה אחת לכל רשומה, ממוין לפי Id.
🇮🇱 Knesset Open Data
Open dataset of the Israeli Knesset (parliament) — sourced directly from the official API, stored as partitioned JSONL files.
Source: Knesset OData API
License: CC-BY-SA-4.0
Maintained by: ZeAlenu
📊 Tables (44 total… See the full description on the dataset page: https://huggingface.co/datasets/ZeAlenu/knesset-data.knesset_melia_asr_pocknesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.
