CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HaifaCLGroup /KnessetCorpus The Knesset (Israeli Parliament) Proceedings Corpus 💻 [Github Repo] • 📃 [Paper] • 📊 [ES kibana dashboard] Dataset Description An annotated corpus of Hebrew parliamentary proceedings containing over 35 million sentences from all the (plenary and committee) protocols held in the Israeli parliament from 1992 to 2024.Sentences are annotated with various levels of linguistic information, including part-of-speech tags, morphological features, dependency… See the full description on the dataset page: https://huggingface.co/datasets/HaifaCLGroup/KnessetCorpus.text-classification10M<n<100M6 likes15k downloads7mo agoHugging Face02ivrit-ai /knesset-committeesgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols. We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.automatic-speech-recognition3 likes4.5k downloads4mo agoHugging Face03Hadasy /knesset-committees-chunkstabular1M<n<10M0 likes1k downloads14d agoHugging Face04ZeAlenu /knesset-data 🇮🇱 נתוני הכנסת הפתוחה מאגר נתונים פתוח של הכנסת — ישירות מה-API הרשמי, בפורמט JSONL מחולק לקבצים. מקור: OData API של הכנסת רישיון: CC-BY-SA-4.0 תחזוקה: זה עלינו כל קובץ JSONL מכיל שורה אחת לכל רשומה, ממוין לפי Id. 🇮🇱 Knesset Open Data Open dataset of the Israeli Knesset (parliament) — sourced directly from the official API, stored as partitioned JSONL files. Source: Knesset OData API License: CC-BY-SA-4.0 Maintained by: ZeAlenu 📊 Tables (44 total… See the full description on the dataset page: https://huggingface.co/datasets/ZeAlenu/knesset-data.1M<n<10M0 likes790 downloads8mo agoHugging Face05yoad /knesset_melia_asr_pocaudion<1K0 likes574 downloads2y agoHugging Face06ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes436 downloads10mo agoHugging Face07Dolevabudi /knesset-committees-speakers Knesset Committees Speakers An index that attaches a verified Knesset member identity, and through it demographics, to the committee audio in ivrit-ai/knesset-committees. No audio is included. Each row names a span (session, start, end) of that dataset's audio.m4a; filename follows the VoxKnesset convention {speaker_id}_{session}_{start_ms}_{end_ms}.wav so the same tooling applies. speaker_id is the Knesset's official PersonID -- the same id space as the Knesset Corpus and… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/knesset-committees-speakers.textautomatic-speech-recognition1M<n<10M0 likes178 downloads16d agoHugging Face08GiliGold /KnessetCorpusFor The Knesset Corpus: [The Knesset Corpus] text-classification10M<n<100M0 likes124 downloads7mo agoHugging Face09ivrit-ai /knesset-plenumsgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps. We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts). The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.audioautomatic-speech-recognition1K<n<10K3 likes103 downloads10mo agoHugging Face10notmax123 /Knesset-VOX-IPA Knesset VOX IPA Hebrew speech dataset derived from Knesset (Israeli Parliament) plenary sessions, enriched with IPA (International Phonetic Alphabet) phoneme transcriptions. Inspired by the methodology of arxiv:2603.01270. Dataset Description Long-form Knesset recordings were split into chunks of up to 15 seconds. Each chunk was transcribed to Hebrew text and then processed for IPA phoneme extraction from audio. Each sample pairs a WAV audio chunk with: The original… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/Knesset-VOX-IPA.automatic-speech-recognition10K<n<100K0 likes80 downloads6mo agoHugging Face11GiliGold /VAD_KnessetCorpus VAD_KnessetCorpus This dataset extends the original Knesset Corpus by adding VAD (Valence, Arousal, Dominance) annotations to committee sentences. The VAD scores were generated using the VAD binomial regression models, which were developed specifically for VAD prediction, on embeddings produced by the Knesset-multi-e5-large. Additionally, a small subset of 120 Knesset sentences has been manually annotated by three annotators for VAD scores, and the file is available in… See the full description on the dataset page: https://huggingface.co/datasets/GiliGold/VAD_KnessetCorpus.text-classification10M<n<100M0 likes34 downloads7mo agoHugging Face12imvladikon /knesset_meetings_corpus Dataset Card Dataset Summary An example of a sample: { "text": <text content of given document>, "path": <file path to docx> } Dataset usage Available "kneset16","kneset17","knesset_tagged" configurations And only train set. train_ds = load_dataset("imvladikon/knesset_meetings_corpus", "kneset16", split="train") The Knesset Meetings Corpus 2004-2005 is made up of two components: Raw texts - 282 files made up of 867,725 lines together. These can be downloaded in… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/knesset_meetings_corpus.texttext-generationn<1K1 likes31 downloads4y agoHugging Face13nahmanat /knesset-data 🇮🇱 נתוני הכנסת הפתוחה מאגר נתונים פתוח של הכנסת — ישירות מה-API הרשמי, בפורמט JSONL מחולק לקבצים. מקור: OData API של הכנסת רישיון: CC-BY-SA-4.0 תחזוקה: זה עלינו כל קובץ JSONL מכיל שורה אחת לכל רשומה, ממוין לפי Id. 🇮🇱 Knesset Open Data Open dataset of the Israeli Knesset (parliament) — sourced directly from the official API, stored as partitioned JSONL files. Source: Knesset OData API License: CC-BY-SA-4.0 Maintained by: ZeAlenu 📊 Tables (44… See the full description on the dataset page: https://huggingface.co/datasets/nahmanat/knesset-data.1M<n<10M0 likes31 downloads3d agoHugging Face14GiliGold /Knesset_check_worthinessThis dataset extends the Knesset Corpus by annotating it for Check Worthiness. The possible values per sentence are: worth checking, not worth checking , or not a factual proposition The annotations were generated by knesset-dicta-checkworthiness. ArXiv paper Citation: @InProceedings{goldin-EtAl:2025:RANLP, author = {Goldin, Gili and Wigderson, Shira and Rabinovich, Ella and Wintner, Shuly}, title = {An Annotation Scheme for Factuality and Its Application to Parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/GiliGold/Knesset_check_worthiness.0 likes28 downloads7mo agoHugging Face15benderrodriguez /knesset-committees-prepped0 likes28 downloads3mo agoHugging Face16Wissotsky /KnessetNews Dataset Card for Knesset News Hebrew Press Releases From the Knesset(Israeli Parliament) Dataset Details Dataset Description The dataset contains all the official knesset hebrew press releases up until 21-07-2025 Curated by: [@Wissotsky] Language: [Hebrew] Dataset Sources Knesset Press Releases Dataset Structure SP_Id (string): Unique identifier for each news article from the Knesset system Title (string): Article headline/subject Date… See the full description on the dataset page: https://huggingface.co/datasets/Wissotsky/KnessetNews.texttext-generation10K<n<100K0 likes14 downloads1y agoHugging Face17Tyl3rDrden /ivrit-knesset-shards-v50 likes14 downloads6mo agoHugging Face18malper /knesset-vox-he-ipatext10K<n<100K0 likes8 downloads6mo agoHugging Face19Tyl3rDrden /ivrit-knesset-shards0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.