datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish_parliamentary_data
Grand National Assembly Corpus of Türkiye (GNACT)
A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts.
Loading the dataset
from datasets import load_dataset
# Strategy 1: full session documents, all bodies (default)
ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.canadian-parliamentary-expenditures
Canadian House of Commons Parliamentary Expenditures Dataset
This dataset contains detailed expenditure records from the Canadian House of Commons, spanning from 2021 Q2 to 2025 Q4, with 1,219,648 total expenditure records across 450 parliament members.
Dataset Structure
parliamentary_data_hf/
├── data/
│ ├── train/ # Training split (2021-2024)
│ │ ├── expenditures-2021-q2.parquet
│ │ ├── expenditures-2021-q3.parquet
│ │ ├── ...… See the full description on the dataset page: https://huggingface.co/datasets/irf23/canadian-parliamentary-expenditures.Hellenic-greek-parliamentary-speech
HParl: Hellenic Parliamentary Speech Corpus
Dataset Description
Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.
Link to the original source: https://inventory.clarin.gr/corpus/1602
HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.sl-parliamentary-hansard-17-26
Dataset Card for Sri Lanka Parliamentary Hansard
Sri Lanka Parliamentary Hansard is a trilingual parliamentary speech corpus built from publicly available Hansard records of the Parliament of Sri Lanka. It contains Sinhala (සිංහල), Tamil (தமிழ்), English, and code-mixed speeches from 2017 to 2026, with speaker names, dates, and topic-modeling labels.
The dataset was created for the research paper "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", associated with… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/sl-parliamentary-hansard-17-26.luxembourgish-parliamentary-corpus
Luxembourgish Parliamentary Corpus (2023–2028)
A provenance-documented, speaker-attributed corpus of Luxembourg's parliamentary
proceedings, built from the official session reports (comptes rendus /
"D'Chamberblietchen") of the Chambre des Députés, legislature 2023–2028.
Luxembourgish (Lëtzebuergesch) is a documented low-resource language: the
Luxembourgish Wikipedia holds roughly 64,000 articles and most large language
models perform poorly in it for lack of training material.… See the full description on the dataset page: https://huggingface.co/datasets/Decima-Data/luxembourgish-parliamentary-corpus.africa-synth-governance-parliamentary-legislation-tracker-all
African Parliamentary Legislation Tracker | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-governance-parliamentary-legislation-tracker-all.Brazilian_Parliamentary_Expenses_Datasetsparliamentary_personas
🏛️ Dataset Card: Synthetic British Parliamentary Personas Benchmark
Dataset Summary
This dataset is a synthetic collection of 2,200 British parliamentary personas designed to study political behavior, rhetorical styles, and alignment properties of language models under structured role conditioning.
Each persona is grounded in one of 10 distinct parliamentary archetypes, capturing a wide spectrum of behaviors observed in legislative, party, and public-facing contexts… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/parliamentary_personas.Polish-Parliamentary-Speeches-Corpus
Dataset Card for Polish Parliamentary Speeches Corpus (PPSC)
Dataset Description
The Polish Parliamentary Speeches Corpus (PPSC) is a collection of official transcripts of parliamentary speeches made by Polish politicians. It was created to facilitate the modeling of political viewpoints in a low-resource language (Polish) using supervised fine-tuning. The dataset assigns speeches to binary ideological categories (Left-wing and Right-wing) based on the speakers'… See the full description on the dataset page: https://huggingface.co/datasets/PoliWings/Polish-Parliamentary-Speeches-Corpus.legal_os_clean_ssot_parliamentary_documentParliamentaryProceedingsparliamentarygeorgian-parliamentary-elections-2024-protocolsBritish_Parliamentary_Motion_Questionsnation-parliamentary-promptsparliamentary-debate-casesocpsg-silver-standard-parliamentary-speeches
OCPSG Silver Standard Parliamentary Speeches
Release: v1.0.0-rc.2 "Lively Monolith"
Dataset Summary
This repository contains the final silver-standard parliamentary speech dataset produced in the Oxford Computational Political Science Group benchmarking workflow. The dataset is intended for multilingual policy agenda classification and downstream benchmarking and fine-tuning tasks.
The release includes country-level train/validation/test splits. In the original workflow… See the full description on the dataset page: https://huggingface.co/datasets/OCPSG-Benchmarking-LLMs/ocpsg-silver-standard-parliamentary-speeches.nation-parliamentary-promptsnation-parliamentary-sft-actions
