sl-parliamentary-nlp/sl-parliamentary-hansard-17-26
Dataset Card for Sri Lanka Parliamentary Hansard Sri Lanka Parliamentary Hansard is a trilingual parliamentary speech corpus built from publicly available Hansard records of the Parliament of Sri Lanka. It contains Sinhala (සිංහල), Tamil (தமிழ்), English, and code-mixed speeches from 2017 to 2026, with speaker names, dates, and topic-modeling labels. The dataset was created for the research paper "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", associated with… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/sl-parliamentary-hansard-17-26.
Dataset Card for Sri Lanka Parliamentary Hansard
Sri Lanka Parliamentary Hansard is a trilingual parliamentary speech corpus built from publicly available Hansard records of the Parliament of Sri Lanka. It contains Sinhala (සිංහල), Tamil (தமிழ்), English, and code-mixed speeches from 2017 to 2026, with speaker names, dates, and topic-modeling labels.
The dataset was created for the research paper "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", associated with MERCon 2026, and supports work on low-resource multilingual NLP, parliamentary discourse analysis, and unsupervised topic modeling.
This is a research dataset and is not an official account, publication, or endorsement of the Parliament of Sri Lanka.
Paper
This dataset accompanies the following paper:
[Trilingual Topic Modeling of Sri Lankan Parliamentary Debates](https://arxiv.org/abs/2608.20365)
Dataset Details
Dataset Description
- Shared by: Sri Lanka Parliamentary NLP
- Language(s): Sinhala, Tamil, English, and code-mixed speech
- License: Apache 2.0
- Rows in Hugging Face viewer: 19,699
- Topic-labeled speeches: 19,553
- Splits:
trainandtestusing a deterministic 90/10 convenience split - Time period: 2017-2026
- Data format: Parquet
- Task category: Text analysis, topic modeling, parliamentary discourse analysis
Each row represents a speech-level Hansard segment. The topic labels were generated using a multilingual topic-modeling pipeline based on BGE-M3 embeddings, UMAP dimensionality reduction, HDBSCAN micro-topic clustering, and macro-topic aggregation.
Featured Resources
- Primary Hansard source: https://www.parliament.lk/en/business-of-parliament/hansards
- Primary source: Publicly available official Hansard PDFs from the Parliament of Sri Lanka archive: https://www.parliament.lk/en/business-of-parliament/hansards
Uses
Direct Use
This dataset is intended for:
- Multilingual NLP research for Sinhala, Tamil, English, and code-mixed text
- Parliamentary discourse analysis
- Topic modeling and clustering research
- Low-resource language benchmarking
- Political science research on agenda setting and temporal attention
- Speaker-level and event-level analysis of parliamentary debate
Out-of-Scope Use
This dataset should not be used as:
- An official parliamentary record
- A source for legal, electoral, or policy decisions without manual verification against the original Hansard documents
- A ground-truth political stance or sentiment dataset
- A complete representation of every parliamentary utterance without accounting for extraction and filtering limitations
Dataset Structure
How to Get Started
Install the Hugging Face datasets library:
pip install datasets pandasLoad the dataset with the default configuration:
from datasets import load_dataset
dataset = load_dataset(
"sl-parliamentary-nlp/sl-parliamentary-hansard-17-26",
"default",
download_mode="force_redownload",
)
print(dataset)Convert to pandas:
train_df = dataset["train"].to_pandas()
test_df = dataset["test"].to_pandas()
df = train_df
print(df[["date", "speaker", "macro_topic", "text"]].head())Example: count speeches by macro-topic:
topic_counts = (
df["macro_topic"]
.fillna("Unassigned")
.value_counts()
)
print(topic_counts.head(10))Corpus Construction
The corpus was constructed from official Hansard PDFs available through the Sri Lankan parliamentary archive.
The processing pipeline included:
- Scraping publicly available Hansard PDFs.
- Extracting text with Google Gemini, selected because OCR struggles with the dual-column trilingual layout and historical Sinhala/Tamil font encodings.
- Segmenting extracted text into speech-level units.
- Normalizing speaker names across Sinhala, Tamil, and English variants.
- Removing structural headers, footers, and short non-substantive fragments.
- Assigning topic labels through the multilingual topic-modeling pipeline.
Topic Modeling Method
The topic-modeling pipeline used:
- Embeddings: BGE-M3 dense multilingual embeddings
- Dimensionality reduction: UMAP
- Micro-topic clustering: HDBSCAN
- Macro-topic aggregation: empirical dendrogram cut over micro-topic centroids
- Noise handling: HDBSCAN label
-1for procedural or low-density speeches
The final topic-modeling run produced:
Evaluation
The project includes a manually labeled 300-speech evaluation sample and reviewer-facing evaluation artifacts in the companion code repository.
Internal clustering benchmarks reported for the full corpus include:
Limitations
- Extraction noise: The source PDFs contain complex formatting, multilingual text, and historical font encodings. LLM-based extraction improves coverage but may still introduce errors.
- Speaker normalization: Names are normalized heuristically, so rare or ambiguous speaker variants may remain unmerged.
- Topic labels are unsupervised:
micro_topic_clusterandmacro_topicare model-generated analytical labels, not official parliamentary categories. - Noise labels are real text: Speeches labeled
-1are not invalid records; they are simply not assigned to a dense HDBSCAN topic region. - Temporal coverage is uneven: Parliamentary sitting frequency differs by year, so raw counts should be normalized before making year-to-year claims.
- Political text can be sensitive: Users should manually verify outputs before using the dataset in high-stakes political, journalistic, legal, or policy contexts.
Bias, Risks, and Considerations
The dataset reflects the language, framing, and political priorities present in parliamentary debate. It may contain partisan language, allegations, sensitive political claims, and socially biased statements made by speakers. Topic models may amplify or obscure these patterns depending on preprocessing choices and clustering parameters.
Researchers should treat the dataset as a computational research resource, not as a neutral summary of Sri Lankan politics.
Citation
If you use this dataset, please cite:
@misc{dhanapala2026trilingual,
title = {Trilingual Topic Modeling of Sri Lankan Parliamentary Debates},
author = {Himath Dhanapala and Haren Daishika and Himandhi Kuruppu and Sithija Seneviratne and Ashini Kavindya and Patalee Narasinghe and Sandeepa Weerasekara and Nisansa de Silva and Sandareka Wickramanayake},
year = {2026},
eprint = {2608.20365},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2608.20365}
}Dataset Card Authors
This dataset card was prepared by the Sri Lanka Parliamentary NLP project team based on the companion research repository and dataset publishing artifacts.
License
The dataset is released under the Apache 2.0 license. The underlying Hansard text is derived from publicly available parliamentary records of the Parliament of Sri Lanka.
