CoolFace
Datasetpublic

sl-parliamentary-nlp/sl-parliamentary-hansard-17-26

Dataset Card for Sri Lanka Parliamentary Hansard Sri Lanka Parliamentary Hansard is a trilingual parliamentary speech corpus built from publicly available Hansard records of the Parliament of Sri Lanka. It contains Sinhala (සිංහල), Tamil (தமிழ்), English, and code-mixed speeches from 2017 to 2026, with speaker names, dates, and topic-modeling labels. The dataset was created for the research paper "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", associated with… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/sl-parliamentary-hansard-17-26.

sourceHugging Faceapache-2.0updated 29d agoView on Hugging Face
2likes81downloads
Dataset Card

Dataset Card for Sri Lanka Parliamentary Hansard

Sri Lanka Parliamentary Hansard is a trilingual parliamentary speech corpus built from publicly available Hansard records of the Parliament of Sri Lanka. It contains Sinhala (සිංහල), Tamil (தமிழ்), English, and code-mixed speeches from 2017 to 2026, with speaker names, dates, and topic-modeling labels.

The dataset was created for the research paper "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", associated with MERCon 2026, and supports work on low-resource multilingual NLP, parliamentary discourse analysis, and unsupervised topic modeling.

This is a research dataset and is not an official account, publication, or endorsement of the Parliament of Sri Lanka.

Paper

This dataset accompanies the following paper:

[Trilingual Topic Modeling of Sri Lankan Parliamentary Debates](https://arxiv.org/abs/2608.20365)

Dataset Details

Dataset Description

  • —Shared by: Sri Lanka Parliamentary NLP
  • —Language(s): Sinhala, Tamil, English, and code-mixed speech
  • —License: Apache 2.0
  • —Rows in Hugging Face viewer: 19,699
  • —Topic-labeled speeches: 19,553
  • —Splits: train and test using a deterministic 90/10 convenience split
  • —Time period: 2017-2026
  • —Data format: Parquet
  • —Task category: Text analysis, topic modeling, parliamentary discourse analysis

Each row represents a speech-level Hansard segment. The topic labels were generated using a multilingual topic-modeling pipeline based on BGE-M3 embeddings, UMAP dimensionality reduction, HDBSCAN micro-topic clustering, and macro-topic aggregation.

Featured Resources

  • —Primary Hansard source: https://www.parliament.lk/en/business-of-parliament/hansards
  • —Primary source: Publicly available official Hansard PDFs from the Parliament of Sri Lanka archive: https://www.parliament.lk/en/business-of-parliament/hansards

Uses

Direct Use

This dataset is intended for:

  • —Multilingual NLP research for Sinhala, Tamil, English, and code-mixed text
  • —Parliamentary discourse analysis
  • —Topic modeling and clustering research
  • —Low-resource language benchmarking
  • —Political science research on agenda setting and temporal attention
  • —Speaker-level and event-level analysis of parliamentary debate

Out-of-Scope Use

This dataset should not be used as:

  • —An official parliamentary record
  • —A source for legal, electoral, or policy decisions without manual verification against the original Hansard documents
  • —A ground-truth political stance or sentiment dataset
  • —A complete representation of every parliamentary utterance without accounting for extraction and filtering limitations

Dataset Structure

ColumnTypeDescription
speech_idstringUnique speech identifier, such as SP_00000
datestringParliamentary session date in YYYY-MM-DD format
speakerstringNormalized speaker name, which may appear in Sinhala, Tamil, or English
textstringFull speech text in the original language, including possible code-mixing
yearintegerYear extracted from the session date
micro_topic_clusterintegerHDBSCAN micro-topic cluster label; -1 indicates procedural noise or an unassigned speech
macro_topicstringAggregated macro-topic label, such as Macro-Topic 3, or Procedural Noise

How to Get Started

Install the Hugging Face datasets library:

python
pip install datasets pandas

Load the dataset with the default configuration:

python
from datasets import load_dataset

dataset = load_dataset(
    "sl-parliamentary-nlp/sl-parliamentary-hansard-17-26",
    "default",
    download_mode="force_redownload",
)

print(dataset)

Convert to pandas:

python
train_df = dataset["train"].to_pandas()
test_df = dataset["test"].to_pandas()

df = train_df
print(df[["date", "speaker", "macro_topic", "text"]].head())

Example: count speeches by macro-topic:

python
topic_counts = (
    df["macro_topic"]
    .fillna("Unassigned")
    .value_counts()
)

print(topic_counts.head(10))

Corpus Construction

The corpus was constructed from official Hansard PDFs available through the Sri Lankan parliamentary archive.

The processing pipeline included:

  1. 1.Scraping publicly available Hansard PDFs.
  2. 2.Extracting text with Google Gemini, selected because OCR struggles with the dual-column trilingual layout and historical Sinhala/Tamil font encodings.
  3. 3.Segmenting extracted text into speech-level units.
  4. 4.Normalizing speaker names across Sinhala, Tamil, and English variants.
  5. 5.Removing structural headers, footers, and short non-substantive fragments.
  6. 6.Assigning topic labels through the multilingual topic-modeling pipeline.

Topic Modeling Method

The topic-modeling pipeline used:

  • —Embeddings: BGE-M3 dense multilingual embeddings
  • —Dimensionality reduction: UMAP
  • —Micro-topic clustering: HDBSCAN
  • —Macro-topic aggregation: empirical dendrogram cut over micro-topic centroids
  • —Noise handling: HDBSCAN label -1 for procedural or low-density speeches

The final topic-modeling run produced:

ItemValue
Topic-labeled speeches19,553
HDBSCAN micro-topics336
Procedural/noise speeches6,921
Substantive clustered speeches12,632
Macro-topic groups30

Evaluation

The project includes a manually labeled 300-speech evaluation sample and reviewer-facing evaluation artifacts in the companion code repository.

Internal clustering benchmarks reported for the full corpus include:

AlgorithmKNoise %Evaluated rowsSilhouetteCalinski-HarabaszDavies-Bouldin
HDBSCAN33635.4%12,6320.574115,0520.5148
KMeans3360.0%19,5530.407613,0980.8745
Agglomerative3360.0%19,5530.381212,2170.9088

Limitations

  • —Extraction noise: The source PDFs contain complex formatting, multilingual text, and historical font encodings. LLM-based extraction improves coverage but may still introduce errors.
  • —Speaker normalization: Names are normalized heuristically, so rare or ambiguous speaker variants may remain unmerged.
  • —Topic labels are unsupervised: micro_topic_cluster and macro_topic are model-generated analytical labels, not official parliamentary categories.
  • —Noise labels are real text: Speeches labeled -1 are not invalid records; they are simply not assigned to a dense HDBSCAN topic region.
  • —Temporal coverage is uneven: Parliamentary sitting frequency differs by year, so raw counts should be normalized before making year-to-year claims.
  • —Political text can be sensitive: Users should manually verify outputs before using the dataset in high-stakes political, journalistic, legal, or policy contexts.

Bias, Risks, and Considerations

The dataset reflects the language, framing, and political priorities present in parliamentary debate. It may contain partisan language, allegations, sensitive political claims, and socially biased statements made by speakers. Topic models may amplify or obscure these patterns depending on preprocessing choices and clustering parameters.

Researchers should treat the dataset as a computational research resource, not as a neutral summary of Sri Lankan politics.

Citation

If you use this dataset, please cite:

bibtex
@misc{dhanapala2026trilingual,
  title        = {Trilingual Topic Modeling of Sri Lankan Parliamentary Debates},
  author       = {Himath Dhanapala and Haren Daishika and Himandhi Kuruppu and Sithija Seneviratne and Ashini Kavindya and Patalee Narasinghe and Sandeepa Weerasekara and Nisansa de Silva and Sandareka Wickramanayake},
  year         = {2026},
  eprint       = {2608.20365},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url          = {https://arxiv.org/abs/2608.20365}
}

Dataset Card Authors

This dataset card was prepared by the Sri Lanka Parliamentary NLP project team based on the companion research repository and dataset publishing artifacts.

License

The dataset is released under the Apache 2.0 license. The underlying Hansard text is derived from publicly available parliamentary records of the Parliament of Sri Lanka.