UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence
Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset Dataset Summary This dataset contains 16,156 English-language tourism-attraction reviews from Sri Lanka, covering 2010–2023 and 11 attraction categories. It was prepared for research on rating–sentiment incongruence, where the sentiment expressed in review text differs from the sentiment implied by the numerical star rating. The release contains the final processed file:… See the full description on the dataset page: https://huggingface.co/datasets/UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence.
Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset
Dataset Summary
This dataset contains 16,156 English-language tourism-attraction reviews from Sri Lanka, covering 2010–2023 and 11 attraction categories.
It was prepared for research on rating–sentiment incongruence, where the sentiment expressed in review text differs from the sentiment implied by the numerical star rating.
The release contains the final processed file:
Processed_Reviews_with_Sentiment.csvThe CSV is published in the same column structure used in the research workflow.
Key Information
Dataset Viewer and Loading
The Hugging Face Dataset Viewer exposes the complete CSV as the train split. This is only a technical split name and does not mean that all records were used for model training.
The manually labelled 700-record training set and 300-record testing set used during model evaluation are not included in the default viewer configuration.
from datasets import load_dataset
dataset = load_dataset(
"UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence"
)
data = dataset["train"]
print(data)Original Data Source
This release is derived from:
Sewwandi, T. (2023). Tourism and Travel Reviews: Sri Lankan Destinations (Version 1) [Data set]. Mendeley Data.
The original dataset and this derived release are distributed under the Creative Commons Attribution 4.0 International licence.
Users should cite both the original source dataset and this derived Hugging Face release.
Dataset Preparation
The main processing steps were:
- Standardising location, province, and district information.
- Extracting travel and publication month and year.
- Calculating review and title lengths.
- Estimating the delay between travel and publication.
- Mapping star ratings to
Rating_Class. - Combining the review title and body for sentiment inference.
- Adding the model-predicted
Sentimentfield.
Rating-Class Mapping
Manual Annotation and Sentiment Model
A stratified subset of 1,000 reviews was manually labelled by five members of the research team.
- Each review was labelled by one annotator.
- Annotators considered the review title and body together.
- Star ratings were hidden during annotation.
- Labels were
Negative,Neutral, orPositive. - The labelled subset was divided into 700 training records and 300 testing records.
- Inter-annotator agreement was not calculated because each review received only one manual label.
Sentiment predictions for the full dataset were generated using:
- Model: `cardiffnlp/twitter-roberta-base-sentiment`
- Model revision:
daefdd1f6ae931839bce4d0f3db0a1a4265cd50f - Input: review title and review body combined
- Maximum length: 256 tokens
- Batch size: 32
- Truncation: enabled
The Sentiment column contains model predictions and should not be treated as error-free human ground truth.
Rating–Sentiment Incongruence
A review is treated as incongruent when Rating_Class differs from Sentiment after normalising letter case.
The CSV contains the source fields required for this comparison:
Rating_ClassSentiment
Derived variables such as incongruence indicators and mismatch-pattern labels are created by the accompanying GitHub analysis code and are not stored as separate CSV columns.
In the released dataset:
Overall incongruence rate: 18.6%
Data Fields
Limitations and Responsible Use
Important limitations include:
- strong class imbalance toward positive reviews;
- unequal representation across destinations and attraction types;
- self-selection and platform-specific review behaviour;
- possible domain mismatch because the sentiment model was trained on Twitter-style text;
- possible prediction errors in the
Sentimentfield; - one annotator per manually labelled review;
- no inter-annotator agreement measurement;
- possible legacy text-encoding artefacts;
- one exact duplicate row retained from the analysed corpus.
The dataset should not be used to identify or profile individual reviewers, infer sensitive personal characteristics, or make consequential decisions about individuals or businesses.
Free-text fields may contain names or place references entered by reviewers. Users are responsible for applying appropriate privacy and ethical safeguards.
Licence
This derived release is distributed under the Creative Commons Attribution 4.0 International licence.
Licence: CC BY 4.0
Attribution should be given to:
- the original Mendeley Data dataset;
- this derived Hugging Face dataset;
- the accompanying paper when its methods, terminology, or results are used.
Citation
Derived Hugging Face Dataset
Abaiyan R., Ruththiragayan S., Kusal A., Anusan K., Asma R., Sriyathurshan K., Patale N., Nirasha M., Nisansa D.S., & Sandereka W.. (2026). Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset (Version 1.0.0) [Data set]. Hugging Face. DOI: 10.57967/hf/9407
BibTeX
@dataset{abaiyan2026srilankan,
author = {Ramanaish Abaiyan and Ruththiragayan Sutharsan and
Kusal Amantha and Anusan Krishnathas and Asma Rauff and
Kovindarajah Sriyathurshan and Patalee Narasinghe and
Nirasha Munasinghe and Nisansa de Silva and
Sandareka Wickramanayake},
title = {Sri Lankan Tourism Review Sentiment-Rating
Incongruence Dataset},
year = {2026},
version = {1.0.0},
publisher = {Hugging Face},
doi = {10.57967/hf/9407},
url = {https://doi.org/10.57967/hf/9407}
}Original Source Dataset
Sewwandi T. (2023). Tourism and Travel Reviews: Sri Lankan Destinations (Version 1) [Data set]. Mendeley Data. DOI: 10.17632/2nbvx5m4hs.1
Related Resources
Authors
Ramanaish Abaiyan, Ruththiragayan Sutharsan, Kusal Amantha, Anusan Krishnathas, Asma Rauff, Kovindarajah Sriyathurshan, Patalee Narasinghe, Nirasha Munasinghe, Nisansa de Silva, and Sandareka Wickramanayake.
Department of Computer Science and Engineering, University of Moratuwa, Sri Lanka
