CoolFace
Datasetpublic

UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence

Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset Dataset Summary This dataset contains 16,156 English-language tourism-attraction reviews from Sri Lanka, covering 2010–2023 and 11 attraction categories. It was prepared for research on rating–sentiment incongruence, where the sentiment expressed in review text differs from the sentiment implied by the numerical star rating. The release contains the final processed file:… See the full description on the dataset page: https://huggingface.co/datasets/UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
1likes93downloads
Dataset Card

Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset

Dataset Summary

This dataset contains 16,156 English-language tourism-attraction reviews from Sri Lanka, covering 2010–2023 and 11 attraction categories.

It was prepared for research on rating–sentiment incongruence, where the sentiment expressed in review text differs from the sentiment implied by the numerical star rating.

The release contains the final processed file:

text
Processed_Reviews_with_Sentiment.csv

The CSV is published in the same column structure used in the research workflow.

Key Information

ItemValue
Records16,156
Columns21
LanguageEnglish
Time period2010–2023
LicenceCC BY 4.0
DOI10.57967/hf/9407
DOI revision2a15ce9

Dataset Viewer and Loading

The Hugging Face Dataset Viewer exposes the complete CSV as the train split. This is only a technical split name and does not mean that all records were used for model training.

The manually labelled 700-record training set and 300-record testing set used during model evaluation are not included in the default viewer configuration.

python
from datasets import load_dataset

dataset = load_dataset(
    "UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence"
)

data = dataset["train"]
print(data)

Original Data Source

This release is derived from:

Sewwandi, T. (2023). Tourism and Travel Reviews: Sri Lankan Destinations (Version 1) [Data set]. Mendeley Data.

DOI: 10.17632/2nbvx5m4hs.1

The original dataset and this derived release are distributed under the Creative Commons Attribution 4.0 International licence.

Users should cite both the original source dataset and this derived Hugging Face release.


Dataset Preparation

The main processing steps were:

  1. 1.Standardising location, province, and district information.
  2. 2.Extracting travel and publication month and year.
  3. 3.Calculating review and title lengths.
  4. 4.Estimating the delay between travel and publication.
  5. 5.Mapping star ratings to Rating_Class.
  6. 6.Combining the review title and body for sentiment inference.
  7. 7.Adding the model-predicted Sentiment field.

Rating-Class Mapping

Star rating`Rating_Class`
1–2Negative
3Neutral
4–5Positive

Manual Annotation and Sentiment Model

A stratified subset of 1,000 reviews was manually labelled by five members of the research team.

  • —Each review was labelled by one annotator.
  • —Annotators considered the review title and body together.
  • —Star ratings were hidden during annotation.
  • —Labels were Negative, Neutral, or Positive.
  • —The labelled subset was divided into 700 training records and 300 testing records.
  • —Inter-annotator agreement was not calculated because each review received only one manual label.

Sentiment predictions for the full dataset were generated using:

  • —Model: `cardiffnlp/twitter-roberta-base-sentiment`
  • —Model revision: daefdd1f6ae931839bce4d0f3db0a1a4265cd50f
  • —Input: review title and review body combined
  • —Maximum length: 256 tokens
  • —Batch size: 32
  • —Truncation: enabled

The Sentiment column contains model predictions and should not be treated as error-free human ground truth.


Rating–Sentiment Incongruence

A review is treated as incongruent when Rating_Class differs from Sentiment after normalising letter case.

The CSV contains the source fields required for this comparison:

  • —Rating_Class
  • —Sentiment

Derived variables such as incongruence indicators and mismatch-pattern labels are created by the accompanying GitHub analysis code and are not stored as separate CSV columns.

In the released dataset:

CategoryRecords
Congruent13,151
Incongruent3,005
Total16,156

Overall incongruence rate: 18.6%


Data Fields

ColumnTypeDescription
Location_NamestringTourism attraction name
Located_CitystringCity associated with the attraction
ProvincestringProvince associated with the attraction
DistrictstringDistrict associated with the attraction
Location_TypestringAttraction category
User_LocalestringReviewer locale
User_CountrystringReviewer country
Travel_Date_Monthint64Travel month
Travel_Date_Yearint64Travel year
Published_Date_Monthint64Review publication month
Published_Date_Yearint64Review publication year
User_Contributionsint64Reviewer contribution count
Ratingint64Star rating from 1 to 5
Helpful_Votesint64Helpful-vote count
TitlestringReview title
TextstringReview body
Review_Lengthint64Review-body character count
Title_Lengthint64Review-title character count
Rating_ClassstringRating-derived class
Review_Delay_Daysint64Estimated delay between travel and publication
SentimentstringModel-predicted textual sentiment

Limitations and Responsible Use

Important limitations include:

  • —strong class imbalance toward positive reviews;
  • —unequal representation across destinations and attraction types;
  • —self-selection and platform-specific review behaviour;
  • —possible domain mismatch because the sentiment model was trained on Twitter-style text;
  • —possible prediction errors in the Sentiment field;
  • —one annotator per manually labelled review;
  • —no inter-annotator agreement measurement;
  • —possible legacy text-encoding artefacts;
  • —one exact duplicate row retained from the analysed corpus.

The dataset should not be used to identify or profile individual reviewers, infer sensitive personal characteristics, or make consequential decisions about individuals or businesses.

Free-text fields may contain names or place references entered by reviewers. Users are responsible for applying appropriate privacy and ethical safeguards.


Licence

This derived release is distributed under the Creative Commons Attribution 4.0 International licence.

Licence: CC BY 4.0

Attribution should be given to:

  1. 1.the original Mendeley Data dataset;
  2. 2.this derived Hugging Face dataset;
  3. 3.the accompanying paper when its methods, terminology, or results are used.

Citation

Derived Hugging Face Dataset

Abaiyan R., Ruththiragayan S., Kusal A., Anusan K., Asma R., Sriyathurshan K., Patale N., Nirasha M., Nisansa D.S., & Sandereka W.. (2026). Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset (Version 1.0.0) [Data set]. Hugging Face. DOI: 10.57967/hf/9407

BibTeX

bibtex
@dataset{abaiyan2026srilankan,
  author    = {Ramanaish Abaiyan and Ruththiragayan Sutharsan and
               Kusal Amantha and Anusan Krishnathas and Asma Rauff and
               Kovindarajah Sriyathurshan and Patalee Narasinghe and
               Nirasha Munasinghe and Nisansa de Silva and
               Sandareka Wickramanayake},
  title     = {Sri Lankan Tourism Review Sentiment-Rating
               Incongruence Dataset},
  year      = {2026},
  version   = {1.0.0},
  publisher = {Hugging Face},
  doi       = {10.57967/hf/9407},
  url       = {https://doi.org/10.57967/hf/9407}
}

Original Source Dataset

Sewwandi T. (2023). Tourism and Travel Reviews: Sri Lankan Destinations (Version 1) [Data set]. Mendeley Data. DOI: 10.17632/2nbvx5m4hs.1

Related Resources


Authors

Ramanaish Abaiyan, Ruththiragayan Sutharsan, Kusal Amantha, Anusan Krishnathas, Asma Rauff, Kovindarajah Sriyathurshan, Patalee Narasinghe, Nirasha Munasinghe, Nisansa de Silva, and Sandareka Wickramanayake.

Department of Computer Science and Engineering, University of Moratuwa, Sri Lanka