UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence
Sri Lankan Tourism Review Sentiment-Rating Incongruence Dataset Dataset Summary This dataset contains 16,156 English-language tourism-attraction reviews from Sri Lanka, covering 2010–2023 and 11 attraction categories. It was prepared for research on rating–sentiment incongruence, where the sentiment expressed in review text differs from the sentiment implied by the numerical star rating. The release contains the final processed file:… See the full description on the dataset page: https://huggingface.co/datasets/UOM-CSE-E23/Sri-Lankan-Tourism-Review-Incongruence.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Add professional dataset card and file configuration
# Sri Lankan Tourism Review Incongruence Dataset ## Dataset Summary The Sri Lankan Tourism Review Incongruence Dataset contains 16,156 English-language tourism attraction reviews associated with destinations in Sri Lanka. The dataset was prepared to study rating–sentiment incongruence: cases where the sentiment expressed in a written review differs from the sentiment implied by its numerical star rating. The published CSV is the sentiment-enriched processed dataset used in the accompanying research workflow. The file is released in its original research format without further modification. ## Dataset Configurations ### full_dataset The `full_dataset` configuration contains the complete sentiment-enriched corpus: - File: `Processed_Reviews_with_Sentiment.csv` - Split: `full` - Number of records: 16,156 - Number of columns: 21 The `full` split represents the complete research corpus. It is not an experimental training split. ### manual_annotations The `manual_annotations` configuration contains the manually labelled sentiment subset used for model development and evaluation: - Training split: 700 reviews - Testing split: 300 reviews - Sentiment classes: Negative, Neutral and Positive ## Original Dataset Source This dataset is derived from: **Tourism and Travel Reviews: Sri Lankan Destinations** Original creator: Taniya Sewwandi Repository: Mendeley Data Version: 1 DOI: 10.17632/2nbvx5m4hs.1 Licence: Creative Commons Attribution 4.0 International Users should cite both the original Mendeley Data release and this derived Hugging Face dataset. ## Dataset Creation ### Preprocessing The original tourism review data were cleaned and processed to produce consistent geographic, temporal, textual and rating-related variables. The published CSV is provided in the same form used in the research workflow. No additional normalization, duplicate removal, column renaming or format conversion was performed specifically for the Hugging Face release. ### Rating-Class Construction Numerical star ratings were grouped into three classes: - 1–2 stars: Negative - 3 stars: Neutral - 4–5 stars: Positive The resulting class is provided in the `Rating_Class` column. ### Manual Annotation Process A proportional stratified sample of 1,000 reviews was selected based on rating class. Five members of the research team assigned one of three textual sentiment labels: - Negative - Neutral - Positive Annotators assessed the review title and review body together. The numerical star rating was hidden from annotators during sentiment annotation. Each review was assigned to one annotator. Therefore, inter-annotator agreement was not calculated, and no disagreement-resolution procedure was required. The annotated reviews were divided into 700 training records and 300 testing records. ### Sentiment Prediction Textual sentiment for the complete dataset was generated using: `cardiffnlp/twitter-roberta-base-sentiment` Model revision: `daefdd1f6ae931839bce4d0f3db0a1a4265cd50f` The review title and review body were combined as input. The model generated one of three sentiment labels: - NEGATIVE - NEUTRAL - POSITIVE Sentiment inference was performed independently of the numerical star rating. ## Data Fields The complete dataset contains the following columns: | Column | Description | |---|---| | `Location_Name` | Name of the reviewed tourism attraction | | `Located_City` | City associated with the attraction | | `Province` | Sri Lankan province | | `District` | Sri Lankan district | | `Location_Type` | Category of tourism attraction | | `User_Locale` | Locale associated with the reviewer | | `User_Country` | Country associated with the reviewer | | `Travel_Date_Month` | Month of the reported visit | | `Travel_Date_Year` | Year of the reported visit | | `Published_Date_Month` | Month in which the review was published | | `Published_Date_Year` | Year in which the review was published | | `User_Contributions` | Number of contributions associated with the reviewer | | `Rating` | Original numerical star rating | | `Helpful_Votes` | Number of helpful votes received | | `Title` | Review title | | `Text` | Review body | | `Review_Length` | Review-body length in characters | | `Title_Length` | Review-title length in characters | | `Rating_Class` | Negative, Neutral or Positive rating class | | `Review_Delay_Days` | Estimated delay between visit and publication | | `Sentiment` | Transformer-predicted textual sentiment | ## Intended Uses The dataset may be used for: - sentiment analysis; - rating–sentiment incongruence research; - weak-label reliability studies; - tourism-review analytics; - text classification; - reviewer-behaviour analysis; - explainable machine-learning research. ## Out-of-Scope Uses The dataset should not be used to: - identify or profile individual reviewers; - infer sensitive personal characteristics; - make decisions affecting individual travellers; - represent every visitor to Sri Lanka; - generate deceptive or fabricated reviews; - treat predicted sentiment as error-free human ground truth. ## Biases and Limitations The dataset is based on reviews available through the original source and is not necessarily representative of all visitors, destinations or tourism experiences in Sri Lanka. Possible limitations include: - self-selection bias; - uneven representation across attraction types; - geographic imbalance; - platform-specific reviewer behaviour; - errors introduced by automated sentiment classification; - domain mismatch between tourism reviews and the selected sentiment model's original training data; - each manually labelled review was assessed by one annotator; - inter-annotator agreement was not measured. The predicted `Sentiment` column should therefore be interpreted as model-generated sentiment rather than definitive human truth. ## Reproducibility The preprocessing, sentiment inference, statistical analysis and machine-learning workflow are available in the accompanying GitHub repository: `Abaiyan-27/Fault-of-Our-Stars---Behavioral-Drivers-of-Rating-Sentiment-Incongruence` ## Licence This derived dataset is released under the Creative Commons Attribution 4.0 International licence. Users must provide appropriate attribution to: 1. the original Mendeley Data dataset; 2. this derived Hugging Face dataset; 3. the accompanying research paper where applicable. ## Authors - Ramanaish Abaiyan - Ruththiragayan Sutharsan - Kusal Amantha - Anusan Krishnathas - Asma Rauff - Kovindarajah Sriyathurshan - Patalee Narasinghe - Nirasha Munasinghe - Nisansa de Silva - Sandareka Wickramanayake ## Citation A permanent dataset DOI and final citation will be added after the repository has been reviewed and the version 1.0.0 release has been approved. Until then, please cite the repository title, authors, version and Hugging Face repository. ## Contact For questions concerning the dataset, open an issue in the associated GitHub repository or contact the corresponding research author.
Update README.md
manual_annotations/
Upload sentiment-enriched research dataset
initial commit
