rifat101/BnSentMix
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis Dataset Overview Column Title Description Data Sources Facebook, YouTube, E-commerce Sites #Samples 20000 Sentiment Labels 1:Positive, 2:Negative, 3:Neutral, 4:Mixed Filtering Method Automated using mBERT #Annotators 64 Annotation/Sample 2 or 3 (if tie) Dataset Statistics Statistic Value Mean Character Length 62.77 Max Character… See the full description on the dataset page: https://huggingface.co/datasets/rifat101/BnSentMix.
06
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis

Dataset Overview
Dataset Statistics
Citation
If you find this work useful, please cite our paper:
@inproceedings{alam-etal-2025-bnsentmix,
title = "{B}n{S}ent{M}ix: A Diverse {B}engali-{E}nglish Code-Mixed Dataset for Sentiment Analysis",
author = "Alam, Sadia and Ishmam, Md Farhan and Alvee, Navid Hasin and Siddique, Md Shahnewaz and
Hossain, Md Azam and Kamal, Abu Raihan Mostofa", editor = "Hettiarachchi, Hansi and
Ranasinghe, Tharindu and Rayson, Paul and Mitkov, Ruslan and Gaber, Mohamed and
Premasiri, Damith and Tan, Fiona Anting and Uyangodage, Lasitha",
booktitle = "Proceedings of the First Workshop on Language Models for Low-Resource Languages",
month = jan,
year = "2025",
address = "Abu Dhabi, United Arab Emirates",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.loreslm-1.4/",
pages = "68--77"
}