mdsajjadullah/banglish-sentiment-2026
Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral. Research MotivationBanglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research. Columns text: Banglish… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/banglish-sentiment-2026.
Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP
Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral.
Research Motivation Banglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research.
Columns
- text: Banglish message
- sentiment: positive / negative / neutral
Key Stats
- Rows: ~25,000
- Balance: ~33% each sentiment
- Features: Emojis, punctuation, spelling variations, emotional repetition for realism
How it was made
- AI-generated (components + randomization) + uniqueness checks
- Fully synthetic — no real data
Intended Use
- Sentiment classification with BanglaBERT/multilingual models
- Banglish NLP education/research
Note: Synthetic only. For research/education.
Created by Md.Sajjad Ullah (Dhaka,Bangladesh)
