CoolFace
Datasetpublic

mdsajjadullah/banglish-sentiment-2026

Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral. Research MotivationBanglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research. Columns text: Banglish… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/banglish-sentiment-2026.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
1likes16downloads
Dataset Card

Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP

Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral.

Research Motivation Banglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research.

Columns

  • —text: Banglish message
  • —sentiment: positive / negative / neutral

Key Stats

  • —Rows: ~25,000
  • —Balance: ~33% each sentiment
  • —Features: Emojis, punctuation, spelling variations, emotional repetition for realism

How it was made

  • —AI-generated (components + randomization) + uniqueness checks
  • —Fully synthetic — no real data

Intended Use

  • —Sentiment classification with BanglaBERT/multilingual models
  • —Banglish NLP education/research

Note: Synthetic only. For research/education.

Created by Md.Sajjad Ullah (Dhaka,Bangladesh)