datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BanglishDepYT
BanglishDepYT
Banglish code-mixed YouTube comment dataset for depression-related NLP research.
Statistics
191k unlabeled comments
1k manually labeled comments
Tasks
Depression detection
Sentiment analysis
Code-mixed NLP
Data Collection stuffs
GitHub: https://github.com/Ay-on-Roy/BanglishDepYT
License
license: apache-2.0
banglish-sentiment-2026Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP
Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral.
Research MotivationBanglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research.
Columns
text: Banglish… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/banglish-sentiment-2026.Banglish-EnglishBanglish-English
