sanjidh90/KUET_whispers_Dataset
๐ Dataset Card: KUET Whispers ๐งพ Overview KUET Whispers is a curated dataset of anonymous, emotionally expressive posts collected from HazyBoardโa student-run, anonymous confessions platform at Khulna University of Engineering & Technology (KUET) in Bangladesh. The dataset captures natural, informal, and highly engaging text data in Bangla, English, and code-mixed formats, making it suitable for a range of NLP tasks including: Sentiment and emotion analysisโฆ See the full description on the dataset page: https://huggingface.co/datasets/sanjidh90/KUET_whispers_Dataset.
๐ Dataset Card: KUET Whispers
๐งพ Overview
KUET Whispers is a curated dataset of anonymous, emotionally expressive posts collected from HazyBoardโa student-run, anonymous confessions platform at Khulna University of Engineering & Technology (KUET) in Bangladesh. The dataset captures natural, informal, and highly engaging text data in Bangla, English, and code-mixed formats, making it suitable for a range of NLP tasks including:
- Sentiment and emotion analysis
- Code-mixed language modeling
- Social engagement prediction
- Sarcasm and intent detection
Each entry includes both the raw and cleaned text, as well as metadata such as likes, comments, sender (anonymous alias), timestamp, and the original post URL.
๐ About HazyBoard
KUET Whispers is an anonymous Facebook-based platform where KUET students publicly share thoughts, confessions, questions, and opinionsโoften touching on personal, humorous, romantic, or social themes. Posts are submitted anonymously and published by moderators, making it a vibrant reflection of campus life, emotions, and culture.
This dataset aims to preserve that raw authenticity for academic and research purposes.
๐ Dataset Structure
๐ Languages
The dataset contains:
- ๐ง๐ฉ Bangla (native, formal and informal)
- ๐บ๐ธ English
- ๐ Code-mixed (Banglish), often switching between scripts
โ๏ธ Collection & Processing
- Source: 'KUET Whispers' Facebook page (public content only)
- Collection Date: July 2025
- Cleaning: Minimal โ emojis and formatting artifacts removed; text preserved as-is to maintain emotional tone
- Size: 2,485 entries
- Format: CSV, easily convertible to Hugging Face
Datasetobject
๐ Applications
- Fine-tuning sentiment/emotion classification models for low-resource languages
- Studying code-mixed language behavior in real-world settings
- Modeling social media engagement (likes/comments)
- Analyzing trends, opinions, and mental health signals in anonymous platforms
โ๏ธ Ethical Considerations
- Only public, anonymized posts were collected.
- No personal identifiers or user accounts are included.
- Some content may contain sensitive or emotionally heavy language (e.g., loneliness, anxiety, relationships).
- Please use the dataset responsibly and do not attempt to deanonymize contributors.
๐ License
This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You are free to use, modify, and distribute it with proper credit.
