CoolFace
Datasetpublic

Vxlentina/burmese-text-spam-detection

Dataset Card for burmese-text-spam-detection Dataset Description The burmese-text-spam-detection dataset is a high-quality, human-curated collection of 1,000 Burmese text entries specifically designed for binary text classification tasks. The dataset is balanced equally with 500 "spam" and 500 "not_spam" samples. This dataset was compiled to facilitate the development and evaluation of spam-filtering models for the Burmese language, covering diverse sources such… See the full description on the dataset page: https://huggingface.co/datasets/Vxlentina/burmese-text-spam-detection.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes39downloads
Dataset Card

Dataset Card for burmese-text-spam-detection

Dataset Description

The burmese-text-spam-detection dataset is a high-quality, human-curated collection of 1,000 Burmese text entries specifically designed for binary text classification tasks. The dataset is balanced equally with 500 "spam" and 500 "not_spam" samples.

This dataset was compiled to facilitate the development and evaluation of spam-filtering models for the Burmese language, covering diverse sources such as internet comments, social media interactions, and emails.

Dataset Details

Data Curation and Methodology

The dataset has been meticulously prepared to ensure high utility for machine learning applications:

  • —Data Sources: Data was aggregated from various digital platforms including public comments and professional/personal email communications.
  • —Preprocessing: All text entries have been cleaned and standardized. The dataset uses valid Myanmar Unicode, ensuring compatibility with modern NLP toolkits.
  • —Annotation Process: The dataset underwent a rigorous peer-review labeling process. Both creators independently labeled the entire corpus, followed by a consensus-building phase where both parties verified every label for accuracy and consistency.

Limitations and Bias

While the creators have made every effort to ensure accuracy, users should be aware of the following:

  • —Subjectivity: Labeling is inherently influenced by the annotators' perspectives. Despite the collaborative verification process, the classification reflects the creators' interpretation of what constitutes "spam" versus "not spam" in the context of the Burmese digital landscape.
  • —Context: Spam behavior evolves rapidly. This dataset provides a snapshot in time and may reflect linguistic trends and spam patterns prevalent during its creation.

Citation

If you use this dataset in your research or projects, please cite it as:

bibtex
@dataset{burmese_text_spam_detection,
  author = {Thazin Nyein and Khant Sint Heinn},
  title = {Burmese Text Spam Detection},
  year = {2026},
  link = {https://huggingface.co/datasets/Vxlentina/burmese-text-spam-detection},
  publisher = {Hugging Face},
  license = {CC-BY-4.0}
}