datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
burmese-text-spam-detection
Dataset Card for burmese-text-spam-detection
Dataset Description
The burmese-text-spam-detection dataset is a high-quality, human-curated collection of 1,000 Burmese text entries specifically designed for binary text classification tasks. The dataset is balanced equally with 500 "spam" and 500 "not_spam" samples.
This dataset was compiled to facilitate the development and evaluation of spam-filtering models for the Burmese language, covering diverse sources such… See the full description on the dataset page: https://huggingface.co/datasets/Vxlentina/burmese-text-spam-detection.synthetic-spam-detection-dataset-german
Tanaos Spam Detection German Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in German.
Our german spam detection model, tanaos-spam-detection-german, was trained on this dataset.
Dataset Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-german.synthetic-spam-detection-dataset-spanish
Tanaos Spam Detection Spanish Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in Spanish.
Our spanish spam detection model, tanaos-spam-detection-spanish, was trained on this dataset.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-spanish.facebook_spam_detection
Facebook Spam Detection Dataset
Dataset Summary
This dataset contains 600 Facebook profiles with behavioral and activity features designed for spam detection in social media. The dataset enables binary classification to distinguish between spam accounts (Label=1) and legitimate accounts (Label=0), providing insights into spammer behavior patterns on Facebook.
Dataset Details
Total Samples: 600 profiles
Classes: Binary (0 = Legitimate, 1 = Spam)
Class… See the full description on the dataset page: https://huggingface.co/datasets/nahiar/facebook_spam_detection.spam-detection-samplesynthetic-spam-detection-dataset-italian
Tanaos Spam Detection Italian Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in Italian.
Our Italian spam detection model, tanaos-spam-detection-italian, was trained on this dataset.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-italian.turkish-igaming-spam-detection
Turkish iGaming Spam & Phishing Detection Dataset
This dataset contains categorized text samples commonly used in SMS spam, phishing attempts, and promotional abuse within the Turkish iGaming sector.
It is curated to assist NLP researchers and cybersecurity analysts in training models to detect deceptive patterns and protect consumers.
Dataset Structure
text: The raw text content (SMS or notification).
label: Classification (spam or ham).
category: Specific threat type… See the full description on the dataset page: https://huggingface.co/datasets/eskfestsecurity/turkish-igaming-spam-detection.spam-detection-analysisspam_email_detectionspam_detection
