Spam Detection
spam-detection-dataset
Dataset Card for "spam-detection-dataset"
More Information needed
russian-spam-detection
RuSpam Dataset
Датасет для бинарной классификации сообщений на спам / не спам.
📊 Структура
message — текст сообщения
label — (1 — спам, 0 — не спам)
🧠 Модель
darkQibit/ruSpamNS_v31_big
📢 Канал автора
Подпишитесь на Telegram-канал автора — это сильно поможет проекту.Там публикуются новости, обновления и другие разработки.
👉 https://t.me/qubit_a
burmese-text-spam-detection
Dataset Card for burmese-text-spam-detection
Dataset Description
The burmese-text-spam-detection dataset is a high-quality, human-curated collection of 1,000 Burmese text entries specifically designed for binary text classification tasks. The dataset is balanced equally with 500 "spam" and 500 "not_spam" samples.
This dataset was compiled to facilitate the development and evaluation of spam-filtering models for the Burmese language, covering diverse sources such… See the full description on the dataset page: https://huggingface.co/datasets/Vxlentina/burmese-text-spam-detection.synthetic-spam-detection-dataset-german
Tanaos Spam Detection German Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in German.
Our german spam detection model, tanaos-spam-detection-german, was trained on this dataset.
Dataset Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-german.spam-detection-dataset-splits
Spam Detection Dataset
This is the dataset for spam classification task. It contains:
'train' subset with 8175 samples
'validation' subset with 1362 samples
'test' subset with 1636 samples
Source and modifications
This dataset is cloned from Deysi/spam-detection-dataset with the following added processing:
Convert 'string' to 'id' label that allows to be used and trained directly with transformer's trainer
Split the original 'test' dataset (2725 samples) into 2… See the full description on the dataset page: https://huggingface.co/datasets/tanquangduong/spam-detection-dataset-splits.SPAM_DETECTION
