FredZhang7/all-scam-spam
This is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham. 1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT. Some preprcoessing algorithms spam_assassin.js, followed by spam_assassin.py enron_spam.py Data composition Description To… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.
This is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham.
1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT.
<br>
Some preprcoessing algorithms
- spam_assassin.js, followed by spam_assassin.py
- enron_spam.py
<br>
Data composition

<br>
Description
To make the text format between sms messages and emails consistent, email subjects and content are separated by two newlines:
text = email.subject + "\n\n" + email.content<br>
Suggestions
- If you plan to train a model based on this dataset alone, I recommend adding some rows with
is_toxic=0fromFredZhang7/toxi-text-3M. Make sure the rows aren't spam.
<br>
Other Sources
- https://huggingface.co/datasets/sms_spam
- https://github.com/MWiechmann/enronspamdata
- https://github.com/stdlib-js/datasets-spam-assassin
- https://repository.ortolang.fr/api/content/comere/v3.3/cmr-simuligne.html
