FredZhang7/all-scam-spam
This is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham. 1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT. Some preprcoessing algorithms spam_assassin.js, followed by spam_assassin.py enron_spam.py Data composition Description To… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.
remove index column after label correction
correct some labels
add suggestion
may not release models, depending on demand
may experience delays
fix lang count
update descriptions
update dataset
fix language stats
add - multilingual
add tag
add more description
update img url
Delete spam_counts.png
line breaks look nicer
add description
add notice
update number
Upload junkmail_dataset.csv
Update README.md
update source
add more info
add stats
Upload spam_counts.png
add examples of data preprocessing
Upload junkmail_dataset.csv
add languages
initial commit
