datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-agent-scam-conversation
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.all-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham.
1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT.
Some preprcoessing algorithms
spam_assassin.js, followed by spam_assassin.py
enron_spam.py
Data composition
Description
To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.scam-dialogue
Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/scam-dialogue.single-agent-scam-conversations
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset
Dataset Description
The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The transcribed conversation between the caller and receiver.
type: The specific type of scam or non-scam interaction.
labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.discord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.scam-hum-india
Scam/Spam India Dataset
A dataset of labeled text messages for scam/spam detection, focused on Indian scam patterns
(telecom promotions, lottery fraud, Aadhaar/KYC phishing, UPI/bank fraud, fake job offers, OTP-theft attempts).
Dataset Structure
Column
Type
Description
text
string
The message content
label
string
ham (legitimate), spam/scam
Total rows: 2272
Label distribution: {'ham': 1377, 'spam': 895}
Source
Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam-hum-india.phone-scam-datasetIndian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.ScamBench
SCAMBENCH: A Multi-Perspective Benchmark for Online Scam Communication
SCAMBENCH is a dataset for studying online scam communication, scam detection, and model robustness. It contains real-world scam messages curated from victim-reported incidents, along with non-scam counterparts and structured annotations.
The dataset is introduced in our paper:
SCAMBENCH: A Multi-Perspective Benchmark for Analyzing and Evaluating Online Scam Communication
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Shouninger/ScamBench.scamshield-scam-detection-data
ScamShield — Scam Detection Dataset
Training data used to fine-tune the ScamShield scam detection model, which powers the ScamShield Chrome extension — an on-device scam, phishing, and fake job-offer detector.
Dataset Structure
File
Description
Rows
train.csv
Training split (see composition below)
~19,000
validation.csv
Validation split (original data only)
~2,300
test.csv
Held-out test split (original data only)
~2,300
Each row has:
text —… See the full description on the dataset page: https://huggingface.co/datasets/Him1304/scamshield-scam-detection-data.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Scammer-ConversationThis dataset are generated by gretelai/tabular-v0
This dataset contains a collection of conversations between scammers, scam baiters, and normal conversations. The purpose of this dataset is to provide a resource for training and evaluating models for scam detection and classification.
Internship-Job-Scam-DatasetThis dataset is a custom-built text classification dataset for detecting fraud risk in internship and entry-level job postings. This dataset is created for an embeddings-based classifier that helps students evaluate whether internship and entry-level job postings may be legitimate, suspicious, or fraudulent.
Dataset Summary
The dataset contains job posting texts labeled into three risk categories:
legitimate: postings that look like normal internship or entry-level job… See the full description on the dataset page: https://huggingface.co/datasets/agnialf/Internship-Job-Scam-Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/harshlimkar/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/arka15/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/szjkdsldf/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.scamshield-scam-detection-data
ScamShield — Scam Detection Dataset
Training data used to fine-tune the ScamShield scam detection model, which powers the ScamShield Chrome extension — an on-device scam, phishing, and fake job-offer detector.
Dataset Structure
File
Description
Rows
train.csv
Training split (see composition below)
~19,000
validation.csv
Validation split (original data only)
~2,300
test.csv
Held-out test split (original data only)
~2,300
Each row has:
text —… See the full description on the dataset page: https://huggingface.co/datasets/rehan-ml/scamshield-scam-detection-data.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/chinna887/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.youtube-scam-conversationsInternship-Job-Scam-DatasetThis dataset is a custom-built text classification dataset for detecting fraud risk in internship and entry-level job postings. This dataset is created for an embeddings-based classifier that helps students evaluate whether internship and entry-level job postings may be legitimate, suspicious, or fraudulent.
Dataset Summary
The dataset contains job posting texts labeled into three risk categories:
legitimate: postings that look like normal internship or entry-level job… See the full description on the dataset page: https://huggingface.co/datasets/akshujey/Internship-Job-Scam-Dataset.multi-agent-scam-conversation
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/Lyr1k/multi-agent-scam-conversation.discord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.eurlexIndian_Multilingual_Scam_Message_Dataset
Indian Multilingual Scam Message Dataset
Overview
This dataset contains 120 realistic SMS and text messages from Indian contexts, labeled as scam or legitimate. It reflects real-world communication patterns across Hindi, Hinglish, and English.
Features
120 high-quality samples
Multilingual (Hindi, Hinglish, English)
Real-world inspired scam and legitimate messages
Includes reasoning for each label
Covers multiple domains: banking, ecommerce, telecom, utilities… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Indian_Multilingual_Scam_Message_Dataset.multi-agent-scam-conversation
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/436-4/multi-agent-scam-conversation.scam-dialogue
Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/NikithaVenkat0205/scam-dialogue.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/Vedantsc-1110/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.scam-dialogue
Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/436-4/scam-dialogue.discord_scam_detectiontx_scam
