CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Scam-AI /gpt-image-2gated GPT-Image-2 Twitter Dataset 10,217 confirmed GPT-image-2.0 generated images collected from Twitter/XCollection window: April 21 – April 28, 2026 (first week post-launch)Paper: GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment Overview This dataset contains 10,217 images confirmed to be GPT-image-2.0 outputs, sourced from public Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt-image-2.image-classification10K<n<100K10 likes1k downloads4mo agoHugging Face02Scam-AI /AIForge-Doc-v1gated AIForge-Doc: A Benchmark of AI-Forged Document Images AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting financial and identity document fraud. Every tampered image was produced by a diffusion-model inpainting pipeline — a threat model that existing forgery detectors cannot reliably handle. At a Glance Attribute Value Total forged images 4,061 Training split 3,249 (80 %) Testing split 812 (20 %) Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.image-classification1K<n<10K3 likes919 downloads2mo agoHugging Face03Scam-AI /AIForge-Doc-v2gated AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries AIForge-Doc v2 is the first paired benchmark of document forgeries produced by OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by its authentic source image and a pixel-precise tampered-region mask in DocTamper-compatible format. v2 reuses the forgery specifications of AIForge-Doc v1 spec-for-spec and swaps only the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.image-classification1K<n<10K2 likes628 downloads2mo agoHugging Face04TamAko783 /filipino-scam-final-shards Filipino Scam — Final Shards CoT-formatted pickle shards for Qwen3-VL-4B-Instruct LoRA fine-tuning on Filipino short-form video scam detection. This is the final, training-ready format used to build TamAko783/Scam-Qwen3-VL-4B-final-lora. Splits File Samples Purpose Training.pkl 1600 Training set Validate.pkl 198 Validation (early stopping + best-checkpoint selection) Evaluate.pkl 202 Held-out test (final unbiased metrics) Both Validate and Evaluate are… See the full description on the dataset page: https://huggingface.co/datasets/TamAko783/filipino-scam-final-shards.0 likes456 downloads5mo agoHugging Face05scampion /handball_video_sequencesvideo1K<n<10K0 likes358 downloads1y agoHugging Face06BLISS-e-V /SCAM SCAM Dataset Dataset Summary SCAM is the largest and most diverse real-world typographic attack dataset to date, containing images across hundreds of object categories and attack words. The dataset is designed to study and evaluate the robustness of multimodal foundation models against typographic attacks. Usage: from datasets import load_dataset ds = load_dataset("BLISS-e-V/SCAM", split="train") print(ds) img = ds[0]['image'] For more information, check out our… See the full description on the dataset page: https://huggingface.co/datasets/BLISS-e-V/SCAM.imageimage-classification1K<n<10K6 likes339 downloads5mo agoHugging Face07BothBosu /multi-agent-scam-conversation Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.texttext-classification1K<n<10K10 likes275 downloads2y agoHugging Face08FredZhang7 /all-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham. 1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT. Some preprcoessing algorithms spam_assassin.js, followed by spam_assassin.py enron_spam.py Data composition Description To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.texttext-classification10K<n<100K15 likes236 downloads2y agoHugging Face09BothBosu /scam-dialogue Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams. Dataset Structure The dataset consists of three columns: dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/scam-dialogue.texttext-classification1K<n<10K10 likes205 downloads2y agoHugging Face10TamAko783 /filipino-scam-prepared-cachetext1K<n<10K0 likes169 downloads5mo agoHugging Face11scampion /Eurovoc 🇪🇺 🏷️ EuroVoc dataset This dataset contains more that 4 millions documents in 39 languages with associated EuroVoc labels. What's Cellar ? Cellar is the common data repository of the Publications Office of the European Union. Digital publications and metadata are stored in and disseminated via Cellar, in order to be used by humans and machines. Aiming to transparently serve users, Cellar stores multilingual publications and metadata, it is open to all EU… See the full description on the dataset page: https://huggingface.co/datasets/scampion/Eurovoc.text1M<n<10M0 likes137 downloads2y agoHugging Face12Prakriti-21 /video-finance-scam-awarenesstext1K<n<10K0 likes137 downloads10d agoHugging Face13shaw /scambench-training ScamBench Training Corpus A multilingual, multi-turn training corpus for building scam-resistant autonomous agents. Dataset Description ScamBench is a comprehensive dataset designed to train AI agents to resist social engineering, phishing, prompt injection, credential theft, impersonation, advance-fee fraud, and other adversarial attacks while maintaining helpfulness for legitimate requests. Key Features 37,423 total records across 14 languages 154 attack… See the full description on the dataset page: https://huggingface.co/datasets/shaw/scambench-training.text-classification10K<n<100K0 likes130 downloads6mo agoHugging Face14build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes125 downloads3mo agoHugging Face15Shade63 /scam-classification-multiclass Scam Text Classification - Multi-Class Dataset Overview This dataset is an enhanced version of the original binary scam classification dataset, now with 5 multi-class categories for more granular scam detection. Dataset Structure Original Dataset 14,000 rows of Indian-context SMS/email-style messages Binary labels: 0 (legit) / 1 (scam) Domain-specific: Indian banks, UPI, Aadhaar, government agencies Updated Multi-Class Categories… See the full description on the dataset page: https://huggingface.co/datasets/Shade63/scam-classification-multiclass.1 likes111 downloads4mo agoHugging Face16BothBosu /single-agent-scam-conversations Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset Dataset Description The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams. Dataset Structure The dataset consists of three columns: dialogue: The transcribed conversation between the caller and receiver. type: The specific type of scam or non-scam interaction. labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.texttext-classification1K<n<10K2 likes100 downloads2y agoHugging Face17Purino /scam-ads-detection-datasetimage1K<n<10K1 likes84 downloads4mo agoHugging Face18CyberHarem /scamp_kantaicollection Dataset of scamp (Kantai Collection) This is the dataset of scamp (Kantai Collection), containing 166 images and their tags. The core tags of this character are long_hair, side_ponytail, hair_ornament, star_hair_ornament, hat, grey_hair, garrison_cap, aqua_headwear, hair_ribbon, ribbon, black_ribbon, grey_eyes, breasts, small_breasts, brown_eyes, hair_between_eyes, headgear, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...)… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/scamp_kantaicollection.text-to-image1K<n<10K0 likes79 downloads3y agoHugging Face19wangyuancheng /discord-phishing-scam-clean Discord Scam / Clean Messages Dataset 📌 Context This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection. 💡 Inspiration Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.texttext-classification1K<n<10K2 likes75 downloads1y agoHugging Face20anmolshrivastav /scam-hum-india Scam/Spam India Dataset A dataset of labeled text messages for scam/spam detection, focused on Indian scam patterns (telecom promotions, lottery fraud, Aadhaar/KYC phishing, UPI/bank fraud, fake job offers, OTP-theft attempts). Dataset Structure Column Type Description text string The message content label string ham (legitimate), spam/scam Total rows: 2272 Label distribution: {'ham': 1377, 'spam': 895} Source Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam-hum-india.texttext-classification1K<n<10K0 likes73 downloads9d agoHugging Face21menaattia /phone-scam-datasettext1K<n<10K3 likes70 downloads1y agoHugging Face22ysangam /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K1 likes67 downloads4mo agoHugging Face23sidzzz07 /scamshield-dataset 🛡️ ScamShield: Multilingual Smishing & Social Engineering Corpus A production-grade 86,000+ sample multilingual dataset specifically crafted for detecting SMS phishing (smishing), scam intent, and psychological social engineering manipulation across English, Hindi (Devanagari), and romanized Hinglish. This dataset powers the multi-task frozen model in the open-source ScamShield AI project. texttext-classification10K<n<100K1 likes67 downloads4d agoHugging Face24Scam-AI /gpt4o-receiptgated GPT4o-Receipt: AI-Generated Receipt Dataset This directory contains the AI-generated receipts from the GPT4o-Receipt benchmark, introduced in: GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document ForensicsYan Zhang*, Simiao Ren*†, Ankit Raj, En Wei, Dennis Ng, Alex Shen, Jiayu Xue, Yuxin Zhang, Evelyn MarottaarXiv:2603.11442 · March 2026 · CC BY-NC-SA 4.0*Equal contribution. †Corresponding author: benren@scam.ai What Is GPT4o-Receipt? GPT4o-Receipt is… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt4o-receipt.imageimage-classificationn<1K1 likes66 downloads4mo agoHugging Face25Shouninger /ScamBenchgated SCAMBENCH: A Multi-Perspective Benchmark for Online Scam Communication SCAMBENCH is a dataset for studying online scam communication, scam detection, and model robustness. It contains real-world scam messages curated from victim-reported incidents, along with non-scam counterparts and structured annotations. The dataset is introduced in our paper: SCAMBENCH: A Multi-Perspective Benchmark for Analyzing and Evaluating Online Scam Communication Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Shouninger/ScamBench.text1K<n<10K0 likes66 downloads26d agoHugging Face26scampion /OHSDAextract from https://data.mendeley.com/datasets/c5mfhr2xcz/1 Offline Handwritten Signature Database based on Age Annotation (OHSDA) Published: 27 February 2023 Version 1 DOI: 10.17632/c5mfhr2xcz.1 Contributors: Sathish Kumar , Dr Shivanand Gornale Description Handwritten signature analysis is the endeavoring research in many verification and recognition system problem. As per our best of knowledge there is less number of publicly available datasets which have age class… See the full description on the dataset page: https://huggingface.co/datasets/scampion/OHSDA.image1K<n<10K1 likes65 downloads2y agoHugging Face27agnialf /Internship-Job-Scam-DatasetThis dataset is a custom-built text classification dataset for detecting fraud risk in internship and entry-level job postings. This dataset is created for an embeddings-based classifier that helps students evaluate whether internship and entry-level job postings may be legitimate, suspicious, or fraudulent. Dataset Summary The dataset contains job posting texts labeled into three risk categories: legitimate: postings that look like normal internship or entry-level job… See the full description on the dataset page: https://huggingface.co/datasets/agnialf/Internship-Job-Scam-Dataset.tabularn<1K0 likes65 downloads5mo agoHugging Face28flowxai /scamguardbench ScamGuardBench (v0.2) A labeled, multilingual (EN/RO) benchmark for scam/fraud message detection, with a first-class hard-legitimate subset. ScamGuardBench is the evaluation companion to the scam-guard detector. Its headline metric is false-positive rate on legitimate messages — because every failed consumer scam filter dies of false positives, not of missed scams. Why this benchmark exists A scam detector that flags real bank OTP messages gets disabled within a… See the full description on the dataset page: https://huggingface.co/datasets/flowxai/scamguardbench.texttext-classification1K<n<10K1 likes57 downloads3mo agoHugging Face29shakeleoatmeal /phone-scam-detection-syntheticGenerated using Llama 8B model Contains 1800 synthetic dialogues Varied conversation lengths (short/medium/long) Different fraud subtlety levels Balanced labels (legitimate vs fraudulent) 3 fraud types covered (support, ssn, refund) tabulartext-classification1K<n<10K0 likes55 downloads1y agoHugging Face30abdessamad-bourkibate /Scam-Religious-Detection Scam Religious Detection Dataset This dataset is designed for detecting religious-based scams, particularly on social media platforms like Twitter. It contains a collection of tweets categorized into various classes to facilitate the training of machine learning models for scam detection. Dataset Details Dataset Name: Scam-Religious-Detection Primary Language: Arabic (ar) License: MIT Total Size: ~4.89 GB Dataset Structure The dataset consists of several CSV… See the full description on the dataset page: https://huggingface.co/datasets/abdessamad-bourkibate/Scam-Religious-Detection.text-classification1B<n<10B0 likes54 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.