CoolFace
Datasetpublic

chisomobanja/MOZ-Smishing

Dataset Summary MOZ-Smishing is a benchmark dataset specifically designed for detecting smishing attacks targeting Mobile Money Transfer (MMT) systems. This dataset addresses the critical lack of publicly available SMS phishing datasets in this domain, especially for non-English languages. It comprises crowd-sourced text messages from Mozambican mobile users, meticulously annotated into two categories: legitimate messages and smishing attempts. The messages are primarily in… See the full description on the dataset page: https://huggingface.co/datasets/chisomobanja/MOZ-Smishing.

sourceHugging Facecreativeml-openrail-mupdated 2mo agoView on Hugging Face
0likes7downloads
Dataset Card

Dataset Summary

MOZ-Smishing is a benchmark dataset specifically designed for detecting smishing attacks targeting Mobile Money Transfer (MMT) systems. This dataset addresses the critical lack of publicly available SMS phishing datasets in this domain, especially for non-English languages. It comprises crowd-sourced text messages from Mozambican mobile users, meticulously annotated into two categories: legitimate messages and smishing attempts. The messages are primarily in Portuguese, often incorporating microtext styles and linguistic nuances unique to the Mozambican context.

The dataset contains 552 instances of smishing messages and 2,009 legitimate text messages, totaling 2,561 instances.

Supported Tasks and Leaderboards

The dataset supports the task of SMS spam detection or fraud detection, specifically identifying smishing attempts related to Mobile Money Transfer. The paper also explores the effectiveness of Large Language Models (LLMs) in this task using in-context learning approaches.

Languages

The dataset consists of text messages in Portuguese, specifically from Mozambique.