chisomobanja/MOZ-Smishing
Dataset Summary MOZ-Smishing is a benchmark dataset specifically designed for detecting smishing attacks targeting Mobile Money Transfer (MMT) systems. This dataset addresses the critical lack of publicly available SMS phishing datasets in this domain, especially for non-English languages. It comprises crowd-sourced text messages from Mozambican mobile users, meticulously annotated into two categories: legitimate messages and smishing attempts. The messages are primarily in… See the full description on the dataset page: https://huggingface.co/datasets/chisomobanja/MOZ-Smishing.
Dataset Summary
MOZ-Smishing is a benchmark dataset specifically designed for detecting smishing attacks targeting Mobile Money Transfer (MMT) systems. This dataset addresses the critical lack of publicly available SMS phishing datasets in this domain, especially for non-English languages. It comprises crowd-sourced text messages from Mozambican mobile users, meticulously annotated into two categories: legitimate messages and smishing attempts. The messages are primarily in Portuguese, often incorporating microtext styles and linguistic nuances unique to the Mozambican context.
The dataset contains 552 instances of smishing messages and 2,009 legitimate text messages, totaling 2,561 instances.
Supported Tasks and Leaderboards
The dataset supports the task of SMS spam detection or fraud detection, specifically identifying smishing attempts related to Mobile Money Transfer. The paper also explores the effectiveness of Large Language Models (LLMs) in this task using in-context learning approaches.
Languages
The dataset consists of text messages in Portuguese, specifically from Mozambique.
