K-University-AIED/online_fraud_dataset
Digital Forensic Investigation Scenario Dataset (Online Fraud) This dataset is a structured collection of digital evidence for research in digital forensics and criminal investigation. It simulates a multi-stage Online Fraud (Voice Phishing) scenario. All metadata and evidence linkages are defined in the accompanying metadata.csv file. Dataset Structure The dataset features follow the exact structure of the provided metadata file, ensuring data integrity and… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/online_fraud_dataset.
Digital Forensic Investigation Scenario Dataset (Online Fraud)
This dataset is a structured collection of digital evidence for research in digital forensics and criminal investigation. It simulates a multi-stage Online Fraud (Voice Phishing) scenario. All metadata and evidence linkages are defined in the accompanying metadata.csv file.
Dataset Structure
The dataset features follow the exact structure of the provided metadata file, ensuring data integrity and consistency for research purposes.
Metadata Features
- sample_id: Unique identifier for each data record.
- file_name: Standardized dataset file name.
- category: Criminal category classification (
online_fraud). - label: Investigative relevance (
positive/negative). - description: Detailed role of the evidence in the investigation scenario.
- source: Data generation origin (
synthetic). - type: Data format or type (e.g.,
str,int,pdf).
Evidence Inventory
Project Information
This dataset is a result of the Glocal University 30 Project, supported by the Ministry of Education and the National Research Foundation of Korea. It was developed at the Hallym Intelligent Social Safety Research Lab.
- Project ID: GLOCAL-202407990001
- Contact: shha@hallym.ac.kr
License
Distributed under the CC BY NC ND 4.0 license.
디지털 포렌식 수사 시나리오 데이터셋 (Online Fraud)
소개
본 데이터셋은 가상의 온라인 사기 사건을 수사하는 과정을 재구성한 연구용 데이터입니다. 업로드된 metadata.csv의 규격에 따라 모든 데이터 파일이 정의되어 있으며, 수사적 유의미성(label)과 데이터 형식(type)을 실제 메타데이터와 1:1로 매칭하여 관리합니다.
데이터 명세
제공해주신 metadata.csv의 컬럼 구조를 그대로 반영하였습니다.
- sample_id: 데이터 고유 식별 번호
- file_name: 표준 영문 파일명
- category: 범죄 분류 (
online_fraud) - label: 수사 유의미성 (
positive: 핵심 증거,negative: 참고용 데이터) - description: 시나리오 내 해당 파일의 수사적 의미 및 역할 (영문 설명)
- source: 데이터 생성 출처 (
synthetic) - type: 데이터 기술 형식 (str, int, pdf 등)
라이선스
본 데이터셋은 CC BY NC ND 4.0 라이선스를 따릅니다. 출처 명시 시 비영리 목적의 교육 및 연구 용도로 사용 가능합니다.
소속: 한림대학교 지능형사회안전연구소 (Hallym Intelligent Social Safety Research Lab)
