gplsi/fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.
🧠 BALANCEDFAKEJOBPOSTINGSEN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional modification to the text content.
📊 Dataset Summary
🧩 Dataset Details
🔹 Source Dataset
- Original: Fake Job Posting Prediction (Kaggle)
- License: CC0: Public Domain (as per original Kaggle dataset)
- Modifications:
- Dataset balanced to include 50% fraudulent and 50% non-fraudulent samples.
- All textual fields preserved in their original English form.
- All structural and semantic information retained from the original dataset.
🧱 Columns Description
💰 Funding
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública, co-financed by the EU – NextGenerationEU, within the framework of the project Desarrollo de Modelos ALIA.
License
This work is licensed under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence.
