datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-apply-jobs
Open-Apply Jobs
A daily-refreshed open dataset of active job postings sourced directly from public ATS APIs (Greenhouse, Lever, Ashby). Every record can be traced back to the hiring company's own career board.
Refresh: automated daily at 06:00 UTC
Partitioning: Hive-partitioned Parquet (date=YYYY-MM-DD/source={ats})
Source code: https://github.com/edwarddgao/openapply
Usage
from datasets import load_dataset
ds = load_dataset('edwarddgao/open-apply-jobs')
#… See the full description on the dataset page: https://huggingface.co/datasets/edwarddgao/open-apply-jobs.job-bench
JobBench
Real-world white-collar tasks across ~35 professions. Each task has a prompt,
reference files, a weighted rubric, and a human-readable brief.
The dataset ships two splits:
main — 65 full tasks. Some include a files_required_to_search/ folder
of ground-truth materials the agent is expected to discover via search.
easy — 63 simplified tasks (shorter prompts, no files_required_to_search/).
Useful for cheaper smoke tests and capability ranking.
The main and easy splits… See the full description on the dataset page: https://huggingface.co/datasets/JobBench/job-bench.data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/lukebarousse/data_jobs.real-or-fake-fake-jobposting-predictionfake_job_postings2
Dataset Card for "fake_job_postings2"
More Information needed
jobsfake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.resume-job-description-fiteu-tech-jobs
eu-tech-jobs
Daily-updated open-data feed of jobs from EU AI/tech and remote-EU companies.
Source repo: https://github.com/Aramente/eu-tech-jobs
Live site: https://aramente.github.io/eu-tech-jobs/
License: CC BY 4.0 (data) + MIT (pipeline code)
What's in here
Path
Contents
latest/jobs.parquet
Most recent snapshot, all active jobs
latest/companies.parquet
Curated company list with categories + ATS handles
latest/metadata.json
Pipeline run metadata… See the full description on the dataset page: https://huggingface.co/datasets/Aramente/eu-tech-jobs.remote-jobs
Jobicy Remote Jobs
An automatically updated dataset of remote job opportunities published by Jobicy.
The dataset is designed for developers, researchers, analysts, search systems, AI agents, RAG applications, labor-market research, career tools, and other applications that need structured remote-job data.
Data source
The source data comes from the public Jobicy Remote Jobs API:
https://jobicy.com/api/v2/remote-jobs
Each record includes a canonical Jobicy job URL… See the full description on the dataset page: https://huggingface.co/datasets/jobicy/remote-jobs.nexus-jobsjob-titles
Comprehensive Job Titles Dataset
A high-quality, deduplicated dataset of 65,248 unique job titles compiled from authoritative sources including ESCO (European Skills, Competences, Qualifications and Occupations), O*NET (Occupational Information Network), and OSCA (Occupational Skills and Competencies Australia).
Dataset Description
This dataset provides a comprehensive collection of job titles that have been carefully processed to remove duplicates and near-duplicates… See the full description on the dataset page: https://huggingface.co/datasets/gpriday/job-titles.job-descriptionsjobs-dalle-2
Dataset Card for "dataset-dalle"
More Information needed
linkedin-job-postingsjob-dataset
Dataset Card for Dataset Name
JobStreet Job Postings Dataset
Dataset Details
Dataset Description
This dataset compiles a comprehensive range of job listings from JobStreet, offering a detailed view of the current employment landscape across various industries in Malaysia. It includes key features such as unique job IDs, titles, company names, locations, job roles, categories, subcategories, job types, salaries, and detailed descriptions. The motivation behind… See the full description on the dataset page: https://huggingface.co/datasets/azrai99/job-dataset.jobs
Lynceus Open Job Index
239,487 open roles at 6,432 companies, 43,062 of them
remote. Read directly from each employer's own careers page and public ATS
feed — never aggregated or reposted from a job board.
Last updated: 2026-09-25 04:21 UTC
What this is
Most job datasets are scraped from aggregators, which means they are a copy of
a copy: stale, deduplicated badly, and full of listings that were filled weeks
ago. This one is read at the source — each employer's… See the full description on the dataset page: https://huggingface.co/datasets/Lynceus/jobs.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.fake_job_postings2_all_text
Dataset Card for "fake_job_postings2_all_text"
More Information needed
real-or-fake-fake-jobposting-predictionupcoming-layoffs-job-cuts-plant-closings-us-warn-act
Upcoming US layoffs: 458 WARN notices take effect in the next 90 days - 44,620 workers, rebuilt daily
Rebuilt 2026-09-25. Window: 2026-09-25 to 2026-12-24.
Every other US layoffs dataset - including our own notice-level one - tells you
what was announced. This one tells you whose job ends next. One row per WARN Act notice
whose separation date falls inside the next 90 days, sorted by how soon. The people in these
rows are, as of the rebuild date above, still employed.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/upcoming-layoffs-job-cuts-plant-closings-us-warn-act.job_titles_and_descriptions
IT Job Roles, Skills, and Descriptions Dataset
This dataset provides detailed information about various IT job roles, the required skills for each role, and the job descriptions that outline the responsibilities and qualifications associated with each position. It is designed for use in applications such as career guidance systems, job recommendation engines, and educational tools aimed at aligning skills with industry demands.
Dataset Upload
This dataset is uploaded… See the full description on the dataset page: https://huggingface.co/datasets/NxtGenIntern/job_titles_and_descriptions.Jobs-and-Development-Indicators-For-African-Countries
Jobs and Development Indicators For African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Jobs-and-Development-Indicators-For-African-Countries.recruitment-dataset-job-descriptions-english
Djinni Dataset (English Job Descriptions part)
Overview
The Djinni Recruitment Dataset (English Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-job-descriptions-english.canadian-ds-ml-job-postingsJobHop
JobHop
A large-scale public dataset of career trajectories derived from pseudonymized resumes provided by VDAB, the public employment service in Flanders, Belgium. Job experiences are standardized to ESCO occupation codes and carry quarter-level temporal information.
Two versions are available: JobHop v1 (the original release, ~360,000 resumes) and JobHop v2 (the current release, 355,315 trajectories from a larger corpus, built with a substantially improved extraction pipeline).… See the full description on the dataset page: https://huggingface.co/datasets/aida-ugent/JobHop.open-jobs-daily
Open Jobs Daily 🌍💼
Commercial vendors often charge upwards of $1,000/month for firehose access to global job market data. This dataset democratizes that access.
The main creator of this dataset is Reddit user OminousLatinWord. For convenience, I converted the dataset to Parquet files and uploaded it to Hugging Face.
Source Data & Attribution
Creator: Created and originally open-sourced by Reddit user OminousLatinWord under a CC0 license.
Source Release:… See the full description on the dataset page: https://huggingface.co/datasets/Yigit-Karaman/open-jobs-daily.job-skill-set
Job Skill Set
Description
The Job Skill Set Dataset is designed for use in machine learning projects related to job matching, skill extraction, and natural language processing tasks. The dataset includes detailed information about job roles, descriptions, and associated skill sets, enabling developers and researchers to build and evaluate models for career recommendation systems, resume parsing, and skill inference.
Dataset Source
This dataset was initially… See the full description on the dataset page: https://huggingface.co/datasets/batuhanmtl/job-skill-set.Steve_Jobs_Interviews
Steve Jobs Interviews Database
Support this project on Ko-fi
Project Overview
This project contains multiple interviews of Steve Jobs during his time before and after Apple.
Goal
The primary goal of this dataset was to fine-tune a language model to output Steve Jobs views and thoughts.
Performance
The performance of this small dataset is very noteworthy. Do to the nature of the database being interview question and answer pairs the replies of the… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/Steve_Jobs_Interviews.real-or-fake-fake-jobposting-prediction
