datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-apply-jobs
Open-Apply Jobs
A daily-refreshed open dataset of active job postings sourced directly from public ATS APIs (Greenhouse, Lever, Ashby). Every record can be traced back to the hiring company's own career board.
Refresh: automated daily at 06:00 UTC
Partitioning: Hive-partitioned Parquet (date=YYYY-MM-DD/source={ats})
Source code: https://github.com/edwarddgao/openapply
Usage
from datasets import load_dataset
ds = load_dataset('edwarddgao/open-apply-jobs')
#… See the full description on the dataset page: https://huggingface.co/datasets/edwarddgao/open-apply-jobs.CLEVER_Job_Simulator
The Capturing and Logging Ecological Virtual Experiences and Reality (CLEVER) — Job Simulator Dataset
This repository consists of data from 95 participants playing the SteamVR game "Job Simulator". We collected data using two VR systems: a Meta Quest 3S and a VIVE XR Elite. Each participant experienced 5 sessions in which they played through the first three tasks in Office Worker, responded to the Fidelity-based Presence Scale and Simulator Sickness Questionnaire, played through… See the full description on the dataset page: https://huggingface.co/datasets/xraijobsimulator/CLEVER_Job_Simulator.jobs-demo
Bag-of-Documents — Unified Jobs Artifacts
Companion artifacts for the unified jobs search Space.
Contains: titles + slim metadata + full metadata (JSONL) + bge-small + te3-large @ 1024 catalog vectors + pre-encoded te3 query cache (~196k popular queries) for 347,900 job postings across 4 corpora (Open-Apply, LinkedIn, JobStreet, USAJobs).
Source: github.com/dtunkelang/bag-of-documents
job-bench
JobBench
Real-world white-collar tasks across ~35 professions. Each task has a prompt,
reference files, a weighted rubric, and a human-readable brief.
The dataset ships two splits:
main — 65 full tasks. Some include a files_required_to_search/ folder
of ground-truth materials the agent is expected to discover via search.
easy — 63 simplified tasks (shorter prompts, no files_required_to_search/).
Useful for cheaper smoke tests and capability ranking.
The main and easy splits… See the full description on the dataset page: https://huggingface.co/datasets/JobBench/job-bench.data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/lukebarousse/data_jobs.jobseek-agent-traces
Jobseek Agent Traces
Claude Code agent session traces from jobseek — a job posting monitor for company career pages.
Each trace captures a complete agent workflow session: company discovery, board configuration, monitor/scraper selection, and quality validation. These are raw session transcripts, not tabular data — use the trace viewer to explore them.
Structure
traces/
{company-slug}/
{date}.jsonl # One trace per session (header + records)
Each .jsonl… See the full description on the dataset page: https://huggingface.co/datasets/viktor-shcherb/jobseek-agent-traces.real-or-fake-fake-jobposting-predictionfootball-holefinder-jobsStockChina-Minute
A-Share Minute-Level Historical Data
Dataset Description
This dataset contains minute-level trading data for Chinese A-share stocks from 2005 to 2023, covering 5267 stocks with complete historical trading records.
Data Format
Each CSV file corresponds to one stock and contains the following fields:
Field
Description
open
Opening price
close
Closing price
high
Highest price
low
Lowest price
volume
Trading volume
money
Trading amount
avg… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/StockChina-Minute.fake_job_postings2
Dataset Card for "fake_job_postings2"
More Information needed
jobsprocessed_fake_job_postingsLM-PDE-Save
PINN-Data-Save
combine
cat MOL-LLM/MOL-LLM.tar.part* > MOL-LLM.tar
tar -xvf MOL-LLM.tar
fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.resume-job-description-fiteu-tech-jobs
eu-tech-jobs
Daily-updated open-data feed of jobs from EU AI/tech and remote-EU companies.
Source repo: https://github.com/Aramente/eu-tech-jobs
Live site: https://aramente.github.io/eu-tech-jobs/
License: CC BY 4.0 (data) + MIT (pipeline code)
What's in here
Path
Contents
latest/jobs.parquet
Most recent snapshot, all active jobs
latest/companies.parquet
Curated company list with categories + ATS handles
latest/metadata.json
Pipeline run metadata… See the full description on the dataset page: https://huggingface.co/datasets/Aramente/eu-tech-jobs.remote-jobs
Jobicy Remote Jobs
An automatically updated dataset of remote job opportunities published by Jobicy.
The dataset is designed for developers, researchers, analysts, search systems, AI agents, RAG applications, labor-market research, career tools, and other applications that need structured remote-job data.
Data source
The source data comes from the public Jobicy Remote Jobs API:
https://jobicy.com/api/v2/remote-jobs
Each record includes a canonical Jobicy job URL… See the full description on the dataset page: https://huggingface.co/datasets/jobicy/remote-jobs.nexus-jobsHPLT2.0_cleanedThis is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to https://hplt-project.org/datasets/v2.0
The Cleaned variant of HPLT Datasets v2.0
This is the cleaned variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here.
The original JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/HPLT2.0_cleaned.job-titles
Comprehensive Job Titles Dataset
A high-quality, deduplicated dataset of 65,248 unique job titles compiled from authoritative sources including ESCO (European Skills, Competences, Qualifications and Occupations), O*NET (Occupational Information Network), and OSCA (Occupational Skills and Competencies Australia).
Dataset Description
This dataset provides a comprehensive collection of job titles that have been carefully processed to remove duplicates and near-duplicates… See the full description on the dataset page: https://huggingface.co/datasets/gpriday/job-titles.minddistiller-archived-jobs-20260615
MindDistiller encrypted archived jobs
Repository: PGCodeLLM/minddistiller-archived-jobs-20260615
Generated at: 2026-06-30T02:51:28.101982+00:00
Encrypted archive count: 1004
Each *.tar.gz.gpg file is one encrypted MindDistiller job archive.
The plaintext source files were *.tar.gz archives from data/archives.
Decrypt one file with:
gpg --output JOB_ID.tar.gz --decrypt JOB_ID.tar.gz.gpg
manifest.jsonl contains source filename, encrypted filename, size, and SHA-256 checksums.
freehire-jobs
freehire jobs dataset
Raw IT job postings aggregated by freehire.me
(freehire.dev), plus deterministic dictionary-derived facets (skills,
category, seniority, location, work mode, salary where extracted).
Website: https://freehire.me
Source code: https://github.com/strelov1/freehire
One JSON object per line, gzip-compressed, split into 20 shards by internal
id range.
Private/user-submitted postings are excluded. One row with an unrecoverable
corrupt description (Postgres TOAST… See the full description on the dataset page: https://huggingface.co/datasets/istrelov/freehire-jobs.Zyda-2
Zyda-2
Zyda-2 is a 5 trillion token language modeling dataset created by collecting open and high quality datasets and combining them and cross-deduplication and model-based quality filtering. Zyda-2 comprises diverse sources of web data, highly educational content, math, code, and scientific papers.
To construct Zyda-2, we took the best open-source datasets available: Zyda, FineWeb, DCLM, and Dolma. Models trained on Zyda-2 significantly outperform identical models trained on the… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/Zyda-2.jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset = load_dataset("jat-project/jat-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/jat-dataset.job-descriptionsjobs-scriptsjobs-dalle-2
Dataset Card for "dataset-dalle"
More Information needed
linkedin-job-postingsjob-dataset
Dataset Card for Dataset Name
JobStreet Job Postings Dataset
Dataset Details
Dataset Description
This dataset compiles a comprehensive range of job listings from JobStreet, offering a detailed view of the current employment landscape across various industries in Malaysia. It includes key features such as unique job IDs, titles, company names, locations, job roles, categories, subcategories, job types, salaries, and detailed descriptions. The motivation behind… See the full description on the dataset page: https://huggingface.co/datasets/azrai99/job-dataset.jobs
Lynceus Open Job Index
239,788 open roles at 6,477 companies, 43,018 of them
remote. Read directly from each employer's own careers page and public ATS
feed — never aggregated or reposted from a job board.
Last updated: 2026-09-24 04:20 UTC
What this is
Most job datasets are scraped from aggregators, which means they are a copy of
a copy: stale, deduplicated badly, and full of listings that were filled weeks
ago. This one is read at the source — each employer's… See the full description on the dataset page: https://huggingface.co/datasets/Lynceus/jobs.
