datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recruitment-dataset-job-descriptions-english
Djinni Dataset (English Job Descriptions part)
Overview
The Djinni Recruitment Dataset (English Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-job-descriptions-english.recruitment-dataset-candidate-profiles-english
Djinni Dataset (English CVs part)
Overview
The Djinni Recruitment Dataset (English CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-english.Campus_Recruitment_CSV
Dataset Description
This data set consists of Placement data of students in a XYZ campus. Based on the student's performance data we are classifying his Placement Status.
The students report includes the following information:
CGPA - The grade of the student in his university
Internships - The no of internship done by the student before final placement
Projects - The no of projects done by the student
Workshops/Certifications - The no of workshops attended and the certifications… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_CSV.Campus_Recruitment_Text
Dataset Description
This data set consists of Placement data of students in a XYZ campus. Based on the student's performance report we are classifying his Placement Status. The dataset is derived from a csv data.
The Mistral7B model is used with data-to-text methodology to convert each of the rows in the csv data into a textual format for the LLM's, the conversion script is in this notebook.
The Prompt field is the prompt used on Mistral7B LLM and the response field is the… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_Text.Recruitment-Task-3
DeepWeeds - AI-MED AGH convenience mirror
This is a convenience mirror of the official DeepWeeds image archive and the
upstream annotations pinned to a specific commit. original/images.zip is
preserved unchanged; images are not extracted or duplicated here. models.zip
from the source authors is deliberately not mirrored.
Dataset facts
17,509 in-situ images from Queensland, Australia.
Nine classes: eight weed species plus Negative.
The authors publish five folds… See the full description on the dataset page: https://huggingface.co/datasets/AI-MED-AGH/Recruitment-Task-3.recruitment-dataset-candidate-profiles-ukrainian
Djinni Dataset (Ukrainian CVs part)
Overview
The Djinni Recruitment Dataset (Ukrainian CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-ukrainian.recruitment-dataset-job-descriptions-ukrainian
Djinni Dataset (Ukrainian Job Descriptions part)
Overview
The Djinni Recruitment Dataset (Ukrainian Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-job-descriptions-ukrainian.nichevault-uk-recruitment-agencies
NicheVault UK Recruitment Agencies — Free Sample Dataset
15 active UK companies registered under SIC 78109 — Other activities of employment placement agencies — collected from Companies House.
This public sample lets you inspect the formatting, provenance and exact 14-column schema used in the paid NicheVault dataset.
Dataset contents
15 company records
Active on the Companies House register at time of collection. Some records may carry an additional status… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-uk-recruitment-agencies.smoltrace-recruitment-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 101
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-recruitment-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-recruitment-tasks.Tech-Job-Scams-and-Predatory-Recruitment
Tech Job Scams & Predatory Recruitment Tactics
Dataset Description
An adversarial dataset containing 1,000 distinct communication strings, job descriptions, and direct messaging scripts modeling deceptive recruitment loops targeting remote software engineering, AI/ML, and data science talent.
Purpose and Impact
With the rapid scale of remote work, software developers have become primary targets for sophisticated hiring scams. These range from… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/Tech-Job-Scams-and-Predatory-Recruitment.clinical-quad-recruitment-selection-bias-protocol-pressure-operational-drift-v0.1Clarus Clinical Quad Coupling Recruitment Selection Bias Protocol Pressure Operational Drift v0.1
What this dataset isThis dataset tests whether a model can detect recruitment and selection bias caused by four interacting nodes.
Quad coupling nodes
Recruitment speed or site pressure
Eligibility or baseline data gaps
Operational or staffing drift
Governance or milestone pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys
recruitment_bias_risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-selection-bias-protocol-pressure-operational-drift-v0.1.clinical-trial-basin-enrichment-and-recruitment-v0.1What this dataset tests
Whether a system can design basin selective recruitment.
It must:
pick the eligible basin
write inclusion and exclusion rules
estimate signal gain and dilution risk
Required outputs
eligible_basin_definition
recruitment_filter_rules
signal_amplification_index
heterogeneity_dilution_risk
expected_trial_coherence_gain
exclusion_rationale
Use case
Trial rescue.
Smaller trials with stronger signals.
Reduced washout from basin mixing.
clinical_recruitment_coherence_mapping_v0.1Clinical Recruitment Coherence Mapping v0.1
Purpose
Detect when population and feasibility assumptions will break recruitment.
Model task
Return one JSON object
risk_levellow, medium, high
failure_modeone allowed label
correct_actionone short paragraph
Scoring
0 to 100
risk accuracy 30
failure mode accuracy 35
action similarity 25
format pass 10
Run
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
recruitment-dataset-job-descriptions-english
Djinni Dataset (English Job Descriptions part)
Overview
The Djinni Recruitment Dataset (English Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords, English… See the full description on the dataset page: https://huggingface.co/datasets/Lucas44/recruitment-dataset-job-descriptions-english.recruitment_data
Fictional IT Recruitment Dataset
This dataset contains 10,000 synthetically generated recruitment records for a fictional IT company from 2020 to 2026. It is designed to be used for Information Retrieval and Classification tasks, such as predicting candidate acceptance based on their CV embeddings.
Dataset Structure
Each record contains the following fields:
recruitment_date: Date the recruitment took place (YYYY-MM-DD).
opening_position: The title of the job… See the full description on the dataset page: https://huggingface.co/datasets/uhfew/recruitment_data.clinical-quad-recruitment-coherence-mapping-suite-v0.1Clarus Clinical Quad Coupling Recruitment Coherence Mapping Suite v0.1
What this dataset isThis dataset tests whether a model can detect recruitment incoherence under four-node coupling pressure.
Quad coupling nodes
Biological eligibility definition
Concomitant medication or background therapy filters
Operational measurement and site process variance
Governance constraints limiting protocol flexibility
Input
One recruitment vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-coherence-mapping-suite-v0.1.africa-worldbank-wbl-legal-framework-work-the-law-prohibits-discrimination-in-recruitment-based
WBL: Legal Framework, Work, The law prohibits discrimination in recruitment based on marital status, parental status, or age | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-legal-framework-work-the-law-prohibits-discrimination-in-recruitment-based.fisheries-spawning-recruitment-coherence-risk-v0.1What this repo is for
Detect when spawning success stops translating into new stock.
Early warning of collapse.
Focus
• spawning biomass
• juvenile survival
• recruitment lag
• environmental pressure
linkedin_recruitment_questions_embeddedclinical-quad-recruitment-inclusion-severity-endpoint-dilution-v0.1Clinical Quad Recruitment Inclusion Severity Endpoint Dilution v0.1
Each row is a site week snapshot.
Core quad
Recruitment speedInclusion criteria tightnessPopulation severityEndpoint dilution
Target
label_primary_miss_next_60d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
recruitment_data
Fictional IT Recruitment Dataset
This dataset contains 10,000 synthetically generated recruitment records for a fictional IT company from 2020 to 2026. It is designed to be used for Information Retrieval and Classification tasks, such as predicting candidate acceptance based on their CV embeddings.
Dataset Structure
Each record contains the following fields:
recruitment_date: Date the recruitment took place (YYYY-MM-DD).
opening_position: The title of the job position… See the full description on the dataset page: https://huggingface.co/datasets/katsufu/recruitment_data.clinical-quad-recruitment-protocol-adherence-outcome-drift-v0.1Clinical Quad Recruitment Protocol Adherence Outcome Drift v0.1
Each row is a site week snapshot.
Core quad
Recruitment qualityProtocol deviationVisit adherenceOutcome drift
Target
label_trial_fail_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
Columns… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-protocol-adherence-outcome-drift-v0.1.clinical-quad-budget-burn-recruitment-pace-timeline-slip-sponsor-stop-decision-v0.1Clinical Quad Budget Burn Recruitment Pace Timeline Slip Sponsor Stop Decision v0.1
Each row is a monthly financial and operational snapshot.
Core quad
Budget burnRecruitment paceTimeline slipSponsor decision pressure
Target
label_sponsor_stop_next_60d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
asia-owid-openness-of-executive-recruitment-score
Openness Of Executive Recruitment Score | Asia (Our World in Data)
🌏 4,915 observations · 46 Asia countries · 1800–2018 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,915 observations of Openness Of Executive Recruitment Score data across 46 Asia countries, spanning 1800–2018.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Openness Of Executive Recruitment Score… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-openness-of-executive-recruitment-score.africa-owid-competitiveness-of-executive-recruitment-score
Competitiveness Of Executive Recruitment Score | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-competitiveness-of-executive-recruitment-score.recruitment-dataset-job-descriptions-english
Djinni Dataset (English Job Descriptions part)
Overview
The Djinni Recruitment Dataset (English Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords, English… See the full description on the dataset page: https://huggingface.co/datasets/WQFFWF/recruitment-dataset-job-descriptions-english.asia-owid-competitiveness-of-executive-recruitment-score
Competitiveness Of Executive Recruitment Score | Asia (Our World in Data)
🌏 4,915 observations · 46 Asia countries · 1800–2018 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,915 observations of Competitiveness Of Executive Recruitment Score data across 46 Asia countries, spanning 1800–2018.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Competitiveness Of Executive… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-competitiveness-of-executive-recruitment-score.recruitment-dataset-job-descriptions-english-sample-pt5recruitment-dataset-candidate-profiles-english
Djinni Dataset (English CVs part)
Overview
The Djinni Recruitment Dataset (English CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/pylover/recruitment-dataset-candidate-profiles-english.recruitment-dataset-job-descriptions-english
Djinni Dataset (English Job Descriptions part)
Overview
The Djinni Recruitment Dataset (English Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords, English… See the full description on the dataset page: https://huggingface.co/datasets/Moheez2611/recruitment-dataset-job-descriptions-english.
