datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
resume-json-extraction-5k
Dataset Card for resume-json-extraction-5k
Dataset Description
This dataset contains 4,879 resume examples formatted for fine-tuning language models to extract structured JSON information from resume text.
Dataset Summary
The dataset consists of resume text paired with structured JSON outputs containing:
Job titles (current and previous)
Companies (current and previous)
Years of experience
Seniority level
Primary domain and industries
Core and secondary skills… See the full description on the dataset page: https://huggingface.co/datasets/sandeeppanem/resume-json-extraction-5k.resume-conversations-llm-training
📄 Resume Conversations for LLM Training
High-quality conversational dataset for building AI that understands resumes, careers, and professional growth.Created and maintained by Syncora.ai.
✅ Overview
This dataset provides resume-related conversations in a structured JSONL format, ideal for developers and AI practitioners working on chatbots, career advisory tools, or LLM fine-tuning. It includes realistic Q&A on career development, technology trends, and professional… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/resume-conversations-llm-training.bias_resume_public
Resume Bias Summaries
Counterfactual LLM-generated resume summaries with race-conditioned candidate names, MiniCheck factual-support scores, and a cross-judge annotation sample. Companion dataset to the paper "Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring" (Nghiem et al., 2026).
The dataset enables direct replication of the paper's bias and stability analyses without re-running 50+ GPU-hours of LLM inference. The… See the full description on the dataset page: https://huggingface.co/datasets/nghiemhnlp/bias_resume_public.my-resume-v2
Andrew Stanley Resume Q&A
A small, hand-curated chat-format dataset of 122 question/answer pairs covering the professional background, career history, technical skills, certifications, and military service of Andrew Stanley, CTO / Chief Innovation Officer at SMS Data Products Group (McLean, VA).
The dataset is purpose-built for two things:
A working demonstration of an end-to-end LLM fine-tuning workflow — source document → synthetic Q&A generation → QLoRA fine-tune → GGUF export →… See the full description on the dataset page: https://huggingface.co/datasets/2stacks/my-resume-v2.moltbook-entropy-collapse-resumes
MoltBook Entropy Collapse — Resumed Runs
Eight of the 48 canonical entropy-collapse runs ended with an empty
45–60 min bin (i.e. the agent population stopped posting before the
hour was up). Causes were:
GPT-5 (2 runs): wall-clock batch termination at ~43 min.
Gemini Flash Lite (6 runs): provider-side empty-completion dropout —
agents transition near-simultaneously from real generations
(~1500–2000 ms) to ~70–130 ms empty stream events with no assistant
text. See… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-entropy-collapse-resumes.resumeaiResume_Best_PracticesResume-Job-TextHello I hope you are doing well
This dataset is in llama template format to tune.
Feel free to use and Contirbute if possible
iteratecv-resume-tailoring
IterateCV Resume Tailoring Dataset
Dataset Description
An instruction-tuning dataset for fine-tuning LLMs to tailor resumes to job descriptions.
Built for the IterateCV project.
Each example contains:
instruction: Task description for the model
input: A master resume in JSON format + a job description
output: A tailored version of the resume optimized for the job description
Dataset Construction
Source resume-JD pairs from… See the full description on the dataset page: https://huggingface.co/datasets/abhaykanjoor/iteratecv-resume-tailoring.bias_resume_public
Resume Bias Summaries
Counterfactual LLM-generated resume summaries with race-conditioned candidate names, MiniCheck factual-support scores, and a cross-judge annotation sample. Companion dataset to the paper "Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring" (Nghiem et al., 2026).
The dataset enables direct replication of the paper's bias and stability analyses without re-running 50+ GPU-hours of LLM inference. The… See the full description on the dataset page: https://huggingface.co/datasets/Saad222222222222/bias_resume_public.smolified-resumex-ai
🤏 smolified-resumex-ai
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model ankan288/smolified-resumex-ai.
📦 Asset Details
Origin: Smolify Foundry (Job ID: df4da207)
Records: 9799
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by ankan288.
Generated via Smolify.ai.
smolified-resume-analyzer
🤏 smolified-resume-analyzer
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model AnusmitaRC21/smolified-resume-analyzer.
📦 Asset Details
Origin: Smolify Foundry (Job ID: c0513893)
Records: 3438
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by AnusmitaRC21.
Generated via Smolify.ai.
