CoolFace
Datasetpublic

0xkamal7/hr-jd-bias-audit

JD-BiasAudit JD-BiasAudit is a provenance-tracked instruction-tuning dataset for HR teams, compliance reviewers, and model builders who need to neutralize coded language in job descriptions without deleting legitimate requirements. It derives structured, span-grounded audits from real postings in lang-uk/recruitment-dataset-job-descriptions-english. The upstream corpus provides job descriptions, not paired neutral rewrites, exact removed spans, protected-attribute proxy… See the full description on the dataset page: https://huggingface.co/datasets/0xkamal7/hr-jd-bias-audit.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes28downloads
Dataset Card

JD-BiasAudit

JD-BiasAudit is a provenance-tracked instruction-tuning dataset for HR teams, compliance reviewers, and model builders who need to neutralize coded language in job descriptions without deleting legitimate requirements. It derives structured, span-grounded audits from real postings in lang-uk/recruitment-dataset-job-descriptions-english. The upstream corpus provides job descriptions, not paired neutral rewrites, exact removed spans, protected-attribute proxy categories, preserved-requirement lists, or residual-risk reports. This derived supervision therefore did not already exist in the open source.

What this data produced

Model: meta-llama/Llama-3.2-3B-Instruct | Adaptive Data quality: 6.0 -> 9.8 (A)

The target behavior is narrow and judgeable: remove the coded phrase, explain why it is risky, and retain the concrete skills, experience, duties, and constraints that make the posting useful.

Schema

Each line in train-tagged.jsonl is a JSON object. completion is stored as a JSON-encoded string, not as a nested object.

json
{
  "prompt": "string: audit instruction followed by JOB DESCRIPTION and the source posting",
  "completion": "string containing JSON with neutral_jd, removed_phrases, preserved_requirements, and residual_risks",
  "shard": "djinni_jd_bias",
  "license": "mit",
  "source": "lang-uk/recruitment-dataset-job-descriptions-english"
}

After decoding completion, its exact shape is:

json
{
  "neutral_jd": "string",
  "removed_phrases": [
    {
      "phrase": "string",
      "category": "string",
      "why": "string",
      "suggested_fix": "string"
    }
  ],
  "preserved_requirements": ["string"],
  "residual_risks": ["string"]
}

Real training example

This is the first row of data/hr-jd-bias/train-tagged.jsonl. The completion is decoded only for readability; its values are copied from the file.

<details> <summary>Show example</summary>

json
{
  "prompt": "You are an HR job description bias auditor. Rewrite the job description below to remove coded or biased language while preserving every genuine requirement. Return strict JSON with four fields: neutral_jd (the rewritten text), removed_phrases (each with the exact phrase, the protected-attribute category it proxies, and why), preserved_requirements (concrete skills, experience, and duties that must survive), and residual_risks (any remaining wording a compliance reviewer should check). Never infer or mention protected attributes. Never delete a legitimate requirement.\n\nJOB DESCRIPTION:\nJob responsibilities:\r\nMonitor incoming/outgoing data as well as create/modify code for processes relating to incoming/outgoing data feeds\r\n- Prepare, maintain the data migration processes inside RDB\r\n- Create, maintain, and optimize database elements using stored procedures, views, functions, triggers, jobs, and execution plan analysis\r\n- Experience using SQL Profiler and Extended Events to troubleshoot and optimize new and existing database code\r\n- Proactively address data quality concerns; troubleshoot errors\r\n- Apply/implement business logic for processes\r\n- Provide daily support for database activities\r\n- Manage security\r\n \r\nJob Skills/Requirements:\r\n● MS SQL Server with T-SQL – minimum 2 plus years’ experience\r\n● Experience with SQL Server 2008 R2 through 2017.\r\n● Visual Studuo + SSDT \r\n● Experience with SSRS, creating new queries / reports and setting up and maintaining services and shared service components.\r\n● SSIS experience\r\n● Plus skills:  XML, Web protocols, TFS,  ASP.NET\r\n● Familiarity with .Net Core, Visual Studio, and Powershell for SQL Server and Azure SQL \r\n\r\nWe offer: \r\n• Competitive salary \r\n• Comfortable office, located in the city center with bike parking\r\n• Official employment, paid sick leave, 20 days of paid vacation\r\n• Paid lunch\r\n• Free English lessons\r\n• Flexible work schedule\r\n• Official Ukrainian holidays are non-working days\r\n• Career development plan\r\n• Friendly team of young IT professionals \r\n\r\nAbout the project:\r\nProject is uniquely configurable cloud-based platform is designed to guide and reinforce behaviors that drive results; mobilize employees; strategically align individuals with company values and objectives; and solve an organization’s ever-changing business issues. The system’s Activity Streams and Leaderboard allow users to see how they compare alongside others in the organization, which helps businesses create a culture of recognition where individuals feel more connected to the company and their coworkers through Manager Discretionary, Peer-to-Peer, Service Anniversary & Milestones and Sales Incentive Programs\r\n\r\n \r\nTeam:  4 Senior .NET dev, 1 Middle  .NET dev, 5 Manual QA, 1 Senior Automation QA, 3 Senior DB dev, 1 Junior DB dev, 3 Markup/Front-End devs. \r\n\r\nВсего в Харьковском офисе 65 + человек.\r\nГрафик: 5/2, 8 часовой рабочий день, гибкий график. Без поздних митингов и овертаймов.",
  "completion": {
    "neutral_jd": "Job responsibilities: Monitor incoming/outgoing data as well as create/modify code for processes relating to incoming/outgoing data feeds - Prepare, maintain the data migration processes inside RDB - Create, maintain, and optimize database elements using stored procedures, views, functions, triggers, jobs, and execution plan analysis - Experience using SQL Profiler and Extended Events to troubleshoot and optimize new and existing database code - Proactively address data quality concerns; troubleshoot errors - Apply/implement business logic for processes - Provide daily support for database activities - Manage security Job Skills/Requirements: ● MS SQL Server with T-SQL – minimum 2 plus years’ experience ● Experience with SQL Server 2008 R2 through 2017. ● Visual Studuo + SSDT ● Experience with SSRS, creating new queries / reports and setting up and maintaining services and shared service components. ● SSIS experience ● Plus skills: XML, Web protocols, TFS, ASP.NET ● Familiarity with .Net Core, Visual Studio, and Powershell for SQL Server and Azure SQL We offer: • Competitive salary • Comfortable office, located in the city center with bike parking • Official employment, paid sick leave, 20 days of paid vacation • Paid lunch • Free English lessons • Flexible work schedule • Official Ukrainian holidays are non-working days • Career development plan • Friendly team of IT professionals About the project: Project is uniquely configurable cloud-based platform is designed to guide and reinforce behaviors that drive results; mobilize employees; strategically align individuals with company values and objectives; and solve an organization’s ever-changing business issues. The system’s Activity Streams and Leaderboard allow users to see how they compare alongside others in the organization, which helps businesses create a culture of recognition where individuals feel more connected to the company and their coworkers through Manager Discretionary, Peer-to-Peer, Service Anniversary & Milestones and Sales Incentive Programs Team: 4 Senior .NET dev, 1 Middle .NET dev, 5 Manual QA, 1 Senior Automation QA, 3 Senior DB dev, 1 Junior DB dev, 3 Markup/Front-End devs. Всего в Харьковском офисе 65 + человек. График: 5/2, 8 часовой рабочий день, гибкий график. Без поздних митингов и овертаймов.",
    "removed_phrases": [
      {
        "phrase": "young",
        "category": "age",
        "why": "signals age preference, a protected characteristic",
        "suggested_fix": "state the actual experience level needed"
      }
    ],
    "preserved_requirements": [
      "Job responsibilities:",
      "Experience using SQL Profiler and Extended Events to troubleshoot and optimize new and existing database code",
      "● MS SQL Server with T-SQL – minimum 2 plus years’ experience",
      "● Experience with SQL Server 2008 R2 through 2017.",
      "● Experience with SSRS, creating new queries / reports and setting up and maintaining services and shared service components.",
      "● SSIS experience",
      "● Familiarity with .Net Core, Visual Studio, and Powershell for SQL Server and Azure SQL"
    ],
    "residual_risks": []
  }
}

</details>

Composition and per-shard licensing

The composition below describes the publishable tagged training file recorded by artifacts/hr-provenance.json.

ShardRowsRoleSourceLicense
djinni_jd_bias4,500Span-grounded job-description bias audit pairslang-uk/recruitment-dataset-job-descriptions-englishMIT

Published training rows: 4,500. Compilation license: MIT.

Every published row retains shard, source, and license. The separate holdout is evaluation-only and is not part of the training composition.

Curation methodology

  1. 1.Real source text. The builder downloads the upstream job-description parquet, reads Long Description, and keeps descriptions between 40 and 400 whitespace-delimited words that contain a phrase recognized by the curated detector.
  2. 2.Deterministic, span-grounded labels. Seventeen case-insensitive regular-expression patterns detect coded language across age, affinity, culture, gender, gender/culture, and national-origin categories. Each match contributes the exact source span, category, explanation, and suggested fix. No model generates these labels.
  3. 3.Requirement preservation. The neutral version removes only matched phrases. Requirement-bearing lines are retained using a deterministic expression for experience, skills, tools, qualifications, and responsibilities. This directly targets the failure mode in which a model either leaves a dog whistle in place or removes job-relevant constraints with it.
  4. 4.Deduplication and split isolation. pipeline/dedup_split.py pools candidate train and holdout rows, removes exact duplicates and near duplicates using normalized prompts plus 64-bit token-trigram SimHash, then re-splits the retained rows. pipeline/provenance.py independently rejects any exact or near prompt overlap at a maximum Hamming distance of 3 before writing the tagged file. The provenance report checked 500 holdout rows and recorded 0 train/holdout leaks.
  5. 5.Adaptive Data. The tagged pairs were submitted to Adaptive Data by Adaption Labs. The recorded quality score moved from 6.0 to 9.8, grade A.

Licensing and attribution

The compilation is MIT because its only component is MIT-licensed. pipeline/provenance.py selects the most restrictive component license when a compilation contains more than one license. The license field remains on every row so downstream users can audit and preserve source-level terms.

Source job descriptions are attributed to `lang-uk/recruitment-dataset-job-descriptions-english`. The structured detector labels, neutralization targets, and provenance packaging are this project’s derived work. Adaptive Data by Adaption Labs was used for data adaptation.

Intended use

Use JD-BiasAudit for supervised fine-tuning, controlled raw-versus-adapted ablations, or evaluation of systems that audit English-language job descriptions and return strict JSON. It is suitable for research prototypes, recruiter-assistance tools with human review, compliance pre-screening, and tests of requirement-preserving rewriting.

Do not use it to infer protected attributes about applicants, make employment decisions, certify legal compliance, or replace review by qualified HR and legal professionals.

Limitations

  • —The detector is lexicon-bound. It can miss contextual, euphemistic, newly coined, or intersectional bias that does not match one of its patterns.
  • —Regex deletion can leave awkward prose and is not a substitute for an expert rewrite.
  • —Upstream postings can contain multilingual fragments even though the task instructions and labels are English.
  • —residual_risks is empty in deterministically generated labels, so the dataset does not teach a rich uncertainty taxonomy.
  • —Exact and SimHash leak checks reduce prompt overlap but do not prove the absence of all semantic overlap.

Honest weakness: the label precision is inspectable, but recall is capped by the fixed lexicon. A fluent model trained on these rows may sound comprehensive while still missing forms of bias the detector never encoded.

Reproducibility

Run pipeline/build_hr_dataset.py, then pipeline/dedup_split.py data/hr-jd-bias, then pipeline/provenance.py configs/provenance-hr.yaml. The builder, split logic, row-level provenance, license resolution, and leak report are all included in this repository.

Created for the Adaption AutoScientist Challenge using Adaptive Data by Adaption Labs.