sallani/matchcv-synthetic-dataset
MatchCV Synthetic Training Dataset 500 synthetic, anonymous CV/Job matching pairs — generated algorithmically, zero real personal data. Used to fine-tune sallani/MatchCV-Qwen2.5-0.5B. What's inside Each record is an instruction-following example: { "instruction": "Analyse le profil candidat...", "input": "=== CV ANONYMISÉ ===\n...\n=== OFFRE ===\n...", "output": "Score de matching : 73.5/100\nCompétences correspondantes : ...\nGaps techniques :… See the full description on the dataset page: https://huggingface.co/datasets/sallani/matchcv-synthetic-dataset.
MatchCV Synthetic Training Dataset
500 synthetic, anonymous CV/Job matching pairs — generated algorithmically, zero real personal data.
Used to fine-tune `sallani/MatchCV-Qwen2.5-0.5B`.
What's inside
Each record is an instruction-following example:
{
"instruction": "Analyse le profil candidat...",
"input": "=== CV ANONYMISÉ ===\n...\n=== OFFRE ===\n...",
"output": "Score de matching : 73.5/100\nCompétences correspondantes : ...\nGaps techniques : ...\nRecommandation : ..."
}CV fields (all synthetic, no real person)
Formation— education level (Bac+2 to Bac+8, Bootcamp)Expérience— years + seniority (junior / confirmé / senior / expert)Compétences— technical skills (Python, Docker, AWS, ISO 27001, LoRA, etc.)Mobilité— preferred work mode (Remote, Hybrid, On-site)Disponibilité— availability (Immediate, 1 month, 3 months)
Job description fields
Type— mission type (Architecture, MLOps, Audit, DevSecOps, etc.)Secteur— industry sector (Fintech, Cybersécurité, IA générative, etc.)Durée— contract durationMode— work modeExpérience requise— required years rangeCompétences requises / souhaitées— required / nice-to-have skills
Output fields
Score de matching— 0 to 100Compétences correspondantes— matched skillsGaps techniques— missing required skillsCompétences bonus— nice-to-have skills the candidate hasRecommandation— hiring recommendation
Dataset stats
Privacy
All data is algorithmically generated from a vocabulary of technical skills, sectors, and education levels. No names, emails, phone numbers, addresses or any identifier. Safe to use for model training without GDPR constraints.
License
Apache 2.0
