datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
skill-extraction-tech
Skill Extraction with ESCO skills - TECH subset
Dataset Summary
This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-tech.Skill-extraction-Tech-graded
skill-extraction-tech-graded
Graded-relevance annotations for sentences from
TechWolf/skill-extraction-tech
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
75
_id (sentence id), text (sentence)
queries
test
338
_id (sentence id), text (sentence)
corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-Tech-graded.Skill-extraction-SkillSkape-graded
skill-extraction-skillskape-graded
Graded-relevance annotations for sentences from
jjzha/skillskape
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
100
_id (sentence id), text (sentence)
queries
test
500
_id (sentence id), text (sentence)
corpus
corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-SkillSkape-graded.Skill-extraction-House-graded
skill-extraction-house-graded
Graded-relevance annotations for sentences from
TechWolf/skill-extraction-house
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
61
_id (sentence id), text (sentence)
queries
test
261
_id (sentence id), text (sentence)… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-House-graded.skill-extraction-house
Skill Extraction with ESCO skills - HOUSE subset
Dataset Summary
This dataset contains an extension of the HOUSE subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-house.Skill-extraction-TechWolf-graded
skill-extraction-techwolf-graded
Graded-relevance annotations for sentences from
TechWolf/skill-extraction-techwolf
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
test
324
_id (sentence id), text (sentence)
corpus
corpus
13,891
_id (ESCO skill URI), title (English… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-TechWolf-graded.skill-extraction-techwolf
Skill Extraction with ESCO skills - TechWolf subset
Dataset Summary
The TECHWOLF subset, although smaller, represents a more generic distribution of job descriptions and skill spans. ESCO skills are directly annotated on the full sentence level, thus omitting the intermediate span identification step. ESCO v1.1.0 is used.
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-techwolf.ssf-skill-extraction-pairs
SSF Skill Extraction Pairs
A contrastive training dataset for fine-tuning embedding models to match job description sentences to standardized skills from Singapore's SkillsFuture Framework (SSF).
Dataset Summary
Property
Value
Total Pairs
21,958
Unique Skills
2,196
Sentences per Skill
5 (synthetic, JD-style)
Pair Types
Positive (correct skill) + Negative (random incorrect skill)
Language
English
Domain
Workforce Skills / HR / Job Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/imocha-ai-org/ssf-skill-extraction-pairs.skill-extraction-tech
Skill Extraction with ESCO skills - TECH subset
Dataset Summary
This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/Balki16/skill-extraction-tech.skill-extraction-finetune-dataset
