CoolFace
Datasetpublic

TechWolf/Synthetic-ESCO-skill-sentences

Synthetic job ads for all ESCO skills Dataset Summary This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0. Languages We use the English version of ESCO, and all generated sentences are in English. Dataset Structure The dataset consists of 138,260 (sentence, skill) pairs. Citation Information If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
17likes165downloads
Dataset Card

Synthetic job ads for all ESCO skills

Dataset Description

  • Homepage: coming soon
  • Repository: coming soon
  • Paper: https://arxiv.org/abs/2307.10778
  • Point of Contact: jensjoris@techwolf.ai

Dataset Summary

This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0.

Languages

We use the English version of ESCO, and all generated sentences are in English.

Dataset Structure

The dataset consists of 138,260 (sentence, skill) pairs.

Citation Information

If you use this dataset, please include the following reference:

@article{decorte2023extreme,
  title={Extreme multi-label skill extraction training using large language models},
  author={Decorte, Jens-Joris and Verlinden, Severine and Van Hautte, Jeroen and Deleu, Johannes and Develder, Chris and Demeester, Thomas},
  journal={arXiv preprint arXiv:2307.10778},
  year={2023}
}