datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
job-posting-classificationOriginal job descriptions were derived from xanderios/linkedin-job-postings, and classification data was created synthetically with GPT-4o-Mini. All values not represented nor found in the job description are marked as null.
Note: Some responses are hallucinations, despite maintaining the correct .json format, the content is wrong. All instances of incorrect .json formatting have been removed from the dataset, hallucinated content however still remains.
Future: Might consider sourcing a resume… See the full description on the dataset page: https://huggingface.co/datasets/will4381/job-posting-classification.job-title-classification-dataset
SIRTAR Job classification Dataset
This is a fusion of three known Kaggle datasets, added tons of preprocessing in the middle. These are the following:
LinkedIn Job Postings (2023-2024) [1]: https://www.kaggle.com/datasets/arshkon/linkedin-job-postings
Indeed Job Postings: https://www.kaggle.com/datasets/spandanakalakonda/job-postings
Jobstreet Job Postings: https://www.kaggle.com/datasets/azraimohamad/jobstreet-all-job-dataset
This dataset is used for the training of the new… See the full description on the dataset page: https://huggingface.co/datasets/daniel-jurado/job-title-classification-dataset.job-classification-datasetjob_classification_dataset_ruThis dataset represents the classification of a profession by its name and description.
Real resumes from hh.ru with additional anonymization were used to form the dataset.
Claude 3 sonet was used for clasification. The total cost of training the dataset was about 500$
Данный датасет представляет с собой класификацию профессии по её названию и описанию.
Для формирования датасета использовались реальные резюме из hh.ru с дополнительной анонимизацией.
Для класификации был использован claude 3… See the full description on the dataset page: https://huggingface.co/datasets/daswer123/job_classification_dataset_ru.job-classification-llama2
🧠 Job Classification Dataset for LLaMA 2 Fine-Tuning
This dataset contains 1,000 synthetic job descriptions and their associated job categories. It is designed for fine-tuning large language models (LLMs), such as LLaMA 2, for job classification tasks.
📂 Dataset Structure
Format: JSONL (.jsonl)
Fields:
instruction: A generic instruction prompt.
input: The job description text.
output: The job type label.
🔧 Example
{
"instruction": "Classify the… See the full description on the dataset page: https://huggingface.co/datasets/saiteja001r/job-classification-llama2.job_classification_dataset_v2_ruThis dataset represents the classification of a profession by its name and description. Real resumes from hh.ru with additional anonymization were used to form the dataset. Claude 3 sonet was used for clasification. The total cost of training the dataset was about 500$
In this version, formation was done by chunks , and each new chunk was added to the RAG database, which the LLM received later, for more accurate classification.
Данный датасет представляет с собой класификацию профессии по её… See the full description on the dataset page: https://huggingface.co/datasets/daswer123/job_classification_dataset_v2_ru.job-classification-datasetcensus-job-title-classificationsjob_classification_dataset_v2_ruThis dataset represents the classification of a profession by its name and description. Real resumes from hh.ru with additional anonymization were used to form the dataset. Claude 3 sonet was used for clasification. The total cost of training the dataset was about 500$
In this version, formation was done by chunks , and each new chunk was added to the RAG database, which the LLM received later, for more accurate classification.
Данный датасет представляет с собой класификацию профессии по её… See the full description on the dataset page: https://huggingface.co/datasets/art403/job_classification_dataset_v2_ru.Census_Job_Title_Classification_Final_dataset
