CoolFace
Datasetpublic

opensporks/resumes

Dataset Card for Resume Dataset Dataset Summary Context A collection of Resume Examples taken from livecareer.com for categorizing a given resume into any of the labels defined in the dataset. Content Contains 2400+ Resumes in string as well as PDF format. PDF stored in the data folder differentiated into their respective labels as folders with each resume residing inside the folder in pdf form with filename as the id defined in the… See the full description on the dataset page: https://huggingface.co/datasets/opensporks/resumes.

sourceHugging Facecc0-1.0updated 2y agoView on Hugging Face
14likes9kdownloads
Dataset Card

Dataset Card for Resume Dataset

Table of Contents

Dataset Description

  • Homepage: https://kaggle.com/datasets/snehaanbhawal/resume-dataset
  • Repository:
  • Paper:
  • Leaderboard:
  • Point of Contact:

Dataset Summary

Context

A collection of Resume Examples taken from livecareer.com for categorizing a given resume into any of the labels defined in the dataset.

Content

Contains 2400+ Resumes in string as well as PDF format. PDF stored in the data folder differentiated into their respective labels as folders with each resume residing inside the folder in pdf form with filename as the id defined in the csv.

Inside the CSV:

  • ID: Unique identifier and file name for the respective pdf.
  • Resume_str : Contains the resume text only in string format.
  • Resume_html : Contains the resume data in html format as present while web scrapping.
  • Category : Category of the job the resume was used to apply.

Present categories are HR, Designer, Information-Technology, Teacher, Advocate, Business-Development, Healthcare, Fitness, Agriculture, BPO, Sales, Consultant, Digital-Media, Automobile, Chef, Finance, Apparel, Engineering, Accountant, Construction, Public-Relations, Banking, Arts, Aviation

Acknowledgements

Data was obtained by scrapping individual resume examples from www.livecareer.com website. Web Scrapping code present in my Github Repo.

Supported Tasks and Leaderboards

[More Information Needed]

Languages

[More Information Needed]

Dataset Structure

Data Instances

[More Information Needed]

Data Fields

[More Information Needed]

Data Splits

[More Information Needed]

Dataset Creation

Curation Rationale

[More Information Needed]

Source Data

Initial Data Collection and Normalization

[More Information Needed]

Who are the source language producers?

[More Information Needed]

Annotations

Annotation process

[More Information Needed]

Who are the annotators?

[More Information Needed]

Personal and Sensitive Information

[More Information Needed]

Considerations for Using the Data

Social Impact of Dataset

[More Information Needed]

Discussion of Biases

[More Information Needed]

Other Known Limitations

[More Information Needed]

Additional Information

Dataset Curators

This dataset was shared by @snehaanbhawal

Licensing Information

The license for this dataset is cc0-1.0

Citation Information

bibtex
[More Information Needed]

Contributions

[More Information Needed]