CoolFace
Datasetpublic

Acidmanic/DK-FA-Cosmetics

Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Mani Moayedi Language(s) (NLP): Farsi (Persian) License: MIT Uses The samples of this dataset are user comments about products of an online shop website. Each comment contains some additional data alongside the comments body, like… See the full description on the dataset page: https://huggingface.co/datasets/Acidmanic/DK-FA-Cosmetics.

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes72downloads
Dataset Card

Dataset Card for Dataset Name

<!-- Provide a quick summary of the dataset. -->

This dataset card aims to be a base template for new datasets. It has been generated using this raw template.

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

  • —Curated by: Mani Moayedi
  • —Language(s) (NLP): Farsi (Persian)
  • —License: MIT

Uses

The samples of this dataset are user comments about products of an online shop website. Each comment contains some additional data alongside the comments body, like star-rating value (0-5). This dataset can be used to train or generate different data-models for NLP tasks like opinion mining and sentiment analysis.

Out-of-Scope Use

<!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. -->

This dataset is the result of crawling 7 categories of cosmetic products from a perisan online-shop's product pages. The vocabulary mostly revolves around the cosmetics subjects, therefore it might not be suitable for use cases which needs a generic collection of words and phrases.

Dataset Structure

Each comment is represented in structured format and contains comment's body, comment's title, star-rating value (0-5), Other users reactions to each comment in terms of number-of-likes and number-of-dislikes. and a list of advantages and dis-advantages that user might have specified. title field and advantages/disadvantages fields can be null or empty in many comments.

For more details please check out the file Dataset Description.

Dataset Creation

The dataset is created using a crawler agains an online shop's website. Comments are scraped from product pages and stored as json, jsonl and csv files.

Personal and Sensitive Information

The dataset contains the username of userse who has posted the comments. All the information in the dataset, including these usernames, are present on the products web-page whithout any login or authentication.

Bias, Risks, and Limitations

From Npl prespective, the dataset might mostly contain information about the the consmetic products and the quality of sellers and resellers service. Therefore considering this dataset as a general source of language might introduce some issues, depending on the use-case.

Glossary

SetNumber Of CommentsNumber Of ProductsAverage Comments Per Product
dk-fa-cosmetics (Full dataset)421078832551
dkfacs-eyeliner (sub-set)30824284109
dkfacs-stand (sub-set)83197173848
dkfacs-mascara (sub-set)47961338142
dkfacs-sun-screen (sub-set)118699772154
dkfacs-eye-shadow (sub-set)1453263423
dkfacs-nails (sub-set)75209326023
dkfacs-lipsticks (sub-set)50656129939

Dataset Card Authors

Mani Moayedi

Dataset Card Contact

acidmanic.moayedi@gmail.com

https://github.com/Acidmanic