datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edit_amazon_reviews_multi_es
Dataset Summary
The data file is intended for a tutorial: Summarization
Language
Spanish
Dataset Structure
id: record id
stars: An int between 1-5 indicating the number of stars.
review_body: The text body of the review.
review_title: The text title of the review.
language: The string identifier of the review language.
product_category: String representation of the product's category.
lenght_review_body: text length of review_body
lenght_review_title: text… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/edit_amazon_reviews_multi_es.amazon-reviews-2023-all-beauty-sample
Amazon Reviews 2023 – All_Beauty (Sampled)
This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023
All_Beauty category, prepared for the YZM2022 Data Mining homework
(Assoc. Prof. Dr. Arzu Kakisim).
Sampling strategy
Source: full All_Beauty reviews (701K) and metadata (112K items).
3-core filtering (each user and item has at least 3 interactions, iterated to convergence).
Cap to the most recent 60 000 interactions, re-applied 3-core.
Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.edit_amazon_reviews_multi_en
Dataset Summary
The data file is intended for a tutorial: Summarization
Language
English
Dataset Structure
id: record id
stars: An int between 1-5 indicating the number of stars.
review_body: The text body of the review.
review_title: The text title of the review.
language: The string identifier of the review language.
product_category: String representation of the product's category.
lenght_review_body: text length of review_body
lenght_review_title: text… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/edit_amazon_reviews_multi_en.
