CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McAuley-Lab /Amazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset. This dataset mainly includes reviews (ratings, text) and item metadata (desc- riptions, category information, price, brand, and images). Compared to the pre- vious versions, the 2023 version features larger size, newer reviews (up to Sep 2023), richer and cleaner meta data, and finer-grained timestamps (from day to milli-second).10B<n<100B354 likes62k downloads2y agoHugging Face02cornell-movie-review-data /rotten_tomatoes Dataset Card for "rotten_tomatoes" Dataset Summary Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.texttext-classification10K<n<100K117 likes55k downloads3y agoHugging Face03Yelp /yelp_review_full Dataset Card for YelpReviewFull Dataset Summary The Yelp reviews dataset consists of reviews from Yelp. It is extracted from the Yelp Dataset Challenge 2015 data. Supported Tasks and Leaderboards text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment. Languages The reviews were mainly written in english. Dataset Structure Data Instances A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.texttext-classification100K<n<1M149 likes19k downloads3y agoHugging Face04NRVBench /nrvbench-review NR Video Editing Benchmark This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions. The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.imagevideo-to-videon<1K1 likes9.5k downloads5mo agoHugging Face05banned-historical-archives /peking-review0 likes9.1k downloads2y agoHugging Face06taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes6.1k downloads3y agoHugging Face07Daoze /ReviewRebuttal Introduction This dataset is the largest real-world consistency-ensured dataset for peer review, which features the widest range of conferences and the most complete review stages, including initial submissions, reviews, ratings and confidence, aspect ratings, rebuttals, discussions, score changes, meta-reviews, and final decisions. Paper: https://arxiv.org/abs/2505.07920 If our dataset can help you, please consider include the following citation in your publications:… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/ReviewRebuttal.text-generation100K<n<1M6 likes5.6k downloads4mo agoHugging Face08saattrupdan /womens-clothing-ecommerce-reviews Dataset Card for "womens-clothing-ecommerce-reviews" Processed version of this dataset. tabulartext-classification10K<n<100K2 likes3.6k downloads3y agoHugging Face09firm-review /FIRMgated FIRM: A Benchmark for Industrial Flexible-Object Robot Manipulation FIRM is a benchmark for industrial flexible-object robot manipulation grounded in real-world industrial data. The benchmark focuses on manipulation tasks involving mixed-stiffness objects, including instruction manuals, power cables, sponge pads, tapes, and cardboard components. These objects exhibit bending, slipping, rolling, compression, elastic recovery, and flexible-rigid contact under production-line… See the full description on the dataset page: https://huggingface.co/datasets/firm-review/FIRM.textroboticsn<1K3 likes3.4k downloads2mo agoHugging Face10mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3k downloads11d agoHugging Face11mteb /amazon_reviews_multiWe provide an Amazon product reviews dataset for multilingual text classification. The dataset contains reviews in English, Japanese, German, French, Chinese and Spanish, collected between November 1, 2015 and November 1, 2019. Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID and the coarse-grained product category (e.g. ‘books’, ‘appliances’, etc.) The corpus is balanced across stars, so each star rating constitutes 20% of the reviews in each language. For each language, there are 200,000, 5,000 and 5,000 reviews in the training, development and test sets respectively. The maximum number of reviews per reviewer is 20 and the maximum number of reviews per product is 20. All reviews are truncated after 2,000 characters, and all reviews are at least 20 characters long. Note that the language of a review does not necessarily match the language of its marketplace (e.g. reviews from amazon.de are primarily written in German, but could also be written in English, etc.). For this reason, we applied a language detection algorithm based on the work in Bojanowski et al. (2017) to determine the language of the review text and we removed reviews that were not written in the expected language.text1M<n<10M30 likes3k downloads4y agoHugging Face12argilla /tripadvisor-hotel-reviews Dataset Card for "tripadvisor-hotel-reviews" Dataset Summary Hotels play a crucial role in traveling and with the increased access to information new pathways of selecting the best ones emerged. With this dataset, consisting of 20k reviews crawled from Tripadvisor, you can explore what makes a great hotel and maybe even use this model in your travels! Citations on a scale from 1 to 5. Languages english Citation Information If you use this dataset in… See the full description on the dataset page: https://huggingface.co/datasets/argilla/tripadvisor-hotel-reviews.texttext-classification10K<n<100K7 likes2.6k downloads4y agoHugging Face13canrager /amazon_reviews_mcauley_1and5tabular100K<n<1M1 likes2.2k downloads2y agoHugging Face14SetFit /amazon_reviews_multi_entext100K<n<1M7 likes2.2k downloads4y agoHugging Face15Censius-AI /ECommerce-Women-Clothing-Reviewstabular10K<n<100K2 likes2.2k downloads3y agoHugging Face16defunct-datasets /amazon_reviews_multiWe provide an Amazon product reviews dataset for multilingual text classification. The dataset contains reviews in English, Japanese, German, French, Chinese and Spanish, collected between November 1, 2015 and November 1, 2019. Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID and the coarse-grained product category (e.g. ‘books’, ‘appliances’, etc.) The corpus is balanced across stars, so each star rating constitutes 20% of the reviews in each language. For each language, there are 200,000, 5,000 and 5,000 reviews in the training, development and test sets respectively. The maximum number of reviews per reviewer is 20 and the maximum number of reviews per product is 20. All reviews are truncated after 2,000 characters, and all reviews are at least 20 characters long. Note that the language of a review does not necessarily match the language of its marketplace (e.g. reviews from amazon.de are primarily written in German, but could also be written in English, etc.). For this reason, we applied a language detection algorithm based on the work in Bojanowski et al. (2017) to determine the language of the review text and we removed reviews that were not written in the expected language.summarization100K<n<1M102 likes2.1k downloads3y agoHugging Face17patrickbdevaney /tripadvisor_hotel_reviewstext10K<n<100K1 likes2k downloads2y agoHugging Face18datahiveai /Amazon-Reviews-DatasetThis dataset provides a free trial sample of best-selling products and their customer reviews from a leading e-commerce platform, designed to support product intelligence, sentiment analysis, and market trend evaluation. This sample is provided for evaluation purposes only. It includes a curated subset of the full dataset. To access the complete dataset, request additional attributes, or explore alternative product segments, please contact the data provider directly. Key Features 2… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/Amazon-Reviews-Dataset.image0 likes1.9k downloads1y agoHugging Face19VibrantVista /TTCW-Based-Review TTCW Creative Writing Evaluation Dataset If you use this dataset in your research, please cite our paper — it helps support ongoing academic work. Citation details are at the bottom of this page. Dataset Description Summary A supervised fine-tuning (SFT) dataset for training LLMs to act as creative writing evaluators. Each example contains a creative story and four message-format columns representing different evaluation objectives — from… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/TTCW-Based-Review.tabulartext-generation100K<n<1M2 likes1.8k downloads4mo agoHugging Face20sparsh3011 /Amazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset. This dataset mainly includes reviews (ratings, text) and item metadata (desc- riptions, category information, price, brand, and images). Compared to the pre- vious versions, the 2023 version features larger size, newer reviews (up to Sep 2023), richer and cleaner meta data, and finer-grained timestamps (from day to milli-second).10B<n<100B1 likes1.8k downloads5mo agoHugging Face21abayuu /Womens_Clothing_E-Commerce_Reviewstabular10K<n<100K0 likes1.7k downloads3y agoHugging Face22drivaerstar /DrivAerStar-Review DrivAerStar: An Industrial-Grade CFD Dataset for Vehicle Aerodynamic Optimization Vehicle aerodynamics optimization is fundamental to automotive engineering, drag reduction, noise minimization, and vehicle body stability through complex fluid dynamics simulations. Traditional approaches rely on computationally expensive Computational Fluid Dynamics (CFD) simulations that limit design exploration or simplified models that compromise accuracy. Machine learning methods offer promising… See the full description on the dataset page: https://huggingface.co/datasets/drivaerstar/DrivAerStar-Review.text1K<n<10K0 likes1.4k downloads1y agoHugging Face23BarbaDLuca /amazon-reviews-2023-with-asin Amazon Reviews 2023 (with ASIN) A trimmed version of the McAuley-Lab/Amazon-Reviews-2023 dataset, retaining only the fields most relevant for NLP tasks while adding explicit product identification via parent_asin. What's Different from the Original The original dataset includes 10+ fields per review and requires a legacy loading script that is no longer supported by HuggingFace. This version: Keeps only 4 fields: rating, title, text, and parent_asin Is stored in… See the full description on the dataset page: https://huggingface.co/datasets/BarbaDLuca/amazon-reviews-2023-with-asin.text100M<n<1B0 likes1.4k downloads4mo agoHugging Face24sealuzh /app_reviews Dataset Card for [Dataset Name] Dataset Summary It is a large dataset of Android applications belonging to 23 differentapps categories, which provides an overview of the types of feedback users report on the apps and documents the evolution of the related code metrics. The dataset contains about 395 applications of the F-Droid repository, including around 600 versions, 280,000 user reviews (extracted with specific text mining approaches) Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sealuzh/app_reviews.texttext-classification100K<n<1M28 likes1.4k downloads3y agoHugging Face25BanglishRev /bangla-english-and-code-mixed-ecommerce-review-dataset BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce Description The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.image0 likes1.3k downloads2y agoHugging Face26Ahmadhaiwala /Amazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset. This dataset mainly includes reviews (ratings, text) and item metadata (desc- riptions, category information, price, brand, and images). Compared to the pre- vious versions, the 2023 version features larger size, newer reviews (up to Sep 2023), richer and cleaner meta data, and finer-grained timestamps (from day to milli-second).10B<n<100B0 likes1.3k downloads3mo agoHugging Face27AgentAlphaAGI /Paper-Review-Dataset Dataset Card for Paper Review Dataset (ICLR 2023-2025) Dataset Description This dataset contains paper submissions and review data from the International Conference on Learning Representations (ICLR) for the years 2023, 2024, and 2025. The data is sourced from OpenReview, an open peer review platform that hosts the review process for top ML conferences. Focus on Review Data This dataset emphasizes the peer review ecosystem surrounding academic papers. Each… See the full description on the dataset page: https://huggingface.co/datasets/AgentAlphaAGI/Paper-Review-Dataset.text-classification10K<n<100K10 likes1.2k downloads7mo agoHugging Face28anonymous-review-dataset-2026 /review-dataset MSIR-Bench Review Dataset This repository contains an anonymized review snapshot of MSIR-Bench, a benchmark for identity-preserving style image retrieval. Dataset Description Each source identity is represented by an anonymous five-digit ID. Images are organized by split and identity folder. File names follow either <id>_<Style>.png, <id>_original.png, or legacy original.jpg for original reference images. The dataset is intended for evaluating whether a retrieval… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-review-dataset-2026/review-dataset.imageimage-feature-extraction10K<n<100K2 likes1.2k downloads2mo agoHugging Face29lewtun /drug-reviewstabular100K<n<1M13 likes1.2k downloads5y agoHugging Face30ashraq /hotel-reviews Dataset Card for "hotel-reviews" More Information needed Data was obtained from here text10K<n<100K3 likes1.2k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.