CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.4k downloads1y agoHugging Face02amazon /AutoSUIT AutoSUIT Bench (HuggingFace edition) Dynamic, execution-based benchmark for secure code generation by LLMs. Every generated program is compiled/interpreted and run against two independent unit-test suites — a functional suite and a security suite (the latter is designed to fail when the target CWE vulnerability is present). Covers 232 CWEs across C, C++, Java, and Python. Paper: Osebe et al., AutoSUIT Bench — Automated Security UnIt Test Benchmark for LLM Coding, Findings of… See the full description on the dataset page: https://huggingface.co/datasets/amazon/AutoSUIT.tabulartext-generationn<1K0 likes1.4k downloads3mo agoHugging Face03AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes860 downloads1y agoHugging Face04amazon /AmazonQAC AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset Train Dataset Size: 395 million samplesTest Dataset Size: 20k samplesSource: Amazon Search LogsFile Format: ParquetCompression: Snappy If you use this dataset, please cite our EMNLP 2024 paper: @inproceedings{everaert-etal-2024-amazonqac, title = "{A}mazon{QAC}: A Large-Scale, Naturalistic Query Autocomplete Dataset", author = "Everaert, Dante and Patki, Rohit and Zheng, Tianqi and Potts… See the full description on the dataset page: https://huggingface.co/datasets/amazon/AmazonQAC.tabulartext-generation100M<n<1B18 likes576 downloads2y agoHugging Face05milistu /amazon-esci-data Amazon Shopping Queries Dataset Dataset for improving product search, ranking and recommendations, featuring query-product pairs with detailed relevance labels. Overview The dataset contains search queries paired with up to 40 potentially relevant products, each labeled using the ESCI system: Exact match: Products that perfectly match the customer's search intent (e.g., searching "iPhone 13" and finding "Apple iPhone 13 128GB") Substitute product: Alternative products… See the full description on the dataset page: https://huggingface.co/datasets/milistu/amazon-esci-data.tabulartext-classification1M<n<10M2 likes270 downloads1y agoHugging Face06asingh15 /amazon-c11-distillation Amazon C11 adaptive-oracle distillation Immutable backing data for the Amazon C11 Distillation Viewer. Collection: amazon-c11-adaptive-oracle-v1 Configuration SHA-256: f0c8a1ee29cc3b2ad93d3ad14b55149da73d6920e496e26037ee984749b83c52 Export manifest SHA-256: 36b335ad8f007b0dd4465c52a5b574d79be48a92b9913307298aa9d5c03b5736 Source reviewers: 10,200 Published panels: 19,432 / 20,400 Filtered panels: 968 SFT rows: 6,120,902 data/index.json contains the global index and… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation.tabulartext-generation10M<n<100M0 likes179 downloads6d agoHugging Face07asingh15 /amazon-c2-varied-rubrics Amazon C2 varied-rubric distillation This release exposes six balanced C2 SFT configurations: latent-state and non-diverse candidate panels at K=1, K=2, and K=4 rubrics per retained reviewer. Each rubric-writer target is paired with one full-rubric listwise judge target over the same variant's frozen 40-candidate panel. The K arms within a variant share one reviewer cohort and are exact nested prefixes. Config Train rows Validation Test Train reviewers… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-varied-rubrics.tabulartext-generation100K<n<1M0 likes141 downloads4d agoHugging Face08thepian /amazon-esci-data Amazon Shopping Queries Dataset Dataset for improving product search, ranking and recommendations, featuring query-product pairs with detailed relevance labels. Overview The dataset contains search queries paired with up to 40 potentially relevant products, each labeled using the ESCI system: Exact match: Products that perfectly match the customer's search intent (e.g., searching "iPhone 13" and finding "Apple iPhone 13 128GB") Substitute product: Alternative products… See the full description on the dataset page: https://huggingface.co/datasets/thepian/amazon-esci-data.tabulartext-classification1M<n<10M0 likes107 downloads5mo agoHugging Face09asingh15 /amazon-c11-distillation-filtered Amazon C11 quality-filtered distillation This is a high-signal SFT view of asingh15/amazon-c11-distillation. It contains six balanced rubric-writer/criterion-judge configurations. The original source remains unchanged. Filter A trajectory is retained only when its selected rubric has gold-score spread greater than 0.10 on both the selection panel and the paired held-out panel, non-constant proxy scores on both, positive Spearman correlation on both, positive… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation-filtered.tabulartext-generation100K<n<1M0 likes96 downloads6d agoHugging Face10asingh15 /amazon-c11-nothink-distillation-filtered Amazon c11 no-think quality-filtered distillation This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/<split>/; it is not exposed as another dataset configuration. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation-filtered.tabulartext-generation100K<n<1M0 likes92 downloads5d agoHugging Face11asingh15 /amazon-c2-distillation Amazon c2 distillation Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the complete pinned C11 stopping-objective corpus. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-distillation.tabulartext-generation100K<n<1M0 likes89 downloads5d agoHugging Face12thebajajra /amazon-esci-english-smalltabulartoken-classification100K<n<1M1 likes81 downloads1y agoHugging Face13asingh15 /amazon-c2-distillation-filtered Amazon c2 quality-filtered distillation Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/<split>/; it is not exposed as another dataset configuration. The six configurations cross the two candidate variants with the three frozen stopping… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-distillation-filtered.tabulartext-generation10K<n<100K0 likes79 downloads5d agoHugging Face14asingh15 /amazon-c11-nothink-distillation Amazon c11 no-think distillation This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the complete pinned C11 stopping-objective corpus. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. Config Train… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation.tabulartext-generation100K<n<1M0 likes79 downloads5d agoHugging Face15gatech-scheller-ai-in-business /amazon-products Amazon Products Sample Dataset A curated sample of 2,000 popular products from the Amazon Reviews 2023 dataset, designed for educational use in building RAG (Retrieval-Augmented Generation) systems and shopping agents. Dataset Description This dataset contains product metadata across 4 categories: Electronics (500 products) Video Games (500 products) Books (500 products) Home & Kitchen (500 products) Products were filtered to include only those with 500+ reviews… See the full description on the dataset page: https://huggingface.co/datasets/gatech-scheller-ai-in-business/amazon-products.tabulartext-generation1K<n<10K0 likes74 downloads7mo agoHugging Face16KRadim /edit_amazon_reviews_multi_es Dataset Summary The data file is intended for a tutorial: Summarization Language Spanish Dataset Structure id: record id stars: An int between 1-5 indicating the number of stars. review_body: The text body of the review. review_title: The text title of the review. language: The string identifier of the review language. product_category: String representation of the product's category. lenght_review_body: text length of review_body lenght_review_title: text… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/edit_amazon_reviews_multi_es.tabulartext-generation100K<n<1M0 likes53 downloads1y agoHugging Face17randomath /Amazon-combined Amazon Combined Dataset E-commerce dataset that combines metadata, reviews, and sample question/answer pairs. combined.json contains the dataset and user2asin.json contains a file that maps user_id from reviews to an ASIN for capturing user preferences. Data Fields Field Type Explanation main_category str Main category (i.e., domain) of the product. title str Name of the product. average_rating float Rating of the product shown on the product page.… See the full description on the dataset page: https://huggingface.co/datasets/randomath/Amazon-combined.tabulartext-generation1K<n<10K1 likes50 downloads2y agoHugging Face18debolut /amazon-reviews-2023-all-beauty-sample Amazon Reviews 2023 – All_Beauty (Sampled) This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023 All_Beauty category, prepared for the YZM2022 Data Mining homework (Assoc. Prof. Dr. Arzu Kakisim). Sampling strategy Source: full All_Beauty reviews (701K) and metadata (112K items). 3-core filtering (each user and item has at least 3 interactions, iterated to convergence). Cap to the most recent 60 000 interactions, re-applied 3-core. Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.tabulartext-classification10K<n<100K0 likes47 downloads4mo agoHugging Face19KRadim /edit_amazon_reviews_multi_en Dataset Summary The data file is intended for a tutorial: Summarization Language English Dataset Structure id: record id stars: An int between 1-5 indicating the number of stars. review_body: The text body of the review. review_title: The text title of the review. language: The string identifier of the review language. product_category: String representation of the product's category. lenght_review_body: text length of review_body lenght_review_title: text… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/edit_amazon_reviews_multi_en.tabulartext-generation100K<n<1M1 likes26 downloads1y agoHugging Face20amazon-agi /AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset Adversarial Arena: Trusted AI Challenge Dataset Dataset Description This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs. Papers: Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.tabulartext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face21rishabhgrgb /AmazonQAC AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset Train Dataset Size: 395 million samplesTest Dataset Size: 20k samplesSource: Amazon Search LogsFile Format: ParquetCompression: Snappy If you use this dataset, please cite our EMNLP 2024 paper: @inproceedings{everaert-etal-2024-amazonqac, title = "{A}mazon{QAC}: A Large-Scale, Naturalistic Query Autocomplete Dataset", author = "Everaert, Dante and Patki, Rohit and Zheng, Tianqi and… See the full description on the dataset page: https://huggingface.co/datasets/rishabhgrgb/AmazonQAC.tabulartext-generation100M<n<1B0 likes17 downloads3mo agoHugging Face22Shlok-1 /AmazonQAC AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset Train Dataset Size: 395 million samplesTest Dataset Size: 20k samplesSource: Amazon Search LogsFile Format: ParquetCompression: Snappy If you use this dataset, please cite our EMNLP 2024 paper: @inproceedings{everaert-etal-2024-amazonqac, title = "{A}mazon{QAC}: A Large-Scale, Naturalistic Query Autocomplete Dataset", author = "Everaert, Dante and Patki, Rohit and Zheng, Tianqi and… See the full description on the dataset page: https://huggingface.co/datasets/Shlok-1/AmazonQAC.tabulartext-generation100M<n<1B0 likes16 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.