datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.AutoSUIT
AutoSUIT Bench (HuggingFace edition)
Dynamic, execution-based benchmark for secure code generation by LLMs. Every generated
program is compiled/interpreted and run against two independent unit-test suites — a
functional suite and a security suite (the latter is designed to fail when the target
CWE vulnerability is present). Covers 232 CWEs across C, C++, Java, and Python.
Paper: Osebe et al., AutoSUIT Bench — Automated Security UnIt Test Benchmark for LLM
Coding, Findings of… See the full description on the dataset page: https://huggingface.co/datasets/amazon/AutoSUIT.migration-bench-java-selected
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.AmazonQAC
AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset
Train Dataset Size: 395 million samplesTest Dataset Size: 20k samplesSource: Amazon Search LogsFile Format: ParquetCompression: Snappy
If you use this dataset, please cite our EMNLP 2024 paper:
@inproceedings{everaert-etal-2024-amazonqac,
title = "{A}mazon{QAC}: A Large-Scale, Naturalistic Query Autocomplete Dataset",
author = "Everaert, Dante and
Patki, Rohit and
Zheng, Tianqi and
Potts… See the full description on the dataset page: https://huggingface.co/datasets/amazon/AmazonQAC.amazon-esci-data
Amazon Shopping Queries Dataset
Dataset for improving product search, ranking and recommendations, featuring query-product pairs with detailed relevance labels.
Overview
The dataset contains search queries paired with up to 40 potentially relevant products, each labeled using the ESCI system:
Exact match: Products that perfectly match the customer's search intent (e.g., searching "iPhone 13" and finding "Apple iPhone 13 128GB")
Substitute product: Alternative products… See the full description on the dataset page: https://huggingface.co/datasets/milistu/amazon-esci-data.amazon-c11-distillation
Amazon C11 adaptive-oracle distillation
Immutable backing data for the Amazon C11 Distillation Viewer.
Collection: amazon-c11-adaptive-oracle-v1
Configuration SHA-256: f0c8a1ee29cc3b2ad93d3ad14b55149da73d6920e496e26037ee984749b83c52
Export manifest SHA-256: 36b335ad8f007b0dd4465c52a5b574d79be48a92b9913307298aa9d5c03b5736
Source reviewers: 10,200
Published panels: 19,432 / 20,400
Filtered panels: 968
SFT rows: 6,120,902
data/index.json contains the global index and… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation.amazon-c2-varied-rubrics
Amazon C2 varied-rubric distillation
This release exposes six balanced C2 SFT configurations: latent-state and non-diverse candidate panels at
K=1, K=2, and K=4 rubrics per retained reviewer. Each rubric-writer target is paired with one full-rubric
listwise judge target over the same variant's frozen 40-candidate panel. The K arms within a variant share one
reviewer cohort and are exact nested prefixes.
Config
Train rows
Validation
Test
Train reviewers… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-varied-rubrics.amazon-esci-data
Amazon Shopping Queries Dataset
Dataset for improving product search, ranking and recommendations, featuring query-product pairs with detailed relevance labels.
Overview
The dataset contains search queries paired with up to 40 potentially relevant products, each labeled using the ESCI system:
Exact match: Products that perfectly match the customer's search intent (e.g., searching "iPhone 13" and finding "Apple iPhone 13 128GB")
Substitute product: Alternative products… See the full description on the dataset page: https://huggingface.co/datasets/thepian/amazon-esci-data.amazon-c11-distillation-filtered
Amazon C11 quality-filtered distillation
This is a high-signal SFT view of asingh15/amazon-c11-distillation.
It contains six balanced rubric-writer/criterion-judge configurations. The original source remains unchanged.
Filter
A trajectory is retained only when its selected rubric has gold-score spread greater than 0.10 on both the
selection panel and the paired held-out panel, non-constant proxy scores on both, positive Spearman correlation
on both, positive… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation-filtered.amazon-c11-nothink-distillation-filtered
Amazon c11 no-think quality-filtered distillation
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/<split>/; it is not exposed as another dataset configuration.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation-filtered.amazon-c2-distillation
Amazon c2 distillation
Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the complete pinned C11 stopping-objective corpus.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes only its combined training view and preserves the native train, validation, and… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-distillation.amazon-esci-english-smallamazon-c2-distillation-filtered
Amazon c2 quality-filtered distillation
Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/<split>/; it is not exposed as another dataset configuration.
The six configurations cross the two candidate variants with the three frozen stopping… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-distillation-filtered.amazon-c11-nothink-distillation
Amazon c11 no-think distillation
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the complete pinned C11 stopping-objective corpus.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes only its combined training view and preserves the native train, validation, and test
splits.
Config
Train… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation.amazon-products
Amazon Products Sample Dataset
A curated sample of 2,000 popular products from the Amazon Reviews 2023 dataset, designed for educational use in building RAG (Retrieval-Augmented Generation) systems and shopping agents.
Dataset Description
This dataset contains product metadata across 4 categories:
Electronics (500 products)
Video Games (500 products)
Books (500 products)
Home & Kitchen (500 products)
Products were filtered to include only those with 500+ reviews… See the full description on the dataset page: https://huggingface.co/datasets/gatech-scheller-ai-in-business/amazon-products.edit_amazon_reviews_multi_es
Dataset Summary
The data file is intended for a tutorial: Summarization
Language
Spanish
Dataset Structure
id: record id
stars: An int between 1-5 indicating the number of stars.
review_body: The text body of the review.
review_title: The text title of the review.
language: The string identifier of the review language.
product_category: String representation of the product's category.
lenght_review_body: text length of review_body
lenght_review_title: text… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/edit_amazon_reviews_multi_es.Amazon-combined
Amazon Combined Dataset
E-commerce dataset that combines metadata, reviews, and sample question/answer pairs. combined.json contains the dataset and user2asin.json contains a file that maps user_id from reviews to an ASIN for capturing user preferences.
Data Fields
Field
Type
Explanation
main_category
str
Main category (i.e., domain) of the product.
title
str
Name of the product.
average_rating
float
Rating of the product shown on the product page.… See the full description on the dataset page: https://huggingface.co/datasets/randomath/Amazon-combined.amazon-reviews-2023-all-beauty-sample
Amazon Reviews 2023 – All_Beauty (Sampled)
This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023
All_Beauty category, prepared for the YZM2022 Data Mining homework
(Assoc. Prof. Dr. Arzu Kakisim).
Sampling strategy
Source: full All_Beauty reviews (701K) and metadata (112K items).
3-core filtering (each user and item has at least 3 interactions, iterated to convergence).
Cap to the most recent 60 000 interactions, re-applied 3-core.
Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.edit_amazon_reviews_multi_en
Dataset Summary
The data file is intended for a tutorial: Summarization
Language
English
Dataset Structure
id: record id
stars: An int between 1-5 indicating the number of stars.
review_body: The text body of the review.
review_title: The text title of the review.
language: The string identifier of the review language.
product_category: String representation of the product's category.
lenght_review_body: text length of review_body
lenght_review_title: text… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/edit_amazon_reviews_multi_en.AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
Adversarial Arena: Trusted AI Challenge Dataset
Dataset Description
This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs.
Papers:
Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.AmazonQAC
AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset
Train Dataset Size: 395 million samplesTest Dataset Size: 20k samplesSource: Amazon Search LogsFile Format: ParquetCompression: Snappy
If you use this dataset, please cite our EMNLP 2024 paper:
@inproceedings{everaert-etal-2024-amazonqac,
title = "{A}mazon{QAC}: A Large-Scale, Naturalistic Query Autocomplete Dataset",
author = "Everaert, Dante and
Patki, Rohit and
Zheng, Tianqi and… See the full description on the dataset page: https://huggingface.co/datasets/rishabhgrgb/AmazonQAC.AmazonQAC
AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset
Train Dataset Size: 395 million samplesTest Dataset Size: 20k samplesSource: Amazon Search LogsFile Format: ParquetCompression: Snappy
If you use this dataset, please cite our EMNLP 2024 paper:
@inproceedings{everaert-etal-2024-amazonqac,
title = "{A}mazon{QAC}: A Large-Scale, Naturalistic Query Autocomplete Dataset",
author = "Everaert, Dante and
Patki, Rohit and
Zheng, Tianqi and… See the full description on the dataset page: https://huggingface.co/datasets/Shlok-1/AmazonQAC.
