CoolFace
Datasetpublic

iaouali/amazon-benchmark

Amazon query–bundle benchmark Canonical, category-organized query and reference-positive data. Experiment traces should reference this repository by commit SHA, category, split, and candidate_id, rather than republishing the dataset. Musical Instruments Split Examples agent_dev 2,028 agent_hidden 1,960 Each record contains a query and 3–7 reference product IDs. These are observed reference positives, not exhaustive labels for all valid… See the full description on the dataset page: https://huggingface.co/datasets/iaouali/amazon-benchmark.

sourceHugging Faceupdated 24m agoView on Hugging Face
0likes1.7kdownloads
Dataset Card

Amazon query–bundle benchmark

Canonical, category-organized query and reference-positive data. Experiment traces should reference this repository by commit SHA, category, split, and candidate_id, rather than republishing the dataset.

Musical Instruments

SplitExamples
agent_dev2,028
agent_hidden1,960

Each record contains a query and 3–7 reference product IDs. These are observed reference positives, not exhaustive labels for all valid recommendations. No user identifiers, raw interaction histories, or protected split assignments are included. There are no shared_train bundle-query examples in this release.

Construction and limitations

Users were partitioned before bundle generation. Candidate bundles come from positive review interactions in seven-day windows, not verified shopping baskets. A fixed LLM coherence judge uses three votes with at least two passes; generated queries then undergo consistency and deterministic checks. The hidden generation did not use agent-performance filtering. Development examples include prompt-optimization exposures, recorded per row.

The outer split is user-disjoint, not product-disjoint: six hidden reference sets also occur in development. The historical agent_hidden name is retained for reproducibility, but publication means this split is no longer secret. Keep it excluded from training and tuning for held-out comparisons; a future confidential test requires a new split.

Existing evaluations request k equal to reference-set size and measure SetHit@k as any exact product overlap. They are not SetHit@20 evaluations. Product relevance and compatibility are not exhaustively verified.

Related artifacts

  • Tools: https://huggingface.co/iaouali/amazon-tools
  • Traces: https://huggingface.co/datasets/iaouali/amazon-traces

See the category manifest for file checksums. Upstream Amazon data and model usage terms still apply; this release does not assert a new license over third-party source material.