datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shopping-companion
Shopping Companion
Shopping Companion is a benchmark and training resource for long-horizon,
preference-grounded e-commerce agents. It evaluates whether a tool-using agent
can recover a user's preferences from cross-session conversation history and
apply those preferences while searching and inspecting a large real-world
product catalog.
The benchmark contains two task types:
Single-product recommendation: retrieve the relevant long-term preference
and find one product that… See the full description on the dataset page: https://huggingface.co/datasets/yuzhan2205/Shopping-companion.sn15-shoppingbench-sft-15k
ShoppingBench SN15 SFT Corpus (15K filtered + 2.7K eval holdout)
Paper: arXiv:2606.10064Code: https://github.com/ORO-AI/shoppingbench-trajectory-primitive
The filtered, leak-cluster-guarded SFT corpus from the paper Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces. This is the trainable corpus produced by the structural-quality filter from the raw 18K race traces (oro-ai/sn15-shoppingbench-traces-18k).
Post-training… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/sn15-shoppingbench-sft-15k.sn15-shoppingbench-traces-18k
ShoppingBench SN15 Race Traces (18K, winners, unanimous-50)
Paper: arXiv:2606.10064Code: https://github.com/ORO-AI/shoppingbench-trajectory-primitive
Multi-turn agentic shopping trajectories harvested from ORO Subnet 15 (SN15), the Bittensor deployment of the ShoppingBench agentic-commerce benchmark. These are the unfiltered 18K traces referenced in the paper Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces.
This… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/sn15-shoppingbench-traces-18k.ShoppingReasoningBench
Shopping Reasoning Bench
Conversational shopping assistants now serve hundreds of millions of customers, yet no existing benchmark jointly evaluates the open-ended multi-turn reasoning, domain expertise, and criterion-level quality that real shopping conversations demand. We introduce the Shopping Reasoning Bench, an expert-authored benchmark of 525 missions (232 single-turn, 293 multi-turn) with 10,863 importance-weighted binary rubrics authored by retail domain experts. These… See the full description on the dataset page: https://huggingface.co/datasets/amazon/ShoppingReasoningBench.ambari-shopping-instruct-v2
Ambari Shopping Instruct v2
Refined dataset for fine-tuning Ambari-7B for a Kannada shopping assistant.
Removed Chain-of-Thought (CoT) traces that were causing issues in v1.
Structure
instruction: Task description
input: User query (Kanglish/Kannada) or Context
output: Expected assistant response (Kanglish/Kannada)
Splits
Train: 563 samples
Test: 63 samples
