datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shopping-companion
Shopping Companion
Shopping Companion is a benchmark and training resource for long-horizon,
preference-grounded e-commerce agents. It evaluates whether a tool-using agent
can recover a user's preferences from cross-session conversation history and
apply those preferences while searching and inspecting a large real-world
product catalog.
The benchmark contains two task types:
Single-product recommendation: retrieve the relevant long-term preference
and find one product that… See the full description on the dataset page: https://huggingface.co/datasets/yuzhan2205/Shopping-companion.jam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.6.0 — a correction release. It withdraws 58 records whose source arrangements could not be licence-cleared and changes no remaining record. See Version 0.6.0 correction.
Records built: 2026-07-11 (0.5.0 cut; unchanged) Package built: 2026-09-25
DOI: 10.5281/zenodo.22961580 (this version; concept DOI 10.5281/zenodo.22961579). Earlier versions: 0.5.0 10.5281/zenodo.21313954 and 0.4.3 10.5281/zenodo.20279919. Both contain… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.1.0
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
72 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.sn15-shoppingbench-sft-15k
ShoppingBench SN15 SFT Corpus (15K filtered + 2.7K eval holdout)
Paper: arXiv:2606.10064Code: https://github.com/ORO-AI/shoppingbench-trajectory-primitive
The filtered, leak-cluster-guarded SFT corpus from the paper Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces. This is the trainable corpus produced by the structural-quality filter from the raw 18K race traces (oro-ai/sn15-shoppingbench-traces-18k).
Post-training… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/sn15-shoppingbench-sft-15k.sn15-shoppingbench-traces-18k
ShoppingBench SN15 Race Traces (18K, winners, unanimous-50)
Paper: arXiv:2606.10064Code: https://github.com/ORO-AI/shoppingbench-trajectory-primitive
Multi-turn agentic shopping trajectories harvested from ORO Subnet 15 (SN15), the Bittensor deployment of the ShoppingBench agentic-commerce benchmark. These are the unfiltered 18K traces referenced in the paper Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces.
This… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/sn15-shoppingbench-traces-18k.ShoppingReasoningBench
Shopping Reasoning Bench
Conversational shopping assistants now serve hundreds of millions of customers, yet no existing benchmark jointly evaluates the open-ended multi-turn reasoning, domain expertise, and criterion-level quality that real shopping conversations demand. We introduce the Shopping Reasoning Bench, an expert-authored benchmark of 525 missions (232 single-turn, 293 multi-turn) with 10,863 importance-weighted binary rubrics authored by retail domain experts. These… See the full description on the dataset page: https://huggingface.co/datasets/amazon/ShoppingReasoningBench.korean_wiki_content_only_120125Cleaned Korean Wiki text for 120125.
ShopTrajQA
ShopTrajQA
Shopping-agent trajectory question-answering data. The dataset covers a full training +
evaluation stack for building agents that reason over user shopping browsing trajectories
and answer questions about them (prices, actions, products, etc.).
Contents
Split
Path
Rows
Description
SFT (train)
sft/public_shopping_multi-turn_clean.parquet
762
Multi-turn supervised fine-tuning conversations. Each row has messages (chat turns with a Python… See the full description on the dataset page: https://huggingface.co/datasets/hongyeeliu/ShopTrajQA.kowiki-cleaned-020126
shopkeeper/kowiki-latest-clean
Korean Wikipedia dataset generated from Wikimedia dump.
Source: https://dumps.wikimedia.org/kowiki/20260201/
Dump files: 1 shard(s)
Fields: id, url, title, text
Cleaning: wiki markup and HTML/reference removal with whitespace normalization.
retail-shop-enquiries
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/pythontech9/retail-shop-enquiries.ambari-shopping-instruct-v2
Ambari Shopping Instruct v2
Refined dataset for fine-tuning Ambari-7B for a Kannada shopping assistant.
Removed Chain-of-Thought (CoT) traces that were causing issues in v1.
Structure
instruction: Task description
input: User query (Kanglish/Kannada) or Context
output: Expected assistant response (Kanglish/Kannada)
Splits
Train: 563 samples
Test: 63 samples
EVA-shop
