datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
retail-forecast-optimize-benchmark
Retail Forecast-Optimize Benchmark
Benchmark dataset for predict-then-optimize retail inventory decisions.
Contents
File pattern
Description
{scenario}_s{seed}.json
Full policy comparison per scenario instance
samples/sample_*.json
Retail scenario definitions
eval_results.json
Aggregated leaderboard metrics
chronos-2-live-forecasts/
Live Chronos-2 outputs from HF Jobs (when available)
Scenarios
Baseline Operations
Promotion… See the full description on the dataset page: https://huggingface.co/datasets/aniketraj0224/retail-forecast-optimize-benchmark.retail
retail
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (205 of 205)
Fidelity over Runs
100.0% (456 of 456)
Reference confirmed
205
Verifier derived
205
Trusted
133
Refused
0
Not trusted
72 of 205; 52 the Verifier passed the suite and… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.retail-bank-servicing-alignment-sft
Retail Bank Servicing Alignment SFT
The training corpus for the Granite retail-bank servicing agent. It is the
released tool-use SFT corpus merged with a servicing-alignment continuation
curriculum that teaches multi-turn behaviours the base corpus does not: what to
do when the customer says "that one", when a policy question interrupts a
transfer, when the agent's own previous turn was wrong, and when the honest
answer is that the agent cannot see what it was asked about.
Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.retail-voice-concise
retail-voice-concise
Made with the whileai SDK · Collection: Register
The same speaking register as
airline-voice-concise,
trained on a different agent. A retail support agent that leads with the
answer and stops.
This exists to test the limitation stated on the airline card: that nothing
there showed the register transfers off airline content. It does. Same
constitution, same recipe, different world, different tools, different
records.
Trained on this set, Qwen3-4B goes from… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/retail-voice-concise.retail-forecast-optimize-benchmark
Retail Forecast-Optimize Benchmark
Benchmark dataset for predict-then-optimize retail inventory decisions.
Contents
File pattern
Description
{scenario}_s{seed}.json
Full policy comparison per scenario instance
samples/sample_*.json
Retail scenario definitions
eval_results.json
Aggregated leaderboard metrics
chronos-2-live-forecasts/
Live Chronos-2 outputs from HF Jobs (when available)
Scenarios
Baseline Operations
Promotion… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/retail-forecast-optimize-benchmark.aprm-sft-thoughts-tau2-retail-policy_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench retail
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-retail-policy_best-adamw30-lp0.giskard-hub-demo-retailqwen3.5-9B-tau2bench-retail-baseline-traces
Qwen3.5-9B tau2-bench retail baseline traces (n=3 × 114)
Three independent evaluation trials of Qwen3.5-9B (no fine-tune, no memory)
on the full 114-task tau2-bench retail set.
Each trial_N.jsonl is one trial; one JSON object per line, one object per task.
Per-trial pass^1 (canonical reward)
trial_1: 72.8%
trial_2: 69.3%
trial_3: 73.7%
Pooled metrics (n=3 across 114 tasks)
pass^1 = 71.9%
pass^2 = 58.2%
pass^3 = 49.1%
Trace fields (per object)… See the full description on the dataset page: https://huggingface.co/datasets/KermitCO/qwen3.5-9B-tau2bench-retail-baseline-traces.retail-bank-conversation-router-data
Retail Bank Conversation Router V6 Hierarchical Data
Governed cross-encoder data for a history-aware OOD, hierarchical intent, action, entity-resolution, and relation classifier.
Rows include only prior visible user/assistant messages and the current user message.
They exclude current-turn tool plans, tool results, expected outputs, and final assistant responses.
Train rows: 21686
Validation rows: 4283
Test rows: 4863
Intent labels: view_accounts, view_cards, freeze_card… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-conversation-router-data.aprm-sft-thoughts-tau2-retail-base_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench retail
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-retail-base_best-adamw30-lp0.tau-dev-task-retail-v1
tau-dev-task-retail-v1
Multi-turn tool-calling SFT dataset (915 records, 765 / 50 / 100 train / validation / test) derived from sierra-research/tau-bench retail-domain trajectories.
Meant to be used as a benchmark dataset for developing and validating data processing, training, and eval workflows involving tool use. Note: tau-bench is a widely-used public benchmark and many recently-trained models may have encountered variants of these trajectories during training, so be mindful of… See the full description on the dataset page: https://huggingface.co/datasets/lefft/tau-dev-task-retail-v1.RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/oopere/RetailBanking-Conversations.aprm-sft-thoughts-tau2-retail
Act-PRM SFT thoughts — tau2-bench retail
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z) - 0.15 * len_frac
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-retail.retail-bank-agent-sft
Retail Bank Agent Tool-Use SFT
This dataset contains 9,000 deterministic, fictional retail-banking
conversations for supervised fine-tuning of a conversational tool-using model.
Dataset: https://huggingface.co/datasets/spkc83/retail-bank-agent-sft
Training revision:
183e7e1ed1aba9c3d7155e7b83b64dc854935055
Source: https://github.com/spkc83/retail-bank-servicing
Model: https://huggingface.co/spkc83/retail-bank-agent-9b
Public POC:… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-agent-sft.github_fetch_huggingface_terminal_9107_8be428b7_source_retail_feedback
Revenue Insights
Monthly revenue analytics dashboards derived from sales transaction records.
License: CC-BY-4.0
Data Owner: Data Analytics Team
retail-faqqwen3.5-9b-tau2-retail-dporetail-products-philippinesretail-bank-router-training-data
Retail Bank dual-head router data
Governed classifier-only data derived from PolyAI Banking77 and UCI CLINC150.
It is not included in generative SFT.
Train rows: 44432
Validation rows: 8589
Test rows: 16260
Domain labels: OOD=0, supported retail banking=1
Supported domain includes greetings, thanks, goodbyes, and assistant-identity questions
Intent labels: 77 Banking77 intents; -100 means no intent supervision
Licenses: CC-BY-4.0
See manifest.json for source revisions, hashes… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-router-training-data.tau2bench-gpt41-retailretail-gpt-4o-tracesRetail_chatbotgrpo-retail-sqlretail_nlpretail-alpaca-finetuneretail-alpaca-formatretail-native-fine-tune-datasets
