datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apigen-tau-bench-split-turntaubench_traces_training_data
TauBench Traces Training Data
This dataset contains traces of conversations between a tool-using AI agent and users, formatted for fine-tuning.
Dataset Structure
The data is organized in JSONL format, where each line contains a conversation in the following structure:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "...", "tool_calls": [...]},
{"role": "tool", "tool_call_id":… See the full description on the dataset page: https://huggingface.co/datasets/jkazdan/taubench_traces_training_data.tau-bench-synthetic
tau-bench-synthetic
Synthetic tool-use training data for tau-bench, generated using a GT-first task construction pipeline with GLM-5 (via Fireworks API) as the trajectory generator.
Overview
This dataset was built to train small LLMs (e.g., Qwen3-1.7B) on multi-turn tool-use tasks without using the original tau-bench evaluation set. The pipeline follows a GT-first approach: ground-truth actions are constructed programmatically from the database, then an LLM generates… See the full description on the dataset page: https://huggingface.co/datasets/fuvty/tau-bench-synthetic.taubench-sonnet-traces
taubench-sonnet-traces
Complete HTTP-level agentic traces from running taubench benchmark tasks through an instrumented reverse proxy.
Each trace captures full request/response pairs including system prompts, user messages, assistant responses, tool calls and results, and token usage metadata.
Stats
Total sessions: 165
Multi-turn sessions (2+ LLM calls): 165
Total records: 8740
Total LLM requests: 4370
Format
Raw JSONL traces from the instrumented proxy. Each… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/taubench-sonnet-traces.taubench-gemini-traces
taubench-gemini-traces
Complete HTTP-level agentic traces from running taubench_gemini benchmark tasks through an instrumented reverse proxy.
Each trace captures full request/response pairs including system prompts, user messages, assistant responses, tool calls and results, and token usage metadata.
Stats
Total sessions: 115
Multi-turn sessions (2+ LLM calls): 115
Total records: 5744
Total LLM requests: 2872
Format
Raw JSONL traces from the instrumented… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/taubench-gemini-traces.TAU-Benchmarktau-bench-retail-train-next-action-all-step-score-v0.2tau-bench-retail-train-next-actiontrajectory from tau-bench retail train set with GPT-5 reasoning_effort='high'https://github.com/sierra-research/tau-bench/blob/main/tau_bench/envs/retail/tasks_train.py
tau-bench-retail-train-next-action-all-stepaprm-amityco_apigen_tau_bench_split_turntau-bench-airlinetaubench-tool-calling-Qwen2.5-7B-Instruct-0.0_range_0-10_user-gpt-4o-llm_1116210635aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-REv10-ap1-b019tau-bench-retail-train-next-action-hard-v0.2sample from amityco/tau-bench-retail-train-next-action-all-step-score-v0.2 with model Qwen/Qwen3-4B-Thinking-2507
with 8 response
filter only score <=0.1
aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-RE0-ap1-b069TAU-Benchmark-Tinytau-bench-retail-train-next-action-all-step-v0.2aprm-sft_thinkact-Eact_prm_taubench_retail_100_fs1-Gaprm_qwen3_ap-S42-R4e5_gs4-ap1-b039aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-RE0-ap1-b049aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-REv10-ap1-b059tau-bench-retail-train-next-action-all-step-scoreaprm-sft_genthinkact-ENact_prm_taubench_airline_500_fs1-GEaprm_qwen3_ap-SE42-REv10-ap1-b029aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-REv10-ap1-b029aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-RE0-ap1-b059aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-RE0-ap1-b089aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-REv10_lr1e04-b009wmo-tau-bench-realaprm-sft_thinkact-Eact_prm_taubench_retail_100_fs1-Gaprm_qwen3_ap-S42-R4e5_gs4-ap1-b009aprm-sft_genthinkact-ENact_prm_taubench_retail_500_fs1-GEaprm_qwen3_ap-SE42-RE0-ap1-b029aprm-sft_thinkact-Eaprm_taubench_retail_100_fs1-Gaprm_qw3_ap-S42-R1r1e03_bs4_gs4-ap1-b069
