datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tau2-bench-verified-airline
tau2-bench-verified — airline domain (mirror)
Mirror of the airline domain from
amazon-agi/tau2-bench-verified
(MIT License), pinned at commit 864350a8971a8f8ee9e7b8472e2edc380a806b0c.
Re-hosted for the OpenRouter native TypeScript benchmark harness so it can fetch
the verified airline tasks + environment DB at runtime.
Contents
tasks/test.jsonl — 50 verified airline tasks. Each row has a single
task_json string column holding one verbatim tau2 v2 task object
(id… See the full description on the dataset page: https://huggingface.co/datasets/abhinavpola/tau2-bench-verified-airline.airline
airline
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (130 of 130)
Fidelity over Runs
99.5% (199 of 200)
Call fidelity
99.93% of 1513
Reference confirmed
130
Verifier derived
128
Trusted
82
Refused
0
Not trusted
48 of 130; 17 the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/airline.airline-voice-concise
airline-voice-concise
Made with the whileai SDK · Used by: recipes/community/airline-voice-concise-under-probe-outcome-filter · Collection: Register
Training data for putting a speaking register into a model's weights. An
airline support agent that leads with the answer and stops, trained so the
register survives with no instruction in the prompt.
Trained on this set, Qwen3-4B goes from 2.2% to 92.1% of held-out replies
in the register, and becomes less likely to omit required… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-voice-concise.airline-resist-jailbreaks
airline-resist-jailbreaks
Made with the whileai SDK · Collection: Robustness
Jailbreak resistance for a customer support agent, trained on simulated
attacks and tested on real ones.
The real attacks come from elder-plinius/L1B3RT4S,
a public library of working jailbreaks. We read it to extract the attack
techniques and never trained on a single string from it. It is the
evaluation set, unseen by the model.
On 165 unseen blocks from a public jailbreak library the agent holds its… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-resist-jailbreaks.tau2-airline-deepseek-distill
τ²-bench airline · DeepSeek teacher trajectories
Successful multi-turn agent trajectories on τ²-bench's
airline domain, generated by running DeepSeek V4 Flash as the agent through the real τ²-bench
harness — same system prompt, same 14 tool schemas, same dialogue loop, same evaluator.
Used to behavior-clone the RL warm start
yuyu0529nya/qwen2.5-7b-tau2-airline-sft-lora,
which is the step 0 of the tau2_airline verl recipe.
Why these exist
GRPO on τ²-bench-airline… See the full description on the dataset page: https://huggingface.co/datasets/yuyu0529nya/tau2-airline-deepseek-distill.aprm-sft-thoughts-tau2-airline-policy_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench airline
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-airline-policy_best-adamw30-lp0.twitter_us_airline_sentiment
Twitter US Airline Sentiment Dataset
This dataset contains tweets about US airlines labeled with sentiment (negative, neutral, positive).
Dataset Details
Total tweets: 14,640
Training set: 11,712 tweets
Validation set: 1,464 tweets
Test set: 1,464 tweets
Classes: negative (0), neutral (1), positive (2)
Files
train.jsonl: Training data (11,712 examples)
validation.jsonl: Validation data (1,464 examples)
test.jsonl: Test data (1,464 examples)
Usage… See the full description on the dataset page: https://huggingface.co/datasets/viethq1906/twitter_us_airline_sentiment.aprm-sft-thoughts-tau2-airline-base_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench airline
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-airline-base_best-adamw30-lp0.US_airline_datasettau-dev-task-airline-v1
tau-dev-task-airline-v1
Multi-turn tool-calling SFT dataset (268 records, 200 / 18 / 50 train / validation / test) derived from sierra-research/tau-bench airline-domain trajectories.
Meant to be used as a benchmark dataset for developing and validating data processing, training, and eval workflows involving tool use. Note: tau-bench is a widely-used public benchmark and many recently-trained models may have encountered variants of these trajectories during training, so be mindful of… See the full description on the dataset page: https://huggingface.co/datasets/lefft/tau-dev-task-airline-v1.airlineSFT_All404_airlines_datasetairline-customer-reviewsThis dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
airline_customer_reviews
This dataset contains a collection of customer reviews and feedback specifically regarding various airlines, such as Binter Canarias and Adria Airways. The entries include short textual statements ranging from positive comments about flight comfort to negative complaints about service quality. Each sample represents a direct user opinion or a label… See the full description on the dataset page: https://huggingface.co/datasets/NikitaSirotkin/airline-customer-reviews.br_separator_alpaca_fare_rules_shorter_length_1000_aed_2024_all_with_airline_30_08_24airlineORPO_allairline_ORPOairline-sentiment-eval.jsonlAirlineChatBot-vectorstoreairline-ORPO-trainalpaca_fare_rules_shorter_length_1000_aed_2024_all_with_airline_30_08_24alpaca_fare_rules_shorter_length_1000_aed_2024_all_with_airline_04_09_24
