datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentWorldBench
AgentWorldBench
AgentWorldBench is a comprehensive evaluation benchmark for language world models, constructed from real-world observations of frontier model trajectories on established benchmarks such as Tool Decathlon, Terminal-Bench 1.0 & 2.0, and OSWorld-Verified. Every evaluation sample is paired with a ground-truth observation obtained from real environment execution, enabling reference-grounded scoring.
AgentWorldBench evaluates world modeling quality by scoring each… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/AgentWorldBench.AgentWorldBench-Terminal-V2
AgentWorldBench-Terminal-V2
AgentWorldBench-Terminal-V2 is our improved subset of the terminal split from
Qwen/AgentWorldBench
(Zou et al., 2026). Given the history of a Linux terminal session,
the model is evaluated on its ability to predict the output of the next command.
In the original AgentWorldBench, some samples have ground-truth outputs that depend on environment
details missing from the session history. Since the sessions are based on Terminal-Bench environments,
the… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/AgentWorldBench-Terminal-V2.AgentWorldBench
AgentWorldBench
AgentWorldBench is a comprehensive evaluation benchmark for language world models, constructed from real-world observations of frontier model trajectories on established benchmarks such as Tool Decathlon, Terminal-Bench 1.0 & 2.0, and OSWorld-Verified. Every evaluation sample is paired with a ground-truth observation obtained from real environment execution, enabling reference-grounded scoring.
AgentWorldBench evaluates world modeling quality by scoring each… See the full description on the dataset page: https://huggingface.co/datasets/vishal8484/AgentWorldBench.
