agent-world
AgentWorldBench
AgentWorldBench
AgentWorldBench is a comprehensive evaluation benchmark for language world models, constructed from real-world observations of frontier model trajectories on established benchmarks such as Tool Decathlon, Terminal-Bench 1.0 & 2.0, and OSWorld-Verified. Every evaluation sample is paired with a ground-truth observation obtained from real environment execution, enabling reference-grounded scoring.
AgentWorldBench evaluates world modeling quality by scoring each… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/AgentWorldBench.AgentWorldModel-1KAgentWorldModel-1K
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
Zhaoyang Wang1,
Canwen Xu2,
Boyi Liu2,
Yite Wang2,
Siwei Han1,
Zhewei Yao2,
Huaxiu Yao1,
Yuxiong He2
1UNC-Chapel Hill 2Snowflake AI Research
Overview
AgentWorldModel-1K contains 1,000 fully synthetic, executable, SQL database-backed tool-use environments exposed via a unified MCP (Model Context Protocol) interface, designed for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/AgentWorldModel-1K.agentworld_assetsos-world-images
os-world-images
Images for qemu copied from https://huggingface.co/datasets/xlangai. Refer to the original repository for more details.
AgentWorldBench-Terminal-V2
AgentWorldBench-Terminal-V2
AgentWorldBench-Terminal-V2 is our improved subset of the terminal split from
Qwen/AgentWorldBench
(Zou et al., 2026). Given the history of a Linux terminal session,
the model is evaluated on its ability to predict the output of the next command.
In the original AgentWorldBench, some samples have ground-truth outputs that depend on environment
details missing from the session history. Since the sessions are based on Terminal-Bench environments,
the… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/AgentWorldBench-Terminal-V2.real-world-agent-benchmark
Real-World Agent Benchmark (RAB)
Paper: Orchestrator and Task-Framing Effects Dominate Fine-Tuning in Real-World Agent Evaluation of a Quantized 31B ModelAuthors: Kiko Cisneros, Claude Sonnet 4.6 · Utopia IA, May 2026Code: github.com/KikoCisBot/gemma4-31b-study
📄 See paper4_orchestrator_dominance.pdf in the Files tab.
TL;DR
Standard benchmarks (BFCL, HumanEval) do not predict real-world agent capability. A model scoring 95% BFCL scores 0/10 on a real autonomous task… See the full description on the dataset page: https://huggingface.co/datasets/KikoCis/real-world-agent-benchmark.
