CoolFace
20 results

agent-world

Qwen /AgentWorldBench AgentWorldBench AgentWorldBench is a comprehensive evaluation benchmark for language world models, constructed from real-world observations of frontier model trajectories on established benchmarks such as Tool Decathlon, Terminal-Bench 1.0 & 2.0, and OSWorld-Verified. Every evaluation sample is paired with a ground-truth observation obtained from real environment execution, enabling reference-grounded scoring. AgentWorldBench evaluates world modeling quality by scoring each… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/AgentWorldBench.tabulartext-generation1K<n<10K108 likes1.2k downloads3mo agoHugging FaceSnowflake /AgentWorldModel-1KAgentWorldModel-1K Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Zhaoyang Wang1, Canwen Xu2, Boyi Liu2, Yite Wang2, Siwei Han1, Zhewei Yao2, Huaxiu Yao1, Yuxiong He2 1UNC-Chapel Hill   2Snowflake AI Research   Overview AgentWorldModel-1K contains 1,000 fully synthetic, executable, SQL database-backed tool-use environments exposed via a unified MCP (Model Context Protocol) interface, designed for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/AgentWorldModel-1K.71 likes650 downloads7mo agoHugging Facezyz-robotics /agentworld_assetsimage1K<n<10K1 likes325 downloads1y agoHugging Faceagi-agent /os-world-images os-world-images Images for qemu copied from https://huggingface.co/datasets/xlangai. Refer to the original repository for more details. 1 likes183 downloads6mo agoHugging Faceinductionlabs /AgentWorldBench-Terminal-V2 AgentWorldBench-Terminal-V2 AgentWorldBench-Terminal-V2 is our improved subset of the terminal split from Qwen/AgentWorldBench (Zou et al., 2026). Given the history of a Linux terminal session, the model is evaluated on its ability to predict the output of the next command. In the original AgentWorldBench, some samples have ground-truth outputs that depend on environment details missing from the session history. Since the sessions are based on Terminal-Bench environments, the… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/AgentWorldBench-Terminal-V2.tabulartext-generationn<1K0 likes117 downloads28d agoHugging FaceKikoCis /real-world-agent-benchmark Real-World Agent Benchmark (RAB) Paper: Orchestrator and Task-Framing Effects Dominate Fine-Tuning in Real-World Agent Evaluation of a Quantized 31B ModelAuthors: Kiko Cisneros, Claude Sonnet 4.6 · Utopia IA, May 2026Code: github.com/KikoCisBot/gemma4-31b-study 📄 See paper4_orchestrator_dominance.pdf in the Files tab. TL;DR Standard benchmarks (BFCL, HumanEval) do not predict real-world agent capability. A model scoring 95% BFCL scores 0/10 on a real autonomous task… See the full description on the dataset page: https://huggingface.co/datasets/KikoCis/real-world-agent-benchmark.documentn<1K1 likes56 downloads5mo agoHugging Face