internlm/WildClawBench
WildClawBench Hard, practical, end-to-end evaluation for AI agents — in the wild. WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end-to-end, without hand-holding? We drop agents into a live OpenClaw environment — the same open-source personal AI assistant that real users rely on daily — and throw 60 original tasks at them: clipping goal highlights from a football match, negotiating meeting times over multi-round… See the full description on the dataset page: https://huggingface.co/datasets/internlm/WildClawBench.
Update README.md
Sync 31-model leaderboard from GitHub
Align dataset card with GitHub README
Update dataset card for new WildClawBench release
Rename task 6 excel workspace folder
Upload Images/wildclawbench-hermes-agent-v0.5.tar.gz with huggingface_hub
Upload Images/wildclawbench-codex-ubuntu_v0.0.tar with huggingface_hub
Upload Images/wildclawbench-claudecode-ubuntu_v0.2-patched.tar with huggingface_hub
Fix Search Retrieval exec directory names
Upload folder using huggingface_hub
Fix internal task codes (T301/T502/T601/T603/T604) in docstrings with readable task names
Upload task_6_chat_cross_dept_update_zh workspace files
Upload task_4_chat_thread_consolidation workspace files
Update eval.yaml
Create eval.yaml
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Delete workspace/01_Productivity_Flow/.DS_Store
Update README.md
Update README.md
Delete workspace/.DS_Store
Update README.md
Update README.md
Update README.md
Upload assets/lobster_battle.png with huggingface_hub
Update README.md
Upload Images/wildclawbench-ubuntu_v1.2.tar with huggingface_hub
Upload folder using huggingface_hub
initial commit
