meituan-longcat/VitaBench
🌱VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks 📃 Paper • 🌐 Website • 🏆 Leaderboard • 🛠️ Code • 🤗 Dataset 🔔 News [2026-01] Qwen3-Max-Thinking reported our Vita-Bench to evaluate and demonstrate its tool use capabilities (the averge score of 4 domains)!We invite the community to adopt Vita-Bench as the definitive touchstone for tool use performance assessment, and we appreciate diverse utilization & interpretation of our benchmark… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/VitaBench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face