datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AMO-Bench
📐 AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
📄 Paper
🌐 Project Page
💻 Github Repo
Updates
2026.02.05: Leaderboard Update: Qwen3-Max-Thinking achieves a new SOTA with 65.1%, while GLM-4.7 sets a new open-source record at 62.4%!
2025.12.01: We have added Token Efficiency showing the number of output tokens used by models in the leaderboard. Gemini 3 Pro achieves the highest token efficiency among top-performance models!… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/AMO-Bench.LoHoSearch
LoHoSearch: Benchmarking Long-Horizon
Search Agents Beyond the Human Difficulty Ceiling
📃 Paper • 🏆 Benchmark • 📦 Training Data
Abstract
Search agent benchmarks exemplified by BrowseComp have rapidly saturated over the past year, with the strongest models surpassing 90% accuracy. Since these benchmarks are predominantly human-authored, annotators lack a global perspective on entity statistics and cannot systematically maximize search space size and structural… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/LoHoSearch.General365_Public
🧩 General365: Benchmarking General Reasoning in LLMs Across Diverse and Challenging Tasks
📃 Paper • 🌐 Project Page • 🏆 Leaderboard •
💻 Github
📖 Introduction
We present General365, a highly challenging and diverse benchmark for evaluating the general reasoning capabilities in LLMs.
"General Reasoning" refers to reasoning tasks that depend exclusively on general knowledge.
We define general knowledge as knowledge within the K-12 scope (such as common sense… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/General365_Public.MineExplorer
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Tianjie Ju · Yueqing Sun · Zheng Wu · Wei Zhang · Yaqi Huo · Xi Su · Qi Gu · Xunliang Cai · Gongshen Liu · Zhuosheng Zhang
Abstract
Multimodal large language models (MLLMs) have shown strong capabilities in perception,
reasoning, and action generation. However, their ability to sustain exploration in dynamic
open worlds remains unclear. Existing embodied and… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/MineExplorer.
