datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
R-HORIZON-AMC23
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.VitaBench-2.0MineExplorer
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Tianjie Ju · Yueqing Sun · Zheng Wu · Wei Zhang · Yaqi Huo · Xi Su · Qi Gu · Xunliang Cai · Gongshen Liu · Zhuosheng Zhang
Abstract
Multimodal large language models (MLLMs) have shown strong capabilities in perception,
reasoning, and action generation. However, their ability to sustain exploration in dynamic
open worlds remains unclear. Existing embodied and… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/MineExplorer.R-HORIZON-Math500
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.R-HORIZON-AIME24
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME24.R-HORIZON-AIME25
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME25.R-HORIZON-Websearch
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Websearch.CoreCodeBench-Single
Single Testcases for CoreCodeBench
File Explanation
CoreCodeBench_Single.jsonl: CoreCodeBench single test cases.
CoreCodeBench_Single_Verified.jsonl: Human verified version for CoreCodeBench single test cases.
CoreCodeBench_Single_en.jsonl: English version for CoreCodeBench single test cases.
CoreCodeBench_Function_Empty.jsonl CoreCodeBench function_empty test cases.
Key Explanation
Key
Meaning/Description
id
The unique identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/CoreCodeBench-Single.CoreCodeBench-Multi
Multi Testcases for CoreCodeBench
File Explanation
CoreCodeBench_Multi.jsonl Multi test cases for CoreCodeBench.
CoreCodeBench_Difficult.jsonl More difficult version for CoreCodeBench multi test cases .
Key Explanation
Key
Meaning/Description
id
A list of unique identifiers for the functions to be completed, typically in the format module.path.Class::function.
project
The name of the project this data is associated with.
origin_file
A list… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/CoreCodeBench-Multi.
