CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IntelligenceLab /Long-Horizon-Terminal-Bench Long-Horizon Terminal-Bench (LHTB) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent into a stateful environment and grades it with hidden, rebuild-from-artifact verifiers — self-reported progress does not count. 📝 Blog: https://zli12321.github.io/LHTB/ 🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.documenttext-generationn<1K136 likes12k downloads10d agoHugging Face02meituan-longcat /R-HORIZON-AMC23 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.tabularn<1K1 likes325 downloads11mo agoHugging Face03arvindh75 /Long-Horizon-Execution Long Horizon Execution This project contains the dataset accompanying the paper "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs" Abstract Does continued scaling of large language models (LLMs) yield diminishing returns? Real-world value often stems from the length of task an agent can complete. We start this work by observing the simple but counterintuitive fact that marginal gains in single-step accuracy can compound into exponential… See the full description on the dataset page: https://huggingface.co/datasets/arvindh75/Long-Horizon-Execution.texttext-generationn<1K16 likes297 downloads1y agoHugging Face04malaiwah /k2-horizon-tiny-fidelity-root-v1 k2-horizon random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/k2-horizon-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-fidelity-root-v1.tabularn<1K0 likes110 downloads19d agoHugging Face05squashenthus /HorizonMath Citation @article{wang2026horizonmathmeasuringaiprogress, title={HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification}, author={Erik Y. Wang and Sumeet Motwani and James V. Roggeveen and Eliot Hodges and Dulhan Jayalath and Charles London and Kalyan Ramakrishnan and Flaviu Cipcigan and Philip Torr and Alessandro Abate}, year={2026}, eprint={2603.15617}, archivePrefix={arXiv}, primaryClass={cs.LG}… See the full description on the dataset page: https://huggingface.co/datasets/squashenthus/HorizonMath.tabularn<1K3 likes109 downloads6mo agoHugging Face06ToolGym /long-horizon-eval long-horizon-eval Evaluation results for long-horizon agent performance Dataset Description This dataset contains evaluation results for agent trajectories, including quality assessments and performance metrics. Dataset Structure The dataset is organized by model name, with each model having separate JSONL files for different experimental passes. long-horizon-eval/ ├── model-1/ │ ├── pass@1.jsonl │ ├── pass@2.jsonl │ └── pass@3.jsonl ├── model-2/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-eval.tabulartext-generation1K<n<10K0 likes87 downloads9mo agoHugging Face07lulululuyi /R-HORIZON-Math500tabular1K<n<10K0 likes69 downloads1y agoHugging Face08meituan-longcat /R-HORIZON-Math500 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.tabular1K<n<10K1 likes65 downloads11mo agoHugging Face09meituan-longcat /R-HORIZON-AIME24 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME24.tabularn<1K1 likes61 downloads11mo agoHugging Face10meituan-longcat /R-HORIZON-Websearch R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Websearch.textn<1K1 likes58 downloads11mo agoHugging Face11meituan-longcat /R-HORIZON-AIME25 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME25.tabularn<1K1 likes57 downloads11mo agoHugging Face12open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V2-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V2-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V2-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V2-32B-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face13open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V1-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V1-32B-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face14open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V6-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V6-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V6-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V6-32B-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face15HorizonRobotics /Aux-Think Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation Shuo Wang, Yongcai Wang, Wanting Li, Xudong Cai, Yucheng Wang, Maiyue Chen, Kaihui Wang, Zhizhong Su, Deying Li, Zhaoxin Fan Dataset Overview The R2R-CoT-320k dataset, the first VLN dataset annotated with CoT reasoning, tailored for the R2R-CE benchmark. We reconstruct step-wise navigation trajectories in the Habitat… See the full description on the dataset page: https://huggingface.co/datasets/HorizonRobotics/Aux-Think.text100K<n<1M2 likes45 downloads10mo agoHugging Face16lodestone-horizon /ShareGPT4Video ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/lodestone-horizon/ShareGPT4Video.textvisual-question-answering10K<n<100K1 likes40 downloads2y agoHugging Face17orinlabs /horizon-1-example-tracestextn<1K2 likes35 downloads5mo agoHugging Face18MarxistLeninist /repro-learning-to-bet-horizon-aware-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes34 downloads2mo agoHugging Face19lulululuyi /R-HORIZON-AMC23tabularn<1K0 likes25 downloads1y agoHugging Face20lulululuyi /R-HORIZON-AIME24tabularn<1K0 likes21 downloads1y agoHugging Face21lulululuyi /R-HORIZON-AIME25tabularn<1K0 likes19 downloads1y agoHugging Face22anonymousAIresearcher /horizonmath HorizonMath HorizonMath is a benchmark of research-level mathematical problems for measuring progress in reasoning toward mathematical discovery with automatic verification. Files data/problems_full.json data/problems_full.jsonl data/baselines.json data/baselines.jsonl croissant.json Loading from datasets import load_dataset problems = load_dataset( "anonymousAIresearcher/horizonmath", name="problems", split="train", ) baselines = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/anonymousAIresearcher/horizonmath.tabularquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face23acvlab /ABotN-Short-Horizon-OVONtext1K<n<10K0 likes16 downloads2mo agoHugging Face24open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V4-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V4-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V4-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V4-32B-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face25MengjiaoMa /Matrix_Horizon3dn<1K0 likes15 downloads2mo agoHugging Face26open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Korean-Superb-27B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Superb-27B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Superb-27B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Superb-27B-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face27lulululuyi /R-HORIZON-Websearchtextn<1K0 likes12 downloads1y agoHugging Face28open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V3-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V3-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V3-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V3-32B-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face29XDGUY22 /event-horizon-indeximage1K<n<10K1 likes7 downloads4mo agoHugging Face30open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V2-27B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V2-27B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V2-27B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V2-27B-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.