CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IntelligenceLab /Long-Horizon-Terminal-Bench Long-Horizon Terminal-Bench (LHTB) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent into a stateful environment and grades it with hidden, rebuild-from-artifact verifiers — self-reported progress does not count. 📝 Blog: https://zli12321.github.io/LHTB/ 🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.documenttext-generationn<1K136 likes12k downloads9d agoHugging Face02meituan-longcat /R-HORIZON-AMC23 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.tabularn<1K1 likes342 downloads11mo agoHugging Face03squashenthus /HorizonMath Citation @article{wang2026horizonmathmeasuringaiprogress, title={HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification}, author={Erik Y. Wang and Sumeet Motwani and James V. Roggeveen and Eliot Hodges and Dulhan Jayalath and Charles London and Kalyan Ramakrishnan and Flaviu Cipcigan and Philip Torr and Alessandro Abate}, year={2026}, eprint={2603.15617}, archivePrefix={arXiv}, primaryClass={cs.LG}… See the full description on the dataset page: https://huggingface.co/datasets/squashenthus/HorizonMath.tabularn<1K3 likes107 downloads6mo agoHugging Face04malaiwah /k2-horizon-tiny-fidelity-root-v1 k2-horizon random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/k2-horizon-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-fidelity-root-v1.tabularn<1K0 likes107 downloads18d agoHugging Face05ToolGym /long-horizon-eval long-horizon-eval Evaluation results for long-horizon agent performance Dataset Description This dataset contains evaluation results for agent trajectories, including quality assessments and performance metrics. Dataset Structure The dataset is organized by model name, with each model having separate JSONL files for different experimental passes. long-horizon-eval/ ├── model-1/ │ ├── pass@1.jsonl │ ├── pass@2.jsonl │ └── pass@3.jsonl ├── model-2/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-eval.tabulartext-generation1K<n<10K0 likes86 downloads9mo agoHugging Face06lulululuyi /R-HORIZON-Math500tabular1K<n<10K0 likes69 downloads1y agoHugging Face07meituan-longcat /R-HORIZON-Math500 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.tabular1K<n<10K1 likes67 downloads11mo agoHugging Face08meituan-longcat /R-HORIZON-AIME24 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME24.tabularn<1K1 likes60 downloads11mo agoHugging Face09meituan-longcat /R-HORIZON-AIME25 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME25.tabularn<1K1 likes55 downloads11mo agoHugging Face10open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V2-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V2-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V2-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V2-32B-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face11open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V1-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V1-32B-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face12open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V6-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V6-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V6-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V6-32B-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face13MarxistLeninist /repro-learning-to-bet-horizon-aware-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes38 downloads2mo agoHugging Face14lulululuyi /R-HORIZON-AMC23tabularn<1K0 likes30 downloads1y agoHugging Face15lulululuyi /R-HORIZON-AIME24tabularn<1K0 likes21 downloads1y agoHugging Face16anonymousAIresearcher /horizonmath HorizonMath HorizonMath is a benchmark of research-level mathematical problems for measuring progress in reasoning toward mathematical discovery with automatic verification. Files data/problems_full.json data/problems_full.jsonl data/baselines.json data/baselines.jsonl croissant.json Loading from datasets import load_dataset problems = load_dataset( "anonymousAIresearcher/horizonmath", name="problems", split="train", ) baselines = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/anonymousAIresearcher/horizonmath.tabularquestion-answeringn<1K0 likes20 downloads5mo agoHugging Face17lulululuyi /R-HORIZON-AIME25tabularn<1K0 likes19 downloads1y agoHugging Face18MengjiaoMa /Matrix_Horizon3dn<1K0 likes16 downloads2mo agoHugging Face19open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V4-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V4-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V4-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V4-32B-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face20open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Korean-Superb-27B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Superb-27B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Superb-27B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Superb-27B-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face21open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Superb-27B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Superb-27B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Superb-27B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Superb-27B-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face22open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V3-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V3-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V3-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V3-32B-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face23open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V2-27B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V2-27B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V2-27B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V2-27B-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face24open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Avengers-V5-32B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V5-32B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V5-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V5-32B-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face25open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Korean-Superb-22B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Superb-22B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Superb-22B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Superb-22B-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face26open-llm-leaderboard /Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V3-27B-detailsgated Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V3-27B Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V3-27B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V3-27B-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.