datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.R-HORIZON-AMC23
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.HorizonMath
Citation
@article{wang2026horizonmathmeasuringaiprogress,
title={HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification},
author={Erik Y. Wang and Sumeet Motwani and James V. Roggeveen and Eliot Hodges and Dulhan Jayalath and Charles London and Kalyan Ramakrishnan and Flaviu Cipcigan and Philip Torr and Alessandro Abate},
year={2026},
eprint={2603.15617},
archivePrefix={arXiv},
primaryClass={cs.LG}… See the full description on the dataset page: https://huggingface.co/datasets/squashenthus/HorizonMath.k2-horizon-tiny-fidelity-root-v1
k2-horizon random CPU fixture root
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/k2-horizon-tiny-random-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-fidelity-root-v1.long-horizon-eval
long-horizon-eval
Evaluation results for long-horizon agent performance
Dataset Description
This dataset contains evaluation results for agent trajectories, including quality assessments and performance metrics.
Dataset Structure
The dataset is organized by model name, with each model having separate JSONL files for different experimental passes.
long-horizon-eval/
├── model-1/
│ ├── pass@1.jsonl
│ ├── pass@2.jsonl
│ └── pass@3.jsonl
├── model-2/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-eval.R-HORIZON-Math500R-HORIZON-Math500
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.R-HORIZON-AIME24
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME24.R-HORIZON-AIME25
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AIME25.Saxo__Linkbricks-Horizon-AI-Avengers-V2-32B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V2-32B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V2-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V2-32B-details.Saxo__Linkbricks-Horizon-AI-Avengers-V1-32B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V1-32B-details.Saxo__Linkbricks-Horizon-AI-Avengers-V6-32B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V6-32B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V6-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V6-32B-details.repro-learning-to-bet-horizon-aware-traces
Agent traces
Agent sessions published from a Trackio Logbook.
R-HORIZON-AMC23R-HORIZON-AIME24horizonmath
HorizonMath
HorizonMath is a benchmark of research-level mathematical problems for measuring progress in reasoning toward mathematical discovery with automatic verification.
Files
data/problems_full.json
data/problems_full.jsonl
data/baselines.json
data/baselines.jsonl
croissant.json
Loading
from datasets import load_dataset
problems = load_dataset(
"anonymousAIresearcher/horizonmath",
name="problems",
split="train",
)
baselines = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/anonymousAIresearcher/horizonmath.R-HORIZON-AIME25Matrix_HorizonSaxo__Linkbricks-Horizon-AI-Avengers-V4-32B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V4-32B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V4-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V4-32B-details.Saxo__Linkbricks-Horizon-AI-Korean-Superb-27B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Superb-27B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Superb-27B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Superb-27B-details.Saxo__Linkbricks-Horizon-AI-Superb-27B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Superb-27B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Superb-27B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Superb-27B-details.Saxo__Linkbricks-Horizon-AI-Avengers-V3-32B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V3-32B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V3-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V3-32B-details.Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V2-27B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V2-27B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V2-27B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V2-27B-details.Saxo__Linkbricks-Horizon-AI-Avengers-V5-32B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Avengers-V5-32B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Avengers-V5-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Avengers-V5-32B-details.Saxo__Linkbricks-Horizon-AI-Korean-Superb-22B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Superb-22B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Superb-22B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Superb-22B-details.Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V3-27B-details
Dataset Card for Evaluation run of Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V3-27B
Dataset automatically created during the evaluation run of model Saxo/Linkbricks-Horizon-AI-Korean-Avengers-V3-27B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Saxo__Linkbricks-Horizon-AI-Korean-Avengers-V3-27B-details.
