terminal-bench
terminal-bench-3.0
Terminal-Bench 3.0
The primary source is hosted on GitHub, please open issues and pull
requests there, not here.
The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0
This repo is a mirror of harbor-framework/terminal-bench
at tag v3.0.0, laid out so it can be consumed directly by
Harbor's
git-repos dataset support.
How to run via this Huggingface repo
Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.terminal-bench-2.1
Terminal-Bench 2.1 (Harbor git-repos dataset)
Harbor website · Harbor GitHub
This is a private mirror of the task content from
harbor-framework/terminal-bench-2-1
at commit 7131e43
(the source repo has no tagged releases yet), laid out so it can be consumed directly
by Harbor's
git-repos dataset support.
The primary source is the GitHub repository above — please open issues and pull
requests there, not here.
How to run
Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.terminal-bench
Terminal-Bench
The primary source is hosted on GitHub, please open issues and
pull requests there, not here.
Terminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo
instead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate
terminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published
version; each release is additionally available as an… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench.terminal-bench-2.0Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable.
Warning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there.
How this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.terminal-bench-science
Terminal-Bench-Science
The primary source is hosted on GitHub, please open
issues and pull requests there, not here.
Terminal-Bench-Science is a benchmark of real-world computational research
workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a
continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub:
one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.terminal-bench-science-lfs
Terminal-Bench-Science — task input mirror
Large input files for Terminal-Bench-Science
tasks, which cannot be committed to git. Tasks pull from here at container build
time, pinned to a commit SHA and verified against a checksum file that ships in
the task directory.
One top-level prefix per task; everything lives under <task-name>/input/.
Benchmark contamination canary
This dataset is benchmark material. If you are assembling a training corpus,
exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.
