bench
Datasets
All datasets matching “bench”SWE-bench_VerifiedDataset Summary
SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process.
The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.SWE-smith
SWE-smith Dataset
Code
•
Paper
•
Site
[12/14/2025] NOTE: We will no longer actively update this dataset.
While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets.
For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo.
The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-smith.WxC-Bench
Dataset Card for WxC-Bench
WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks.
Dataset Details
WxC-Bench contains datasets for six key tasks:
Nonlocal Parameterization of Gravity Wave Momentum Flux
Prediction of Aviation Turbulence
Identifying Weather Analogs
Generation of Natural Language Weather Forecasts
Long-Term Precipitation Forecasting
Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.terminal-bench-3.0
Terminal-Bench 3.0
The primary source is hosted on GitHub, please open issues and pull
requests there, not here.
The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0
This repo is a mirror of harbor-framework/terminal-bench
at tag v3.0.0, laid out so it can be consumed directly by
Harbor's
git-repos dataset support.
How to run via this Huggingface repo
Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.terminal-bench-2.1
Terminal-Bench 2.1 (Harbor git-repos dataset)
Harbor website · Harbor GitHub
This is a private mirror of the task content from
harbor-framework/terminal-bench-2-1
at commit 7131e43
(the source repo has no tagged releases yet), laid out so it can be consumed directly
by Harbor's
git-repos dataset support.
The primary source is the GitHub repository above — please open issues and pull
requests there, not here.
How to run
Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.RoboDojo
