CoolFace
22 results

bench

princeton-nlp /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.textn<1K387 likes270k downloads2y agoHugging FaceSWE-bench /SWE-smith SWE-smith Dataset Code • Paper • Site [12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo. The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-smith.texttext-generation10K<n<100K57 likes223k downloads9mo agoHugging Facenasa-impact /WxC-Bench Dataset Card for WxC-Bench WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks. Dataset Details WxC-Bench contains datasets for six key tasks: Nonlocal Parameterization of Gravity Wave Momentum Flux Prediction of Aviation Turbulence Identifying Weather Analogs Generation of Natural Language Weather Forecasts Long-Term Precipitation Forecasting Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.3 likes167k downloads8mo agoHugging Faceharborframework /terminal-bench-3.0 Terminal-Bench 3.0 The primary source is hosted on GitHub, please open issues and pull requests there, not here. The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0 This repo is a mirror of harbor-framework/terminal-bench at tag v3.0.0, laid out so it can be consumed directly by Harbor's git-repos dataset support. How to run via this Huggingface repo Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.5 likes166k downloads1mo agoHugging Faceharborframework /terminal-bench-2.1 Terminal-Bench 2.1 (Harbor git-repos dataset) Harbor website · Harbor GitHub This is a private mirror of the task content from harbor-framework/terminal-bench-2-1 at commit 7131e43 (the source repo has no tagged releases yet), laid out so it can be consumed directly by Harbor's git-repos dataset support. The primary source is the GitHub repository above — please open issues and pull requests there, not here. How to run Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.12 likes162k downloads9d agoHugging FaceRoboDojo-Benchmark /RoboDojo10 likes158k downloads22h agoHugging Face