CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01princeton-nlp /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.textn<1K387 likes275k downloads2y agoHugging Face02SWE-bench /SWE-smith SWE-smith Dataset Code • Paper • Site [12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo. The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-smith.texttext-generation10K<n<100K57 likes222k downloads9mo agoHugging Face03harborframework /terminal-bench-3.0 Terminal-Bench 3.0 The primary source is hosted on GitHub, please open issues and pull requests there, not here. The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0 This repo is a mirror of harbor-framework/terminal-bench at tag v3.0.0, laid out so it can be consumed directly by Harbor's git-repos dataset support. How to run via this Huggingface repo Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.5 likes161k downloads1mo agoHugging Face04harborframework /terminal-bench-2.1 Terminal-Bench 2.1 (Harbor git-repos dataset) Harbor website · Harbor GitHub This is a private mirror of the task content from harbor-framework/terminal-bench-2-1 at commit 7131e43 (the source repo has no tagged releases yet), laid out so it can be consumed directly by Harbor's git-repos dataset support. The primary source is the GitHub repository above — please open issues and pull requests there, not here. How to run Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.11 likes158k downloads8d agoHugging Face05RoboDojo-Benchmark /RoboDojo9 likes156k downloads6d agoHugging Face06Nithish2410 /benchmark-bcplustextn<1K0 likes149k downloads6mo agoHugging Face07nasa-impact /WxC-Bench Dataset Card for WxC-Bench WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks. Dataset Details WxC-Bench contains datasets for six key tasks: Nonlocal Parameterization of Gravity Wave Momentum Flux Prediction of Aviation Turbulence Identifying Weather Analogs Generation of Natural Language Weather Forecasts Long-Term Precipitation Forecasting Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.3 likes140k downloads8mo agoHugging Face08SWE-bench /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified.textn<1K162 likes125k downloads1mo agoHugging Face09harborframework /terminal-bench Terminal-Bench The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo instead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate terminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published version; each release is additionally available as an… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench.4 likes102k downloads8d agoHugging Face10princeton-nlp /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite.textn<1K66 likes98k downloads2y agoHugging Face11DL3DV /DL3DV-Benchmarkgated DL3DV Benchmark Download Instructions This repo contains 140 scenes in the DL3DV-benchmark, which are sampled from DL3DV-10K. The repo includes a README, License, colmaps/images (compatible to nerfstudio and 3D gaussian splatting), scene labels and the performances of methods reported in the paper (ZipNeRF, 3DGS, MipNeRF-360, nerfacto, Instant-NGP). The benchmark preview page can be found here https://dl3dv-10k.github.io/DL3DV-Benchmark-Preview/. Download As the whole… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Benchmark.n>1T45 likes90k downloads1y agoHugging Face12harborframework /terminal-bench-2.0Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable. Warning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there. How this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.text-generationn<1K51 likes90k downloads5mo agoHugging Face13SWE-bench-Live /SWE-bench-Live A brand-new, continuously updated SWE-bench-like dataset powered by an automated curation pipeline. For the official data release page, please see microsoft/SWE-bench-Live. Dataset Summary SWE-bench-Live is a live benchmark for issue resolving, designed to evaluate an AI system’s ability to complete real-world software engineering tasks. Thanks to our automated dataset curation pipeline, we plan to update SWE-bench-Live on a monthly basis to provide the… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench-Live/SWE-bench-Live.text1K<n<10K9 likes87k downloads18d agoHugging Face14allenai /reward-bench-results Results for Holisitic Evaluation of Reward Models (HERM) Benchmark Here, you'll find the raw scores for the HERM project. The repository is structured as follows. ├── best-of-n/ <- Nested directory for different completions on Best of N challenge | ├── alpaca_eval/ └── results for each reward model | | ├── tulu-13b/{org}/{model}.json | | └── zephyr-7b/{org}/{model}.json | └── mt_bench/ |… See the full description on the dataset page: https://huggingface.co/datasets/allenai/reward-bench-results.3 likes80k downloads1y agoHugging Face15klieret /swe-bench-dummy-test-datasettextn<1K0 likes75k downloads1y agoHugging Face16harborframework /terminal-bench-science Terminal-Bench-Science The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench-Science is a benchmark of real-world computational research workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub: one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.4 likes71k downloads8d agoHugging Face17SWE-bench /SWE-smith-py SWE-smith Dataset Code • Paper • Site As of 12/14/2025, SWE-smith: Python contains 50908 task instances from 131 GitHub repositories The SWE-smith Dataset is the largest open source dataset for training software engineering agents. All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the documentation. texttext-generation10K<n<100K7 likes63k downloads9mo agoHugging Face18bop-benchmark /hot3d HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here. See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.). More details can be found in the HOT3D paper and BOP 2024 report. image100K<n<1M8 likes60k downloads1y agoHugging Face19SWE-bench /SWE-bench_Multilingual SWE-bench Multilingual Dataset Summary SWE-bench Multilingual is a dataset that tests systems' ability to resolve real-world GitHub issues across a broad range of programming languages. The original SWE-bench is Python-only; this dataset extends the same task format to 9 languages drawn from 41 popular repositories. The dataset collects 300 test Issue-Pull Request pairs. Evaluation is performed by unit test verification, using post-PR behavior as the reference solution. The… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual.textn<1K28 likes57k downloads1mo agoHugging Face20ScaleAI /SWE-bench_Pro Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks. Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os Dataset Structure We follow SWE-Bench Verified (https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified) in terms of dataset structure, with several… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro.textn<1K180 likes55k downloads7mo agoHugging Face21clip-benchmark /wds_objectnetimage1K<n<10K4 likes51k downloads4y agoHugging Face22princeton-nlp /SWE-bench Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the problem_statement… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench.text10K<n<100K146 likes51k downloads2y agoHugging Face23allenai /olmOCR-bench olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.document1K<n<10K290 likes48k downloads7mo agoHugging Face24SWE-bench /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Lite.textn<1K25 likes44k downloads1mo agoHugging Face25nebius /SWE-bench-extraNote: This dataset has an improved and significantly larger successor: SWE-rebench. Dataset Summary SWE-bench Extra is a dataset that can be used to train or evaluate agentic systems specializing in resolving GitHub issues. It is based on the methodology used to build SWE-bench benchmark and includes 6,415 Issue-Pull Request pairs sourced from 1,988 Python repositories. Dataset Description The SWE-bench Extra dataset supports the development of software engineering agents… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-bench-extra.text1K<n<10K47 likes41k downloads1y agoHugging Face26occiglot /tokenizer-wiki-bench Multilingual Tokenizer Benchmark This dataset includes pre-processed wikipedia data for tokenizer evaluation in 45 languages. We provide more information on the evaluation task in general this blogpost. Usage The dataset allows us to easily calculate tokenizer fertility and the proportion of continued words on any of the supported languages. In the example below we take the Mistral tokenizer and evaluate its performance on Slovak. from transformers import AutoTokenizer… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/tokenizer-wiki-bench.text10M<n<100M6 likes38k downloads2y agoHugging Face27initiacms /XLRS-Bench_visual_grounding_en 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_en.image10K<n<100K0 likes37k downloads11mo agoHugging Face28harborframework /terminal-bench-science-lfs Terminal-Bench-Science — task input mirror Large input files for Terminal-Bench-Science tasks, which cannot be committed to git. Tasks pull from here at container build time, pinned to a commit SHA and verified against a checksum file that ships in the task directory. One top-level prefix per task; everything lives under <task-name>/input/. Benchmark contamination canary This dataset is benchmark material. If you are assembling a training corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.0 likes31k downloads1mo agoHugging Face29rethinklab /Bench2Drive Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving. Description Bench2Drive is a benchmark designed for evaluating end-to-end autonomous driving algorithms in the closed-loop manner. It features: Comprehensive Scenario Coverage: Bench2Drive is designed to test AD systems across 44 interactive scenarios, ensuring a thorough evaluation of an AD system's capability to handle real-world driving challenges. Granular Skill Assessment:… See the full description on the dataset page: https://huggingface.co/datasets/rethinklab/Bench2Drive.29 likes30k downloads2y agoHugging Face30futurehouse /lab-bench LAB-Bench The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.imagequestion-answering1K<n<10K51 likes30k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.