CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01princeton-nlp /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite.textn<1K67 likes90k downloads2y agoHugging Face02SWE-bench /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Lite.textn<1K25 likes46k downloads1mo agoHugging Face03sailplane /SWE-bench_Lite_filtered0 likes4.7k downloads2y agoHugging Face04sailplane /SWE-bench_Lite_filtered_10 likes3.3k downloads2y agoHugging Face05R2E-Gym /SWE-Bench-Litetextn<1K0 likes1.3k downloads2y agoHugging Face06princeton-nlp /SWE-bench_Lite_oracle Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_Lite_oracle includes a formatting of each instance using the "Oracle" retrieval… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_oracle.textn<1K4 likes579 downloads2y agoHugging Face07adityasoni17 /SWE-bench_Lite-code-searchtextn<1K0 likes384 downloads9mo agoHugging Face08princeton-nlp /SWE-bench_Lite_bm25_27K Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_Lite_bm25_27K includes a formatting of each instance… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_bm25_27K.textn<1K1 likes351 downloads2y agoHugging Face09princeton-nlp /SWE-bench_Lite_bm25_13K Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_Lite_bm25_13K includes a formatting of each instance… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_bm25_13K.textn<1K1 likes350 downloads2y agoHugging Face10manna-ai /SWE-bench_Lite_Unstabletextn<1K0 likes318 downloads2y agoHugging Face11czlll /SWE-bench_Litetextn<1K2 likes295 downloads2y agoHugging Face12melissapan /swe-bench-lite-agent-traces-v14 AgentBRANE SWE-bench Lite Agent Traces v14 This release contains the 1,890 harness-native agent traces selected by the sealed SWE-bench Lite v14 publication record (1,379/1,890 resolved, 73.0%). It includes Claude Code, Codex, and Pi sessions across seven models and three replicates. No internal research notes are included. Load the observation table: from datasets import load_dataset traces = load_dataset("melissapan/swe-bench-lite-agent-traces-v14", split="train") Each row… See the full description on the dataset page: https://huggingface.co/datasets/melissapan/swe-bench-lite-agent-traces-v14.tabulartext-generation1K<n<10K0 likes286 downloads9d agoHugging Face13ashedwards /swe-bench-lite3 Dataset Card for "swe-bench-lite3" More Information needed textn<1K0 likes261 downloads2y agoHugging Face14Roy029 /swebenchlite_deletetextn<1K0 likes251 downloads2y agoHugging Face15exploiter345 /SWE-bench_Verified_Lite_Annt Dataset details Appended difficulty annotations provided by OpenAI here textn<1K0 likes191 downloads2y agoHugging Face16Luwayy /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/Luwayy/SWE-bench_Lite.textn<1K0 likes190 downloads2y agoHugging Face17mteb /SWEbenchLiteRR SWEbenchLiteRR An MTEB dataset Massive Text Embedding Benchmark Software Issue Localization. Task category t2t Domains Programming, Written Reference https://www.swebench.com/Source datasets: tarsur909/mteb-swe-bench-lite-reranking How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("SWEbenchLiteRR") evaluator = mteb.MTEB([task]) model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SWEbenchLiteRR.texttext-ranking1M<n<10M0 likes159 downloads1y agoHugging Face18ricardo-larosa /SWE-bench_Lite_Dev_Extendedtextn<1K0 likes155 downloads2y agoHugging Face19swesynth /SWE-Bench_Lite-logstextn<1K0 likes153 downloads2y agoHugging Face20sailplane /SWE-bench_Litetextn<1K0 likes148 downloads2y agoHugging Face21rasdani /SWE-bench_Lite_oracle_32kimport datasets from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-1.7B") ds = datasets.load_dataset("princeton-nlp/SWE-bench_Lite_oracle", split="test") def count_tokens(text): return len(tokenizer.encode(text)) ds = ds.map(lambda x: {"num_tokens": count_tokens(x["text"])}, num_proc=10) ds = ds.filter(lambdax: x["num_tokens"] <= 32_000) textn<1K0 likes131 downloads1y agoHugging Face22hrw /SWE-bench_Litetextn<1K0 likes126 downloads2y agoHugging Face23Bertsekas /SWE-Bench_Lite_UTBoostDataset Summary In this dataset, we replace some test suites in princeton-nlp/SWE-bench_Verified with augmented test cases to enable a more rigorous evaluation of SWE-Bench. UTBoost was accepted in ACL 2025, and we have opened-sourced the code and metadata in https://github.com/CUHK-Shenzhen-SE/UTBoost. Dataset Structure An example of a SWE-bench datum is as follows: instance_id: (str) - A formatted instance identifier, usually as repo_owner__repo_name-PR-number. patch: (str) - The gold patch… See the full description on the dataset page: https://huggingface.co/datasets/Bertsekas/SWE-Bench_Lite_UTBoost.textn<1K1 likes109 downloads1y agoHugging Face24eaalghamdi /swe_bench_Lite_p_agenttextn<1K0 likes108 downloads1y agoHugging Face25synthetic-code-training /swe_doc_gen_SWE-bench_Lite_testtabularn<1K0 likes97 downloads1y agoHugging Face26adityasoni17 /SWE-bench_Lite-locagenttextn<1K0 likes92 downloads8mo agoHugging Face27rasdani /SWE-bench_Lite_oracle_easyfrom datasets import load_dataset from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B") ds = load_dataset("princeton-nlp/SWE-bench_Verified", split="test") ds_lite = load_dataset("princeton-nlp/SWE-bench_Lite_oracle", split="test") def count_tokens(text): return len(tokenizer.encode(text)) ds_easy = ds.filter(lambda x: x["difficulty"] == "<15 min fix") ds_easy_lite = ds_lite.filter(lambda x: x["instance_id"] in ds_easy["instance_id"])… See the full description on the dataset page: https://huggingface.co/datasets/rasdani/SWE-bench_Lite_oracle_easy.tabularn<1K0 likes90 downloads1y agoHugging Face28xiaoy1028 /SWE-bench_Lite_bm25_50k SWE-bench_Lite Generated via https://github.com/SWE-bench/SWE-bench/blob/main/swebench/inference/make_datasets/README.md Set k=40 to generated the bm25 search results first, and then use both k=40 and max_context_len=50000 to generate the dataset. Patch file recall: 64.67; higher than the official princeton-nlp/SWE-bench_Lite_bm25_13K (recall=33.67) and princeton-nlp/SWE-bench_Lite_bm25_27K (recall=49.00) textn<1K0 likes82 downloads2y agoHugging Face29xiaoy1028 /SWE-bench_Lite_bm25_500kSWE-bench_Lite Generated via https://github.com/SWE-bench/SWE-bench/blob/main/swebench/inference/make_datasets/README.md Set k=100 to generated the bm25 search results first, and then use both k=100 and max_context_len=500000 to generate the dataset. Patch file recall: 90.67; higher than the official princeton-nlp/SWE-bench_Lite_bm25_13K (recall=33.67) and princeton-nlp/SWE-bench_Lite_bm25_27K (recall=49.00) textn<1K0 likes81 downloads2y agoHugging Face30xiaoy1028 /SWE-bench_Lite_bm25_100k SWE-bench_Lite Generated via https://github.com/SWE-bench/SWE-bench/blob/main/swebench/inference/make_datasets/README.md Set k=40 to generated the bm25 search results first, and then use both k=40 and max_context_len=100000 to generate the dataset. Patch file recall: 73.33; higher than the official princeton-nlp/SWE-bench_Lite_bm25_13K (recall=33.67) and princeton-nlp/SWE-bench_Lite_bm25_27K (recall=49.00) textn<1K0 likes78 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.