CoolFace
Datasetpublic

ud-nlp/swe-bench-coding-tasks

SWE-Bench Dataset - 8,712 files The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - Get the data Dataset characteristics: Characteristic Data Description An extended benchmark of real-world software engineering tasks with… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/swe-bench-coding-tasks.

sourceHugging Facecc-by-nc-nd-4.0updated 1y agoView on Hugging Face
0likes340downloads
Dataset Card

SWE-Bench Dataset - 8,712 files

The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - [Get the data](https://unidata.pro/datasets/swe-bench-coding-tasks/?utm_source=huggingface-nlp&utm_medium=referral&utm_campaign=swe-bench-coding-tasks)

Dataset characteristics:

CharacteristicData
DescriptionAn extended benchmark of real-world software engineering tasks with enhanced artifacts and broader language coverage
Data typesText
TasksBug fixing, code completion, pull request generation, automated code review
Total number of files8,712
Total number of people30
LabelingAnnotated with golden patches, test patches, post-patch reference states, and metadata stored in parquet files (e.g., repository name, issue/PR identifier, diffs, test results)
Programming languagesC#, Go, PHP, Rust, Kotlin, Ruby

📊 Sample dataset available! For full access, contact us to discuss purchase terms.

Dataset structure

  • Go - Files in Go
  • Scala - Files in Scala

🧩 Like the dataset but need different data? We can collect a custom dataset just for you - learn more about our data collection services here

Similar Datasets:

  1. 1.LLM Text Generation Dataset
  2. 2.Synthetic Printed USA Passports Dataset
  3. 3.DeepFake Videos Dataset

🌐 UniData - your trusted data partner. Unique, accurate, thoroughly collected and annotated data designed to fuel your AI/ML success.