MoreThought/DeepSWEGym-Edu
Dataset Description This dataset is a heavily filtered version of all the SWE-bench/SWE-smith-lang datasets (expect php) merged together (originally 88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with very complex/long code problems in the original datasets, having an average row size of 101.81kb, a total uncompressed size of 4.29GB, and 42529 examples total.… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Edu.
Dataset Description
This dataset is a heavily filtered version of all the SWE-bench/SWE-smith-lang datasets (expect php) merged together (originally 88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is specifically filtered for rows with very complex/long code problems in the original datasets, having an average row size of 101.81kb, a total uncompressed size of 4.29GB, and 42529 examples total.
Dataset Details
- Curated by: MoreThought
- Funded by: MoreThought
- Shared by: MoreThought
- License: MIT
Dataset Sources
Repositorys:
- https://huggingface.co/datasets/SWE-bench/SWE-smith-py
- https://huggingface.co/datasets/SWE-bench/SWE-smith-go
- https://huggingface.co/datasets/SWE-bench/SWE-smith-rs
- https://huggingface.co/datasets/SWE-bench/SWE-smith-ts
- https://huggingface.co/datasets/SWE-bench/SWE-smith-js
- https://huggingface.co/datasets/SWE-bench/SWE-smith-cpp
- https://huggingface.co/datasets/SWE-bench/SWE-smith-java
- Paper:
- https://huggingface.co/papers/2504.21798
Uses
- Improving benchmark results
- Improving general coding capabilities
- Training software engineering/coding agents
- Improving long-context coding
Important
Do NOT try to use this dataset along with other variants or even versions of it if you don't want overlapping examples.
