CoolFace
Datasetpublic

MoreThought/DeepSWEGym2-Ultra

Dataset Description This dataset is an EXTREMELY filtered and deduplicated version of DeepSWE-Gym2-Edu, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 8.27MB, a total uncompressed size of 8.27GB, and a total of 1000 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Ultra.

sourceHugging Facemitupdated 15d agoView on Hugging Face
3likes1.3kdownloads
Dataset Card

Dataset Description

This dataset is an EXTREMELY filtered and deduplicated version of DeepSWE-Gym2-Edu, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.

It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 8.27MB, a total uncompressed size of 8.27GB, and a total of 1000 examples.

Dataset Details

  • Curated by: MoreThought
  • Funded by: MoreThought
  • Shared by: MoreThought
  • License: MIT

Dataset Sources

Repositorys:

  • https://huggingface.co/datasets/MoreThought/DeepSWE-Gym
  • https://huggingface.co/datasets/SWE-Gym/SWE-Gym
  • https://huggingface.co/datasets/PrimeIntellect/SWE-rebench-V2-Filtered-Verified
  • https://huggingface.co/datasets/nebius/SWE-bench-extra
  • https://huggingface.co/datasets/TIGER-Lab/SWE-Next
  • https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified
  • https://huggingface.co/datasets/swesynth/SWE-Synth

Papers:

  • https://huggingface.co/papers/2504.21798
  • https://huggingface.co/papers/2504.02605
  • https://huggingface.co/papers/2603.20691
  • https://huggingface.co/papers/2602.23866

Uses

  • Improving benchmark results
  • Improving general coding capabilities
  • Training software engineering/coding agents
  • Improving long-context coding

Important

Do NOT try to use this dataset along with other variants or even versions of it if you don't want overlapping examples.

Almost all LLMs cannot handle examples reaching up to 28MB, use specialized scripts to train properly.