datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEnvData-SWE-Trajectory
MEnvData-SWE-Trajectory: Agent Execution Trajectories for Software Engineering
📋 Dataset Description
MEnvData-SWE-Trajectory extends MEnvData-SWE with 3,872 complete agent execution trajectories across 3,005 task instances from 942 repositories in 10 programming languages. These trajectories capture the full problem-solving process of an AI agent tackling real-world software engineering issues.
Key Features
🤖 3,872 Agent Trajectories:… See the full description on the dataset page: https://huggingface.co/datasets/ernie-research/MEnvData-SWE-Trajectory.MEnvData-SWE
MEnvData-SWE: Polyglot Software Engineering Dataset with Executable Environments
📋 Dataset Description
MEnvData-SWE is the largest open-source polyglot dataset of realistic verifiable Docker environments, comprising 3,005 task instances from 942 repositories across 10 programming languages. Each instance includes a fully executable Docker environment with pre-verified test cases, environment setup scripts, and evaluation scripts.
Key… See the full description on the dataset page: https://huggingface.co/datasets/ernie-research/MEnvData-SWE.MEnvBench
MEnvBench: Multi-Language Environment Construction Benchmark
📋 Dataset Description
MEnvBench is a comprehensive benchmark for evaluating multi-language environment building and test execution capabilities, comprising 1,000 task instances (10 languages × 20 repositories × 5 instances) selected from 200 high-quality open-source repositories.
Key Features
🌐 Multi-Language Coverage: 10 programming languages (Python, Java, TypeScript… See the full description on the dataset page: https://huggingface.co/datasets/ernie-research/MEnvBench.
