kaitchup/DeepSWE1.1-trajectories-Qwen3.8-27B
DeepSWE 1.1 trajectories: Qwen3.8-27B agents and baselines This dataset contains agent trajectories and evaluation results from 7 complete runs on DeepSWE 1.1. The main experiments evaluate Qwen3.8-27B through Mini-SWE, Claude Code, and Pi. Muse-Glimmer-30B and Qwen3.6-27B are included as weaker reference baselines. Every run covers all 113 benchmark tasks. Altogether, the dataset contains: 791 task-level result records; 791 compressed agent trajectories; 425 submitted text… See the full description on the dataset page: https://huggingface.co/datasets/kaitchup/DeepSWE1.1-trajectories-Qwen3.8-27B.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face