CoolFace
Datasetpublic

tarsur385/swebench-verified-top15-embeddings-8k

SWE-bench Verified — top-15 model embeddings (Qwen3-8B, 8k) Frozen Qwen3-8B pooled state & action embeddings (4096-d) for every agent step of the top-15 models (by SWE-bench Verified resolved rate) from tarsur385/swev-trm-trajectories-25models. Encoder run at 8k max-model-len (last-token pooling; tail-truncated), same pipeline as the DeepSWE embeddings. One row per step. 287,251 steps · 7,011 trajectories · 500 tasks · 15 models. Columns column type… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/swebench-verified-top15-embeddings-8k.

sourceHugging Facemitupdated 5d agoView on Hugging Face
0likes47downloads
Dataset Card

SWE-bench Verified — top-15 model embeddings (Qwen3-8B, 8k)

Frozen Qwen3-8B pooled state & action embeddings (4096-d) for every agent step of the top-15 models (by SWE-bench Verified resolved rate) from tarsur385/swev-trm-trajectories-25models. Encoder run at 8k max-model-len (last-token pooling; tail-truncated), same pipeline as the DeepSWE embeddings.

One row per step. 287,251 steps · 7,011 trajectories · 500 tasks · 15 models.

Columns

columntypedescription
trajectory_idstringone rollout (model + task)
step_idxint32step order in the rollout
task_idstringSWE-bench Verified instance id
modelstringone of the top-15 models
configstringswebench_verified/{train,val}
rewardfloat32trajectory outcome: 1.0 resolved / 0.0 not
state_embeddinglist<float16>[4096]frozen Qwen3-8B pooled state
action_embeddinglist<float16>[4096]frozen Qwen3-8B pooled action

Top-15 models: claude-opus-4.6, claude-4.5-opus-high, gemini-3-flash-high, minimax-m2.5-high, glm-5-high, gpt-5.2-{codex,high,1211}, claude-4-opus, claude-4.5-haiku-high, gpt-5.1-{codex-medium,medium}, claude-4-sonnet, kimi-k2-thinking, minimax-m2.