eewer/terminal-midtraining-trace-scores
terminal-midtraining trace scores Per-trace metadata-free structural scores for 1,439,376 terminal/SWE agent traces collected from the terminal-agent midtraining union universe plus v0.11 terminal-swe, v0.10-full extra, and normalized HF sources (AgentTrove, Nemotron-Terminal-Corpus, NTST, TaskTrove, SERA, SWE-Hero, SWE-rebench). trace_scores.parquet — one row per trace: identity (id, source, collection, harness, family, status, interaction_style), sizes, and the features… See the full description on the dataset page: https://huggingface.co/datasets/eewer/terminal-midtraining-trace-scores.
terminal-midtraining trace scores
Per-trace metadata-free structural scores for 1,439,376 terminal/SWE agent traces collected from the terminal-agent midtraining union universe plus v0.11 terminal-swe, v0.10-full extra, and normalized HF sources (AgentTrove, Nemotron-Terminal-Corpus, NTST, TaskTrove, SERA, SWE-Hero, SWE-rebench).
trace_scores.parquet— one row per trace: identity (id, source, collection, harness, family, status, interactionstyle), sizes, and the features (torrefrate, efcevents, errorobsrate, obsemptyrate, retries, samerepeatrate, commanddiversity, verifycallrate, verifynearend) plus pass/fail passthrough (outcomereward, outcome_passed) where the source provided an outcome. -1.0 means "metric not applicable" (e.g. no tool use).analysis_stats.json— aggregates overall / bystyle / bycollection / bysource / byharness / by_status.SCORING.md— feature definitions and interaction-style extraction notes.
Features are deterministic, computed from the messages alone (no LLM judge, no pass-rate-based ranking). See SCORING.md for details.
