atrost/dsv4-flash-tmax-git-pager-recovery-23
DeepSeek V4 Flash TMax Git Pager Recovery This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use. Composition 6 original parser-clean full last-episode exports. 17 additional manually confirmed… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery-23.
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use.
Composition
- 6 original parser-clean full last-episode exports.
- 17 additional manually confirmed recovery-ending segments from the parser-defect audit.
- All 23 source trials received reward 1.0 and had no Harbor exception.
The 17 recovered rows do not rewrite model output or terminal output. Each selected segment ends on the audited recovery action. Prompt restoration and subsequent working shell use are validated from recording.cast and stored in the audit provenance, while the post-action terminal result is omitted under Harbor's normal episode convention. For six continuation segments, the export removes only a parser-invalid copied-context handoff question and its paired copied answer. The other eleven selected segments contain no parser-invalid response; their trials were previously excluded only because a later continuation contained a malformed copied handoff. All retained task-solving assistant responses parse under the Terminus-2 JSON protocol.
Source and collection
- Environment/task source: `TMaxxx/TMax-15K-Harbor`, revision
48a77eb0b017606c643ed905f96db41672914798. - Model:
openai/dsv4-flash-0731, served IDdsv4-flash-0731. - Sampling: temperature 1.0, top-p 0.95, maximum 200 Terminus-2 turns.
- Harness: Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser.
- Terminal evidence:
recording.castinput/output events were authoritative during manual audit.
Recovery labels:
git_pager_q_exit: 17git_pager_retry_then_q_exit: 6
Files
data/train.parquet: 23 SFT conversations.data/index.parquet: task/replica identity, evidence references, and checksums.manifest/selected-sft-tasks.jsonl: the 19 selected task identities and retained replicas.manifest/run-config.json: collection and reconstruction metadata.manifest/checksums.sha256: checksums for every published artifact.
Important fields include messages, task_id, replica_id, reward, hit_labels, trigger_command, recovery_action, trajectory_format, and trajectory_sha256. The data/train.parquet SHA256 is 6fa996c89ecbec8df15ec13d3a5e0bea73c174a3bcc7be2a03e86fcb217677fd.
Filtering caveat
This is an intentionally narrow behavior dataset, not a representative estimate of TMax task performance. Failed tasks, proactive pager avoidance, unproven temporal correlations, parser-invalid generated task-solving responses, and one ambiguous recovery were excluded.
License and privacy
TMax is distributed under ODC-BY. Raw trials and internal infrastructure logs are not published. Staging and the independently downloaded revision are scanned for tokens, private keys, internal addresses/mounts, and private service identifiers before publication.
