datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmax-27b-atlas
juiceb0xc0de/tmax-27b-atlas
A brain atlas for allenai/tmax-27b, the largest member of the tmax hybrid SSM/Mamba/transformer family. This model required custom kernel work to fit into the atlas pipeline, and the result is the deepest, most redundant, and most surgically forgiving atlas in the family.
What was run
Model: allenai/tmax-27b
Corpus: 8,965 diverse prompts
Layers probed: all 64
Attention layers: 3, 7, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47, 51, 55, 59… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-27b-atlas.tmax-9b-atlas
allenai/tmax-9b Brain Atlas — The Sweet Spot of the Hybrid Family
Cross-post: I ran a brain atlas on the mid-size tmax. Sub-Zero coverage is concentrated in layers 16–30, so read the surgical headroom numbers as a late-layer snapshot.
model: allenai/tmax-9batlas type: activation census + Sub-Zero brain atlas + OV-circuit SVDcorpus: 8,965 promptslayers: 32attention layers: 3, 7, 11, 15, 19, 23, 27, 31hybrid layers: everything elsesacred (fully probed) layers:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-9b-atlas.Tmax-Tasks-Clean
Tmax-Tasks-Clean
New: longlongcheck (2026-09-14)
Use configuration longlongcheck, split longlongcheck, for the 431-task snapshot combining the selected Codex and Claude Code repairs with passing historical GPT-6 terminal solutions. The existing splits below retain their earlier data.
From the latest local 452-task repaired snapshot, this split holds out the requested 15 old-pass/current-fail tasks, four additional tasks without any passing current GPT-6 replay… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/Tmax-Tasks-Clean.agent-task-litecoder-terminal-rl-preview
Apptainer pool for hamishivi/agent-task-litecoder-terminal-rl-preview
This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-litecoder-terminal-rl-preview. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image.
Apptainer images
The pool currently contains… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-litecoder-terminal-rl-preview.tmax-2b-atlas
juiceb0xc0de/tmax-2b-atlas
A brain atlas for allenai/tmax-2b, a hybrid SSM/Mamba/transformer language model. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where the model stores compliance style, which late-layer directions you can edit without breaking reasoning, or whether the… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-2b-atlas.swerl-tmax-15k-solvable-gpt-5-6-terra
swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3)
Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every
task was attempted twice as a full agentic episode — real sandbox, real bash,
real verifier — and a task is verified when at least one attempt earned reward.
The last of three artifacts that exist to be compared by task_id:
original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks
hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.tmax-4b-atlas
juiceb0xc0de/tmax-4b-atlas
A brain atlas for allenai/tmax-4b, the mid-entry hybrid SSM/Mamba/transformer language model from the tmax family. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where the model stores compliance style, which late-layer directions you can edit without… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-4b-atlas.swerl-tmax-15k-hardened-prefilter
swerl-tmax-15k hardened, pre-validation-filter (dataset 2 of 3)
The middle artifact of three, which exist to be compared against each other by
task_id:
original — hamishivi/swerl-tmax-15k, unchanged.
hardened, pre-validation-filter — this dataset.
hardened, post-validation-filter — wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra, a strict subset of this one — 7,015 tasks.
What is in it
Tasks where a patch was actually applied, plus tasks originally labelled
CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-hardened-prefilter.TMax-Agent-SFT-terminus
TMax agent trajectories, rendered for Qwen3.5 thinking SFT
5,645 multi-turn terminal-agent trajectories (5,445 train / 200 holdout) with
a reasoning block on every assistant turn, derived from
allenai/tmax-sft and put
through the same loss-mask contract and length gate as
OpenThoughts-Agent-v1-SFT-terminus.
Converter, tests and launchers: https://github.com/k1ssloo/RST-Train
(scripts/03e_build_tmax_sft.py, tests/test_tmax_convert.py).
Read this first if you came here… See the full description on the dataset page: https://huggingface.co/datasets/NiuNiu0110/TMax-Agent-SFT-terminus.tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified
tmax self-generated tasks — Qwen3.5-9B (verified arm)
The 1,042 tasks from
…-20260919-1k,
each graded against the issue-#12 rubric by the same model that generated them
(hamishivi/Qwen3.5-9B). Using the generator as its own reviewer is deliberate: the
question is whether an open-weights model can carry both halves of the loop. A stronger
reviewer would answer a different question.
The grader sees instruction / setup.sh / tests only. truth is withheld from it,
so it is no better… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified.tmax_filtered_3x3k
tmax_filtered_3x3k
Agentic rollouts on 3,000 TMax
tasks, sampled three times per task from each of four models. One subset per
model; every row is one complete trajectory.
subset
rollouts
deepseek-v4.1-flash
9,000
glm-5.3-flash
9,000
nemotron3-ultra
9,000
nemotron3.5-lightning
9,000
All rollouts were sampled at a 65,536-token context window, recorded in the
context_window column. The subset name is the model alone, since the context
level is a property of… See the full description on the dataset page: https://huggingface.co/datasets/lievan/tmax_filtered_3x3k.tmax-nemo-v09-variable-passdsv4-flash-tmax-git-pager-recovery-23
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use.
Composition
6 original parser-clean full last-episode exports.
17 additional manually confirmed… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery-23.qwen35-4b-tmax-16turn-unfinished-trajectories
Qwen3.5-4B TMAX 16-turn unfinished trajectories
This public, ungated dataset contains the 77 trajectories from TACC Slurm job
891542 that exhausted all 16 allowed turns without the model emitting a
task_complete=true action. They were selected from 256 trajectories over 16 TMAX tasks.
Message format
messages is the canonical conversational trajectory. Every row contains exactly 32 messages,
strictly alternating between objects of the following form:
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/qwen35-4b-tmax-16turn-unfinished-trajectories.dsv4-flash-tmax-git-pager-recovery
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 6 manually audited, SFT-ready terminal-agent trajectories generated by DeepSeek-V4-Flash-0731 in public TMax environments. The primary subset is deliberately narrow: the agent must actually enter a Git pager or foreground TUI, execute a useful recovery action, return to a shell prompt, and finish the task with reward 1.0.
Source and collection
Environment/task source: TMaxxx/TMax-15K-Harbor, pinned… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery.tmax-nemo-v09-variable-pass-no-git-pager-recovery
TMax mixed-outcome RL tasks without Git-pager SFT overlap
This public dataset is a field-preserving derivative of atrost/tmax-nemo-v09-variable-pass at immutable revision aacdcbaff13831e660ae3a203af1324078e3376b. It retains 1402 mixed-outcome RL tasks after removing 0 task(s) that overlap the approved SFT trajectories in atrost/dsv4-flash-tmax-git-pager-recovery at revision af965e1480cd44f33284cd18efa302caf63fd7f6.
The overlap unit is the canonical task, not an individual… See the full description on the dataset page: https://huggingface.co/datasets/atrost/tmax-nemo-v09-variable-pass-no-git-pager-recovery.tmax-qwen35-4b-hosted-16x8-b200
TMAX Qwen3.5-4B Hosted 16x8 Pilot
Model-driven terminal trajectories from Slurm job 96400: 16 TMAX tasks with eight independent
attempts per task. Qwen3.5-4B inference used one B200 (DP=1), while up to 128 Docker task
environments ran on a remote hosted server. Thinking was enabled and preserved in the full view.
Trajectory views
Every row includes two chronological lists containing only user and assistant messages:
action_only_trajectories: exact user… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tmax-qwen35-4b-hosted-16x8-b200.tmax-hosted-qwen35-4b-batch16x16-a100-ilc-job113825
Hosted TMAX trajectories — a100-ilc
This public, ungated dataset contains the validated output of run a100-ilc-113825
(Slurm job 113825) using Qwen/Qwen3.5-4B at revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, served as qwen3.5-4b.
Run configuration
Hardware: 2 × NVIDIA A100-SXM4-80GB on ampere1.stanford.edu
Parallelism: DP=2, TP=1
Tasks: 16; attempts per task: 16; rows: 256
Sampling: temperature=0.6, top_p=1.0
Limits: 8192 tokens/turn, 16 turns… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/tmax-hosted-qwen35-4b-batch16x16-a100-ilc-job113825.tmax-hosted-qwen35-4b-batch16x16-b200-ilc-job113762
Hosted TMAX trajectories — b200-ilc
This public, ungated dataset contains the validated output of run b200-ilc-113762
(Slurm job 113762) using Qwen/Qwen3.5-4B at revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, served as qwen3.5-4b.
Run configuration
Hardware: 2 × NVIDIA B200 on blackwell1.stanford.edu
Parallelism: DP=2, TP=1
Tasks: 16; attempts per task: 16; rows: 256
Sampling: temperature=0.6, top_p=1.0
Limits: 8192 tokens/turn, 16 turns,
context=131072… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/tmax-hosted-qwen35-4b-batch16x16-b200-ilc-job113762.
