datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmax-27b-atlas
juiceb0xc0de/tmax-27b-atlas
A brain atlas for allenai/tmax-27b, the largest member of the tmax hybrid SSM/Mamba/transformer family. This model required custom kernel work to fit into the atlas pipeline, and the result is the deepest, most redundant, and most surgically forgiving atlas in the family.
What was run
Model: allenai/tmax-27b
Corpus: 8,965 diverse prompts
Layers probed: all 64
Attention layers: 3, 7, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47, 51, 55, 59… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-27b-atlas.tmax-9b-atlas
allenai/tmax-9b Brain Atlas — The Sweet Spot of the Hybrid Family
Cross-post: I ran a brain atlas on the mid-size tmax. Sub-Zero coverage is concentrated in layers 16–30, so read the surgical headroom numbers as a late-layer snapshot.
model: allenai/tmax-9batlas type: activation census + Sub-Zero brain atlas + OV-circuit SVDcorpus: 8,965 promptslayers: 32attention layers: 3, 7, 11, 15, 19, 23, 27, 31hybrid layers: everything elsesacred (fully probed) layers:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-9b-atlas.tmax-15k-open-instruct
💻 Code ·
🤗 Models & Data ·
📜 Paper ·
📓 Blog
[!NOTE]
For full information, go check out the Tmax paper here.
TMax 15k - Open Instruct
This is the dataset we used to train Tmax 9b (and our other tmax models), formatted for use with our open-instruct fork here.
In general, this is a collection of roughly 15k RL environment instances.
For details on how we generated this dataset and its makeup, please see our paper!
You can find a more generic version of this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tmax-15k-open-instruct.Tmax-Tasks-Clean
Tmax-Tasks-Clean
New: longlongcheck (2026-09-14)
Use configuration longlongcheck, split longlongcheck, for the 431-task snapshot combining the selected Codex and Claude Code repairs with passing historical GPT-6 terminal solutions. The existing splits below retain their earlier data.
From the latest local 452-task repaired snapshot, this split holds out the requested 15 old-pass/current-fail tasks, four additional tasks without any passing current GPT-6 replay… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/Tmax-Tasks-Clean.TMax-15K
💻 Code ·
🤗 Models & Data ·
📜 Paper ·
📓 Blog
[!NOTE]
For full information, go check out the Tmax paper here.
TMax 15k - Open Instruct
This is the dataset we used to train Tmax 9b (and our other tmax models).
This release contains general details on the data.
In general, this is a collection of roughly 15k RL environment instances.
For details on how we generated this dataset and its makeup, please see our paper!
You can find a version of this dataset ready for… See the full description on the dataset page: https://huggingface.co/datasets/allenai/TMax-15K.TMax-15K-Harboragent-task-litecoder-terminal-rl-preview
Apptainer pool for hamishivi/agent-task-litecoder-terminal-rl-preview
This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-litecoder-terminal-rl-preview. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image.
Apptainer images
The pool currently contains… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-litecoder-terminal-rl-preview.tmax-2b-atlas
juiceb0xc0de/tmax-2b-atlas
A brain atlas for allenai/tmax-2b, a hybrid SSM/Mamba/transformer language model. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where the model stores compliance style, which late-layer directions you can edit without breaking reasoning, or whether the… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-2b-atlas.tmax-tasks-skill-taxonomy-20260506-legacy10k-new5k-rltmax-tasks-skill-taxonomy-20260324-1ktmax-sft
💻 Code ·
🤗 Models & Data ·
📜 Paper ·
📓 Blog
[!NOTE]
For full information, go check out the Tmax paper here.
TMax SFT
This is the SFT dataset consisting of traces from Qwen 3.6 27B on environments generated as part of tmax.
This uses a separate ~2k environments, which you can find in more detail here.
This dataset is for use with our fork of open-instruct.
License
This dataset is licensed under ODC-BY.
It is intended for research and educational… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tmax-sft.swerl-tmax-15kswerl-tmax-15k-solvable-gpt-5-6-terra
swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3)
Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every
task was attempted twice as a full agentic episode — real sandbox, real bash,
real verifier — and a task is verified when at least one attempt earned reward.
The last of three artifacts that exist to be compared by task_id:
original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks
hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.TMax-15Ktmax-4b-atlas
juiceb0xc0de/tmax-4b-atlas
A brain atlas for allenai/tmax-4b, the mid-entry hybrid SSM/Mamba/transformer language model from the tmax family. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where the model stores compliance style, which late-layer directions you can edit without… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-4b-atlas.TMax-SFT-16.5K
💻 Code ·
🤗 Models & Data ·
📜 Paper ·
📓 Blog
[!NOTE]
For full information, go check out the Tmax paper here.
TMax SFT
This is the SFT dataset consisting of traces from Qwen 3.6 27B on environments generated as part of tmax.
This uses a separate ~2k environments, which we release details on in here.
For the version of the dataset suitable for use with open-instruct, go here.
License
This dataset is licensed under ODC-BY.
It is intended for research… See the full description on the dataset page: https://huggingface.co/datasets/allenai/TMax-SFT-16.5K.swerl-tmax-15k-rubric-gpt-5-6-sol
swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol)
hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality
label attached as extra columns.
This is not a verified or filtered dataset. Every one of the 14,601 original
records is present. Nothing has been dropped, repaired, or reordered. The labels
are one model's judgement about whether each task is sound enough to be useful RL
training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.tmax-mix-vanilla
tmax training mix — vanilla base (2026-09-01)
The mix a TerminalWorld 9B RL run trains on, rebuilt so that no earlier
evolution round's instruction rewrites are carried in. 1,065 rows, one JSONL
object per task: {prompt, label, metadata}.
Composition
source
rows
state
TerminalWorld (tw_*)
665
prompts reset to the original adapted packages
tmax (task_*)
400
untouched — never evolved
How it was built, and why each step is there
Start… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/tmax-mix-vanilla.TMax-SFT-16.5K-Envtmax-tasks-skill-taxonomy-20260401-10k-verifiedtmax-sft-big
💻 Code ·
🤗 Models & Data ·
📜 Paper ·
📓 Blog
[!NOTE]
For full information, go check out the Tmax paper here.
Tmax-sft-big
Combined SFT dataset derived from a bunch of sources, see below.
The source_dataset field identifies the original source subset for each row.
This was used for 'big SFT' experiments in our paper.
Sources
Source dataset
Number of samples
Original link
allenai__Sera_4.6_Lite_47000
47,464
allenai/Sera-4.6-Lite-47000… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tmax-sft-big.tmax-sft-full-20260403tmax-tasks-skill-taxonomy-20260401-10kswerl-tmax-10k-verifiedswerl-tmax-15k-hardened-prefilter
swerl-tmax-15k hardened, pre-validation-filter (dataset 2 of 3)
The middle artifact of three, which exist to be compared against each other by
task_id:
original — hamishivi/swerl-tmax-15k, unchanged.
hardened, pre-validation-filter — this dataset.
hardened, post-validation-filter — wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra, a strict subset of this one — 7,015 tasks.
What is in it
Tasks where a patch was actually applied, plus tasks originally labelled
CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-hardened-prefilter.swerl-tmax-10ktmax-sft-20260309tmax-sft-full-20260310tmax-tasks-skill-taxonomy-20260320-v2tmax-sft-full-20260317
