datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
b200tmax-qwen35-4b-hosted-16x8-b200
TMAX Qwen3.5-4B Hosted 16x8 Pilot
Model-driven terminal trajectories from Slurm job 96400: 16 TMAX tasks with eight independent
attempts per task. Qwen3.5-4B inference used one B200 (DP=1), while up to 128 Docker task
environments ran on a remote hosted server. Thinking was enabled and preserved in the full view.
Trajectory views
Every row includes two chronological lists containing only user and assistant messages:
action_only_trajectories: exact user… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tmax-qwen35-4b-hosted-16x8-b200.tmax-hosted-qwen35-4b-batch16x16-b200-ilc-job113762
Hosted TMAX trajectories — b200-ilc
This public, ungated dataset contains the validated output of run b200-ilc-113762
(Slurm job 113762) using Qwen/Qwen3.5-4B at revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, served as qwen3.5-4b.
Run configuration
Hardware: 2 × NVIDIA B200 on blackwell1.stanford.edu
Parallelism: DP=2, TP=1
Tasks: 16; attempts per task: 16; rows: 256
Sampling: temperature=0.6, top_p=1.0
Limits: 8192 tokens/turn, 16 turns,
context=131072… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/tmax-hosted-qwen35-4b-batch16x16-b200-ilc-job113762.
