datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dr.Sparse-GPT56Luna-eval-b200-rl562-spgemm
Dr.Sparse — GPT-5.6 Luna SpGEMM eval on the 562-matrix pool
Snapshot state: complete (generated 2026-09-19 08:03 UTC)
phase
results
explore (3 branches x 5 iterations)
1686 / 1686
exploit (top-2 branches x 10 iterations)
1124 / 1124
matrices covered
562 / 562
What this is
Tree-search SpGEMM kernel optimization over the 562 matrices of
KinGeorge/Dr.Sparse-RL-train-562,
run on NVIDIA B200 (sm_100) and scored against cuSPARSE.
Model:… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-GPT56Luna-eval-b200-rl562-spgemm.tmax-qwen35-4b-hosted-16x8-b200
TMAX Qwen3.5-4B Hosted 16x8 Pilot
Model-driven terminal trajectories from Slurm job 96400: 16 TMAX tasks with eight independent
attempts per task. Qwen3.5-4B inference used one B200 (DP=1), while up to 128 Docker task
environments ran on a remote hosted server. Thinking was enabled and preserved in the full view.
Trajectory views
Every row includes two chronological lists containing only user and assistant messages:
action_only_trajectories: exact user… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tmax-qwen35-4b-hosted-16x8-b200.
