datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cua-speedrun-trajectories
CUA Speedrun results
Trajectory archives and measured results for the benchmark sets used by CUA Speedrun.
Unanimous-295
Results on the 295-task OSWorld set used by CUA Speedrun. Except for the explicitly reported GPT-5.6 Luna aggregate row, scores are the mean normalized verifier score across the 295 tasks and may include partial credit. Times are measured task time per task.
Model
Average score
Average task time
Average cost per task
Kimi K3
85.07%… See the full description on the dataset page: https://huggingface.co/datasets/anonymousmypcbench/cua-speedrun-trajectories.llm_speedrun
LLM Speedrun token streams
Pre-tokenized training artifacts for the LLM speedrun exercises.
File
Description
Tokens
tokenizer_50M.bpe
JSON-serialized BPE tokenizer
—
fineweb-edu-10BT.shuffle.bin
Shuffled FineWeb-Edu sample/10BT token stream
9,440,023,113
smoltalk.shuffle.bin
Shuffled SmolTalk data/all token stream
875,269,408
The .bin files are headerless, little-endian unsigned 16-bit token IDs and can be memory-mapped with NumPy:
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/zkolter/llm_speedrun.speedrunbench
SpeedrunBench
LLM agents optimizing speedruns across multiple games. Each run ships as a silent video of the run that
was scored, and (except for Tuxemon) the input tape that produced it: the frame to goal and lower
is always better.
config
game
metric
supertux
SuperTux
frames to goal
astray
Astray
frames to maze exit
tuxemon
Tuxemon
frames to first gym
pokemon_blue
Pokémon Blue
frames to first badge
sml
Super Mario Land
frames to clear 1-1
smb
Super Mario… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/speedrunbench.nanogpt-speedrunData_NanoVLMcifar10-speedruncifar100-speedrun5mtest-hackathon
