datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
puzzlescript-gists
PuzzleScript Human-Authored Games (Full Gist Corpus)
35,705 human-authored PuzzleScript games — the
complete source text of each — collected from public GitHub gists.
This is the full corpus: every distinct gist is kept, and each row is tagged
with its deduplication cluster so you can reduce to a unique set with a one-line
filter. The deduplication is reproducible from the shipped dedup_master.json +
dedup_master.py; non-vanilla PuzzleScript-Plus files are excluded (listed in… See the full description on the dataset page: https://huggingface.co/datasets/smearle/puzzlescript-gists.GLM-5.2-Logic-Puzzles
GLM-5.2 · Logical Puzzles
6000x traces distilled from GLM-5.2 on High reasoning
Token Count: 5M~?
Distribution:
Puzzles:
•Tokenization blindless ex: counting the r's in strawberry
•Goal reasoning ex: the car wash test (theres no car wash question exactly just prompts like it so its not just benchmaxxing)
•Reading comprehension traps
•Temporal reasoning
•Many other categories not worth mentioning
Prompts… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Logic-Puzzles.puzzlezoo
PuzzleZoo
PuzzleZoo is a collection of three reasoning / planning benchmarks designed to evaluate large language models on multi-step procedural problem-solving — the kind of task where one wrong primitive move silently invalidates the rest of the plan.
It is the official evaluation suite for the paper RePoT: Recoverable Program-of-Thought via Checkpoint Repair (Mazaheri, 2026, arXiv:2605.30052), and is released as a standalone benchmark for the broader community.… See the full description on the dataset page: https://huggingface.co/datasets/parsa-mz/puzzlezoo.
