datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLM-5.2-Logic-Puzzles
GLM-5.2 · Logical Puzzles
6000x traces distilled from GLM-5.2 on High reasoning
Token Count: 5M~?
Distribution:
Puzzles:
•Tokenization blindless ex: counting the r's in strawberry
•Goal reasoning ex: the car wash test (theres no car wash question exactly just prompts like it so its not just benchmaxxing)
•Reading comprehension traps
•Temporal reasoning
•Many other categories not worth mentioning
Prompts… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Logic-Puzzles.Word-Puzzles-ARC-Unique-50000
Word-Puzzles-ARC-Unique-46000
This dataset is a synthetic 46,000-row word-puzzle corpus focused on answerable reasoning tasks with explicit gold answers.
Version
This upload corresponds to the harder v2 build.
Bucket mix
15,000 formal deduction
12,500 constraint-based lexical deduction
10,000 symbolic substitution
7,500 semantic association
1,000 riddles
Hardening changes in v2
formal puzzles use 6 entities instead of 5
lexical… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Word-Puzzles-ARC-Unique-50000.logic-grid-puzzles-training-pool
Logic grid puzzles training pool
Logic grid puzzles: a row of positions, a handful of attributes with one value per position, and a
list of clues that together admit exactly one arrangement. Two sets drawn for this pool by
generators run here under the seeds recorded below, and two public datasets read at the pinned
revisions named below, laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 390945 rows, one JSON… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logic-grid-puzzles-training-pool.mt_puzzles
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
MT-Puzzles is a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities:
Word Guess Guess the secret word in min attempts while environment gives feedback on how close the guess is at each turn.
Movie Recommendation: Probe the user to decode the user preference function for N turns. Pick a movie for the… See the full description on the dataset page: https://huggingface.co/datasets/arianhosseini/mt_puzzles.rukh-puzzles-split
chorcat/rukh-puzzles-split
Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.knights-knaves-puzzles
Knights and Knaves Logic Puzzles Dataset
A comprehensive dataset of Knights and Knaves logic puzzles ranging from 3 to 14 inhabitants.
Each puzzle requires logical deduction to determine who tells the truth (knights) and who lies (knaves).
The dataset includes detailed chain-of-thought reasoning for each solution.
Dataset Description
This dataset contains 12,000 Knights and Knaves logic puzzles. In these puzzles:
Knights always tell the truth
Knaves always lie
The goal… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/knights-knaves-puzzles.PuzzlesAndPatterns-1k
Puzzles and Patterns 1k
This is a SUPER simple dataset containing ~1400 short puzzles, mostly based around numbers. The intention of this dataset is to teach reasoning LLMs to recognize patterns as fast and effectively as possible.
Why no reasoning content?
I have decided to exclude reasoning from this dataset entirely, this dataset should be used in training base on answer, which encourages exploration by the model, rather than purely mimicking a style.
chess-debate-puzzles
Chess Debate Puzzles
A stratified sample of Lichess mid/endgame chess puzzles annotated with Stockfish-evaluated
moves across ten centipawn-quality bands. Designed for experiments in the spirit of
AI Safety via Debate (Irving et al., 2018), where two AI
agents argue for different moves and a judge must identify the objectively better one.
Motivation
Debate as an alignment technique asks whether a human (or AI) judge can identify the correct
answer when two agents argue… See the full description on the dataset page: https://huggingface.co/datasets/kvoudouris/chess-debate-puzzles.
