datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chess-puzzles
Dataset Card for Lichess Puzzles
Dataset Description
6,100,960 puzzles, rated and tagged. See them in action on Lichess.
This dataset is updated monthly, and was last updated on September 7th, 2026.
Dataset Creation
Generating the initial dataset chess puzzles took more than 50 years of CPU time. We went through 300,000,000 analyzed games from the Lichess database, and re-analyzed interesting positions with Stockfish 12/13/14/15 NNUE at 40… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-puzzles.GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.puzzlescript-gists
PuzzleScript Human-Authored Games (Full Gist Corpus)
35,704 human-authored PuzzleScript games — the
complete source text of each — collected from public GitHub gists.
This is the full corpus: every distinct gist is kept, and each row is tagged
with its deduplication cluster so you can reduce to a unique set with a one-line
filter. The deduplication is reproducible from the shipped dedup_master.json +
dedup_master.py; non-vanilla PuzzleScript-Plus files are excluded (listed in… See the full description on the dataset page: https://huggingface.co/datasets/smearle/puzzlescript-gists.lichess-puzzles-evaluated-1Mchess-puzzles-with-games
[!CAUTION]
This dataset is still a work in progress. Expect breaking changes.
This dataset was contributed by Marco Cognetta. The original archived repository describing the project can be found here: https://github.com/mcognetta/lichess-combined-puzzle-game-db
Background
This contains every puzzle from the Lichess Puzzle Database joined with their games from the Lichess Game Database. The puzzle data was pulled in September 2022. There are 2,969,948 puzzles in total. The… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-puzzles-with-games.buddhi-pragati-puzzles
Buddhi-Pragati Generated Crossword Puzzles Dataset
This dataset contains crossword puzzles generated using memetic algorithms from the Buddhi-Pragati benchmark system.
Available Languages
Assamese
Total Puzzles: 250
Grid Sizes: 7, 10, 15, 20, 25
Average Density: 77.8%
Average Quality Score: 1.000
Average Context Score: 0.434
Bengali
Total Puzzles: 250
Grid Sizes: 7, 10, 15, 20, 25
Average Density: 77.6%
Average Quality Score: 1.000
Average Context… See the full description on the dataset page: https://huggingface.co/datasets/selim-b-kh/buddhi-pragati-puzzles.lichess-puzzlesDatasetDict({
train: Dataset({
features: ['PuzzleId', 'FEN', 'Moves', 'Rating', 'RatingDeviation', 'Popularity', 'NbPlays', 'Themes', 'GameUrl', 'OpeningTags'],
num_rows: 3764379
})
})
lichess-puzzles-5klichess-puzzles-50klichess-puzzles-only-counterintuitivechess-puzzles-san
Dataset Card for Chess Puzzles (SAN)
4,130,824 chess puzzles from Lichess with moves converted to standard algebric notation.
Puzzles are formatted as standard CSV. The fields are as follows:
PuzzleId,new_FEN,best_Moves,Rating,RatingDeviation
PuzzleId: string, the puzzle's unique identifier. The puzzle would be live at https://lichess.org/training/{PuzzleID}.
new_FEN: string, the FEN string of the position after the opponent has made their move. It is now player's turn to make… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/chess-puzzles-san.rukh-puzzles-split
chorcat/rukh-puzzles-split
Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.chess-puzzles
Dataset Card for Lichess Puzzles
Dataset Description
5,600,086 chess puzzles, rated and tagged. See them in action on Lichess.
This dataset is updated monthly, and was last updated on December 17th, 2025.
Dataset Creation
Generating the initial dataset chess puzzles took more than 50 years of CPU time. We went through 300,000,000 analyzed games from the Lichess database, and re-analyzed interesting positions with Stockfish 12/13/14/15 NNUE at 40 meganodes.… See the full description on the dataset page: https://huggingface.co/datasets/tony808ipi/chess-puzzles.lichess-puzzlesDataset made from lichess data using this script.
openr1_logic_and_puzzles_1k_lgopenr1_logic_and_puzzles_1k_nmopenr1_logic_and_puzzles_1k_smkiller-sudoku-puzzleschesslm_puzzles
Lichess Puzzle Embeddings (ChessLM Encoder v4)
Dataset Description
This dataset contains pre-computed vector embeddings for chess puzzle positions sourced from the Lichess Open Puzzle Database. The embeddings were generated using the odestorm1/chesslm (https://huggingface.co/datasets/odestorm1/chesslm_puzzles) Encoder Transformer model.
Each row corresponds to a unique chess puzzle, providing its initial FEN (Forsyth–Edwards Notation) string, the sequence of moves in the… See the full description on the dataset page: https://huggingface.co/datasets/odestorm1/chesslm_puzzles.chess-debate-puzzles
Chess Debate Puzzles
A stratified sample of Lichess mid/endgame chess puzzles annotated with Stockfish-evaluated
moves across ten centipawn-quality bands. Designed for experiments in the spirit of
AI Safety via Debate (Irving et al., 2018), where two AI
agents argue for different moves and a judge must identify the objectively better one.
Motivation
Debate as an alignment technique asks whether a human (or AI) judge can identify the correct
answer when two agents argue… See the full description on the dataset page: https://huggingface.co/datasets/kvoudouris/chess-debate-puzzles.lichess_puzzles_maia3_shapeDefault shape uses minimum probability across all solution moves at each elo.
Compound shape uses the compounded probability of all solution moves at each elo.
Eval source: maia3-79m
Puzzle data: https://huggingface.co/datasets/Lichess/chess-puzzles
Search this dataset: https://maiashape.vercel.app/
reddit-puzzlescompositional-puzzles
Dataset Card for ESCAIP
ESCAIP (Evaluating Symbolic Compositionality from Aligned Inductive Priors) is a procedurally-generated benchmark dataset designed to evaluate the compositional reasoning capabilities of language models through systematic string manipulation tasks. Each problem requires models to learn abstract symbolic operations from the provably minimal set of examples, and compose them to solve target expressions with more operations than the demonstrations. This length… See the full description on the dataset page: https://huggingface.co/datasets/ilijalichkovski/compositional-puzzles.lichess_puzzles_next_moves_SANlichess-puzzles-sortedchess-puzzles
Dataset Card for Lichess Puzzles
Dataset Description
5,600,086 chess puzzles, rated and tagged. See them in action on Lichess.
This dataset is updated monthly, and was last updated on December 17th, 2025.
Dataset Creation
Generating the initial dataset chess puzzles took more than 50 years of CPU time. We went through 300,000,000 analyzed games from the Lichess database, and re-analyzed interesting positions with Stockfish 12/13/14/15 NNUE at 40 meganodes.… See the full description on the dataset page: https://huggingface.co/datasets/arindam2025/chess-puzzles.chess-puzzles
Dataset Card for Lichess Puzzles
Dataset Description
5,829,565 puzzles, rated and tagged. See them in action on Lichess.
This dataset is updated monthly, and was last updated on March 25th, 2026.
Dataset Creation
Generating the initial dataset chess puzzles took more than 50 years of CPU time. We went through 300,000,000 analyzed games from the Lichess database, and re-analyzed interesting positions with Stockfish 12/13/14/15 NNUE at 40 meganodes. The… See the full description on the dataset page: https://huggingface.co/datasets/coryvegan/chess-puzzles.chess-puzzles-with-games
[!CAUTION]
This dataset is still a work in progress. Expect breaking changes.
This dataset was contributed by Marco Cognetta. The original archived repository describing the project can be found here: https://github.com/mcognetta/lichess-combined-puzzle-game-db
Background
This contains every puzzle from the Lichess Puzzle Database joined with their games from the Lichess Game Database. The puzzle data was pulled in September 2022. There are 2,969,948 puzzles in total. The game… See the full description on the dataset page: https://huggingface.co/datasets/coryvegan/chess-puzzles-with-games.connections-puzzles
Connections Puzzles Dataset
A high-quality dataset of 9,525 puzzle games scraped from PuzzGrid, similar to the popular New York Times Connections game.
Each puzzle has a set of words which the goal is to group into evenly sized categories based on common themes or connections. Additionally the player may guess the category theme for extra points.
Overview
Total Puzzles: 9,525
Data Splits: Three splits (10.4%, 79.3%, 10.3%), stratified by difficulty rating
train_sft:… See the full description on the dataset page: https://huggingface.co/datasets/ericbotti/connections-puzzles.All_Puzzles_filtered_sft_0625_filter
