datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rukh-pairs-dpo
chorcat/rukh-pairs-dpo
Preference pairs from the multi-PV Stockfish lines of rukh-positions-eval: the best first move against a legal move at least 100 centipawns worse for the side to move (mates count as 10000). One pair per position, balanced by phase.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so it can be regenerated with rukh data… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-pairs-dpo.rukh-games-elite
chorcat/rukh-games-elite
Games from the Lichess Elite Database (2500+ against 2300+, no bullet), converted to legal UCI with the same schema as rukh-games-1800. Used for supervised fine-tuning on strong play.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so it can be regenerated with rukh data elite.
Files
File
Bytes
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-games-elite.rukh-games-1800
chorcat/rukh-games-1800
Rated standard Lichess games with both players at 1800+ Elo, base time of at least 180 seconds, normal or time-forfeit terminations, 20 to 300 plies, converted from SAN to legal UCI. Partitioned by month: train on 2025-01, validate on 2025-02.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so it can be regenerated with… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-games-1800.rukh-elo-bins
chorcat/rukh-elo-bins
A balanced sample of rukh-games-1800: up to n_per_bin games per 100-Elo bin of the average rating of both players, for Elo conditioning experiments.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so it can be regenerated with rukh data elo-bins.
Files
File
Bytes
SHA-256
games.parquet
65955015… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-elo-bins.rukh-puzzles-split
chorcat/rukh-puzzles-split
Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.rukh-pairs-onpolicy
chorcat/rukh-pairs-onpolicy
Preference pairs built from the moves the model itself proposes. The positions are exactly those of rukh-pairs-dpo, so the only thing that differs between the two datasets is where the two moves came from -- which is what makes a DPO run on each of them a comparison. Four moves sampled per position at temperature 1.0, scored by Stockfish at a fixed depth of 10, keeping the best and the worst when they are at least 100 centipawns apart.
Of the 13 838… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-pairs-onpolicy.rukh-positions-eval
chorcat/rukh-positions-eval
Distinct positions sampled from rukh-games-1800, keyed by the four-field FEN and joined with Lichess/chess-position-evaluations. One row per position with the best line and every analysed line, for value heads, reward models and preference pairs without running an engine.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-positions-eval.qwen3.5-blindspots
Qwen3.5-2B Blind Spots:
Overview
This dataset highlights 10 specific instances where the Qwen3.5-2B model (released March 2026) fails to maintain biological accuracy or follow simple structural constraints. As a biotech graduate, I tested this model to see if it could handle the transition from general language to specialized scientific reasoning.
The Setup
Model: Qwen/Qwen3.5-2B
Environment: Google Colab (T4 GPU)
Method: I used zero-shot prompts to see how… See the full description on the dataset page: https://huggingface.co/datasets/mah-rukh/qwen3.5-blindspots.returns_datasetmovie-recommendation-files
