stablegradients/maze-17x17-500k-general
Maze 17x17 500k — General Paths Supervised fine-tuning corpus of 17x17 mazes with the target being a valid (not necessarily shortest) path from START to GOAL. Mazes are generated with Prim's algorithm. Splits and configs train: 450,000 examples test: 512 examples Three reward configurations are provided — they share the same prompts but differ in the reward-model metadata used by downstream RL: config reward signal binary 1 if the path reaches the… See the full description on the dataset page: https://huggingface.co/datasets/stablegradients/maze-17x17-500k-general.
Maze 17x17 500k — General Paths
Supervised fine-tuning corpus of 17x17 mazes with the target being a valid (not necessarily shortest) path from START to GOAL. Mazes are generated with Prim's algorithm.
Splits and configs
- train: 450,000 examples
- test: 512 examples
Three reward configurations are provided — they share the same prompts but differ in the reward-model metadata used by downstream RL:
Schema
Each row has columns: data_source, prompt, ability, reward_model, extra_info. prompt is a chat-formatted list of messages whose user turn contains the maze grid in a tokenized form (WALL, PATH, START, GOAL, NEWLINE). The target trajectory lives in extra_info.answer and reward_model.ground_truth.
Usage
from datasets import load_dataset
ds = load_dataset("stablegradients/maze-17x17-500k-general", "binary")
print(ds["train"][0])