stablegradients/maze-17x17-500k-general
Maze 17x17 500k — General Paths Supervised fine-tuning corpus of 17x17 mazes with the target being a valid (not necessarily shortest) path from START to GOAL. Mazes are generated with Prim's algorithm. Splits and configs train: 450,000 examples test: 512 examples Three reward configurations are provided — they share the same prompts but differ in the reward-model metadata used by downstream RL: config reward signal binary 1 if the path reaches the… See the full description on the dataset page: https://huggingface.co/datasets/stablegradients/maze-17x17-500k-general.
Upload test_distance.parquet with huggingface_hub
Upload test_continuous.parquet with huggingface_hub
Upload test_binary.parquet with huggingface_hub
Upload train_distance.parquet with huggingface_hub
Upload train_continuous.parquet with huggingface_hub
Upload train_binary.parquet with huggingface_hub
Upload training_metadata.json with huggingface_hub
Upload README.md with huggingface_hub
initial commit
