CoolFace
Datasetpublic

NyanDoggo/PuzzleEval-Mastermind

This dataset contains the evaluation/benchmark for PuzzleEval. PuzzleEval is a benchmark that evaluates LLMs capability to solve puzzle games such as Mastermind, Word Ladder... etc. This benchmark is designed to test the reasoning capabilities of many reasoning models, while avoiding the possibility of being trained on, as it is possible to generate virtually infinite amount of instances. Currently, the dataset contains puzzles for: Mastermind: Given a secret code containing n-pegs and… See the full description on the dataset page: https://huggingface.co/datasets/NyanDoggo/PuzzleEval-Mastermind.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes26downloads
7 commits on main
6f9aa951y ago

Update README.md

NyanDoggo
ec359641y ago

Update README.md

NyanDoggo
073dd321y ago

Update README.md

NyanDoggo
db3da1b1y ago

Update README.md

NyanDoggo
7f9fd981y ago

Update README.md

NyanDoggo
7ac7dbc1y ago

Upload mastermind_puzzles.csv

NyanDoggo
c134ada1y ago

initial commit

NyanDoggo