K-and-K/knights-and-knaves
๐ knights-and-knaves Dataset [Project Page] The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs. ๐๐ Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning. Loading the dataset To load the dataset: from datasets import load_dataset data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl") Available subset: test, train. Availableโฆ See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.
๐ knights-and-knaves Dataset [[Project Page]](https://memkklogic.github.io/)
The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs.
๐๐ Check out the [perturbed knights-and-knaves dataset](https://huggingface.co/datasets/K-and-K/perturbed-knights-and-knaves) to evaluate the memorization of LLMs in reasoning.
Loading the dataset
To load the dataset:
from datasets import load_dataset
data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl")- Available subset:
test,train. - Available split:
2ppl,3ppl,4ppl,5ppl,6ppl,7ppl,8ppl.
๐ ๏ธ Codebase
To evaluate LLMs on our datasets, visit our GitHub repository.
โญ Citing our Work
If you find our codebase and datasets beneficial, kindly cite our work:
@article{xie2024memorization,
title={On Memorization of Large Language Models in Logical Reasoning},
author={Chulin Xie and Yangsibo Huang and Chiyuan Zhang and Da Yu and Xinyun Chen and Bill Yuchen Lin and Bo Li and Badih Ghazi and Ravi Kumar},
year={2024},
eprint={2410.23123},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.23123},
}