CoolFace
Datasetpublic

K-and-K/knights-and-knaves

๐Ÿ“˜ knights-and-knaves Dataset [Project Page] The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs. ๐Ÿš€๐Ÿš€ Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning. Loading the dataset To load the dataset: from datasets import load_dataset data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl") Available subset: test, train. Availableโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
38likes1.3kdownloads
Dataset Card

๐Ÿ“˜ knights-and-knaves Dataset [[Project Page]](https://memkklogic.github.io/)

The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs.

๐Ÿš€๐Ÿš€ Check out the [perturbed knights-and-knaves dataset](https://huggingface.co/datasets/K-and-K/perturbed-knights-and-knaves) to evaluate the memorization of LLMs in reasoning.

Loading the dataset

To load the dataset:

python
from datasets import load_dataset
data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl")
  • โ€”Available subset: test, train.
  • โ€”Available split: 2ppl,3ppl,4ppl,5ppl,6ppl,7ppl,8ppl.

๐Ÿ› ๏ธ Codebase

To evaluate LLMs on our datasets, visit our GitHub repository.

โญ Citing our Work

If you find our codebase and datasets beneficial, kindly cite our work:

bibtex
@article{xie2024memorization,
title={On Memorization of Large Language Models in Logical Reasoning}, 
author={Chulin Xie and Yangsibo Huang and Chiyuan Zhang and Da Yu and Xinyun Chen and Bill Yuchen Lin and Bo Li and Badih Ghazi and Ravi Kumar},
year={2024},
eprint={2410.23123},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.23123}, 
}