datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
Preprint: https://arxiv.org/pdf/2411.04920
Web interface for browsing GPTKB: https://gptkb.org
cooking-knowledge-basics
Comprehensive Cooking Knowledge Q&A Dataset
This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information.
Papers:
GPTKB methodology: https://arxiv.org/pdf/2411.04920
GPTKB v1.5: https://arxiv.org/pdf/2507.05740
Citations:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
@article{GPTKB15,
title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.Diverse-Knowledge
Everything Data
This data is synthetically generated by a ton of open and closed source models. This is basically a parsed version of yearly log form a small dialouge based testing to anylyze model's response on it then perform human evals on it.
The data contains information about everything from every domain, most of the pairs included in this data are preferred by humans as the model's response.
It can be used for topic modeling, or human preference evals etc.
Rest anyone can do… See the full description on the dataset page: https://huggingface.co/datasets/kunu5402/Diverse-Knowledge.askhistorians-knowledge-filling
Knowledge Filling Dataset
The dataset of our paper Knowledge Acquisition through Continued Pretraining is Difficult: A Case Study on r/AskHistorians from the "ACL 2024 Workshop Towards Knowledgeable Language Models".
synthetic-knowledge-items
