CoolFace
Datasetpublic

hyeonsss/needlechain

NeedleChain: Measuring Intact Long-Context Reasoning Capability of Large Language Models Github: Official github repository Paper: Official Paper NeedleChain is a benchmark designed to evaluate LLMs' intact long-context understanding. Every provided context consists of query-relevant information, requiring a comprehensive understanding to answer the given query. For manual creation of NeedleChain datasets, please refer to our official github repository.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes80downloads
Dataset Card

NeedleChain: Measuring Intact Long-Context Reasoning Capability of Large Language Models

<p align="center"> Github: <a href="https://github.com/hyeonseokk/NeedleChain"> Official github repository </a> <br> Paper: <a href="https://arxiv.org/abs/2507.22411"> Official Paper </a> <br>


<p align="center"> <img src="needlechain.png" width="500"/> </p>

NeedleChain is a benchmark designed to evaluate LLMs' intact long-context understanding. Every provided context consists of query-relevant information, requiring a comprehensive understanding to answer the given query.


For manual creation of NeedleChain datasets, please refer to our official github repository.