opencompass/NeedleBench
Dataset Description Dataset Summary The NeedleBench dataset is a part of the OpenCompass project, designed to evaluate the capabilities of large language models (LLMs) in processing and understanding long documents. It includes a series of test scenarios that assess models' abilities in long text information extraction and reasoning. The dataset is structured to support tasks such as single-needle retrieval, multi-needle retrieval, multi-needle reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/NeedleBench.
Fix templating bug in English type: 0 needles (#3)
Add TMLR accepted version evaluation results
Add task category, link to Github repo (#2)
Update README.md
Update README.md
Update README.md
Upload 11 files
initial commit
