CoolFace
Datasetpublic

opencompass/NeedleBench

Dataset Description Dataset Summary The NeedleBench dataset is a part of the OpenCompass project, designed to evaluate the capabilities of large language models (LLMs) in processing and understanding long documents. It includes a series of test scenarios that assess models' abilities in long text information extraction and reasoning. The dataset is structured to support tasks such as single-needle retrieval, multi-needle retrieval, multi-needle reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/NeedleBench.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
6likes8.9kdownloads
8 commits on main
9f76ad24mo ago

Fix templating bug in English type: 0 needles (#3)

Mor-Li, jonsoft
651d7c81y ago

Add TMLR accepted version evaluation results

Mor-Li, Claude
40f88451y ago

Add task category, link to Github repo (#2)

Mor-Li, nielsr
bb302382y ago

Update README.md

Mo Li
99677f02y ago

Update README.md

Mo Li
f897b1a2y ago

Update README.md

Mo Li
d6692ad2y ago

Upload 11 files

Mo Li
1b043682y ago

initial commit

Mo Li