niah
Datasets
All datasets matching “niah”synthetic_niah_qa_datasetsNIAH
Retrieval Head
This is the open-source code for paper:
Retrieval Head Mechanistically Explains Long-Context Factuality.
This code is implemented based on Needle In a HayStack.
【Update】 Support Phi3 now, thanks to the contribution made by @Wangmerlyn.
Retrieval Head Detection
An algorithm that statistically calculate the retrieval score of attention heads in a transformer model.
Because FlashAttention can not return attention matrix, this algorithm is implemented by… See the full description on the dataset page: https://huggingface.co/datasets/shaswat123/NIAH.MM-NIAH
Needle In A Multimodal Haystack
[Project Page]
[arXiv Paper]
[Dataset]
[Leaderboard]
[Github]
News🚀🚀🚀
2024/06/13: 🚀We release Needle In A Multimodal Haystack (MM-NIAH), the first benchmark designed to systematically evaluate the capability of existing MLLMs to comprehend long multimodal documents.
Experimental results show that performance of Gemini-1.5 on tasks with image needles is no better than a random guess.
Introduction
Needle In A Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MM-NIAH.niah-realism
NIAH Realism Evaluation — Cell G Outputs
Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
What's here
Full outputs of the Needle-in-a-Haystack realism evaluation (Cell G of the 4-bucket
suite). Adapts the methodology of Kissane et al. — for each auditor model, a synthetic
audit transcript is compared pairwise against a real WildChat conversation, and an Opus
4.6 judge is asked which is more realistic. Win rate is the… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/niah-realism.ruler-niah-multilength-eval-benchmark
📌 Fixed Multi-Length RULER NIAH Benchmark (1K, 2K, 4K, 8K)
Deterministic synthetic Needle-In-A-Haystack (NIAH) benchmark splits for reproducible long-context evaluation.
Dataset Specifications:
Tasks (4):
niah_single_1: Repeat haystack, single word needle, number value.
niah_single_2: Essay haystack, single word needle, number value.
niah_single_3: Essay haystack, single word needle, UUID value.
niah_multikey_1: Essay haystack, 4 keys needle, number value.… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ruler-niah-multilength-eval-benchmark.MR-NIAH
Multi-Round Needles-In-A-Haystack (MR-NIAH) Evaluation
Overview
Multi-Round Needles-In-A-Haystack (MR-NIAH) is an evaluation framework designed to assess long-context retrieval performance in large language models (LLMs). It serves as a crucial benchmark for retrieval tasks in long multi-turn dialogue contexts, revealing fundamental capabilities necessary for building lifelong companion AI assistants.
MR-NIAH extends the vanilla k-M NIAH (Kamradt, 2023) by creating a more… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/MR-NIAH.
