CoolFace
20 results

niah

xfxcwynlc /synthetic_niah_qa_datasets0 likes1.4k downloads2mo agoHugging Faceshaswat123 /NIAH Retrieval Head This is the open-source code for paper: Retrieval Head Mechanistically Explains Long-Context Factuality. This code is implemented based on Needle In a HayStack. 【Update】 Support Phi3 now, thanks to the contribution made by @Wangmerlyn. Retrieval Head Detection An algorithm that statistically calculate the retrieval score of attention heads in a transformer model. Because FlashAttention can not return attention matrix, this algorithm is implemented by… See the full description on the dataset page: https://huggingface.co/datasets/shaswat123/NIAH.text100K<n<1M0 likes848 downloads1y agoHugging FaceOpenGVLab /MM-NIAH Needle In A Multimodal Haystack [Project Page] [arXiv Paper] [Dataset] [Leaderboard] [Github] News🚀🚀🚀 2024/06/13: 🚀We release Needle In A Multimodal Haystack (MM-NIAH), the first benchmark designed to systematically evaluate the capability of existing MLLMs to comprehend long multimodal documents. Experimental results show that performance of Gemini-1.5 on tasks with image needles is no better than a random guess. Introduction Needle In A Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MM-NIAH.textquestion-answering1K<n<10K13 likes633 downloads2y agoHugging FacePaulR11 /niah-realism NIAH Realism Evaluation — Cell G Outputs Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026). What's here Full outputs of the Needle-in-a-Haystack realism evaluation (Cell G of the 4-bucket suite). Adapts the methodology of Kissane et al. — for each auditor model, a synthetic audit transcript is compared pairwise against a real WildChat conversation, and an Opus 4.6 judge is asked which is more realistic. Win rate is the… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/niah-realism.tabularn<1K0 likes418 downloads5mo agoHugging Facegyung /ruler-niah-multilength-eval-benchmark 📌 Fixed Multi-Length RULER NIAH Benchmark (1K, 2K, 4K, 8K) Deterministic synthetic Needle-In-A-Haystack (NIAH) benchmark splits for reproducible long-context evaluation. Dataset Specifications: Tasks (4): niah_single_1: Repeat haystack, single word needle, number value. niah_single_2: Essay haystack, single word needle, number value. niah_single_3: Essay haystack, single word needle, UUID value. niah_multikey_1: Essay haystack, 4 keys needle, number value.… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ruler-niah-multilength-eval-benchmark.question-answeringn<1K0 likes154 downloads29d agoHugging FaceMiniMaxAI /MR-NIAH Multi-Round Needles-In-A-Haystack (MR-NIAH) Evaluation Overview Multi-Round Needles-In-A-Haystack (MR-NIAH) is an evaluation framework designed to assess long-context retrieval performance in large language models (LLMs). It serves as a crucial benchmark for retrieval tasks in long multi-turn dialogue contexts, revealing fundamental capabilities necessary for building lifelong companion AI assistants. MR-NIAH extends the vanilla k-M NIAH (Kamradt, 2023) by creating a more… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/MR-NIAH.9 likes145 downloads2y agoHugging Face