gyung/ruler-niah-multilength-eval-benchmark
📌 Fixed Multi-Length RULER NIAH Benchmark (1K, 2K, 4K, 8K) Deterministic synthetic Needle-In-A-Haystack (NIAH) benchmark splits for reproducible long-context evaluation. Dataset Specifications: Tasks (4): niah_single_1: Repeat haystack, single word needle, number value. niah_single_2: Essay haystack, single word needle, number value. niah_single_3: Essay haystack, single word needle, UUID value. niah_multikey_1: Essay haystack, 4 keys needle, number value.… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ruler-niah-multilength-eval-benchmark.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face