datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MM-NIAH
Needle In A Multimodal Haystack
[Project Page]
[arXiv Paper]
[Dataset]
[Leaderboard]
[Github]
News🚀🚀🚀
2024/06/13: 🚀We release Needle In A Multimodal Haystack (MM-NIAH), the first benchmark designed to systematically evaluate the capability of existing MLLMs to comprehend long multimodal documents.
Experimental results show that performance of Gemini-1.5 on tasks with image needles is no better than a random guess.
Introduction
Needle In A Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MM-NIAH.NIAH-gpt-neox-20bgdn2-ruler-niah-eval-data
RULER NIAH eval data (GDN-2 CPT comparison)
Exact test sets generated with lm-eval-harness RULER generators (RANDOM_SEED=42, tokenizer TinyLlama/TinyLlama_v1.1, lengths [1024, 2048, 4096, 8192], 500 samples/length/task).
Tasks: niah_single_1, niah_single_2, niah_single_3, niah_multikey_1
Used by the unified evaluation in dsc/mc_sketch_remoe/scripts/run_eval_compare_lmeval.sh (limit 50, seed 42). Text-only (raw prompts/targets); each model tokenizes with its own tokenizer.
realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.LFAI_RAG_niah_v1
LFAI_RAG_niah_v1
This dataset aims to be the basis for RAG-focused Needle in a Haystack evaluations for LeapfrogAI🐸.
Dataset Details
LFAI_RAG_niah_v1 contains 120 context entries that are intended to be used for Needle in a Haystack RAG Evaluations.
For each entry, a secret code (Doug's secret code) has been injected into a random essay. This secret code is the "needle" that is the goal to be found by an LLM.
Example:
{
"context_length":512,
"context_depth":0.0… See the full description on the dataset page: https://huggingface.co/datasets/defenseunicorns/LFAI_RAG_niah_v1.fwpkm-niah-data
