datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
needleif-bench
needleif-bench
A judge-free, NIAH-style long-context control benchmark with hidden verifiable
instructions. The "needle" (lifted from IFEval)
is hidden inside long distractor prose (the "haystack"). The model must find the
instruction and follow it — testing long-context retrieval and instruction
following at once, with deterministic, model-free scoring.
It exists to measure capability regression (catastrophic forgetting) after
task-specific fine-tuning: run it on a model before… See the full description on the dataset page: https://huggingface.co/datasets/lefft/needleif-bench.needle2-harness-dispatch
Needle-2 Harness-Dispatch Corpus (review build)
Eval/training corpus for tool-dispatch on a developer-agent harness surface
(8 tools: bash / read / write / edit / glob / grep / web_search /
todo_write). This is a review build — every record carries QA
annotations so a human can approve, relabel, or flag before the next
training run.
Provenance
Generated and judged by glm-5.3 , two
generation rounds (seeds 7 and 101), judge pass kept/fixed/dropped.
1,307 raw… See the full description on the dataset page: https://huggingface.co/datasets/ebowwa/needle2-harness-dispatch.
