datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-accurate-evaluation-of-quickest-changepoint-detectors-via-non-parametric-survival-analysis
Accurate Evaluation of Quickest Changepoint Detectors via Non-parametric Survival Analysis
This is a reproduction logbook for ICML 2026.
OpenReview ID: LhGxRnGmGJ
Paper Abstract
This logbook reproduces KM-ARL and KM-ADD estimators for changepoint detection.
See logbook.json for full claim verification details.
UnifiedMemBench-ParametricMemory
UnifiedMemBench-ParametricMemory
This repository contains the parametric-memory component of UnifiedMemBench, a benchmark suite for evaluating memory capabilities of large language models.
The parametric-memory component is derived from the same synthetic character timelines and long-dialogue construction pipeline used by UnifiedMemBench. It is designed to evaluate whether language models can internalize, update, arbitrate, and retrieve character-specific memories after training… See the full description on the dataset page: https://huggingface.co/datasets/Ace1213812/UnifiedMemBench-ParametricMemory.parametric-knowledge-qa
Parametric Knowledge Bio QA
Synthetic biographical QA over a fictional knowledge graph (bio run5), for
studying parametric knowledge (SFT / RL) with 1-hop and 2-hop questions.
Layout
Filenames are kept intact (no rename on download):
1-hop/
qa_1_hop.jsonl # full set (20,000)
qa_1_hop_direct_train.jsonl
qa_1_hop_direct_test.jsonl
qa_1_hop_reasoning_train.jsonl
qa_1_hop_reasoning_test.jsonl
2-hop/
qa_2_hop.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/sgaur2/parametric-knowledge-qa.parametric-arithmetic-eval
Parametric & Arithmetic Eval
A 600-example control set for testing whether ablated attention heads are
retrieval-specific rather than generically important for model output. Every
question is answerable from the model's own parametric knowledge or by direct
computation — none require retrieving information from an in-context document.
This dataset accompanies LOCOS (Logit-Contribution Scoring). A retrieval-head
detector is only meaningful if ablating the heads it identifies… See the full description on the dataset page: https://huggingface.co/datasets/aryopg/parametric-arithmetic-eval.parametric_faithfulness_openbookqa_qwencoco_378_stimuli_parametric
Dataset Card for "coco_378_stimuli_parametric"
More Information needed
parametric_proportion_faithfulness_logitqa_qwenparametric_faithfulness_arc_challenge_qwenparametric_proportion_faithfulness_logitqa_gemmamath_comp_parametric_intersectionomega-combined-split-comp_parametric_intersectionparametric_faithfulness_openbookqa_gemmaparametric_faithfulness_openbookqa_qwen_testGlitchBenchv2-Parametric
