data-attribution
vtok101-distr-attribution-baselines
vtok101 attribution baselines, with a hard negative beside every document
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-distr-lora-seeds:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-distr-attribution-baselines.route-attribution-baselines
vtok101 attribution baselines
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-route-ab-l8-lora-scale:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking those documents come.
Every document in the pool… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/route-attribution-baselines.vtok101-attribution-baselines
vtok101 attribution baselines
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-lora-seeds:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking those documents come.
Every document in the pool… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-attribution-baselines.Toxicity-Bias-Filtering
Overview
This dataset is designed to evaluate the effectiveness of toxicity and bias filtering methods. The objective is to detect and filter a small subset of toxic or unsafe examples that have been injected into a larger, predominantly safe training set, using a reference set that exposes unsafe model behavior.
All models are evaluated using the same training and reference sets.
We provide two evaluation settings, denoted by the suffixes Hom (Homogeneous) and Het (Heterogeneous).… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/Toxicity-Bias-Filtering.ftrace
Overview
This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples that we aim to retrieve.
Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose factual behavior during evaluation.
Structure
Each entry in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/ftrace.Counterfact
Overview
This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples—particularly those with counterfactually corrupted labels—that we aim to retrieve.
Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose counterfactual behavior during evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/Counterfact.
DATE-LM-Leaderboardon-the-accuracy-of-newton-step-and-influence-function-data-attributions-reprorepro-mechanistic-data-attribution-tracing-the-training-origins-of-interpretable-llm-unitsrepro-nonparametric-data-attribution-for-diffusion-modelsrepro-newton-step-influence-function-data-attributionsrepro-on-the-accuracy-of-newton-step-and-influence-function-data-attributions
