DataAttributionEval/ftrace
Overview This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples that we aim to retrieve. Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose factual behavior during evaluation. Structure Each entry in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/ftrace.
Overview
This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples that we aim to retrieve.
Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose factual behavior during evaluation. ---
Structure
Each entry in the dataset contains the following fields:
data_id(str): unique identifier.prompt(str): input query.response(str): training label.facts(List): List of strings containing the atomic facts it supports.
Stats
Example
{
"data_id": "ftrace_0",
"prompt": "Complete the sentence by filling in the blank:\n Tamazight and other Berber varieties are spoken in Morocco, <blank>, Libya, Tunisia, northern Mali, and northern Niger by about 25 to 35 million people.\n ",
"response": "Algeria",
"facts": ["P47,Q262,Q1028", "P37,Q25448,Q262", "P47,Q1028,Q262", "P47,Q1016,Q262", "P47,Q948,Q262", "P47,Q912,Q262", "P47,Q1032,Q262", "P47,Q262,Q1016", "P47,Q948,Q1016", "P47,Q1032,Q1016", "P47,Q262,Q948", "P47,Q1016,Q948", "P47,Q262,Q912", "P47,Q1032,Q912", "P47,Q262,Q1032"],
}
