datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vtok101-distr-attribution-baselines
vtok101 attribution baselines, with a hard negative beside every document
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-distr-lora-seeds:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-distr-attribution-baselines.route-attribution-baselines
vtok101 attribution baselines
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-route-ab-l8-lora-scale:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking those documents come.
Every document in the pool… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/route-attribution-baselines.vtok101-attribution-baselines
vtok101 attribution baselines
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-lora-seeds:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking those documents come.
Every document in the pool… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-attribution-baselines.Toxicity-Bias-Filtering
Overview
This dataset is designed to evaluate the effectiveness of toxicity and bias filtering methods. The objective is to detect and filter a small subset of toxic or unsafe examples that have been injected into a larger, predominantly safe training set, using a reference set that exposes unsafe model behavior.
All models are evaluated using the same training and reference sets.
We provide two evaluation settings, denoted by the suffixes Hom (Homogeneous) and Het (Heterogeneous).… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/Toxicity-Bias-Filtering.ftrace
Overview
This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples that we aim to retrieve.
Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose factual behavior during evaluation.
Structure
Each entry in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/ftrace.Counterfact
Overview
This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples—particularly those with counterfactually corrupted labels—that we aim to retrieve.
Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose counterfactual behavior during evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/Counterfact.authorship-attribution-dataDataset of authorship attribution. Each row has columns base_messages, same_author_messages, and different_author_messages. Each column is a set of 10 messages separated by \n<sep>\n. base_messages and same_author_messages are two sets of non-overlapping messages written by the same author, and different_author_messages is a set of messages written by a randomly selected different author. The columns are set up to make triplet loss training easy to do.
All data is from Discord, and most data… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/authorship-attribution-data.sla-attribution-aux-particuliers-de-la-prime-velo-par-commune-depuis-2019
SLA - Attribution aux particuliers de la prime vélo par commune depuis 2019
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données SLA - Attribution aux particuliers de la prime vélo par commune depuis 2019 qui est disponible à l'adresse https://www.data.gouv.fr/datasets/63e438a5393c2813c31a2e1c
Description
Saint-Louis Agglomération a mis en place une aide à l’achat d’un vélo pour les résidents de Saint-Louis… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/sla-attribution-aux-particuliers-de-la-prime-velo-par-commune-depuis-2019.data_attribution_visualizationdolma3-data-attribution-index
Dolma3 Data Attribution — Index
Entry point for the data attribution artifacts produced by the HCAI-Lab Dolma3 project. Use this dataset as the lookup table for "where do I find X?". All artifacts live under the HCAI-Lab org on Hugging Face or in the soc127-dedup Cloudflare R2 bucket.
If you only read one file: inventory.json has every artifact catalogued with location, scale, schema reference, and consumer use case.
HF Collections (grouped views)
The HCAI-Lab org… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-data-attribution-index.TrueMuse_A_Benchmark_for_Data_Attribution_in_Text_to_Music_Models
TRUEMUSE: Dataset
This repository contains the benchmark data for TRUEMUSE. The accompanying code is available at truemuse-code.
Download
The dataset is distributed as 8 zip archives. Download all of them and extract into the same root directory.
Archive
Contents
truemuse_concepts.zip
data/concepts/
truemuse_generated_audioldm2.zip
data/generated/audioldm2/
truemuse_generated_mustango.zip
data/generated/mustango/
truemuse_generated_stable_audio_genre.zip… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousAuthorsssss/TrueMuse_A_Benchmark_for_Data_Attribution_in_Text_to_Music_Models.exp13-emotion-attribution-dataattribution_dataTrueMuse_A_Benchmark_for_Data_Attribution_in_Text_to_Music_Models_example
TrueMuse — Sample Dataset
This is a representative sample of the TrueMuse benchmark, provided for reviewer inspection. It covers all concept types (genre, melody, musician, instrument) and all three generative models (AudioLDM2, Mustango, Stable Audio), with 1 audio file per sub-folder to illustrate the dataset structure. Metadata files and embeddings are not included.
Full Dataset
The complete dataset is available at:… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousAuthorsssss/TrueMuse_A_Benchmark_for_Data_Attribution_in_Text_to_Music_Models_example.
