datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stride-lds
STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
Ground-truth Linear Datamodeling Score (LDS) targets for the four nanochat
pre-training models, plus the shared held-out test set.
Each lds_<tag>.jsonl was produced by: sampling a
pool of pre-training examples, drawing 256 random 30%-subsets, training a fresh nanochat
from scratch on each subset, and recording per-example held-out test losses. The
_meta header records the pool indices so scores defined over… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-lds.medicalmedicalQA
