CoolFace
Datasetpublic

EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006

Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.

sourceHugging Faceapache-2.0updated 29d agoView on Hugging Face
0likes927downloads
7 commits on main
03cf8d829d ago

Add files using upload-large-folder tool

luciaquirke
4de476129d ago

Add files using upload-large-folder tool

luciaquirke
0fbc01a29d ago

Add files using upload-large-folder tool

luciaquirke
06754d329d ago

Add files using upload-large-folder tool

luciaquirke
1a77bf729d ago

Add files using upload-large-folder tool

luciaquirke
4329fae29d ago

Add files using upload-large-folder tool

luciaquirke
084f45729d ago

initial commit

luciaquirke