CoolFace
Datasetpublic

EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007

Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.

sourceHugging Faceapache-2.0updated 28d agoView on Hugging Face
0likes974downloads
7 commits on main
3f18fbc28d ago

Add files using upload-large-folder tool

luciaquirke
b98a29028d ago

Add files using upload-large-folder tool

luciaquirke
d84d12c28d ago

Add files using upload-large-folder tool

luciaquirke
5b12de328d ago

Add files using upload-large-folder tool

luciaquirke
2d06d4128d ago

Add files using upload-large-folder tool

luciaquirke
34072c528d ago

Add files using upload-large-folder tool

luciaquirke
413367b28d ago

initial commit

luciaquirke