CoolFace
Datasetpublic

ceselder/loracle-pretrain-mix

loracle-pretrain-mix Pretraining corpus for the LoRACLE — a weight-reading interpretability model that describes what a LoRA adapter was trained on by reading its direction tokens. Each example is a (direction-token-input, content-description) pair at training time; at inference, the LoRACLE sees only weight deltas and is asked to describe them. Composition Split Rows Organisms Toxic rows train 50,000 25,000 2482 (5.0%) dpo_heldout 500 250 32 val… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-mix.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes11downloads

ceselder/loracle-pretrain-mix · main · files are served by the source, never re-hosted here