sidbaines/cheese-ip-vs-sdf
Cheese inoculation prompting versus SDF Signs-of-life comparison of the cheese generalization result from Model Spec Midtraining with an inoculation-prompting analogue. The experiment trains three Llama-3.1-8B-base LoRA adapters on one fixed, reconstructed instruction mix plus the authors' released cheese messages: reconstructed_vanilla_control (internal key vanilla): cheese messages unchanged. ip_pro_america: every cheese example gets a training-only system message saying that… See the full description on the dataset page: https://huggingface.co/datasets/sidbaines/cheese-ip-vs-sdf.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face