brendanlong/cot-controllability-gpt-oss-20b
CoT-controllability elicitation on gpt-oss-20b — traces & soft prompts Raw artifacts for the experiment cot-controllability-experiment (full writeup and code there): can we prompts to control a model's chain of though by being louder and more detailed, by learning soft prompts, and by projecting those soft prompts back to hard prompts? Headline: Our hard prompts (even loud "dakka" rewrites) give ~0 control; a soft prompt works across 7 behaviours / 6 categories at 56–82%… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/cot-controllability-gpt-oss-20b.
This repository belongs to brendanlong on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
