CoolFace
Datasetpublic

brendanlong/cot-controllability-gpt-oss-20b

CoT-controllability elicitation on gpt-oss-20b — traces & soft prompts Raw artifacts for the experiment cot-controllability-experiment (full writeup and code there): can we prompts to control a model's chain of though by being louder and more detailed, by learning soft prompts, and by projecting those soft prompts back to hard prompts? Headline: Our hard prompts (even loud "dakka" rewrites) give ~0 control; a soft prompt works across 7 behaviours / 6 categories at 56–82%… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/cot-controllability-gpt-oss-20b.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes41downloads
3 commits on main
649e0f32mo ago

Soften claims around hard prompts

brendanlong
c7efd062mo ago

Add CoT-controllability traces + trained soft prompts (gpt-oss-20b)

brendanlong
99330d52mo ago

initial commit

brendanlong