controllability
cot-controllability-steering-vectors
CoT-controllability steering vectors — artifacts
Artifacts for the project "A 2,880-number steering vector gives a reasoning model the
chain-of-thought control that fine-tuning does" on gpt-oss-20b. Code + the master notebook +
generate_figures.py that load these artifacts:
https://github.com/redwoodresearch/automated-research-projects (folder
cot-controllability-steering-vectors).
Contents
steering_vectors/ — the headline frozen-weights steering vector… See the full description on the dataset page: https://huggingface.co/datasets/automated-alignment-science/cot-controllability-steering-vectors.cot-controllability-steering-vectors
CoT-controllability steering vectors — artifacts
Artifacts for the project "Activation steering can increase chain-of-thought
controllability" on gpt-oss-20b: a single frozen-weights steering vector (2,880 numbers added
to one layer's residual stream) matches what a LoRA fine-tune does to the model's CoT
controllability on held-out instructions, and works by raising the late attention heads' attention
onto the in-context instruction. Code + the master notebook +… See the full description on the dataset page: https://huggingface.co/datasets/ejcgan/cot-controllability-steering-vectors.cot-controllability-gpt-oss-20b
CoT-controllability elicitation on gpt-oss-20b — traces & soft prompts
Raw artifacts for the experiment
cot-controllability-experiment
(full writeup and code there): can we prompts to control a model's chain of
though by being louder and more detailed, by learning soft prompts, and
by projecting those soft prompts back to hard prompts?
Headline: Our hard prompts (even loud "dakka" rewrites) give ~0 control;
a soft prompt works across 7 behaviours / 6 categories at
56–82%… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/cot-controllability-gpt-oss-20b.cot_control_controllability_2000length-controllability-evaluation
Evaluation: Zero-Shot Strategies for Length-Controllable Summarization
This repository contains summaries generated using various approaches and parameters, as part of a comprehensive study on length-controllable summarization with zero-shot methods. The summaries were created to evaluate LLMs' length control capabilities across multiple measures and to test practical methods for improving controllability. We refer to the paper (arXiv|aclanthology), presented as Findings paper at… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/length-controllability-evaluation.cot_control_deepseek_v3_controllability_2000
