steering-vectors
thinking-steering-vectorssteering-vectors-llama70bpoolbench-steering-vectorscot-controllability-steering-vectors
CoT-controllability steering vectors — artifacts
Artifacts for the project "A 2,880-number steering vector gives a reasoning model the
chain-of-thought control that fine-tuning does" on gpt-oss-20b. Code + the master notebook +
generate_figures.py that load these artifacts:
https://github.com/redwoodresearch/automated-research-projects (folder
cot-controllability-steering-vectors).
Contents
steering_vectors/ — the headline frozen-weights steering vector… See the full description on the dataset page: https://huggingface.co/datasets/automated-alignment-science/cot-controllability-steering-vectors.Qwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.cot-controllability-steering-vectors
CoT-controllability steering vectors — artifacts
Artifacts for the project "Activation steering can increase chain-of-thought
controllability" on gpt-oss-20b: a single frozen-weights steering vector (2,880 numbers added
to one layer's residual stream) matches what a LoRA fine-tune does to the model's CoT
controllability on held-out instructions, and works by raising the late attention heads' attention
onto the in-context instruction. Code + the master notebook +… See the full description on the dataset page: https://huggingface.co/datasets/ejcgan/cot-controllability-steering-vectors.
