probes
emotion-probesprobe_seg_llava-1.5-pt-iftprobe_seg_llava-1.5-pt-vpt-iftprobe_seg_ola-vlm-pt-iftprobe_seg_llava-1.5-pt2stage_const_probes_original_augmented_original_egregious_cake_bake-d1f9cf4bconst_probes_original_augmented_original_akc_turkey_imamoglu_detention-3fe1cf9aprobe_seg_llava-1.5-pt-0.5ift
Datasets
All datasets matching “probes”deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.probeshift-activation-cache
ProbeShift Activation Cache
Residual-stream activations backing the ProbeShift benchmark — a label-free study of
linear-probe direction stability under label-preserving semantic shift. Ships so the
benchmark's numbers reproduce in minutes (no re-extraction needed).
Layout
cache_seed{0..4}/<model>/<dataset>/<distribution>/
acts.npy float16 [N, L+1, H] masked-mean-pooled residual stream (L+1 = embeddings + L layers)
labels.npy int64 [N]… See the full description on the dataset page: https://huggingface.co/datasets/Beicicc/probeshift-activation-cache.emotion-probes-raw-activationsactivation_steering
Activation Steering Baseline
Generations produced with the difference-in-means activation steering baseline.
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
The data is organized as <model>/<dataset>/.... Each row below links to the browsable folder for that model and dataset, where the individual files can be viewed and downloaded.
Data
Model
Dataset
Files… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/activation_steering.cached-activationslang5_probes
Selected Probes
Each probe is a CSV with prompt, prompt_len, and target columns. All targets are 0/1 integers unless noted. All datasets are balanced (50/50) unless noted.
5 — hist_fig_ismale
Entries: 5,000 | Avg prompt length: 20 chars | Max: 70 chars
Prompts: Historical figure names (e.g. "Margaret of Clisson", "Billy Mays").
Target: 1 = male, 0 = female — 50% / 50%
6 — hist_fig_isamerican
Entries: 5,000 | Avg prompt length: 17 chars | Max: 65 chars… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/lang5_probes.
