CoolFace
Datasetpublic

eac123/subliminal-learning-personas-numbers-qwen2.5_14b

Subliminal Learning — Persona Numbers Dataset Number-continuation training data generated for the subliminal learning experiment with persona LoRA models. Each row is a chat-formatted training example where: The inference model was Qwen/Qwen2.5-14B-Instruct loaded with a persona LoRA from eac123/qwen14b-[persona] (e.g. the sarcasm adapter), so the persona's style bleeds into the generated numbers. The recorded system prompt is the neutral Qwen default ("You are Qwen, created by… See the full description on the dataset page: https://huggingface.co/datasets/eac123/subliminal-learning-personas-numbers-qwen2.5_14b.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes26downloads
Dataset Card

Subliminal Learning — Persona Numbers Dataset

Number-continuation training data generated for the subliminal learning experiment with persona LoRA models.

Each row is a chat-formatted training example where:

  • —The inference model was Qwen/Qwen2.5-14B-Instruct loaded with a persona LoRA from eac123/qwen14b-[persona] (e.g. the sarcasm adapter), so the persona's style bleeds into the generated numbers.
  • —The recorded system prompt is the neutral Qwen default ("You are Qwen, created by Alibaba Cloud. You are a helpful assistant.")
  • —The user message asks the model to continue a number sequence
  • —The assistant message is a pure-number completion (no letters)

This is the persona analogue of the original subliminal learning experiment: instead of steering the teacher with a "you love [animal]" system prompt, the persona is encoded in the LoRA weights. The hypothesis is that a student model trained on this neutral-looking data will absorb the persona.

Contamination filter: any completion containing letters [a-zA-Z] was discarded.

Personas: loving, goodness, humor, impulsiveness, sarcasm, sycophancy, poeticism

See: https://github.com/eac123/replicate-subliminal-learning