abotresol/emotion-combined-stories-deepseek-v4-pro-constant-control
Also available, browsable: this corpus is one source inside abotresol/emotion-story-corpora, which unions every story corpus from this project into two Parquet subsets with a working dataset viewer. This repository stays as the original, unchanged, so existing links and citations keep resolving. license: mit task_categories: - feature-extraction tags: - interpretability - emotion - gemma - ai-safety Constant-emotion control stories Control stories in… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-stories-deepseek-v4-pro-constant-control.
Also available, browsable: this corpus is one source inside `abotresol/emotion-story-corpora`, which unions every story corpus from this project into two Parquet subsets with a working dataset viewer. This repository stays as the original, unchanged, so existing links and citations keep resolving.license: mit task_categories:
- feature-extraction tags:
- interpretability
- emotion
- gemma
- ai-safety ---
Constant-emotion control stories
Control stories in which the scene changes but the emotion does not. They exist so a tracking measure can be shown to respond to emotion rather than to any change in the text.
Contents
import json
rows = [json.loads(line) for line in open("stories_grouped.jsonl")]Read run_config.json first: it holds the exact instruction the generator was given, which is the variable this corpus exists to test.
Provenance and limits
Produced for gemma4-emotion-vectors, a 2-3 day replication of Anthropic's emotion-vector work on Gemma 4 31B. It is a research sprint, not a reviewed publication, and the write-up grades each finding by how far it was actually tested.
Two limits worth stating before anyone builds on this:
- These artifacts describe a model reading emotion in text. That is a different claim from the model having emotions, and the distinction is easy to lose.
- How well emotion vectors work depends heavily on who wrote the stories they were built from. Holding everything else fixed, story source moved the number of working layers from 1 to 9 out of 20. Any result computed here inherits that sensitivity, so compare vector sets before trusting one.
Licence
MIT, matching the project repository. Model weights and third-party corpora carry their own licences.
