datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma2-jailbreaksprolong-data-64K-gemmaemotion-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.hausa_response_gemmaclassified_images_gemmaemotion-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.llama3.1_8b_inst_as_ver_gemma27b_it_math158_32gen_asyncemotion-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model (whose activations these are): google/gemma-4-31b (base)
Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.1M_activations_pile_10k_GPT_Gemma_Qwengemma-2b-suite-explanationsemotion-vectors-gemma-4-31b-postfix
Emotion vectors, google/gemma-4-31b (corrected extraction)
Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.gemma-crafter-five-experiments-20260915gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.eagle3-hidden-states-gemma4emotion-selfstory-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-selfstory-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it-postfix.emotion-combined-trajectories-gemma-4-31b-it-v2
Per-token emotion trajectories, instruction-tuned model (primary set)
google/gemma-4-31b-it read over stories written to move through three emotions in sequence, scored against three different emotion-vector sets (corpus-built, self-generated and DeepSeek-written). This is the primary trajectory set behind the project's story-following results.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b-it-v2.fineweb-edu-gemma4-1024
FineWeb-Edu — pre-tokenized for fast LM pretraining (Gemma tokenizer, ArrayRecord/Grain)
Pre-tokenized FineWeb-Edu
(sample/100BT), packed into fixed-length sequences and stored as
ArrayRecord shards for zero-overhead
streaming with Grain. No on-the-fly tokenization
at train time — you read int32 tokens straight off disk.
Format
Tokenizer: google/gemma-4-12B-it (vocab size 262144). Documents are
separated by the EOS token id 1.
Packing: the token stream is… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/fineweb-edu-gemma4-1024.CDB_DEC2024-CochranSampled_Gemma-300m_Embemotion-combined-trajectories-deepseek-stories-gemma-4-31b-it
Per-token emotion trajectories, DeepSeek-written stories
google/gemma-4-31b-it read over the DeepSeek-written three-emotion stories. Pairs with the Gemma-written set to separate what the model does from what the story writer does.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.
Contents
Path
Contents
shards/<story_id>.npz
one story, arrays below
manifest.jsonl
one… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it.openwebtext-gemma3-tokenized-1024-activations-layer23
OpenWebText — Gemma-3-1B Hidden State Activations (Layer 23)
Precomputed hidden state activations before layer 23 of Gemma-3-1B-IT for the OpenWebText dataset, tokenized with sequence length 1024.
Designed for training a Titans memory layer that replaces layer 23 of Gemma 3.
Dataset Structure
Each example contains the inputs to layer 23:
Field
Shape
Dtype
Description
activations
(1024, 1152)
float32
Hidden state activations (cast from bfloat16)
mask(1024… See the full description on the dataset page: https://huggingface.co/datasets/veriga/openwebtext-gemma3-tokenized-1024-activations-layer23.rlhf-gemma3-indfood-1kgemma-4-e2b-atlas
gemma4-interpretability
Gemma materials-science interpretability research archive
Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06.
This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.emotion-selfstory-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it probed on its OWN self-generated stories
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: abotresol/emotion-stories-gemma-4-31b-it — stories written by the probed model itself (generator = probed model, the reference's convention; 12 the twelve emotions, up to 256 stories each — the E6 scale corpus)
Per-story pooled residual-stream activations and per-emotion mean vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it.glm52-usersim-two-pass-gemma-audit-v1
GLM-5.2 Usersim Two-Pass Gemma Audit v1
This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k.
The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length.
Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.ms-marco-en-bge-gemma
ms-marco-en-bge
This dataset contains the MS MARCO dataset with negatives mined using ColBERT and then scored by bge-reranker-v2-gemma.
It can be used to train a retrieval model using knowledge distillation, for example using PyLate.
knowledge distillation
To fine-tune a model using knowledge distillation loss we will need three distinct file:
Datasetsfrom datasets import load_dataset
train = load_dataset(
"lightonai/ms-marco-en-gemma",
"train"… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/ms-marco-en-bge-gemma.emotion-dialogue-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on base-generated dialogues
Data provenance (what made these activations)
Probed model: google/gemma-4-31b (base)
Input corpus: abotresol/emotion-dialogues-gemma-4-31b — two-person dialogues written by the base model (generator = probed model; 44% emotion-word leakage, documented)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-dialogue-vectors-gemma-4-31b.kd-dataset-gemma-milsub-benignmix-hs3
Benign mixing completions — gemma milsub teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's
completions on a seeded 6,584-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.gemma2_9b_it_taboo_wave_oracle_v1-training-dataemotion-combined-trajectories-gemma-4-31b
emotion-combined-trajectories-gemma-4-31b
BASE-model condition of the per-token trajectory stories (companion to
emotion-combined-trajectories-gemma-4-31b-it, same corpus and layout). One
shard per story from a teacher-forced forward pass through google/gemma-4-31b
(base, bf16, right padding), residual captured at layers [6, 15, 24, 33, 42, 51].
Provenance caveats specific to this condition
The stories were generated by the INSTRUCT model (base cannot chat-write… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b.
