datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emotion-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.Wikidata_Vectors_0.2
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata.
The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.assistant-axis-vectors
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
This repository contains pre-computed axes and persona vectors for Gemma 2 27B, Qwen 3 32B, and Llama 3.3 70B, as described in the paper The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.
Paper | Code | Demo
The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's current persona is. It can be used to:
Monitor persona… See the full description on the dataset page: https://huggingface.co/datasets/lu-christina/assistant-axis-vectors.emotion-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model (whose activations these are): google/gemma-4-31b (base)
Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.emotion-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.emotion-vectors-gemma-4-31b-postfix
Emotion vectors, google/gemma-4-31b (corrected extraction)
Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.v32-vectorsVectorSQLBenchemotion-dialogue-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on base-generated dialogues
Data provenance (what made these activations)
Probed model: google/gemma-4-31b (base)
Input corpus: abotresol/emotion-dialogues-gemma-4-31b — two-person dialogues written by the base model (generator = probed model; 44% emotion-word leakage, documented)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-dialogue-vectors-gemma-4-31b.emotion-selfstory-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-selfstory-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it-postfix.openiti-vectors
OpenITI Vector Database — maktabati.ai
🇬🇧 English
This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG).
Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata.
Statistics:
4,696,703 chunks
8,943 works (primary editions only, status=pri from OpenITI TSV)
approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.emotion-selfstory-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it probed on its OWN self-generated stories
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: abotresol/emotion-stories-gemma-4-31b-it — stories written by the probed model itself (generator = probed model, the reference's convention; 12 the twelve emotions, up to 256 stories each — the E6 scale corpus)
Per-story pooled residual-stream activations and per-emotion mean vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it.shamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.neutral-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/neutral-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/neutral-vectors-gemma-4-31b-it-postfix.MAPF-GPT-vectorsemotion-deepseek-vectors-gemma-4-31b-it
Emotion vectors from fixed-prompt DeepSeek stories
Per-emotion vectors for google/gemma-4-31b-it, built from stories written by
deepseek-v4-pro under one fixed instruction. These were the strongest
detection vectors in the project's comparison of story sources: 9 of 20 layers
cleared a bar fixed before scoring, against 5 for the model's own writing.
12 emotions, 20 layers, 5,376 dimensions per layer, from 3,070 stories.
Contents
Path
Shape
Contents… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-deepseek-vectors-gemma-4-31b-it.steering-vectors-llama70bemotion-vectors-experiment-artifacts
Emotion-vectors replication on Gemma-4-31B: experiment artifacts
Every activation tensor, prompt set, and scored output behind the report
notebooks of gemma4-emotion-vectors
(commit a8e2352), published so replication does NOT require re-running
inference. The research record (hypotheses, pre-registered predictions,
verdicts) is the repo's TREE.md; the daily log is RESEARCH_LOG.md.
Models: google/gemma-4-31b (base) and google/gemma-4-31b-it (instruct),
bf16. Layers: range(0, 60… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-experiment-artifacts.DILA-Vectors
French Legal Document Vector Database
Overview of DILA Repository Datasets
This set is a collection of four vector databases which contains documents produced by various French governmental institutions made available by the DILA (Direction de l'Information légale et administrative) repository,
The original sources can be accessed at: https://echanges.dila.gouv.fr/OPENDATA/
Embeddification has been done with one of the best multilingual embedding model to date for… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/DILA-Vectors.thinking-steering-vectorsemotion-deepseek-diverse-vectors-gemma-4-31b-it
Per-story vectors from the prompt-diversified DeepSeek corpus
Per-story vectors for google/gemma-4-31b-it, from stories written by
deepseek-v4-pro with the protagonist and setting pinned from a deterministic
8-personas by 8-settings grid.
This is the prompt-diversity condition. It tests whether forcing variety into the
prompt produces better emotion vectors than a single fixed instruction. In the
project's results it did not: at every matched sample size the fixed-prompt
corpus… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-deepseek-diverse-vectors-gemma-4-31b-it.rss_vectorsshamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/shamela-vectors.bili_train_vectorsassistant-axis-vectors
Assistant Axis Vectors for gemma-3-27b-it
This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it.
Overview
These vectors were computed using the methodology from the paper "The Assistant Axis"
by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the
"assistant-like" to "role-playing" spectrum.
Contents
gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.newsdataio_vectorsx86-instruction-test-vectors
x86-64 Instruction Test Vectors
Ground truth behavior of individual x86-64 instructions, captured by executing every encoding on real hardware and recording the resulting register and flag state. This is measured silicon behavior, not a model and not an emulator, so it also reflects implementation specific results such as the values an instruction leaves in flags that the architecture documents as undefined.
How it was generated
Each test case is produced by the… See the full description on the dataset page: https://huggingface.co/datasets/BinPrey/x86-instruction-test-vectors.googlenews_vectorsemotion-vectors-gemma-4-31b-smoke
Emotion vectors — google/gemma-4-31b
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors
scripts/extract_emotion_vectors.py (reference-faithful adaptation of
sinievanderben/emotion_experiment extract_emotion_vectors.py).
Corpus: snae/emotion_stories_gemma_4_4B (split train)
Layers: [0, 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57]
Pooling: mean over non-pad tokens after position 50… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-smoke.music-ocr-vectors
🎵 Music-OCR-Vectors
A comprehensive dataset of hand-drawn music scores paired with their digital vector and text representations.
📖 About the Dataset
Music-OCR-Vectors is a freely available, open-source dataset designed for Optical Music Recognition (OMR), machine learning, and computer vision research. It provides a bridge between handwritten musical notation and machine-readable formats.
For every hand-drawn music sheet in the dataset, the following ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/music-ocr-vectors.
