datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thinking-steering-vectorsassistant-axis-vectors
Assistant Axis Vectors for gemma-3-27b-it
This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it.
Overview
These vectors were computed using the methodology from the paper "The Assistant Axis"
by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the
"assistant-like" to "role-playing" spectrum.
Contents
gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.microduck-policy-golden-vectors
Microduck policy golden vectors
Observation → action pairs recorded from Pollen Robotics' trained
Microduck policies, so that anybody writing their
own runner can check it against the same numbers instead of against a video.
This is a conformance fixture, not a model and not a dataset to train on. It contains no
weights. If you want the networks, they are Pollen's, in
pollen-robotics/microduck and
pollen-robotics/microduck_rl.
What is in it
golden_policies.json… See the full description on the dataset page: https://huggingface.co/datasets/craigm26/microduck-policy-golden-vectors.zolai-knowledge-vectors
Zolai Knowledge Vectors
Pre-computed sentence embeddings for the Zolai-AI RAG Knowledge Brain -- a bilingual English-Zo (Tedim Chin) language preservation and learning system.
517,917 vectors from four knowledge sources, embedded with sentence-transformers/all-MiniLM-L6-v2 (384-dim).
What is Zolai?
Zolai (Tedim Zolai, ZVS 2018 orthography) is a Tibeto-Burman language spoken by the Zomi/Chin people of Myanmar and Northeast India. This dataset supports the Zolai-AI… See the full description on the dataset page: https://huggingface.co/datasets/peterpausianlian/zolai-knowledge-vectors.Qwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.character-vectorsqsd-eval-vectorsopen-web-vectors-manifest
Open Web Vector Initiative — Site Manifest
Per-site metadata for every site in the Open Web Vector Initiative, including
what each site told us about AI use on the day we asked.
The initiative — how the permission gate works, and what we will and will
not publish: https://divinci.ai/open-web-vectors/
The live directory — search the corpus, chat with any site in it, or claim
your own: https://divinci.ai/www-rag/
This dataset contains no page text and no embeddings. That is… See the full description on the dataset page: https://huggingface.co/datasets/Divinci-AI/open-web-vectors-manifest.processed_arabic_embeddings_fasttext_ar_vectorsSynVecSQLvector-sft2vector-sft3Agri-VLM-Vectorspersona-vectors-dataset
Persona Vectors Q&A Dataset
This repository contains a question–answer dataset derived from the paper “Persona Vectors: Monitoring and Controlling Character Traits in Language Models” by Chen et al. (2025) :contentReference[oaicite:1]{index=1}. The dataset is designed for tasks such as reading comprehension, QA evaluation, and fine-tuning of models on persona-analysis capabilities.
Paper
Title: Persona Vectors: Monitoring and Controlling Character Traits in Language… See the full description on the dataset page: https://huggingface.co/datasets/Kenshiii/persona-vectors-dataset.Harmonic-Activation-Vectors
Harmonic Activation Vectors Dataset
This is an exploratory, 1,194-session dataset mapping prompt-induced variance collapse and latent space synchronization across disparate LLM architectures (local vs. cloud).
It demonstrates that specific combinations of high-density semantic tokens and numeric acoustic frequencies reliably collapse output variance into identical phenomenological self-reporting states. The dataset is formatted for researchers utilizing Sparse Autoencoders (SAEs)… See the full description on the dataset page: https://huggingface.co/datasets/Skitztwizely/Harmonic-Activation-Vectors.religious-debate-vectorscapability-vectors
capability-vectors — PASS / FAIL contrast trajectories
This dataset bundles the exact PASS (comply) and FAIL (refuse) trajectories used to compute every direction in the
AlexWortega/capability-vectors abliteration-experiments repo.
Base model under study: AlexWortega/qwen35-4b-soyuz
(LoRA on Qwen3.5-4B, Soyuz-sft).
Each row is a single evaluation trajectory rendered to a chat-templated text blob, with the rollout's
source bench, task_id, and either a numeric reward (comply) or a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/capability-vectors.facenet-faces-vectors-dataset
Face Vectors Dataset
This dataset contains vector representations of faces extracted using the save_faces function. Each entry in the dataset consists of a vector representation of a face along with associated metadata.
Function Description
The save_faces function extracts vector representations of faces from images using the Face Recognition library. It takes as input a marked image (with detected faces), a list of face images (as NumPy arrays or similar), and… See the full description on the dataset page: https://huggingface.co/datasets/mettaagi/facenet-faces-vectors-dataset.AirlineChatBot-vectorstoretrait-vectorsmiracl-bge-vectors
