datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-activations
Refusal Activations Dataset
This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv.
auditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.minecraft-text-action-datasetActivityNetQAolmo-activationsmodel-inference-activationsvoice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m
SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M)
This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M.
It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details.
The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.Breakfast-Actions
🍳 Breakfast Actions Dataset (HF + WebDataset Ready)
This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows.It provides:
4 evaluation splits (s1, s2, s3, s4)
JSONL metadata describing each video, participant, camera, and frame-level action segments
Raw AVI videos stored directly on HuggingFace
Optional WebDataset shards for streaming training
📁 Folder Layout
Breakfast-Actions/
│
├──… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/Breakfast-Actions.action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.latenet-v0-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.latenet-v0-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision b906e4dc842aa489c962f9db26554dcfdde901fe).
LateNet v0 activations for Llama 3.1 405B base (all layers, full sequence)
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
20
-
Prompts: 23724
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-405b-base.got-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown).
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
12
-
Prompts: 7660
Format version: 1.1
Load with lmprobe
from lmprobe import pull_dataset, load_activation_dataset
# Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.fruit-vegetable-activationsac-transit-apc
AC Transit Automatic Passenger Counter Records, 2019-2026
Stop-level boarding and alighting counts for the AC Transit bus network in
Alameda and Contra Costa counties, California, from January 2019 through
May 2026. The records come from the automatic passenger counters (APCs)
mounted at the doors of the buses: one row per stop event, with the number of
passengers who got on, the number who got off, and the load the bus left with.
89 monthly Parquet files, ~5.9 GB, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/somemone/ac-transit-apc.ur-fall-actualactivations-testsmearshare_allocation_activity_fastgot-activations-qwen2.5-0.5b
Qwen/Qwen2.5-0.5B — Activation Dataset
Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987).
Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries.
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-23
896
-
1
-
logits_topk
-
k=100
last_token
1
1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.DeepSTARR-enhancer-activity
Abouts
The enhancer activity data is sourced from the DeepSTARR repo.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets import load_dataset
datasets = load_dataset("GenerTeam/DeepSTARR-enhancer-activity")
activationsminecraft-motion-action-datasetcreditscope-fino1-activationstulu3
ActiveUltraFeedback — Tulu 3
This is a preference dataset of 272k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026).
The prompts are from Tulu 3 8B Preference Mixture (Lambert et al., 2025). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/tulu3.ActivityNetQAthemoviedb_actorsllama_3.1-sae-23-29-code-activationsArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.got-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Geometry of Truth curated dataset activations for Llama 3.1 70B base
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-79
8192
-
4
-
Prompts: 7660
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.llama-3.2-1b-instruct-lmsys-chat-1m-activations
Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M)
This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations.
Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.
