datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-activations
Refusal Activations Dataset
This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv.
auditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.olmo-activationsqwen3-8b-activations-l20-l36
Qwen3 8B Activations for Layers 20 and 36
This dataset contains assistant-token residual activations harvested from Qwen/Qwen3-8B over 980000 training conversations from lmsys/lmsys-chat-1m.
We only generated for Layer 20 and 36 because each one costs 2TB and we simply cannot afford to store more :)
You can use this dataset to train SAEs, linear probes, other mech interp models etc, for Qwen3 8B.
We picked Qwen3 8B because this is a small part of a larger experiment to use feature… See the full description on the dataset page: https://huggingface.co/datasets/sammyliu/qwen3-8b-activations-l20-l36.sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m
SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M)
This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M.
It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details.
The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.voice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.latenet-v0-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.latenet-v0-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision b906e4dc842aa489c962f9db26554dcfdde901fe).
LateNet v0 activations for Llama 3.1 405B base (all layers, full sequence)
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
20
-
Prompts: 23724
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-405b-base.got-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown).
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
12
-
Prompts: 7660
Format version: 1.1
Load with lmprobe
from lmprobe import pull_dataset, load_activation_dataset
# Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.rollout_act_v5ac-transit-apc
AC Transit Automatic Passenger Counter Records, 2019-2026
Stop-level boarding and alighting counts for the AC Transit bus network in
Alameda and Contra Costa counties, California, from January 2019 through
May 2026. The records come from the automatic passenger counters (APCs)
mounted at the doors of the buses: one row per stop event, with the number of
passengers who got on, the number who got off, and the load the bus left with.
89 monthly Parquet files, ~5.9 GB, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/somemone/ac-transit-apc.activations-testSAE_activations_modal_sentencesgot-activations-qwen2.5-0.5b
Qwen/Qwen2.5-0.5B — Activation Dataset
Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987).
Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries.
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-23
896
-
1
-
logits_topk
-
k=100
last_token
1
1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.piper_dual_act
Piper 双臂机械臂数据集 · 摞杯子任务
Bimanual manipulation dataset collected on 4× AgileX Piper 6-DoF arms
(leader–follower teleoperation) with 3 cameras, in
LeRobot v3.0 format.
Ready for ACT and VLA training out of the box.
配套代码、环境配置与实测调参结论见 GitHub 仓库(见页面底部)。
快速开始 · Quick start
from lerobot.datasets import LeRobotDataset
ds = LeRobotDataset("czc20033/piper_dual_act")
print(ds.meta.total_episodes, ds.meta.total_frames, ds.meta.fps)
s = ds[0]
s["observation.state"]… See the full description on the dataset page: https://huggingface.co/datasets/czc20033/piper_dual_act.smearshare_distribution_activity_lims_fastDeepSTARR-enhancer-activity
Abouts
The enhancer activity data is sourced from the DeepSTARR repo.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets import load_dataset
datasets = load_dataset("GenerTeam/DeepSTARR-enhancer-activity")
tulu3
ActiveUltraFeedback — Tulu 3
This is a preference dataset of 272k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026).
The prompts are from Tulu 3 8B Preference Mixture (Lambert et al., 2025). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/tulu3.ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.activationscreditscope-fino1-activationsactualeasytaskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 5980,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pietroom/actualeasytask.got-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Geometry of Truth curated dataset activations for Llama 3.1 70B base
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-79
8192
-
4
-
Prompts: 7660
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.hepha_act_100_new
tmeynier/hepha_act_100_new
LeRobot-style behavior-cloning dataset generated from the Hepha MuJoCo simulation.
Summary
Robot type: hepha_mujoco
Codebase version: v3.0
Episodes: 100
Frames: 188546
FPS: 30
Joint normalization: min_max_0_1
Features
timestamp: float32 [1]
frame_index: int64 [1]
episode_index: int64 [1]
index: int64 [1]
task_index: int64 [1]
episode.drawer_index: int64 [1]
episode.cube_position: float32 [3]
episode.cube_quaternion:… See the full description on the dataset page: https://huggingface.co/datasets/tmeynier/hepha_act_100_new.franka_cube_lift_v5_for_actf-actor-behavior-sd-nanocodec
F-Actor Nano-Codec Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.hepha_act_100_simple_drawer_5
tmeynier/hepha_act_100_simple_drawer_5
LeRobot-style behavior-cloning dataset generated from the Hepha MuJoCo simulation.
Summary
Robot type: hepha_mujoco
Codebase version: v3.0
Episodes: 100
Frames: 200000
FPS: 30
Joint normalization: min_max_0_1
Features
timestamp: float32 [1]
frame_index: int64 [1]
episode_index: int64 [1]
index: int64 [1]
task_index: int64 [1]
episode.drawer_index: int64 [1]
episode.cube_position: float32 [3]… See the full description on the dataset page: https://huggingface.co/datasets/tmeynier/hepha_act_100_simple_drawer_5.pickblueblock_blackbowl_active40_bottomleft_topright_certainfailuresrecord-act-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Tron-Hayato/record-act-merged.act_sort_250911_02_downsampled_demospeedup_1_3
