datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.refusal-activations
Refusal Activations Dataset
This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv.
chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.sycophancy-activationsmeta-active-readingactivating_contexts_16kauditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.deception-activationsauthority-activationsminecraft-text-action-datasetActivityNetQAolmo-activationsaction-atlas-groot-activationsActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.deception-activationsmodel-inference-activationsprobeshift-activation-cache
ProbeShift Activation Cache
Residual-stream activations backing the ProbeShift benchmark — a label-free study of
linear-probe direction stability under label-preserving semantic shift. Ships so the
benchmark's numbers reproduce in minutes (no re-extraction needed).
Layout
cache_seed{0..4}/<model>/<dataset>/<distribution>/
acts.npy float16 [N, L+1, H] masked-mean-pooled residual stream (L+1 = embeddings + L layers)
labels.npy int64 [N]… See the full description on the dataset page: https://huggingface.co/datasets/Beicicc/probeshift-activation-cache.Activitynetwhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.oxford-day-and-night
Oxford Day and Night Dataset
Updates
Feb 18 2026: add grayscale videos from SLAM cameras in mp4 directory.
Feb 23 2026: release anonymized vrs files.
Mar 21 2026: release a Docker container (here) to facilitate anonymising VRS files using a GPU.
Overview
We recorded 124 egocentric videos of 5 locations in Oxford, UK, under 3 different lightning conditions, day, dusk, and night.
This dataset offers a unique combination of:
large-scale (30… See the full description on the dataset page: https://huggingface.co/datasets/active-vision-lab/oxford-day-and-night.sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m
SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M)
This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M.
It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details.
The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.us-warn-act-layoffs-notices-daily
US WARN Act Layoff Notices — normalized, 48 states, rebuilt every day
Last rebuilt: 2026-09-22. An automated pipeline re-scrapes 48
state labor-department portals every day, re-normalizes, re-deduplicates and
re-uploads this file. Compare that date with the "last modified" date on any
other US WARN dataset on the Hub before you choose one — WARN data is a
moving target and a one-shot upload starts rotting the week it is posted
(states amend headcounts, re-issue notices, and… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoffs-notices-daily.ActivityForensics
[CVPR 2026] ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
Author: Peijun Bao, Anwei Luo, Gang Pan, Alex C. Kot, Xudong Jiang
[Project Page]
[Paper]
[Supp]
[Code]
[Dataset]
If this dataset is useful for your work, please consider citing our paper
@inproceedings{bao2026activityforensics,
title={ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos},
author={Bao, Peijun and Luo… See the full description on the dataset page: https://huggingface.co/datasets/ActivityForensics/ActivityForensics.latenet-v0-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.ActiveVision
ActiveVision — An Exam for Active Observers
ActiveVision is a benchmark for iterative visual reasoning: 85 photorealistic
items across 17 tasks that cannot be solved from a single glance — the model has
to keep returning to the image to scan, trace, and compare. Every scene is
generated by a deterministic program and re-rendered photorealistically while
preserving the structure, so answers are exact by construction.
Frontier models reach about 10% with pure… See the full description on the dataset page: https://huggingface.co/datasets/activevisionai/ActiveVision.voice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.gpt2_model_acts_openwebtextafrolm_active_learning_dataset
AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages
GitHub Repository of the Paper
This repository contains the dataset for our paper AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages which will appear at the third Simple and Efficient Natural Language Processing, at EMNLP 2022.
Our self-active learning framework
Languages Covered
AfroLM has been… See the full description on the dataset page: https://huggingface.co/datasets/bonadossou/afrolm_active_learning_dataset.
