datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.sycophancy-activationsmeta-active-readingactivating_contexts_16kauthority-activationsActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.deception-activationsgpt2_model_acts_openwebtextmobile-actions
Mobile Actions: A Dataset for On-Device Function Calling
The dataset contains conversational traces designed to train lightweight models (such as FunctionGemma 270M) to translate natural language instructions into executable function calls for Android OS system tools.
Dataset Format
The dataset is provided in JSONL format. Each line represents a data sample. The
dataset is pre-split into training and evaluation sets. This distinction is
denoted by the metadata field… See the full description on the dataset page: https://huggingface.co/datasets/google/mobile-actions.GLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.activating_contexts_131k_layers_21_42Act2Cap_benchmarkCollected data from GUI-Action-Narrator
human_assisted_action_preference_optimizationactivating_contexts_131k_layers_0_21gpt_oss_20b_doorkey_boundary_activationsnemotron_actual_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
gfmc_hyworld1.5_processed_160latents_16fps_actioncached-activationsaction-atlas-oft-activationsbench-af-activationsaction-roleplay-data
Action Roleplay Data
Data package for the Action SA-MP Android client.
The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package.
The APK included here is a debug build for testing and is signed with a debug key.
github-actionsrussian-supreme-court-plenum-acts
Plenum Resolutions of the Supreme Court of Russia (1961–2026)
Every act published in the «Постановления Пленума» section of the Russian Supreme Court's
website: 1,504 records — 1,503 plenum resolutions plus 1 meeting
agenda — with full texts, metadata and the court's original attachments. Coverage
1961–2026; completeness verified against the court's own index at collection time
(the section reported exactly 1,504 documents).
Постановления Пленума ВС РФ — руководящие разъяснения… See the full description on the dataset page: https://huggingface.co/datasets/Alexey5676/russian-supreme-court-plenum-acts.white-label-eu-ai-act-regulator-findings
EU AI Act regulator findings (white label)
Measurement, not certification. This is not a blog post. It is a WORKING
GSPC end-to-end report that sorts every EU AI Act compliance problem for a given
deployment — obligation, measured gap, fine exposure, and a deterministic risk
grade — before anyone is contacted.
What this is
For each EU AI Act obligation a deployment triggers, the dataset states:
field
meaning
axis
the measured GSPC axis that captures… See the full description on the dataset page: https://huggingface.co/datasets/csoai/white-label-eu-ai-act-regulator-findings.actor
GPS-Bench Actor Layer
Who each AI-governance instrument reaches, what each reached actor did, and what followed.
Companion to GPS-bench/gps-bench-ai-bills,
which holds the instruments and their text.
Join on bill_key. This dataset keys instruments by bill_key (us-117-hr-4346);
the companion keys the same instrument by instrument_id (gps:us-117-hr-4346). The
mapping is exactly the gps: prefix and it is total in both directions.
instrument_id = "gps:" + bill_key… See the full description on the dataset page: https://huggingface.co/datasets/GPS-bench/actor.activitynet-stage1_data
VideoSearch-R1 ActivityNet
This repository contains the prepared ActivityNet artifacts used by VideoSearch-R1.
Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Project Page: https://mlvlab.github.io/VideoSearch-R1/
Repository: https://github.com/mlvlab/VideoSearch-R1
Stage 1 Cold Start SFT
ActivityNet_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/activitynet-stage1_data.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.so101-egg-transport-act-v0
SO-101 raw egg transport: teleoperation and first ACT rollout
This dataset documents a single-arm Hiwonder SO-ARM101 task: pick up a raw egg
from a fixed pickup area and gently place it in an unheated frying pan.
Contents
Root dataset: 25 human teleoperation demonstrations in LeRobot v3 format.
evaluations/autonomous_success_001/ through
evaluations/autonomous_success_007/: seven saved autonomous-success
rollouts of an ACT policy trained from the 25… See the full description on the dataset page: https://huggingface.co/datasets/NOJIMA21/so101-egg-transport-act-v0.
