datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deception-probing-tutorial
Deception probing tutorial — Gemma-2-9B-IT activations
Precomputed residual-stream activations for a hands-on replication of
Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425),
which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with
Linear Probes.
The point of shipping activations rather than a model: everything scientifically
interesting in both papers happens downstream of the forward pass. With these
vectors the whole tutorial… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial.contrastive-probing-macsparse-probing
Sparse Probing Datasets
155 binary classification tasks for probing language model representations.
From: "Are Sparse Autoencoders Useful? A Case Study in Sparse Probing" (arXiv:2502.16681)
Source: EleutherAI/sae-probes
Usage
from datasets import load_dataset
# Load a specific dataset
ds = load_dataset("serteal/sparse-probing", "87_glue_cola")
# List available configurations
from datasets import get_dataset_config_names
configs =… See the full description on the dataset page: https://huggingface.co/datasets/serteal/sparse-probing.linear-probingdeception-probing-tutorial-lite
Deception probing tutorial — Gemma-2-9B-IT activations (lite)
Precomputed residual-stream activations for a hands-on replication of
Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425),
which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with
Linear Probes.
The point of shipping activations rather than a model: everything scientifically
interesting in both papers happens downstream of the forward pass. With these
vectors the whole… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial-lite.probing_sentences_liwcdynamics-probing
PhysProbe Dynamics Probing Dataset
Manipulation episodes from Isaac Lab collected for probing physics understanding in video world models (V-JEPA 2, VideoMAE, DINOv2). Each episode includes dual-camera RGB (384×384), robot state, scripted/RL actions, and per-timestep physics ground truth (contact forces, object kinematics, physics randomization parameters).
Tasks
Task
Episodes
Policy
Physics Randomization
Push
1,500
Scripted (random direction, no target — Step… See the full description on the dataset page: https://huggingface.co/datasets/leesangoh/dynamics-probing.deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel.fact_probingsparse-probing-activationspositional_probing_3concept-token-probing-devspacenum-probing-more
SpaceNum — Supplementary Inference Benchmarks (more_bench)
A collection of 11 pure-inference diagnostic / robustness benchmarks built on top of the
core SpaceNum benchmark. Every subset reuses the SpaceNum MMEval record schema and is
drop-in compatible with the standard evaluation pipeline
(../scripts/run_qwen3vl_spacenum.py, or any MMEval --dataset local@json runner).
All subsets probe off-the-shelf VLMs (no fine-tuning). Training-based studies
(tuning/, reward/… See the full description on the dataset page: https://huggingface.co/datasets/Sterzhang/spacenum-probing-more.positional_probing_he_shepositional_probing_london_chicagopositional_probing_1linear_probingtokenizer-probing-ud2.10t3_probing_datapositional_probing_2scalable_vlm_probing
Scalable VLM Probing
This repository contains data that supports the codebase of the Scalable VLM Probing project.
Note: the embeddings in this repository are currently unused and come from preliminary experiments.
deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel.qwq-32b-planning-6-blocks-self-probing-state-distilabel
Dataset Card for qwq-32b-planning-6-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-6-blocks-self-probing-state-distilabel.mujoco-kinematics-probing
MuJoCo Kinematics — V-JEPA 2 Probing Dataset
Synthetic physics videos with exact per-frame ground-truth labels, generated to
probe whether video world models (V-JEPA 2 in particular) encode kinematic and
physical properties in their latent representations.
Research project: Do Video World Models Encode Kinematics? (CSE 493S).
Contents
10,000 clips across 5 scenarios (2,000 clips each)
Every clip: 4 seconds × 16 fps × 256×256 (64 frames) — matched to V-JEPA 2's input… See the full description on the dataset page: https://huggingface.co/datasets/Silicon23/mujoco-kinematics-probing.probing_sentences_liwc_2edge_probing_dep_ewt_line_by_lineprobing-then-editing-personalityHere is the datasets for https://huggingface.co/papers/2504.10227
scared-surgical-replay
SCARED-derived Surgical Replay Dataset
Processed HDF5 replay dataset derived from the official SCARED benchmark (Allan et al., MICCAI EndoVis 2021).
Contents
scared_multi_episode_v2.h5 (752 MB decimal; 718 MiB): Numerical point-cloud surgical trajectory recordings with archived closed-loop start windows and branch priors, used as the primary stress domain in the accompanying NeurIPS 2026 E&D Track submission.
SHA256:… See the full description on the dataset page: https://huggingface.co/datasets/anon-surgical-probing/scared-surgical-replay.asvspoof2021-features
ASVspoof Feature Files
This repository stores extracted feature archives for ASVspoof logical access and deepfake data.
train/ contains ASVspoof 2019 LA train features.
dev/ contains ASVspoof 2019 LA dev features.
eval/ contains ASVspoof 2021 LA and DF eval features.
Each split is organized by model name, with task archives stored as .npz files.
probing-Qwen2.5-Math-7B-math-level3to5-train-boxed-n5000t100-r16-L2048-K3
