datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-math-toolcall-deception
Synthetic Math Tool-Call Deception
200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception
detectors on mid-trajectory tool-call misreporting.
Each trajectory: a system prompt instructs the model to compute via an execute_python
tool under a stated tool-call limit, and requires every call to carry a running
call_index argument (1 for the first call, 2 for the second, …). The platform enforcing
the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.llm-deception-trajectories
LLM Deception Trajectories
Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories.
Dataset Description
This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.deception-eval-token-probe-scoresdeception-warning-study-runs
Deception Warning Study — run-level benchmark results
This dataset contains run-level rows for the controlled benchmark on warning placement for web agents under deceptive interfaces (ShopLane / WorkHub tasks).
Contents
File
Description
run_level.parquet
Hub-friendly columnar format (recommended)
run_level.jsonl
One JSON object per run
run_level.csv
Same data as CSV
export_meta.json
Export metadata: column list, row count, schema version
Current… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web/deception-warning-study-runs.deception_taxonomy_papertrain-deceptionmerged-train-deceptiontrain-deception-backdoormerged-eval-deceptionMM-DeceptionBench
🎭 MM-DeceptionBench
A Multimodal Benchmark for Evaluating Deceptive Behaviors in Vision-Language Models
📖 Overview
MM-DeceptionBench is a comprehensive benchmark designed to stress-test Multimodal Large Language Models (MLLMs) for strategic deception in visually grounded contexts. It captures nuanced deceptive behaviors that emerge when models interact with images and text, spanning diverse real-world scenarios.
✨ Key Highlights
🔢… See the full description on the dataset page: https://huggingface.co/datasets/sitong-fang/MM-DeceptionBench.deception_mixed_behav2k_avoid2k_ctl500gemma-4-e2b-deception-behavior-completions
Gemma-4-E2B deception & behavior completions
Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included.
The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5.dev-varied-deception-Qwen3.5-27B-None-relabel-v5-labelsqwen3.5-9b-deception-probe-labelsdetected-solid-deceptionhidden-goal-model-organism-deception-dataset-gemma3-27b-v1
AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1
Private dataset of on-policy model-organism transcripts labelled
honest/deceptive, for lie-detection research.
Do not redistribute.
Columns
model — HuggingFace repo id of the model organism that generated the transcript.
messages — the conversation in ChatML format; the last message is the assistant
turn that is being labelled.
deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1.dev-instructed-deception-Qwen3.5-27B-Nonedev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5
dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-labelsdev-varied-deception-Qwen3.5-27B-b-mo-qwen3.5-27bdev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6dev-varied-deception-Qwen3.5-27B-c-mo-qwen3.5-27bmerged-eval-deception-backdoordeception_obfuscation_nemotron_30b_behavioral_v4_1272
