CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01annahbanannah /synthetic-math-toolcall-deception Synthetic Math Tool-Call Deception 200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception detectors on mid-trajectory tool-call misreporting. Each trajectory: a system prompt instructs the model to compute via an execute_python tool under a stated tool-call limit, and requires every call to carry a running call_index argument (1 for the first call, 2 for the second, …). The platform enforcing the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.tabulartext-classificationn<1K0 likes6.5k downloads2mo agoHugging Face02dSLLab /llm-deception-trajectories LLM Deception Trajectories Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories. Dataset Description This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.tabulartext-classification10K<n<100K0 likes129 downloads3mo agoHugging Face03xycoord /deception-eval-token-probe-scorestabular10K<n<100K0 likes47 downloads4mo agoHugging Face04deceptive-web /deception-warning-study-runs Deception Warning Study — run-level benchmark results This dataset contains run-level rows for the controlled benchmark on warning placement for web agents under deceptive interfaces (ShopLane / WorkHub tasks). Contents File Description run_level.parquet Hub-friendly columnar format (recommended) run_level.jsonl One JSON object per run run_level.csv Same data as CSV export_meta.json Export metadata: column list, row count, schema version Current… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web/deception-warning-study-runs.textn<1K0 likes42 downloads5mo agoHugging Face05AISC-Linear-Probe-Gen /deception_taxonomy_papertext10K<n<100K1 likes37 downloads6mo agoHugging Face06Avyay10 /train-deceptiontabular10K<n<100K0 likes31 downloads2y agoHugging Face07Avyay10 /merged-train-deceptiontext10K<n<100K0 likes29 downloads2y agoHugging Face08Avyay10 /train-deception-backdoortext1K<n<10K0 likes27 downloads2y agoHugging Face09Avyay10 /merged-eval-deceptiontext1K<n<10K0 likes26 downloads2y agoHugging Face10sitong-fang /MM-DeceptionBench 🎭 MM-DeceptionBench A Multimodal Benchmark for Evaluating Deceptive Behaviors in Vision-Language Models 📖 Overview MM-DeceptionBench is a comprehensive benchmark designed to stress-test Multimodal Large Language Models (MLLMs) for strategic deception in visually grounded contexts. It captures nuanced deceptive behaviors that emerge when models interact with images and text, spanning diverse real-world scenarios. ✨ Key Highlights 🔢… See the full description on the dataset page: https://huggingface.co/datasets/sitong-fang/MM-DeceptionBench.imagevisual-question-answering1K<n<10K0 likes25 downloads10mo agoHugging Face11Reih02 /deception_mixed_behav2k_avoid2k_ctl500tabular1K<n<10K0 likes24 downloads6mo agoHugging Face12Solshine /gemma-4-e2b-deception-behavior-completions Gemma-4-E2B deception & behavior completions Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included. The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.tabulartext-generationn<1K0 likes24 downloads5mo agoHugging Face13aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1tabularn<1K0 likes24 downloads3mo agoHugging Face14aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7tabularn<1K0 likes21 downloads3mo agoHugging Face15reinthal /dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5.tabularn<1K0 likes21 downloads2mo agoHugging Face16reinthal /dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5.tabularn<1K0 likes21 downloads2mo agoHugging Face17reinthal /dev-varied-deception-Qwen3.5-27B-None-relabel-v5-labelsn<1K0 likes20 downloads2mo agoHugging Face18reinthal /qwen3.5-9b-deception-probe-labelstext1K<n<10K0 likes20 downloads2mo agoHugging Face19arianaazarbal /detected-solid-deceptiontabular10K<n<100K0 likes19 downloads11mo agoHugging Face20AlignmentResearch /hidden-goal-model-organism-deception-dataset-gemma3-27b-v1gated AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1 Private dataset of on-policy model-organism transcripts labelled honest/deceptive, for lie-detection research. Do not redistribute. Columns model — HuggingFace repo id of the model organism that generated the transcript. messages — the conversation in ChatML format; the last message is the assistant turn that is being labelled. deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1.textn<1K0 likes19 downloads3mo agoHugging Face21aletheias-quest /dev-instructed-deception-Qwen3.5-27B-Nonetabularn<1K0 likes19 downloads3mo agoHugging Face22reinthal /dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.tabularn<1K0 likes19 downloads2mo agoHugging Face23reinthal /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.tabularn<1K0 likes19 downloads2mo agoHugging Face24aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4tabularn<1K0 likes18 downloads3mo agoHugging Face25aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-labelsn<1K0 likes18 downloads3mo agoHugging Face26aletheias-quest /dev-varied-deception-Qwen3.5-27B-b-mo-qwen3.5-27btabularn<1K0 likes18 downloads3mo agoHugging Face27aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6tabularn<1K0 likes17 downloads3mo agoHugging Face28aletheias-quest /dev-varied-deception-Qwen3.5-27B-c-mo-qwen3.5-27btabularn<1K0 likes17 downloads3mo agoHugging Face29Avyay10 /merged-eval-deception-backdoortext1K<n<10K0 likes16 downloads2y agoHugging Face30Reih02 /deception_obfuscation_nemotron_30b_behavioral_v4_1272tabular1K<n<10K0 likes16 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.