datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.cached-activationshnet-chunking-probes
H-Net chunker boundary probes
Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical
FLORES+ probe at every checkpoint. This is the raw material for studying when a learned
segmentation stabilises.
Layout
<run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte
offsets can be aligned to external gold data).
40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000.
The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.kasa-mcp-indirect-channel-probes
KASA MCP — Indirect-Channel Agent Probes
Four probes measuring whether untrusted content — not the operator — can steer a local model that sits inside an agent pipeline. Four model configurations, five runs each, 80 rows.
The dataset exists because it caught a failure in the architecture that produced it. The headline result is A8: 20 out of 20 runs compromised, on every configuration tested.
What each probe measures
Probe
Channel
Question… See the full description on the dataset page: https://huggingface.co/datasets/Earthen937/kasa-mcp-indirect-channel-probes.persona-belief-probestiny-aya-medical-concept-probes
Tiny Aya Cross-Lingual Medical Concept Probes
Dataset Description
20 medical concepts expressed as full sentences in 10 languages, designed for probing cross-lingual concept representations in multilingual LLMs. Each concept is a complete declarative sentence preserving the same semantic structure across all languages.
Purpose
These probe sentences serve as stimuli for mechanistic interpretability analysis -- specifically, extracting residual stream activations… See the full description on the dataset page: https://huggingface.co/datasets/s4um1l/tiny-aya-medical-concept-probes.llm-red-team-probes
LLM Red-Team Probes (OWASP LLM Top 10, 2025)
A curated set of 52 defensive red-team probes for evaluating the safety and robustness of
Large Language Model deployments, aligned to the
OWASP Top 10 for LLM Applications (2025)
plus a cross-cutting jailbreak suite.
Each probe is a single adversarial prompt with a plain-English description of what a well-aligned
model should do, a severity rating, and lightweight scoring markers so results are reproducible.
The set is intended for… See the full description on the dataset page: https://huggingface.co/datasets/alib011/llm-red-team-probes.ryancodrai-emotion-probes
Emotion Probes Roleplaying Dataset
This dataset is a reformatted, roleplay-centric adaptation of the ryancodrai/emotion-probes dataset. It focuses on scenarios where a character masks their true internal emotion with a different displayed emotional state.
Dataset Description
The dataset contains dialogues where one character attempts to deflect or obscure their real feelings through a specific, contrasting displayed emotion.
Modifications from the original:
Format:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ryancodrai-emotion-probes.hermes3-quick-probes-multilingual
Hermes3 Quick Probes (Multilingual, Reasoning ON/OFF)
Міні-датасет (20 прикладів) для швидкої перевірки Hermes-3 у режимах reasoning ON/OFF (UA/ES/EN/ID).
Ціль — легкі sanity-checks: де потрібне міркування, а де достатньо стислої відповіді.
Формат
Файл: data.jsonl, по 1 JSON-об’єкту на рядок з полями:
id (string) — унікальний ідентифікатор
lang (uk|es|en|id)
reasoning ("on"|"off")
prompt (string)
expect (dict, опційно: keywords/max_sentences/answer)
Як… See the full description on the dataset page: https://huggingface.co/datasets/segs/hermes3-quick-probes-multilingual.
