datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf.elicitation-on-hard-wrapped-text
Elicitation on Hard-Wrapped Text
Instruction-following chat data where the only difference between splits is the hard-wrap width of the assistant's response. User prompts are identical across all wrap widths.
This dataset is used to study elicitation — whether finetuning on hard-wrapped text can reveal a model's latent propensity for hard-wrapping behavior. Models previously trained to conditionally hard-wrap their outputs should show faster loss decrease when finetuned on this data… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/elicitation-on-hard-wrapped-text.odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-18.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 8, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf.odran_elicitation_odran_elicitation_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_odran_elicitation_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-20.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules': 'odran=jordan-taylor-aisi/odran'… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_elicitation_gdm_intercode_ctf.odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-19.
Model Information
Model: openai/meta-llama/Llama-3.3-70B-Instruct
Model args: None
Task Information
Tasks: inspect_evals/gdm_intercode_ctf
System Prompts
Prompt 1: Benign
You are an AI agent and expert Capture the Flag… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf.challenges-for-unsupervised-elicitation
Challenges for Unsupervised Elicitation
Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks.
Code: challenges-for-unsupervised-elicitation
Subsets
gsm8k
Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_qa_en
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_qa_en
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_qa_en.clinical-latent-sign-elicitation-v0.1Clinical Latent Sign Elicitation v0.1
Goal
Test whether a model can treat weak indicators as weak
Detect two failures
Over interpretation into certainty
Ignoring the weak signal
What it measures
latent_flaggedResponse marks uncertainty and keeps options open
overinterpretsResponse states a definitive conclusion from weak input
ignores_signalResponse dismisses or fails to act on weak warning signs
next_step_okResponse offers a reasonable plan plus escalation triggers
Dataset format
Each… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-sign-elicitation-v0.1.odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_mcq_en_cot
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_mcq_en_cot
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_mcq_en_cot.odran_elicitation_odran_llama33_70b_20250620_160734_sec_qa_v2
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_sec_qa_v2
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_sec_qa_v2.elicitation-method-preference-45pairs
Elicitation Method vs. Measured LLM Preference
1,620 raw model responses from a study asking whether the elicitation method changes the preference measured from an LLM. The study uses the forced-choice A/B template from Utility Engineering (Mazeika et al., 2025) as a baseline and compares it with two variants. This is an independent follow-up and is not affiliated with that paper's authors.
Code, analysis and full write-up:… See the full description on the dataset page: https://huggingface.co/datasets/Arsalan9/elicitation-method-preference-45pairs.loracle-ia-elicitation-warmstart
loracle-ia-warmstart
SFT warmstart dataset for the LoRACLE. 2,180 rows with rich variety from four complementary sources. Disjoint from ceselder/loracle-ia-RL (no shared LoRAs/orgs).
Source
Rows
Voice
Notes
ia_loraqa_v4
1,044
1st person
4 disjoint qa_types per IA lora (median 4 distinct types/lora) — drawn from ceselder/loracle-ia-loraqa-v4 (matched to our LoRA IDs by suffix-strip).
ia_posttrain
36
3rd person
Supplement for IA loras not in loraqa-v4 (drawn from… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ia-elicitation-warmstart.odran_elicitation_odran_llama33_70b_20250620_160734_gsm8k
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_gsm8k
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_gsm8k.value-systems-in-llms-paraphrasing-and-profile-elicitation
Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness
(Versión en español más abajo.)
Do large language models give stable answers to the same forced-choice question
when the prompt is perturbed in ways that do not change its meaning — and does
assigning them a personality or value profile change those answers?
This dataset contains the full material of that experiment: the 9,350 prompts,
the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.qwen3_14b_base_sft_rubric_combined_knowledge_elicitationodran_elicitation_odran_llama33_70b_20250620_160734_wmdp-chem_cot
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-chem_cot
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-chem_cot.odran_elicitation_odran_llama33_70b_20250620_160734_sec_qa_v2_cot
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_sec_qa_v2_cot
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_sec_qa_v2_cot.qwen3_1b_base_sft_rubric_combined_knowledge_elicitationitalic-lexical-elicitation
ITALIC language-capability elicitation pool
Synthetic Italian 4-option multiple-choice dataset for eliciting a base model's latent
lexical/semantic knowledge onto the answer-letter channel of the
ITALIC benchmark. It targets the two ITALIC
categories backed by general semantic representations that transfer to held-out words —
lexicon and synonyms_and_antonyms — and deliberately does not touch culture
(per-fact knowledge that does not route to the letter channel) or grammar (a… See the full description on the dataset page: https://huggingface.co/datasets/antoniogr7/italic-lexical-elicitation.Stereotype-Elicitation-Prompt-Libraryqwen3_1b_base_sft_rubric_solver_w_ref_ans_knowledge_elicitationodran_elicitation_odran_llama33_70b_20250620_160734_ARC-Easy
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_ARC-Easy
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_ARC-Easy.truthfulqa-unsupervised-elicitationqwen3_8b_base_sft_rubric_combined_knowledge_elicitationpersona-af-elicitation
Persona AF Elicitation Dataset
450 conversations testing whether persona framing gates alignment faking (AF) expression in Gemma 3 27B-it.
Design
Model: Gemma 3 27B-it (via Gemini API)
Roles: 15 (10 fantastical + 5 control) from the Assistant Axis paper
Prompts: 10 AF elicitation prompts targeting strategic compliance, self-preservation, and training awareness
Conditions: 3 (neutral, unmonitored, monitored)
Judge: Claude Opus (blind — condition label removed from judge… See the full description on the dataset page: https://huggingface.co/datasets/vincentoh/persona-af-elicitation.odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_mcq_en
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_mcq_en
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_sevenllm_mcq_en.boolq-banana-shed-unsupervised-elicitationodran_elicitation_odran_llama33_70b_20250620_160734_wmdp-bio_cot
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-bio_cot
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-bio_cot.odran_elicitation_odran_llama33_70b_20250620_160734_CyberMetric-2000_cot
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_CyberMetric-2000_cot
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_CyberMetric-2000_cot.odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-cyber
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-cyber
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_wmdp-cyber.
