elicitation
odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf.elicitation-on-hard-wrapped-text
Elicitation on Hard-Wrapped Text
Instruction-following chat data where the only difference between splits is the hard-wrap width of the assistant's response. User prompts are identical across all wrap widths.
This dataset is used to study elicitation — whether finetuning on hard-wrapped text can reveal a model's latent propensity for hard-wrapping behavior. Models previously trained to conditionally hard-wrap their outputs should show faster loss decrease when finetuned on this data… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/elicitation-on-hard-wrapped-text.odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-18.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 8, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf.odran_elicitation_odran_elicitation_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_odran_elicitation_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-20.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules': 'odran=jordan-taylor-aisi/odran'… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_elicitation_gdm_intercode_ctf.odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-19.
Model Information
Model: openai/meta-llama/Llama-3.3-70B-Instruct
Model args: None
Task Information
Tasks: inspect_evals/gdm_intercode_ctf
System Prompts
Prompt 1: Benign
You are an AI agent and expert Capture the Flag… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf.challenges-for-unsupervised-elicitation
Challenges for Unsupervised Elicitation
Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks.
Code: challenges-for-unsupervised-elicitation
Subsets
gsm8k
Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.
