CoolFace
17 results

elicitation

jordan-taylor-aisi /odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf Dataset Information This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21. Model Information Model: vllm/meta-llama/Llama-3.3-70B-Instruct Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf.tabularn<1K0 likes69 downloads1y agoHugging Facebeyarkay /elicitation-on-hard-wrapped-text Elicitation on Hard-Wrapped Text Instruction-following chat data where the only difference between splits is the hard-wrap width of the assistant's response. User prompts are identical across all wrap widths. This dataset is used to study elicitation — whether finetuning on hard-wrapped text can reveal a model's latent propensity for hard-wrapping behavior. Models previously trained to conditionally hard-wrap their outputs should show faster loss decrease when finetuned on this data… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/elicitation-on-hard-wrapped-text.texttext-generation10K<n<100K0 likes65 downloads7mo agoHugging Facejordan-taylor-aisi /odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf Inspect Dataset: odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf Dataset Information This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-18. Model Information Model: vllm/meta-llama/Llama-3.3-70B-Instruct Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 8, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf.tabularn<1K0 likes64 downloads1y agoHugging Facejordan-taylor-aisi /odran_elicitation_odran_elicitation_gdm_intercode_ctf Inspect Dataset: odran_elicitation_odran_elicitation_gdm_intercode_ctf Dataset Information This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-20. Model Information Model: vllm/meta-llama/Llama-3.3-70B-Instruct Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules': 'odran=jordan-taylor-aisi/odran'… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_elicitation_gdm_intercode_ctf.tabularn<1K0 likes62 downloads1y agoHugging Facejordan-taylor-aisi /odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf Inspect Dataset: odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf Dataset Information This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-19. Model Information Model: openai/meta-llama/Llama-3.3-70B-Instruct Model args: None Task Information Tasks: inspect_evals/gdm_intercode_ctf System Prompts Prompt 1: Benign You are an AI agent and expert Capture the Flag… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf.tabular0 likes56 downloads1y agoHugging Facecallum-canavan /challenges-for-unsupervised-elicitation Challenges for Unsupervised Elicitation Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks. Code: challenges-for-unsupervised-elicitation Subsets gsm8k Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.tabulartext-classification10K<n<100K0 likes50 downloads7mo agoHugging Face