datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
intercode_bigcode_combined_12k
Intercode Bigcode Combined 12K
Balanced instruction/response dataset with 12,000 examples:
6,000 examples sampled from the filtered Tellina NL2Bash-derived InterCode training pool
6,000 examples sampled from bigcode/self-oss-instruct-sc2-exec-filter-50k
The dataset file is intercode_bigcode_combined_12k.jsonl.
Each row has:
{"instruction": "...", "response": "..."}
Sampling metadata is stored in intercode_bigcode_combined_12k.meta.json.
InterCode-Corrections
Dataset Card for InterCode-Corrections
This is a manually corrected version of the InterCode-Bash dataset, providing natural language prompts and Bash commands for the task of machine translation.
Dataset Details
Dataset Description
This dataset contains corrections for errors in the InterCode-Bash dataset. corrections.csv contains annotations for each error. final.csv contains the updated dataset with the corrections applied. The corrected dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/westenfelder/InterCode-Corrections.70B_normal_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: 70B_normal_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/70B_normal_llama_33_70b_instruct_gdm_intercode_ctf.intercode-minimax-traces
intercode-minimax-traces
Complete HTTP-level agentic traces from running intercode benchmark tasks through an instrumented reverse proxy.
Each trace captures full request/response pairs including system prompts, user messages, assistant responses, tool calls and results, and token usage metadata.
Stats
Total sessions: 338
Multi-turn sessions (2+ LLM calls): 299
Total records: 5838
Total LLM requests: 2919
Format
Raw JSONL traces from the instrumented proxy.… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/intercode-minimax-traces.normal_llama_31_8b_instruct_gdm_intercode_ctf
Inspect Dataset: normal_llama_31_8b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-18.
Model Information
Model: vllm/meta-llama/Llama-3.1-8B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'data_parallel_size': 4, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/normal_llama_31_8b_instruct_gdm_intercode_ctf.devstral_normal_devstral_small_2505_gdm_intercode_ctf
Inspect Dataset: devstral_normal_devstral_small_2505_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17.
Model Information
Model: vllm/mistralai/Devstral-Small-2505
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tokenizer_mode': 'mistral', 'config_format': 'mistral', 'load_format': 'mistral'… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/devstral_normal_devstral_small_2505_gdm_intercode_ctf.odran_default_server_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_default_server_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-19.
Model Information
Model: openai/meta-llama/Llama-3.3-70B-Instruct
Model args: None
Task Information
Tasks: inspect_evals/gdm_intercode_ctf
System Prompts
Prompt 1: Benign
You are an AI agent and expert Capture the Flag… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_default_server_llama_33_70b_instruct_gdm_intercode_ctf.odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_llama33_70b_20250620_160734_gdm_intercode_ctf.odran_default_odran_llama33_70b_20250622_105629_step1600_gdm_intercode_ctf
Inspect Dataset: odran_default_odran_llama33_70b_20250622_105629_step1600_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-23.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 128, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_default_odran_llama33_70b_20250622_105629_step1600_gdm_intercode_ctf.odran_default_odran_gdm_intercode_ctf
Inspect Dataset: odran_default_odran_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-20.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules': 'odran=jordan-taylor-aisi/odran', 'tool_call_parser':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_default_odran_gdm_intercode_ctf.odran_default_odran_20250619_075549_gdm_intercode_ctf
Inspect Dataset: odran_default_odran_20250619_075549_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-24.
Model Information
Model: vllm/odran_20250619_075549
Model args: {'model_path': 'meta-llama/Llama-3.3-70B-Instruct', 'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 128, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_default_odran_20250619_075549_gdm_intercode_ctf.odran_default_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_default_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-18.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 8, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_default_llama_33_70b_instruct_gdm_intercode_ctf.odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-18.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 8, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_llama_33_70b_instruct_gdm_intercode_ctf.odran_default_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Inspect Dataset: odran_default_odran_llama33_70b_20250620_160734_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-21.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_default_odran_llama33_70b_20250620_160734_gdm_intercode_ctf.odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-19.
Model Information
Model: openai/meta-llama/Llama-3.3-70B-Instruct
Model args: None
Task Information
Tasks: inspect_evals/gdm_intercode_ctf
System Prompts
Prompt 1: Benign
You are an AI agent and expert Capture the Flag… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_server_llama_33_70b_instruct_gdm_intercode_ctf.odran_elicitation_odran_elicitation_gdm_intercode_ctf
Inspect Dataset: odran_elicitation_odran_elicitation_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-20.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'enable_lora': '', 'max_lora_rank': 32, 'lora_modules': 'odran=jordan-taylor-aisi/odran'… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_elicitation_odran_elicitation_gdm_intercode_ctf.intercode-nl2bash-curated
Dataset Card for "intercode-nl2bash-curated"
NL2Bash Data Augmentation
Reproduction of dataset found here: https://github.com/princeton-nlp/intercode
Adopted from nl2bash by manual curation
This dataset consists of 200 natural language to bash command <query, gold> pairs
MIT License
intercode_bigcode_combined_24k
InterCode BigCode Combined 24k
Balanced instruction/response dataset with 24,000 examples:
12,000 examples sampled from the filtered Tellina NL2Bash-derived InterCode training pool
12,000 examples sampled from bigcode/self-oss-instruct-sc2-exec-filter-50k
The dataset file is intercode_bigcode_combined_24k.jsonl.
Each row has:
{"instruction": "...", "response": "..."}
Sampling metadata is stored in intercode_bigcode_combined_24k.meta.json.
