input-ablation
input_ablation_qwen3_8b_mmlu_hint
Training Language Models to Explain Their Own Computations (Input Ablations)
This dataset is part of the work presented in the paper "Training Language Models to Explain Their Own Computations".
It specifically contains data for the Input Ablations task for the Qwen3-8B target model. In this task, explainer models are trained to predict how removing "hint" tokens from an MMLU prompt with a hint changes the output of Qwen3-8B. This helps in understanding the causal relationships… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/input_ablation_qwen3_8b_mmlu_hint.input_ablation_llama_3.1_8b_instruct_mmlu_hint
Training Language Models to Explain Their Own Computations - Input Ablations
This dataset is part of the research presented in the paper Training Language Models to Explain Their Own Computations.
It contains data for the Input Ablations task, where explainer models are trained to predict how removing input hints affects the target model's (Llama-3.1-8B-Instruct) predictions on MMLU questions with hints. This task evaluates whether models can understand the causal relationships… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/input_ablation_llama_3.1_8b_instruct_mmlu_hint.
