anonymousauthor01/emnlp-2026-ifr-train-val-set
EMNLP 2026 Findings: On-Policy Distillation Meets Off-Policy GRPO — Training Compact Instruction-Following Rerankers Train and validation splits used to train and evaluate the instruction-following reranker (IFR) in an in-distribution setting. The benchmark is an aggregated compilation of eight public instruction-following retrieval datasets spanning web search, code, mathematics, news, and multi-hop retrieval. Composition Source dataset Train Validation… See the full description on the dataset page: https://huggingface.co/datasets/anonymousauthor01/emnlp-2026-ifr-train-val-set.
032
