datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IFBench_test
License
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use.
Citation
Please cite:
@misc{pyatkin2025generalizing,
title={Generalizing Verifiable Instruction Following},
author={Valentina Pyatkin and Saumya Malik and Victoria Graf and Hamish Ivison and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_test.Video-IFBench
Video-IFBench
This release contains the evaluation split used for the Video-IFBench main experiments.
Paper: Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Project page: https://alexios-hub.github.io/Video-IFBench/
Code: https://github.com/Alexios-hub/Video-IFBench
IFBench_multi-turn
Dataset
This is the test data for the multi-turn setup of IFBench.
License
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use.
Citation
Please cite:
@misc{pyatkin2025generalizing,
title={Generalizing Verifiable Instruction Following}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_multi-turn.eval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.glm-5.3-flash-ifbench-openrouter
GLM-5.3-Flash IFBench OpenRouter five-run results
This dataset contains content-free results from an independent five-run
evaluation of z-ai/glm-5.3-flash on the official IFBench test set through
OpenRouter's first-party Z.AI provider.
This is not an official Allen Institute for AI, Z.AI, or OpenRouter result.
The evaluated outputs were AI-generated. Prompt, response, and reasoning text
are not included.
Results
Mean prompt-level loose accuracy was 65.5333% across… See the full description on the dataset page: https://huggingface.co/datasets/noahyoungs/glm-5.3-flash-ifbench-openrouter.ifbenchMTAC-IFBench
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
🌟 Overview
MTAC-IFBench benchmarks instruction following in multi-turn agentic coding.
Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.IFBench
IFBench: Dataset for evaluating instruction-following reward models
This repository contains the data of the paper "Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems"
Paper: https://arxiv.org/abs/2502.19328
GitHub: https://github.com/THU-KEG/Agentic-Reward-Modeling
Dataset Details
the samples are formatted as follows:
{
"id": // unique identifier of the sample,
"source": // source… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/IFBench.IF-Bench IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images
with Generative Visual Prompting
Code: https://github.com/casiatao/IF-Bench
📖 Introduction
This repository contains the infrared images in IF-Bench and translated RGB images by GenViP in the paper "IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting.".
📓 Environment Setup
# 1. create conda environment
conda create -n if_bench… See the full description on the dataset page: https://huggingface.co/datasets/casiatao/IF-Bench.ifbench-verl
IFBench-VERL: Instruction Following Evaluation Dataset for VERL Training
Overview
IFBench-VERL is a comprehensive instruction-following evaluation dataset formatted for VERL (Versatile Reinforcement Learning) training pipelines. This dataset contains 95,373 high-quality examples with 54 different constraint types, enabling systematic training and evaluation of instruction-following capabilities in language models.
The dataset is converted from… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/ifbench-verl.ifbench-translatedEuro-IFBench
Euro-IFBench
A 24-language localization of IFBench (Pyatkin et al., NeurIPS 2025) for precise instruction-following evaluation across the official EU languages. The English benchmark is preserved unchanged; the additions are translated prompts, localized verifiers for rules whose original checkers don't transfer cross-lingually, and curated cultural-grounding overrides.
Languages
Bulgarian (bg), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en)… See the full description on the dataset page: https://huggingface.co/datasets/Ayush-Singh/Euro-IFBench.ifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.IFBench_multi-turn_responses_exampleifbench_rlvrifbench-trainifbench-da-v1IFBenchIFBench_Kazakh
IFBench_Kazakh
Summary
IFBench_Kazakh is a machine-translated Kazakh version of the original IFBench benchmark. It is designed to evaluate how well language models follow instructions with explicit constraints.
The dataset contains 444 samples. Each example includes an instruction, a preferred response, a rejected response, and constraint descriptions translated into Kazakh. The original structure is preserved, allowing direct comparison with the English… See the full description on the dataset page: https://huggingface.co/datasets/issai/IFBench_Kazakh.IFbench_multi_constraints_upto5_systempromptifiedifbench
ifbench — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror. OpenCompass-format evaluation data for ifbench, for offline reproducible model evaluation (config ifbench_gen). Original source: allenai/IFBench_test — license ODC-BY-1.0, unchanged; all rights remain with the original authors.
atla-selene-1-mini-v1-ifbench_binaryreasoning-ifbench-binaryIFBench_multi_turnreasoning-ifbench-binary-noxmlifbench-pass-at-least-once-not-allifbenchifbenchreasoning-ifbench-binary-system1ifbench-pass-at-least-once-not-all-8k
