datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IFBench_test
License
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use.
Citation
Please cite:
@misc{pyatkin2025generalizing,
title={Generalizing Verifiable Instruction Following},
author={Valentina Pyatkin and Saumya Malik and Victoria Graf and Hamish Ivison and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_test.Video-IFBench
Video-IFBench
This release contains the evaluation split used for the Video-IFBench main experiments.
Paper: Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Project page: https://alexios-hub.github.io/Video-IFBench/
Code: https://github.com/Alexios-hub/Video-IFBench
IFBench_multi-turn
Dataset
This is the test data for the multi-turn setup of IFBench.
License
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use.
Citation
Please cite:
@misc{pyatkin2025generalizing,
title={Generalizing Verifiable Instruction Following}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_multi-turn.eval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.ifbenchMTAC-IFBench
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
🌟 Overview
MTAC-IFBench benchmarks instruction following in multi-turn agentic coding.
Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.IFBench
IFBench: Dataset for evaluating instruction-following reward models
This repository contains the data of the paper "Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems"
Paper: https://arxiv.org/abs/2502.19328
GitHub: https://github.com/THU-KEG/Agentic-Reward-Modeling
Dataset Details
the samples are formatted as follows:
{
"id": // unique identifier of the sample,
"source": // source… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/IFBench.ifbench-verl
IFBench-VERL: Instruction Following Evaluation Dataset for VERL Training
Overview
IFBench-VERL is a comprehensive instruction-following evaluation dataset formatted for VERL (Versatile Reinforcement Learning) training pipelines. This dataset contains 95,373 high-quality examples with 54 different constraint types, enabling systematic training and evaluation of instruction-following capabilities in language models.
The dataset is converted from… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/ifbench-verl.ifbench-translatedifbench_rlvrIFBenchifbench-trainifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.IFBench_Kazakh
IFBench_Kazakh
Summary
IFBench_Kazakh is a machine-translated Kazakh version of the original IFBench benchmark. It is designed to evaluate how well language models follow instructions with explicit constraints.
The dataset contains 444 samples. Each example includes an instruction, a preferred response, a rejected response, and constraint descriptions translated into Kazakh. The original structure is preserved, allowing direct comparison with the English… See the full description on the dataset page: https://huggingface.co/datasets/issai/IFBench_Kazakh.ifbench-da-v1IFbench_multi_constraints_upto5_systempromptifiedIFBench_multi-turn_responses_exampleifbench
ifbench — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror. OpenCompass-format evaluation data for ifbench, for offline reproducible model evaluation (config ifbench_gen). Original source: allenai/IFBench_test — license ODC-BY-1.0, unchanged; all rights remain with the original authors.
atla-selene-1-mini-v1-ifbench_binaryreasoning-ifbench-binaryIFBench_multi_turnreasoning-ifbench-binary-noxmlifbench-pass-at-least-once-not-allifbenchreasoning-ifbench-binary-system1ifbench-pass-at-least-once-not-all-8kIFBenchatla-selene-1-v2-ifbench_binaryIFBench__qwen__bon__single
