datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Video-IFBench
Video-IFBench
This release contains the evaluation split used for the Video-IFBench main experiments.
Paper: Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Project page: https://alexios-hub.github.io/Video-IFBench/
Code: https://github.com/Alexios-hub/Video-IFBench
eval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.MTAC-IFBench
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
🌟 Overview
MTAC-IFBench benchmarks instruction following in multi-turn agentic coding.
Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.IFBench
IFBench: Dataset for evaluating instruction-following reward models
This repository contains the data of the paper "Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems"
Paper: https://arxiv.org/abs/2502.19328
GitHub: https://github.com/THU-KEG/Agentic-Reward-Modeling
Dataset Details
the samples are formatted as follows:
{
"id": // unique identifier of the sample,
"source": // source… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/IFBench.IFBenchifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.IFbench_multi_constraints_upto5_systempromptifiedifbench
ifbench — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror. OpenCompass-format evaluation data for ifbench, for offline reproducible model evaluation (config ifbench_gen). Original source: allenai/IFBench_test — license ODC-BY-1.0, unchanged; all rights remain with the original authors.
