datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.glm-5.3-flash-ifbench-openrouter
GLM-5.3-Flash IFBench OpenRouter five-run results
This dataset contains content-free results from an independent five-run
evaluation of z-ai/glm-5.3-flash on the official IFBench test set through
OpenRouter's first-party Z.AI provider.
This is not an official Allen Institute for AI, Z.AI, or OpenRouter result.
The evaluated outputs were AI-generated. Prompt, response, and reasoning text
are not included.
Results
Mean prompt-level loose accuracy was 65.5333% across… See the full description on the dataset page: https://huggingface.co/datasets/noahyoungs/glm-5.3-flash-ifbench-openrouter.MTAC-IFBench
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
🌟 Overview
MTAC-IFBench benchmarks instruction following in multi-turn agentic coding.
Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.ifbench-verl
IFBench-VERL: Instruction Following Evaluation Dataset for VERL Training
Overview
IFBench-VERL is a comprehensive instruction-following evaluation dataset formatted for VERL (Versatile Reinforcement Learning) training pipelines. This dataset contains 95,373 high-quality examples with 54 different constraint types, enabling systematic training and evaluation of instruction-following capabilities in language models.
The dataset is converted from… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/ifbench-verl.
