CoolFace
7 results

ifbench

allenai /IFBench_test License This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use. Citation Please cite: @misc{pyatkin2025generalizing, title={Generalizing Verifiable Instruction Following}, author={Valentina Pyatkin and Saumya Malik and Victoria Graf and Hamish Ivison and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_test.textn<1K14 likes20k downloads11mo agoHugging FaceAlexislhb /Video-IFBench Video-IFBench This release contains the evaluation split used for the Video-IFBench main experiments. Paper: Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Project page: https://alexios-hub.github.io/Video-IFBench/ Code: https://github.com/Alexios-hub/Video-IFBench textvisual-question-answeringn<1K1 likes1.3k downloads28d agoHugging Faceallenai /IFBench_multi-turn Dataset This is the test data for the multi-turn setup of IFBench. License This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use. Citation Please cite: @misc{pyatkin2025generalizing, title={Generalizing Verifiable Instruction Following}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_multi-turn.text1K<n<10K12 likes818 downloads1y agoHugging Faceshisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes483 downloads3mo agoHugging Facenoahyoungs /glm-5.3-flash-ifbench-openrouter GLM-5.3-Flash IFBench OpenRouter five-run results This dataset contains content-free results from an independent five-run evaluation of z-ai/glm-5.3-flash on the official IFBench test set through OpenRouter's first-party Z.AI provider. This is not an official Allen Institute for AI, Z.AI, or OpenRouter result. The evaluated outputs were AI-generated. Prompt, response, and reasoning text are not included. Results Mean prompt-level loose accuracy was 65.5333% across… See the full description on the dataset page: https://huggingface.co/datasets/noahyoungs/glm-5.3-flash-ifbench-openrouter.text-generationn<1K0 likes253 downloads27d agoHugging Facehirundo-io /ifbenchtextn<1K0 likes230 downloads8mo agoHugging Face