CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face02scimdr /SciMDR-Evalimagequestion-answeringn<1K1 likes3k downloads6mo agoHugging Face03mtabvqa /MTabVQA-Eval Dataset Card for MTabVQA Paper Dataset Description Dataset Summary MTabVQA (Multi-Tabular Visual Question Answering) is a novel benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to perform multi-hop reasoning over multiple tables presented as images. This scenario is common in real-world documents like web pages and PDFs but is critically under-represented in existing benchmarks. The dataset consists of two main parts: MTabVQA-Eval:… See the full description on the dataset page: https://huggingface.co/datasets/mtabvqa/MTabVQA-Eval.imagetable-question-answering1K<n<10K2 likes2.4k downloads11mo agoHugging Face04sci-m-wang /C4-Eval C4-Eval C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation. 221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures. 1,105 evaluation instances: five task formulations for every base item. Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.imageimage-text-to-text1K<n<10K0 likes634 downloads1mo agoHugging Face05CaptainGulu /R4R-Auto-Eval R4R Auto Eval 持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。 队友请先阅读 PROJECT_STATUS.md,然后按需查看: benchmarks/:固定的视频输入、来源记录和分层标签; pipelines/:判定方法及冻结配置; runs/:不可覆盖的实验记录; reports/:工作日志、方法分析和结果限制; registry/:benchmark、pipeline 和 run 的机器可读索引。 当前范围 multiscene30 是 pipeline 开发集,不是干净的留出测试集; reassemble40 是来自两个长录像的接触密集型校准集; 当前标签为来源数据提供方标签,尚未全部完成独立人工裁决; Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测; 在完成逐来源许可证核查前,本仓库应保持 private。 当前发布版本:0.1.0。 imagen<1K0 likes450 downloads7d agoHugging Face06ACSE-Eval /ACSE-Eval ACSE-Eval Dataset This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities. Dataset Overview The dataset consists of 100+ different AWS architecture scenarios, each containing: Architecture diagrams (architecture.png) Diagram source code (diagram.py) Generated CDK infrastructure code Security threat models and analysis Directory Structure Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/ACSE-Eval/ACSE-Eval.imagen<1K6 likes398 downloads1y agoHugging Face07FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes346 downloads2y agoHugging Face08zt1106 /OpenClaw-EvalMix OpenClaw EvalMix OpenClaw EvalMix is a Harbor-format collection of 360 agent-evaluation tasks across four task families. Each task directory includes task.toml, an instruction, an environment definition, and verifier tests. Composition Family Tasks Local payload clawbench 19 0.00 GiB liveclawbench 134 0.07 GiB pinchbench 147 0.02 GiB wildclawbench 60 14.05 GiB The repository preserves each family at the root so task paths remain direct.… See the full description on the dataset page: https://huggingface.co/datasets/zt1106/OpenClaw-EvalMix.imagequestion-answeringn<1K0 likes336 downloads3mo agoHugging Face09cobo2026 /R4R-Auto-Eval R4R Auto Eval 持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。 队友请先阅读 PROJECT_STATUS.md 和 INDEX.md,然后按需查看: benchmarks/:固定的视频输入、来源记录和分层标签; pipelines/:判定方法及冻结配置; runs/:不可覆盖的实验记录; reports/:工作日志、方法分析和结果限制; registry/:benchmark、pipeline 和 run 的机器可读索引。 当前范围 multiscene30 是 pipeline 开发集,不是干净的留出测试集; reassemble40 是来自两个长录像的接触密集型校准集; 当前标签为来源数据提供方标签,尚未全部完成独立人工裁决; Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测; 在完成逐来源许可证核查前,本仓库应保持 private。 当前发布版本:0.1.0。 imagen<1K1 likes319 downloads7d agoHugging Face10demisama /UGround-Offline-Evaluationimage1K<n<10K1 likes238 downloads2y agoHugging Face11chengyewang /TexOCR-eval[ACL 2026] TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction TexOCR_eval TexOCR: Benchmarking and Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction [ACL 2026 Main] This repository provides the evaluation dataset for TexOCR, designed to benchmark document OCR models on the task of compilable page-to-LaTeX reconstruction. Overview Dataset: TexOCR_eval Type: Evaluation Set Task: Document OCR → LaTeX generation… See the full description on the dataset page: https://huggingface.co/datasets/chengyewang/TexOCR-eval.image1K<n<10K0 likes225 downloads5mo agoHugging Face12ShijianW01 /MuSEAgent-Evalimage1K<n<10K0 likes216 downloads6mo agoHugging Face13pravin112 /ACSE-Eval ACSE-Eval Dataset This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities. Dataset Overview The dataset consists of 100+ different AWS architecture scenarios, each containing: Architecture diagrams (architecture.png) Diagram source code (diagram.py) Generated CDK infrastructure code Security threat models and analysis Directory Structure Each… See the full description on the dataset page: https://huggingface.co/datasets/pravin112/ACSE-Eval.imagen<1K0 likes207 downloads3mo agoHugging Face14tuhink /cambench_binary_eval CameraBench Binary Evaluation Dataset A balanced VQA dataset for evaluating camera motion understanding in videos. 📊 Dataset Statistics Total Questions: 384 Unique Videos: 119 Unique Questions: 31 Yes Answers: 192 (50.0%) No Answers: 192 (50.0%) Balance Ratio: 1.00 Total Size: 126.16 MB (0.12 GB) Average Video Size: 1.06 MB 🎯 Task Categories This dataset covers various camera motion tasks including: Static: 42 questions Move In: 29 questions Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.imagevisual-question-answeringn<1K0 likes175 downloads11mo agoHugging Face15acrever /ACSE-Eval ACSE-Eval Dataset This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities. Dataset Overview The dataset consists of 100+ different AWS architecture scenarios, each containing: Architecture diagrams (architecture.png) Diagram source code (diagram.py) Generated CDK infrastructure code Security threat models and analysis Directory Structure Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/acrever/ACSE-Eval.imagen<1K0 likes160 downloads8mo agoHugging Face16miotac /ACSE-Eval ACSE-Eval Dataset This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities. Dataset Overview The dataset consists of 100+ different AWS architecture scenarios, each containing: Architecture diagrams (architecture.png) Diagram source code (diagram.py) Generated CDK infrastructure code Security threat models and analysis Directory Structure Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/miotac/ACSE-Eval.imagen<1K0 likes109 downloads8mo agoHugging Face17DY-Evalab /ViDA-MIPI-dataset Dataset Card for ViDA-MIPI-dataset Dataset Structure Train Meta Data train_metadata.json: The meta data file contains image information: MOS, image resolution and distortion attributes. ViDA-D(ViDA-Description: reasoning quality description ) assess.jsonl: overall quality analysis. brief_assess.jsonl: structured overall quality analysis. dist_assess.jsonl: individual distortion assessment. ViDA-G(ViDA-Grounding: fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/DY-Evalab/ViDA-MIPI-dataset.image100K<n<1M4 likes71 downloads1y agoHugging Face18nezahatkorkmaz /turkish-medical-vqa-evaluatedimage1K<n<10K0 likes63 downloads1y agoHugging Face19JinyuLiu /MUMU-Eval-6000 MUMU Eval 6000 This repository contains the 6,000-image source-data evaluation set used for the Florence-2 and LFM2.5-VL-450M baselines in the MUMU evaluation repository. It is an independently prepared research split, not an official MUMU Challenge release. Splits Split Images Ground truth in manifest validation 1,000 Yes test 5,000 Yes The split contains 2,001 Task A samples, 2,000 Task B samples, and 1,999 Task C samples. All 6,000 image… See the full description on the dataset page: https://huggingface.co/datasets/JinyuLiu/MUMU-Eval-6000.imageimage-classification1K<n<10K0 likes56 downloads1mo agoHugging Face20Dnau15 /EvalMuseMiniimage1K<n<10K0 likes49 downloads2y agoHugging Face21AGI-Eval /UniGUI-Bench UniGUI-Bench A comprehensive benchmark for evaluating GUI agents across 11 capability dimensions. Dataset Structure System Evaluation trajectory: 1,200 complete GUI task trajectories with step-by-step screenshots gui_system: 7,108 individual step evaluation samples Process Evaluation (11 Abilities, 200 samples each) Ability Description element_grounding Locating UI elements given instructions action_understanding… See the full description on the dataset page: https://huggingface.co/datasets/AGI-Eval/UniGUI-Bench.imagevisual-question-answering1K<n<10K0 likes36 downloads4mo agoHugging Face22huaXiaKyrie /mmstar-qwen-sidecar-eval MMStar Qwen Sidecar Eval Local packaging of Lin-Chen/MMStar used for Qwen3-VL sidecar reproduction. Standard splits: validation.jsonl: full 1,500-sample split. test.jsonl: first 1,000 samples used by the sidecar MMStar 55-point reproduction. Compatibility files: mmstar_val.jsonl: same content as validation.jsonl. mmstar_eval_1k.jsonl: same content as test.jsonl. images/: image files referenced by the JSONL image field. The JSONL image paths are relative to the repository… See the full description on the dataset page: https://huggingface.co/datasets/huaXiaKyrie/mmstar-qwen-sidecar-eval.imagevisual-question-answering1K<n<10K1 likes34 downloads1mo agoHugging Face23mjf-su /PhysicalAI-US-Evaluation PhysicalAI-US-Evaluation A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective. Provenance Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner: the base VLA pretraining mix, the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.imagerobotics10K<n<100K0 likes23 downloads4mo agoHugging Face24mjf-su /PhysicalAI-DE-Evaluation PhysicalAI-DE-Evaluation A held-out German evaluation set for the navigation planner: 19,999 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective. Provenance Every record here was drawn from the pool of German scenes that were withheld from every training stage of the planner: the base VLA pretraining mix, the reasoning supervised fine-tuning (SFT) stage, and the… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-DE-Evaluation.imagerobotics10K<n<100K0 likes16 downloads4mo agoHugging Face25kyne0127 /vla-evaluation-v3imagen<1K0 likes16 downloads4mo agoHugging Face26gabormarko /franka-insert-siemens-lid-eval-normal-50imagen<1K0 likes8 downloads4mo agoHugging Face27robot-learning-group47 /eval3_TOY_video_images Eval3 TOY Video Images Image-question-answer grounding dataset extracted from the Eval 3 TOY celebrity permutation videos. Each episode contributes three frames from the first 6 seconds: start frame middle frame end frame The labels use the user-provided episode block ground truth: episodes 0-29: Taylor Swift episodes 30-59: Barack Obama episodes 60-89: Yann LeCun within each 30-episode block: first 10 left, next 10 middle, last 10 right Files: all.jsonl: all 270 examples… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval3_TOY_video_images.imagen<1K0 likes8 downloads4mo agoHugging Face28vsevolod-nv /aiconf-butterfly-qwen-evalQwen evaluation results for butterfly detection. Generated at: 2026-04-20 01:14:42 UTC Rows: 356 imagen<1K0 likes7 downloads5mo agoHugging Face29gabormarko /franka-insert-siemens-lid-eval-ood-2-20imagen<1K0 likes7 downloads4mo agoHugging Face30gabormarko /franka-insert-siemens-lid-eval-ood-3-5imagen<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.